Source Tracker

SkillDatabases & data

Persistent citation database for multi-session research. Add URLs as they're cited, dedup variants, tag by topic, check link health, and export bibliographies in Markdown/BibTeX/CSV/JSON.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Source Tracker skill

What this skill tells your AI

The instructions your AI receives, as published by moonlight-lupin/agent-skills in research/source-tracker/SKILL.md and read by ahel’s review.

Overview

Source Tracker is a small persistent citation manager for research workflows. Research sources often scatter across browser tabs, notes, temporary files, and separate sessions. URLs get cited once and then disappear, duplicates accumulate as http/https or www variants, and bibliography export becomes a manual cleanup job at the end.

This skill fixes that by keeping a portable SQLite database of every cited URL: add sources as soon as they are cited, tag them by topic, deduplicate URL variants, re-check stale links, and export bibliographies in Markdown, BibTeX, CSV, or JSON.

The database path is configurable with --db-path or the SOURCE_TRACKER_DB environment variable. If neither is set, commands use ./sources.db relative to the current working directory.

Quick Start

From the skill directory:

python scripts/source_db.py add --url URL --topic TOPIC

A practical example:

python scripts/source_db.py add \
  --url "https://example.org/report" \
  --topic "battery supply chain" \
  --title "Example Battery Report" \
  --notes "Baseline market-size estimate" \
  --type report

Use a persistent project database when research spans folders or sessions:

export SOURCE_TRACKER_DB="$HOME/research/sources.db"
python scripts/source_db.py add --url "https://example.org" --topic "market scan"

Or pass the path explicitly:

python scripts/source_db.py --db-path "$HOME/research/sources.db" add --url "https://example.org" --topic "market scan"

Workflows

1. Add Sources

Add every URL at the moment it becomes a cited source. Always include a topic. Use a source type when the material is not a normal web page.

python scripts/source_db.py add \
  --url "https://example.com/article#section" \
  --topic "ai regulation" \
  --title "Example Article" \
  --notes "Defines the enforcement timeline" \
  --type news \
  --session-id "research-2026-07-06"

If --title is omitted, the script tries a best-effort GET request and parses <title>. If that fetch fails, the source is still added with a blank title.

Completion criterion: the command prints a JSON object with inserted: true for a new row or inserted: false for an exact canonical URL already in the DB.

2. Search Sources

Search by topic and optional filters. Results are JSON so they can be piped into other tools.

python scripts/source_db.py search --topic "ai regulation"

Filter by date range:

python scripts/source_db.py search \
  --topic "ai regulation" \
  --from 2026-01-01 \
  --to 2026-12-31

Filter by source type and verification status:

python scripts/source_db.py search \
  --topic "ai regulation" \
  --type report \
  --verified true

Completion criterion: the JSON array contains only records matching the topic and filters.

3. Deduplicate URL Variants

Run dedup after a research burst or before export. It merges common URL variants such as http vs https, www. vs bare host, fragment-only differences, and trailing slash differences.

python scripts/source_db.py dedup

Completion criterion: the command prints a JSON summary with merged_groups, removed_rows, and removed_ids. The survivor keeps the earliest accessed_at and combines non-empty notes from all merged rows.

4. Export Bibliographies

Export a topic bibliography. If --output is omitted, content prints to stdout. If --output is set, provide the file path to the user.

Markdown:

python scripts/source_db.py export \
  --topic "ai regulation" \
  --format markdown \
  --output ai-regulation-sources.md

BibTeX:

python scripts/source_db.py export \
  --topic "ai regulation" \
  --format bibtex \
  --output ai-regulation-sources.bib

CSV:

python scripts/source_db.py export \
  --topic "ai regulation" \
  --format csv \
  --output ai-regulation-sources.csv

JSON:

python scripts/source_db.py export \
  --topic "ai regulation" \
  --format json \
  --output ai-regulation-sources.json

Completion criterion: the output contains every source for the topic, sorted by topic/title/date, in the requested format.

5. Health Check

Run link health checks manually before a final report or from a cron scheduler. The checker uses HEAD requests, marks HTTP 200-399 as alive, updates last_checked, and flags dead links by setting verified = 0.

python scripts/url_health.py --stale-days 30 --timeout 10 --batch-size 50

With an explicit database path:

python scripts/url_health.py \
  --db-path "$HOME/research/sources.db" \
  --stale-days 7 \
  --timeout 5 \
  --batch-size 100

Completion criterion: the command prints Checked N URLs: M alive, K dead and exits with code 0. Dead links are recorded in the database; they are not treated as runtime errors.

6. Stats and Topic Inventory

Show source counts by topic, type, and verification status:

python scripts/source_db.py stats

List all topic tags:

python scripts/source_db.py list-topics

Completion criterion: stats print as JSON and topic inventory prints a JSON array of distinct topic strings.

URL Normalization Rules

Source Tracker uses two levels of URL handling:

  1. Canonical storage URL

    • Lowercase scheme and host.
    • Strip URL fragments such as #section.
    • Strip trailing slash on non-root paths.
    • Preserve query strings.
    • Preserve www. because it may be the URL the user expects to see.
  2. Dedup comparison key

    • Applies all canonical storage rules.
    • Treats http and https variants as the same source.
    • Strips a leading www. prefix from the host for comparison only.

Examples:

InputCanonical URLDedup comparison
HTTPS://WWW.Example.com/Page/#introhttps://www.example.com/Page//example.com/Page
http://example.com/http://example.com///example.com/
https://www.example.comhttps://www.example.com///example.com/

Source Types

Use one of these values with --type:

TypeUse for
webStandard web pages, documentation pages, blog posts, landing pages.
pdfDirect PDF URLs or pages where the source of record is a PDF.
apiAPI endpoints, JSON/XML data endpoints, machine-readable service output.
datasetData downloads, CSV/Parquet repositories, public data catalogs.
bookOnline books, book chapters, scans, or bibliographic pages for books.
newsNews articles, wire reports, interviews, live blogs.
reportWhite papers, government reports, analyst reports, institutional reports.

Research Workflow Integration

Deep Research

During iterative search/extract/synthesize work, add a source immediately after it passes the quality filter and before using it in the synthesis:

python scripts/source_db.py add \
  --url "SOURCE_URL" \
  --topic "PROJECT_TOPIC" \
  --title "SOURCE_TITLE" \
  --notes "Supports sub-question: ..." \
  --session-id "SESSION_ID"

Before the final report:

python scripts/source_db.py dedup
python scripts/source_db.py export --topic "PROJECT_TOPIC" --format markdown --output bibliography.md

Notebook-style Source Vaults

When building a source vault, add each source as it enters the vault. Store the vault or corpus identifier in --session-id, and use notes for coverage labels such as primary evidence, background, or contradiction.

python scripts/source_db.py add \
  --url "SOURCE_URL" \
  --topic "VAULT_TOPIC" \
  --notes "Vault source 007; primary evidence" \
  --session-id "vault-007"

Entity Research

For entity dossiers, tag sources with the entity name plus the research lens. Use source types to distinguish official registries, adverse media, reports, and datasets.

python scripts/source_db.py add \
  --url "SOURCE_URL" \
  --topic "Acme Ltd adverse media" \
  --type news \
  --notes "Allegation source; not independently verified"

Cron Usage

Run weekly health checks with the cron scheduler. Example crontab entry:

0 9 * * 1 cd /path/to/source-tracker && SOURCE_TRACKER_DB=/path/to/sources.db python scripts/url_health.py --stale-days 30 --batch-size 100 >> /path/to/source-health.log 2>&1

Use small batches for very large databases to avoid long scheduler jobs:

0 9 * * 1 cd /path/to/source-tracker && python scripts/url_health.py --db-path /path/to/sources.db --stale-days 30 --batch-size 50

Output Formats

  • Markdown bibliography — grouped by topic, with each source rendered as - [Title](URL) — notes (accessed_at).
  • BibTeX@misc{key, title={...}, url={...}, note={...}, urldate={...}}.
  • CSVid,url,title,topic,source_type,accessed_at,notes,verified,last_checked.
  • JSON — array of full source objects, including session_id and normalized URL fields.

See references/export-formats.md for concrete examples.

Common Pitfalls

  1. Invented URLs — never add a URL unless it came from a web search tool, web extraction tool, user-provided source, or another verifiable source channel. If a URL was guessed, verify it before adding it.
  2. Missing topic tags--topic is required for a reason. Use consistent topic names so search, stats, and export do not fragment across near-duplicates.
  3. Stale health checksverified = 1 means the source was alive at its last check, not that it is alive forever. Run url_health.py before final delivery for long-running projects.
  4. Duplicate variantshttps://www.example.com/, http://example.com, and https://example.com#intro can refer to the same source. Run dedup before export.
  5. Blank titles — title extraction is best-effort. For important citations, pass --title explicitly instead of relying on a remote page fetch.
  6. Relative database confusion — the default ./sources.db depends on the current working directory. Use SOURCE_TRACKER_DB or --db-path for long-running projects.

Verification Checklist

  • Every cited URL was observed through a real source channel before adding.
  • Each added row has a meaningful --topic and appropriate --type.
  • python scripts/source_db.py dedup was run before bibliography export.
  • python scripts/url_health.py was run recently for final deliverables.
  • Exported bibliography uses the requested format and output path.
  • Dead or unverified links are reviewed before final citation use.
  • The database path is explicit for multi-session work (--db-path or SOURCE_TRACKER_DB).

Signals

GitHub stars
65
Forks
11
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
source-tracker
Source
github.com/moonlight-lupin/agent-skills