Source Tracker
SkillDatabases & dataPersistent citation database for multi-session research. Add URLs as they're cited, dedup variants, tag by topic, check link health, and export bibliographies in Markdown/BibTeX/CSV/JSON.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Source Tracker skill
What this skill tells your AI
The instructions your AI receives, as published by moonlight-lupin/agent-skills in research/source-tracker/SKILL.md and read by ahel’s review.
Overview
Source Tracker is a small persistent citation manager for research workflows.
Research sources often scatter across browser tabs, notes, temporary files, and
separate sessions. URLs get cited once and then disappear, duplicates accumulate
as http/https or www variants, and bibliography export becomes a manual
cleanup job at the end.
This skill fixes that by keeping a portable SQLite database of every cited URL: add sources as soon as they are cited, tag them by topic, deduplicate URL variants, re-check stale links, and export bibliographies in Markdown, BibTeX, CSV, or JSON.
The database path is configurable with --db-path or the
SOURCE_TRACKER_DB environment variable. If neither is set, commands use
./sources.db relative to the current working directory.
Quick Start
From the skill directory:
python scripts/source_db.py add --url URL --topic TOPIC
A practical example:
python scripts/source_db.py add \
--url "https://example.org/report" \
--topic "battery supply chain" \
--title "Example Battery Report" \
--notes "Baseline market-size estimate" \
--type report
Use a persistent project database when research spans folders or sessions:
export SOURCE_TRACKER_DB="$HOME/research/sources.db"
python scripts/source_db.py add --url "https://example.org" --topic "market scan"
Or pass the path explicitly:
python scripts/source_db.py --db-path "$HOME/research/sources.db" add --url "https://example.org" --topic "market scan"
Workflows
1. Add Sources
Add every URL at the moment it becomes a cited source. Always include a topic. Use a source type when the material is not a normal web page.
python scripts/source_db.py add \
--url "https://example.com/article#section" \
--topic "ai regulation" \
--title "Example Article" \
--notes "Defines the enforcement timeline" \
--type news \
--session-id "research-2026-07-06"
If --title is omitted, the script tries a best-effort GET request and parses
<title>. If that fetch fails, the source is still added with a blank title.
Completion criterion: the command prints a JSON object with inserted: true for
a new row or inserted: false for an exact canonical URL already in the DB.
2. Search Sources
Search by topic and optional filters. Results are JSON so they can be piped into other tools.
python scripts/source_db.py search --topic "ai regulation"
Filter by date range:
python scripts/source_db.py search \
--topic "ai regulation" \
--from 2026-01-01 \
--to 2026-12-31
Filter by source type and verification status:
python scripts/source_db.py search \
--topic "ai regulation" \
--type report \
--verified true
Completion criterion: the JSON array contains only records matching the topic and filters.
3. Deduplicate URL Variants
Run dedup after a research burst or before export. It merges common URL variants
such as http vs https, www. vs bare host, fragment-only differences, and
trailing slash differences.
python scripts/source_db.py dedup
Completion criterion: the command prints a JSON summary with merged_groups,
removed_rows, and removed_ids. The survivor keeps the earliest accessed_at
and combines non-empty notes from all merged rows.
4. Export Bibliographies
Export a topic bibliography. If --output is omitted, content prints to stdout.
If --output is set, provide the file path to the user.
Markdown:
python scripts/source_db.py export \
--topic "ai regulation" \
--format markdown \
--output ai-regulation-sources.md
BibTeX:
python scripts/source_db.py export \
--topic "ai regulation" \
--format bibtex \
--output ai-regulation-sources.bib
CSV:
python scripts/source_db.py export \
--topic "ai regulation" \
--format csv \
--output ai-regulation-sources.csv
JSON:
python scripts/source_db.py export \
--topic "ai regulation" \
--format json \
--output ai-regulation-sources.json
Completion criterion: the output contains every source for the topic, sorted by topic/title/date, in the requested format.
5. Health Check
Run link health checks manually before a final report or from a cron scheduler.
The checker uses HEAD requests, marks HTTP 200-399 as alive, updates
last_checked, and flags dead links by setting verified = 0.
python scripts/url_health.py --stale-days 30 --timeout 10 --batch-size 50
With an explicit database path:
python scripts/url_health.py \
--db-path "$HOME/research/sources.db" \
--stale-days 7 \
--timeout 5 \
--batch-size 100
Completion criterion: the command prints Checked N URLs: M alive, K dead and
exits with code 0. Dead links are recorded in the database; they are not treated
as runtime errors.
6. Stats and Topic Inventory
Show source counts by topic, type, and verification status:
python scripts/source_db.py stats
List all topic tags:
python scripts/source_db.py list-topics
Completion criterion: stats print as JSON and topic inventory prints a JSON array of distinct topic strings.
URL Normalization Rules
Source Tracker uses two levels of URL handling:
-
Canonical storage URL
- Lowercase scheme and host.
- Strip URL fragments such as
#section. - Strip trailing slash on non-root paths.
- Preserve query strings.
- Preserve
www.because it may be the URL the user expects to see.
-
Dedup comparison key
- Applies all canonical storage rules.
- Treats
httpandhttpsvariants as the same source. - Strips a leading
www.prefix from the host for comparison only.
Examples:
| Input | Canonical URL | Dedup comparison |
|---|---|---|
HTTPS://WWW.Example.com/Page/#intro | https://www.example.com/Page | //example.com/Page |
http://example.com/ | http://example.com/ | //example.com/ |
https://www.example.com | https://www.example.com/ | //example.com/ |
Source Types
Use one of these values with --type:
| Type | Use for |
|---|---|
web | Standard web pages, documentation pages, blog posts, landing pages. |
pdf | Direct PDF URLs or pages where the source of record is a PDF. |
api | API endpoints, JSON/XML data endpoints, machine-readable service output. |
dataset | Data downloads, CSV/Parquet repositories, public data catalogs. |
book | Online books, book chapters, scans, or bibliographic pages for books. |
news | News articles, wire reports, interviews, live blogs. |
report | White papers, government reports, analyst reports, institutional reports. |
Research Workflow Integration
Deep Research
During iterative search/extract/synthesize work, add a source immediately after it passes the quality filter and before using it in the synthesis:
python scripts/source_db.py add \
--url "SOURCE_URL" \
--topic "PROJECT_TOPIC" \
--title "SOURCE_TITLE" \
--notes "Supports sub-question: ..." \
--session-id "SESSION_ID"
Before the final report:
python scripts/source_db.py dedup
python scripts/source_db.py export --topic "PROJECT_TOPIC" --format markdown --output bibliography.md
Notebook-style Source Vaults
When building a source vault, add each source as it enters the vault. Store the
vault or corpus identifier in --session-id, and use notes for coverage labels
such as primary evidence, background, or contradiction.
python scripts/source_db.py add \
--url "SOURCE_URL" \
--topic "VAULT_TOPIC" \
--notes "Vault source 007; primary evidence" \
--session-id "vault-007"
Entity Research
For entity dossiers, tag sources with the entity name plus the research lens. Use source types to distinguish official registries, adverse media, reports, and datasets.
python scripts/source_db.py add \
--url "SOURCE_URL" \
--topic "Acme Ltd adverse media" \
--type news \
--notes "Allegation source; not independently verified"
Cron Usage
Run weekly health checks with the cron scheduler. Example crontab entry:
0 9 * * 1 cd /path/to/source-tracker && SOURCE_TRACKER_DB=/path/to/sources.db python scripts/url_health.py --stale-days 30 --batch-size 100 >> /path/to/source-health.log 2>&1
Use small batches for very large databases to avoid long scheduler jobs:
0 9 * * 1 cd /path/to/source-tracker && python scripts/url_health.py --db-path /path/to/sources.db --stale-days 30 --batch-size 50
Output Formats
- Markdown bibliography — grouped by topic, with each source rendered as
- [Title](URL) — notes (accessed_at). - BibTeX —
@misc{key, title={...}, url={...}, note={...}, urldate={...}}. - CSV —
id,url,title,topic,source_type,accessed_at,notes,verified,last_checked. - JSON — array of full source objects, including
session_idand normalized URL fields.
See references/export-formats.md for concrete examples.
Common Pitfalls
- Invented URLs — never add a URL unless it came from a web search tool, web extraction tool, user-provided source, or another verifiable source channel. If a URL was guessed, verify it before adding it.
- Missing topic tags —
--topicis required for a reason. Use consistent topic names so search, stats, and export do not fragment across near-duplicates. - Stale health checks —
verified = 1means the source was alive at its last check, not that it is alive forever. Runurl_health.pybefore final delivery for long-running projects. - Duplicate variants —
https://www.example.com/,http://example.com, andhttps://example.com#introcan refer to the same source. Rundedupbefore export. - Blank titles — title extraction is best-effort. For important citations,
pass
--titleexplicitly instead of relying on a remote page fetch. - Relative database confusion — the default
./sources.dbdepends on the current working directory. UseSOURCE_TRACKER_DBor--db-pathfor long-running projects.
Verification Checklist
- Every cited URL was observed through a real source channel before adding.
- Each added row has a meaningful
--topicand appropriate--type. -
python scripts/source_db.py dedupwas run before bibliography export. -
python scripts/url_health.pywas run recently for final deliverables. - Exported bibliography uses the requested format and output path.
- Dead or unverified links are reviewed before final citation use.
- The database path is explicit for multi-session work (
--db-pathorSOURCE_TRACKER_DB).
Signals
- GitHub stars
- 65
- Forks
- 11
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
source-tracker- Source
- github.com/moonlight-lupin/agent-skills