Pincite

SkillDocs & knowledge

ALWAYS use when a law review manuscript's footnotes need page numbers, or when a page number already in one looks wrong, 'add pincites', 'pincite this draft', 'the editors want pin cites', 'my footnotes have no page cites', 'fill in the at-page for these cites', 'find the page that supports this claim', 'check my pincites', 'what page does this source say that'. HAND REVIEW: 'review my pincites by hand', 'show me each footnote with its PDF', 'let me enter the page numbers myself', 'apply the pincites I entered'. STATUS: 'how many footnotes still need pincites', 'which cites are still missing page numbers', 'what's left to pincite', 'how many pin-cites are left'. WRONG PIN: 'the page number is wrong', 'this pin is off by a page', 'this pincite points at preprint pagination', 'why did this cite land on the wrong page'. Use proactively before any law review submission whose footnotes carry bare cites. NEGATIVE ROUTING: whether a PDF is the version of record, a preprint or proof, or the wrong document entirely is bib-manage, hand off rather than pinning against it, even mid-run; whether a cited source exists or says anything at all is source-verify; Bluebook FORM of a citation is bluebook or bluebook-audit; archiving URLs is permacc.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Pincite skill

What this skill tells your AI

The instructions your AI receives, as published by edwinhu/workflows in skills/pincite/SKILL.md and read by ahel’s review.

What this skill carries — grep references/ for any subject the names below miss: !d=${CLAUDE_SKILL_DIR}; command -v skill-toc >/dev/null 2>&1 && exec skill-toc "$d"; s=$HOME/.claude/skills/plugin-utils/bin/skill-toc; [ -x "$s" ] && exec "$s" "$d"; echo "(skill-toc unavailable: references and scripts are NOT listed here — install the plugin-utils plugin, or start a new session so its bin/ reaches PATH)"

Supplies the page number for each footnote in a law review manuscript, and then verifies it — the model's own page number is never believed.

candidates  parse the .typ → a KIND per footnote, and for the `pdf` ones the claim
            they support plus the source PDF resolved through the BIBLIOGRAPHY
triage      per kind, what the author must actually do — including "nothing"
run         ask Gemini per candidate (Files API) for a page AND a verbatim quote
verify      re-find the quote in the PDF, derive the printed page from the
            document's own numbering, compare
report      the CONFIRMED pincites, ready to paste — plus a separate
            CONFLICTS section for rows that ALREADY carry a pin

And a hand-review loop for what the model could not settle:

review_build.py  one row per citation SITE → <root>/scratch/review-data.json,
                 plus the page copied to <root>/scratch/review/index.html: each
                 footnote with THIS site marked, beside its PDF opened at the
                 best-guess page
apply            the page's pincites.json export → the manuscript. `set` inserts
                 `, at N`; `keep` and `skip` edit nothing. DRY RUN by default

Kinds

Every footnote gets a kind. Classification runs on the footnote's own text through one table at the top of scripts/pincite.py (KIND_PATTERNS, first match wins) — except that resolution succeeding beats every label: a footnote whose source is on disk is pdf whatever its prose looks like.

kinddisposition
pdfthe Gemini pipeline — run, verify, report
legislativehearings, GAO, CRS — fetchable from govinfo as a PDF, then re-run candidates
caseout of scope: needs Westlaw/Lexis STAR PAGINATION, which no PDF yields. Flag and stop
bookscripts/gbooks.py — Google Books' print pagination, off a VERBATIM quote. See Books
webnothing to fetch — the perma.cc link IS the pin; a paragraph cite only if the journal asks
statutenothing to fetch — the section number IS the pin
regulationnothing to fetch — the section or Fed. Reg. page IS the pin (a Fed. Reg. PDF in --fedreg-dir resolves and becomes pdf)
short-formnothing here — inherits the parent cite's pin; Id. at N follows the preceding note
noneexplanatory text or an internal cross-reference (See infra Section II.B.)

none and short-form are NOT gaps. Neither is web, statute or regulation. On the first manuscript those five kinds were 138 of 262 footnotes — more than half the document — and a tool that reports them as missing pincites is reporting work that does not exist.

triage is the command that answers "what is left to do".

Already pinned — two forms, three traps

A footnote can already carry a pincite in either of two shapes, and only one of them says "at". has_pincite() in scripts/pincite.py covers both.

formexamplepinned at
at-formKahan & Rock, supra note 8, at 12791279
journal-form51 Rev. Econ. Stud. 393, 394--99 (1984)394–99
journal-form68 Fed. Reg. 6585, 6587--88 (Feb. 7, 2003)6587–88

The journal form is the STANDARD Bluebook shape — <start page>, <pin page>, no marker word anywhere. Recognising only at N inflates the candidate set and a bulk apply then double-pins footnotes that were already correct.

Detecting the journal form is harder than it looks, because a naive \d+,\s*\d+ also matches three things that are not pincites:

trapexampleguard
thousands separator90 Fed. Reg. 58,503 reads as 58, 503a number with separators is ONE token, matched atomically; the pin separator is a comma plus WHITESPACE, which a thousands separator never has
release / order numberExec. Order No. 14,366, Release No.~89,372reject a pair preceded by No. / No.~; the pin must also be ≥ the start
dateMar. 23, 2024, (Feb.~9, 2018)reject a 4-digit second number in 1900–2030

Two further guards: the pin is never less than the start page, and the two must be close (PIN_MAX_GAP, 1000) — numbers thousands apart are unrelated. Never "fix" a miss by stripping commas globally: that is precisely what manufactures the 58, 503 false pair the first guard exists to prevent.

candidates records the detected pin on the row (pin), because cite is truncated at 300 characters and a Fed. Reg. pin can sit past the cut — on the first manuscript two did.

All four stages read and write one state file (--state), so a rerun resumes rather than re-uploading.

CLI

P="${CLAUDE_SKILL_DIR}/scripts/pincite.py"

python3 "$P" candidates --root . --body paper/typst/opv-body.typ --state scratch/pincite.json
python3 "$P" triage    --root . --state scratch/pincite.json
python3 "$P" run       --root . --state scratch/pincite.json          # needs GOOGLE_API_KEY
python3 "$P" verify    --root . --state scratch/pincite.json
python3 "$P" report    --root . --state scratch/pincite.json

Hand review

python3 "${CLAUDE_SKILL_DIR}/scripts/review_build.py" --root .
python3 "${CLAUDE_SKILL_DIR}/assets/review/serve.py" 8765 .   # the REPO ROOT, or the PDFs 404
# http://localhost:8765/scratch/review/index.html — every edit autosaves to
# <root>/pincites.json; `e` still exports it if the status line says NOT SAVED.

python3 "$P" apply --root . --from pincites.json             # dry run
python3 "$P" apply --root . --from pincites.json --confirm   # writes

Both scratch/percite-classified.json and scratch/pincite.json are optional inputs to the build: without them there is no decision column and the best-guess page is the bibliography's start page. review-data.json and scratch/review/index.html are OUTPUT — overwritten every build, never edited in place.

apply identifies a site by (fn, citekey, occurrence), asserts its match string occurs exactly once inside that footnote's own span, and refuses the WHOLE run if any assertion fails. It records every decision — set, keep and skip alike — in the existing scratch/percite-classified.json, marking a hand-entered pin "source": "author" with the author's page as verified_page. That record is what a provenance gate traces; a pin applied without it traces to nothing and should fail.

Every path is a flag, so the tool works on any manuscript: --body, --bib, --pdf-dir, --fedreg-dir. --bio-offset (default 3) is how many leading author-identification footnotes render as *, , and take no Arabic number — get it wrong and every reported footnote number is off by a constant. --only 12,40 and --limit scope a run; --redo re-asks footnotes already answered.

run requires GOOGLE_API_KEY. So does verify --vision-fallback, which is opt-in because it costs API calls — use it only on rows verify refuses with "page numbering not readable".

The gate on the page-number derivation is a script, not a judgement:

python3 "${CLAUDE_SKILL_DIR}/tests/test_page_offset.py" [--corpus /path/to/repo] [--vision]

It asserts 11 measured offsets against real PDFs and exits non-zero on a real failure; with no corpus it skips and exits 0. The fixtures are paywalled and are not vendored here.

apply has its own suite, which needs no corpus — every fixture is built inline: cd "${CLAUDE_SKILL_DIR}" && uvx --with pytest pytest (no system pytest here).

Books

A book has no PDF, so no offset route reaches it. Google Books indexes the print pagination of the edition IT scanned, and scripts/gbooks.py reads a page off it.

G="${CLAUDE_SKILL_DIR}/scripts/gbooks.py"
python3 "$G" volume --volume-id <VOL>                        # edition record
python3 "$G" page   --volume-id <VOL> --quote "<verbatim>"   # or --isbn / --title
python3 "$G" calibrate --pins pins.json --volume-id <VOL>

The input is a VERBATIM quote, never the claim. This searches a text index; a manuscript's paraphrase matches nothing. A Kindle highlight already IS the right input — a verbatim quote with the book attached — so Readwise is the feeder:

python3 "$G" highlight --highlights hl.json --location 1874 --volume-id <VOL>
python3 "$G" calibrate-highlights --highlights hl.json --pins pins.json --volume-id <VOL>

The librarian agent runs readwise and writes hl.json; this skill only ever reads it — never shell out to readwise from here or from the main thread.

Readwise's stored text carries its own artifacts, so the key is built, not taken: em dashes flattened to hyphens (three people-Monks), spaces lost at line breaks (office onM Street), and ~20 of one book's 92 records truncated mid-sentence with a trailing . search_key normalises, steps the window around the artifacts, and strips the truncation — and reports it as evidence: prefix, which is weaker than a full sentence because the page is the prefix's, not the whole highlight's. calibrate_highlights counts those as weak.

Keep quotes to 4–8 words. The scan hyphenates across line breaks (consid- erations), so a long quote spanning one silently fails to match while its tail fragment finds the page.

calibrate is the gate and it is computed, not judged: re-find quotes whose pages are already trusted. reproduces → the book's new pages are usable; offset → a different printing, and k IS the finding; inconsistent → unusable, and it says so; blocked → no answer either way, which is never a pass. A page inside a cited range (145--48 → 148) is a match, not a three-page miss.

Google refuses scripted clients on both hosts, so the transport drives the real browser over CDP (9250, then 9222) and throttles to REQUEST_GAP. A burst of searches costs the whole host its Google access for an hour — the user's own browsing included.

Facts

  • "No local PDF" is not a gap — it is five different dispositions. On the first manuscript that one bucket held 153 of 262 footnotes, and the largest group in it (short-form, 60) needs nothing at all. Classifying instead: 109 pdf, 60 short-form, 36 none, 21 web, 12 statute, 10 case, 9 regulation, 4 legislative, 1 book. Of those, 138 need no action whatsoever.

  • Keep the book pattern narrow and keep it last. A book cite's shape is Author, Title Page (Year) — a bare page against a parenthetical year — which is one token away from a journal cite's Volume Reporter Page (Year). The pattern therefore requires the book shape AND the absence of any volume-reporter-page string. A misfire here sends a real article to the "needs a scan" pile, where nobody will ever look for it again.

  • at N is the minority pincite form, not the only one. Matching only \bat\s+\*?\d missed 10 of 29 already-pinned footnotes on the first manuscript — every one of them a standard journal or Fed. Reg. <start>, <pin> cite. Three of the ten had also been asked and confirmed, and two of those three came back DISAGREEING with the pin already in the text (fn91 existing 2968--72 vs computed 2955; fn133 existing 850--52 vs computed 881). A bulk apply would have double-pinned all three and overwritten two author page choices with a machine's.

  • Never trust the model's page number. It reports a page and a quote; only the quote can be checked mechanically. verify finds the quote in the PDF and then derives the printed page from the document's own numbering — the modal printed − pdf_index across every page, since one page's header cannot be told from a year or a footnote number. Accepting the reported page caught nothing; deriving it caught two real errors on the first manuscript.

  • The modal offset's own confidence cannot tell a thin-but-right answer from a thin-and-wrong one — the bib's start page can. Measured on 11 OPV sources, fisch1994's correct offset held 15 votes of 42 (conf 0.36, refused) while choi2009's garbage offset of 23 held 2 of 56 — the same symptom, six hundred pages apart. Loosening the threshold trades one error for the other. Instead page_offset takes pages = {...} from the bib: under offset k the work's first printed page is k+1, so the cited start page must sit at pdf index start − k, a small number inside the front matter (FRONT_MAX, 12 — enough for a HeinOnline cover and for a bib filed at 1012 on a work that prints 1009). fisch1994 puts it on pdf page 4; choi2009's 23 puts it on page 626 of a 56-page PDF. Requiring two pages to agree (MIN_CORROB_HITS) is what makes it safe: a CONSTANT number repeated across pages yields a different k on every page and can never accumulate, so only a real folio counts twice. This route runs FIRST and promoted 3 more footnotes; corroboration is independent evidence, the vote count is not.

  • A dominant offset whose printed range does not contain the bib's start page is numbering a different version of the work. hayne2019 is a working paper paginated −1–57, cited to J. Acct. Res. 57:969; the old gate confirmed "page 6" — a wrong pincite, now refused. The one exception is real and narrow: Elsevier records an ARTICLE NUMBER in pages (154 (2024) 103810) and paginates the version of record 1–N, so a pages value over ARTICLE_NUMBER_FLOOR is treated as no start page at all. Without that exemption three correct pincites read as preprints.

  • Strip [Vol. 82:649 before reading folios. A verso running head carries the volume and the work's START page, and that trailing number is not this page's folio — left in, it mints a spurious candidate on every verso and, on fisch1994, was most of the runner-up mode that sank the real offset.

  • Vision is the LAST resort, and only for the offset. When a scan has lost every folio from the text layer (choi2010: the best text offset held 1 vote of 52) --vision-fallback renders two pages and reads the folios off the images. The numbers are still not believed: they must advance in lockstep with the pdf index AND the offset they imply must corroborate against the bib start. On choi2010 that returned 867 — matching the Hein cover arithmetic — and on choi2009 it independently reproduced the text route's 647.

  • Resolve footnote → source through the bibliography, never through PDF filenames. Each bib entry carries file = {...}; the manuscript cites by citekey. Matching a footnote's prose against filenames sent 21 of 88 footnotes to the wrong source, and one of those verified CONFIRMED while pointing at the wrong paper — a false positive, which is worse than a miss.

  • A footnote is usually a string cite; the claim belongs to the source cited FIRST. The rest are "see also" support. For a full cite carrying no citekey label, an author-surname match requires the entry's year within ~220 characters of the name — a footnote-wide year test pairs one work's author with another work's date across a string cite.

  • A NO answer is usually MIS-RESOLUTION, not a bad citation. Check which PDF the footnote resolved to (pdf in the state file) before reporting a citation error to the author. All three NOs on the first manuscript were the wrong file: a Calluzzo & Kedia claim had been handed Albuquerque, a Li, Maug & Schwartz-Ziv claim had been handed Malenko. Reporting them as citation errors would have sent three false corrections to coauthors while leaving eighteen wrong pincites in place. Reporting "your source doesn't support this" about a source the author never cited is worse than silence.

  • Quote matching must tolerate text injected INTO the quote. Publisher watermarks are injected MID-SENTENCE into the text stream (Downloaded from https://academic.oup.com/…, This content downloaded from … JSTOR), so a real quote splits in half. Strip them, then score coverage — the sum of difflib matching blocks over the quote's length. Never fixed-width window similarity: it rejected 11 of 17 real quotes on the first run. Coverage tolerates insertion but not reordering, which is the right discrimination. Two cases are neither a hit nor a miss but "check by hand": a model page given as a span (1231-1232) compares on its first number, and a short EXCERPT of a book has no derivable numbering at all — that is not "quote not found".

  • Page numbers sit in running headers beside other text (2008] THE HANGING CHADS OF CORPORATE VOTING 1279) and run to six digits, so a bare-number-line regex finds nothing. \d{1,4} silently excludes Federal Register pages, which are five digits — the range test then reports every FR document as unpaginated.

  • Use Flash, not Pro. gemini-batch measured Pro as the most CONSERVATIVE extractor — 47% detection vs Flash's 70% on identical documents. Here under-extraction is the silent failure: a missed pincite reads as "the source doesn't support it". A weaker quote is not silent — verify catches it.

  • A footnote citing Fed. Reg. pagination needs a Fed. Reg. page. Fetch from govinfo's link service, https://www.govinfo.gov/link/fr/<vol>/<page>?link-type=pdf, into --fedreg-dir as fedreg-<vol>-<page>.pdf. The SEC's own release PDF of the same document has different pagination. federalregister.gov bot-blocks. Pre-2000 volumes return the whole issue (200+ MB) — trim with qpdf.

  • A preprint gives the WRONG page, and so can the publisher's own PDF. Check the PDF's derived page range against the bib's published start page. Two exceptions produced false alarms: Elsevier article-number journals (154 (2024) 103810) are paginated 1–N and ARE the version of record; and a publisher endpoint can serve an ADVANCE-ACCESS PROOF — OUP returned v 00 n 0 2021, internal name RFS-OP-…tex, pages 1–49 instead of 5581+, authenticated and still useless for a pincite. Re-run the range check on anything fetched; a successful fetch never means the pagination problem is solved.

Red flags — STOP

About toWhy wrongDo instead
Bulk-apply report outputIts CONFLICTS section is footnotes that are ALREADY pinned, where the computed page can disagree with the author's own. A wrong pincite is worse than a missing oneRead CONFLICTS first and hand every disagreement to the author; paste only the clean list
Chase a pincite for every footnote without a PDFA third of them need none — and case needs star pagination this tool cannot giveRun triage first
Paste the model's page into a footnoteIt is unverified; two were wrong on the first manuscriptRun verify, paste only what report prints
Match a footnote to a PDF by filename21 of 88 went to the wrong source, one silently CONFIRMEDResolve through file = {...} in the bib
Report "the source doesn't support this" on a NONO is usually mis-resolutionRead the row's pdf first
Switch --model to a Pro model for "better" answersPro under-extracts (47% vs 70%) and the failure is silentStay on Flash
Use the SEC release PDF for a Fed. Reg. citeDifferent pagination from the one the cite usesFetch via govinfo link/fr/<vol>/<page>
Trust a page from a PDF paginated 1–NPreprint — the published pages differNothing: page_offset now checks the bib start itself (Elsevier article-number journals excepted)
Loosen the offset threshold to rescue a refused footnoteA weak-and-WRONG offset looks identical from inside the vote count; choi2009's garbage was 600 pages offCorroborate against the bib's start page — and run tests/test_page_offset.py, which fails if 23 or 30 is ever accepted
Take a folio a vision read returnsIt is one unverified number, the thing this tool exists to refuseTwo pages, advancing in lockstep, corroborated against the bib start — what --vision-fallback already does
Search Google Books with the manuscript's CLAIMIt is a paraphrase; the index holds the book's own words. Keyword queries return a page, not the page — on the first book that produced a 37-page "offset" that was pure query noiseSearch a 4–8 word verbatim quote, and read the returned snippet before believing the page
Read a Google "no results" page as "not in the book"The rate-limit interstitial has no results in it either, so a block reads as a clean negativeis_blocked() — a blocked lookup reports blocked, and calibrate refuses to rule
Pin a book whose calibration came back inconsistentGoogle paginates the printing it scanned, not necessarily the cited oneReport it as unpinnable; do not average the deltas
Match a block marker against raw HTMLA healthy SERP ships the script that DETECTS a /sorry redirect, so every good page reads as blockedis_blocked matches rendered text, and /sorry only as an href
Call readwise to feed this routeStanding rule, and the CLI is not on the main thread's allowlistHave the librarian agent write the dump; pass it in with --highlights
Hand-edit the state file to "fix" a pageThe next candidates re-parse drops answers whose PDF or claim changed, so the edit is silently lost or silently kept staleFix the resolution or the bib, then --redo

Signals

GitHub stars
21
Forks
4
Last commit
Sep 2026

ahel review

  • K6low
    bundled executables the agent is told to run

Automated review, not a security audit. Ruleset v1+k2.

Advanced
Catalog kind
skill
Gateway key
pincite
Source
github.com/edwinhu/workflows