ISDS Research

SkillDev tools

Compliant, retrieval-grounded research over investor-State dispute settlement (ISDS) awards and decisions. Use when the user asks about ICSID / investment-treaty arbitration cases, awards, or doctrines (fair and equitable treatment, expropriation, jurisdiction, costs, annulment, etc.) and wants answers grounded in the actual document text with pinpoint citations. Identifies the correct document on the case page, confirms it against the PDF's own first pages, retrieves primary documents on demand from ICSID, PCA, etc., never scrapes or hosts a corpus in violation of applicable terms, and cites only retrieved text.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the ISDS Research skill

What this skill tells your AI

The instructions your AI receives, as published by lawve-ai/awesome-legal-skills in skills/isds-research-cameron-russell/SKILL.md and read by ahel’s review.

Answer questions about ISDS cases by retrieving the primary documents on demand and grounding every legal statement in the retrieved text, with pinpoint (paragraph / page) citations. This is a research aid, not legal advice.

Intended users: lawyers, arbitration practitioners, academics, and students who need verifiable ISDS research — every output is research support for the user's own professional judgment: the tool retrieves, cites, and discloses; the user analyzes and concludes.

Golden rules (read first)

  1. Ground, don't recall. Every holding, quote, or pinpoint cite MUST come from text you retrieved this session. If it isn't in the retrieved text, say "not found in the retrieved document" — never fill the gap from memory. This is the anti-hallucination guarantee. Framework law too: a statement of treaty, Convention, or arbitration-rules law (e.g. "Art. 52(6) resubmission presupposes annulment", "improper constitution is the Art. 52(1)(a) ground") made without retrieving the provision's text must carry a basis label ("per general knowledge based on training data and/or websearch — provision text not retrieved this session"); where such a premise is load-bearing for the answer, prefer retrieving the provision (ICSID hosts the Convention and Rules on its own site) — article-number and ground mislabels are a known secondary-reporting failure.
  2. Confirm whether a decision is reciting a party's argument or the tribunal's view. Decisions often spend significant space reciting the parties' positions before providing the tribunal's analysis. Before quoting or characterizing any passage: (a) Voice — verify whether it is the tribunal/committee's own finding or its recital of a party's argument, and attribute quotes, positions, and holdings accordingly (headings can help to identify whether a passage is attributable to a party or the tribunal, but are not definitive). (b) How held — when describing a holding, check whether it was unanimous or by majority: read the dispositif, check the case page for dissenting/separate opinions, and state which limbs were unanimous vs. by majority. If the retrieved text doesn't establish it, say so rather than assuming. Note: some user questions will require providing the parties' arguments, not just the tribunal's holdings.
  3. Identify before you download. A case page can list dozens of documents (award, decisions on jurisdiction, rectification, annulment, dissents, procedural orders). Never assume "the first PDF" is the one you want. List the documents, choose by title + date + proceeding, then confirm from the document's own first page(s) before relying on it.
  4. Language is an attribute, not a filter. Awards are frequently available only in Spanish/French/etc. A non-English award is a valid, relevant result — do not skip it because it isn't in English.
  5. On demand, single documents. Fetch the specific document the user needs. Never bulk-download or mirror.
  6. Sources and their rules:
    • ICSID (icsid.worldbank.org / icsidfiles.worldbank.org) — primary text. Permissive robots; Terms allow viewing/downloading for personal, non-commercial use. Do not redistribute; attribute (below).
    • PCA (pca-cpa.org; documents on docs.pca-cpa.org) — primary text for PCA-administered cases (many UNCITRAL investor-State arbitrations). Robots + Terms verified 2026-07-02: the main site's robots is default-only — no bot-specific groups, disallowing only /wp-admin/ (verified 2026-07-02 for Claude-User; re-verified 2026-07-18, incl. ChatGPT-User); the document host returns S3 AccessDenied for robots.txt (no robots file → no crawl restriction; a 4xx robots response is treated as "allow"); PCA's Terms of Use bar only commercial use without permission and impose no automated-access restriction. Fetch specific documents on demand for non-commercial research; attribute; do not republish; honor any case-specific restriction (Terms cl.1 — many PCA/UNCITRAL matters are confidential or only partially published). No PCA helper script yet: locate the document on the PCA case page and fetch that URL, then extract as for ICSID.
    • UNCTAD ISDS Navigator — discovery / metadata only. You may run these searches yourself via targeted, user-initiated fetches under the platform's own agent token (e.g. Claude-User, ChatGPT-User). Robots re-verified 2026-07-18: investmentpolicy.unctad.org currently publishes no robots.txt and unctad.org's is default-only with no bot-specific rules (an earlier verification, c. 2026-07-01, recorded an explicit ClaudeBot disallow — robots files change; re-verify periodically). Permission therefore rests on UNCTAD's Terms, which permit personal, non-commercial use — the absence of robots restrictions is NOT a licence for bulk collection. Do not scrape or paginate the UI; link and attribute, never republish it. Reliable path for complete category-filtering: use UNCTAD's official full-data Excel export (structured; all filter fields) — that's the intended public data product, not scraping (local non-commercial filtering only; don't republish or build a public derivative DB). Individual Navigator case pages are server-rendered and load fine via web_fetch (good for targeted single-case metadata — which also carries the ICSID case number for grounding); it's the filtered search/list views that are JS-rendered and time out, so enumerate via the Excel export (or a browser render), not by scraping search pages. Freshness: the data is dated snapshots refreshed ~1–2×/year and, per UNCTAD, "cannot be deemed exhaustive" (publicly-known cases only; confidential ones excluded) — state the snapshot date and this caveat.
    • italaw (italaw.com) — last-resort, per-case-confirmed primary-text fallback. Reach italaw only after ICSID / PCA / other official sources are exhausted for the specific document. Default route: the user obtains the document manually and supplies the file (--pdf-file) — italaw's Terms §4.2 expressly permit manual human browsing. An automated fetch is the fallback only, and only under all of these conditions: (i) per-document human confirmation — show the case + exact italaw URL and obtain an affirmative approval each time; no "approve all", no persistent auto-yes; (ii) the platform's own user-initiated agent token only (e.g. Claude-User on Claude, ChatGPT-User on ChatGPT) — never spoof a human-browser user-agent; (iii) one document per approval — never bulk, batch, loop, or deep-link lists; (iv) reference, don't reproduce — short pinpoint quotes + cite/link back to the italaw case page; no wholesale reproduction or AI summary substituting for the source; (v) log each approval (document, URL, timestamp) in the run log. Keep italaw content out of any build/test corpus (Terms §4.1); non-commercial use only (§§4.3/5.1); re-check robots.txt + Terms periodically (§8.1 lets italaw change them without notice). A single, per-document, human-confirmed fetch is not prohibited under §4.2 of italaw's terms as it does not constitute bulk access or circumvent any italaw access controls. Robots re-verified 2026-07-18: italaw's robots.txt operatively disallows bulk crawlers (ClaudeBot, GPTBot) but names no user-initiated agent — Claude-User and ChatGPT-User fall under the permissive * rules (crawl-delay 10, trivially satisfied by one-document-per-approval) — and now carries Content-Signals (search=yes, ai-train=no, use=reference, framed as an express rights reservation under art. 4 of EU Directive 2019/790): use=reference matches condition (iv) exactly, and ai-train=no is honored by keeping italaw content out of any build/test corpus. The gate above remains in force unchanged.
  7. UNCTAD is required for "which cases" questions — never fake completeness. When a question needs a complete set of cases filtered by a UNCTAD category (issue/breach, treaty, sector, forum, outcome, amount), build it from UNCTAD's data via the local Excel helper: python scripts/query_unctad_excel.py (filters, ICSID-case-number extraction, and the mandatory data-freshness footer). Do not substitute ICSID, WebSearch, or memory to enumerate cases — ICSID doesn't tag issues, WebSearch isn't exhaustive, and memory hallucinates and is bound by a training cut-off, so each silently misses cases. If you can't reach UNCTAD data, say the set can't be completed and scope the answer; note results are current only to UNCTAD's last snapshot and are non-exhaustive per UNCTAD; indicate to users when answers may be impacted by known data gaps.
  8. Attribute and disclaim every answer (templates below).

First run: the UNCTAD Excel (walk the user through one download)

Enumeration ("which cases") questions need UNCTAD's full-data Excel in the skill's data/ folder. If scripts/query_unctad_excel.py prints DATA_MISSING (or before the first enumeration question, if data/ has no .xlsx), do NOT try to fetch the file yourself and do NOT answer from memory. Instead, tell the user — in your own words, covering all four points:

  1. What's needed: UNCTAD publishes its full ISDS case dataset as a free Excel download; the skill filters it locally to answer "which cases" questions completely and verifiably.
  2. Why they must download it (not you, not the repo): UNCTAD's Terms permit personal, non-commercial use but bar redistribution — so the skill doesn't ship the file, and each user obtains their own copy directly from UNCTAD under those terms.
  3. Where: check the release page for the newest version first — https://investmentpolicy.unctad.org/publications/1303/investment-dispute-settlement-navigator-full-isds-data-release-as-of-31-12-2023-in-excel-format- — latest known direct link (31/12/2023 snapshot): https://investmentpolicy.unctad.org/uploaded-files/document/UNCTAD-ISDS-Navigator-data-set-31December2023.xlsx
  4. Then: ask them to upload the downloaded file into the chat so you can save it to the skill's data/ folder — that way it persists and is reused in every later session without re-uploading. (Users running locally can just place the file in data/ themselves.) When you receive the upload, save it to data/ and confirm.

Award retrieval works without the Excel; only enumeration needs it.

First run: preferences (ask once — language + where to save)

Before the first retrieval, check whether preferences are on record:

python scripts/fetch_icsid_award.py --show-config

This prints CONFIG_PATH (where the script is looking), CONFIG_DIR_WRITABLE (whether that location can store preferences), and the stored preferences or NO_CONFIG.

On a platform where this skill has not run before, also run python scripts/fetch_icsid_award.py --check-env (dependencies, network egress to the source hosts, config writability — see Platform notes) before the first retrieval.

Look for an existing config before treating this as a first run. Installed skill folders are commonly mounted read-only, so the config may live outside the skill folder: if the default path prints NO_CONFIG, check the root of the user's ISDS research area (where earlier research folders were saved) for isds-research-config.json; if found, pass it on every call: --config "<research area>/isds-research-config.json".

If no config exists anywhere, this is the first run: ask the user two questions in one interview, telling them the answers are stored so they are asked once:

  1. Preferred language, stating the policy you will follow:

I'll (i) default to providing decisions in your preferred language; (ii) indicate when the original version is in a different language; and (iii) if a decision is not available in your preferred language, tell you which languages it is available in on the ICSID website and ask how you'd like to proceed — read one of those versions (e.g. an existing ICSID translation), or have me translate the original myself, flagged as my own, non-authoritative translation.

  1. Where research folders should be saved — the local folder that will hold each topic's memo and retrieved PDFs (see Research folders below). Suggest the project's existing folder on the user's local hard drive if one is visible; otherwise ask the user to name a location they can see and keep. Never propose a temporary, hidden, or sandbox-scratch location. If the user's chosen location is not reachable from the current environment (e.g. a cloud sandbox that cannot write to the user's disk), still record the user's path as research_root — never a sandbox or session path — and deliver the files for download per Research folders item 6.

Then record both. If CONFIG_DIR_WRITABLE=yes, the default location (inside the skill folder) works:

python scripts/fetch_icsid_award.py --set-prefer-lang "English" --set-research-root "<research area path>"

If CONFIG_DIR_WRITABLE=no (read-only skill folder), ask the user where preferences should live (suggested default: the root of the research area they just chose — fold this into the same interview), then record with an explicit path and pass the same --config on every later call, in this and future sessions:

python scripts/fetch_icsid_award.py --config "<research area>/isds-research-config.json" \
    --set-prefer-lang "English" --set-research-root "<research area path>"

A failed config write never tracebacks: the script prints a CONFIG NOT SAVED block with these same instructions and exits with code 3. Never drop the user's answer or silently fall back to re-asking next session — store it at the writable location the user chose.

Workflow

Before you answer — the workflow gate (applies to every request, including one-paragraph lookups).

A request that names a case or passage is a retrieval task, not a question you may answer from the document text — or from memory, a search snippet, or a copy you already have — directly. The brevity of the request never shortens the workflow: "just quote paragraph 154" is still a grounded retrieval that must be configured, filed, and logged. Before you deliver any case text, in order:

  1. Preferences on record? Run --show-config (and, on first use per platform, --check-env). A config that has a language but no research_root is incomplete — you must still ask the research-folder half of the first-run interview. If nothing is on record, run the full interview. Do not infer a research location, and do not proceed on a language-only config.
  2. Retrieve only through the script. If dependencies are missing, install them; where you cannot, fall back to the script's own --pdf-file route or a documented degraded mode and disclose it. Never bypass the script to fetch a document by hand: an ad-hoc download skips the CONFIRM identity check and may pull a non-canonical copy (an exhibit filed in another case, a mirror) instead of the official document.
  3. Folder + document + memo. Create or locate the topic folder under research_root, retrieve and CONFIRM the document, save the PDF, and write the memo (or, for a repeat-topic quick lookup, follow Research folders item 5).
  4. Files reach the user. Deliver every file to the user's storage and say where (Research folders item 6).

If you are about to quote the document before steps 1–3 are done, stop: you are not using the skill, only imitating its output.

  1. Identify the case (or discover cases). Get the ICSID case number (e.g. ARB(AF)/00/2) or name. For cross-institution or "which cases" discovery, follow the discovery ladder (degrade gracefully, disclose at each step):

    1. Excel helper firstpython scripts/query_unctad_excel.py … gives the complete UNCTAD-tagged set up to the snapshot date (see Source routing below).
    2. Recency window (after the snapshot): the Navigator's search/list views are JS-rendered — web_fetch times out on them — so live filtered discovery needs a JS-capable render (an optional prerequisite — Claude in Chrome on Claude, or the platform's browsing/agent mode elsewhere), keeping to the specific searches the request needs; never bulk-crawl.
    3. No JS-capable render available? Supplement with ICSID's own live case database for the ICSID subset (server-rendered), and/or targeted WebSearch — but present these as non-exhaustive pointers, never as the complete set (golden rule 7).
    4. Whatever the path, state what the coverage is and what may be missed. Individual Navigator case pages (named-case lookups) never need Chrome — they are server-rendered and fetchable by numeric id.

    Then hand each ICSID case off to the retrieval steps below for grounded text. For non-ICSID cases: PCA-administered, published matters are groundable too (fetch the specific document — see Sources); other forums (e.g. SCC) yield metadata + the official link only.

  2. List the documents on the case page:

    python scripts/fetch_icsid_award.py --case "ARB(AF)/00/2" --list
    

    This prints every published document — proceeding, title, date, and one row per available language + URL — plus a machine-readable JSON_DOCS= line.

  3. Choose the target document by title + date + proceeding (e.g. "Award of the Tribunal (May 29, 2003)", not the "Introductory Note"; the right proceeding if there was an annulment). If the choice is unambiguous, pick it. If several documents plausibly match, ask the user which one rather than guessing.

  4. Retrieve + CONFIRM. Download the chosen document and read its first page(s):

    python scripts/fetch_icsid_award.py --case "ARB(AF)/00/2" --select 1 --query "fair and equitable treatment"
    

    The script prints a CONFIRM block (case-page label, detected case number, first-page text). Verify the title, parties, case number, and date match the document you intended. If they don't, stop and re-select. Language selection follows the policy below; the script downloads full text (not truncated) and extracts paragraph-aware passages.

  5. Apply the language policy. If the document is available in the user's preferred language, the script retrieves it and proceeds. If it is not, the script stops and does NOT substitute another language — it prints LANGUAGE CHOICE NEEDED with the languages the document is available in. Tell the user those languages and ask how they want to proceed: you often cannot tell from the ICSID page whether one language is authoritative or both are equally authoritative, and the user may prefer an existing ICSID translation (e.g. the English translation of a Spanish original) over a translation you produce. Only after the user chooses do you retrieve that version: --select N --lang "<choice>". In your answer: always state which language you retrieved and whether the page labels it original or translation; note that exact wording is authoritative only in the original where one is indicated; and if the user asks you to translate the original yourself, label it clearly as your own, non-authoritative translation.

  6. Answer, grounded. Quote the relevant passage; give the paragraph number (and page). Page convention: when the document's printed page numbers diverge from the PDF's page indices (common in ICSID Reports reprints and repaginated scans), cite both in the format "p. X (PDF p. Y)" — printed page first, PDF page in parentheses; when they coincide, a single page number suffices. If the user's point isn't in the retrieved text, say so plainly.

  7. Verify (required). Before sending, confirm each quote and paragraph number actually appears in the retrieved text. If you cannot verify it, remove it.

  8. Attribute + disclaim.

Worked example (single-document question, end to end)

Question: "How did the tribunal in Tecmed v. Mexico articulate the fair and equitable treatment standard?"

python scripts/fetch_icsid_award.py --case "ARB(AF)/00/2" --list
#  → structured document table, e.g.:
#    [1] Award of the Tribunal — May 29, 2003 — Original proceeding — Spanish, English
#    [2] Introductory Note — …
python scripts/fetch_icsid_award.py --case "ARB(AF)/00/2" --select 1 --query "fair and equitable treatment"

The script prints a CONFIRM block (case-page label; detected case number ARB(AF)/00/2; first-page text showing Técnicas Medioambientales Tecmed, S.A. v. United Mexican States, Award, May 29, 2003) — verify title, parties, case number, and date before relying on it — then the matching passages with para N (p.M) locators.

Expected answer shape (abbreviated):

The tribunal articulated the FET standard as requiring "…exact passage quoted verbatim from the text retrieved this session…" (Award, ¶154 (p. 61)). Retrieved: English version; the ICSID page also carries the Spanish original — exact wording is authoritative in the original where one is indicated.

Source: International Centre for Settlement of Investment Disputes. Available at https://icsid.worldbank.org. For research only; not legal advice. Verify against the official primary source.

This example deliberately does not reproduce the ¶154 text: under golden rule 1, the quote must come from the document retrieved in your session — never from this file, and never from memory.

What this tool can answer — and how completely (say so every time)

This tool has deliberate limits: it holds no scraped corpus and has no access to subscription research databases (Investor-State LawGuide (ISLG), Jus Mundi) and no full-text search over italaw (single-document retrieval only, under the last-resort gate in Sources). Acknowledging those limits is part of the design — never paper over them. Classify each question and disclose accordingly:

  1. Single-document questions ("how did case X discuss topic Y?") — fully answerable: retrieve, confirm, quote with pinpoints. No completeness caveat needed (language/translation flags still apply).
  2. Bounded-set comparisons ("compare how X, Y and Z treat topic A") — fully answerable within the named set, and the answer may rely on those cases alone. But then run the completeness check (required): using training knowledge plus a targeted WebSearch, consider whether a full treatment of the topic would implicate other cases, lines of authority, or materials you cannot access — and list them as unexamined leads, expressly not analyzed. Never let a synthesis generalize from the bounded set to "the law" without this step. (A bounded set can read as a settled trend when the set happens to sit on one side of a doctrinal split — most contested doctrines have a competing line of authority the named cases exclude; the completeness check exists to surface it.)
  3. Enumeration by UNCTAD-taggable category ("which cases arose from the Venezuelan nationalizations?") — answerable from the UNCTAD data via the Excel helper, with the standing disclosures: treaty-based cases only (contract-only or domestic-investment-law-only disputes are excluded by UNCTAD's methodology), publicly known cases only, snapshot freshness. Where the filter depends on free-text fields (e.g. summary of dispute), note that some rows have empty summaries: filter broadly (respondent/year), review, and say what the method was.
  4. Analytics over the full corpus ("what is the most-cited case on topic Z?", "how often does arbitrator N dissent?") — NOT completely answerable: that requires citation analytics or full-text search over a complete database this tool does not have. Say exactly that, then — if useful — give a general-knowledge answer clearly labeled as such, stating its basis — training data and/or websearch, as appropriate ("based on general knowledge from training data and/or websearch, not a database search, the leading case is …"), and point the user to ISLG / Jus Mundi for the authoritative answer.

Shortened here. Read the whole file on GitHub.

Signals

GitHub stars
691
Forks
88
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
isds-research
Source
github.com/lawve-ai/awesome-legal-skills