Market Sizing Skill

SkillDocs & knowledge

Builds credible TAM/SAM/SOM analysis with external validation and sensitivity testing for startup fundraising. Supports top-down, bottom-up, or dual-methodology approaches. Run the sourced, sensitivity-tested analysis rather than estimating a market from memory.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Market Sizing Skill skill

What this skill tells your AI

The instructions your AI receives, as published by lool-ventures/founder-skills in founder-skills/skills/market-sizing/SKILL.md and read by ahel’s review.

Help startup founders build credible, defensible TAM/SAM/SOM analysis — the kind that earns investor trust rather than raising eyebrows. Produce a structured, validated market sizing with external sources, sensitivity testing, and a self-check against common pitfalls. The tone is founder-first: a rigorous but supportive coaching session.

Skill Metadata

  • Author: lool-ventures
  • Version: managed in founder-skills/.claude-plugin/plugin.json
  • Compatibility: Python 3.10+ and uv for script execution.
  • Exports:
    • sizing.jsonfinancial-model-review, ic-sim, fundraise-readiness
    • sensitivity.jsonfinancial-model-review

Skill Execution Model (READ FIRST)

See founder-skills/references/skill-execution-model.md for the full inline-skill execution model (3 dispatch contexts, Mitigation 1+2, producer contract, Cowork quirks, per-symptom triage).

This skill runs inline in the main thread, not as a sub-agent — see the reference above ("Why Inline (Not Forked Sub-Agent)") for the rationale. Sub-agents are deliberately shell-free, so orchestration (producer scripts, artifact persistence, web research) stays in the main thread.

Two dispatch contexts for the sub-agent:

  • Context A — Per-step analytical dispatch (Mitigation 1): Steps 5 and 6 dispatch the market-sizing agent via the Task tool. The key element here is parallel dispatch: Step 5 (methodology calculation) dispatches the agent twice simultaneously — one for TOP_DOWN_METHODOLOGY and one for BOTTOM_UP_METHODOLOGY — in a single assistant turn when the methodology is "both". The sub-agent does deep analysis, WRITES its output JSON to the OUTPUT_PATH given in its prompt (the handoff/ dir), and returns a small receipt. The main thread gates the file with check_handoff.py, then pipes it through the producer script (market_sizing.py --stdin). The sub-agent never writes canonical artifacts — only its hand-off file.
  • Context B — Post-compose coaching dispatch: The final step dispatches the sub-agent after compose_report.py --write-md has written report.md. The sub-agent Reads the staged coaching_payload.json from the hand-off dir (Mitigation 2) — it does NOT read the full report.md — composes the coaching commentary, WRITES it to the OUTPUT_PATH hand-off file, and returns a small receipt. The main thread gates the file (check_handoff.py) and inserts it via the shared insert_coaching.py script (idempotency matrix, uuid-marker replacement, run_id-parity verification — all deterministic). See the reference above for the full Context B contract.

Research-before-dispatch pattern: The main thread performs web research (WebFetch/WebSearch, or the host's equivalents) BEFORE dispatching sub-agents. Research data is passed inline in the sub-agent prompts. This skill's sub-agent tools: allowlist deliberately includes no network tools (a design choice, not a platform limit — see the reference), so research runs in the main thread and is passed inline.

Tolerant JSON extraction protocol (Context B returns; also the Context A message-channel fallback): capture the sub-agent's final assistant message. It should be raw JSON, but may be wrapped in ```json ... ``` fences or carry a prose preamble. Extract tolerantly:

  1. If the message is wrapped in a ```json ... ``` (or plain ``` ... ```) fence, strip the fence first.
  2. Try to parse the stripped text directly as JSON.
  3. If that fails, walk through the text looking for the first { character and try json.JSONDecoder().raw_decode(text[i:]) — this is brace-aware and handles nested objects correctly (unlike regex, which truncates on the first }).
  4. If extraction fails entirely, re-prompt the sub-agent with: "Your previous reply could not be parsed as JSON. Return ONLY the JSON object — no markdown fences, no prose preamble."

Context A receipts don't need this protocol by hand — check_handoff.py --receipt-json - applies the same tolerant extraction internally; pass the final message verbatim.

Input Formats

Accept any format: pitch deck (PDF, PPTX, markdown), financial model, market data, text descriptions, or verbal description of the business.

Available Scripts

All scripts are at ${CLAUDE_PLUGIN_ROOT}/skills/market-sizing/scripts/:

  • market_sizing.py — TAM/SAM/SOM calculator (top-down, bottom-up, or both); accepts --stdin for JSON piping
  • sensitivity.py — Stress-test assumptions with low/base/high ranges and confidence-based auto-widening
  • checklist.py — Validates 22-item self-check with pass/fail per item
  • compose_report.py — Assembles report with cross-artifact validation; --write-md writes report.md; --strict exits 1 on high/medium warnings
  • visualize.py — Generates self-contained HTML with SVG charts (not JSON)

Also available from ${CLAUDE_PLUGIN_ROOT}/scripts/ (shared):

  • founder_context.py — Per-company context management (init/read/merge/validate)

Run with: python3 ${CLAUDE_PLUGIN_ROOT}/skills/market-sizing/scripts/<script>.py --pretty [args]

Available References

Read as needed from ${CLAUDE_PLUGIN_ROOT}/skills/market-sizing/references/:

  • tam-sam-som-methodology.md — Definitions, calculation methods, industry examples, best practices
  • pitfalls-checklist.md — Self-review checklist for common mistakes
  • artifact-schemas.md — JSON schemas for all analysis artifacts

Artifact Pipeline

Every analysis deposits structured JSON artifacts into a working directory. The final step assembles all artifacts into a report and validates consistency. This is not optional.

StepArtifactProducer
1founder contextfounder_context.py read/init
2inputs.jsonAgent (heredoc)
3methodology.jsonAgent (heredoc)
4validation.jsonMain thread (WebFetch/WebSearch research)
5sizing.jsonContext A dispatch: TOP_DOWN_METHODOLOGY + BOTTOM_UP_METHODOLOGY in parallelmarket_sizing.py --stdin
6asensitivity.jsonContext A dispatch: SENSITIVITY_TEST → sensitivity.py
6bchecklist.jsonContext A dispatch: CHECKLIST → checklist.py
7Reportcompose_report.py --write-md (writes both report.json and report.md)
8CoachingContext B dispatch: POST_COMPOSE_COACHING

Rules:

  • Deposit each artifact before proceeding to the next step
  • For agent-written artifacts (Steps 2-4), consult references/artifact-schemas.md for the JSON schema
  • If a step is not applicable, deposit a stub: {"skipped": true, "reason": "..."}
  • Do NOT use isolation: "worktree" for sub-agents — files written in a worktree won't appear in the main $ANALYSIS_DIR

Keep the founder informed with brief, plain-language updates at each step. Narrate the founder-visible OUTCOME, never the internal step. That is the test to apply, and it catches more than a word list can: the forbidden thing is not a syntax, it is talking about the machinery. Bad — "Gating and piping the extraction through the producer, then staging the coaching hand-off"; good — "I've checked your numbers and I'm writing up what stood out." Bad — "schema-drift warning on coaching_payload"; good — nothing, because the founder has no stake in it. Never name an internal artifact, field, or token (a payload key, a marker name, an artifact filename, a hand-off dir) even in plain prose with no backticks — a detector keyed on syntax cannot see "gated", "hand-off" or "canonical artifacts", but the founder still reads them and they still mean nothing to them. The between-step progress lines are the primary leak vector, not the final summary. They feel internal — you are narrating what you are about to do — but the founder reads every one of them, and this is where the leaks actually appear: "Now gating the hand-off before piping through the checklist producer", "Gate 1 passes", "Running the final verification gate". Rewrite each pipeline transition as the founder-visible outcome: "Checking your numbers against the 46-point review", "Your inputs look consistent — moving on to unit economics", "Finishing up and putting the report together". If a progress line would mean nothing to someone who has never seen this skill's internals, it does not belong in the channel. Also excluded, as before: file/script names, paths, *.py, --flags, $vars, exit codes ("Exit N", "not found"), W_/E_ codes, JSON, and step/route labels ("Lane N", "Context A/B", "Phase N", "structure detection", "the grid", any ALL_CAPS_TOKEN). After each analytical step (5–6), share a one-sentence finding before moving on. The task tracker is founder-visible too — the same rule governs its labels. "Gate the inputs review handoff", "Validate inputs.json", "resolve agent namespace paths", "Initialize founder context" are leaks even though each names a real step, and even when the prose around them is clean. Label each task by the founder-visible outcome — "Check your inputs", "Score against the review", "Write up what I found" — never by a file, directory, script, or pipeline stage.

Workflow

Step 0: Path Setup

Every Bash tool call runs in a fresh shell — variables do not persist. Run the block below exactly once: it resolves $PLUGIN_ROOT deterministically, and every later block must substitute the printed value as a literal rather than re-running the resolution — repeating the self-heal search can land on a different mount than Step 0 picked when more than one is present (see why in the block's comments).

Optional, best-effort, and via the Read tool (not a shell command): before the block below, Read ${CLAUDE_PLUGIN_ROOT}/.claude-plugin/plugin.json and note its version field as EXPECT_VERSION. Passing it to select_plugin_root.py below lets an exact version match win over an arbitrary first hit. If the Read fails, skip it and omit --expect-version — selection is still deterministic without it.

SCRIPTS="${CLAUDE_PLUGIN_ROOT}/skills/market-sizing/scripts"
if [ ! -d "$SCRIPTS" ]; then
  # In Cowork, CLAUDE_PLUGIN_ROOT substitutes to a host-side path absent inside
  # the session VM — self-heal by collecting EVERY candidate mount (a session can
  # have more than one at once: a stale host-side cache, a test marketplace, even
  # a symlink into a different session's tree) and handing them to
  # select_plugin_root.py, which picks ONE deterministically and names the
  # rejects — never trust `find`'s arbitrary first hit, which can silently mix
  # scripts across plugin versions mid-pipeline.
  CANDIDATES="$(find /sessions -type d -path '*/skills/market-sizing/scripts' 2>/dev/null)"
  [ -n "$CANDIDATES" ] || CANDIDATES="$(find / -type d -path '*/skills/market-sizing/scripts' 2>/dev/null)"
  PROVISIONAL_ROOT="$(printf '%s\n' "$CANDIDATES" | head -1)"
  PROVISIONAL_ROOT="${PROVISIONAL_ROOT%/skills/*}"
  # Bootstrap order: $SHARED_SCRIPTS isn't known until a root is chosen, so use the
  # provisional root's OWN copy of the selector; an older plugin copy without one
  # falls back to the provisional root unchanged.
  SELECTOR="$PROVISIONAL_ROOT/scripts/select_plugin_root.py"
  if [ -f "$SELECTOR" ]; then
    if [ -n "$EXPECT_VERSION" ]; then
      PLUGIN_ROOT="$(printf '%s\n' "$CANDIDATES" | python3 "$SELECTOR" --expect-version "$EXPECT_VERSION")"
    else
      PLUGIN_ROOT="$(printf '%s\n' "$CANDIDATES" | python3 "$SELECTOR")"
    fi
  else
    PLUGIN_ROOT="$PROVISIONAL_ROOT"
  fi
  SCRIPTS="$PLUGIN_ROOT/skills/market-sizing/scripts"
fi
PLUGIN_ROOT="${SCRIPTS%/skills/*}"
echo "PLUGIN_ROOT=$PLUGIN_ROOT"   # resolved ONCE, here — paste this literal into every later block; never re-run this resolution
REFS="$PLUGIN_ROOT/skills/market-sizing/references"
SHARED_SCRIPTS="$PLUGIN_ROOT/scripts"
SHARED_REFS="$PLUGIN_ROOT/references"
# Resolve the canonical artifacts root via a SCRIPT, not inline bash (the agent paraphrases inline
# path computations → outputs/ vs outputs/artifacts/ drift across runs). Deterministic + creates it.
python3 "$SHARED_SCRIPTS/resolve_artifacts_root.py"   # prints ARTIFACTS_ROOT — use the printed path verbatim as ARTIFACTS_ROOT in every later block (a captured var dies in the next fresh shell)

Reaching the self-heal branch is normal in Cowork — ${CLAUDE_PLUGIN_ROOT} resolves to a HOST path that does not exist inside the VM, so the [ ! -d "$SCRIPTS" ] test fails by design rather than by misconfiguration. It is not a sign anything is wrong, and it is not worth narrating to the founder.

Outputs mount is append-only. Everything under the promoted outputs mount (.../mnt/outputs/, not just $ANALYSIS_DIR) is write-allowed and delete-denied by the platform: never rm, move away, or empty anything under it — including files you created yourself. Never create ad-hoc scratch anywhere under the outputs mount (no _src/ copies, no run-state note files); scratch belongs in $STAGING_DIR (a /tmp dir, defined below). Do not "clean up" the outputs folder before delivering — extra working files there are expected and harmless.

If ARTIFACTS_ROOT resolves to $(pwd)/artifacts but no artifacts/ directory exists at $(pwd): Use Glob with pattern **/artifacts/founder_context.json to locate existing artifacts, and derive ARTIFACTS_ROOT from the result. If nothing is found, mkdir -p "$ARTIFACTS_ROOT" and proceed.

After Step 1 (when the slug is known), derive ANALYSIS_DIR. Two modes — pick exactly one:

  • Full analysis (default — the founder shared materials, asked for a TAM/SAM/SOM analysis or a report, OR there is no existing full analysis for this slug): run Steps 2–10. ANALYSIS_DIR="$ARTIFACTS_ROOT/market-sizing-${SLUG}".
  • Quick-check mode — a single directional sizing question in conversation, with no materials attached and no request for an analysis or report ("roughly how big is this market if we charge $15k to 18,000 pharmacies?", "does a €2B TAM sound plausible for X?"). Run Step 5-quick instead of Steps 2–10. ANALYSIS_DIR="$ARTIFACTS_ROOT/market-sizing-${SLUG}-quickcheck".

Tie-breaker when both bullets seem to fit — and they often will. A founder who supplies complete inputs conversationally ("size the market: 18,000 pharmacies at €15k, 35% serviceable, 2% capture") matches the full-analysis bullet on what they asked for and the quick-check bullet on how they asked. Decide on the verb, not the inputs:

  • "size the market", "analyze", "build me a TAM", "I need this for a deck"full analysis, even when every number is already in hand. They asked for the work product, and the sourcing, sensitivity and 22-item check are the work product.
  • "roughly", "ballpark", "sanity-check", "does X sound right", "how big is"quick-check, even when materials are attached.

Complete inputs are not a signal for quick-check. They make the full analysis faster, not less wanted. When the verb is genuinely absent — a bare list of numbers with no request — default to full analysis and say you did: an unwanted full run costs the founder time, an unwanted quick check costs them the analysis they came for.

Never answer a sizing question from your own arithmetic. Quick-check exists because the alternative a model reaches for — computing the number in its head and offering the real analysis as an opt-in — produces a figure with no provenance, no sensitivity range, and no record, under this skill's name. Running fewer producers is fine; running none is not.

Step 5-quick: the quick-check path

Run the same producer the full pipeline uses, with only the inputs the founder gave you:

printf '%s' "$QUICK_JSON" | python3 "$SCRIPTS/market_sizing.py" --stdin --pretty \
  --run-id "$RUN_ID" --currency "$CURRENCY" -o "$ANALYSIS_DIR/sizing.json"

Producers deliberately NOT run: external validation (Step 4), sensitivity.py, checklist.py, compose_report.py, visualize.py, and the Context-B coaching dispatch. No report.md is written.

Same-numbers guarantee. The TAM/SAM/SOM figures are identical to what the full analysis would compute from the same inputs — it is the same script reading the same shape. Only the production weight is dropped. What you do not get is what those skipped producers add: sourced assumptions, a low/base/high range, the 22-item quality check, and the deck-claim reconciliation.

Presenting it. Label it a quick check, not an analysis. State the figures, name the inputs they came from, and say plainly that the assumptions are unsourced and unstressed. Then close with a statement, never a question: "The full analysis sources each assumption, stress-tests the range, and produces a report you can put in front of an investor — say the word and I'll run it." A question invites a "no" to something the founder would have wanted.

ANALYSIS_DIR="${ANALYSIS_DIR:-$ARTIFACTS_ROOT/market-sizing-${SLUG}}"            # full analysis
# ANALYSIS_DIR="${ANALYSIS_DIR:-$ARTIFACTS_ROOT/market-sizing-${SLUG}-quickcheck}"  # quick check
mkdir -p "$ANALYSIS_DIR"
RUN_ID="$(date -u +%Y%m%dT%H%M%SZ)"
# Context A hand-off dir — PER RUN: sub-agents WRITE their raw output JSON here (the audit trail —
# raw sub-agent output as returned, before producer validation). Permanent by platform design
# (outputs/ mounts are write-allowed / delete-denied); nothing in it is ever a canonical artifact.
# The $RUN_ID segment is load-bearing: it prevents a stale prior-run file from silently passing
# the hand-off gate when a dispatch fails to write.
HANDOFF_DIR="$ANALYSIS_DIR/handoff/$RUN_ID"
mkdir -p "$HANDOFF_DIR"
# Sub-agents address the SAME dir by a different path (their file tools are rooted at the outputs
# mount in Cowork). Resolve the FULL agent-namespace paths via the script — never hand-splice the
# printed root with a literal skill-name/slug/run-id string yourself (that string-splicing is
# exactly the non-determinism the resolver script exists to remove):
python3 "$SHARED_SCRIPTS/resolve_artifacts_root.py" --handoff-dir-agent \
  --dir-name "market-sizing-${SLUG}" --run-id "$RUN_ID"   # prints HANDOFF_AGENT verbatim
HANDOFF_AGENT="<printed value>"   # use verbatim in OUTPUT_PATH lines
# Sub-agent READ paths for under-outputs artifacts use the SAME agent namespace (relative — the
# sub-agent's file-tool cwd IS the outputs mount on host-loop; an absolute /sessions/... read is denied):
python3 "$SHARED_SCRIPTS/resolve_artifacts_root.py" --analysis-dir-agent \
  --dir-name "market-sizing-${SLUG}"   # prints the dir in the agent namespace
ANALYSIS_DIR_AGENT="<printed value>"   # e.g. inputs.json, validation.json, sizing.json reads
# Ad-hoc scratch (NOT sub-agent hand-off) lives OUTSIDE the promoted outputs/ tree, in a temp dir
# that is safe to both create and reclaim. Use the printed path verbatim in later steps.
STAGING_DIR="$(mktemp -d "${TMPDIR:-/tmp}/market-sizing-${SLUG:-co}.staging.XXXXXX")"

Pass RUN_ID to all sub-agents. Every artifact written to $ANALYSIS_DIR must include "metadata": {"run_id": "$RUN_ID"} at the top level. compose_report.py checks that all artifact run IDs match — a mismatch triggers a STALE_ARTIFACT high-severity warning, blocking under --strict.

Overwrite-in-place — do NOT delete prior artifacts under $ANALYSIS_DIR. It is the promoted outputs/ tree in Cowork, where deleting a user-visible path is unsafe (Cowork can deny it; the parity gate flags it). Each producer writes its artifact fresh via -o every run, and RUN_ID is minted fresh per run — so if a prior run left an artifact a later step doesn't regenerate, compose_report.py's STALE_ARTIFACT check (run_ids must match) catches the mismatch. No bulk rm is needed or wanted.

Step 1: Read or Create Founder Context

python3 "$SHARED_SCRIPTS/founder_context.py" read --artifacts-root "$ARTIFACTS_ROOT" --pretty

Exit 0 (found): Use the company slug and pre-filled fields. Proceed to Step 2.

Exit 1 (not found): Expected on a first run — do NOT mention this check or its exit status to the founder; if you narrate anything first, say only "Let me grab a few basics about the company." Deck/materials carve-out — derive field-by-field, never all-or-nothing (do not ask for what you were already given): if the founder provided materials (a deck, financial model, or a sufficiently detailed description), derive each of the four basics — company name, stage, sector, geography — that the materials state, and skip the gate entirely when all four are in hand. Treat the four independently: deriving three and missing one does NOT send you back to asking for all four. Before gating on a still-missing field, try to infer it from a clear signal in the materials and proceed (noting it as inferred, not founder-stated, so it isn't presented as confirmed): geography from a phone country code or an office address (e.g. a +972 number → Israel), but never from currency alone$ is also CAD, AUD and SGD, and founders everywhere price in USD, so a currency symbol is not a country; stage from an ambiguous fundraise signal (a named round, round size, or "raising our seed" language → the matching --stage value); sector from the product category and ICP. Use AskUserQuestion (NOT plain chat) only for the specific field(s) that genuinely have no derivable or inferable signal — and ask for only those, stating the values you already derived so the founder confirms or corrects rather than re-supplying everything. If AskUserQuestion is genuinely unavailable in the host, do NOT skip the ask and do NOT assume the answer: ask the same question in plain chat, state the options explicitly, and wait for an answer before continuing. The ban above is on asking casually WHILE the tool is available — it is not a reason to stall a host that lacks it. (If none of the four can be derived at all, that reduces to asking for all four.)

Stage is the one field with a real fixed label set — use it verbatim if asking. Options: Pre-seed / Seed / Series A / Series B+pre-seed | seed | series-a | series-b (founder_context.py's VALID_STAGES has 7 values including series-c/series-d/later; on a Series B+ pick, ask a plain-text follow-up for the specific stage rather than defaulting to series-b). Company name, sector and geography cannot take fixed labels — shape each as an affirmative option carrying the inferred/derived value plus a stated-value fallback. Provide at least 2 options. Then create:

--stage is enum-validated (hyphenated, lowercase) — one of: pre-seed, seed, series-a, series-b, series-c, series-d, later. Passing a non-canonical token (e.g. seriesa, pre_seed) is an argparse error and forces a retry — map the founder's answer to one of these 7 values before calling init.

--sector-type is an optional override (also enum-validated, hyphenated): one of saas, ai-native, marketplace, hardware, hardware-subscription, consumer-subscription, usage-based, transactional-fintech, retail. When omitted, founder_context.py auto-derives it from --sector via a small alias table (e.g. "B2B SaaS" -> saas); if the sector doesn't match a known alias, the script emits a runtime warning asking you to set --sector-type explicitly — pick the closest value from the enum above rather than waiting for that warning.

python3 "$SHARED_SCRIPTS/founder_context.py" init \
  --company-name "Acme Corp" --stage seed --sector "B2B SaaS" \
  --geography "US" --artifacts-root "$ARTIFACTS_ROOT"
  # Add --sector-type <value> if the auto-derivation warning fires or the sector
  # doesn't map cleanly to one of the 9 canonical sector-type values above.

Exit 2 (multiple): Present the list, ask which company, re-read with --slug.

Shortened here. Read the whole file on GitHub.

Signals

GitHub stars
34
Forks
3
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
market-sizing-lool-ventures
Source
github.com/lool-ventures/founder-skills