Structured-Output-Picker — 在解码处约束,还是在校验处约束?

SkillDev tools

This skill helps your AI decide how to keep model answers in a fixed, predictable structure when other code needs to read them. It compares three ways to enforce that structure: constraining the model as it writes (Outlines), checking the output and retrying (Instructor), and grammar rules (Guidance). It also helps pick what happens when output comes out wrong: stop with a hard failure (Assert) or retry softly (Suggest).

Available today. Use it from your connected AI after setup.

After adding it, describe a case where your AI's answers are read by other code and ask which enforcement approach and failure stance fit best.

Then ask your AI: use the Structured-Output-Picker — 在解码处约束,还是在校验处约束? skill

What your AI can do with it

  • Compare Outlines, Instructor, and Guidance as ways to enforce structured output
  • Recommend whether to constrain output while the model writes or validate and retry afterward
  • Choose a failure stance: hard-fail with Assert or soft-retry with Suggest
  • Advise when model output is parsed or typed by code that depends on it

What this skill tells your AI

The instructions your AI receives, as published by agentsope/skillalchemy in skills/agentsop-structured-output-picker/SKILL.md and read by ahel’s review.

One-liner: Three local libraries (Outlines, Instructor, Guidance) plus provider-native structured outputs all "make the model emit valid structure", but they enforce at different points and fail differently. Pick by how costly a malformed output is and whether you control the decoder. Then pick the failure stance — Assert (hard-fail + retry) vs Suggest (soft nudge, degrade gracefully) — borrowed from DSPy's constraint primitives.

This is an ENHANCE overlay. The four enforcement mechanisms each have a working local skill; what no single one provides is the cross-library *which-one

  • how-to-handle-failure* decision. That gap is hit every time an LM output is consumed by code. For what shape the content should take (code vs JSON vs prose), descend first to [[agentsop-output-format-by-model]] — this skill assumes the shape is already chosen and asks only how to enforce it.

1. 何时激活 (When to activate)

Activate after you have decided the output shape (via [[agentsop-output-format-by-model]]) and the answer was "a typed/validated object", and before you write the parsing code.

TriggerSignal
LM output feeds a parserjson.loads(resp) / Model.model_validate(...) is in the next line of code
You picked a typed shapeformat-by-model said "JSON / Pydantic / typed field" — now: who enforces it?
Repeated parse failuresJSONDecodeError, ValidationError, truncated/extra-prose responses in logs
Library is already installedoutlines, instructor, or guidance is in the env and you must choose between them
An enum / regex / range must holdoutput must be one of N labels, a valid date, a bounded int
You must decide failure stance"if the model returns garbage, do I crash, retry, or accept-and-flag?"

Anti-triggers (skip this skill):

  • The content is code / multi-step reasoning / long prose — go back to [[agentsop-output-format-by-model]]; enforcing a JSON grammar on code is the headline anti-pattern there (Aider 61%→20%). Enforcement strength is the wrong question when the shape is wrong.
  • No code consumes the output yet (one-shot exploration).
  • The shape is one token / one number — any parser works; no library needed.

2. 核心心智模型 (Core mental model)

2.1 The axis: where is the constraint applied?

   PROMPT ──────► DECODE ──────► RAW TEXT ──────► VALIDATE ──────► TYPED OBJECT
     │              │                                │
     │         constraint at DECODE             constraint at VALIDATE
     │         (grammar masks tokens)           (parse, check, retry on fail)
     │              │                                │
  ask nicely    Outlines / Guidance /          Instructor / DSPy Suggest /
  (weakest)     provider strict-mode           hand-rolled retry loop
                  ↑                                  ↑
            CANNOT emit invalid               CAN emit invalid, then
            structure — masked at             catches it and re-asks
            the logit level                   with the error injected

The pick is governed by one question: how costly is a malformed output?

  • A malformed output is cheap to recover from (one extra API round-trip is fine; the model is strong; failures are rare) → constrain at validate (Instructor-style retry). Simpler, model writes more naturally, no decoder access needed.
  • A malformed output is expensive or impossible to recover from (no retry budget, hard real-time, the output must be in a fixed enum/grammar, a single bad token corrupts a batch job) → constrain at decode (Outlines / Guidance grammar, or provider strict-mode). The model cannot emit invalid structure.

2.2 Two prerequisites that gate the choice

  1. Do you control the decoder? Grammar/token-masking (Outlines, Guidance) requires logit access — i.e. local/open weights (Transformers, vLLM, llama.cpp) [[outlines]] [[guidance]]. Closed API models (GPT, Claude) expose only their own native structured-output / tool_use; you cannot bolt Outlines onto them. Instructor works on top of the API by parse-and-retry [[instructor]].
  2. Is the content code-shaped? If yes, stop — see §1 anti-triggers and [[agentsop-output-format-by-model]]. Grammar-constraining code yields valid JSON containing degraded code: enforcement cannot buy back content quality.

2.3 Failure stance is orthogonal — Assert vs Suggest

Independent of which library, you choose what happens on a constraint violation. DSPy names the two stances [dspy-sop-skill §Constraint primitives; arxiv.org/pdf/2312.13382]:

StanceBehavior on violationUse when
Assert (hard)retry up to N; then raise / haltDev-time bug-catching; downstream cannot tolerate a bad value; correctness > availability
Suggest (soft)retry with error injected; then log + continue with best effortProduction; partial result beats no result; availability > strict correctness

Decode-time grammar (Outlines/Guidance) is itself a hard guarantee on shape — but semantic checks (range, cross-field, business rules) still need an Assert/Suggest stance layered on top.


3. SOP 工作流 (Decision workflow)

Three ordered steps. Each presupposes [[agentsop-output-format-by-model]] already said "typed/validated object".

Step 1 — Decide enforcement strength (how costly is malformed?)

malformed output cost?
   │
   ├── cheap to recover (retry OK, strong model, rare failures)
   │        └──► VALIDATE-time enforcement   (weakest sufficient — prefer this)
   │
   ├── must never escape (enum/regex/grammar is a hard contract)
   │        └──► DECODE-time grammar          (only if you control the decoder)
   │
   └── no retry budget / hard real-time / batch where 1 bad token poisons many
            └──► DECODE-time grammar          (or provider strict-mode if API)

Heuristic: start at the weakest layer that meets the cost constraint. Validate

  • retry is cheaper to build, keeps the model's generation natural, and needs no decoder access. Escalate to decode-time grammar only when retry is too costly or the contract is absolute.

Step 2 — Pick the library (gated by decoder access + provider)

You haveOutput goes toPick
Closed API model (GPT/Claude) + Pydantic schematyped object, retries acceptableInstructor — validate-time, Pydantic + auto-retry [[instructor]]
Closed API model + provider supports nativetyped object, want zero extra depsprovider native structured outputs / tool_use (then Instructor or hand-parse on top)
Local/open weights + must guarantee JSON/regex/enumhard contract, no retry budgetOutlines — decode-time grammar/regex/JSON-schema [[outlines]]
Local/open weights + interleaved gen + control flowmulti-step / fill-in-the-middle / mixed text+constrainedGuidance — decode-time + Pythonic control flow [[guidance]]
Any model + only semantic checks needed (shape already valid)range / cross-field / business rulesretry loop with Assert/Suggest stance (DSPy or hand-rolled)

Step 3 — Choose the failure stance (Assert vs Suggest)

  • Default Suggest in production (degrade gracefully; log the violation).
  • Use Assert in dev/test and on values whose corruption is unacceptable downstream (IDs that index a DB, money amounts, irreversible actions).
  • Decode-time grammar covers shape for free; still wrap semantic checks in a stance. Set a finite retry cap on either stance — an unbounded retry loop is an outage.

4. 操作模型 (Picker table — task → mechanism + stance)

#TaskMechanismStanceRationale / Evidence
1Extract invoice fields from GPT-4o, retry on missInstructor (Pydantic + max_retries)SuggestAPI model → validate-time; auto-retry on ValidationError [[instructor]]
2Classify into a fixed 4-label enum on a local LlamaOutlines regex/choiceAssertenum is a hard contract; one masked decode guarantees it [[outlines]]
3Local model must emit schema-valid JSON, no retry budget (batch)Outlines JSON-schemaSuggest (log shape always holds)decode-time guarantee; a bad token in a 100k batch is too costly to catch later [[outlines]]
4Interleave reasoning text + a constrained {action, arg} block, localGuidance (control flow + grammar)SuggestGuidance interleaves free text and constrained spans in one program [[guidance]]
5Closed API, want typed output with zero new depsprovider native structured outputs / tool_useAssert on parseprovider strict-mode guarantees schema validity (not content quality)
6Output shape already valid; need 0 <= score <= 1 and start < endretry loop, no grammarAssert (dev) / Suggest (prod)semantic, not shape — grammar can't express cross-field; needs a check + stance [dspy 2312.13382]
7Local model, must match a date/email regexOutlines regexAssertregex is a decode-time native; cheaper than parse-and-retry [[outlines]]
8Streaming partial Pydantic objects from an API as they arriveInstructor streamingSuggestInstructor streams partial validated objects [[instructor]]

5. 困境决策案例 (Dilemma cases / worked examples)

Case A — "Instructor keeps retrying and burning tokens; should I switch to Outlines?"

Trigger: An extraction service on GPT-4o uses Instructor with max_retries=5. Latency p95 spiked; logs show 15% of calls retry ≥2× on the same nested schema.

Constraints:

  • Closed API model — no decoder access, so Outlines is not available for GPT-4o [[outlines]] [[instructor]].
  • The schema is deeply nested with several free-text fields.
  • The cost being paid is retries, i.e. malformed output is currently expensive.

Decision steps:

  1. Outlines is off the table for an API model (Step-1 prerequisite §2.2(1)). The real lever is reducing the validate-time failure rate.
  2. Inspect what fails. If the model emits valid JSON but wrong content, no enforcement layer helps — that's a prompt/format problem; check whether a free-text field is being squeezed into JSON ([[agentsop-output-format-by-model]] mixed-content trap).
  3. If the failure is shape (extra prose, truncation): switch the API call to provider native structured outputs / tool_use (strict-mode guarantees schema validity at the API), then keep Instructor only for the Pydantic typing layer. Retries collapse because shape is now guaranteed upstream.
  4. Lower max_retries to 2 and flip the stance to Suggest with a logged fallback object, so a stubborn case degrades instead of inflating p95.

Outcome: Don't "switch to Outlines" (impossible here). Move shape-enforcement to the provider's native layer, keep Instructor for typing, cap retries, Suggest.


Case B — "Local model, 200k-row batch extraction; a few rows come back malformed"

Trigger: Nightly batch over a local vLLM-served model. ~0.3% of rows produce JSON that fails to parse, poisoning the downstream load.

Constraints:

  • Local weights → decoder access available (Outlines/Guidance both viable).
  • Batch job, no interactive retry budget — re-running the whole batch is the only "retry", which is hugely expensive. Malformed output is very costly.
  • The structure is a flat JSON schema; no interleaved free text.

Decision steps:

  1. Step-1: malformed output is expensive and there is no per-row retry budget → decode-time grammar. This is exactly the case validate-time loses.
  2. Step-2: local weights + pure JSON shape + no interleaving → Outlines JSON-schema. (Guidance would be the pick only if rows needed interleaved free-text + constrained spans, Case D-style.) [[outlines]] [[guidance]]
  3. Shape is now guaranteed per row — the 0.3% parse failures go to zero by construction. Stance for shape becomes moot.
  4. Layer a Suggest semantic check (e.g. amount >= 0) and write violators to a quarantine table rather than failing the batch — availability of the 99.7% beats halting on outliers.

Outcome: Decode-time Outlines grammar removes the shape failures the retry model couldn't afford to catch; a Suggest semantic check quarantines the rest.


6. 反模式与边界 (Anti-patterns & boundaries)

Anti-patterns

  1. Grammar-constraining when a retry suffices. Reaching for Outlines/Guidance on a strong API model with rare failures buys complexity (and may be impossible — no decoder access) when an Instructor retry would have been two lines. Start at the weakest sufficient layer (§3 Step 1).
  2. Assert where Suggest is enough → over-rejection. Hard-failing the whole request because one optional field violated a soft preference throws away a usable answer. In production, default to Suggest; reserve Assert for values whose corruption is unacceptable [dspy 2312.13382].
  3. Enforcing structure on code/prose content. The headline cross-skill anti-pattern: a JSON grammar produces valid JSON containing degraded code. Enforcement cannot recover content quality — fix the shape first via [[agentsop-output-format-by-model]].
  4. Assuming provider strict-mode fixes content. Native structured outputs guarantee schema validity, not field correctness — same lesson as Aider's strict-mode JSON test [[agentsop-output-format-by-model]].
  5. Unbounded retry loop. Validate-time enforcement without a finite cap turns a flaky field into an availability outage. Always cap N.
  6. Stacking Outlines on a closed API model. Token-masking needs logits; GPT/Claude expose none. Mixing the mental models wastes a debugging cycle (§2.2(1)).
  7. Using a grammar for cross-field/semantic rules. Grammars constrain token shape, not relationships like start < end or "id exists in DB". Those need a validate-time check + stance, regardless of how shape was enforced.
  8. Picking the library before deciding strength + stance. The library is the third decision, after "how costly is malformed" and "Assert or Suggest".

Boundaries (when this skill doesn't apply)

  • Output is code / reasoning / prose[[agentsop-output-format-by-model]], not here.
  • One token / one number / yes-no → any parser; no enforcement library.
  • No code consumes the output → one-shot exploration, defer.
  • Provider contract is fixed (must emit a specific webhook JSON) → enforcement is mandatory by definition; the only remaining choice is the failure stance.
  • The fix is prompt-level (model emits valid-but-wrong values) → no enforcement layer helps; this is a prompt/format problem.

7. 跨框架对照 (Ecosystem cross-reference)

Where each mechanism sits on the decode vs validate axis, and what it guarantees. Cross-link [[agentsop-output-format-by-model]] for the prior shape decision.

MechanismConstraint pointNeeds decoder access?Works on closed API?Native failure modelBest for
Outlines [[outlines]]Decode (token mask: regex / CFG / JSON-schema)Yes (Transformers/vLLM/llama.cpp)NoCannot emit invalid shape — no failure to handle for shapeLocal models; hard enum/regex/JSON guarantee; batch w/o retry budget
Instructor [[instructor]]Validate (parse → Pydantic → auto-retry)NoYesCatches ValidationError, re-asks with error; streaming partialsAPI models; Pydantic typing + graceful retry
Guidance [[guidance]]Decode (grammar) + interleaved control flowYesNoConstrained spans can't be invalid; free spans unconstrainedLocal; interleaved text + constrained blocks; multi-step programs
Provider native (OpenAI structured outputs / Anthropic tool_use)Decode (provider strict-mode)N/A (provider-side)Yes (that provider only)Schema-valid guaranteed; content quality notAPI models; zero extra deps; tool-call args
DSPy Assert/Suggest [dspy 2312.13382]Validate (assertion + backtrack)NoYesAssert raises after N; Suggest logs + continuesThe stance layer on top of any of the above
            DECODE-TIME                         VALIDATE-TIME
   (guarantee shape, need logits)        (catch + retry, model-agnostic)
   ┌─────────────┬──────────────┐        ┌──────────────┬─────────────┐
   │  Outlines   │   Guidance   │        │  Instructor  │ DSPy Assert/ │
   │ (regex/CFG/ │ (grammar +   │        │ (Pydantic +  │   Suggest    │
   │  JSON-sch.) │  control flow│        │  auto-retry) │ (stance)     │
   └─────────────┴──────────────┘        └──────────────┴─────────────┘
   ┌────────────────────────────┐
   │ Provider native strict-mode│  (decode-time, but provider-side; API only)
   └────────────────────────────┘

How the two skills compose:

[[agentsop-output-format-by-model]]  →  decides the SHAPE   (code? JSON? prose? typed?)
                                       │
                  shape == "typed/validated object"
                                       ▼
[[agentsop-structured-output-picker]] (this) →  decides ENFORCEMENT
                                       │
              Step1 strength → Step2 library → Step3 Assert/Suggest

[[agentsop-output-format-by-model]] already lists Outlines/Guidance under "grammar libs" and provider tool_use under its §7 — it says choose the format first, then let a lower layer enforce it. This skill is that lower layer made into a decision.


Quick reference card

┌────────────────────────────────────────────────────────────────────┐
│                  STRUCTURED-OUTPUT ENFORCEMENT CARD                 │
├────────────────────────────────────────────────────────────────────┤
│ 0. Shape already "typed object"? If not → [[agentsop-output-format-by-model]]│
│ 1. How costly is a malformed output?                               │
│      cheap to recover  → VALIDATE-time (Instructor / retry)        │
│      must never escape → DECODE-time grammar (need decoder access) │
│ 2. Pick library:                                                   │
│      API model + Pydantic + retry ok   → Instructor                │
│      API model, zero deps              → provider native           │
│      local + hard enum/regex/JSON      → Outlines                  │
│      local + interleaved text+constr.  → Guidance                  │
│      only semantic/cross-field checks  → retry loop + stance       │
│ 3. Failure stance:                                                 │
│      prod default → Suggest (log + continue, capped retries)       │
│      dev / unrecoverable value → Assert (raise after N)            │
├────────────────────────────────────────────────────────────────────┤
│ NEVER:                                                             │
│  • Bolt Outlines/Guidance onto a closed API model (no logits)      │
│  • Grammar-constrain code/prose (valid JSON, degraded content)     │
│  • Assume strict-mode fixes content correctness                    │
│  • Run an uncapped retry loop                                      │
│  • Assert where Suggest suffices (over-rejection)                  │
└────────────────────────────────────────────────────────────────────┘

引用源 (Citations)

Source lib skills (local):

  • [[outlines]] — decode-time grammar/regex/JSON-schema; local models (Transformers/vLLM/llama.cpp); ~/.claude/skills/outlines/SKILL.md.
  • [[instructor]] — validate-time Pydantic + auto-retry + streaming partials; OpenAI/Anthropic; ~/.claude/skills/instructor/SKILL.md.
  • [[guidance]] — decode-time grammar + Pythonic multi-step control flow; local; ~/.claude/skills/guidance/SKILL.md.

Constraint-stance source:

  • DSPy Assert vs Suggestdspy-sop-skill/SKILL.md §Constraint primitives; DSPy Assertions paper [arxiv.org/pdf/2312.13382]; [dspy.ai/learn/programming/7-assertions/].

Sibling Phase-D skill (cross-link, not duplicated):

  • [[agentsop-output-format-by-model]]d-output-format-by-model-skill/SKILL.md. Decides the shape; this skill decides enforcement. Anchors the "strict-mode ≠ content quality" and "don't grammar-constrain code" claims (Aider code-in-JSON 61%→20%).

Provider docs (named, re-verify before pasting code, May 2026):

  • OpenAI structured outputs: [platform.openai.com/docs/guides/structured-outputs].
  • Anthropic tool_use: [docs.anthropic.com/en/docs/agents-and-tools/tool-use].

All source SKILLs read on 2026-05-20. This overlay introduces no API absent from the sources; see references/R1-source-evidence.md.

Signals

GitHub stars
398
Forks
21
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
agentsop-structured-output-picker
Source
github.com/agentsope/skillalchemy