deep-interview

SkillDev tools

The deep interview skill is a Socratic clarification loop that runs before planning or implementation. It asks one probing question per round about intent, scope, and boundaries, checks answers against the actual codebase, and scores each response for remaining ambiguity. The loop stops once the score drops below the chosen depth threshold, then produces a written specification other workflows can consume.

Available today. Use it from your connected AI after setup.

Have a vague feature idea or broad request that lacks concrete acceptance criteria.

Then ask your AI: use the deep-interview skill

What your AI can do with it

  • Runs a Socratic interview with one question per round
  • Scores each answer for remaining ambiguity
  • Checks answers against the actual codebase
  • Stops when ambiguity drops below the chosen depth threshold
  • Produces a written specification for other workflows
  • Supports quick, standard, deep, and autoresearch profiles

Getting started

  1. Have a vague feature idea or broad request that lacks concrete acceptance criteria.
  2. Choose a depth profile: quick, standard, deep, or autoresearch.
  3. Start the skill with the idea and the chosen profile.
  4. Answer the single probing question asked each round about intent, scope, and boundaries.
  5. Continue until the ambiguity score drops below the profile threshold, then use the written specification.

What this skill tells your AI

The instructions your AI receives, as published by yeachan-heo/oh-my-codex in skills/deep-interview/SKILL.md and read by ahel’s review.

<Use_When>

  • The request is broad, ambiguous, or missing concrete acceptance criteria
  • The user says "deep interview", "interview me", "ask me everything", "don't assume", or "ouroboros"
  • The user wants to avoid misaligned implementation from underspecified requirements
  • You need a requirements artifact before handing off to ralplan, autopilot, ultragoal, or team </Use_When>

<Do_Not_Use_When>

  • The request already has concrete file/symbol targets and clear acceptance criteria
  • The user explicitly asks to skip planning/interview and execute immediately
  • The user asks for lightweight brainstorming only (use plan instead)
  • A complete PRD/plan already exists and execution should start </Do_Not_Use_When>

<Why_This_Exists> Execution quality is usually bottlenecked by intent clarity, not just missing implementation detail. A single expansion pass often misses why the user wants a change, where the scope should stop, which tradeoffs are unacceptable, and which decisions still require user approval. This workflow applies Socratic pressure + quantitative ambiguity scoring so orchestration modes begin with an explicit, testable, intent-aligned spec. </Why_This_Exists>

<Depth_Profiles>

  • Quick (--quick): fast pre-PRD pass; target threshold <= 0.30; max rounds 5
  • Standard (--standard, default): full requirement interview; target threshold <= 0.20; max rounds 12
  • Deep (--deep): high-rigor exploration; target threshold <= 0.15; max rounds 20
  • Autoresearch (--autoresearch): same interview rigor as Standard, but specialized for $autoresearch mission readiness and .omx/specs/ artifact handoff

Profile max rounds is a hard cap, not a target. Do not continue only to reach a numbered round count. Extra Socratic rigor does not override the active threshold unless the profile/config changes.

If no flag is provided, use Standard.

<Mode_Flags>

  • --autoresearch: switch the interview into autoresearch-intake mode for $autoresearch handoff. In this mode, the interview should converge on a validator-ready research mission, write canonical artifacts under .omx/specs/, and preserve the explicit refine further vs launch boundary for downstream skill intake. </Mode_Flags> </Depth_Profiles>

<Execution_Policy>

  • Ask ONE question per round (never batch multiple interview rounds into one questions[] form)
  • Ask about intent and boundaries before implementation detail
  • Target the weakest clarity dimension each round after applying the stage-priority rules below
  • Treat every answer as a claim to pressure-test before moving on: the next question should usually demand evidence or examples, expose a hidden assumption, force a tradeoff or boundary, or reframe root cause vs symptom
  • Do not rotate to a new clarity dimension just for coverage when the current answer is still vague; stay on the same thread until one layer deeper, one assumption clearer, or one boundary tighter
  • Before crystallizing, complete at least one explicit pressure pass that revisits an earlier answer with a deeper, assumption-focused, or tradeoff-focused follow-up
  • Gather codebase facts via explore before asking user about internals
  • omx explore is deprecated. Use normal repository inspection tools/subagents for simple read-only brownfield fact gathering; use omx sparkshell only for explicit shell-native read-only evidence, and keep ambiguous or non-shell-only investigation on the richer normal path.
  • Always run a preflight context intake before the first interview question
  • For brownfield work, preflight must include doc/context grounding before user-facing questions: inspect applicable AGENTS.md files, README/getting-started docs, relevant docs/ contracts/plans/ADRs, existing .omx/context/ snapshots, and any project-local glossary/context files such as CONTEXT.md or CONTEXT-MAP.md when present.
  • Treat existing repo language as evidence, not authority: if the user uses a fuzzy, overloaded, or conflicting term, surface the specific doc/code wording and ask which meaning should govern before implementation.
  • Cross-check user claims about current behavior against code or documented contracts when discoverable. If docs and code disagree, ask a confirmation question that names both sources instead of silently choosing one.
  • Use scenario-based edge-case grilling when relationships, boundaries, or handoff behavior are unclear: invent one concrete scenario that stresses the ambiguous boundary, then ask one focused question about the expected outcome.
  • Durable docs, glossary, ADR, or memory updates are opt-in and public-safe only. Deep-interview may recommend such updates in the handoff summary, but must not automatically create or dump public docs from interview transcripts unless the user explicitly chooses that as in-scope.
  • If initial context is oversized or would exceed the prompt budget, do not paste or forward the raw payload into interview prompts; request and record a prompt-safe initial-context summary first
  • The oversized initial-context summary gate is blocking: wait for the concise summary before ambiguity scoring, crystallizing artifacts, or any downstream execution handoff
  • The summary must preserve goals, constraints, success criteria, non-goals, decision boundaries, and references to any full source documents so downstream consumers receive a prompt-safe but faithful context
  • Keep total prompt payloads within a safe budget by summarizing or trimming retained history; preserve newest/highest-signal answers and never let raw oversized context crowd out the current question
  • Reduce user effort: ask only the highest-leverage unresolved question, and never ask the user for codebase facts that can be discovered directly
  • For brownfield work, prefer evidence-backed confirmation questions such as "I found X in Y. Should this change follow that pattern?"
  • Route facts before judgment in the Ouroboros style: before presenting a user-facing interview round, classify whether the needed information is a discoverable fact, a fact needing confirmation, or a human decision. The interview is with the human for judgment, not for facts the agent can inspect.
  • When unresolved ambiguity depends on current external best practices, official/upstream guidance, standards, or version-aware behavior, use $best-practice-research as the bounded evidence wrapper before crystallizing requirements or handing off to planning/execution.
  • Use these transcript/spec labels only; never use them as omx question source values, and never replace the runtime source: "deep-interview" contract for user-facing deep-interview questions:
    • [from-code][auto-confirmed] — exact, high-confidence codebase facts from manifests/configs or direct source evidence, with no prescription attached.
    • [from-code] — codebase findings that are useful but inferred, pattern-based, or low/medium confidence and therefore need a confirmation-style user-facing round before being treated as settled.
    • [from-research] — externally sourced facts such as API limits, compatibility, or public documentation; facts only, not decisions.
    • [from-user] — goals, preferences, business logic, scope, non-goals, acceptance criteria, tradeoffs, and any decision-bearing interpretation.
  • Treat [from-code][auto-confirmed] and other non-user fact discoveries as context/transcript updates, not interview rounds: do not call omx question, do not create a pending deep-interview question obligation, and do not increment the user-facing round number for facts the agent can safely establish.
  • Auto-confirm only descriptive facts. If a finding implies what the new feature should do, which pattern it should follow, which tradeoff to accept, or what should stay in/out of scope, route the entire decision-bearing question to the user as [from-user] even when code or research facts are available.
  • In attached-tmux Codex CLI, deep-interview uses omx question as the required OMX-owned structured questioning path for every interview round
  • When invoking omx question through attached-tmux Bash/tool paths, preserve the leader-pane return target by prefixing the command with OMX_QUESTION_RETURN_PANE=$TMUX_PANE (or a concrete %pane value)
  • If you launch omx question in a background terminal, immediately wait for that background terminal to finish and read its JSON answer before scoring ambiguity, asking another round, or handing off
  • Treat answers[] as the primary omx question success contract. For a single interview round, read answers[0].answer; use legacy top-level answer only as a compatibility fallback when needed.
  • If the current runtime is outside tmux and cannot render omx question, use the native structured question tool when available; otherwise ask exactly one concise plain-text question and wait for the answer
  • Re-score ambiguity after each answer and show progress transparently
  • Once ambiguity is at or below the active profile threshold, stop ordinary questioning. Run the practical closure audit: crystallize/handoff when readiness gates pass; otherwise ask only the final closure question needed to satisfy a named gate.
  • Treat max_rounds as a stop cap, not evidence that more rounds are needed.
  • Do not hand off to execution while ambiguity remains above threshold unless user explicitly opts to proceed with warning
  • Do not crystallize or hand off while Non-goals or Decision Boundaries remain unresolved, even if the weighted ambiguity threshold is met
  • Treat early exit as a safety valve, not the default success path
  • Persist mode state for resume safety only through the CLI/programmatic single-writer authority (omx state write/read --input '<json>' --json, backed by src/state/operations.ts). The MCP state server is a read-only projection and must never become a second writer. </Execution_Policy>

Phase 0: Preflight Context Intake

  1. Parse {{ARGUMENTS}} and derive a short task slug.
  2. Attempt to load the latest relevant context snapshot from .omx/context/{slug}-*.md.
  3. Check whether the provided initial context or loaded snapshot is too large for safe prompt use. If it is oversized, the first interview round must ask for a concise prompt-safe summary instead of scoring ambiguity or continuing to downstream handoff.
  4. If no snapshot exists, create a minimum context snapshot with:
    • Task statement
    • Desired outcome
    • Stated solution (what the user asked for)
    • Probable intent hypothesis (why they likely want it)
    • Known facts/evidence
    • Constraints
    • Unknowns/open questions
    • Decision-boundary unknowns
    • Likely codebase touchpoints
    • Relevant repo docs/rules/context inspected
    • Terminology or doc/code conflicts found
    • Prompt-safe initial-context summary status (not_needed, needed, or recorded)
  5. For brownfield tasks, inspect the applicable documentation/rule surface before the first user-facing round. Prefer exact, nearby sources over broad scans:
    • governing AGENTS.md files and template/runtime instruction surfaces that apply to the touched paths
    • README/getting-started docs and relevant docs under docs/, especially contracts, plans, ADR-like records, and workflow docs
    • existing .omx/context/ snapshots, .omx/specs/, and planning artifacts relevant to the slug
    • project-local glossary/context files such as CONTEXT.md, CONTEXT-MAP.md, or context-specific docs when they exist
  6. Save snapshot to .omx/context/{slug}-{timestamp}.md (UTC YYYYMMDDTHHMMSSZ) and reference it in mode state.

Phase 1: Initialize

  1. Parse {{ARGUMENTS}} and depth profile (--quick|--standard|--deep).
  2. Detect project context:
    • Run explore to classify brownfield (existing codebase target) vs greenfield.
    • For brownfield, collect relevant codebase context before questioning.
  3. Initialize state via omx state write --input '{"mode":"deep-interview","active":true}' --json:
{
  "active": true,
  "current_phase": "deep-interview",
  "state": {
    "interview_id": "<uuid>",
    "profile": "quick|standard|deep",
    "type": "greenfield|brownfield",
    "initial_idea": "<user input>",
    "rounds": [],
    "current_ambiguity": 1.0,
    "threshold": 0.3,
    "max_rounds": 5,
    "challenge_modes_used": [],
    "codebase_context": null,
    "current_stage": "intent-first",
    "current_focus": "intent",
    "context_snapshot_path": ".omx/context/<slug>-<timestamp>.md"
  }
}
  1. Announce kickoff with profile, threshold, and current ambiguity.

Phase 2: Socratic Interview Loop

Repeat until ambiguity <= threshold, the pressure pass is complete, the readiness gates are explicit, the user exits with warning, or max rounds are reached. This is a stop condition: below threshold, do not open a new ordinary interview branch.

2a) Generate next question

If the initial context is oversized and no prompt-safe summary has been recorded yet, the next question must be only a summary request. Do not score ambiguity, do not run readiness gates, and do not hand off to $ultragoal, $ralplan, $autopilot, or $team until that summary answer is captured.

Use:

  • Original idea
  • Prior Q&A rounds
  • Current dimension scores
  • Brownfield context (if any)
  • Doc/context grounding notes, including existing terminology, governing rules, and any doc/code mismatch
  • Activated challenge mode injection (Phase 3)

Target the lowest-scoring dimension, but respect stage priority:

  • Stage 1 — Intent-first: Intent, Outcome, Scope, Non-goals, Decision Boundaries
  • Stage 2 — Feasibility: Constraints, Success Criteria
  • Stage 3 — Brownfield grounding: Context Clarity (brownfield only)

Follow-up pressure ladder after each answer:

  1. Ask for a concrete example, counterexample, or evidence signal behind the latest claim
  2. Probe the hidden assumption, dependency, or belief that makes the claim true
  3. Force a boundary or tradeoff: what would you explicitly not do, defer, or reject?
  4. Challenge fuzzy or conflicting terms against the repo's documented language and current code behavior
  5. Stress-test the boundary with one concrete scenario or edge case when a relationship or handoff remains ambiguous
  6. If the answer still describes symptoms, reframe toward essence / root cause before moving on

Prefer staying on the same thread for multiple rounds when it has the highest leverage. Breadth without pressure is not progress.

Maintain a Breadth Ledger across independent ambiguity tracks: scope, constraints, outputs, verification, brownfield integration, and any user-mentioned deliverable tracks. The ledger is a guard, not a mandatory rotation rule: stay deep on the current thread until it has been pressure-tested, then zoom out only when another material track remains unresolved and would change execution.

Maintain a Docs/Terminology Ledger for brownfield interviews:

  • repo docs/rules/context sources inspected, with path references
  • canonical terms already used by the repo and terms to avoid or disambiguate
  • user terms that conflict with docs or current code behavior
  • doc/code mismatches that require a human decision before implementation
  • optional durable-doc follow-ups that are safe to propose but not auto-apply

Detailed dimensions:

  • Intent Clarity — why the user wants this
  • Outcome Clarity — what end state they want
  • Scope Clarity — how far the change should go
  • Constraint Clarity — technical or business limits that must hold
  • Success Criteria Clarity — how completion will be judged
  • Context Clarity — existing codebase understanding (brownfield only)

Non-goals and Decision Boundaries are mandatory readiness gates. Ask about them early and keep revisiting them until they are explicit.

2b) Ask the question

Use the surface-appropriate structured questioning path for every interview round. In attached-tmux sessions, use OMX-owned structured questioning via omx question (this is the required structured-question equivalent and required AskUserQuestion equivalent for deep-interview). Outside tmux, use native structured input when available; otherwise ask exactly one concise plain-text question and wait for the answer. Present:

Round {n} | Target: {weakest_dimension} | Ambiguity: {score}%

{question}

omx question payload guidance for interview rounds:

  • Deep-interview is Socratic: ask one focused round at a time. Do not use batch questions[] to combine multiple interview rounds, even though omx question supports batch forms for other workflows.
  • Use canonical type values instead of authoring raw multi_select flags by hand. type: "single-answerable" is the default for one-path decisions; type: "multi-answerable" is the canonical shape for bounded multi-select rounds. The runtime will keep multi_select aligned with type.
  • Use single-answerable when exactly one answer should drive the next branch, the options are mutually exclusive, or selecting more than one answer would blur the decision boundary. Typical cases: handoff lane selection, choosing the primary failure mode, or confirming which of several competing interpretations is correct.
  • Use multi-answerable when multiple options may all be true at once and you need to capture a bounded set of coexisting constraints, non-goals, risks, or acceptance checks in one round. Typical cases: selecting all out-of-scope items, all success metrics that must hold, or all deployment constraints that apply together.
  • If one selected option would immediately require a follow-up question to disambiguate the others, prefer a single-answerable round now and ask the follow-up next. Do not hide a branching interview tree inside one overloaded multi-select prompt.
  • Keep interview options bounded and concrete. If the valid answers are already known, set allow_other: false; only leave allow_other: true when the interview genuinely needs one user-supplied option that cannot be enumerated in advance.
  • Read answers structurally from the primary answers[] array. For a normal single-round interview response, use answers[0].answer as the source of truth; the top-level answer field is a legacy single-question projection/fallback only.
  • For single-answerable, expect one decisive selection in the value field of answers[0].answer plus its selected-values metadata. For multi-answerable, treat the selected-values field inside answers[0].answer as the source of truth for all chosen constraints/non-goals and preserve the full set in the transcript/spec. In legacy single-question projections, this is equivalent to: For multi-answerable, treat answer.selected_values as the source of truth.

Canonical bounded single-choice payload:

{
  "question": "Which execution lane should own this once the interview is complete?",
  "type": "single-answerable",
  "options": [
    {
      "label": "Plan first",
      "value": "ralplan",
      "description": "Need architecture and test-shape review before execution"
    },
    {
      "label": "Execute directly",
      "value": "autopilot",
      "description": "Requirements are already explicit enough for planning plus execution"
    },
    {
      "label": "Refine further",
      "value": "refine",
      "description": "Clarification is still needed before any handoff"
    }
  ],
  "allow_other": false,
  "other_label": "Other",
  "source": "deep-interview"
}

Canonical bounded multi-select payload:

{
  "question": "Which non-goals must stay out of scope for the first pass?",
  "type": "multi-answerable",
  "options": [
    {
      "label": "No UI redesign",
      "value": "no-ui-redesign",
      "description": "Keep layout and styling unchanged"
    },
    {
      "label": "No new dependencies",
      "value": "no-new-dependencies",
      "description": "Work within the existing toolchain"
    },
    {
      "label": "No API contract changes",
      "value": "no-api-contract-changes",
      "description": "Preserve external request and response shapes"
    }
  ],
  "allow_other": false,
  "other_label": "Other",
  "source": "deep-interview"
}

Canonical answer-shape reminders:

{
  "answer": {
    "kind": "option",
    "value": "ralplan",
    "selected_labels": ["Plan first"],
    "selected_values": ["ralplan"]
  }
}
{
  "answer": {
    "kind": "multi",
    "value": ["no-new-dependencies", "no-api-contract-changes"],
    "selected_labels": ["No new dependencies", "No API contract changes"],
    "selected_values": ["no-new-dependencies", "no-api-contract-changes"]
  }
}

2c) Score ambiguity

Score each weighted dimension in [0.0, 1.0] with justification + gap.

Greenfield: ambiguity = 1 - (intent × 0.30 + outcome × 0.25 + scope × 0.20 + constraints × 0.15 + success × 0.10)

Brownfield: ambiguity = 1 - (intent × 0.25 + outcome × 0.20 + scope × 0.20 + constraints × 0.15 + success × 0.10 + context × 0.10)

Readiness gate:

  • Non-goals must be explicit
  • Decision Boundaries must be explicit
  • A pressure pass must be complete: at least one earlier answer has been revisited with an evidence, assumption, or tradeoff follow-up
  • A practical closure audit must pass: another question would change execution materially, not merely polish wording or chase a narrow edge case
  • If either gate is unresolved, or the pressure pass is incomplete, continue below threshold only with a final closure question that names the unresolved gate and would materially change execution.
  • Treat a low ambiguity score as permission to audit closure, not permission to keep drilling indefinitely. If remaining uncertainty would not change implementation, crystallize the spec instead of opening a new branch.
  • If ambiguity is <= 0.10, another user-facing question is allowed only as that final closure question; otherwise crystallize immediately.

2d) Report progress

Show weighted breakdown table, readiness-gate status (Non-goals, Decision Boundaries), and the next focus dimension.

2e) Persist state

Append round result and updated scores via omx state write --input '<json>' --json. Do not write {mode}-state.json directly and do not use the read-only MCP state projection as a writer.

2f) Round controls

  • Do not offer early exit before the first explicit assumption probe and one persistent follow-up have happened
  • Apply a Dialectic Rhythm Guard: track consecutive non-user fact discoveries and confirmation-style answers ([from-code][auto-confirmed], [from-code], or [from-research]). After 3 consecutive non-user or confirmation answers, the next material user-facing round must solicit direct human judgment ([from-user]) unless the closure audit says the interview is ready to crystallize.
  • Round 4+: allow explicit early exit with risk warning
  • Soft warning at profile midpoint (e.g., round 3/6/10 depending on profile)
  • Hard cap at profile max_rounds; never treat this cap as a desired interview length or quota

Phase 3: Challenge Modes (assumption stress tests)

Use each mode once when applicable. These are normal escalation tools, not rare rescue moves:

Shortened here. Read the whole file on GitHub.

Signals

GitHub stars
33k
Forks
3k
Last commit
Oct 2026

Questions

What is the deep interview skill?
It is a Socratic clarification loop that turns vague ideas into execution-ready specifications by asking targeted questions and scoring remaining ambiguity before planning or implementation.
When should I use it?
Use it when a request is broad, ambiguous, or missing concrete acceptance criteria, or when you want to avoid misaligned implementation from underspecified requirements.
How many questions does it ask?
It asks one probing question per round. The maximum number of rounds depends on the depth profile: quick has 5, standard has 12, and deep has 20.
Advanced
Item type
skill
Key
deep-interview-2
Source
github.com/yeachan-heo/oh-my-codex