Analyzing narrated recordings with talkthrough

SkillFiles & storage

Analyze narrated screen recordings and audio files through the talkthrough MCP server — triage feedback into findings, extract specs/backlogs/action items from recordings, and correlate spoken remarks with logs via wall-clock timestamps. Use when the user mentions a screen recording, screencast, narrated video/audio file, or asks to "watch" a recording and act on it.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Analyzing narrated recordings with talkthrough skill

What this skill tells your AI

The instructions your AI receives, as published by korovin-aa97/talkthrough-mcp in .agents/skills/talkthrough/SKILL.md and read by ahel’s review.

The talkthrough MCP server turns a local video/audio file into queryable structured data: timestamped transcript segments, scene keyframes, OCR'd on-screen text, and wall-clock anchoring. No LLM inside — you bring the reasoning; it brings the evidence. Everything is lazy and token-budgeted: never ask for more than the moment you are analyzing.

Prerequisite

The talkthrough MCP server must be connected (tools like process_media / get_transcript are visible). If not, tell the user to install it: claude mcp add -s user talkthrough -- uvx --python ">=3.11,<3.14" "talkthrough-mcp[diarization,url]" (see the repository README for other clients).

Core workflow

  1. Ingest once: process_media(path) — idempotent by content hash; re-calls on the same file return instantly. Given a public video/audio URL instead of a file, call process_url(url): the source is downloaded once (the only network step; YouTube needs the [url] extra) and kept inside the job, then everything below is identical and local — never download twice; a repeat call serves the stored job. Long videos take minutes and stream progress. The summary gives you job_id, counts, wall-clock, and a transcript preview — do NOT dump anything else eagerly. Multi-person recording (meeting/interview)? Add diarize=true — even when the ask is just "summarize", speaker structure is part of meeting analysis — and — whenever the headcount is known — num_speakers=N (the main accuracy lever): segments get S1/S2/… labels and the summary a talk-time roster. On an already-processed job the amend re-runs ONLY diarization (no re-transcription) — still minutes on long recordings.
  2. Orient: get_transcript(job_id) (paginate via next_start_ms when truncated) or search(job_id, "<distinctive word>") to jump straight to the relevant moments (searches speech AND on-screen OCR text). Multi-word search defaults to match_mode="all_words"; use "any_word" for broader lexical recall.
  3. Evidence per remark: get_moment(job_id, t0-2000, t1+2000) — one call returns the transcript slice + up to 3 unique frames + their OCR text + the wall-clock range. This is the workhorse; describe observed from the returned pixels, never from imagination.
  4. Precision when needed: get_frames(at_ms=...) for nearby keyframes; extract_frame(job_id, at_ms, crop={x,y,w,h}) for an exact instant at native resolution (keyframes capture scene changes + a 1 fps floor, so sub-second moments can fall between them).
  5. Keep verified names: after proving an anonymous label's identity, call label_speakers(job_id, labels={"S1":"Name"}, evidence={"S1":"intro or frame proof"}). Saved names appear in later transcript, moment, and search calls while raw S1/S2 labels remain. If a diarization amend changes labels, those names move to speaker_names_pending_review and stop being identities. Use the stored old-roster context anchors to re-check them. A pending label still in the roster can be confirmed, replaced, or removed; a stale pending label can only be removed with labels={"Sx":null}. Never use a pending name in minutes or search as though it were active. A full force=true rebuild of a job with active or pending names must also use diarize=true; it rebuilds safely and moves every old identity to pending review, while omitting diarization is refused without changing the stored job.
  6. Recall across sessions: list_jobs() — the store persists; a file processed yesterday (even via CLI) is queryable by job_id today.

Timestamps

Every timestamped result carries t_ms (video-relative) and, when the recording start is known, t_wall (ISO 8601 real time). Copy t_wall VERBATIM from the payload — never compute it from t_ms yourself (hand-derived wall-clocks drift by whole hours). Use t_wall to correlate remarks with server/app logs (±30 s grep window). If wall_clock is null or low-confidence, ask the user when the recording started and re-anchor: process_media(path, recorded_at="<ISO 8601>", force=true); when the job already has speaker identities, include diarize=true as required by the safe-rebuild contract.

Packaged workflows (server prompts)

Prefer the server prompts when the task matches — they encode the full method: bug (one recording → evidence-backed GitHub issue draft; silent, narration-free recordings welcome), triage-recording (screencast → findings JSON per the contract in examples/output-contract.schema.json), spec-from-workshop, backlog-from-demo, meeting-actions (audio-only friendly), correlate-with-logs.

Rules of thumb

  • Audio-only jobs (.m4a/.mp3/…): transcript tools work; frame tools error by design — that error is expected, not a failure.
  • Speaker labels are anonymous (S1/S2, ordered by first voice). Mapping them to names is YOUR job: self-introductions, vocatives, the attendees list — and on video jobs the screen check is MANDATORY: for every label you map, get_frames(at_ms=<that label's longest_turn_at_ms from the roster>) and read the meeting-app name plates, the recording's title card, the active-speaker highlight BEFORE asserting the mapping. STT homophones lie about name spellings (spoken "profit" vs on-screen "Prophet") — trust OCR/frames over the transcript for names. State the mapping explicitly and mark unmapped labels "unidentified". Roster name_candidates are raw OCR hints, not identities: they may be UI text, a job title, or somebody else's name. Inspect the cited frame and persist only defensible mappings with label_speakers; never auto-save a candidate. A name_candidates_note on a pre-0.3.1 video job explains that its legacy flat OCR may not yield hints. The job remains readable; regenerate only when useful, with force=true, diarize=true, so old identities become pending review instead of being lost. diarize=true needs the [diarization] extra — its absence produces an actionable install-hint error.
  • Findings/quotes must cite the narrator's exact words + t_ms (+ t_wall when known) + the frame files you actually inspected.
  • Low STT/vision confidence → surface a question; never silently guess.
  • Any narration language works (Whisper auto-detects; the summary reports language + language_probability). Garbled transcript or low/wrong detection → re-call process_media(path, model="large-v3-turbo", force=true) (best multilingual quality) or pin language="…"; domain jargon → pass vocabulary="Term1, Term2".
  • Write digests/summaries for the recording author in the narrator's language; keep quotes verbatim in the original — translate in your own prose only, never inside a quote.

Signals

GitHub stars
29
Forks
4
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
talkthrough-korovin-aa97
Source
github.com/korovin-aa97/talkthrough-mcp