AI Video Editing with Timeline Studio

SkillWeb & browsing

Analyze images, video, speech, motion, products, and websites; route local vision, audio, depth, tracking, matting, identity, and restoration models; auto-edit, replicate, enhance, caption, voice, assemble, annotate, validate, and export editable Timeline Studio projects and videos. Use for reference-video remakes, filter and repeated-shot reconstruction, subject-aware reframing, clip splitting, source-time speed curves, Color Wheels grading, ramps and holds, person or product cutout, person or object outline, authorized face swap, shot and timing reconstruction, timeline markers, chapters, beat cues, review notes and ranges, supplied or web-sourced footage, AI video platform selection, cleanup, highlights, product promotion, website walkthroughs, image-to-video assisted edits, optical-flow editing, depth/2.5D/transition finishing, AI voiceover, subtitles, short-form production, .timeline automation, or editor evaluation.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the AI Video Editing with Timeline Studio skill

What this skill tells your AI

The instructions your AI receives, as published by martindelophy/ai-video-editor in skills/edit-timeline-studio/SKILL.md and read by ahel’s review.

Turn the user's exact editorial request and media into reversible Timeline Studio edits. Keep the editable timeline as the source of truth; never replace it with an opaque one-shot render.

Choose the execution path

  1. On first local use after installation, read references/host-environment.md. If Node.js is unavailable, start with the zero-dependency Shell or PowerShell bootstrap; otherwise run node scripts/setup-host.mjs --check. Agent-driven Chinese and mixed Chinese/English voiceover uses Timeline Studio's owned browser-local Hojo TTS Light 80M two-voice bundle and does not require a separate Python voiceover capability. If language runtimes or dependencies are missing, show the exact installation plan and obtain explicit user approval before install mode; never treat Skill installation as permission to modify the host or download models.
  2. Treat local project-file processing as the default for deterministic editing: inspect media locally, modify the portable .timeline through the command layer or local archive services, and render and verify decoded output when the task changes rendered media. Do not open a browser merely because the editor has a UI.
  3. Treat https://video-editor.ai-creator.top/ as the canonical hosted editor only when the user explicitly asks to use the website, provides no local repository or project path, or requires a hosted-only capability.
  4. When this repository is available, prefer its Agent command layer and local media tools. Start the local server and browser only for a verified UI-only operation that the local project pipeline cannot express and the user has not required a local-only workflow. Read the actual server URL from process output; never assume port 5173.
  5. Prefer the bundled Timeline Studio MCP tools when the host exposes them. Read references/mcp-integration.md, inspect the project, and always call timeline_project_diff before timeline_project_apply with the same revision and operations. The MCP server is a transport over the repository command runner, not a separate editing implementation.
  6. When MCP is unavailable, inspect package.json for an Agent command script. Do not use npm run ... --if-present as capability detection because it can succeed silently. If the command runner exists, read references/command-contract.md, inspect the project, build a versioned plan, run the structural validator, and use project.diff as the authoritative semantic dry run before project.run.
  7. If a required operation is missing from the local contract, state the exact gap. For repository-development work, implement the smallest shared local operation and renderer support before falling back to UI. Otherwise ask before switching to the browser workflow.
  8. Do not claim deterministic or idempotent execution when only UI automation was available. State the limitation and preserve an editable project archive when the UI supports it.

Workflow

For a marker-only request, read references/timeline-markers.md and follow its inspect → plan → validate → diff → apply → inspect workflow through annotation handoff. Markers, chapter cues, ranges, and notes are project annotations; adding them requires no narration, model download, browser session, or video render. The media-production steps below apply only when the user also requests a media edit.

1. Inspect before editing

  • Preserve the user's prompt verbatim as the creative brief.
  • Resolve every referenced asset to an explicit path or URL. Never sweep a directory without approval.
  • Inspect duration, dimensions, audio presence, and media type.
  • Read the current project summary before changing an existing project.
  • Before planning long videos, chapters, musical beat cues, review notes, or edit ranges, read references/timeline-markers.md. Inspect existing annotations and use evidence-backed markers to preserve editorial anchors and decisions where useful. Preserve user annotations and distinguish Agent observations from supplied feedback.
  • Ask only when an unresolved choice materially changes the edit, such as the desired output duration or aspect ratio.
  • For an automatic-editing request, read references/auto-edit-workflow.md. Inspect first, classify the content, goal, and delivery with an explicit confidence level, then ask only the minimum category-specific questions that can change the cut. Never ask for facts discoverable from the media.
  • For a request to reproduce, imitate, recreate, or reverse-engineer a reference video, read references/replication-workflow.md. Classify it as editing-style replication, AI-generation replication, or a hybrid; reconstruct filters, repetitions, source splits, speed curves, transitions, shots, and timing before building; and explicitly resolve whether the authorized original audio track must be retained. Do not start editing until the replication analysis-completeness gate passes. Use current web search to compare AI video platforms only when generation is required, and use lawful web-sourced footage only when the user has not supplied adequate material.
  • Before loading or downloading a model for media analysis, read references/local-model-routing.md. Inventory the repository's existing local and pinned mirrored capabilities, choose the minimum model chain needed for the evidence gap, run inference locally without driving the visible editor UI, and record exact model/runtime/fallback provenance. Never load every available model by default or create a duplicate cache.
  • When the user needs web-sourced footage or asks where downloadable material can be found, read references/web-footage-sourcing.md. Give current, task-specific platform suggestions from live search and rank them by source legitimacy, explicit download support, usage rights, visual fit, quality, and provenance. Keep the skill provider-neutral; never hard-code one platform or brand as the permanent route.
  • For every completed automatic edit, read references/professional-editing-workflow.md. Analyze images directly; analyze video with representative frames, speech/OCR, semantics, and global plus subject-region optical flow. Stabilize before tracking or enhancement. Ask about image-to-video or image-to-image models only after inspection proves that generation is materially useful.
  • When a video plan could benefit from temporal emphasis, shot-matched color shaping, foreground/background separation, a product or person hero treatment, an editable subject outline, or identity replacement, read references/curves-and-subject-effects.md. Consider these capabilities during shot design instead of waiting for the user to name a control. Analyze in source time before retiming, use speed curves only for motivated timing changes, use Color Wheels to preserve skin and product identity while shaping the look, derive cutouts and outlines from verified masks, require explicit authorization for face swap, and validate preview/export parity and temporal stability before delivery.
  • When remote generation is materially required, read references/remote-video-generation.md. Search current official documentation, compare only providers that fit the shot blueprint, obtain approval before any paid or privacy-sensitive job, normalize the asynchronous task and provenance, download expiring output bytes, and add verified results to My assets without automatic timeline placement.
  • For product, brand, marketing-commerce, website-promotion, or other promotional edits, read references/promotion-narrative-workflow.md. Proactively construct an ambitious evidence-backed umbrella narrative rather than a feature list or kinetic-typography montage. Unless the user explicitly requests a teaser, build a complete problem-to-transformation-to-proof-to-CTA arc and actively find several visually distinct cases—normally at least three—each with its own setup, product action, visible result, and connection to the final payoff. Never invent customers, outcomes, metrics, or product behavior to make the story feel larger.
  • For highlight edits and reference replications where emphasis or dramatic impact matters, read references/highlight-tension-workflow.md. Treat saliency as candidate evidence, assign setup/rise/pre-impact/peak/aftershock roles, design a non-flat tension envelope, protect the decisive hero frame, and reject completion when accurate cutting still lacks a dominant payoff.
  • For a website walkthrough or promotional recording, use a supported browser-control skill to inspect and rehearse the authorized journey before capture. Read that browser skill completely before browser actions, then follow references/website-promo-workflow.md. Build and complete a page/flow coverage manifest before drawing product conclusions. When required pages are gated, ask the user to sign in themselves in the selected browser; never request credentials or describe inaccessible behavior as verified. Confirm any consequential external action separately, protect signed-in and personal data, and never claim a real screen recording was captured when only screenshots or static assets were available.
  • Treat tutorial-demo, vlog-event, marketing-commerce, and narrative-documentary as narrated-by-default categories. Preserve and reuse authorized source speech when it already carries the required story; otherwise author the minimum complete narration needed for context, progression, visible result, consequence, and closure, then synthesize it before timing picture. Do not ask whether narration should exist unless the user explicitly requests a source-only, music-only, or silent treatment; ask only for choices that materially affect language, casting, claims, or delivery.
  • For narrated edits, read references/voiceover-workflow.md. Prefer Timeline Studio's owned, pinned browser-local voices for every supported language. Route Chinese and mixed Chinese/English narration through the owned Hojo TTS Light 80M browser bundle and cast one of its two stable built-in reference profiles—晴岚 / zh_f_qinglan or 若溪 / zh_f_ruoxi—for each narrator or character; never install or use Hojo 40M or MeloTTS for this route. Use a remote service, operating-system voice, or unowned runtime only when the requested language or voice is unavailable locally and the user explicitly approves that fallback. Unless the user requests another delivery, choose the warmest natural storyteller-like match from the eligible local profiles and direct a close, conversational performance with meaningful phrasing; never default to a flat, metallic, or mechanical system-voice effect. Generate narration as separate short breath-group segments, not as one monolithic performance to split afterward. Treat a comma as a sufficient default synthesis boundary, prefer several short phrase clips over one long sentence clip, and use the shared desktop/H5 0.4s gap between adjacent voice clips. Lock the complete segmented audio spine before finalizing scene durations, motion, transitions, captions, or picture cuts; adapt and trim visuals to the measured speech and pauses, never the other way around. Treat runtime only as an outcome measurement and do not target, chase, or align to a preset number of seconds.

2. Plan at the supported fidelity

  • For an automatic edit, preserve the prompt and normalize inferred, confirmed, defaulted, and unresolved decisions into an editable brief. Build a source-time decision record with keep/remove/shorten/reorder decisions, reasons, confidence, caption expectations, audio-continuity constraints, and protected content before changing the timeline.
  • For explanatory product, tutorial, or website beats, choose an explicit attention treatment for the single named target: magnify small or dense evidence, underline exact text or numbers, or frame the exact boundary of a control, card, or result. Use at most one supporting treatment with a camera move, and apply the underline or frame only after the camera has stopped.
  • For subject-led shots, choose the smallest justified treatment: use cutout to change context or layer depth, outline to focus attention without removing context, and face swap only when identity replacement is explicitly requested and authorized. Apply a speed curve to a meaningful action phase or editorial beat, not as continuous decoration. Preserve the analyzed source-time mapping when retiming.
  • With the command runner, express edits as declarative operations with stable IDs, seconds, revisions, operation IDs, and preconditions. Run scripts/validate_edit_plan.mjs <plan.json> for transport-shape errors, then run npm run agent -- project.diff <plan.json> to reject unsupported operations and invalid project-specific edits before applying anything.
  • With browser UI only, write a short ordered checklist of visible user intents and expected UI outcomes. Prefer named controls and clip labels; use coordinates only as a last-resort fallback grounded in a current screenshot.
  • Keep main Visuals contiguous. Treat captions, stickers, source audio, voiceover, music, and overlays as timed clips.
  • Apply a one-way caption-to-speech rule: if the project configures or enables any caption, every visible caption must correspond to audible speech. Link transcription captions to the existing spoken source clip, and generate a voiceover for every new narration, explanatory, promotional, or text-led caption. Existing source speech satisfies this rule and must not receive a duplicate voiceover. When captions are absent, narration is still the default for tutorial-demo, vlog-event, marketing-commerce, and narrative-documentary; it remains optional for other categories unless the brief requires it. If a configured caption has no authorized speech route, omit it or stop with the editable project preserved.
  • For narrated edits, synthesize each short breath group into its own physical archived audio file from the outset and bind each caption to exactly one matching audioClipId; do not generate one monolithic narration file and split it after synthesis. Split at sentence-ending punctuation and, by default, at commas, semicolons, colons, em dashes, or another clear spoken pause. Keep a boundary joined only when splitting would create a meaningless fragment or break a proper name, number, URL, or intended bilingual phrase. For free-script generation, keep the first clip at the explicit playhead, append every later clip after the current voiceover-track end, and never reuse an unchanged playhead or 0s; caption-scoped generation stays anchored to its caption. Place adjacent narration clips with the shared desktop/H5 default 0.4s gap, then derive caption timing and picture timing from the accepted audio sequence.
  • Preserve media identity and source-time mapping when moving or trimming clips.

3. Apply safely

  • Save a project version or export a .timeline archive before a destructive batch.
  • Apply one transaction per user-visible intent. Fail the whole transaction when a precondition fails.
  • Never silently substitute missing media, voices, models, fonts, or effects.
  • Keep every result undoable and editable in the normal UI.
  • Do not start a paid or remote generation job without a clear user request.
  • Do not put output.render in a command plan or claim that project.run renders video. Use the separate versioned project.render request for its documented portable subset, and use the browser editor for AI generation or unsupported composition features.
  • For a completed video-editing request, resolve an explicit absolute output directory and create both a portable .timeline project and the rendered result video there. Planning, diagnosis, annotation-only work, and an explicit editor-only handoff are exemptions. Annotation-only delivery needs a newly written, inspected .timeline archive; do not render an unchanged video merely to deliver markers.

4. Verify the result

  • Re-read the timeline summary and compare it with the requested ordering, track placement, and enabled states. For narrated edits, confirm the final project duration follows the locked audio spine rather than an earlier runtime estimate.
  • Preview the opening, every cut or transition, caption boundaries, overlays, and the final frame.
  • Play the timeline continuously across every visual, caption, and audio boundary. The timeline clock must advance monotonically; reject any boundary that stalls, jumps backward, repeats a clip tail, or activates both adjacent half-open clips at once.
  • Check audible behavior, not just visible tracks. Distinguish embedded video audio from explicitly separated source-audio clips and verify mute/link state.
  • When placing stereo or multichannel audio with FFmpeg, apply every intended offset to every channel explicitly. For adelay, use adelay=<milliseconds>:all=1 or provide one delay value per channel; a single value with the default all=false delays only the first channel and can pile every later clip into the other channel at time zero. Before delivery, compare left/right activity in the opening window and around every scheduled speech boundary. Reject channel-only early speech, multiple narration clips stacked at the opening, or undocumented interchannel onset skew.
  • Verify every visible caption resolves to one audible speech clip for its complete active interval. Reject orphan captions, silent linked clips, captions extending beyond speech, duplicate source-speech plus voiceover, or text-only caption delivery.
  • Verify adjacent Agent-generated narration clips are separate physical assets, use short phrase-level breath groups with commas treated as valid default split points, and retain the shared desktop/H5 0.4s of intentional timeline space. Inspect the isolated speech bus at the opening and reject repeated free-script generations that share 0s, reuse an unchanged playhead, or overlap before their scheduled starts. Also reject overlong multi-clause synthesis, a monolithic narration that was merely cut into ranges, meaningless micro-fragments, accidental overlaps, clipped breath/release tails, or picture timing that forced the accepted speech out of its natural cadence.
  • Listen to the complete narration at normal playback speed. Reject cold or mechanical timbre, flat pitch and energy, synthetic word-by-word delivery, rigidly equal pauses, rushed cadence, clipped pauses, unnatural pronunciation, segment-level speed changes, unexplained loudness jumps, or narration that was globally accelerated merely to hit a target duration. Require a warm, human, storyteller-like result with restrained pitch variation, phrase-level emphasis, and natural breath space unless the user explicitly requests another character. For phrase-scoped narration, measure every final stem after all processing; by default target -18 LUFS integrated and no higher than -2 dBTP, require the loudest-to-quietest segment spread to stay within 1 LU, and keep segment LRA within 5 LU unless an intentional exception is documented. Never accept a narration mix from full-program loudness alone, and do not rely on one-pass normalization of short clips as proof of consistency.
  • For final export, verify container, dimensions, duration, decoded frames, visible overlays/captions, and a real audio track.
  • Treat every artifact produced before a product or workflow fix as stale for delivery. A code or Skill edit is not a regenerated result: rerun the affected flow from a fresh project, export a new .timeline and result video under distinct paths, record their modification times and SHA-256 hashes, and compare them with the prior artifacts. Never present, relink, rename, or copy an old render as evidence of the fix.
  • For repeated free-script voice generation, run a minimum three-segment regression through the real UI. Reopen the exported .timeline and require exactly one segment at the initial playhead, every later start to equal the preceding end plus the planned gap within timeline precision, and no overlaps. Then inspect or decode the final rendered audio at the same boundaries; project structure alone is insufficient proof that the mix is correct.
  • When a review build must be unmistakably distinguishable from an earlier render, add a brief explicit verification slate or build marker and disclose it. Remove it from the clean final only after the new artifact has passed freshness, timeline, decode, and hash checks.
  • Reopen and verify the .timeline in Timeline Studio, not only with structural inspection: the main Visuals track must be visible, the first frame must render in Preview, archived media must resolve, and captions/audio/track state must match. Fully decode and verify the rendered video, then return both absolute paths.

Interpret underspecified requests conservatively

  • For “try it,” “open it,” or “let me edit” requests without an editorial brief, start the editor, import only the explicitly named assets, verify automatic placement, and hand off the live editable workspace.
  • Do not invent trims, captions, aspect-ratio changes, AI generation, or exports.
  • Treat an explicit request to “automatically edit,” “clean up,” “condense,” “make highlights,” or equivalent wording as permission to make reversible editorial decisions within the confirmed brief. State consequential defaults, protect category-specific content, and report the decisions; do not treat that request as a mere handoff.
  • Treat persistent onboarding completion, model downloads, remote generation, and destructive reset as separate user decisions.

Learn from every real run

For editor evaluation, regression work, or any run that exposes friction, read references/e2e-evaluation.md. For automatic-editing evaluation, also read references/auto-edit-scenarios.md and use its fixed category cards, clarification checks, hard gates, and adjacent stress variants. Capture the attempted action, observed result, evidence, fallback, and verification. Classify the finding as product, browser-control, environment, or skill guidance. Update the smallest relevant skill instruction or reference, validate the skill, reinstall the local copy, and rerun the affected scenario plus adjacent smoke tests. Never weaken an assertion merely to make a test pass.

Capability boundaries

Read references/host-environment.md for host dependency checks and approved installation, and references/voiceover-workflow.md before Agent-driven narration or pre-voiceover generation.

Shortened here. Read the whole file on GitHub.

Signals

GitHub stars
808
Forks
101
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
edit-timeline-studio
Source
github.com/martindelophy/ai-video-editor