Scenario Sprite Animation

SkillMedia

Lets your agent animate 2D game sprites into walk cycles, attacks, effects like fire or smoke, and seamless looped GIFs.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the Scenario Sprite Animation skill

About this skill

Use when animating 2D game art with Scenario: character walk cycles, idle or attack animations, sprite sheets, animated GIFs, frame-by-frame pixel art animation, VFX loops such as fire or smoke, looping a sprite seamlessly, animating an existing character image, a full animation sheet from one promp

What this skill tells your AI

The instructions your AI receives, as published by scenario-labs/skills in skills/scenario-sprite-animation/SKILL.md and read by ahel’s review.

Overview

A bare prompt to a generic image model for "a sprite sheet, walk cycle, 8 frames" returns a picture of one; a written grid spec on a high-adherence image model returns the sheet (the single-sheet row below). Frame sequences come from routes picked by the control the task needs, not sprite size alone. This skill routes between lanes and carries the frame math; depth lives with siblings: scenario-video (video model families), scenario-consistency and scenario-model-training (one character across an action set), scenario-game-assets (statics, pixel cleanup, upscaling; a sheet whose cells are one character's frames stays here), scenario-refine-loop (iterate to brief). Connection and the core loop: the scenario skill. If a sibling skill named here is missing from your available skills, ask the user to install it (npx skills add scenario-labs/skills --skill <name>); unattended, proceed from tool schemas and flag the gap.

Quick reference

You needRoute
Pixel cycles or VFX, style-locked canvasPurpose-trained animation model (recommend with the sprite-animation need in the user's own words, non-deprecated pick): GIF preview, same-seed spritesheet re-run, slice
An exact grid, chosen frame count, scripted beatsHigh-adherence general image model (recommend with the grid spec in the user's own words; expect ask_user against the purpose-trained pick): one model_run, slice, snap per tile, assemble
A seamless loop from a stillFirst+last-frame video model (recommend with the loop requirement): the same flattened still at both ends plus explicit motion language, extract, seam-check
Poses at chosen momentsKeyframe video model (recommend with the keyframe-pinning need): pin pose stills to frame indices or timestamps, extract
Animate this sprite, larger or styledImage-to-video on the flattened still, prompt anchored in a sprite-animation style with scripted motion beats; model per scenario-video
One character across an action setTurnaround plus reference-to-video per scenario-consistency; reuse references and settings for every action
Frames to sheet, sheet to framesTool models via model_run: model_scenario-video-to-image-seq, model_scenario-grid-maker, model_scenario-image-slicer, model_scenario-image-seq-to-video (fixed first-party ids: each is Scenario's single deterministic tool for its step)

The schema is the contract, read with model_schema_get on the pick: re-discover the generative members each session rather than hardcode them (the tool ids above are fixed). Per-lane schema notes, covering the Retro Diffusion canvas lock and returnSpritesheet seed pairing on the pixel lane, the single-sheet schema shape and prices, the video-lane control surfaces, the alpha routes, and how to read a deprecation tag: references/lane-schemas.md.

The single-sheet lane trades the pixel lane's native pixels, locked canvas and seed pairing for a grid honored as written; the pixel look is emulated, so the snapping pass is part of the lane. Unattended, the task decides between the table's first two rows: an exact grid, frame count or scripted beats in the brief takes the top-ranked general model over specialty.model_id, whose when_general_better names exactly those needs. Write the spec for the pipeline that will read it: rows and columns within the slicer's 6x6 (16 frames as 4x4, never 2 by 8, which is a local cut), a canvas the grid divides evenly, one scale, camera, facing and ground baseline, the beats per frame range with the last frame matching the first, generous margins, no labels. Leave GIF production lines out: an image model returns stills, and a hold is the same pose across extra frames. A background enum names no color, so a solid field is prompt wording; an existing character goes through the reference-image field; with no seed a re-run is a new character, so candidates come from the sample-count parameter and one sheet is approved before slicing. Read the delivered sheet before slicing: count the cells, read the cell edges, and count distinct poses, since a beat range naming a single end state can come back as one held pose. Then slice, seam-check the tiles per step 5, align the body across tiles, snap every kept tile with the same color count and seed, check that the snapped tiles share one size, and assemble. Alignment is a post step, not a prompt: the body drifts between rows (in one run the grounded feet sat 87 px higher in the last row than in the first, and neither anchoring language nor a layout template passed as reference image moved that), so measure each tile's body box, shift it onto one baseline and one horizontal center on a wider canvas so nothing clips, then upload_asset the re-laid sheet and slice it, since frames edited locally never become assets.

For reusable layouts, read the grid manifest and template usage guide. Resolve uploaded assets and publish new versions through the scenario skill's shared asset lifecycle. When none is accessible, the grid builder, run by a maintainer or an agent needing local geometry, creates the selected preset for MCP upload; it does not generate animation. Preserve the user's frame count and canvas over any preset. Templates do not replace the alignment, edge, and loop checks above.

The frames pipeline is tool models run with model_run, not separate MCP tools: a video-to-image-sequence extractor (extractAllFrames, or every Nth frame via frameInterval; the order of its frames is verified, never assumed, see below), a grid maker (images, up to 100, kept in input order; columns, default 3, and rows, computed from the count when unset, so a delivered 2 by 8 is the finished frames at columns: 8; backgroundColor takes a hex string, transparent, black or white), a grid slicer (image plus xSubdivisions and ySubdivisions, 1 to 6 each and defaulting to 2 when omitted, so 6x6 is the cap; no output-size parameter; tiles return in row-major order but not always at cell size, so check one with asset_get, and a resized tile or a sheet past 6x6 means cutting the sheet locally on the grid you counted), a sequence-to-video assembler (model_scenario-image-seq-to-video, a first-party id stable enough to name; fps 1 to 120, loopCount, pingpong, outputFormat gif or mp4), and a pixel snapper (re-detects the art's grid and quantizes to the requested color count, the one palette control in the cleanup chain; priced per tile, so probe one before the set). GIF frame delays quantize to 10 ms, so ask the assembler for an fps that divides 100 (10, 20 or 25; 8 lands at 125 ms and 12 at 83 ms, neither on the grid), or it re-times and duplicates frames to hold the duration. Fps is the only timing lever after generation, so plan frame counts up front and count what came back. Analysis is fine locally: differencing, seam math, edge reads, and window searches need pixels on disk.

Frame order is verified, never assumed. Extracted frames share one createdAt and carry no index in metadata, so their order exists only in the job row's assetIds, and that list has come back out of time order (a confirmed defect at authoring time). Before any stride math, slice, or assembly, pass the ids to the grid maker in returned order and read the contact sheet against the clip's free firstFrame and lastFrame (asset_get) and its motion: a frame that breaks continuity is out of place, and the corrected order is carried as your own id list from then on. Never rebuild order from listings, timestamps, or search.

Worked example: a seamless idle loop from a still

  1. Create a dedicated collection first (collection_create, tool catalog) and collection_add_assets (collection_id, asset_ids) every kept output as it lands. With no still to hand, generate and approve one first: it anchors every later step, so check subject-critical detail against the brief before animating, and budget it apart from the loop (references/lane-schemas.md). Flatten the character onto a plain field (a local edit is fine; assume video inputs drop transparency as reference inputs do) and upload_asset it. Confirm lane and budget with the user; unattended, take them from the task instructions, else the cheapest lane that fits, and dry_run before every paid run.
  2. recommend with the loop requirement in the user's own words, prefer a non-deprecated pick, then model_schema_get. The schema decides the levers: duration and fps enums, a camera lock such as static, an audio toggle, and a native loop flag where present, replacing the end-frame anchor: never send both (dry_run accepts combinations that still fail at run time).
  3. model_run with the same still as first and last frame and a prompt that names the motion explicitly ("weight shifts, cloak sways; the character stays in place"). jobs_wait re-called with pending_job_ids while the job runs.
  4. If the sprite needs alpha, expect no video lane to emit it, so budget a removal pass; verify alpha survived on one downloaded frame before batching: inline display composites it away, and a transparent request can come back in a container without alpha (references/lane-schemas.md). Extract with the video-to-image-sequence tool and verify the frame order first (above). Confirm what was delivered before computing any stride (requests are not always honored): asset_get reports properties.duration, frameRate, nbFrames, width, and height; nbFrames is the delivered frame count the math needs, so read those and leave the rest of the record alone. At authoring time the extractor exposed only extractAllFrames and frameInterval, so confirm with model_schema_get: no start offset, end, or range. Any non-uniform selection (trimmed window, stride off an offset, hand-picked subset) means extractAllFrames: true then subsetting by asset id; a frameInterval that cannot express the window is never a reason to leave the platform. Frame math: extracted count = delivered frames / frameInterval, endpoint-inclusive; where the schema exposes fps, generating at sprite cadence (8 to 12 fps: a cycle typically needs 8 to 16 frames) beats decimating after.
  5. Seam-check against the sequence's own motion, not by eye: a seam reads as a loop only relative to how much the frames already move. Compute the mean absolute pixel delta between consecutive frames (the motion baseline) and the delta across the seam; the loop closes when the seam sits at or below the baseline of the window being looped, not the whole sequence's, since per-step delta swings by orders of magnitude across a scene loop. On a near-static idle, asset_display of first and last frame is enough; on locomotion (hop, walk, attack) eyeballing passes visibly broken loops. If the seam is wide, do not re-run first: image-to-video models re-render off the input still, so frame 0 often diverges from every later frame (a large 0 to 1 delta is the tell) and no re-run fixes a loop anchored to it. Search start and end pairs inside the model's own rendering, excluding frame 0, and take the tightest closing window that still carries real motion. Trim to where the motion is: a duration longer than the action leaves a static tail, and the stride comes from the active window, not the delivered length. Only after a window search fails is a re-run (new seed where the schema has one, else reworded motion) worth the credits. Unattended, allow one re-run, then flag and stop.
  6. Finish and package: a pixel-art target needs a snapping pass, which has traps of its own and is sometimes the wrong move (references/lane-schemas.md). Sheet via the grid maker, transparent background or a color when the brief is opaque; preview via sequence-to-video as GIF; the assembler has no background parameter, so a deliverable GIF takes opaque frames already on its field. asset_download frames with format="png" and the preview with format="gif", and omit format on a video asset, which rejects mp4.

Common mistakes

  • Identical first and last frames with a bare prompt: that asks for a freeze, not a loop; the motion must be named.
  • Treating margin or wording as containment for a traveling effect: what crosses the cell edge is gone, margin gets filled and crossed anyway, and containment wording cancels the throw. Keep the frames whose delivered edge band reads empty; the detached projectile is its own sprite.
  • Assuming every model has a seed: many video schemas and several high-adherence image schemas have none, so reproducibility there is per-asset, not per-run.
  • Counting on an interpolation filter: none exists in the catalog, and fps is the timing lever. Pixelating a non-pixel image is a different job from cleaning pixel art: the catalog carries a pixelate tool (grid size, palette) for the first and the snapper for the second, both found with search (target: "models", public: true, filters tags: ["tool"], query: "pixel" returns both; the descriptions tell them apart).
  • Photo upscalers on pixel frames: they invent texture; use the pixel-preserving routes in scenario-game-assets.
  • Shipping a sheet unstepped: per-frame steps return new ids and order lives only in the lists you pass, so carry the frame-to-output mapping (never rebuild it from listings or timestamps), then compare sheet to preview before delivery.
  • Assembling deliverables locally because the frames are already on disk: analysis pulls the clip down, local assembly then looks free, and the frames never become assets, leaving the collection holding a sheet nobody else can reach.
  • Ping-pong or loop flags set blind: watch the downloaded GIF and match the assembler to it.
  • Softening an attack brief after a moderation block: the filter is the provider's, so try another general image model first per scenario-moderation, grid spec unchanged; rewording the beats loses what the lane exists to honor.

Signals

GitHub stars
681
Forks
82
Last commit
Sep 2026
Advanced
Item type
skill
Key
scenario-sprite-animation
Source
github.com/scenario-labs/skills