Autocue

SkillMedia

Lets your agent record a narrated demo video or screenshots of a live app by driving it through the steps itself.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the Autocue skill

About this skill

Record a demo of the running app that the agent drives itself, a `.mdl` QA test or an ad-hoc walkthrough, as a captioned `.webm` (or a screenshot), trimmed of dead air and ready to attach. Use when asked to demo a feature, show a flow working in the real app, produce a video or screenshots of the

What this skill tells your AI

The instructions your AI receives, as published by dxos/dxos in .agents/skills/autocue/SKILL.md and read by ahel’s review.

"Cue" is the verb. "Cue the projects demo" means: find the committed flow whose name or doc comment matches, serve the app its @app line names, and list its steps. Cueing never opens the browser or runs. "Go" then does all the off-camera prep and replies ready once the start ring is up; clicking it runs the take to the end (see "The protocol").

The agent drives the real app one gesture at a time and the session is recorded. Three things make this different from a Playwright spec, and they are the reasons to reach for it:

  • The next gesture can depend on what the last one rendered. A spec fixes every step up front.
  • Steps with no operation behind them are performable. plugin-chess QA-1 step 1 has no invoke: at all — [op:startGame] has no runtime counterpart, and enabling the plugin is a UI action that manager.enable() from the debug port does not do. A spec cannot execute that flow; this can.
  • A failure is a finding, not a crash. A failed gesture returns an error and the browser stays open, so the next probe can ask why.

It is not a test. Nothing here belongs in CI: there are no assertions, and the output is for a human to watch. When you want a regression test, write a spec — see browser-e2e-tests.

Not to be confused with packages/apps/composer-app/demos/, which arranges several headed windows in a grid on a real desktop for a human to drive via robotjs (its own TODO calls that abandoned). It records nothing and exposes no control channel, so it cannot serve an agent or produce an artifact.

Decide what to produce first

A screenshot is often enough, and always cheaper. Reach for one when the thing being shown is a state rather than a sequence: a layout or styling change, a fixed empty state, a form's validation message, a before/after pair. A video earns its size only when the motion is the point — a drag, a transition, a streaming response, a multi-step flow where order matters. The driver takes screenshots with the same screenshot op, so this is a choice about what to send, not a different setup.

When in doubt, record the session anyway (it costs nothing extra while you are driving) and send only the stills if the video adds nothing.

Manual mode: the user records

The user asks for this; never pick it yourself. The user is recording their own screen and is in the loop at every step. You still drive, and the script is still the .mdl test or walkthrough, but the user decides when each part runs and can change how it runs.

node .agents/skills/autocue/scripts/driver.mjs --mode manual \
  --port 7333 --url http://localhost:4173 --out /tmp/demo

Run it in the background. --mode manual changes the driver in five ways:

  • Headed window, no recorder. The browser opens in the foreground at --width×--height, and the page follows the window, so the user can resize it for their capture. Nothing is encoded, the boot is not cut, and §4, §4b and §5 do not apply.
  • The cursor points before it clicks. Every visible gesture moves the cursor to its target and rests there for a second before the ripple and the click, so the person watching sees what is about to be chosen. --dwell <ms> sets it (default 1000 in manual mode, 0 when recording); a command's own dwell overrides it. Rely on it rather than adding a sleep before each click in a flow; hud: false gestures (off-camera prep) skip it.
  • Cursor, but no pills and no banners. The virtual cursor and click ripple still show what is being clicked. The action feed and the caption banner are suppressed, so they don't compete with the product. --pills on or --captions on brings either back when the user asks.
  • stop leaves the browser open. It answers ok and the driver keeps serving. Closing the window ends the driver. Never close the browser yourself in this mode unless the user asks.
  • The browser profile persists. Every manual session opens the same Chromium profile, ~/.local/state/dxos/autocue/profile by default (--profile <dir> for another). The app's identity, its spaces, and any first-run UI already dismissed carry over, so setup done in one session is not redone on camera in the next. Chromium allows one process per profile. If the driver exits with profile … is in use, the last session's window is still open: close it, or keep using it if its driver is still serving. A clean slate means a new --profile directory. Never delete the default one without asking, since it holds the user's prepared state.
  • A local display is required. The cloud sandbox has none, so this is for a session on the user's machine.

Drive it from a flow script

In manual mode, write the QA flow as a script and run it, rather than issuing one op per turn. Each op over HTTP costs a full agent turn, so a ten-step flow driven op by op leaves the user watching dead air between gestures. A script runs at the app's speed, and you stay the control panel: you choose what to run, edit the script when the user steers, and read the result.

Check for a committed flow first: one that has already run end to end lives in an autocue/ folder at the root of the package it exercises (packages/apps/composer-app/autocue/, or a plugin's own), and running it in place beats rewriting it.

Otherwise the .mdl test comes first, and the script second. A flow binds a spec; it is not one. The test carries the intent a reviewer checks — given, each step's do:/expect:, and every product fact the run surfaces as a note: (a known defect it works around, a model that stayed silent). The script carries only mechanics: selectors, retries, waits, pacing, setup. Writing the test from the script instead loses nothing that matters but duplicates it, and the copy drifts unreviewed. So:

  1. Find the test the demo walks, in the PLUGIN.mdl of the plugin under test. If there is none, draft a test QA-n there (the composer-qa format) before writing a line of script, and show it to the user with the step list.
  2. Copy scripts/flow.example.mjs to /tmp/demo/flow.mjs and write one entry per test step. Each step gets demo, whose methods take the same arguments as the HTTP ops (so the cursor behaves the same), and page, the raw Playwright page. Give each step the do: text as its name, and end it with a wait on what expect: says should appear, so a step that did nothing fails instead of passing.
C '{"op":"steps","file":"/tmp/demo/flow.mjs"}'   # numbered step names; `next` is where `run` resumes
C '{"op":"run","until":4}'                         # runs from `next` through step 4, then stops
C '{"op":"run"}'                                   # carries on to the end
C '{"op":"run","from":3,"until":3}'                # re-runs step 3 alone
C '{"op":"abort"}'                                 # stops the run now, mid-step; `next` stays on that step
  • run re-imports the file every time. Edit the script and run again; the browser and the app's state stay as they are.
  • from/until take a step number (1-based) or a step name. With no from, run resumes where the last one stopped. pace (default 800 ms in manual mode, --pace to change it) is the gap between steps.
  • Every step leaves a screenshot, as <out>/steps/NN-<name>.png, whether it passed or failed. The reply lists each step with ok, the error if it threw, and the screenshot path.
  • On a failure, run stops and leaves next on the failed step. Read that step's screenshot before changing anything. It shows what the page actually held, which is usually a dialog in the way, a collapsed sidebar, or a renamed control. Fix the script, then run again to retry from that step.
  • Background a long run (run_in_background), so the user can interrupt you mid-flow and you can send abort. Give its request no client timeout (no curl --max-time): a timed-out request loses the reply, though the run carries on. status answers where it is at any time — state is setup, cued (play button up), running, done, failed or aborted, with the step number and, once it ends, the same result the run reply carries.

Setup steps and the countdown

Mark off-camera preparation with setup: true (enabling a plugin, picking a model). When run reaches the first step after the setup ones it plays a countdown in the page: a closed ring with a play triangle, then a 3-2-1 inside the ring as it unwinds. In manual mode the ring waits for a click, so the user starts their recorder, clicks it, and the take begins on cue. "countdown": false on run skips it, "wait": false plays it without the click, and the countdown op plays one on demand. It is react-ui-experimental's Countdown (story ui/react-ui-experimental/Countdown): its DOM half, play-countdown.ts, has no imports, and the overlay transpiles and injects that same file rather than keeping a copy, so change the look there.

Browser logs

The driver streams the app's @dxos/log output from the page and its dedicated workers to <out>/app.log, in the NDJSON shape scripts/query-logs.mjs reads (--log <file> to move it, --log off to skip). A bundled app has no vite-plugin-log sink, so this is the only log a vite preview session leaves. HTTP responses with an error status are written alongside it under f: "driver/http". When a step fails for a reason the screen does not explain, such as an agent that never answers, read this before guessing.

Restarting from a step

A flow is written to be picked up at any step, not only replayed from the top. A step that changes app state carries a done check next to its run. done is a quick, read-only test that answers whether the step's outcome already holds, such as whether the project exists or the plugin is on.

C '{"op":"run","from":4,"restart":true}'        # reload the app, bring it to step 4 off camera, run 4 on
C '{"op":"run","from":4,"until":4,"replay":true}' # same without the reload, for a page already loaded
  • restart: true reloads the app, waits for it to be ready (--ready/--settle, as for the boot), then replays. Use it after the page is wedged, after a driver restart, or when a later step needs a clean UI to start from.
  • replay: true only replays. It runs every step before from off camera: no cursor, no pills, no captions, no pacing. A step whose done answers true is skipped, so work already in the profile is not redone. The reply lists replayed steps as ok or skipped, and a replay step that throws stops the run with its screenshot, like any other failure.
  • Write done for every step that creates or changes something. A step without one is always replayed, which is fine for navigation but duplicates a create. Keep done read-only and fast: page.evaluate against app state, or a locator count(), never a gesture.
  • The persistent profile is why this matters. State outlives the session, so the next session usually starts partway through a flow. Restarting from the step the user names is what gets them back on camera quickly.

Single ops are still the right tool for a one-off correction the user asks for on the spot ("click that again"), and for probing a failure before you fix the script.

The protocol

The user's part is two words and one click: "go", then the play button. Everything else is yours.

  1. Cue: stage and wait. On "cue " (or when the user asks for a manual recording), find the committed flow — or, for a new demo, the .mdl test and then the script — and serve the app its @app line names. List the steps by importing the script with node (the driver's steps op needs a driver, and the browser opens only on "go"). Show the numbered list and wait. Nothing runs yet.
  2. "Go": do all the prep, then reply ready. Start the driver (the browser opens) and goto if it is not up, then send run with no bounds, backgrounded and with no client timeout — it holds until the take ends. Poll status until state is cued: every setup step has run off camera and the play button is up, waiting. Then reply with exactly ready and nothing else. If a setup step fails instead (state: failed), say in one line what failed and fix it; there is no take yet.
  3. Play runs to the end. The user starts their recorder and clicks play; the flow runs through its last step unattended, with no more "go"s. Don't poll on a short interval while it runs — the backgrounded run returns when the take ends. Then report each step as passed or failed, in a line or two.
  4. Take steering as it comes. "Do step 3 with a longer title", "skip the settings part", or "go back and open it again" become an edit to the script (or a from/until), then another cue. "Start again from step 4" is run with from: 4 and restart: true. "Run until 4" still works for a user who asks to stop partway. A change the user asks for is not a divergence to report.
  5. Handle failures quietly. Don't narrate passing checks. When a step fails, look at its screenshot, say in a line what went wrong and what you'll change, then fix the script. Retry only when the user says so, since the retry happens on camera.
  6. Finish with the window open. Send stop, and tell the user the browser is still open and that closing it ends the driver.
  7. Commit the flow once it has run end to end. Save it as autocue/<name>.mjs in the package it exercises — the plugin under test, or composer-app for a flow that spans plugins — and commit it, so the next session runs it instead of rediscovering every selector. Commit again whenever a later session fixes it.

Flow metadata

Every committed flow opens with a doc comment that says where it came from and what it needs, so a reader can tell whether it is still in step with its spec:

/**
 * Start a chess game and play the opening.
 *
 * @mdl packages/plugins/plugin-chess/PLUGIN.mdl test QA-1
 * @app composer-app via `moon run composer-app:serve` on :5173
 */
  • @mdl names the .mdl file and the test (or flow) the steps were written from. It is required: a flow with no test behind it is a spec nobody reviews, so write the test first.
  • @app says how to serve the app the flow runs against, including any build flags and environment it depends on.
  • Name each step after its do: text, so the flow and the spec can be read side by side.

1. Get the app running

DX_PWA=false VITE_DX_DISABLE_ANIMATIONS=true moon run composer-app:serve -- --port 4173

VITE_DX_DISABLE_ANIMATIONS=true is not optional here. It turns off animation that runs without a user gesture — the tour's carousel auto-advancing every 10s is the one that bites — and unattended motion defeats §4 entirely: every frame differs from the last, so the trimmer finds no still runs to cull and a 13-minute session stays 13 minutes. Check it took effect the same way §4 does, from --report: a session that sat idle should be almost all stillSeconds.

Wait for ready in. In the cloud sandbox, first read the cloud-sandbox skill — the dev server needs a full dependency build (moon run composer-app:build), and Chromium needs the proxy flags that driver.mjs already applies.

Storybook, when the demo is a component

DX_STORIES=plugins/plugin-assistant,stories/stories-assistant VITE_DX_DISABLE_ANIMATIONS=true \
  moon run storybook-react:serve

DX_STORIES narrows which packages are crawled (see .storybook/main.ts); unset it and the whole monorepo is served from source, which is slower to boot and re-optimizes mid-session. The serve task builds the full package closure first, so budget the same 10+ minutes as an app.

Always record storybook demos in isolation mode. Drive /iframe.html?id=<story-id>&viewMode=story — the story fills the 1728×1080 frame, and the recording carries the component instead of a sidebar, a Controls table and a toolbar that mean nothing to the person watching. The manager is worth one establishing shot at most; it is never where the feature gets demonstrated.

C '{"op":"goto","url":"http://localhost:9009/iframe.html?id=plugins-plugin-assistant-components-chatactivity--sequence&viewMode=story"}'

The id is the CSF path: title lowercased with every / and space turned into -, then --, then the export name in kebab-case (ConnectingMcp → connecting-mcp). Read it off the manager's URL if in doubt.

Four things about storybook that cost a cycle each:

  • Warm the preview bundle before the first isolation load. A cold iframe.html pays the whole Vite dep-scan and holds a light-theme spinner for 15s+ — which lands in the recording. Load the manager once, wait for the story to render, and only then drive iframe.html.
  • eval runs in the manager, not the story. page.evaluate sees the top document, so a story selector resolves to nothing. Reach in explicitly: document.querySelector('#storybook-preview-iframe').contentDocument.querySelector(...). Same for click/text — a Playwright locator does not cross into the iframe.
  • HMR is your edit loop. A source edit re-renders the live story in about 4s, so verify a fix by re-reading the DOM rather than restarting anything.
  • Remount by re-selecting, not reloading. A story with a timer or an animation restarts when it mounts; clicking another story and back is instant, whereas a reload re-bundles.
  • Scope a selector that the thread also matches. A chat's composer and every message already in the thread are all .cm-content, so the bare selector's .first() types into the transcript — silently, since the op still answers ok. Anchor on the container's testid ([data-testid="assistant.prompt"] .cm-content) and read the value back before submitting.
  • clearCaption before touching anything at the bottom of the page. The banner is pinned there, so it sits over a composer or a footer toolbar and swallows the click. Clear it, interact, caption again.

The native desktop app (Tauri)

The same driver drives the desktop app with --target tauri: every op, the overlay, captions, cuts, flow scripts and the trimmer work unchanged. Playwright cannot attach to a WebKitGTK or WKWebView webview, so the driver speaks W3C WebDriver instead — tauri-driver hands the session to WebKitWebDriver, whose pointer and key actions arrive in the page as trusted platform events (menus open, CodeMirror takes typed text). scripts/tauri/page.mjs is a Playwright-shaped page over that session and scripts/tauri/selectors.mjs evaluates Playwright's selector syntax in the page (>>, nth=, text=, role=…[name=…], :visible, :has-text(), :text-is(), :has()), so a flow written for Chromium runs as is. What it does not have: open shadow roots are not pierced, the log tap misses entries logged before the first drain, and there is no page.on('console' | 'response').

Linux (and the cloud sandbox) only. macOS has no WebDriver for WKWebView, so there is no route there; Windows would work through msedgedriver but is untested.

One-time setup:

apt-get install -y libwebkit2gtk-4.1-dev webkit2gtk-driver xvfb ffmpeg bubblewrap socat
cargo install tauri-driver --locked
node .agents/skills/autocue/scripts/tauri/smoke.mjs   # the adapter against stock MiniBrowser, ~5s

Build the app — a release build, and with custom-protocol: a bare cargo build leaves Tauri in cfg(dev), which embeds no frontend and serves blank pages. The frontend is out/composer, so bundle it first:

DX_ENVIRONMENT=dev DX_PWA=false VITE_DX_DISABLE_ANIMATIONS=true moon run composer-app:bundle
moon run composer-app:stage-sandbox-helper        # dx-sandbox beside the binary, for local sandboxes
(cd packages/apps/composer-app/src-tauri && cargo build --release --features tauri/custom-protocol)
node .agents/skills/autocue/scripts/driver.mjs --target tauri --out /tmp/demo
  • The app boots where it always does, its channel's http://localhost:<port>, so goto with no url waits for the ready plank and cuts the boot without navigating; restart reloads that origin.
  • With no DISPLAY the app gets its own Xvfb screen sized to the window in device pixels (--width, --height, --scale), and scripts/tauri/recorder.mjs grabs it with ffmpeg: H.264 while recording, VP9 once on stop, from the last cut. --headed on uses your $DISPLAY instead.
  • --theme works through GTK (GTK_THEME=Adwaita:dark), and the cloud sandbox's proxy through GLib's https_proxy, which the launcher sets; WebKit needs neither of Chromium's TLS flags.
  • The app keeps a profile (~/.local/share/org.dxos.composer), like manual mode's Chromium profile: its identity and spaces carry over between runs. Delete that directory for a first-run take.
  • --app <binary> drives another build; --driver-port moves tauri-driver off 4444.
  • Bundle with VITE_DX_STORAGE=memory for Linux. WebKitGTK cannot hand a worker an OPFS sync access handle (its file-handle IPC is Cocoa-only), so Composer's SQLite store cannot open there and the app stops at a System Error. The app switches the needed WebKit features on itself (src-tauri/src/webkit_features.rs), but the handle is not a feature. With the memory store every launch — and every reload — is a new identity; only localStorage (plugin settings, layout) persists, so a flow that reloads re-selects the first space.
  • Nothing on Xvfb may disable the DMA-BUF renderer. WEBKIT_DISABLE_DMABUF_RENDERER=1 makes the app segfault in AcceleratedBackingStore::update on the first composited frame; the launcher sets LIBGL_ALWAYS_SOFTWARE=1 instead. For a crash, run the app under gdb via a wrapper passed as --app, with libwebkit2gtk-4.1-0-dbgsym from ddebs.ubuntu.com for symbols.
  • Restart the driver after editing scripts/tauri/*: the adapter loads once. Flow scripts reload per run.

2. Start the driver

node .agents/skills/autocue/scripts/driver.mjs \
  --port 7333 --url http://localhost:4173 --out /tmp/demo &

It launches Chromium, opens one recording context, and listens on loopback behind a per-process token (printed at startup and written to <out>/token). Loopback alone is not access control — any page the browser has open can POST here cross-origin in no-cors mode, and the command would run even though the response is opaque to it; a token in a non-safelisted header cannot be set by such a request. Every gesture is one HTTP call, so each is a separate agent turn:

C() { curl -sS -H "x-demo-token: $(cat /tmp/demo/token)" localhost:7333/cmd -d "$1"; echo; }
C '{"op":"goto"}'
C '{"op":"click","selector":"[data-testid=\"treeView.pluginRegistry\"]"}'
C '{"op":"screenshot","name":"01-registry.png"}'
C '{"op":"stop"}'          # closes the context — this is what writes the video

Ops: goto cut click fill type press keys hover drag waitFor text count eval invoke caption clearCaption sleep screenshot run steps status abort stop. In manual mode stop leaves the browser open, and run executes a flow script (see "Drive it from a flow script"). invoke takes key, input and an optional spaceId, and runs the operation through composer.invoke. selector takes any Playwright selector; text selects by visible text instead. Every op answers {ok:true,...} or {ok:false,error} and never kills the driver.

Shortened here. Read the whole file on GitHub.

Signals

GitHub stars
525
Forks
49
Last commit
Sep 2026

ahel review

  • K1binfo
    installs-packages
  • K6info
    bundled executables the agent is told to run
  • K1binfo
    installs-packages (in scripts/tauri/launch.mjs)
  • K1binfo
    installs-packages (in scripts/trim-static.mjs)

Automated review, not a security audit. Ruleset v1+k2.

Advanced
Item type
skill
Key
autocue
Source
github.com/dxos/dxos