SDK e2e Test Creation

SkillAI & models

Plans and scaffolds e2e tests in packages/sdk/e2e for a new or changed public SDK API. Use when adding or modifying SDK functionality that is exposed to consumers. Enforces happy / sad / error coverage, deterministic model-output assertions, desktop/mobile/Electron consumer coverage, smoke-suite selection, and local validation with run:local.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the SDK e2e Test Creation skill

What this skill tells your AI

The instructions your AI receives, as published by tetherto/qvac in .agents/skills/qv-sdk-e2e-create/SKILL.md and read by ahel’s review.

Plan and scaffold e2e tests in packages/sdk/e2e for a new or changed SDK feature exposed through the public API.

When to use this skill

Applies to SDK changes in packages/sdk/ that touch the public API surface.

Use when:

  • Adding a new public SDK function, model type, or capability.
  • Changing an existing public SDK API in a way that affects runtime behaviour.
  • User invokes /qv-sdk-e2e-create.
  • User asks to "add e2e tests for " or similar.

Do NOT use for:

  • Internal refactors that don't change the public surface.
  • Unit tests inside packages/sdk/ (this skill covers only the e2e suite under packages/sdk/e2e).

Approach

Investigate first, then propose a concrete plan. Only ask the user for information that cannot be recovered from code or context.

  1. Read the feature. Identify the new/changed exports, inputs, return type, model dependencies, and any existing examples under packages/sdk/examples/ or tests under packages/sdk/e2e/tests/.
  2. Find comparable tests. Look for an analogous existing feature in tests/test-definitions.ts and its executor. Mirror its style unless there's reason to deviate.
  3. Determine model-output testability (see §"Model-output strategy"). Propose a specific validator and a specific prompt/input that makes the output deterministic enough to assert.
  4. Draft happy / sad / error cases as a concrete test-definition sketch.
  5. Decide executor placement, consumer registration, and mobile constraints (see §"Placement and consumer constraints").
  6. Select at most one smoke candidate (see §"Smoke policy").
  7. Present the plan to the user. Include: feature summary, chosen validators with rationale, test definitions sketch, placement decision, desktop/mobile/Electron registrations, platform concerns, smoke pick. Ask clarifying questions only where genuine ambiguity remains (e.g. expected model behaviour on an edge case, preferred tolerance for a numeric-range).
  8. After approval, scaffold the files and prompt the user to run locally with the exact commands for every consumer. Always run run:local:electron --filter <feature>- to verify either the handler or an intentional skip for Electron/Snap.

Model-output strategy

For any feature that invokes a model, do not default to shape-only checks. type validation proves nothing about model correctness and must be a last resort.

Pick the strongest achievable strategy:

  1. Exact-ish output — constrain the prompt + deterministic params (temperature: 0, fixed seed, top_k: 1) so a known token must appear. Assert with contains-all. Example: prompt "Reply with only the word APPLE." → assert result contains APPLE.
  2. Closed-set — enum-style prompt (known set of valid answers) → contains-any.
  3. Numeric range — a score, similarity, duration, or embedding magnitude with known bounds → numeric-range. Pick bounds tolerant to minor model drift.
  4. Regex structure — structured output (JSON keys, date format, language tag) → regex. Keep the pattern anchored and stable.
  5. Shape-only fallbacktype with minLength. Flag as weak coverage in the plan.
  6. Error paththrows-error with a substring that is stable across SDK versions.
  7. Custom function — use for deterministic but non-trivial checks like cosine similarity against a reference vector.

If the model's output is inherently non-deterministic and cannot be constrained, say so in the plan and justify why shape-only or range-based coverage is the best achievable — do not silently ship a weak assertion.

Happy / sad / error minimum

Every public-API feature MUST have at minimum:

  • Happy path — valid input, canonical output, strongest assertion achievable.
  • Sad path — boundary or edge case that must still succeed (empty input, minimum ctx, longest accepted input, unusual but valid locale, streaming vs non-streaming).
  • Error path — invalid/malformed input, missing asset, or exceeded constraint. Must throw with a matchable message via throws-error.

More cases are encouraged for multi-branch features.

Placement and consumer constraints

Executor placement (from packages/sdk/e2e/AGENTS.md):

  • Pure SDK API, no Node stdlib, no RN APIs → tests/shared/executors/.
  • Needs node:fs, node:path, process.cwd(), or other Node-only APIs → tests/desktop/executors/.
  • Needs RN Platform, bundled assets, or anything specific to React Native → tests/mobile/executors/.

Never import node:* from tests/shared/ or tests/mobile/.

Placement under tests/shared/executors/ does not register an executor automatically. Explicitly check and update every compatible consumer:

  • tests/desktop/consumer.ts
  • tests/mobile/consumer.ts
  • tests/electron/consumer.ts — also used by the Snap consumer

If a consumer cannot support the tests, add a documented SkipExecutor; do not leave scheduled tests without a matching handler, which fails at runtime with No handler found.

Mobile concerns to address in the plan:

  • Memory — can the target device RAM hold the model? If not, propose a smaller model variant on mobile, or a SkipExecutor entry.
  • Filesystemnode:fs is unavailable. Assets must be bundled via qvac-test.config.jsconsumers.mobile.assets.patterns.
  • Platform-specific limitations — known iOS/Android issues (OOM, missing native lib, backend unsupported). Add a SkipExecutor at the top of tests/mobile/consumer.ts with a clear reason.

If the feature cannot run on mobile at all, document the skip reason. Evaluate Electron/Snap compatibility independently rather than treating desktop-only coverage as automatic.

Smoke policy

  • Only tag suites: ["smoke"] if the feature has no existing smoke coverage.
  • Cap at 1-2 smoke tests per feature.
  • Pick the happy path with the most meaningful assertion (not shape-only).
  • Must be deterministic, fast, and stable on both desktop and mobile. Verify before tagging.
  • If no test meets the bar, do not tag any and flag it explicitly in the plan.

Scaffolding templates

Test definition (tests/<feature>-tests.ts)

import type { TestDefinition } from "@qvac/test-suite";

export const <feature>Tests: TestDefinition[] = [
  {
    testId: "<feature>-happy",
    params: { /* canonical input */ },
    expectation: { validation: "contains-all", contains: ["EXPECTED_TOKEN"] },
    suites: ["smoke"], // only if this test qualifies
    metadata: { category: "<feature>", estimatedDurationMs: 10_000 },
  },
  {
    testId: "<feature>-edge",
    params: { /* boundary case */ },
    expectation: { validation: "type", expectedType: "string" },
    metadata: { category: "<feature>", estimatedDurationMs: 10_000 },
  },
  {
    testId: "<feature>-error",
    params: { /* invalid input */ },
    expectation: { validation: "throws-error", errorContains: "specific message" },
    metadata: { category: "<feature>", estimatedDurationMs: 2_000 },
  },
];

Register in tests/test-definitions.ts:

import { <feature>Tests } from "./<feature>-tests.js";
// ...
export const allTests: TestDefinition[] = [
  // ...
  ...<feature>Tests,
];

Executor

Extend AbstractModelExecutor (base: tests/shared/executors/abstract-model-executor.ts) or use createExecutor with TestHandler for ad-hoc cases. Bind handlers per testId, and use ResourceManager.ensureLoaded("<resource-name>") to obtain model IDs.

Register the new executor in every compatible handlers: [...] array: desktop, mobile, and Electron (which also covers Snap). Add a documented SkipExecutor for intentionally unsupported consumers.

Local validation (required before landing)

After scaffolding, provide the exact commands for every compatible local consumer. Do not mark the task complete until the user confirms those tests pass locally.

cd packages/sdk/e2e

# If packages/sdk (outside e2e) or packages/inference changed
npm run install:build:full

# Otherwise (only test code in packages/sdk/e2e changed)
npm run install:build

npx qvac-test run:local:desktop --filter <feature>-
npx qvac-test run:local:electron --filter <feature>- # verifies the handler or intentional skip

Always use install:build:full once packages/sdk itself changed, even if packages/inference has no diff lines of its own — install:build:sdk trusts the published @qvac/inference range, which can already be behind the checked-out packages/inference source from unrelated merged work.

For mobile verification of a smoke candidate (required before tagging suites: ["smoke"]):

npx qvac-test run:local:android --filter <feature>-     # or run:local:ios

Expectation reference

ValidationUse forNotes
contains-all / contains-anyKeyword or closed-set answersPreferred over type when achievable
regexStructured output (JSON keys, date, lang ID)Keep pattern anchored and stable
numeric-rangeScores, latencies, embedding magnitudePick bounds tolerant to minor model drift
type (+ minLength)Last-resort shape checkShallow; flag as weak coverage
throws-errorEvery error patherrorContains must be stable across bumps
functionComplex deterministic checks

Quality checklist

Before presenting the plan:

  • Feature surface understood from code; any genuine gaps raised as targeted clarifying questions.
  • Model-output strategy picked and justified — not defaulted to type.
  • Happy, sad, and error cases drafted.
  • Executor placement chosen; desktop, mobile, Electron/Snap registration checked; mobile memory / filesystem / platform concerns addressed.
  • Smoke candidate selected or explicitly skipped with reason.
  • Local validation command prepared with the correct --filter prefix.

Before marking scaffolding complete:

  • Test definitions aggregated in tests/test-definitions.ts.
  • Executors registered or explicitly skipped in every relevant consumer entry, including Electron when compatible (which also covers Snap).
  • User has confirmed the new tests pass through each compatible local consumer command.

References

  • Executor placement, smoke policy, rebuild flow → packages/sdk/e2e/AGENTS.md and packages/sdk/e2e/README.md.
  • Expectation schema → @qvac/test-suite dist/schemas/expectations.js.
  • Existing examples:
    • Strong output assertion: packages/sdk/e2e/tests/translation-salamandra-tests.ts (contains-any over expected Spanish tokens).
    • Error path: packages/sdk/e2e/tests/vision-tests.ts (throws-error with errorContains).
    • Shape fallback: packages/sdk/e2e/tests/completion-tests.ts (type: "string").

Signals

GitHub stars
601
Forks
111
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
qv-sdk-e2e-create
Source
github.com/tetherto/qvac