qikly

MCP serverAI & models

Lets your agent generate tests from acceptance criteria and refine code against them without seeing the tests.

Unavailable. This server has no hosted endpoint yet, so ahel can't serve it.

Add to setup to save this item as a reference. ahel cannot run it, and signing in will not install it.

About this server

Generates tests from acceptance criteria, then converges code with an agent that never sees them.

Getting started

  1. Save this item in Your setup as a reference.
  2. Read the source or reference documentation for its setup requirements. Saving it here does not connect it to your AI.
  3. Check this page for availability before trying to install it through ahel.

From the project's README

As published by gal-a/qikly in README.md.

The problem: Your AI writes both the code and its tests. How do you know the tests are really valid?

The solution: two agents. One turns the acceptance criteria into tests. The other writes the code and never sees the acceptance criteria.

Who it is for: a developer or team pointing an AI coding agent at a self-contained Python module that transforms data, for example an ETL step, a merge, a calculation or a validation routine, who does not want to trust a green suite when the same agent wrote both the code and the tests. It suits one module at a time in small to mid-sized repositories: when a test fails, only the files that failure names are loaded, so runs stay small and quick. It fits most naturally where verification already has to be independent, such as automotive, medical devices, fintech and defence: ADAS_HEADWAY, a bundled example, checks following distance from forward-radar samples. See What it is for.

Just want to try it on your own data? Quick start on your own data, five steps from your module to a first run.

Just want to see how it works, free and with no API key? Run pip install qikly, then qikly --explain CALC_TAX: it prints what each agent is shown, and the difference. More on the free commands.

The idea

Imagine a student who writes the exam paper, writes the answer key, and then sits the exam. They pass, and nobody would accept that as evidence they know the material. That is what happens when one model gets a specification containing the acceptance criteria and writes both the code and the suite that checks it: everything goes green, and the green means nothing.

qikly takes the answer key away from the student. It generates a test suite from the acceptance criteria, then writes an implementation and repairs it against that suite until every test passes or a retry budget runs out, recording every failure, every piece of reasoning and every diff.

The part that makes the result mean something: the coding agent never sees acceptance_criteria. It gets the specification with that section stripped out, the same vague brief a developer works from, while test generation gets it in full. When a test fails, the agent sees the failure message and never the rule it broke. Without that asymmetry both sides read the same spec identically and every test passes first try, which proves nothing.

Purple is what the coding agent can see. Teal is what the standard is written from. They never touch. A run that never converges is still worth having: it exits non-zero, names the blocking tests, and keeps the same complete record. The purple arrows are the repair loop, and that is where almost all of a run happens: a failing suite sends the agent the failure text and nothing else, it produces a FIX and a PATCH, and the suite runs again, until the stage passes or the retry budget runs out. It never sees the rule it broke, so it cannot write code shaped to a criterion it was shown. tests/test_withholding.py fails the build if any call site lets one through.

A run works through three stages, integration then system then unit:

  1. Integration and system tests are generated first, from the spec alone, before any code exists. They cannot see an implementation because there is not one yet.
  2. The coding agent writes the implementation, from the spec minus the criteria.
  3. pytest runs. On failure the model produces a FIX (failure summary, root cause, plan, and the files it intends to touch) and then a PATCH (a unified diff of only those files), applied all or nothing. Repeat until the stage passes or the budget is spent.
  4. Unit tests are generated last, once real code exists for them to name. This is the only stage allowed to see the implementation.
  5. Clearing a stage re-runs the earlier ones, so a later fix cannot silently break something that already passed.

Every arrow back into FIX carries the pytest error text and nothing else. Unit tests come last because they are the only ones that need to name real functions, which makes them the only stage allowed to read the implementation. Re-running the earlier stages after each success is what stops a later repair quietly breaking something that already passed.

Test generation sees the requirements, the input and output contract, and every acceptance criterion in full. It writes integration, system and unit tests against the standard.

The coding agent sees the same specification with the criteria section removed, plus the text of whatever test just failed. The same vague brief a developer usually works from.

Two ways to get a test suite, and what each can prove

Code-derived suitemost commercial test generators, and qikly's own unit testsSpec-derived suiteqikly's integration and system tests
Written from the code as it is todayWritten from the acceptance criteria you wrote
You supply nothing but the repositoryYou supply a written statement of what correct means
Catches behaviour changing tomorrowCatches behaviour being wrong today
Cannot catch the code being wrong now: today's bug becomes tomorrow's assertionCannot catch anything nobody wrote down
Right choice when nobody wrote the intent down and you need a safety netRight choice when the intent exists in a ticket, a spec page or a Gherkin file

Both are useful and they answer different questions. qikly is not purely one or the other: integration and system tests are written from the criteria before any code exists, the unit stage is written last from the code that just passed them, and --refine-criteria reads a converged implementation to propose criteria the first draft missed. Each of those reads the code on purpose, and none of them can question it.

What makes this different

The tests come from the standard, not from the code. This is the one that matters most. Every other AI test generator in this space writes its tests from an implementation that already exists, so it can only describe what the code already does. That makes an excellent regression harness, and it cannot tell you the code is wrong. qikly writes the integration and system suites from the acceptance criteria before any implementation exists, so the standard cannot have been shaped by the thing it judges.

The withholding is a mechanism you can watch. Not a prompt asking a model to ignore a section, and not a convention someone has to remember. One command prints what each side is given and the difference between them, offline and free:

qikly --explain <MY_TASK>     # e.g. qikly --explain MERGE_SALES

Eleven criteria go to test generation. Twelve lines are removed before the coding agent sees the same file. tests/test_withholding.py fails the build if any call site ever lets one through, including one added next year by someone who has never read this. It is a property of the code, and it takes thirty seconds to check.

What you get is an executable suite you keep. The output is pytest files and JUnit XML. Read them, run them, put them in CI, and when one fails in six months it fails for a reason you can inspect and argue with. A suite is a durable asset in a way a model's verdict is not: a verdict cannot be re-run against tomorrow's commit.

It helps you write the standard, not just check against it. --init and --scaffold turn existing code into a task, --criteria-from lifts criteria out of a ticket you already wrote, --generate-criteria drafts a first bar from requirements alone, and --check-criteria looks for two statements anywhere in the specification that no implementation could satisfy at once, including two acceptance criteria that disagree with each other.

Every run is reproducible, and the whole trail is kept. A run records the provider, the model, the settings and the version that produced it, next to every failing test, every FIX with its stated root cause, and every PATCH as a diff. You can read back exactly why a line of code exists: which assertion forced it, what the model concluded, and what it changed. The record survives a run that never converges, which is when you most want it.

"Why not just use two different models?"

It is the first thing most people ask, and it does help a little. It does not reach the underlying issue, because both models still read the same criteria and so both still write to them: the code is still built to satisfy the standard it is about to be judged by, and changing who types it does not change what they were shown. It is also a habit rather than a mechanism, and nothing checks the two stayed different.

The two compose nicely, incidentally, since qikly picks a provider and model per agent role. You can withhold and use two models.

"Why not just add a reviewer agent?"

The newer version of the same question, and the one worth answering carefully, because independent verification steps are now shipping in mainstream coding agents: a second agent, often from a different model family, reviews what the first one produced.

It helps, and it does not reach this. A reviewer given the same specification has read the same acceptance criteria, and resolves the same ambiguity the same way. It will catch a mistake that is visible from that context: an inconsistency, a requirement plainly skipped, an obvious bug. It cannot catch the case this tool is built for: a line that could be read two ways, read once, with both the code and the standard written from that single reading. Nobody is wrong, so nothing looks wrong.

The problem was never that nothing was checking. It is that everything checking had already seen the answer key. Withholding is what makes the check structural rather than one more opinion drawn from the same context, and it is enforced by a test rather than by an arrangement someone has to remember to keep.

There is a second difference, and it outlasts the run: a reviewer emits a verdict, and this emits a pytest suite you still have in six months.

"Doesn't a failing test give the criteria away?"

It gives away one case, and that is by design. When a test fails, the coding agent sees the test name and the assertion error: in the case study, that a tax rate of 150 was accepted when it should not have been. It never sees the criterion behind it, and never sees the tests it has not failed yet.

That does not undo the separation, because independence is a property of how the suite was written, not of how much feedback the code's author receives afterwards. The integration and system suites are generated from the criteria before any implementation exists, and nothing the coding agent learns later can reshape a test that is already written. A repair that games the one visible failure still has the rest of the suite in its way, and earlier stages run again each time a later one clears.

It is the position a developer is in when CI goes red: they see what broke, not the test plan. The agent in the case study wrote tax_rate > 100 only because a test told it 150 was wrong. Had it been handed the criteria, it would have written the bound first time, and the green would have proved nothing.

Design rationale, and the harder problem of where acceptance_criteria comes from in the first place: docs/design_1_case_study.md, the first of three parts.

What it is for

Built for self-contained Python modules that transform data, not for a large existing repository: ETL, merges, calculations, validation. That is the layer where a wrong answer looks like a right answer, and where a test written from the rule is the only thing that catches it.

Where it does not fit today: an existing large repository. PATCH prompts load only the files a FIX names, and while a large file is now excerpted rather than loaded whole, there is no cross-file index. See Where it fits today.

How to use the tools in this project

Four ways in, and the table under the quick start says which command each one needs:

  1. Verify code you did not write. Supply an implementation through seed: and the suite is written from your acceptance criteria by an agent that never reads that code. A suite generated from the same context as the code is a model agreeing with itself.

    The limit is worth saying plainly: qikly cannot know what the author of supplied code saw. Withholding is a property of a run qikly performed, not of a file you hand it. If the same person or model wrote that code with the criteria open, this gives you an independent suite, not an independent author. What it does give, always, is a suite that was not derived from the code, which is the half a code-derived generator cannot give you at all.

  2. Start from a spec. No code yet: get a first implementation and the suite that justifies it, in outputs/, never in your source tree.

  3. Bring your own tests. Seed any stage and the loop becomes a repair procedure rather than a generator.

  4. Run a catalog unattended. Non-zero exit on any non-convergence, so a scheduler or CI job can run many specs and keep the reports.

Try it without spending anything

Three commands that make no model call, need no API key, and cost nothing.

qikly --explain MERGE_SALES   # what each side is shown, and the difference
qikly --validate              # check your task files: YAML, criteria, fixtures
qikly --explain MERGE_SALES --json
qikly --explain MERGE_SALES --html   # the same, as a page to share

--explain is the one worth running first. It prints the acceptance criteria that test generation receives, then the same task file as the coding agent receives it, then the diff: on MERGE_SALES, all eleven acceptance criteria are removed, along with the acceptance_criteria: key they hang off. It builds those strings through the same function a real run uses, so it shows the mechanism rather than a description of it.

Add --html and it also writes qikly_explain_MERGE_SALES.html: both views side by side with every withheld line highlighted, in one file that loads nothing from anywhere. Attach it to a pull request, put it in a slide, or screenshot it for a post.

--validate reads your task files and nothing else: that they parse, that acceptance_criteria is a list rather than one long string, that fixture paths resolve, and that criteria name values instead of adjectives. It is also available as a pre-commit hook, qikly-validate, deliberately the free check rather than the paid one.

Quick start

pip install qikly
export GEMINI_API_KEY=...     # PowerShell: $env:GEMINI_API_KEY = "..."
qikly --demo

Needs Python 3.10+ and GNU patch; on macOS run brew install gpatch first. The demo runs a bundled task end to end in a throwaway folder, in about thirty seconds, and writes nothing outside it.

That is a real run on gemini-3.5-flash-lite, 38 seconds, not sped up.

To try it on your own code and data, start at docs/QUICK_START_ON_YOUR_OWN_DATA.md: five steps from qikly --scaffold your_module.py to a first run, and the table of which command fits what you already have.

Tried it? Tell us what happened, whether it worked, stalled or never got past install.

Running

python run.py                            # every task found
python run.py --tasks ETL_ADDRESS        # one
python run.py --tasks ETL_ADDRESS,ETL_EMAIL
python run.py --demo --tasks MERGE_STOCK # isolated, any task

Each task runs in its own process, concurrently, with console output prefixed [task_id]. Exit code is non-zero if any task did not fully converge.

Expect some runs to stall, by design. Roughly 8 runs in 10 finish with the code passing every integration and system test. Roughly 6 in 10 pass everything including unit tests. Nearly the whole gap between those two figures is the unit stage.

Those are round numbers because they were measured three times: a 427-run sweep, a 140-run sweep sixteen days later on the same tasks and settings, and a 400-run sweep after correcting the benchmark itself, when eight of the ten tasks turned out to be carrying acceptance criteria that no input row could trigger. All three landed inside each other's intervals. All three used gemini-3.5-flash-lite, a small cheap model chosen to make repeated sweeps affordable, so treat them as a floor. Three sweeps agreeing is worth more than any one of them's decimal places, so the decimal places are not quoted.

A run that exhausts its budget exits non-zero, names the tests that blocked it, and keeps the full record. It never reports success on code its own tests reject.

Thirteen example tasks ship with the tool across five domains, listed in docs/design_3_mechanism.md.

Measuring rather than producing. One run is an artifact, not a rate: the same task with the same seed converges on some runs and not others. To claim how often anything converges, repeat the sweep and read the interval:

python -m qikly.orchestrator.run_all --repeat 10

That writes an aggregate report with confidence intervals and groups the non-converging runs by what they got stuck on. Details in docs/design_3_mechanism.md.

When a run does not converge

A stall is a normal outcome, not a broken tool: the run exits non-zero, names the blocking tests, keeps the whole record, and ships nothing.

The first thing to try is a stronger model, which moves convergence more than any setting in this file and costs one environment variable. After that, in triage order: read the timeline report, check for a collection error (nothing ran at all), rule out a forgotten setting with --trends, look for the same patch repeating (a criterion fighting the model's priors), and run --check-criteria, --validate and propose_fixtures.

Each of those, with the signature to look for and the fix: docs/TROUBLESHOOTING.md.

Questions people ask before they start, including whether it can use the classes you already have and whether anything leaves your machine: docs/FAQ.md.

Output

Everything is namespaced by task_id so concurrent runs never collide:

Shortened here. Read the whole README on GitHub.

Signals

GitHub stars
15
Last commit
Sep 2026
Advanced
Delivery
qikly MCP server → your ahel connector (mcp.ahel.ai) → your AI.
Item type
mcp-server
Key
io-github-gal-a-qikly
Source
github.com/gal-a/qikly