Mutation Testing (pitest)
SkillDev toolsBootstrap pitest via the info.solidsoft.pitest Gradle plugin with Kotlin-sane defaults, run mutation tests scoped to changed classes, interpret surviving mutants from mutations.xml, triage likely-equivalent mutants out of the kill queue, and drive a kill-survivor workflow. Use when the user asks to add mutation testing, check mutation score, investigate surviving mutants, or strengthen test quality beyond line coverage.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Mutation Testing (pitest) skill
What this skill tells your AI
The instructions your AI receives, as published by jvm-skills/jvm-skills in .junie/skills/mutation-testing/SKILL.md and read by ahel’s review.
Mutation testing for Kotlin/Java Gradle projects using pitest via the info.solidsoft.pitest plugin. Non-intrusive: mutates bytecode, requires no production-code changes.
When to use this skill
- User asks: "add mutation testing", "check mutation score", "why does this mutant survive", "strengthen tests", "is this class well-tested".
- A survivor appeared in a pitest report and needs killing.
- A project wants a CI gate on mutation score.
Prerequisites
- Gradle project (Kotlin DSL
build.gradle.ktsprimary, Groovybuild.gradlesupported via analogous blocks). - Tests already exist. If not, run
/ralph-coverageor/kotest-createfirst — mutation testing on an empty suite only producesNO_COVERAGEnoise. - Clean or near-clean working tree.
Preflight — run before any pitest invocation
Pitest fails cryptically (silent UNKNOWN_ERROR from the coverage minion) when the project itself doesn't compile or when the test runtime has misaligned JUnit jars. Always run these checks before ./gradlew pitest so failures surface with their real cause, not pitest's.
# 1. Test sources must compile — pitest can't help if the project itself is broken.
./gradlew compileTestKotlin compileTestJava --quiet
# 2. Find the project's actual junit-platform-launcher version.
./gradlew dependencyInsight --configuration testRuntimeClasspath \
--dependency org.junit.platform:junit-platform-launcher 2>/dev/null \
| head -1
# 3. Check that Kotlin bytecode keeps line-number debug info (pitest needs it).
grep -RnE 'Xno-source-debug-extension|-g:none' build.gradle.kts settings.gradle.kts 2>/dev/null
If step 1 fails, abort and surface the compile error — do not attempt pitest. If step 2 reports a version, use it to pick junit5PluginVersion (see next section). If step 3 finds either flag set, pitest will produce bogus line numbers or silently fail on mutators that need source info — remove the flag or exclude those modules.
Capabilities
1. Bootstrap
Detect whether pitest is already configured:
grep -n "info.solidsoft.pitest" build.gradle.kts settings.gradle.kts **/build.gradle.kts 2>/dev/null
If not configured, propose this diff to the root build.gradle.kts:
plugins {
id("info.solidsoft.pitest") version "1.19.0"
}
pitest {
pitestVersion.set("1.19.0")
targetClasses.set(setOf("<group.package>.*")) // derive from rootProject.group
targetTests.set(setOf("<group.package>.*"))
mutators.set(setOf("STRONGER")) // do NOT use ALL — NPE- and equivalent-mutation-prone
features.set(listOf(
"+FLOGCALL", // built-in: silence many logger-call mutations
// `+fkotlin` is auto-enabled by the junit5 plugin when detected — filters Kotlin bytecode-junk mutations
))
avoidCallsTo.set(setOf(
// FLOGCALL and avoidCallsTo overlap partially but neither fully covers the other on real Kotlin+SLF4J code.
// Verified empirically: removing these re-introduces 7+ logger-call survivors per scoped run.
"kotlin.jvm.internal", "kotlin.Metadata",
"org.slf4j.Logger",
"org.apache.logging.log4j.Logger",
"java.util.logging.Logger",
))
excludedMethods.set(setOf(
// Kotlin data-class synthesized members — generated, behavior guaranteed
"component*", "copy", "hashCode", "equals", "toString",
))
junit5PluginVersion.set("1.2.1") // JUnit 5 projects only — pin to project's Platform version
outputFormats.set(setOf("HTML", "XML"))
threads.set(Runtime.getRuntime().availableProcessors())
timestampedReports.set(false) // stable path for parsing
fullMutationMatrix.set(true) // names covering tests for SURVIVED mutants too — see §3
exportLineCoverage.set(true) // emits coverage.xml alongside mutations.xml
historyInputLocation.set(file("build/pitest/history.bin")) // incremental analysis — read + write same file to skip re-running killed/survived mutants on unchanged classes
historyOutputLocation.set(file("build/pitest/history.bin")) // the gradle-pitest-plugin exposes explicit paths, not the `withHistory` CLI flag
jvmArgs.set(listOf("-Xmx2g")) // avoid MEMORY_ERROR noise on forked minion; raise to 4g for large classpaths
mutationThreshold.set(60)
coverageThreshold.set(60)
testStrengthThreshold.set(60)
}
Key points when proposing:
- Derive
targetClassesfromrootProject.group(e.g.com.example.foo→"com.example.foo.*"), show the detected value in the diff and ask for confirmation. - Omit
junit5PluginVersionfor JUnit 4–only projects. Detect by searching forjunit-jupiterin dependencies. - Pin
junit5PluginVersionto the project's JUnit Platform line. Pitest-junit5-plugin bundles its ownjunit-platform-launcher. If it targets a different Platform than the one intestRuntimeClasspath, the minion dies withOutputDirectoryCreator not available(a silentUNKNOWN_ERRORat the plugin level). Rough matrix (verify against the plugin's release notes before pinning):Project junit-platform-launcherSafe junit5PluginVersion1.8.x 1.0.01.9.x 1.1.01.10.x 1.2.01.11.x 1.2.11.12.x 1.2.2Newer Platform / Jupiter (6.x / Platform 2.x) may require the pitestconfiguration to also pin a matching launcher explicitly:dependencies { pitest("org.junit.platform:junit-platform-launcher:<version>") } - Respect existing config — if a
pitest { }block already exists, only patch missing fields (e.g. addXMLtooutputFormatsif onlyHTMLis set). Never overwrite thresholds ortargetClasses. - Groovy DSL — emit the equivalent
pitest { ... }using=assignment instead of.set(). - Forward
tasks.testsettings. Pitest's minion does not inheritjvmArgs,systemProperty(...),systemProperties(...),minHeapSize, ormaxHeapSizefromtasks.test. Inspect the test task and mirror whatever it sets intopitest { jvmArgs = [...] }— system properties become-Dentries, heap becomes-Xmx/-Xms. Missing this often manifests as the silentUNKNOWN_ERROR(e.g. JUnit can't instantiate a custom@TestClassOrderbecause the test task sets it via system property and the minion doesn't see it).// If tasks.test has: // maxHeapSize = "4096m" // systemProperty("junit.jupiter.testclass.order.default", "com.example.MyOrderer") // // Then pitest needs: pitest { jvmArgs.set(listOf( "-Xmx4g", "-Dkotlin.jupiter.testclass.order.default=com.example.MyOrderer", )) } - Turn on
verbose.set(true)during bootstrap. Silent minion crashes produce actionable output only when verbose is on. Can be switched off after the first clean run.
After applying the diff, run ./gradlew pitest once to generate the baseline. First-run failure modes and what they usually mean:
| Symptom | Likely cause |
|---|---|
OutputDirectoryCreator not available ... unaligned versions of the junit-platform-engine and junit-platform-launcher | junit5PluginVersion doesn't match the project's junit-platform-launcher — see matrix above |
Silent UNKNOWN_ERROR with no stack | Enable verbose.set(true) and rerun — real cause will then surface in PIT >> SEVERE : MINION lines |
NoClassDefFoundError / custom orderer / listener fails | tasks.test system property or jvmArg not forwarded to pitest |
Unresolved reference in test sources | Pre-existing compile error — run the preflight compileTestKotlin check first |
2. Scoped run (on a change)
Derive target classes from git diff and pass them to the wired pitestScope property (see §5 for the wiring). The info.solidsoft.pitest plugin does NOT honor -Ppitest.targetClasses natively — a raw CLI flag with that name will silently be ignored and pitest will run against whatever the pitest { } block has hard-coded.
BASE=$(git merge-base HEAD origin/main 2>/dev/null || git merge-base HEAD main)
CHANGED=$(git diff --name-only "$BASE" HEAD -- '*.kt' '*.java' \
| grep '^src/main/' \
| sed -E 's#^src/main/(kotlin|java)/##; s#\.(kt|java)$##; s#/#.#g' \
| paste -sd, -)
if [ -z "$CHANGED" ]; then
echo "No production classes changed — aborting scoped run."
exit 1
fi
./gradlew pitest -PpitestScope="$CHANGED"
Aborting on empty match is deliberate — never silently fall back to whole-codebase on a scoped run.
For Groovy-DSL projects, the wiring and flag are identical (same findProperty API).
For multi-module, prefix with the module path: ./gradlew :module-name:pitest -PpitestScope=.... Wire the property in each module's pitest { } block.
3. Interpret survivors
Pitest writes build/reports/pitest/mutations.xml (stable path because timestampedReports.set(false)). Parse it:
# Quick survivor extraction
xmllint --xpath '//mutation[@status="SURVIVED"]' build/reports/pitest/mutations.xml 2>/dev/null
Each <mutation> element carries:
status—SURVIVED/KILLED/NO_COVERAGE/TIMED_OUT/MEMORY_ERROR/RUN_ERROR/NON_VIABLE<sourceFile>,<mutatedClass>,<mutatedMethod>,<lineNumber>,<mutator>,<description><killingTest>— present and populated only for KILLED mutantsnumberOfTestsRun(attribute) — how many tests reached this line
Important: with the default pitest configuration, SURVIVED mutants do not list the names of the tests that reached them — only a count via numberOfTestsRun. The killing-test field in the XML is empty for anything that survived. To identify which test to strengthen, the bootstrap sets fullMutationMatrix.set(true), which adds a <killingTests> / <succeedingTests> matrix to each <mutation> element. Without this flag, the skill can only say "2 tests ran against this line and missed" — it cannot name them.
Side-effect: fullMutationMatrix = true increases report size and slightly increases runtime. Acceptable on scoped runs (one package / one class); consider turning off for whole-codebase CI runs where survivor diagnosis isn't the goal.
exportLineCoverage.set(true) emits a separate coverage.xml listing which tests reach each line. Combined with mutations.xml, you can reconstruct the covering-test set even without fullMutationMatrix, but it requires joining two files. Enabled by default in the bootstrap because it is cheap and useful for the aggregator.
Print a compact table:
STATUS FILE:LINE MUTATOR DESCRIPTION
SURVIVED Calculator.kt:42 MATH replaced + with -
covered by: CalculatorTest.`positive number is positive`
SURVIVED Calculator.kt:43 CONDITIONALS_BOUNDARY changed > to >=
covered by: CalculatorTest.`zero is not positive`
NO_COVERAGE PriceCalc.kt:88 VOID_METHOD_CALLS removed call to log.debug
Distinguish SURVIVED (test exists but missed the behavior) from NO_COVERAGE (no test reaches the line) and TIMED_OUT (loop mutant caught by pitest's timeout heuristic — not a real survivor).
4. Triage — classify survivors before trying to kill them
This step runs between interpretation and kill-a-survivor. Its job: separate real test-quality gaps from equivalent mutants — mutations that change the bytecode but not the observable behavior in any test a human would reasonably write. Without triage, a kill-a-survivor loop wastes iterations (and, under unattended automation, produces tautological tests) trying to kill things that shouldn't be killed.
Classify each SURVIVED mutant into one of three buckets:
| Bucket | Meaning | Fed to kill-survivor / autoresearch? |
|---|---|---|
LIKELY_KILLABLE | Behavioral change a sensible test could catch | Yes |
LIKELY_EQUIVALENT | No realistic test can distinguish mutant from original | No. Written to suspected-equivalent.md for the record |
AMBIGUOUS | Can't tell without reading more context | Surface for human decision; do not auto-feed to autoresearch |
Read the source line at file:line for each survivor and match against these archetypes (expand via the project overlay):
Archetype A — logger-gate conditionals. The line is an if (<numeric> <comparator> <constant>) { ... } whose body contains only calls to loggers (logger.info, .warn, .error, .debug, .trace) or trace/metric sinks. Any ConditionalsBoundary / RemoveConditional_ORDER_* / RemoveConditional_EQUAL_* survivor on such a line is LIKELY_EQUIVALENT.
if (deleted > 0) { // <— mutations on `> 0` all survive
logger.info("...") // because tests don't capture log output
}
Archetype B — unreachable loop bounds. The line is inside a while / do..while / for condition whose opposing operand is a constant larger than any realistic test input (threshold: default 1000; configurable via overlay). ConditionalsBoundary / RemoveConditional_ORDER_* survivors on such lines are LIKELY_EQUIVALENT — the mutation is theoretically killable but only with infeasible data volumes.
} while (batch.isNotEmpty() && batchCount < maxBatches) // maxBatches = 10, batch = 10000 rows
Archetype C — null-elvis on non-nullable-in-practice values. The line is a <value> ?: throw <X> pattern where <value> is the non-nullable-by-contract result of a DB primary-key fetch, an Optional.get(), or similar. RemoveConditional_EQUAL_IF / RemoveConditional_EQUAL_ELSE survivors are LIKELY_EQUIVALENT — null is unreachable in any test a DB schema permits.
Kotlin-specific note: pitest's NULL_RETURNS mutator already skips @NotNull-annotated methods, and Kotlin compiles non-nullable return types to @NotNull in bytecode. So null-return mutants only surface on nullable Kotlin returns (Foo?) — archetype C is tuned to that reality and does not over-trigger on plain non-nullable returns.
eventCleanup.hardDeleteEvent(eventId ?: throw IllegalStateException("eventId is null"))
// ^-- eventId is jOOQ-fetched primary key, never null in practice
Archetype D — branch-identity (partial-kill residual). An if (cond) { X } else { Y } where X and Y produce the same observable output when cond is false. Classic example: if (list.isNotEmpty()) { repo.fetch(list).associateBy { it.id } } else { emptyMap() } — when list is empty, fetch([]) returns empty and associateBy on empty yields emptyMap, so the IF-branch is indistinguishable from the ELSE-branch for the "empty" case. RemoveConditional_EQUAL_IF mutations on this line survive any test because flipping the branch doesn't change the observable output.
Detecting this archetype statically is hard — it requires knowing that the IF-branch operation is idempotent over the identity input of the ELSE-branch. The script can't match it reliably from source alone. This archetype is surfaced dynamically: when a single strengthened test kills the EQUAL_ELSE mutation on a line but leaves the EQUAL_IF mutation (or vice versa) surviving, the surviving mutation is a strong candidate for LIKELY_EQUIVALENT. Promote it after 2 failed targeted kill attempts (not 3, since the signal is stronger than blind failure) and skip it in future iterations.
Anything that matches no archetype is AMBIGUOUS by default — bias toward ambiguity, not toward equivalence, so real gaps don't get silently excluded.
Implementation. A ready-to-use triage script lives at scripts/triage.py in this skill directory. Run:
python3 .claude/skills/mutation-testing/scripts/triage.py \
build/reports/pitest/mutations.xml \
.
It reads the mutations XML, opens each survivor's source line, applies archetypes A/B/C, and writes triage.md + triage.json next to the mutations file.
Output: triage.md (human) and triage.json (machine) next to mutations.xml. Each survivor carries the source pointer, archetype reason, and the names of the tests that covered but didn't kill it (extracted from <succeedingTests> / <coveringTests>). Covering-test names require fullMutationMatrix=true in the bootstrap block; otherwise the script prints a hint to enable it.
Shape:
# Triage of <N> survivors
## LIKELY_KILLABLE (<count>)
- `UserCleanupService.kt:58` **VoidMethodCall** — removed call to deleteS3Files
- in `UserCleanupService.hardDeleteUser()`
- reason: no archetype matched
- covering tests:
- `UserCleanupServiceTest.hardDeleteUser removes user and all owned data`
## LIKELY_EQUIVALENT (<count>)
- `AnonymousUserCleanupJob.kt:70` **ConditionalsBoundary** — changed > to >=
- reason: archetype A — logger-gate conditional
- `AnonymousUserCleanupJob.kt:68` **ConditionalsBoundary** — changed < to <=
- reason: archetype B — loop-bound comparison
The suspected-equivalent.md set is durable — it accumulates across sessions, not per-run. A mutant that's triaged as equivalent once stays in that list unless the source line changes (detected by re-running triage when the line's content differs).
Project overlay extensions. references/project.md can add project-specific archetypes:
- Custom logger types beyond the stdlib set
- Known-non-null value types beyond DB primary keys (e.g. "values returned from
MyContext.requireUser()") - Loop-bound constants specific to the project
- Explicit allowlist of mutants always classified as equivalent (by file:line:mutator key)
Do not skip triage before autoresearch. The mutation-autoresearch skill requires a current triage output as its input — it only iterates LIKELY_KILLABLE. Running autoresearch on the raw survivor list would burn iterations on equivalents and produce tautological tests under pressure.
5. Kill a survivor
Inverted TDD cycle. Unlike a normal /tdd-task RED → GREEN cycle, a mutation-kill test is written against production code that is already correct. The test should pass on the first run. The RED → GREEN proof comes from the scoped pitest rerun (see verify step below): the mutant flips from SURVIVED to KILLED. A /tdd-task agent that tries to force a red-first state here will waste iterations — pass this expectation in the prompt.
For each survivor, delegate to /tdd-task with a prompt like:
Kill mutation:
<mutator>at<file>:<line>in<mutatedClass>.<mutatedMethod>. Description:<description>. Currently covered by:<coveringTests>(from triage.md). Spell out both behaviors: what the original code returns and what the mutated code would return for the chosen input. The assertion must distinguish the two. Do not assert the raw mutated value or operator — assert the observable behavior on a specific input. Prefer strengthening assertions in the covering test; if that isn't natural, add a new sibling test case. The test is expected to pass against current production code — don't force a red-first state. RED proof comes from the pitest rerun below.
After /tdd-task returns, verify with a scoped pitest rerun. Override scope via Gradle -P flags — do not edit build.gradle.kts for per-iteration scope changes. The info.solidsoft.pitest plugin does NOT read -Ppitest.targetClasses natively; you must wire a custom property into the pitest { } block. Pick a name like pitestScope:
pitest {
pitestVersion.set("1.19.0")
val scopeProp: String? = project.findProperty("pitestScope") as String?
val testsProp: String? = project.findProperty("pitestTests") as String?
val defaultScope = setOf("com.example.*")
val scopeFqns = scopeProp?.split(",")?.map { it.trim() }?.filter { it.isNotEmpty() }
targetClasses.set(
// Auto-expand each FQN to `{FQN, FQN$*}` so Kotlin-emitted nested
// and synthetic classes (`Foo$Page`, `Foo$methodName$1`) are included.
scopeFqns?.flatMap { listOf(it, "$it\$*") }?.toSet() ?: defaultScope
)
targetTests.set(
testsProp?.split(",")?.map { it.trim() }?.filter { it.isNotEmpty() }?.toSet()
?: scopeFqns?.toSet()
?: defaultScope
)
// ... rest of the block stays stable
}
Then the per-iteration rerun is:
./gradlew pitest -PpitestScope='com.example.pkg.MediaSubmission'
Scoping rules:
- Include nested/synthetic classes. Kotlin emits synthetic bytecode for nested data classes (
Foo$Page) and lambdas (Foo$methodName$1,.let { },.map { }). The bare FQN alone doesn't match them — without$*you will miss mutations inside inner classes and lambda bodies. - Don't use the prefix glob
FQN*(without$). It also matches sibling top-level classes (MediaSubmissionService) and inflates scope. - Keep the
pitest { }block pinned to a stable, broad default (the module-level package you CI against). Treat per-iteration scopes as transient CLI overrides that never land in git.
./gradlew pitest
Parse build/reports/pitest/mutations.xml directly — do not rely on the Gradle exit code. Pitest exits non-zero when line coverage falls below coverageThreshold (typical on narrow single-class scopes), but the XML is still valid. The skill only cares about three facts:
- The target mutation's
statusis nowKILLED. - No previously-
KILLEDmutation in the same class flipped toSURVIVED. - The class's
killed_countstrictly increased.
If all three hold, delegate to /commit. If the new test asserts a tautology (literally references the mutated operator/constant, or only checks a value trivially equal to the mutation's replacement), reject and retry with a stronger prompt. If the rerun shows the mutation still alive, read the committed assertions and ask: did the test actually exercise an input where mutated vs original diverge? Tightening the input data usually kills it.
A quick awk / xmllint / Python one-liner on mutations.xml is enough — pitest 1.19 writes attributes with single quotes (status='SURVIVED'), so grep patterns using double quotes will silently match nothing.
Fallbacks when dependencies aren't installed
- No
/tdd-task: follow the inline inverted-TDD checklist — (1) identify the behavioral input where original and mutated code diverge, (2) add a test asserting the correct behavior on that input (it passes immediately against current code), (3) run./gradlew test --tests <pattern>and confirm green, (4) run scoped pitest and confirm the mutant moved SURVIVED → KILLED. The initial RED step from classic TDD does not apply — production code is correct; RED proof comes from the pitest rerun. - No
/commit:git add src/test && git commit -m "test(mutation): kill <mutator> at <file>:<line>". - No
/test-gradle:./gradlew test --tests <fully-qualified-test-pattern>.
Kotlin-specific gotchas
Shortened here. Read the whole file on GitHub.
Signals
- GitHub stars
- 138
- Forks
- 26
- Last commit
- Aug 2026
Advanced
- Catalog kind
- skill
- Gateway key
mutation-testing-jvm-skills- Source
- github.com/jvm-skills/jvm-skills