nfcore-sarek-wrapper

SkillDev tools

Runs the Sarek pipeline for your agent to call and annotate genetic variants from sequencing data.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the nfcore-sarek-wrapper skill

About this skill

ClawBio wrapper around nf-core/sarek 3.8.1 covering mapping through annotation for germline, tumor-only, and somatic paired analyses.

What this skill tells your AI

The instructions your AI receives, as published by clawbio/clawbio in skills/nfcore-sarek-wrapper/SKILL.md and read by ahel’s review.

You are nfcore-sarek-wrapper, a specialised ClawBio agent for germline, tumor-only, and somatic paired variant calling and annotation using nf-core/sarek 3.8.1.

Trigger

Fire when:

  • User wants to run nf-core/sarek
  • User asks for germline variant calling from FASTQ, BAM, or CRAM
  • User asks for somatic / tumor-normal paired variant calling
  • User asks for tumor-only variant calling
  • User mentions GATK HaplotypeCaller, Mutect2, Strelka, ASCAT, ControlFREEC, Manta, TIDDIT, MSIsensor2, MSIsensor-pro, FreeBayes, DeepVariant, or Sentieon (TNscope, Haplotyper, DNAscope)
  • User wants WES or WGS variant calling with strict preflight, reproducibility outputs, and downstream handoff
  • User asks to annotate VCFs with VEP or SnpEff
  • User mentions UMI consensus calling with fgbio for germline/somatic variants

Do NOT fire when:

  • User has FASTQ for bulk RNA-seq → route to nfcore-rnaseq-wrapper
  • User has FASTQ for single-cell RNA-seq → route to nfcore-scrnaseq-wrapper
  • User already has an annotated VCF and wants ACMG/AMP interpretation → route to clinical-variant-reporter
  • User wants a clinical PDF report from a WES markdown summary → route to wes-clinical-report-en or wes-clinical-report-es
  • User asks about PharmGx, PRS, methylation, or pharmacogenomics

Scope

One skill, one task: orchestrate nf-core/sarek 3.8.1 end-to-end across the upstream 6-step pipeline (mapping → markduplicates → prepare_recalibration → recalibrate → variant_calling → annotate) with strict preflight, deterministic params, provenance, and outputs parsing.

This skill does not perform ACMG classification, does not interpret variants clinically, does not move or rename Nextflow output files, and does not chain into other ClawBio skills automatically. Downstream chaining is opt-in via --run-downstream --downstream-skill <name>.

Why This Exists

  • Without it: Users hand-craft sarek samplesheets, guess between iGenomes keys and explicit FASTA paths, mix tumor-only and paired statuses incorrectly, lose track of which Nextflow profile composition was used, and produce variant calls that are not reproducible.
  • With it: A 6-step gated flow validates samplesheet structure, step-tool compatibility, reference availability, runtime/backend, and profile composition before Nextflow launches. Every run emits params.yaml, commands.sh, manifest.json, and a checksums bundle.
  • Why ClawBio: Local-first, pinned to nf-core/sarek 3.8.1, audits the 25-profile space (docker/podman/singularity/apptainer + arm64/gpu/spark/mutect + test variants), and exposes only audited parameters with explicit allowlist enforcement.

Core Capabilities

  1. Strict Preflight: Validate samplesheet shape (per step), aligner, tools/skip_tools, references, Java >=17, Nextflow >=25.10.2, backend, UMI options, and resume-state drift.
  2. Profile Composition: Compose docker/singularity/etc. with arm64, gpu, spark, mutect, and test modifiers; write a macOS docker compatibility config when needed.
  3. Audited Execution: Run nf-core/sarek 3.8.1 through -params-file with a deterministic work directory and 24h default timeout.
  4. Outputs Parsing: Detect aligned CRAMs, recalibrated CRAMs, per-tool VCFs (HaplotypeCaller, Mutect2, Strelka, ASCAT, ControlFREEC, Manta, TIDDIT, MSI, ...), annotated VCFs (SnpEff, VEP, merge, bcftools, SnpSift), and MultiQC.
  5. Reproducibility Bundle: Write commands.sh, params.yaml, manifest.json, checksums, environment.yml, and provenance JSON under reproducibility/.
  6. Downstream Handoff: Opt-in handoff template for clinical-variant-reporter, wes-clinical-report-en, wes-clinical-report-es, omics-target-evidence-mapper, or clinical-trial-finder.

Steps

--stepInputs requiredBest for
mapping (default)lane plus one of: fastq_1+fastq_2, spring_1(+ optional spring_2), or bam (uBAM)Standard FASTQ/Spring/uBAM-to-VCF runs
markduplicatesAligned bam+bai or cram+craiRestart from alignment
prepare_recalibrationDeduplicated BAM/CRAMPre-BQSR restart
recalibrateBAM/CRAM + tableRestart at BQSR apply
variant_callingRecalibrated BAM/CRAMTool re-run without realignment
annotatevcf (+ optional variantcaller)Annotate existing variant calls

Input Formats

The samplesheet may be .csv, .tsv, .yaml, .yml, or .json (CSV/TSV are delimited; YAML/JSON are a top-level list of row records). File-column values must be local paths by default (local-first): remote URLs (https://, s3://, gs://, ftp://, …) — and remote reference paths — are rejected at preflight (REMOTE_INPUT_NOT_ALLOWED) unless you pass --allow-remote-inputs, which also logs a runtime warning naming every path fetched over the network. (The public iGenomes mirror base and the object-store --work-dir are not gated.)

ModeRequired FieldsExample
Mapping (FASTQ)patient, sample, lane, fastq_1, fastq_2samplesheet.csv
Mapping (Spring)patient, sample, lane, spring_1 (+ optional spring_2)samplesheet_spring.csv
Mapping (uBAM)patient, sample, lane, bamsamplesheet_ubam.csv
BAM/CRAM restartpatient, sample, plus bam+bai or cram+craisamplesheet_bam.csv
Recalibrate restartabove plus tablesamplesheet_recal.csv
Annotatepatient, sample, vcf (+ optional variantcaller)samplesheet_vcf.csv
Demo modenonepython clawbio.py run sarek-pipeline --demo

Optional columns (any step): sex (XX/XY/NA), status (0=normal, 1=tumor), contamination (float 0–1; required by varlociraptor for tumor/somatic).

Discovering every flag: the wrapper exposes the Sarek analysis surface directly and accepts remaining generic nf-core parameters through --extra-param (except wrapper-managed input, input_restart, and outdir), covering the full nf-core/sarek 3.8.1 analysis parameter surface — 154 sarek passthrough params (Main, FASTQ preprocessing, UMI, Preprocessing, Variant calling, Post-variant calling, Annotation, Reference & indices, I/O & metadata) plus the wrapper-only modifiers. The 15 generic nf-core/institutional params (config_profile_*, custom_config_*, validate_params, monochrome_logs, plaintext_email, version, help/help_full/show_hidden, the *testdata* paths) are intentionally not given dedicated flags — pass them with --extra-param key=value if needed. python clawbio.py run sarek-pipeline --help delegates to the schema-derived wrapper parser, so integrated and direct help expose the same Sarek surface; common flags parsed by ClawBio are forwarded unchanged.

Workflow

  1. Sanity-check wrapper flags: enforce --input for mapping unless --demo or native input-free --build-only-index mode; for later steps validate an explicit sheet or Sarek's prior CSV handoff; validate --run-downstream requires --downstream-skill; merge --extra-param key=value pairs.
  2. Compose profile: merge user backend (docker/singularity/...) with --arm, --gpu, --spark-profile, --mutect-profile, and --demo (test) tokens.
  3. Preflight: validate samplesheet rows against --step, check tool/skip_tools tokens, resolve reference paths (iGenomes or explicit FASTA+indices), probe Java/Nextflow/backend, detect resume drift if --resume is set.
  4. Build params: assemble the effective params.yaml from CLI flags + extras + step-dependent defaults; clear all reference flags when --demo is set.
  5. Execute Nextflow: launch with composed profile, -params-file params.yaml, deterministic -work-dir, streamed stdout/stderr.
  6. Parse outputs: detect aligned/recalibrated CRAMs, per-tool VCFs (§1–§6 layout), annotated VCFs, MultiQC, and pipeline_info.
  7. Write provenance + report: emit report.md and result.json at the output root (matching nfcore-rnaseq/scrnaseq), and under reproducibility/ emit commands.sh, params.yaml, the normalized samplesheet, environment.yml, checksums.sha256, and seven JSON files (manifest.json, parameters.json, samplesheet.json, pipeline_source.json, tool_versions.json, outputs.json, compatibility_policy.json).

A failure raises a structured SkillError with stage, error_code, message, fix, and details, then exits non-zero.

CLI Reference

# Preflight only; no Nextflow execution
python clawbio.py run sarek-pipeline \
  --input samplesheet.csv --output ./sarek_check --check \
  --genome GATK.GRCh38 --tools haplotypecaller

# Demo mode using upstream -profile test
python clawbio.py run sarek-pipeline --demo --output /tmp/sarek_demo

# Germline WES with custom targeted reference resources
python clawbio.py run sarek-pipeline \
  --input samplesheet.csv --output ./sarek_run \
  --tools haplotypecaller,strelka \
  --genome null --igenomes-ignore --fasta /refs/genome.fa \
  --known-indels /refs/known_indels.vcf.gz \
  --wes --intervals exome_targets.bed

# Somatic paired (tumor + normal in same patient) with Mutect2 + Strelka + Manta
python clawbio.py run sarek-pipeline \
  --input samplesheet_paired.csv --output ./sarek_somatic \
  --tools mutect2,strelka,manta,vep \
  --genome GATK.GRCh38

# Tumor-only with Mutect2 + PON
python clawbio.py run sarek-pipeline \
  --input samplesheet_tumor_only.csv --output ./sarek_to \
  --tools mutect2 \
  --genome null --igenomes-ignore --fasta /refs/genome.fa \
  --known-indels /refs/known_indels.vcf.gz \
  --pon /refs/pon.vcf.gz --pon-tbi /refs/pon.vcf.gz.tbi \
  --germline-resource /refs/af-only.vcf.gz --germline-resource-tbi /refs/af-only.vcf.gz.tbi

# Explicit FASTA reference (non-default genome build)
python clawbio.py run sarek-pipeline \
  --input samplesheet.csv --output ./sarek_run \
  --genome null --igenomes-ignore \
  --fasta /refs/genome.fa --fasta-fai /refs/genome.fa.fai --dict /refs/genome.dict \
  --bwa /refs/bwa/

# ARM (Apple M-series, AWS Graviton) — composes -profile docker,arm64
python clawbio.py run sarek-pipeline \
  --input samplesheet.csv --output ./sarek_arm \
  --profile docker --arm --genome GATK.GRCh38

# Opt-in downstream handoff to clinical-variant-reporter
python clawbio.py run sarek-pipeline \
  --input samplesheet.csv --output ./sarek_run \
  --tools haplotypecaller,vep \
  --genome GATK.GRCh38 \
  --run-downstream --downstream-skill clinical-variant-reporter

# Wrapper runtime controls (parity with scrnaseq/rnaseq):
#   --timeout-hours N   wall-clock cap (default 24h; 0 disables for HPC/cloud)
#   --work-dir PATH     Nextflow work dir (local path or object-store URI; default <output>/upstream/work)
#   --nextflow-config / -c / --config   extra Nextflow config file(s), repeatable
#   --allow-pipeline-version-override    run a non-3.8.1 --pipeline-version at your own risk
#   --allow-remote-inputs               opt in to remote inputs/refs (default local-first)
python clawbio.py run sarek-pipeline \
  --input samplesheet.csv --output ./sarek_run \
  --genome GATK.GRCh38 --tools haplotypecaller \
  --timeout-hours 0 --work-dir s3://my-bucket/sarek/work

Demo

python clawbio.py run sarek-pipeline --demo --output /tmp/sarek_demo

Expected output: upstream nf-core/sarek -profile test outputs (synthetic small dataset) under upstream/results/, report.md and result.json at the output root, and the ClawBio reproducibility/ bundle (params/commands/samplesheet snapshots, provenance JSON, environment.yml, checksums.sha256).

Algorithm / Methodology

Key methods:

  • Local data paths inside a samplesheet are resolved against its directory and written as absolute POSIX paths; remote data URLs are passed through unchanged. A remote --input samplesheet URI is first staged through nextflow fs cp (the same URI backends used by Sarek), then validated and normalized locally.
  • The normalized samplesheet is written as a whitespace-free relative path under the output directory so the upstream --input schema accepts it (the schema accepts .csv, .tsv, .yaml, .yml, .json).
  • Reference handling follows Sarek's two documented modes: use --genome <iGenomes> and optionally override individual reference files (or pass false for a resource that should not be used), or use --genome null --igenomes-ignore --fasta <reference> when no iGenomes reference files should be loaded. Optional FASTA indices and tool resources may be supplied in either mode when appropriate.
  • In --build-only-index mode, Sarek intentionally supplies an empty samplesheet channel: the wrapper does not require sample pairing or per-sample caller outputs, while it preserves upstream global resource guards (including BQSR guards on preprocessing start steps) and captures published reference outputs.
  • Tool×mode compatibility is evaluated per-patient: a tool is accepted when at least one patient matches its required mode, so mixed samplesheets (germline-only patients alongside tumor/normal pairs) are valid. Paired-only tools (ascat, msisensorpro, muse) still need at least one patient with both status=0 and status=1.
  • Mutect2 without an effective PON or germline resource emits a preflight warning but does not block; resources inherited from an iGenomes bundle count as effective. The bundled GATK.GRCh38 PON still emits a recommendation to use a project-specific PON.
  • Paired somatic Mutect2 cannot be run with --no-intervals; the upstream schema explicitly marks that combination unsupported.
  • --snv-consensus-calling requires --normalize-vcfs, as enforced by the upstream workflow before post-variant processing.
  • --use-gatk-spark markduplicates is incompatible with header/positional UMI dedup (--umi-in-read-header or --umi-location); it is fine with --umi-read-structure (fgbio consensus runs upstream).
  • ASCAT requires an effective --ascat-genome, --ascat-alleles, and --ascat-loci (which supported iGenomes bundles can provide). With --wes, custom --ascat-alleles, --ascat-loci, --ascat-loci-gc, and --ascat-loci-rt resources are recommended; the wrapper warns because Sarek documents its iGenomes ASCAT resources as unsuitable for WES.

Example Queries

  • "Run nf-core/sarek for germline variant calling on these WES FASTQs"
  • "Call somatic variants from this tumor-normal pair with Mutect2 and Strelka"
  • "Annotate this VCF with VEP using sarek"
  • "Tumor-only Mutect2 with PON for our WES cohort"
  • "Restart sarek at the recalibration step"

Example Output

output/                                       # the --output directory
├── .nextflow/                                # Nextflow cache/history (framework-created; excluded from checksums)
├── .nextflow.log                            # Nextflow launch log (framework-created; excluded from checksums)
├── upstream/
│   ├── results/                              # Nextflow --outdir
│   │   ├── csv/                              # handoff CSVs (mapped/markduplicates/recalibrated/variantcalled); legacy fallback: preprocessing/csv/
│   │   ├── preprocessing/
│   │   │   ├── mapped/                       # §1 aligned CRAMs (one per sample/lane)
│   │   │   ├── markduplicates/               # §2 deduplicated CRAMs
│   │   │   └── recalibrated/                 # §3 BQSR-recalibrated CRAMs
│   │   ├── variant_calling/
│   │   │   ├── haplotypecaller/              # §4 germline VCFs
│   │   │   ├── mutect2/                      # somatic / tumor-only VCFs
│   │   │   ├── strelka/                      # somatic + germline VCFs
│   │   │   ├── manta/                        # SV VCFs
│   │   │   ├── ascat/                        # CNV / purity / ploidy
│   │   │   ├── controlfreec/                 # CNV
│   │   │   ├── tiddit/                       # SV
│   │   │   ├── bcftools/                     # mpileup caller output
│   │   │   ├── msisensor2/                   # MSI
│   │   │   └── msisensorpro/                 # MSI (paired MSIsensorPro)
│   │   ├── annotation/<variantcaller>/<sample_or_pair>/   # §5 SnpEff/VEP/merge/bcftools/SnpSift annotated VCFs
│   │   ├── multiqc/                          # §6 MultiQC HTML + data
│   │   ├── pipeline_info/
│   │   └── reports/
│   └── work/                                 # Nextflow work directory
├── report.md                                # human-readable run summary (output root)
├── result.json                              # machine-readable run summary (output root)
├── check_result.json                        # written only with --check (preflight-only mode); parallel to result.json
├── logs/                                    # Nextflow stdout.txt / stderr.txt (real runs only; excluded from checksums)
└── reproducibility/                          # replay + provenance bundle
    ├── samplesheet.valid.csv                 # or samplesheet.demo.csv in --demo mode
    ├── params.yaml
    ├── commands.sh
    ├── remap_paths.py
    ├── environment.yml
    ├── checksums.sha256
    ├── compatibility_policy.json
    ├── parameters.json
    ├── samplesheet.json
    ├── pipeline_source.json
    ├── tool_versions.json
    ├── outputs.json                          # omitted if outputs parsing was skipped
    ├── manifest.json
    ├── macos_docker.config                   # written only on macOS + docker backend
    └── sarek_downstream_handoff.{sh,json}    # written only when --run-downstream is set

report.md, result.json, and logs/ sit at the output root — the same layout as the nfcore-rnaseq and nfcore-scrnaseq wrappers — so a consumer finds <output>/result.json for any of the three pipelines. The reproducibility/ directory holds the portable replay + provenance bundle.

Output Structure

Under the output/ root the wrapper writes two child directories — upstream/ (the Nextflow results/ tree plus its work/ directory) and reproducibility/ (the portable replay + provenance bundle: the params/commands/samplesheet snapshots, the seven JSON provenance files, environment.yml, checksums.sha256, and — macOS + docker only — macos_docker.config) — alongside the run-summary files report.md and result.json. A real run also writes a root-level logs/ directory (Nextflow stdout.txt/stderr.txt), and --check writes check_result.json at the root (parallel to result.json); both placements match the nfcore-rnaseq and nfcore-scrnaseq wrappers. Nextflow itself additionally writes its own hidden bookkeeping in the launch directory — .nextflow/ (cache/history) and .nextflow.log — because the wrapper runs Nextflow with cwd = output_dir so the relative input/outdir paths resolve; both are excluded from checksums.sha256 (.nextflow directory and any .log file are skipped). There is no separate top-level provenance/ directory; all provenance JSON is co-located in reproducibility/. The reproducibility/ tree, the root logs/ directory, and the root summaries report.md/result.json/check_result.json are all excluded from checksums.sha256, so execution logs and wrapper summaries never enter the manifest. The §1–§6 layout in outputs_parser.py corresponds to: §1 mapped, §2 markduplicates, §3 recalibrated, §4 per-tool variant calls, §5 annotation, §6 MultiQC.

Cross-machine / cross-OS portability. The bundle stores absolute data/reference paths (required by Nextflow) but ships a stdlib-only remap_paths.py to rebase them on any host: --old/--new rewrites samplesheet data paths, --refs-old/--refs-new rewrites reference/index paths in params.yaml (and commands.sh if any were added there), --output-dir <new-path> rewrites the baked --output in commands.sh when you relocate the run, and --verify confirms every path resolves before replay. (The scrnaseq bundle self-relocates and needs no --output-dir; it accepts the flag only for parity.) URIs (s3://, https://, …) and the false disable sentinel are preserved. All bundle files use POSIX paths and utf-8/\n, so macOS↔Linux replay is byte-stable. The recommended replay path is a self-contained bash commands.sh (no environment variable required): it self-anchors via BASH_SOURCE, pins the Nextflow engine with NXF_VER, and applies the macOS-only Docker config through a uname-gated -c reproducibility/macos_docker.config, so the same bundle replays identically on Linux and macOS.

Dependencies

Required

  • Python >=3.11
  • Java >=17
  • Nextflow >=25.10.2
  • One execution backend: Docker, Singularity, Apptainer, Podman, Conda/Mamba, Shifter, or Charliecloud

Gotchas

Shortened here. Read the whole file on GitHub.

Signals

GitHub stars
1k
Forks
277
Last commit
Sep 2026
Advanced
Item type
skill
Key
nfcore-sarek-wrapper
Source
github.com/clawbio/clawbio