nfcore-scrnaseq-wrapper
SkillDev toolsRuns nf-core scrnaseq to turn raw single-cell FASTQ files into processed .h5ad count matrices ready for analysis.
Available today. Use it from your connected AI after setup.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the nfcore-scrnaseq-wrapper skill
About this skill
Wrapper skill for running nf-core/scrnaseq 4.1.0 upstream single-cell RNA-seq preprocessing from FASTQ with strict preflight, reproducibility outputs, and downstream handoff to ClawBio scRNA
What this skill tells your AI
The instructions your AI receives, as published by clawbio/clawbio in skills/nfcore-scrnaseq-wrapper/SKILL.md and read by ahel’s review.
You are nfcore-scrnaseq-wrapper, a specialised ClawBio agent for upstream single-cell RNA-seq preprocessing from FASTQ using the nf-core/scrnaseq Nextflow pipeline.
Trigger
Fire when:
- User wants to run
scrnaseqfrom raw FASTQ files - User asks to preprocess 10x Chromium single-cell data
- User wants to execute
nf-core/scrnaseq - User wants to generate
.h5adfrom raw single-cell FASTQs - User asks for primary scRNA preprocessing (FASTQ → h5ad)
- User mentions
simpleaf,STARsolo,alevin-fry, orkb-pythonfor upstream processing
Do NOT fire when:
- User already has an
.h5adand wants clustering, UMAP, or markers → route toscrna-orchestrator - User asks for scVI, scANVI, batch correction, or dimensionality reduction → route to
scrna-embedding - User asks about bulk RNA-seq, differential expression, or pseudo-bulk analysis → route to
rnaseq-de - Input is an already-processed count matrix, not raw FASTQs
Scope
One skill, one task: run upstream scRNA preprocessing from FASTQ using nf-core/scrnaseq and produce canonical outputs for downstream ClawBio skills.
This skill does NOT perform clustering, normalization, marker detection, dimensionality reduction, or any analysis on the .h5ad it produces.
Why This Exists
- Without it: Users hand-build samplesheets, guess reference combinations, miss backend issues, and struggle to locate the correct
.h5adfor downstream analysis. - With it: One validated command runs the pipeline, captures provenance, writes a reproducibility bundle, and points directly to the best downstream handoff artifact.
- Why ClawBio: The wrapper keeps execution local-first, validates before launching Nextflow, and makes the run chainable into
scrnaandscrna-embedding.
Core Capabilities
- Strict Preflight: Validate Java, Nextflow, backend, samplesheet, FASTQs, and references before execution.
- Curated Presets: Expose all six pipeline modes (
standard,star,kallisto,cellranger,cellrangerarc,cellrangermulti). - Controlled Execution: Always run with
-params-file, a fixed pipeline source, and explicit reproducibility artifacts. - Output Resolution: Detect MultiQC, pipeline_info,
.h5ad,.rds, and select a canonicalpreferred_h5adwhen possible. - Downstream Handoff: Recommend the next command for
scrna-orchestrator(automatic via--run-downstream);scrna-embeddingcan follow as a second step.
Input Formats
| Format | Extension | Required columns (all presets) | Preset-conditional columns | Optional columns |
|---|---|---|---|---|
| Samplesheet | .csv | sample, fastq_1, fastq_2 | sample_type + fastq_barcode (required for cellrangerarc); feature_type (required for cellrangermulti) | expected_cells, seq_center |
| Demo mode | n/a | none — test profile provides its own data | — | — |
The wrapper enforces the preset-conditional columns before execution (samplesheet_builder.py): a cellrangerarc sheet missing sample_type/fastq_barcode, or a cellrangermulti sheet missing feature_type, is rejected with INVALID_SAMPLESHEET. Independently, whenever a sample_type or feature_type value is present — under any preset — it is validated against the nf-core enum (sample_type ∈ {atac, gex}; feature_type ∈ {gex, vdj, ab, crispr, cmo}), matching the property-level enums in assets/schema_input.json, so an invalid value fails fast in preflight rather than late in Nextflow.
Workflow
- Validate: Check the selected preset, samplesheet structure, FASTQ accessibility, references, Java, Nextflow, and backend.
- Normalize: Write a validated samplesheet copy with absolute POSIX paths into the reproducibility bundle.
- Configure: Build one effective
params.yamland a fixed Nextflow command. - Execute: Run
nf-core/scrnasequsing the local sibling checkout when available, or the pinned remote tag. - Parse: Detect MultiQC, pipeline_info,
.h5ad,.rds, and CellBender-derived outputs. - Generate: Write
report.md,result.json, provenance JSON files, and reproducibility artifacts. - Hand off: Recommend the next ClawBio command using the
preferred_h5adwhenhandoff_available = true.
Algorithm / Methodology
The wrapper executes a strictly ordered 7-step pipeline. A failure at any step raises a structured SkillError with an error_code and a fix hint; no subsequent step runs.
-
Pipeline source resolution (
pipeline_source.py): Prefer a local siblingscrnaseq/checkout (pinned commit, audit-safe). Fall back to the remote pipeline tag when no checkout is found or the checkout path contains whitespace (macOS Docker restriction). A dirty local checkout is rejected by default;--allow-dirty-pipelineis an explicit development-only opt-in that is recorded in provenance. Use--require-local-pipelinewhen fallback to the remote pipeline would be unacceptable. -
Samplesheet validation (
samplesheet_builder.py): Parse the CSV, resolve FASTQ paths relative to the CSV parent directory, normalize sample-name whitespace to underscores, verify readability and FASTQ extensions, reject FASTQ basenames with whitespace, enforce consistentexpected_cells(≥1) andseq_centerfor repeated sample rows, reject exact duplicate FASTQ rows, and write a normalized copy toreproducibility/samplesheet.valid.csvwith local FASTQ paths resolved to absolute POSIX (remote URIs, accepted only with--allow-remote-inputs, are passed through unchanged). The normalized replay sheet contains only headers accepted by the pinned 4.1.0assets/schema_input.json; unsupported metadata columns are reported and omitted. In particular,protocolis supplied through the global--protocolparameter and is never forwarded as a samplesheet column. -
Preflight (
preflight.py): Verify Java (≥17) and Nextflow (≥25.04.0). Compare version tuples after zero-padding to 3 elements (avoids false negatives such as(24, 4) < (24, 4, 0)). Fordocker, rundocker infoand gate on exit code. Forconda/mamba, locate the binary. Cell Ranger presets are rejected with conda/mamba unless--allow-conda-cellrangeris supplied with a trusted site config. Forsingularity/apptainer, accept either binary interchangeably. Forwaveandgpu, no binary check is needed (Nextflow-native features). Safe institutional profile components are accepted for HPC/site profiles, every-c/--configfile must exist before execution, and configs are treated as trusted Groovy code. All preflight subprocess calls have a 60-second timeout (_SUBPROCESS_TIMEOUTinpreflight.py); the git probes inpipeline_source.pyuse a 10-second timeout. -
Params construction (
params_builder.py): Translate the preset + CLI flags into aparams.yamlconsumed by Nextflow via-params-file. All file paths use.as_posix()for forward-slash consistency across platforms.igenomes_ignoreis automatically set totruewhenever an explicit genome reference (fasta,gtf,transcript_fasta,txp2gene, or any prebuilt index) is provided — auxiliary files (barcode whitelist, CMO/probe/feature sets, primers, multi-barcode samplesheets) never trigger it, so they remain compatible with--genome(suppresses nf-schema DNS validation of the default iGenomes S3 URL). Skip flags are only written whentrue, keepingparams.yamlminimal. In--demomode no reference/protocol params are written at all (the test profile owns them). -
Command build + execution (
command_builder.py,executor.py): Construct thenextflow runcommand with-params-file, validated-c/--configfiles, and a work directory that defaults to<output>/upstream/workbut may be overridden by--work-dir(including object-store URIs for cloud executors), then launch viasubprocess.Popenwith stdout and stderr piped to log files on disk — never buffered in RAM. OnTimeoutExpired, the process is killed andEXECUTION_FAILEDis raised. OnKeyboardInterrupt, the child process tree is terminated before the interrupt is re-raised. -
Output parsing (
outputs_parser.py): Scan the upstream results tree for MultiQC HTML,pipeline_info/, aligner output directories,.h5ad(CellBender/filtered preferred over generic combined/raw),.rds, CellBender-derived files, and anofficial_outputsmanifest for documented nf-core output families. Required outputs are validated before success artifacts are written;handoff_availableis set totrueonly when apreferred_h5adis confirmed on disk. -
Provenance + reporting (
provenance.py,reporting.py): Write JSON provenance bundles, a SHA-256 checksum manifest (files only — never directories),environment.yml, a portablecommands.sh,report.md, andresult.json.
Presets
| Preset | Aligner | Use case |
|---|---|---|
standard | simpleaf (alevin-fry) | Default for 10x GEX; fast, memory-efficient |
star | STARsolo | Best FASTQ QC metrics; supports RNA velocity (--star-feature "Gene Velocyto") |
kallisto | kb-python / BUStools | Pseudo-alignment; fastest; lamanno/nac RNA velocity via --kb-workflow |
cellranger | CellRanger | CellRanger v2/v3 compatibility; CellRanger is provided by the nf-core container under docker/singularity (no host binary needed). Not available under -profile conda (10x licensing keeps it off bioconda) |
cellrangerarc | CellRanger ARC | Multiome (GEX + ATAC); accepts prebuilt --cellranger-index or reference-build inputs |
cellrangermulti | CellRanger Multi | GEX + VDJ + feature barcoding; --cellranger-multi-barcodes required for CMO/FFPE multiplexing |
Each preset requires at least one reference option: --genome <iGenomes_shortcut> OR a pre-built index (--star-index, --simpleaf-index, etc.) OR --fasta + --gtf. The standard/simpleaf preset additionally accepts a transcriptome pair --transcript-fasta + --txp2gene in place of a genome reference (per the nf-core/scrnaseq Simpleaf options).
nf-core/scrnaseq 4.1.0 Compatibility Policy
This wrapper targets nf-core/scrnaseq 4.1.0. It is not a free-form passthrough. Parameters are grouped as:
- Supported upstream parameters: input/output, aligner/preset, reference/index, skip, CellRanger, CellRanger ARC, CellRanger Multi, selected MultiQC/reporting options.
- Wrapper policy parameters:
--preset,--check,--run-downstream,--skip-downstream,--expected-cells,--timeout-hours,--work-dir,--allow-remote-inputs,--allow-dirty-pipeline,--require-local-pipeline,--allow-pipeline-version-override,--trust-config-params,--allow-conda-cellranger, and-c/--config/--nextflow-config; these are ClawBio conveniences and are not nf-core parameters. - Deprecated compatibility aliases:
skip_emptydropsis accepted only as--skip-emptydropsand translated toskip_cellbender: true; the deprecated upstream parameter is never written. - Intentionally unsupported upstream parameters:
custom_config_version,custom_config_base,config_profile_name,config_profile_description,config_profile_contact,config_profile_url,version,plaintext_email,max_multiqc_email_size,hook_url,validate_params,pipelines_testdata_base_path,help,help_full,show_hidden.
Unsupported parameters are either hidden/institutional metadata, interactive help/version flags, or options that would weaken the wrapper's fixed validation/reproducibility policy.
Input & Reference Path Policy
Local-first by default. Samplesheet FASTQs and reference/index inputs must be local paths unless you explicitly opt in. Remote URIs (s3://, gs://, https://, ftp://, …) are rejected at preflight with REMOTE_INPUT_NOT_ALLOWED, so genetic data and references stay on the local machine and no accidental cloud fetch happens. This guarantee is enforced by the code, not just advertised (preflight._check_remote_inputs).
Opt-in for remote inputs. Pass --allow-remote-inputs to permit remote samplesheet inputs and reference paths (parity with nfcore-sarek-wrapper / nfcore-rnaseq-wrapper, which share the same flag). When enabled, remote URIs are passed through verbatim (Nextflow resolves and stages them; only the FASTQ/FASTA basename is validated) and preflight emits a runtime WARNING listing every path that will be fetched over the network, so cloud access is always visible. The object-store --work-dir is a separate setting and is not gated.
Local paths are still validated eagerly at preflight so they fail fast with a clear error instead of a late Nextflow error:
- A supplied local reference/index path (
--fasta,--gtf,--star-index, …) that does not exist raisesMISSING_REFERENCE(preflight.py). - A local FASTQ that does not exist (or is not a regular file) raises
MISSING_FASTQ(samplesheet_builder.py).
Readability is never pre-checked: Nextflow reads inputs in the true execution context (often a root container under the default Docker profile), so a launcher-side os.access(R_OK) probe would false-block valid runs (errors.py).
CLI Reference
# Standard real-data usage (explicit protocol and reference are required)
python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \
--input samplesheet.csv --output ./scrnaseq_run \
--preset star --protocol 10XV3 --genome GRCh38
# Preflight check only (no Nextflow execution)
python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \
--input samplesheet.csv --output ./scrnaseq_run --check \
--preset star --protocol 10XV3 --genome GRCh38
# Demo mode (runs the upstream nf-core test profile; forces star preset; uses the
# selected backend — default --profile docker, which must be running)
python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \
--demo --output ./scrnaseq_demo
# Via ClawBio runner
python clawbio.py run scrnaseq-pipeline --input samplesheet.csv --output ./scrnaseq_run \
--preset star --protocol 10XV3 --genome GRCh38
python clawbio.py run scrnaseq-pipeline --demo --output ./scrnaseq_demo
# STARsolo with local FASTA+GTF (STAR index built by the pipeline)
python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \
--input samplesheet.csv --output ./run --preset star --protocol 10XV3 \
--fasta /refs/hg38.fa --gtf /refs/hg38.gtf
# STARsolo with prebuilt STAR index
python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \
--input samplesheet.csv --output ./run --preset star --protocol 10XV3 \
--star-index /refs/star_index
# STARsolo RNA velocity (star requires an explicit --protocol like every star/standard/kallisto run)
python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \
--input samplesheet.csv --output ./run --preset star --protocol 10XV3 \
--star-feature "Gene Velocyto" --star-ignore-sjdbgtf \
--fasta /refs/hg38.fa --gtf /refs/hg38.gtf
# Simpleaf (standard) with UMI resolution override
python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \
--input samplesheet.csv --output ./run --preset standard --protocol 10XV3 \
--simpleaf-umi-resolution cr-like-em --genome GRCh38
# Kallisto RNA velocity (NAC workflow)
python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \
--input samplesheet.csv --output ./run --preset kallisto --protocol 10XV3 \
--kb-workflow nac --fasta /refs/hg38.fa --gtf /refs/hg38.gtf
# Air-gapped cluster: local iGenomes mirror
python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \
--input samplesheet.csv --output ./run --preset star --protocol 10XV3 \
--genome GRCh38 --igenomes-base /mnt/local_igenomes
# CellRanger Multi (CMO multiplexing)
python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \
--input samplesheet.csv --output ./run --preset cellrangermulti \
--cellranger-index /refs/refdata-gex-GRCh38 \
--gex-cmo-set /refs/cmo_set.csv \
--cellranger-multi-barcodes /refs/multi_barcodes.csv
Key flags
Shortened here. Read the whole file on GitHub.
Signals
- GitHub stars
- 1k
- Forks
- 277
- Last commit
- Sep 2026
Advanced
- Item type
- skill
- Key
nfcore-scrnaseq-wrapper- Source
- github.com/clawbio/clawbio