Single-Cell RNA-seq Gold Chain (single-cell-rna-qc)

SkillDev tools

scverse scRNA gold chain on .h5ad/.h5, inspect, convert, MAD QC, optional scanpy.pp.scrublet, preprocess, Harmony/ComBat, PCA-UMAP-Leiden, Wilcoxon markers, subset, pseudobulk, pydeseq2, stable plots. Use when the user has scRNA-seq counts. Does not assign cell-type labels. rank_genes_groups is not condition DE. Do not call local doublet/ambient helpers SoupX, CellBender, or scDblFinder.

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the Single-Cell RNA-seq Gold Chain (single-cell-rna-qc) skill

What this skill tells your AI

The instructions your AI receives, as published by herry423/bionexus in skills/single-cell-rna-qc/SKILL.md and read by ahel’s review.

[!NOTE] Gold Reference Plugin: This skill serves as the canonical reference implementation for all BioNexus analytical tools. All contributors must follow this architecture.

The single-cell-rna-qc skill implements the official scverse community standard workflow for single-cell RNA sequencing data. It strictly enforces the distinction between Execution Fidelity and Scientific Evidence Quality, providing multi-dimensional EvidenceCard validation, input distribution audits, and reproducible provenance sidecars.


⚡ Canonical Workflow (Command-Line Interface)

The canonical execution path is the modular scverse gold chain:

# 0. Verify environment backend (scanpy, anndata)
python scripts/doctor.py --require-scverse

# 1. Inspect data semantics, sparsity, library size, and log-transformation
python skills/single-cell-rna-qc/scripts/scrna_inspect.py raw.h5ad

# 2. Run the complete canonical pipeline (MAD QC -> norm/log1p -> HVG -> PCA -> Leiden -> markers)
python skills/single-cell-rna-qc/scripts/scrna_pipeline.py raw.h5ad -o clustered.h5ad

# 3. Official doublet detection on raw count layer (scanpy.pp.scrublet)
python skills/single-cell-rna-qc/scripts/scrna_scrublet.py raw.h5ad -o raw_scrub.h5ad

# 4. Pseudobulk aggregation by biological replicate × condition
python skills/single-cell-rna-qc/scripts/scrna_pseudobulk.py clustered.h5ad -o pb.csv --by sample condition --design pb_design.tsv

# 5. Condition Differential Expression (Wald test via PyDESeq2)
python skills/single-cell-rna-qc/scripts/scrna_deseq.py pb.csv --design pb_design.tsv --condition condition --reference control --contrast-level treated -o de.csv

# 6. Generate standardized exploratory figures
python skills/single-cell-rna-qc/scripts/scrna_plot.py clustered.h5ad -o figures/ --color leiden

# 7. Multi-donor DE Evidence Audit (before lab meetings, manuscript submission, or sharing)
bionexus audit-de clustered.h5ad --de-table de.csv -o audit_report.md

🧬 Canonical Scripts & Architecture Matrix

StepScriptCanonical BackendEvidence GradeOutput Artifacts
Inspectscrna_inspect.pyanndata + bionexus.integrityAJSON summary (sparsity, log-scale check)
Convertscrna_convert.pyscanpy.read_*AStandardized .h5ad
QC (MAD)qc_core.pyscanpy / Median Absolute DeviationAFiltered .h5ad + QC metrics
Doubletsscrna_scrublet.pyscanpy.pp.scrublet onlyA.h5ad with doublet scores (Refuse if missing)
Preprocessscrna_preprocess.pyscanpy.pp.normalize_total, log1p, highly_variable_genesAPreprocessed .h5ad
Integratescrna_integrate.pyharmonypy / scanpy.pp.combatABatch-corrected PCA space
Clusterscrna_reduce_cluster.pyscanpy.tl.pca, neighbors, umap, leidenAClustered .h5ad (Numeric labels only)
Markersscrna_markers.pyscanpy.tl.rank_genes_groups (Wilcoxon)ACluster marker gene rankings table
Plotscrna_plot.pyscanpy.pl / matplotlibAumap_leiden.png, dotplot_markers.png, violin_qc.png
Subsetscrna_subset.pyanndata slice & stale embedding dropASubsampled .h5ad
Pseudobulkscrna_pseudobulk.pySum raw counts over replicate groupsAPseudobulk count matrix pb.csv + pb_design.tsv
Condition DEscrna_deseq.pypydeseq2 (Wald test)ADifferential expression table de.csv (Refuse if missing)

🛡️ Scientific Honesty Invariants & Non-Negotiables

  1. Numeric Cluster Labels Only:
    • The pipeline writes numeric cluster identities (leiden: "0", "1", "2").
    • Strictly Forbidden: Guessing or hallucinating biological cell-type labels (e.g. "T-cell", "Macrophage") without validated reference annotations or orthogonal experimental ground truth.
  2. Marker Genes vs Condition DE:
    • Exploratory marker gene identification (rank_genes_groups) discovers cluster-specific expression within a single dataset.
    • Strictly Forbidden: Publishing exploratory marker p-values as experimental condition treatment effect p-values. Condition DE requires pseudobulk replicate aggregation and pydeseq2.
  3. No Masquerading Heuristics:
    • Local fallback scripts (ambient_rna.py, doublet_detection.py, qc_analysis.py) are legacy Grade C heuristics.
    • Strictly Forbidden: Claiming or labeling local heuristics as official community algorithms like SoupX, CellBender, or scDblFinder.
  4. Deterministic Refusal:
    • If a gold-standard backend (scanpy, pydeseq2) is missing, the tool must cleanly return refuse() with EvidenceGrade.ABSTAIN.

📊 Scientific EvidenceCard Contract

Every analytical run produces a structured EvidenceCard evaluating the 7 quality dimensions:

{
  "method": "scanpy_gold_chain",
  "backend": "scanpy",
  "evidence_grade": "A",
  "conclusion_status": "SUPPORTED",
  "evidence_card": {
    "execution_fidelity": "A",
    "input_integrity": "A",
    "assumption_validity": "A",
    "statistical_support": "A",
    "parameter_robustness": "B",
    "cross_method_concordance": "UNTESTED",
    "external_validation": "UNTESTED"
  },
  "limitations": [
    "This plugin does not assign cell-type identity. Leiden/KMeans labels are numeric only.",
    "Research-use only. Not a clinical diagnostic, not CLIA/CAP validated, and not an authorized medical device."
  ]
}

⚠️ Deprecated / Legacy Scripts (Grade C Heuristics)

The following scripts are maintained solely for backward compatibility. They are not part of the default canonical path:

  • qc_analysis.py (Replaced by scrna_pipeline.py)
  • doublet_detection.py (Replaced by scrna_scrublet.py)
  • ambient_rna.py (Local NNLS heuristic; not SoupX/CellBender)

Signals

GitHub stars
31
Forks
4
Last commit
Sep 2026
Advanced
Item type
skill
Key
single-cell-rna-qc
Source
github.com/herry423/bionexus