DnaSP

SkillDev tools

Run population genetics stats like diversity, neutrality tests and linkage disequilibrium on aligned DNA sequences.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the DnaSP skill

About this skill

Population genetics of pre-aligned DNA sequences or multi-sample VCFs

What this skill tells your AI

The instructions your AI receives, as published by clawbio/clawbio in skills/dnasp/SKILL.md and read by ahel’s review.

Trigger

Fire when a user requests population-genetic analysis of aligned DNA, a supported VCF, or a DnaSP-compatible statistic listed below. Do NOT fire for sequence alignment, read mapping, haplotype phasing, clinical advice, or significance tests other than the coalescent P-values --n-sim gives for D, R2 and Fs.

Scope

Analyse genetic variation in supplied alignments using 16 selected DnaSP methods. This is not a complete replacement for the DnaSP GUI or all its analysis modes. Read the statistical reference for definitions, exclusions, source conventions, file formats, examples and release validation evidence.

Core Capabilities

  • Polymorphism: S, Eta, haplotypes, Hd, VarHd, pi, k, theta-W, G+C, Tajima's D, Fu and Li D*/F*, Ramos-Onsins and Rozas R2.
  • LD: D, D', R2, ZnS, Za, ZZ and original-column pair labels; recombination Rm.
  • Mismatch distribution with unbiased variance over unordered pairs and Sokal & Rohlf corrected CV; Harpending (1994) raggedness; Model 1 diallelic InDel diversity.
  • Divergence and Hudson Fst between populations; outgroup Fu and Li D/F.
  • Two-locus HKA, McDonald-Kreitman, Nei-Gojobori Ka/Ks, Fu's Fs and SFS.
  • Opt-in coalescent P-values (--n-sim) for Tajima's D, R2 and Fu's Fs.
  • Ts/Tv, codon counts/RSCU including stops, named per-sequence ENC and its synonymous-codon-weighted summary.
  • Sliding windows mirror DnaSP's Gaps in Sliding Window = considered mode: starts advance from 1 by the step, capped at alignment length; ends are capped there too, and the first window reaching the end terminates the loop. Empty windows are retained. The not-considered mode is not implemented.
  • Window midpoints are the alignment position of the ceil(net/2)-th gap-free column, or the window start if none exists; exported in TSV, JSON and report and used for the plot. VCF windows use retained SNP indices, including the final partial window.
  • Raw per-site Fay-Wu H and theta-L minus theta-W; normalised Hn/ZE are absent.

Fu and Li D*/F* and outgroup D/F mirror Data > Segregating Sites/Mutations = Segregating sites, using DnaSP's v5-style panels. The rp49 Eta-setting D*/F* figures can be reproduced by substituting eta for S, but this is not a general conversion: FULI.vb also changes singleton and external-mutation capping (SingleMut and ExternaMut subtraction). No Eta-mode switch is implemented.

Workflow

  1. Confirm the input is pre-aligned and identify the scientific comparison: ingroup, explicit outgroup, populations, coding interval and genetic code. Do not infer an outgroup from record order or silently select a code.
  2. Use the full repository checkout or supplied validation code snapshot so the shared reproducibility helpers are available. Install the declared dependencies.
  3. Choose an empty output directory. For alignment input use --input; for VCF use --vcf; for two-locus HKA alone use --hka-file --analysis hka.
  4. Select actual implemented names with --analysis. Supply --outgroup, --input2 or --pop-file when needed. Use --genetic-code vertebrate-mitochondrial for the matching mitochondrial table. The default is standard. Coding intervals must be preselected and divisible by three.
  5. Execute the CLI. An explicit analysis that cannot run returns a non-zero code. --analysis all is opportunistic: inspect its completed/skipped manifest.
  6. Read stderr diagnostics, report.md and reproducibility/manifest.json. Distinguish a failed analysis, an undefined statistic and an excluded site.
  7. Interpret the chosen statistic within its documented assumptions. Do not turn a signed value or a threshold into a significance claim; only --n-sim P-values assess significance, and a rejection does not by itself identify its cause.
  8. Retain the input archive, settings, hashes and environment with the report. commands.sh replays on the recorded host/code path into a new output folder; moving a run requires the code and dependencies as well as its input archive.
  9. For Windows GUI comparison, follow the separate validation checklist and capture raw DnaSP output. Source-derived expectations are not GUI observations.

Example Output

python skills/dnasp/dnasp.py --demo --output new_demo_run
python skills/dnasp/dnasp.py --input alignment.fas --analysis polymorphism,ld --output new_run
python skills/dnasp/dnasp.py --input coding.fas --outgroup OutSeq --genetic-code vertebrate-mitochondrial --analysis mk,kaks,codon --output new_coding_run
python skills/dnasp/dnasp.py --vcf samples.vcf --analysis polymorphism,sfs,fufs --output new_vcf_run

The synthetic demo contains 10 ingroup sequences, one outgroup and 300 sites:

QuantityDemo value
S / haplotypes5 / 8
Hd / Tajima D0.9556 / 0.6789
MK Pn / Ps / Dn / Ds2 / 3 / 2 / 1
MK alpha0.6667
Ka / Ks / omega0.010239 / 0.030291 / 0.3380
Ts / Tv4 / 1

The demo runs 15 modules; HKA uses a separate two-locus input. The regression suite, rather than the printed banner alone, checks the expected figures.

Output Structure

output/
  report.md
  results.tsv                  # polymorphism and windows
  summary.json                 # module summaries, window midpoints, named ENC/null
  result.json                  # ClawBio envelope: headline summary, summary.json payload, artifacts
  ld_pairs.tsv                 # when LD pairs exist
  figures/                     # when matplotlib is available
  reproducibility/
    inputs/
    commands.sh
    environment.yml
    manifest.json
    checksums.sha256

Dependencies

Python 3.10 or later. Core estimators use the standard library. Plotting uses matplotlib; the repository's shared reproducibility package also imports NumPy, pandas and OpenTelemetry. Use environment.yml or the validation package's pinned requirements. The Windows validation runner targets Python 3.12.

Gotchas

  • The model will want to count IUPAC symbols as alleles. Do not. Supported ambiguity symbols are missing data and the nucleotide mask excludes their columns.
  • The model will want to analyse an unknown outgroup as an ingroup-only run. Do not. Identifiers must be unique and an explicit outgroup must resolve.
  • The model will want to call VCF diversity per-base diversity. Do not. VCF columns represent retained variant records; invariant callable bases are absent.
  • The model will want to use any coding annotation in a NEXUS file. Do not. This implementation requires a preselected coding alignment and does not read CHARSET coding annotations. In COII, use positions 1-681 for codon usage.
  • The model will want to apply one stop-codon rule everywhere. Do not. Selected stops count in RSCU and as family 21 in MK/KaKs; ENC and coding G+C use sense codons.
  • The model will want to read raw H/E as normalised DnaSP Hn/ZE. Do not. These outputs are explicitly different, and no significance test is supplied.
  • The model will want to overwrite an earlier output folder. Do not. Choose a fresh folder; runs reject non-empty destinations to preserve prior artefacts.
  • The model will want to infer genomic-scale performance from the tiny bundled VCF examples. Do not. LD enumerates all biallelic-site pairs and uses quadratic storage; read the measured limits in the reference before a large run.

Known Differences from DnaSP 6.12.03

Four values compared with DnaSP 6.12.03 are unresolved. They come from the DnaSP VCF examples with a single variant site: the phased diploid Scaffold_2 (n = 20) and the haploid Region_MSA_2 (n = 10). The skill follows the DnaSP 6 source code and the result files shipped with DnaSP 6.0.60; the captures cannot tell whether DnaSP's build or its analysis route changed.

StatisticSkillDnaSP 6.0.60DnaSP 6.12.03
Pi, Scaffold_20.10.1n.a.
R2, Scaffold_20.2179450.2179450.259808
Pi, Region_MSA_20.20.2n.a.
R2, Region_MSA_20.30.30.346410

Differences of setting or definition, not errors:

  • Fu and Li's tests follow DnaSP's Segregating sites setting; DnaSP's default Eta setting gives different D*, F*, D and F.
  • DnaSP prints two F* forms; the skill reproduces the "DnaSP v5" form (Simonsen et al. 1995), not the biallelic-positions form (Achaz 2009) that DnaSP's multi-alignment output reports.
  • A coding region whose final codon is a stop in every sequence is analysed without that codon; DnaSP keeps it if the user declines its prompt.
  • The folded SFS excludes multiallelic sites; DnaSP's segregating-sites spectrum places them in a frequency class.
  • Sliding windows with no segregating site report Tajima's D as undefined; DnaSP prints 0.0000.
  • Ts/Tv has no DnaSP 6 counterpart, and raw Fay and Wu H and Zeng E are not DnaSP's normalised Hn and ZE.

Version History

Current version 0.6.0. Every change that alters a result, and the version compared with DnaSP 6.12.03 (0.5.2), is recorded in docs/version_history.md.

Safety

All sequence analysis and output remain local. ClawBio is a research and educational tool. It is not a medical device and does not provide clinical diagnoses. Consult a healthcare professional before making any medical decisions.

Agent Boundary

The agent selects documented inputs/options, executes the skill and explains reported results. The code computes the statistics. Neither the agent nor the skill may fabricate GUI validation, P-values or missing results.

Integration with Bio Orchestrator

CLI alias: dnasp. Use --analysis for module selection, not invented flags such as --pi, --kaks or --tajima. The repository dispatcher permits the implemented options, including VCF, populations and genetic-code selection.

Route here when the user asks for a statistic this skill computes (for example nucleotide or haplotype diversity, Tajima's D, Fu and Li's tests, Fu's Fs, McDonald-Kreitman, Ka/Ks, HKA, the mismatch distribution, InDel polymorphism or codon usage bias) on an aligned FASTA or NEXUS file or a multi-sample VCF. INTENTS.json publishes the dnasp aliases to ClawBio's intent planner and plans the demo only when the user explicitly asks for a demo; a real analysis needs the user's file, and --analysis selects modules other than the default polymorphism summary.

Chaining Partners

  • Alignment upstream. Sequences must be aligned before this skill. phylogenetics-builder aligns with MAFFT, MUSCLE and other aligners. Versions of that skill that save their alignment write alignment/aligned.fasta; use that untrimmed file rather than the trimAl output, because trimming removes gapped or poorly aligned columns and so changes S, pi and the InDel results. Otherwise align the sequences separately and pass the aligned FASTA.
  • fastreer. Builds distance trees from the same multi-sample VCF: run this skill for diversity and neutrality statistics, then fastreeR for sample relationships.
  • claw-ancestry-pca. Gives population-structure context before samples are grouped in a --pop-file for Fst or divergence.
  • equity-scorer. Reports HEIM heterozygosity and Fst from VCF or ancestry data with its own estimators; this skill's Hudson Fst and diversity statistics follow DnaSP 6, so the two sets of values are not interchangeable.
  • VCF preparation. A filtering or phasing workflow may prepare input, but every filter and phasing choice must be recorded; unphased heterozygotes stay unresolved here.
  • Reporting downstream. Use report.md, summary.json, result.json and the TSV files; the TSV does not contain every module's results.

Maintenance

Recheck source/help-derived regression cases and the 170 historical comparison fixtures after changes to formulas, masks or parsers. Review GUI differences when the target DnaSP build or analysis mode changes. Keep this file, the method reference, CLI metadata, version and catalogue consistent. New Windows GUI observations must be reviewed before changing published concordance counts.

Signals

GitHub stars
1k
Forks
277
Last commit
Sep 2026
Advanced
Item type
skill
Key
dnasp
Source
github.com/clawbio/clawbio