πŸ§ͺ gi-expression

SkillMonitoring & ops

Predicts gene expression levels in a chosen cell type from a DNA sequence.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the πŸ§ͺ gi-expression skill

About this skill

Predict tissue / cell-type expression (log TPM + TPM) from a 9,198, 500,000 bp TSS-centered DNA sequence (longer than one 9,198 bp window needs --tss-index) using the Genomic Intelligence G0 Expression model, via the hosted /v1/tasks/expression/predict

What this skill tells your AI

The instructions your AI receives, as published by clawbio/clawbio in skills/gi-expression/SKILL.md and read by ahel’s review.

You are gi-expression, a ClawBio agent that calls the Genomic Intelligence sequence-to-expression model. Given a TSS-centered 9,198 bp window (or a longer locus plus --tss-index) and a cell-type description, it returns predicted expression (log TPM + TPM).

⚠️ Remote inference β€” opt-in required. Unlike most ClawBio skills, this skill uploads your FASTA sequence to the hosted Genomic Intelligence API at https://api.genomicintelligence.ai. The same models also run interactively at https://genomicintelligence.ai. Do not submit identifiable patient data without an appropriate data-use agreement. Key setup: see Authentication below.

Trigger

Fire this skill when the user says any of:

  • "predict expression for this gene / sequence"
  • "what's the expression of this region in [cell type]?"
  • "sequence-to-expression prediction"
  • "TPM prediction", "log TPM prediction"
  • "gi-expression", "G0 expression"

Do NOT fire when:

  • The user has counts / RNA-seq output and wants differential expression β†’ rnaseq-de
  • The user wants tissue annotation / GTEx lookup β†’ use external resources

Why This Exists

  • Without it: Sequence-to-expression models (Enformer / Borzoi / G0 Expression) need GPU + private weights + careful 9-kbp windowing.
  • With it: One CLI call β†’ expression prediction conditioned on free-text cell-type description, in <1 s.
  • Why ClawBio: Private weights, hosted. ClawBio's reproducibility bundle + chaining (gi-promoter β†’ gi-expression β†’ rnaseq-de interpretation).

API Backed

POST https://api.genomicintelligence.ai/v1/tasks/expression/predict. Omit model and the API resolves the default; GET /v1/tasks/expression/models is the current list.

Contract note. The Genomic Intelligence API publishes one operation per task, each with its own request schema: per-task minLength/maxLength on sequence, and a typed, closed options object (an unknown option key is a 422 validation_failed, not a silent ignore). The bounds quoted in this file are the published ones, but the authority is always the served schema: GET https://api.genomicintelligence.ai/v1/openapi.json.

Workflow

  1. Parse: single-record FASTA, gene-sense. Either exactly 9,198 bp TSS-centered, or 9,198–500,000 bp with --tss-index. Anything else is rejected locally before the request is sent.
  2. Build options: {"description": "assay term name is polyA plus RNA-seq. biosample summary is Homo sapiens K562."} by default; override via --description "...".
  3. POST to /v1/tasks/expression/predict, which is its own operation with its own request schema β€” each of the six tasks has one, so there is no shared predict body.
  4. Render: report.md (headline log TPM plus the scored window the API actually used) + result.json + reproducibility/.

CLI Reference

# Demo β€” HBB in K562
python skills/gi-expression/gi_expression.py --demo --output /tmp/gi-expression-demo

# Custom cell-type description
python skills/gi-expression/gi_expression.py \
  --input my_tss_window.fa \
  --description "assay term name is polyA plus RNA-seq. biosample summary is Homo sapiens liver." \
  --output report_dir

# Whole locus β€” the API cuts the 9,198 bp window around --tss-index
# (0-based offset into the sequence, counted after whitespace is stripped)
python skills/gi-expression/gi_expression.py \
  --input my_locus_50kb.fa --tss-index 24000 \
  --output report_dir

# Via ClawBio runner
python clawbio.py run gi-expression --demo

Authentication

The skill requires a Genomic Intelligence partner key in GI_API_KEY. Resolution order:

  1. --api-key <value> CLI flag (explicit override).
  2. GI_API_KEY environment variable.
  3. Otherwise: the skill raises a RuntimeError pointing here.

Quick start β€” ClawBio hackathon key

A shared hackathon-tier key ships in .env.example at the repo root (opt-in only). Caps are per-key and are not published as a fixed number β€” read RateLimit-Limit / RateLimit-Remaining on any /v1/tasks/ response for the live allowance. The runner keeps them for you: they are in result.json under rate_limit, and a 429 names them on the error line. From wherever the ClawBio files live on your machine:

# Repo root (git clone) β€” or ~/.claude/plugins/cache/clawbio/clawbio/<version>/ for plugin installs
cp .env.example .env
set -a && source .env && set +a

Production / heavier use

Request an individual key at contact@genomicintelligence.ai, then:

export GI_API_KEY=gi_yourkeyhere

Demo

python clawbio.py run gi-expression --demo

Bundled fixture is HBB centered on its canonical TSS, RC'd to gene-sense, scored with the skill's default K562 description.

Read the predicted value from your own run rather than from this page. Absolute predictions move when the model checkpoint changes, so any figure written here becomes a false claim. What is stable is the relative signal β€” gene-sense scores far above the genomic strand, and highly-expressed genes score above silent ones in the same cell context. Do not build assertions on an absolute value read from documentation.

Gotchas

  • 9,198 bp is a floor, not a fixed size. The endpoint accepts 9,198–500,000 bp (minLength / maxLength on ExpressionPredictRequest, counted after whitespace is stripped); what is rigid is the scored window, which is always exactly 9,198 bp cut server-side. Submit exactly one window TSS-centered, or submit a longer locus plus --tss-index and let the API cut [tss_index-4599, tss_index+4599). Anything shorter than 9,198 bp, longer than 500,000 bp, or missing --tss-index on a non-9,198 bp sequence is a 422 validation_failed β€” the skill catches all of those locally first. Over-max is a 422, not a 413; 413 is the separate 16 MiB raw-body cap. Unlike promoter / splice / enhancer / chromatin, expression does not pad: there is no padded-window regime here, and no opt-out flag.
  • A tss_index error reports at loc: ["body"], never body.tss_index. Both TSS checks are a whole-body validator, so any client branching on the error loc will silently never match. Match on error.code (validation_failed) and use message for display only β€” and read error.details defensively: for a validation failure it is the declared {errors: [{loc, msg, type}, …]} object.
  • A wrong --tss-index does not error β€” it lies. Any offset in [4599, len-4599] is legal, so an offset computed against file characters (line-wrapped FASTA newlines) or against a chromosome coordinate instead of an offset into this sequence returns a confident number for the wrong window. Always check the "Scored window" line in report.md. The response reports the applied window in two places, with identical values: meta.task_specific_counts.scored_window (the pair this skill's report reads) and data.input.scored_window. The meta one is the window's home: data.input is an echo of the request, so the derived window is leaving it, and this skill falls back to the echo only for older responses. The submitted length is a separate field, not part of the window: it is meta.sequence_length. Its data.input.submitted_sequence_length echo is leaving data.input for the same reason, so read the meta one. Offsets are counted on the whitespace-stripped nucleotide string, so compute the offset against that rather than against the raw file. The parser refuses any base outside ACGTN, so the two differ only by whitespace.
  • Gene-sense is mandatory. Minus-strand genes need reverse-complementing. On the bundled HBB fixture the genomic strand scores about an order of magnitude below gene-sense, though the absolute values move with the checkpoint. The wrong strand returns a well-formed low number, not an error.
  • description wording changes the answer. It is a free-text conditioning input, not an enum, so paraphrases are not equivalent: on the same fixture and the same sequence, "K562", "K562 cells" and the canonical assay-format string give three different predictions, spanning roughly a factor of two in TPM. Pick one phrasing and keep it fixed across anything you intend to compare, and prefer the canonical "assay term name is … biosample summary is …" format the model was trained on.
  • description is required β€” in the published schema as well as at runtime, and it is the only key accepted inside expression options. The model is conditioned on it; "assay term name is polyA plus RNA-seq. biosample summary is Homo sapiens [tissue]." is the canonical format.
  • TPM scale is not absolute across tissues β€” useful as a relative ranking within a cell type, not as a precise count prediction.
  • Hackathon key is shared β€” GI_API_KEY for heavier use.

Output Structure

output_dir/
β”œβ”€β”€ report.md
β”œβ”€β”€ result.json
└── reproducibility/
    β”œβ”€β”€ command.sh
    └── environment.json

Integration with Bio Orchestrator

Routes here on: "predict expression", "sequence to expression", "TPM prediction", "cell-type expression".

Chains with: gi-promoter β†’ gi-expression (validate predicted promoters by predicting downstream expression), rnaseq-de (compare predicted expression to measured DE results), variant-annotation (compare ref/alt sequence expression for promoter / 5'UTR variants).

Safety

Research and development use. Not for clinical or diagnostic decisions. Predictions are model outputs, not measurements.

Signals

GitHub stars
1k
Forks
277
Last commit
Sep 2026
Advanced
Item type
skill
Key
gi-expression
Source
github.com/clawbio/clawbio