Training Phase B1 — Dataset Studio

SkillDatabases & data

Covert Coder, a local-first sovereign AI development environment for governed models, workflows, execution, evidence, and verification.

Use Training Phase B1 — Dataset Studio in Claude, ChatGPT or Ahel Desktop

Free. Sign in, add Training Phase B1 — Dataset Studio and connect your AI. About a minute.

Also: Claude Code · Cursor · Codex

Then ask your AI: use the Training Phase B1 skill

Details

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Training Phase B1 — Dataset StudioStart free

What this skill tells your AI

The instructions your AI receives, as published by anonymousnomad/covert-coder in skills/packs/aide-training-phase-b1-dataset-studio/SKILL.md and read by Ahel’s review.

What

A local dataset manager for fine-tuning: import/export ChatML-format JSONL, dedup (exact-hash then normalized), length filtering (tokenized 95th-percentile guidance), chat-template render preview (5-sample eyeball rule), and a locked train/validation split that cannot be silently mutated after eval artifacts exist.

Why

2026 consensus across all researched guides: dataset quality beats hyperparameter tuning; silent format failures and template drift are the top run-killers; contamination between train and eval invalidates every claim a run makes. The 5-sample render preview catches malformed roles before a multi-hour job, not after.

Code Plan

  • daemon/dataset-store.mjs (or node/src service): datasets live under <workspace>/.aide/datasets/<id>/ — data.jsonl, meta.json {schema_version, format:'chatml', counts, sha256 of each file, split_seed, split_ratios, locked:boolean}.
  • Operations: import(jsonlPath) validates every line against ChatML message schema (roles, non-empty content); dedup by exact sha256 of rendered text, then lowercase/whitespace-normalized second pass; length stats via tokenizer chars/4 heuristic (consistent with existing estimateTokens); split(seed) writes train.jsonl/val.jsonl + hashes; preview(n=5) renders through the target model's real GGUF chat template (reuse probeGguf template extraction) so what you see is what trains.
  • Lock rule: once any eval artifact references the split hashes, meta.locked=true → mutation rejected with explicit error. Unlock requires deleting dependent artifacts (auditable).
  • Routes: /api/datasets/* CRUD + preview + split; UI view reusing hub-style list/detail patterns. Contracts regen.

Dependencies

gguf.ts (template extraction), estimateTokens, file containment pattern, contracts/OpenAPI drift gate. Doctrine: anti-trash-data rules (dedup, validation, no unverified docs) apply to user data too.

Threat Matrix

ThreatControl
Train/eval contaminationsplit written once with recorded seed+hashes; lock enforced on dependent artifacts
Silent format failureimport-time per-line schema validation + mandatory template preview before a job can reference the dataset
Oversized/path-traversal importssize cap (e.g., 200MB), basename sanitization, workspace containment
Duplicate-heavy data inflating metricsexact + normalized dedup pass with reported removal counts
Privacylocal-only storage; export is explicit user action

Issues / Bugs Watchlist

  • Very long single lines: stream-parse JSONL, don't readFile whole file for validation.
  • Template mismatch warnings must name BOTH formats when base model template ≠ dataset rendering assumption.

Signals

GitHub stars
43
Forks
14
Last commit
Oct 2026
Advanced
Item type
skill
Key
aide-training-phase-b1-dataset-studio
Source
github.com/anonymousnomad/covert-coder