mdvs — Markdown Validation & Search
SkillSearchUse `mdvs search` for any content lookup in a markdown directory, semantic / hybrid / SQL-filtered, beats Grep / Glob for finding by meaning. `mdvs init` infers a schema from existing markdown, `mdvs check` validates frontmatter, `mdvs update` evolves the schema as the KB grows. Activate whenever the project contains markdown with frontmatter, whether or not `mdvs.toml` exists yet.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the mdvs — Markdown Validation & Search skill
What this skill tells your AI
The instructions your AI receives, as published by edochi/mdvs in crates/mdvs/scaffolding/skill/SKILL.md and read by ahel’s review.
A CLI that treats a markdown directory as a database: schema inference, typed frontmatter validation, and semantic / full-text / hybrid search with SQL filters. Single binary, no external services. Full documentation at https://edochi.github.io/mdvs/.
Usage
mdvs init [path] # infer schema, write mdvs.toml
mdvs init [path] --dry-run # preview inference
mdvs init [path] --force # overwrite existing config
mdvs init [path] --from-jsonschema <file> # import schema from JSON Schema 2020-12
mdvs init [path] --ignore-bare-files # exclude files with no frontmatter
mdvs check [path] # validate frontmatter against mdvs.toml
mdvs check [path] --jsonschema <file> # override [fields] for this run
mdvs check [path] --no-update # skip auto-update before validating
mdvs update [path] # detect and add new fields
mdvs update reinfer <field> [path] # re-infer type and constraints
mdvs update reinfer <field> [path] --dry-run # preview reinfer
mdvs update reinfer <field> [path] --with=<categorical|range|none>
mdvs build [path] # validate + chunk + embed → Lance index
mdvs build [path] --force # full rebuild (ignore incremental cache)
mdvs search "<query>" [path] # hybrid search (default)
mdvs search "<query>" [path] --mode <semantic|fulltext|hybrid>
mdvs search "<query>" [path] --where "<SQL>" # frontmatter filter
mdvs search "<query>" [path] --limit <N> # default 10
mdvs search "<query>" [path] -v # show matching chunk text
mdvs search "<query>" [path] --no-build # fail if no index exists
mdvs search "<query>" [path] --no-update # skip auto-update before build
mdvs info [path] # config + index status
mdvs info [path] -v # full field detail
mdvs clean [path] # delete .mdvs/
mdvs export-jsonschema [path] # emit canonical JSON Schema
mdvs export-jsonschema [path] --format <json|toml>
mdvs export-jsonschema [path] --output-file <file>
mdvs scaffold skill [--platform <name>] # this skill file
mdvs scaffold snippet [--platform <name>] # AGENTS.md / CLAUDE.md snippet
mdvs scaffold hook --platform <name> # PostToolUse hook config (JSON snippet)
All commands take --output <pretty|markdown|json>. For agent context, pass
--output markdown explicitly or set default_output_format = "markdown" in
mdvs.toml. <path> defaults to . for every command.
What mdvs is for
mdvs treats a markdown directory as a database: it infers a typed schema from
frontmatter, validates that schema on every edit, and searches the content
semantically. The schema (mdvs.toml) is the source of truth; everything else
(search index, validation reports) is derived. The schema is meant to evolve
with the KB as conventions emerge — mdvs makes deviations visible without
freezing them.
What you must do when invoked
Step 1 — Detect mdvs context
Look for mdvs.toml in the current working directory or any ancestor.
- If found, that directory (or its parent containing the file) is an mdvs vault. Treat any markdown file under it as living under that schema.
- If not found but the project has markdown files with frontmatter, mdvs is
still relevant — propose bootstrapping with
mdvs init <path>before reaching for Grep on the markdown.
Step 2 — For any content lookup, use mdvs search first
When the user (or your own task) needs to find something in markdown content — a
note about a topic, a project status, a person, an experiment, anything —
default to mdvs search "<query>", not Grep / Glob. mdvs is built exactly
for this:
--mode hybrid(default) — semantic + BM25 reranked; best general-purpose mode--mode semantic— vector only; best for "find notes about X" where X is a concept, not a phrase--mode fulltext— BM25 only; for known literal phrases--where "field = 'value'"— filter by frontmatter (SQL syntax)--limit N— cap result count (default 10)-v— show the best matching chunk text per result
Only fall back to Grep when (a) you need a literal substring match and already
know which file to look in, or (b) mdvs search returns no results and you
suspect missing content rather than missing relevance.
Step 3 — For frontmatter changes, check then evolve
After any edit that touches frontmatter (yours or the user's):
- Run
mdvs checkto validate against the schema. - If violations appear: see Step 4 (the schema-evolution loop).
- If new fields appear (present in files but not in
mdvs.toml): they show up as informational, not violations. Runmdvs updateto add them, ormdvs update reinfer <field>to refresh a specific field's constraints.
Step 4 — When a hook surfaces a violation, follow the schema-evolution loop
If a markdown block lands in your context from a PostToolUse hook listing
MissingRequired / WrongType / Disallowed / InvalidCategory /
OutOfRange violations, that's mdvs talking to you via the validation hook.
Claude Code surfaces it under additionalContext; other harnesses use their own
channel (the wiring is harness-specific). The hook is non-blocking by design:
the edit already landed; the warning is for you to act on next.
Your job:
- Read the violation block: which file, which field, which rule, expected vs actual.
- Decide: mistake (typo, wrong type by accident, dropped required field) or intentional (KB is evolving, category needs a new variant, field shifting type)?
- Mistake → fix the file in the next turn. Acknowledge briefly so the user knows the loop is working.
- Intentional → surface the deviation to the user and propose updating
mdvs.toml(mdvs update,mdvs update reinfer <field> --with=<kind>, or a manual edit). Do not silently fix the file. The user decides whether the schema or the file is the source of truth.
A worked example of the intentional path is in the Examples section.
Rules
- Prefer
mdvs searchover Grep / Glob in any markdown corpus. Inside an mdvs vault, the default content-search tool is mdvs. - The validation hook is a warning, not a block. Treat violations as a prompt for discussion, not as edits to be reverted.
- Never silently fix an intentional deviation. Propose a schema update; let the user decide.
- The schema is meant to evolve. A schema that never changes is a schema that's wrong. Enforcement follows the KB's shape; it does not freeze it.
buildalways runscheckfirst. Validation gates the index build.mdvs.tomlis the only source of truth..mdvs/is derived — gitignored, recreatable withmdvs build.- There is no lock file. Schema changes flow through
mdvs.tomlonly. - Frontmatter formats are auto-detected per file (YAML / TOML / JSON). A single vault can mix all three.
- Use
--output markdownfor any output you intend to read. It's the format LLMs parse most fluently, and the format the validation hook surfaces back through the harness's model-context channel.
Two layers
mdvs has two independent layers:
- Validation (
init,check,update) — works immediately, no model download, no build step. Reads markdown and validates frontmatter againstmdvs.toml. - Search (
build,search) — downloads an embedding model, chunks markdown content, builds a local LanceDB index in.mdvs/.
Validation stands alone. You never need to build an index just to validate.
Key files
mdvs.toml— schema config, committed to version control. Source of truth for field types, allowed/required paths, constraints..mdvs/— build artifacts (the Lance dataset underindex.lance/plus a cached model). To be gitignored. Recreatable withmdvs build. Never edit directly.
Command reference
mdvs init
Scans markdown files, infers a typed schema from frontmatter, writes
mdvs.toml.
--force— overwrite an existingmdvs.toml(deletes.mdvs/too)--dry-run— show what would be inferred without writing--ignore-bare-files— exclude files that have no frontmatter--from-jsonschema PATH— import schema from an external JSON Schema 2020-12 document. Round-trips withmdvs export-jsonschema.
Use init --force to start over. Use update to incrementally add new fields.
mdvs check
Validates all frontmatter against mdvs.toml. Reports violation kinds:
MissingRequired— required field is absent from a fileWrongType— value doesn't match declared typeDisallowed— field appears in a path not covered by itsallowedglobsInvalidCategory— value is not in the declared category listOutOfRange— numeric value outside declaredmin/max
New fields (in files but not in mdvs.toml) are reported separately as
informational — no non-zero exit. Run update to add them.
--jsonschema PATH— override[fields]inmdvs.tomlfor this run
Violation output is deterministic: sorted by (field, kind, rule) and path
within.
mdvs update
Re-scans files; adds newly discovered fields to mdvs.toml. Doesn't remove or
change existing fields by default.
mdvs update— detect and add new fieldsmdvs update reinfer <field>— re-infer type and constraintsmdvs update reinfer <field> --dry-run— previewmdvs update reinfer <field> --with=categorical— force categoricalmdvs update reinfer <field> --with=range— infer min/maxmdvs update reinfer <field> --with=none— strip all constraints
Use reinfer when a field's type has changed or you want to refresh its
constraints. --with requires a named field.
mdvs build
Validates, then chunks markdown, generates embeddings, writes the Lance dataset
to .mdvs/.
--force— full rebuild (ignore incremental cache)- Incremental by default — only re-embeds new or edited files
- Aborts if
checkfinds violations
First build downloads the default embedding model
minishlab/potion-multilingual-128M (~480 MB, 101 languages). Subsequent builds
reuse it.
mdvs search
Searches the indexed notes — semantic (vector), full-text (BM25), or hybrid (RRF reranker). Auto-builds the index if needed.
mdvs search "<query>" [path] [--mode <m>] [--where "<SQL>"] [--limit N] [-v]
--where operates on any column in the Lance index: frontmatter fields
(auto-discovered from mdvs.toml, referenced by bare name) and the
always-present filepath column. Field names with spaces need double-quote
escaping: --where "\"lab section\" = 'Photonics'". Filtering on Array(Float)
fields is rejected up front (Lance can't safely decode them); store as parallel
scalar arrays.
Scalar frontmatter — equality, inequality, comparison
--where "author = 'Federica Bianchi'" # string equality
--where "year >= 2020" # numeric comparison
--where "year != 2025" # inequality
--where "year BETWEEN 2018 AND 2024" # closed range
--where "year IN (2021, 2022, 2023)" # discrete set
--where "year NOT IN (2020, 2024)" # exclusion
Strings — pattern matching with LIKE
% matches any sequence, _ matches one character.
--where "title LIKE 'Async%'" # starts with "Async"
--where "title LIKE '%network%'" # contains "network"
--where "title NOT LIKE '%Tutorial%'" # excludes the word
--where "lower(title) LIKE '%rust%'" # case-insensitive (via the lower() function)
Null filters
--where "author IS NOT NULL" # field must be set
--where "url IS NULL" # field must be absent
Array fields — auto-rewritten to array_has(...)
The = / != / IN / NOT IN operators against an array field auto-rewrite
to array_has(...) so element-containment "just works". The search output shows
the rewrite as a one-line Note at the top.
--where "tags = 'rust'" # has 'rust' as one of its tags
--where "tags != 'archived'" # does NOT have 'archived' as a tag
--where "tags IN ('rust', 'python', 'go')" # has at least one of these
--where "tags NOT IN ('archived', 'draft')" # has NONE of these
--where "tags = 'rust' AND tags = 'async'" # has BOTH (two array_has, AND'd)
--where "tags = 'rust' OR tags = 'python'" # equivalent to IN, longer form
--where "array_has(tags, 'rust')" # the explicit form — bypasses the rewrite
Date fields
Date literals use the date '...' keyword form. RFC 3339 datetimes use
timestamp '...'.
--where "published > date '2024-01-01'"
--where "created BETWEEN date '2024-01-01' AND date '2024-12-31'"
--where "published >= date '2024-01-01' AND published < date '2025-01-01'"
Path filtering — the always-present filepath column
The filepath column stores the path relative to the project root — the last
component is the filename (e.g. articles/long-essay-2024.md → filename is
long-essay-2024.md).
--where "filepath LIKE 'articles/%'" # everything under articles/
--where "filepath LIKE '%/notes/%'" # any directory called notes/
--where "filepath LIKE '%-postmortem.md'" # filename suffix
--where "filepath = 'articles/foo.md'" # exact path
Combining filters — AND, OR, NOT, parentheses
--where "year > 2020 AND tags = 'rust'"
--where "year > 2020 AND (tags = 'rust' OR tags = 'python')"
--where "(author = 'Federica Bianchi' OR author = 'Lorenzo Conti') AND year > 2020"
--where "tags = 'rust' AND filepath LIKE 'technologies/%'" # array + path
--where "concepts != 'CRDT' AND filepath LIKE 'technologies/%'"
Functions
Most DataFusion scalar functions work — handy when the raw value doesn't quite match.
--where "length(title) > 50" # long titles
--where "lower(author) LIKE '%bianchi%'" # case-insensitive contains
--where "year + 1 > 2025" # arithmetic
Other internal columns exist (start_line, end_line, built_at,
chunk_text) but are rarely useful for filtering — semantic / fulltext search
handles those concerns better.
mdvs info
Shows current config and index status: scan settings, field definitions, build
metadata (model, chunk size, file counts). -v for full field detail.
mdvs clean
Deletes .mdvs/. Doesn't touch mdvs.toml.
mdvs export-jsonschema
Translates [fields] into a canonical JSON Schema 2020-12 document. Useful for
sharing with other tools, or round-tripping through
mdvs init --from-jsonschema.
--format json|toml— output format (defaultjson)--output-file FILE— write to file instead of stdout
mdvs scaffold
Emits the artifacts that integrate mdvs with an agent harness (Claude Code, Codex, OpenCode, Cursor, Antigravity). Each subcommand prints to stdout; pipe it into the right location for your harness.
mdvs scaffold skill [--platform <name>]— this skill file. Default destination:.agents/skills/mdvs/SKILL.mdfor Codex / OpenCode / Cursor / Antigravity;.claude/skills/mdvs/SKILL.mdfor Claude Code.mdvs scaffold snippet [--platform <name>]— the project-rules snippet forAGENTS.md/CLAUDE.md/.cursor/rules/mdvs.mdc.mdvs scaffold hook --platform claude-code— thePostToolUsehook config (a JSON snippet to merge into.claude/settings.json). The emitted snippet'scommand:fields callmdvs hook handledirectly — no shell scripts, nojqdependency. Currently the only verified hook integration. The other platforms refuse with a pointer to their per-platform mdbook page; wiringmdvs hook handleinto Codex / Cursor / OpenCode / Antigravity is possible by following each harness's own hooks documentation.
Agent-harness integration
mdvs ships as a CLI; integrating it with an agent harness is wiring rather than installing. Three artifacts cover the three integration points:
| Artifact | Purpose | Coverage |
|---|---|---|
Skill file (SKILL.md, this file) | Activated by harnesses implementing the Agent Skills open standard. Loaded on demand; agent reads procedure + reference. | Works in any harness that loads .md skills (Claude Code, Codex, Cursor, OpenCode, Antigravity, …). |
| Project-rules snippet | Always-on text in AGENTS.md / CLAUDE.md / .cursor/rules/mdvs.mdc. Short — names the KB, the search-vs-Grep preference, and the warning-loop rule. | Works in any harness that reads AGENTS.md / CLAUDE.md / .cursor/rules. |
PostToolUse hook | Calls mdvs hook handle after every Edit / Write on a markdown file inside the vault. mdvs walks up to find mdvs.toml, runs check, and surfaces violations to you as non-blocking model-context. | Shipped for Claude Code only. Other harnesses: follow their documented hook system to wire mdvs hook handle in — see https://edochi.github.io/mdvs/recipes/agent-harnesses/. |
Output format
Three formats; default is pretty.
--output pretty— box-drawing tables for terminal display. Adapts to width.--output markdown— GFM tables and##headers. Best for agent consumption and the format the validation hook surfaces back through the harness's model-context channel.--output json— structured JSON for programmatic extraction (pipe throughjqor any JSON tool).
Priority chain when --output is omitted: CLI flag > default_output_format in
mdvs.toml > hard default (pretty). Same command always produces the same
output regardless of TTY state. -v (verbose) adds per-step pipeline output
with timings.
Exit codes
- 0 — success (no violations)
- 1 — violations found (
checkandbuild) - 2 — error (bad config, missing files, model mismatch)
Hook scripts use || true to mask the exit-1 from check — the hook is
intentionally non-blocking. Don't change that contract.
Things to know
- Field types inferred automatically:
String,Integer,Float,Boolean,Date(YYYY-MM-DD),DateTime(RFC 3339 with mandatory timezone), andArray(<scalar>)for any scalar type. The on-disk grammar isScalar | Array(Scalar)only. Mixed scalar types widen (Integer + String → String). - Nested frontmatter uses dotted-name leaves.
calibration.baseline.wavelength: 850.0becomes a[[fields.field]]named"calibration.baseline.wavelength"of typeFloat. Top-levelObjectis rejected; nestedArray(Object{...})is also rejected — represent arrays of structured items as parallel scalar arrays (e.g.measurement_timestamps: Array(String)+measurement_values: Array(Float)). SQL filters use dot notation:--where "calibration.baseline.wavelength > 800". - Preprocessors opt into widening. Each field carries a
preprocessarray. Built-ins:coerce_to_string(accepts non-string scalars on aStringfield) andwiden_int_to_float(accepts integers on aFloatfield). Inference auto-populates these when widening was observed.preprocess = []means strict. - Categorical detection is automatic for low-cardinality repeated values.
Out-of-category values →
InvalidCategory. - Constraint kinds:
categories(closed-set enum, mutually exclusive with everything else),min/max(numeric range),min_length/max_length(string and array length),pattern(regex on strings). Range / length / pattern are not auto-inferred; add manually or viaupdate reinfer <field> --with=range. init --forcerewrites the config from scratch.updatepreserves existing config and only adds new fields.update reinferre-infers specific fields.- Model identity is tracked: changing the model in
mdvs.tomlrequiresbuild --forceto confirm a full re-embed. checkauto-runsupdatefirst by default (unless--no-update, or--jsonschemais given, or[check].auto_update = falseis set inmdvs.toml).searchauto-runsupdateandbuildif needed (unless--no-update/--no-build)..mdvsignoreand.gitignoreare both honored when scanning. mdvs reuses theignorecrate (same matcher git uses), so the per-directory.gitignorerules you already have apply automatically. For mdvs-only exclusions (e.g. "don't index this draft directory, but keep it tracked in git"), drop a.mdvsignorenext to the files using the same syntax as.gitignore. Both apply toinit/update/check/build/search. To stop honoring.gitignore(for example, to index files that are deliberately untracked but should still be searchable), set[scan].skip_gitignore = trueinmdvs.toml. Hidden files and dotfiles are NOT skipped — only.md/.markdownextensions are scanned to begin with.
Examples
Setting up a new vault from scratch
mdvs init # scan current directory, infer schema, write mdvs.toml
mdvs init ~/notes # or point at another directory
mdvs init --dry-run # preview without writing
User added a new field to some files
mdvs check # → reports "category" as a new field (informational)
mdvs update # → detects "category", adds it to mdvs.toml
mdvs check # → clean
Fixing violations after check
mdvs check
# →
# MissingRequired: "title" missing in blog/drafts/untitled.md
# WrongType: "priority" expected Integer, got String in projects/alpha.md
# InvalidCategory: "status" got "wip", expected one of [draft, published, archived]
# OutOfRange: "rating" got 11, expected min=1, max=5
Resolution per kind:
Shortened here. Read the whole file on GitHub.
Signals
- GitHub stars
- 27
- Forks
- 2
- Last commit
- Aug 2026
ahel recommends instead
Advanced
- Catalog kind
- skill
- Gateway key
mdvs-edochi- Source
- github.com/edochi/mdvs