user-corpus-explorer — 用户资料库索引

SkillDatabases & data

Use near the end of EZ_math_model intake when external/user-corpus contains user-provided papers, notes, PDFs, datasets, or examples that should become a local AGENTS.md reference index.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the user-corpus-explorer — 用户资料库索引 skill

What this skill tells your AI

The instructions your AI receives, as published by woodfishhhh/ez_math_model in skills/ez-math-model/tools/user-corpus-explorer/SKILL.md and read by ahel’s review.

何时使用

  • external/user-corpus/ 存在用户自带参考材料。
  • pipeline 00 对该域决策不是 skip
  • pipeline 01 intake 结束前需要生成用户资料索引。

设计原则

  • 由 subagent 执行,不污染主对话上下文。
  • 每次覆盖生成 external/user-corpus/AGENTS.md
  • 失败不打断主 pipeline。
  • 不上传全文到外部服务。

扫描范围

递归扫描 external/user-corpus/,跳过:

  • .gitkeep
  • README.md
  • AGENTS.md
  • .corpus_index.json
  • .git/.cache/、以 . 开头的目录
  • 大于 200MB 的文件只记录路径和大小

读取策略

扩展名策略
.md .txt读全文,超长读首尾
.pdfMinerU → pdf → pdfplumber,长文仅读首 15 页和末 5 页
.docxdocx 提取文本
.htmlJina Reader 或 BeautifulSoup
图片视觉描述
其他二进制只记录元信息

输出

生成:

external/user-corpus/AGENTS.md
external/user-corpus/.corpus_index.json

AGENTS.md 包含 inventory、per-file index、cross-cutting topics、recommendations、limitations。

下游衔接

  • modeler 必读 recommendations,并在 modeling_plan.md 标注参考来源。
  • writer 可优先把 corpus 中可验证 DOI 的论文列入参考候选。
  • coder 不直接读 corpus,除非 modeler 在计划中转述其方法。

失败诊断

情况处理
corpus 目录为空写空索引,pipeline 继续
单文件读取失败在 limitations 记录
MinerU 不可用降级 pdf → pdfplumber
总耗时超过 5 分钟写已完成索引,剩余标 unread

Signals

GitHub stars
41
Forks
1
Last commit
Jul 2026
Advanced
Catalog kind
skill
Gateway key
user-corpus-explorer
Source
github.com/woodfishhhh/ez_math_model