dataset — 公开数据集发现

SkillDatabases & data

Entry point for discovering public datasets. Covers Kaggle / UCI / HuggingFace / Tianchi. Invoke when the task requires "finding data on your own" or "supplementing with external data"; do **not** call it when the attachments already contain data.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the dataset — 公开数据集发现 skill

What this skill tells your AI

The instructions your AI receives, as published by woodfishhhh/ez_math_model in skills/ez-math-model/tools/dataset/SKILL.md and read by ahel’s review.

何时使用

  • 题目附件没有数据集,但题面要求"查找类似公开数据"。
  • 需要历史基准数据集(如 MNIST / Iris / 波士顿房价)做模型对比。
  • 不要在已有附件数据时启用本子 skill。

入口

数据源入口配置
Kagglekaggle datasets list -s <kw> / kaggle datasets download -d <user/name>~/.kaggle/kaggle.json(注册免费下载)
UCI ML直接 HTTPS pd.read_csv(url)
HuggingFacefrom datasets import load_datasetEZMM_HF_TOKEN(私有数据集)
天池浏览器 + 手动下载

命令模板

# Kaggle 搜索 + 下载
kaggle datasets list -s "vegetable retail price"
kaggle datasets download -d <user/dataset-name> -p workdir/.../attachments/external/kaggle --unzip

# HuggingFace
python -c "from datasets import load_dataset; ds = load_dataset('squad', split='train[:1%]')"

# UCI(直接 URL)
python -c "import pandas as pd; df = pd.read_csv('https://archive.ics.uci.edu/...'); df.to_csv('workdir/.../attachments/external/uci/iris.csv', index=False)"

落盘规范

外部下载的数据放在:

workdir/{task_id}/attachments/external/<source>/<dataset>/

并在同目录写 SOURCES.md

- 数据集名: <name>
- 来源: <kaggle url>
- License: <CC0 / CC-BY / Apache-2.0 / 其他>
- 下载时间: <ISO timestamp>
- 用途说明: <一句话>

失败诊断

情况处理
Kaggle token 未配置提示用户 ~/.kaggle/kaggle.json 配置
License 不允许商用数据可用于学术建模报告;论文中标明出处 + license
文件 > 1GBcoder 阶段用 chunksize 处理(参考 prompts/coder.md
网络受限写诊断;建议用户手动下载放入 attachments/

Signals

GitHub stars
41
Forks
1
Last commit
Jul 2026
Advanced
Catalog kind
skill
Gateway key
dataset-woodfishhhh
Source
github.com/woodfishhhh/ez_math_model