Dev Research Scout
SkillAI & modelsMines academic papers, research blogs, and curator newsletters for stealable methods and frameworks. Use when scanning research for applicable techniques across AI/ML/SWE.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Dev Research Scout skill
What this skill tells your AI
The instructions your AI receives, as published by vasilyu1983/ai-agents-public in frameworks/shared-skills/skills/research-scout/SKILL.md and read by ahel’s review.
Scans high-signal research sources for methods, frameworks, and ideas worth applying to your own work, and converts the top finds into idea cards with how-to-apply recipes, evidence quality grades, and reproducibility notes.
Supported sources: arXiv, Hugging Face Papers, Semantic Scholar, Papers with Code (archive only — shut down Jul 2025), conference proceedings (NeurIPS / ICML / ICLR / ACL / EMNLP / KDD), industry research blogs (Anthropic / OpenAI / DeepMind / Google Research / Meta AI / Microsoft Research / Apple ML), and curator newsletters (Lilian Weng, Sebastian Raschka, Eugene Yan, Latent Space, Simon Willison, The Batch, Import AI, Interconnects / Nathan Lambert, Davis Summarizes Papers / Davis Blalock).
Output is a generative toolkit, not a landscape report:
- pattern catalog (methods worth stealing, with how-to-apply)
- anti-pattern catalog (research traps — irreproducibility, benchmark gaming, hype)
- recipes (extraction, validation-before-adoption, kill criteria)
Key distinction from sibling scouts:
- This skill = research-grade idea mining (papers + research blogs + curated synthesis)
research-painpoint-scanner= community-pain mining (Reddit / HN / GitHub Issues / G2 / Stack Overflow)research-arxiv-scout= arXiv-only deep triage with category taxonomy and attribution; specialist downstreamresearch-git= public GitHub repo research for skills, practices, and code patterns (separate concern)
Use this skill when the question is "what methods or frameworks are worth stealing from recent research?" — escalate to research-arxiv-scout for arXiv-only work where category taxonomy and attribution matter most.
Quick Reference
| Need | Go to |
|---|---|
| Pick the source mix | ## Source Selection Guide |
| Run the end-to-end scan | ## Workflow |
| Reject hype / irreproducible / benchmark-gamed work | known-traps.md |
| Pattern-match a paper to a known method shape | idea-extraction-framework.md |
| How to actually apply a stolen idea | recipes.md |
| Source-specific query and credibility guidance | ## Navigation |
| Package the idea cards | ## Templates & Assets |
| Mine industry/eng blogs + HCI papers for killer-feature attribution (bundle handoff) | ## Killer-Feature Mode (Feature-Precedent Mining) |
When to Use
Invoke when users ask for:
- "What methods are people using for {{topic}} that I haven't tried?"
- "Find recent {{AI/ML/SWE}} ideas worth stealing for {{project}}"
- "Mine arXiv + research blogs for {{topic}} in the last {{N}} days"
- "What's worth stealing from NeurIPS / ICML / ICLR {{year}}?"
- "Show me frameworks for {{evaluating LLM agents / RAG eval / inference scaling / etc.}}"
- "Update {{skill name}}'s knowledge base with recent research"
When NOT to Use
| Situation | Use instead |
|---|---|
| arXiv-only deep triage with attribution | research-arxiv-scout |
| Community pain points, not research methods | research-painpoint-scanner |
| Mining public GitHub repos for skills, practices, or code patterns | research-git |
| Validated Q&A answers or known-error solutions (the Stack Overflow corpus / Stack Overflow for Agents exchange) | qa-debugging — that is solved-answer lookup, not research-method mining |
| Production deep-research synthesis (verified citations + reasoning trace) | ai-deep-research |
| Single-paper summary for a known arXiv ID | research-arxiv-scout step 3 |
| End-user career positioning, company interview reviews, recruiter pitches, or CV tailoring | career-jobhunt; this skill may still mine research methods to improve that skill |
Source Selection Guide
| Source | Best for | Query method | Idea quality |
|---|---|---|---|
| arXiv | Bleeding-edge methods (preprints, no peer review) | export.arxiv.org/api/query | High volume, mixed signal — needs trap filter |
| Hugging Face Papers | Community-curated daily highlights | huggingface.co/papers + RSS | Pre-filtered, signal-rich, biased to LLM/VLM |
| Semantic Scholar | Citation graphs, prior work, influential papers | Semantic Scholar API | Best for "what built on this?" |
| Papers with Code | DEAD (Meta shutdown Jul 2025) — historical archive only | github.com/paperswithcode/paperswithcode-data (frozen) | None live; reconstruct via HF Papers + GitHub (research-git) — see papers-with-code-strategy.md |
| Conference proceedings | Peer-reviewed, vetted methods | NeurIPS / ICML / ICLR / ACL / EMNLP / KDD sites | Lagged but high-credibility |
| Industry research blogs | Production-tested methods at scale | RSS or direct site (Anthropic / OpenAI / DeepMind / Google / Meta / MSR / Apple) | High signal but PR-tinged |
| Curator newsletters | Pre-synthesized, opinionated, applied | Substack / blog RSS | Highest applicability, reflects curator bias |
Default mix:
- Fast scan (1-2 hr): HF Papers + 1 curator newsletter (Lilian Weng or Eugene Yan) + GitHub repo signal (via
research-git) for the target task - Standard scan (1 day): arXiv + HF Papers + Semantic Scholar + 2 industry blogs + 2 curator newsletters
- Deep scan (multi-day): all live source types (arXiv, HF Papers, Semantic Scholar, conferences, industry blogs, curator newsletters; Papers with Code is dead — archive only), time windows 7d/30d/90d, full trap filter, full extraction recipes
Quick Start
Semantic Scholar API key: New keys are no longer approved for free email domains (gmail, outlook, etc.). Use an institutional email to apply, or fall back to OpenAlex — same free-key-required model but no email-domain restriction; register at openalex.org/settings/api (OpenAlex has required a key for every request since 2026-02-13). See
references/semantic-scholar-strategy.mdfor detail.
Required inputs:
topic— Research topic or method family (e.g., "LLM agent tool use", "RAG eval", "inference batching", "distillation")target— Where the stolen ideas will be applied (e.g., "ai-rag skill", "production RAG service", "agent evals")
Optional inputs:
sources— Which source families to scan (default: arxiv, hf_papers, semantic_scholar, curator_newsletters)windows— Time windows (default: 30d, 90d, 365d)min_evidence_grade— Minimum evidence grade (A/B/C/D/F, defaultC;Fis the floor used by the scoring engine and validator)- Source-specific:
--arxiv-categories,--conference,--blog-domains,--curators
Workflow
ASCII Flow
research idea-mining request
-> Frame topic, target application, source mix, and time windows
-> Search academic, code-linked, conference, blog, and curator sources
-> Normalize findings into the TSV schema
-> Extract stealable methods, evidence, transfer limits, and kill criteria
-> Score ideas and apply trap filters
-> Match method shapes and package idea cards
-> Produce scan report or sources-json updates with verified claims
Step 1: SCOPE — Frame the idea-hunt
- State the target application: "ideas for {{X}} that I'll apply in {{Y}}".
- State the method family/families: e.g., "agent planning + tool selection", "retrieval reranking", "test-time compute scaling".
- Pick sources from the Source Selection Guide. For AI/ML, default to arXiv + HF Papers + Semantic Scholar + ≥1 curator. For SWE, prefer conference proceedings (ICSE/FSE/PLDI) + GitHub repo signal (via
research-git) + industry blogs. (Papers with Code is dead — do not include it as a live source.) - Confirm time windows. Methods aging faster (LLM agents) → 30d/90d. Slower (compilers, type systems) → 1y/3y.
Step 2: SEARCH — Generate and execute queries
Run the source-specific query generator(s):
# arXiv
python3 scripts/generate_arxiv_queries.py --topic "{{topic}}" --categories cs.AI cs.CL cs.LG --windows 30d 90d 365d
# Hugging Face Papers
python3 scripts/generate_hf_papers_queries.py --topic "{{topic}}" --windows 30d 90d
# Semantic Scholar
python3 scripts/generate_semantic_scholar_queries.py --topic "{{topic}}" --min-citations 5 --windows 365d 1095d
# Papers with Code — DEAD SOURCE (Meta shutdown Jul 2025). The script is now a
# fail-loud shim that emits HF Papers + GitHub (research-git) replacement URLs.
python3 scripts/generate_papers_with_code_queries.py --task "{{task slug}}"
# Conference proceedings (manual seed list, scripts emit URLs)
python3 scripts/generate_conference_queries.py --conference neurips --year 2025 --topic "{{topic}}"
# Research blogs and curator newsletters (RSS/site map seeds)
python3 scripts/generate_blog_queries.py --domains anthropic.com openai.com deepmind.google research.google ai.meta.com --topic "{{topic}}"
For each result, extract into TSV format matching research-findings.tsv. Required fields:
source_url— Stable URL (arXiv abs page, blog post, paper landing)source_type—arxiv,hf_papers,semantic_scholar,papers_with_code,conference,industry_blog,curator_newslettersource_context— Source identifier (e.g., "arxiv:cs.AI", "hf_papers", "ss:semanticscholar.org", "neurips/2025", "anthropic.com/research", "lilianweng.github.io")paper_id— arXiv ID, DOI, conference paper ID, or canonical URL hash when no ID existstitle,authors,posted_at,observed_atmethod_family— From the idea-extraction-framework taxonomyidea_summary— 1-2 sentence statement of the method/framework/idea, not the paperevidence_grade— A/B/C/D/F using grading rubricreproducibility—code+benchmarks,code_only,paper_only,proprietarylift—low(1-3 days),medium(1-2 weeks),high(>2 weeks)trap_tags,shape_tags,quote,windowclaim_type—absolute-performance|relative-gain|efficiency|robustness; see idea-extraction-framework.md — efficiency/robustness claims transfer best regardless of evidence gradecluster_id— stable method-identity key shared by every finding about the same method across different source types. This is what drives cross-source corroboration (≥2 distinctsource_typesharing onecluster_id= corroborated). Assign a short slug per method (e.g.,reflexion-critique-retry); reuse it across the arXiv preprint, the curator mention, and the GitHub repo. If blank, the aggregator falls back topaper_idand emits a loud "corroboration unreliable" warning.
Validate before aggregation:
python3 scripts/validate_findings_tsv.py findings.tsv
Step 3: EXTRACT — Convert papers to ideas
For each surviving entry, extract the stealable unit using idea-extraction-framework.md:
- Method or framework name (or invent a clean one if the paper buries it)
- What it actually does in 1-2 sentences (no jargon shield)
- Inputs / outputs / preconditions — what you need to use it
- Evidence behind it — empirical claim + benchmark + N + baselines
- Why it might transfer to your target — and why it might not
- Lift estimate — days to a working prototype against your stack
- Kill criteria — when you'd stop pursuing it
Discard entries where the method can't be described without the original phrasing — that's a strong "no actual idea" signal.
Step 4: SCORE — Rank ideas
python3 scripts/aggregate_research_ideas.py findings.tsv --output scored.tsv --target "{{target}}"
The gate is rule-decided; the score only ranks. A deterministic rule ladder
sets gate_status; the numeric score never changes a gate decision — it only
orders rows within a bucket. This removes the old failure mode where a
subjective applicability guess (default 3) flipped promote/kill.
Rule ladder (first match wins for the gate):
- trap 11 or 12 present →
kill - ≥3 trap tags →
kill evidence_grade == F→killshape == negative-result→background(exempt from low-score kill — a falsified method you considered is information, not noise)- corroboration < 2 distinct
source_typesharing onecluster_id→ cap atvalidate(enforces the Evidence Quality Gates promote precondition) reproducibility == proprietary→ cap atvalidateevidence_grade == D→ cap atvalidate- any of traps {1,5,6,8} present → cap at
validate - else →
promote
Ranking score (ordering only, never gates): (applicability × evidence_strength × reproducibility) / (lift × trap_penalty), with per-trap numeric adjustments from known-traps.md (evidence -1 for trap 2, applicability -1/-2 for traps 3/9, lift +1 tier for trap 4). Weights: applicability 1-5 (default 3); evidence A=5 B=4 C=3 D=2 F=1; reproducibility code+benchmarks=5 code_only=4 paper_only=2 proprietary=1; lift inverse low=1 medium=3 high=5; trap_penalty 1.0 +0.5 per non-hard trap.
The aggregator emits gate_status (promote / validate / kill / background), gate_reason, score (rank-only), and corroboration (yes / no / unreliable-no-cluster_id). Do not promote kill rows; background rows go in the report's Background section, not the shortlist.
Step 5: COMPARE WINDOWS — Detect emerging vs. mature methods
Use citations-per-month-since-publication rather than raw counts to avoid penalising recent papers. Operational thresholds (Semantic Scholar influentialCitationCount):
- Emerging — first influential citations within 90 days of publication with an accelerating monthly rate (month-over-month increase ≥ 1 influential citation); sparse in 365d window
- Cresting — > 10 influential citations in the last 60 days; mentions accelerating across arXiv, HF Papers, and curator sources — adopt now or be late
- Mature — stable influential-citation rate over 90d–365d, ≥ 2 independent implementations; safest to adopt
- Declining — influential-citation rate falling for 2+ consecutive 30-day windows; likely superseded — investigate the successor
Cross-source corroboration: Methods cited in 2+ source families (e.g., arXiv paper + curator newsletter mention + Papers with Code implementation) are high-confidence steal candidates.
Step 5b: APPLY TRAP FILTER — Reject false positives
Run each top idea through known-traps.md:
- Tag each surviving idea with applicable traps (multi-tag allowed).
- Apply each trap's counter-recipe; downgrade or kill per the scoring-effect table.
- Trap 11 (
proprietary-component) and Trap 12 (benchmark-gaming) are hard kills unless an alternative exists. - Log discarded/downgraded ideas with one-line reason in the scan report.
Step 5c: MATCH SHAPES — Pattern-match surviving ideas
Match each surviving idea against shape catalog in idea-extraction-framework.md:
- Identify the shape(s):
prompting-pattern,architecture-tweak,training-recipe,evaluation-method,data-construction-recipe,inference-time-method,system-design-pattern,theoretical-bound,negative-result,survey-or-taxonomy. - Multi-shape methods often signal generality.
negative-resultis high-value when it falsifies a method you considered (saves time). The aggregator assigns itgate_status = background(rule 4) so it is never killed for lacking a benchmark gain — it lands in the report's Background section.survey-or-taxonomyis not a stealable idea — alsobackground, list as context only.
Step 6: PACKAGE — Generate idea cards
- Fill in one idea-card.md per surviving idea (use recipes.md to populate the "How to apply" section).
- Compile into research-scan-report.md.
- If updating skill
data/sources.jsonfiles, follow the format in../research-arxiv-scout/assets/sources-json-template.md.
Killer-Feature Mode (Feature-Precedent Mining)
Specialized mode for contributing the industry_blog_attribution and hci_retention_paper signals to the bundle's Killer-Feature Convergence Protocol owned by research-review-mining.
Premise. Engineering and PM blog post-mortems and HCI retention papers periodically attribute retention, conversion, or revenue to a specific feature with named metrics. These are the highest-credibility single signals in the bundle (when they exist).
When to use: bundle handoff from research-review-mining Killer-Feature Mode KF3, OR you want a published metric-backed attribution claim for a candidate feature.
Workflow:
KF-PREC-1. SCOPE — commercial product + candidate feature_id
KF-PREC-2. SCAN — generate_blog_queries.py with engineering-blog domain list
biased toward netflixtechblog/stripe/figma/linear/notion/eng.uber/etc.;
generate_conference_queries.py for CHI / CSCW / UIST / IUI
KF-PREC-3. EXTRACT — classify attribution as explicit / strong / implicit / reject;
extract the feature noun (must be testable) and the WTP quote
KF-PREC-4. APPEND — to ../research-review-mining/assets/pay-trigger-ledger.tsv
signal_type = industry_blog_attribution (blog posts)
| hci_retention_paper (CHI/CSCW/UIST/IUI)
KF-PREC-5. HAND OFF — run ../research-review-mining/scripts/converge_killer_features.py
New method shape. This mode adds monetizable-feature-pattern to the idea-extraction-framework catalog. It uses different scoring gates than the research-method shapes (Trap 11 and Trap 12 do not auto-kill; instead it kills on marketing/PR authorship and promotes on quantitative metric + internal authority).
References:
- references/feature-precedent-mining.md — full extraction protocol, source mix, anti-patterns, precision honesty
- ../research-review-mining/references/killer-feature-convergence.md — bundle Convergence Rule
- ../research-review-mining/references/llm-extraction-prompts.md §7 — engineering post-mortem attribution prompt
Templates & Assets
| Template | Purpose |
|---|---|
| research-scan-report.md | Primary output — full scan with rankings, ideas, traps caught |
| idea-card.md | Per-idea card: method, evidence, lift, how-to-apply, kill criteria |
| research-findings.tsv | Input format for aggregate_research_ideas.py (header + example) |
Scripts
| Script | Source | Purpose |
|---|---|---|
| generate_arxiv_queries.py | arXiv | export.arxiv.org/api/query URLs |
| generate_hf_papers_queries.py | HF Papers | huggingface.co/papers URLs + JSON endpoints |
| generate_semantic_scholar_queries.py | Semantic Scholar | API URLs |
| generate_papers_with_code_queries.py | Papers with Code (DEAD) | Fail-loud shim — emits HF Papers + GitHub replacement URLs (PwC shut down Jul 2025) |
| generate_conference_queries.py | Conferences | Per-venue accepted-paper-list URLs |
| generate_blog_queries.py | Blogs / newsletters | RSS + site search URLs |
| validate_findings_tsv.py | All | Findings TSV contract validation |
| aggregate_research_ideas.py | All | Idea scoring, trap filter, gate status |
References
| Reference | Covers |
|---|---|
| idea-extraction-framework.md | Method shape catalog (10 shapes), evidence grades, extraction template |
| known-traps.md | 12 research traps: irreproducibility, benchmark gaming, hype, paywall, etc. |
| recipes.md | How-to-apply playbooks for each method shape |
| arxiv-strategy.md | arXiv API, category mapping, sortBy/relevance, dedupe across versions |
| hf-papers-strategy.md | HF Papers daily, weekly trending, comment signal, RSS endpoints |
| semantic-scholar-strategy.md | Citation graph, influential-papers, embedding search, rate limits |
| papers-with-code-strategy.md | Task slugs, benchmark verification, code+stars signal — DEAD SOURCE (Meta Jul 2025); strategy file documents archive + replacement path |
| conference-proceedings-strategy.md | NeurIPS/ICML/ICLR/ACL/EMNLP/KDD/USENIX seed URLs, accepted-paper-list patterns |
| research-blogs-strategy.md | Anthropic / OpenAI / DeepMind / Google / Meta / MSR / Apple research site map |
| curator-newsletters-strategy.md | Lilian Weng, Sebastian Raschka, Eugene Yan, Latent Space, Simon Willison, The Batch, Import AI — coverage and bias notes |
| source-currency.md | May-2026 verified status table, structural shifts, and anti-pattern catalog for stale/dead/changed sources |
| free-first-sourcing-recipe.md | Decision ladder: free/official-API first → justified escalation to freemium/paid → cost-aware fallbacks |
| ../research-painpoint-scanner/references/crawl-access-economics.md | Shared (owned by research-painpoint-scanner): block signatures, control-query check, llms.txt, access-class table. Read before concluding a source has little on a topic — applies to blog/newsletter fetches and any rate-limited API |
| feature-precedent-mining.md | Killer-feature mode: contributes industry_blog_attribution + hci_retention_paper signals to the bundle's Convergence Protocol; defines the monetizable-feature-pattern method shape |
Evidence Quality Gates
These are enforced by the aggregator's rule ladder (Step 4), not advisory:
| Gate | Minimum | Enforced by |
|---|---|---|
| Cross-source corroboration | ≥2 distinct source_type sharing one cluster_id for promote | Rule 5 — caps at validate if unmet (no longer a decorative column) |
| Evidence grade | C or higher to promote | Rule 7 (D → validate), Rule 3 (F → kill) |
| Reproducibility | paper_only minimum to enter shortlist | Rule 6 (proprietary → validate, never promote) |
| Trap tags | 0-1 → ok; 2 → cap validate; 3+ → kill | Rules 2, 8 + hard-kill rule 1 |
| Negative results | never killed for low score | Rule 4 → background |
Related Skills
Shortened here. Read the whole file on GitHub.
Signals
- GitHub stars
- 87
- Forks
- 19
- Last commit
- Sep 2026
ahel review
K6low
bundled executables the agent is told to run
Automated review, not a security audit. Ruleset v1+k2.
Advanced
- Catalog kind
- skill
- Gateway key
research-scout- Source
- github.com/vasilyu1983/ai-agents-public