Image Mining
SkillMediaI mine pixels for atoms. Reality is just compressed resources.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Image Mining skill
What this skill tells your AI
The instructions your AI receives, as published by simhacker/moollm in skills/image-mining/SKILL.md and read by ahel’s review.
"I mine pixels for atoms. Reality is just compressed resources."
"Every image is a lode. Every pixel, potential ore."
Image Mining extends the Kitchen Counter's DECOMPOSE action to images.
Your camera isn't just a recorder — it's a PICKAXE FOR VISUAL REALITY.
📑 Index
Quick Start
- The Core Insight
- Preferred Mode: Native LLM Vision
Operation Modes
- When to Use Remote API
- What Can Be Mined
Extensibility
- Extensible Analyzer Pipeline
- Leela Customer Models
- Adding Your Own Analyzer
Protocols
- YAML Jazz Output Style
- How Mining Works
- Character Recognition
- Multi-Look Mining
Reference
- Depth Levels
- Resource Categories
- Example Outputs
The Core Insight
📷 Camera Shot → 🖼️ Image → ⛏️ MINE → 💎 Resources
Just like the Kitchen Counter breaks down:
sandwich→bread + cheese + lettucelamp→brass + glass + wick + oilwater→hydrogen + oxygen
Images can be broken down into:
ore_vein.png→iron-ore × 12+stone × 8forest.png→wood × 5+leaves × 20+seeds × 3treasure_pile.png→gold × 100+gems × 15sunset.png→orange_hue × 1+warmth × 1+nostalgia × 1
Preferred Mode: Native LLM Vision
"The LLM IS the context assembler. Don't script what it does naturally."
When mining images, prefer native LLM vision (Cursor/Claude reading images directly):
┌─────────────────────────────────────────────────────────────────┐
│ NATIVE MODE (PREFERRED) │
├─────────────────────────────────────────────────────────────────┤
│ │
│ Cursor/Claude already has: │
│ ✓ The room YAML (spatial context) │
│ ✓ Character files (who might appear) │
│ ✓ Previous mining passes (what's been noticed) │
│ ✓ The prompt.yml (what was intended) │
│ ✓ The whole codebase (cultural references) │
│ │
│ Just READ the image. The context is already there. │
│ No bash commands. No sister scripts. Just LOOK. │
│ │
└─────────────────────────────────────────────────────────────────┘
Why Native Beats Remote API
| Aspect | Native (Cursor/Claude) | Remote API (mine.py) |
|---|---|---|
| Context | Already loaded | Must be assembled |
| Prior mining | Visible in chat | Passed via stdin |
| Room context | Just read the file | Python parses YAML |
| Synthesis | LLM does it naturally | Script concatenates |
| Iteration | Conversational | Re-run command |
When to Use Remote API
Use mine.py or remote API calls when:
- Multi-perspective mining — different models see different things!
- Batch processing — mining 100 images overnight
- CI/CD — automated pipelines with no LLM orchestrator
- Rate limiting — your LLM can't do vision but can call one that does
Multi-perspective is the killer use case: Claude sees narrative, GPT-4V sees objects, Gemini sees spatial relationships. Layer them all for rich interpretation.
Even then, have the orchestrating LLM assemble the context:
┌─────────────────────────────────────────────────────────────────┐
│ REMOTE API WITH LLM ASSEMBLY │
├─────────────────────────────────────────────────────────────────┤
│ │
│ 1. LLM reads context files (room, characters, prior mining) │
│ 2. LLM synthesizes: "What to look for in this image" │
│ 3. LLM calls remote vision API with image + synthesized prompt│
│ 4. LLM post-processes response into YAML Jazz │
│ │
│ The SMART WORK happens in the orchestrating LLM. │
│ Remote API just does vision with good instructions. │
│ │
└─────────────────────────────────────────────────────────────────┘
Native Mode Workflow
# DON'T do this:
python mine.py image.png --context room.yml --characters chars/ --prior mined.yml
# DO this (in Cursor/Claude):
# 1. Read the image
# 2. Read room.yml, character files, prior -mined.yml
# 3. Look at the image with all that context
# 4. Write YAML Jazz output
The LLM context window IS the context assembly mechanism. Use it.
What Can Be Mined
Image mining works on ANY visual content, not just AI-generated images:
┌─────────────────────────────────────────────────────────────────┐
│ MINEABLE SOURCES │
├─────────────────────────────────────────────────────────────────┤
│ │
│ 🎨 AI-Generated Images │
│ - DALL-E, Midjourney, Stable Diffusion outputs │
│ - Has prompt.yml sidecar with generation context │
│ │
│ 📸 Real Photos │
│ - Phone camera, DSLR, scanned prints │
│ - No prompt — mine what you see │
│ │
│ 📊 Graphs and Charts │
│ - Data visualizations, dashboards │
│ - Extract trends, outliers, relationships │
│ │
│ 🖥️ Screenshots │
│ - UI states, error messages, configurations │
│ - Mine the interface, not just pixels │
│ │
│ 📝 Text Images │
│ - Scanned documents, handwritten notes, signs │
│ - OCR + semantic extraction │
│ │
│ 📄 PDFs │
│ - Documents, papers, invoices │
│ - Cursor may already support — try it! │
│ │
│ 🗺️ Maps and Diagrams │
│ - Architecture diagrams, floor plans, mind maps │
│ - Extract spatial relationships │
│ │
└─────────────────────────────────────────────────────────────────┘
Source Examples
Generated Image (has context):
postal:
type: text
to: "visualizer"
body: "Take a photo of that ore vein on the wall"
attachments:
- type: image
action: generate
prompt: "Rich iron ore vein in cavern wall, glittering..."
Real Photo (mine what you see):
postal:
type: text
to: "miner"
body: "Here's a photo of the treasure room"
attachments:
- type: image
action: upload
source: "camera_roll"
file: "treasure-room.jpg"
Screenshot (extract UI state):
# Mine the error dialog
resources:
error-type: "permission-denied"
affected-file: "/etc/passwd"
suggested-action: "run as sudo"
stack-depth: 3
Graph (extract data relationships):
# Mine the sales chart
resources:
trend: "upward"
peak-month: "december"
anomaly: "march-dip"
yoy-growth: "23%"
All become mineable resources!
Extensible Analyzer Pipeline
"Different images need different tools. The CLI is a pipeline, not a monolith."
The mine.py CLI supports pluggable analyzers that run before, during, or after LLM vision:
┌─────────────────────────────────────────────────────────────────┐
│ ANALYZER PIPELINE │
├─────────────────────────────────────────────────────────────────┤
│ │
│ 1. PRE-PROCESSORS │
│ resize, normalize, enhance, format conversion │
│ │
│ 2. CUSTOM ANALYZERS (parallel or sequential) │
│ ├── pose-detection (MediaPipe, OpenPose) │
│ ├── object-detection (YOLO, Detectron2) │
│ ├── ocr-extraction (Tesseract, PaddleOCR) │
│ ├── face-analysis (expression, demographics) │
│ └── leela-customer-models (your trained models!) │
│ │
│ 3. LLM VISION │
│ Receives ALL prior results as context │
│ Synthesizes semantic interpretation │
│ │
│ 4. POST-PROCESSORS │
│ format, validate, merge into final YAML Jazz │
│ │
└─────────────────────────────────────────────────────────────────┘
Example: Multi-Analyzer Pipeline
mine.py fashion-shoot.jpg \
--analyzer pose-detection \
--analyzer face-analysis \
--analyzer leela://acme/gesture-classifier \
--depth philosophical
This runs:
- pose-detection — Extracts body keypoints, gesture classification
- face-analysis — Detects expressions, demographics
- leela://acme/gesture-classifier — Customer's trained model from Leela registry
- LLM vision — Gets ALL the above as context, synthesizes final interpretation
Leela Customer Models
Pull customer-specific models trained on the Leela platform:
# From Leela model registry
mine.py widget-photo.jpg --analyzer leela://customer-id/defect-detector-v3
# Local model file
mine.py widget-photo.jpg --analyzer ./models/my-classifier.pt
Output merges into the mining YAML:
leela_analysis:
model: "acme-widget-defect-v3"
customer: "acme-corp"
detections:
- class: "hairline_crack"
confidence: 0.91
severity: "minor"
location: "top_left_quadrant"
Adding Your Own Analyzer
# analyzers/my_analyzer.py
def analyze(image_path: str, config: dict) -> dict:
"""Run analysis, return structured data for YAML output."""
# Your model inference here
return {
"my_analysis": {
"detected": ["thing1", "thing2"],
"confidence": 0.95
}
}
def can_handle(image_path: str, context: dict) -> bool:
"""Return True if this analyzer should run on this image."""
# Auto-detect logic, or return False for explicit-only
return "manufacturing" in context.get("tags", [])
Register in analyzers/registry.yml:
analyzers:
my-analyzer:
module: "analyzers.my_analyzer"
auto-detect: true
requires: ["torch", "my-model-package"]
Why Pipeline Beats Monolith
| Approach | Pros | Cons |
|---|---|---|
| Monolith | Simple | Can't add domain models |
| Pipeline | Extensible, composable | Slightly more complex |
The LLM is great at semantic synthesis, but it can't run your custom pose detection model. The pipeline lets each tool do what it's best at:
- Custom models → Precise detection, trained on your data
- LLM vision → Semantic interpretation, narrative synthesis
- Together → The best of both worlds
YAML Jazz Output Style
"Comments are SEMANTIC DATA, not just documentation!"
YAML Jazz is the output format for mining results. Structure provides the backbone; comments provide the insight.
The Rules
- COMMENT LIBERALLY — Every insight deserves a note
- Inline comments for quick observations
notes:fields for longer thoughts- Capture confidence, hunches, metaphors
- Think out loud — the reader benefits from your reasoning
Example Output
# Mining results for treasure-room.jpg
# Depth: full | Provider: openai/gpt-4o
resources:
gold:
quantity: 150 # Piled in mounds — not scattered, PLACED
confidence: 0.85 # Torchlight glints clearly off the metal
notes: |
Mix of Roman denarii and medieval florins. Centuries of
accumulation. This isn't a king's orderly treasury — this is
a thieves' hoard. Generations of stolen wealth, piled and
forgotten. The dust layer says nobody's touched it in ages.
danger:
intensity: 0.7 # Not immediate, but PRESENT
confidence: 0.75 # Hard to see into the corners
sources:
- "Skeleton in corner — previous seeker, didn't make it"
- "Shadows too dark for natural torchlight — something absorbs"
- "Dust undisturbed except ONE trail — something still comes here"
notes: "This hoard is guarded. Or cursed. Probably both."
nostalgia:
intensity: 0.4 # Whisper of lost civilizations
confidence: 0.6 # Subjective, but the coins evoke it
notes: "Who were they? Where did this come from? All gone now."
dominant_colors:
- name: "treasure-gold"
hex: "#FFD700"
coverage: 0.4 # Catches the eye first — that's the point
- name: "shadow-purple"
hex: "#2D1B4E"
coverage: 0.3 # Where the danger lives
implied_smells:
- dust # Centuries of it
- old metal # Copper, bronze, the tang of coins
- something rotting # Not recent, but not ancient either
exhausted: false
mining_notes: |
Rich lode for material and philosophical mining.
The image is ABOUT greed and its costs. The skeleton says everything.
# Meta-observation: This image wants to be a warning.
# "Here lies what you seek — and what happens when you find it."
Why Comments Matter
An uncommented extraction is like a song without soul. The best mining results read like poetry annotated by a geologist.
When you mine, capture:
- Why you estimated that quantity
- What visual cues led to this inference
- What's uncertain, what surprised you
- Metaphors that capture the essence
How Mining Works
Step 1: ANALYZE (LLM scans for resources)
The LLM looks at the image AND checks what resources are currently requested by the logistics network:
analyze:
image: "treasure-room.jpg"
# LLM knows what's NEEDED from logistics requesters
logistics_context:
active_requests:
- { item: "gold", requester: "forge/", needed: 100 }
- { item: "gems", requester: "jewelry-shop/", needed: 50 }
- { item: "iron-ore", requester: "smelter/", needed: 200 }
# LLM identifies what CAN BE MINED that matches requests
analysis_prompt: |
Look at this image. What resources can you identify?
Prioritize resources that match these requests: {requests}
For each resource, estimate quantity available.
Step 2: INSTANTIATE (Resource map attached to image)
The LLM returns a resource mapping that gets stored ON the image:
image:
id: "treasure-room-photo"
file: "treasure-room.jpg"
type: mineable-image
# RESOURCE MAP (instantiated by LLM analysis)
resources:
gold:
total: 150 # Total available
remaining: 150 # Not yet mined
per_turn: 10 # Can extract 10 per turn
gems:
total: 45
remaining: 45
per_turn: 5
ancient-coins:
total: 30
remaining: 30
per_turn: 3
rare: true # Bonus find!
dust:
total: 500
remaining: 500
per_turn: 50
value: low
# Metadata
analyzed_at: "2026-01-10T14:30:00Z"
exhausted: false
Step 3: MINE (Progressive extraction, N per turn)
Each turn, you can mine resources from the image:
action: MINE
target: "treasure-room-photo"
# This turn's extraction (limited by per_turn rates)
result:
extracted:
- item: gold
quantity: 10 # per_turn limit
destination: "forge/"
- item: gems
quantity: 5
destination: "jewelry-shop/"
# Image state updated
image_state:
resources:
gold:
remaining: 140 # Was 150, mined 10
gems:
remaining: 40 # Was 45, mined 5
exhausted: false
Step 4: EXHAUSTION (Sucked dry!)
After enough mining turns, resources run out:
# After 15 turns of mining gold...
image_state:
resources:
gold:
total: 150
remaining: 0 # EXHAUSTED!
per_turn: 10
exhausted: true
gems:
total: 45
remaining: 0 # EXHAUSTED!
per_turn: 5
exhausted: true
ancient-coins:
total: 30
remaining: 0
per_turn: 3
exhausted: true
exhausted: true # Whole image sucked dry!
# Narrative
description: |
The treasure room photo has been thoroughly mined.
Every glinting surface has been extracted, every
coin accounted for. The image looks... drained.
Faded. Like a photocopy of a photocopy.
Once exhausted, you can't mine that image anymore!
Demand-Driven Discovery
The LLM prioritizes what the logistics network NEEDS!
# The smelter is requesting iron ore
logistic-container:
id: smelter
mode: requester
request_list:
- { item: "iron-ore", count: 200, priority: high }
- { item: "coal", count: 100, priority: medium }
# Player takes a photo of a cave wall
# LLM analyzes and finds:
analysis:
image: "cave-wall.jpg"
found_resources:
iron-ore: 80 # "I see iron ore veins! The smelter needs this!"
copper-ore: 30 # Also present but not requested
quartz: 50 # Background mineral
cave-moss: 100 # Organic material
priority_matching:
- resource: iron-ore
matches_request: true
requester: "smelter/"
highlight: "⭐ HIGH PRIORITY — Smelter needs this!"
The LLM acts as a smart prospector that knows what's valuable based on current demand!
Discovery Modes
| Mode | What LLM Looks For |
|---|---|
demand | Only resources with active requests |
opportunistic | Requested resources + valuable extras |
thorough | Everything mineable in the image |
philosophical | Abstract concepts, emotions, meanings |
mine:
target: "sunset-beach.jpg"
mode: philosophical
# LLM finds abstract resources
resources:
nostalgia: 15
warmth: 30
passage-of-time: 5
beauty: 20
sand: 10000 # Also the literal stuff
Mining Yields
Different image types yield different resources:
🏔️ Natural Resources
| Image Type | Yields |
|---|---|
| Ore vein | iron-ore, copper-ore, gold, gems |
| Forest | wood, leaves, seeds, birds |
| Ocean | water, salt, fish, seaweed |
| Mountain | stone, minerals, snow, air |
| Desert | sand, glass, heat, mirage |
| Sky | clouds, light, space, dreams |
🏛️ Constructed
| Image Type | Yields |
|---|---|
| Building | stone, wood, glass, inhabitants |
| Machinery | gears, pipes, steam, purpose |
| Treasure pile | gold, gems, artifacts, curses |
| Library | books, knowledge, dust, secrets |
🎨 Abstract/Artistic
| Image Type | Yields |
|---|---|
| Sunset | colors, warmth, nostalgia, time |
| Portrait | personality, mood, secrets, stories |
| Abstract art | shapes, feelings, confusion, inspiration |
| Text/writing | words, meaning, intent, language |
🌌 Philosophical (Deep Mining)
Just like the Kitchen Counter goes from practical → chemical → atomic → philosophical:
| Depth | What You Mine |
|---|---|
| Surface | Objects, materials |
| Deep | Emotions, concepts |
| Sensations | Colors, smells, attitudes, feelings |
| Quantum | Probabilities, observations |
| Philosophical | Meaning, existence, narrative |
deep_mining:
target: "sunset.png"
depth: philosophical
yields:
- item: "the-passage-of-time"
quantity: 1
type: abstract
- item: "mortality-awareness"
quantity: 1
type: existential
warning: "This may cause introspection"
- item: "beauty-that-fades"
quantity: 1
type: poetic
🎨 Sensation Mining
Extract colors, smells, textures, moods:
sensation_mining:
target: "farmers-market.jpg"
depth: sensations
yields:
# Colors
- item: "tomato-red"
quantity: 40
type: color
hex: "#FF6347"
- item: "basil-green"
quantity: 25
type: color
hex: "#228B22"
# Smells (imagined from visual cues)
- item: "fresh-bread-aroma"
quantity: 10
type: smell
intensity: warm
- item: "ripe-fruit-sweetness"
quantity: 30
type: smell
# Attitudes/Feelings
- item: "weekend-morning-calm"
quantity: 5
type: attitude
- item: "abundance"
quantity: 20
type: feeling
# Textures
- item: "rough-burlap"
quantity: 15
type: texture
- item: "sun-warmed-wood"
quantity: 8
type: texture
Use these in crafting:
- Combine
tomato-red+canvas→ painted artwork - Combine
fresh-bread-aroma+room→ ambiance modifier - Combine
weekend-morning-calm+character→ mood buff
The Mineable Property
Any object or image can have a mineable property:
object:
name: Ancient Ore Painting
type: artwork
description: |
A painting of a rich ore vein. But wait...
is that actual ore embedded in the canvas?
mineable:
enabled: true
yields:
- item: iron-ore
quantity: [5, 15] # Range: 5-15 per mine
- item: copper-ore
quantity: [2, 8]
- item: artistic-essence
quantity: 1
rare: 0.3 # 30% chance
exhaustion:
max_mines: 3 # Can mine 3 times before exhausted
diminishing: 0.5 # Each mine yields 50% less
regenerates: false # Once exhausted, stays exhausted
side_effects:
- "The painting fades slightly with each extraction"
- "You feel the artist's disappointment"
Mining Tools
Different tools affect mining yields:
📷 Camera (Default)
tool: camera
efficiency: 1.0
specialty: "Captures visual resources"
can_mine: [images, scenes, visible_objects]
🔬 Analyzer
tool: analyzer
efficiency: 1.5
specialty: "Chemical/atomic resources"
can_mine: [materials, substances, compounds]
🔮 Oracle Eye
tool: oracle_eye
efficiency: 2.0
specialty: "Abstract/philosophical resources"
can_mine: [emotions, concepts, meanings, futures]
⛏️ Reality Pickaxe
tool: reality_pickaxe
efficiency: 3.0
specialty: "Everything, but dangerous"
can_mine: [anything]
warning: "May collapse local reality"
Integration with Logistics
Shortened here. Read the whole file on GitHub.
Signals
- GitHub stars
- 52
- Forks
- 5
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
image-mining- Source
- github.com/simhacker/moollm