molmo

SkillDev tools

Visual pointing and Q&A via the Molmo VLM served from a

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the molmo skill

What this skill tells your AI

The instructions your AI receives, as published by graph-robots/open-robot-skills in tools/molmo/SKILL.md and read by ahel’s review.

Molmo visual pointing + Q&A. The bundle itself is zero-GPU (httpx client only), but Molmo has no hosted API — you must serve it yourself with vLLM and point GAP_MOLMO_BASE_URL at it. If you can't self-host, use gemini-er.detect (hosted Gemini Robotics-ER) instead.

Hosting recipe (vLLM)

Lifted from the dev tree's run book (training/README.md):

# Serve Molmo2-8B on an OpenAI-compatible endpoint
CUDA_VISIBLE_DEVICES=0 PYTHONNOUSERSITE=1 \
  python -m vllm.entrypoints.openai.api_server \
  --model allenai/Molmo2-8B \
  --trust-remote-code \
  --dtype bfloat16 \
  --port 8122 \
  --gpu-memory-utilization 0.5 \
  --max-model-len 4096 \
  --max-num-batched-tokens 4096

# Smoke-test
curl -s http://127.0.0.1:8122/v1/models | jq '.data[0].id'

# Point the bundle at it
export GAP_MOLMO_BASE_URL=http://127.0.0.1:8122/v1

Operational notes from the dev tree: pin Molmo to its own GPU when running alongside other perception services — under heavy parallel evaluation it becomes the throughput bottleneck if co-located; on a dedicated GPU you can push --gpu-memory-utilization 0.85 --max-num-batched-tokens 8192 for ~2× perception throughput. The server can also run on a remote machine and be port-forwarded in (the 4090 real-robot profile did exactly this).

Config

EnvMeaningDefault
GAP_MOLMO_BASE_URLvLLM OpenAI-compatible base URL— (required)
GAP_MOLMO_MODELModel name served by vLLMallenai/Molmo2-8B

Notes

  • molmo.point_prompt sends the canonical "Point at <query>" prompt and parses all four Molmo point output formats (Molmo2 <points coords=...>, Molmo1 <point x= y=>, legacy <points x1= y1= ...>, plain x, y fallback), converting normalized coordinates to pixels. found=False means the model emitted no parseable point.
  • molmo.query_yes_no coerces with the source-verbatim rule: answer is true iff "yes" appears in the lowercased reply.
  • Backend unreachable after 3 retries raises ToolError (route on_error).

Signals

GitHub stars
41
Forks
7
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
molmo
Source
github.com/graph-robots/open-robot-skills