sam3
SkillAI & modelsSegment Anything 3 — text-, point-, and box-prompted instance
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the sam3 skill
What this skill tells your AI
The instructions your AI receives, as published by graph-robots/open-robot-skills in tools/sam3/SKILL.md and read by ahel’s review.
The SAM3 image servicer + video-tracker servicer as in-process tools. Images
are RGB uint8 [H, W, 3] numpy arrays; masks come back as gap Mask
(uint8 [H, W], 0 background / 255 foreground), score-sorted best-first.
When to use
segment_textfor open-vocabulary "find the X" masks (one mask per instance; checkscores[0]— callers typically reject below ~0.3).segment_boxafter a detector (e.g.grounding-dino.detect) for a pixel-accurate mask inside the detection box; add the point prompt (use_point=True) when a pointing model supplies one.tracker_init/tracker_update/tracker_closeto follow a single target across an observation stream (e.g. for visual servoing).
Install
uv sync --extra sam3 # torch + torchvision + the upstream sam3 package
# (pip: pip install -e ".[sam3]")
Model weights download on first model build. Device is taken from
GAP_SAM3_DEVICE (default cuda); the image model also runs on cpu
(slow), the video tracker is CUDA-only in practice.
Gotchas (carried over from the servicers)
- Lazy singletons: the image model and the video predictor each load on first call and stay resident; importing the bundle never imports torch.
segment_textcaps results atmax_results=5by default — cluttered scenes emit 100+ instances (~1 MB/mask at 720p) and downstream consumes only the top mask. Passmax_results<=0for everything.- The video tracker JIT-compiles Triton NMS kernels via the
CCenv var; a staleCC(e.g. a Ray env pointing at a non-existent gcc-13) surfaces asFileNotFoundErrorinsidetracker_init. The bundle forcesCCto a real compiler before tracker use (_ensure_cc_compiler). - Tracker prompt precedence is box > point > text; a point prompt is converted to a small (10% of image) box because the predictor's box path is more reliable for init than a single point.
- The tracker is built with
apply_temporal_disambiguation=False— the default hotstart heuristics silently delete the masklet around frame 3 in streaming mode (no fresh text re-detection per frame). - Drift handling in
tracker_update: a mask-area jump >1.5x the running median or confidence <0.30 keeps the LAST GOOD mask and reportsconfidence=0.0withobject_present=True(skip this frame); after 5 consecutive drift hitsobject_present=False— re-init the tracker. - Sessions idle longer than 120 s are evicted lazily on the next tracker
call; an evicted/unknown
tracker_idraisesToolError. tracker_initreturnsobject_present=Falsewith an emptytracker_id(no exception) when the initial detection finds nothing.
Signals
- GitHub stars
- 41
- Forks
- 7
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
sam3- Source
- github.com/graph-robots/open-robot-skills