sam3

SkillAI & models

Segment Anything 3 — text-, point-, and box-prompted instance

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the sam3 skill

What this skill tells your AI

The instructions your AI receives, as published by graph-robots/open-robot-skills in tools/sam3/SKILL.md and read by ahel’s review.

The SAM3 image servicer + video-tracker servicer as in-process tools. Images are RGB uint8 [H, W, 3] numpy arrays; masks come back as gap Mask (uint8 [H, W], 0 background / 255 foreground), score-sorted best-first.

When to use

  • segment_text for open-vocabulary "find the X" masks (one mask per instance; check scores[0] — callers typically reject below ~0.3).
  • segment_box after a detector (e.g. grounding-dino.detect) for a pixel-accurate mask inside the detection box; add the point prompt (use_point=True) when a pointing model supplies one.
  • tracker_init / tracker_update / tracker_close to follow a single target across an observation stream (e.g. for visual servoing).

Install

uv sync --extra sam3       # torch + torchvision + the upstream sam3 package
# (pip: pip install -e ".[sam3]")

Model weights download on first model build. Device is taken from GAP_SAM3_DEVICE (default cuda); the image model also runs on cpu (slow), the video tracker is CUDA-only in practice.

Gotchas (carried over from the servicers)

  • Lazy singletons: the image model and the video predictor each load on first call and stay resident; importing the bundle never imports torch.
  • segment_text caps results at max_results=5 by default — cluttered scenes emit 100+ instances (~1 MB/mask at 720p) and downstream consumes only the top mask. Pass max_results<=0 for everything.
  • The video tracker JIT-compiles Triton NMS kernels via the CC env var; a stale CC (e.g. a Ray env pointing at a non-existent gcc-13) surfaces as FileNotFoundError inside tracker_init. The bundle forces CC to a real compiler before tracker use (_ensure_cc_compiler).
  • Tracker prompt precedence is box > point > text; a point prompt is converted to a small (10% of image) box because the predictor's box path is more reliable for init than a single point.
  • The tracker is built with apply_temporal_disambiguation=False — the default hotstart heuristics silently delete the masklet around frame 3 in streaming mode (no fresh text re-detection per frame).
  • Drift handling in tracker_update: a mask-area jump >1.5x the running median or confidence <0.30 keeps the LAST GOOD mask and reports confidence=0.0 with object_present=True (skip this frame); after 5 consecutive drift hits object_present=False — re-init the tracker.
  • Sessions idle longer than 120 s are evicted lazily on the next tracker call; an evicted/unknown tracker_id raises ToolError.
  • tracker_init returns object_present=False with an empty tracker_id (no exception) when the initial detection finds nothing.

Signals

GitHub stars
41
Forks
7
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
sam3
Source
github.com/graph-robots/open-robot-skills