multimodal

PackMedia

Gives your agent access to image, audio, and video AI models like Whisper, CLIP, and Stable Diffusion.

Unavailable. Delivery for this kind is on the roadmap — not serving yet.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

About this app

Vision, audio, and multimodal models including CLIP, Whisper, LLaVA, BLIP-2, Segment Anything, Stable Diffusion, AudioCraft, Cosmos Policy, OpenPI, and OpenVLA-OFT. Use when working with images, audio, multimodal tasks, or vision-language-action robot policies.

Signals

GitHub stars
13k
Forks
931
Last commit
Jun 2026
Advanced
Item type
plugin
Key
orchestra-research-ai-research-skills-multimodal
Source
github.com/orchestra-research/ai-research-skills