multimodal
PackMediaGives your agent access to image, audio, and video AI models like Whisper, CLIP, and Stable Diffusion.
Unavailable. Delivery for this kind is on the roadmap — not serving yet.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
About this app
Vision, audio, and multimodal models including CLIP, Whisper, LLaVA, BLIP-2, Segment Anything, Stable Diffusion, AudioCraft, Cosmos Policy, OpenPI, and OpenVLA-OFT. Use when working with images, audio, multimodal tasks, or vision-language-action robot policies.
Signals
- GitHub stars
- 13k
- Forks
- 931
- Last commit
- Jun 2026
Advanced
- Item type
- plugin
- Key
orchestra-research-ai-research-skills-multimodal- Source
- github.com/orchestra-research/ai-research-skills