Audio Tools
SkillMediaLets your agent transcribe audio files and convert speech into text.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Audio Tools skill
About this capability
Audio expert. ALWAYS invoke this skill when the user asks to transcribe, recognize, or convert speech/audio to text.
What this skill tells your AI
The instructions your AI receives, as published by yaoapp/yao in tools/skills/yao-audio/SKILL.md and read by ahel’s review.
Use these tools to transcribe audio files to text using speech-to-text models.
audio_transcribe
Transcribe an audio file to text.
tai tool audio_transcribe --audio_path /path/to/meeting.m4a
tai tool audio_transcribe --audio_path /path/to/recording.wav --language en --provider llm.my-openai:whisper-1
| Parameter | Type | Required | Description |
|---|---|---|---|
| audio_path | string | yes | Audio file path. Supported: mp3, m4a, wav, webm, mp4, mpeg, mpga |
| language | string | no | ISO 639-1 language code (e.g. en, zh, ja). Auto-detected if omitted |
| provider | string | no | STT provider connector ID. If omitted, uses the default STT provider |
audio_providers
List available speech-to-text providers and models.
List STT providers (default):
tai tool audio_providers
| Parameter | Type | Required | Description |
|---|---|---|---|
| capability | string | no | Filter by capability (default: audio) |
Returns a list of providers with their available models and connector IDs that can be passed to audio_transcribe.
Constraints
Only use the parameters listed above for each tool. Do not pass unsupported parameters — they will be ignored or cause errors.
Signals
- GitHub stars
- 8k
- Forks
- 708
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
yao-audio- Source
- github.com/yaoapp/yao