Audio Tools

SkillMedia

Lets your agent transcribe audio files and convert speech into text.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Audio Tools skill

About this capability

Audio expert. ALWAYS invoke this skill when the user asks to transcribe, recognize, or convert speech/audio to text.

What this skill tells your AI

The instructions your AI receives, as published by yaoapp/yao in tools/skills/yao-audio/SKILL.md and read by ahel’s review.

Use these tools to transcribe audio files to text using speech-to-text models.

audio_transcribe

Transcribe an audio file to text.

tai tool audio_transcribe --audio_path /path/to/meeting.m4a
tai tool audio_transcribe --audio_path /path/to/recording.wav --language en --provider llm.my-openai:whisper-1
ParameterTypeRequiredDescription
audio_pathstringyesAudio file path. Supported: mp3, m4a, wav, webm, mp4, mpeg, mpga
languagestringnoISO 639-1 language code (e.g. en, zh, ja). Auto-detected if omitted
providerstringnoSTT provider connector ID. If omitted, uses the default STT provider

audio_providers

List available speech-to-text providers and models.

List STT providers (default):

tai tool audio_providers
ParameterTypeRequiredDescription
capabilitystringnoFilter by capability (default: audio)

Returns a list of providers with their available models and connector IDs that can be passed to audio_transcribe.

Constraints

Only use the parameters listed above for each tool. Do not pass unsupported parameters — they will be ignored or cause errors.

Signals

GitHub stars
8k
Forks
708
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
yao-audio
Source
github.com/yaoapp/yao