Audio Jingle Skill

SkillFiles & storage

audio-jingle is a skill that lets an AI agent generate short audio files: jingles, background music beds, voiceovers, and sound effects. It plans the piece first, routes the request to a matching music, speech, or SFX model, and saves the result as one MP3 or WAV file in the project folder.

Available today. Use it from your connected AI after setup.

Have a project with audio metadata such as audioKind, audioModel, audioDuration, and for speech a voice value.

Then ask your AI: use the Audio Jingle Skill skill

What your AI can do with it

  • Generates music such as jingles and beds via Suno V5, Udio, or Lyria 2
  • Produces voiceover and speech through MiniMax TTS, Fish, or ElevenLabs V3
  • Creates sound effects using ElevenLabs SFX or AudioCraft
  • Plans each piece first: genre, tempo, script, voice, or texture
  • Binds requested duration directly to the API parameter
  • Saves the finished audio as one MP3/WAV file in the project folder

Getting started

  1. Have a project with audio metadata such as audioKind, audioModel, audioDuration, and for speech a voice value.
  2. Add the audio-jingle skill to the agent so its triggers (music, jingle, bed, voiceover, tts, sound effect) are available.
  3. Ask the agent for the audio you want, for example a 30-second jingle or a narrated script.
  4. For MiniMax speech, supply a valid voice_id if you want a specific voice; otherwise the default voice applies.
  5. Find the finished MP3 or WAV file in the project folder.

What this skill tells your AI

The instructions your AI receives, as published by nexu-io/open-design in design-templates/audio-jingle/SKILL.md and read by ahel’s review.

Three sub-modes. The active project's audioKind decides which one runs:

audioKindModels we route toPlan focus
musicSuno V5 (default), Udio, Lyria 2genre + tempo + instrumentation
speechMiniMax TTS (default), Fish, ElevenLabs V3script + voice + pacing
sfxElevenLabs SFX (default), AudioCrafttexture + impact + duration

Resource map

audio-jingle/
├── SKILL.md
└── example.html

Workflow

Step 0 — Read the project metadata

audioKind, audioModel, audioDuration (seconds), and (for speech) voice. Branch by known values and use them verbatim. Missing metadata is not an instruction to ask: infer a safe default when possible, and emit a clarifying form only when the missing answer would materially change the requested output or prevent generation.

Important: voice is provider-specific. For minimax-tts, --voice must be a valid MiniMax voice_id (for example male-qn-qingse), not a natural-language description. If you only have a prose voice brief ("warm female narrator", "neutral Mandarin"), keep that in your plan but omit --voice so the daemon's default voice id applies, or ask the user to choose a specific id.

Step 1 — Plan

Music

  • Genre + reference artists (1-2)
  • Tempo (BPM) + key
  • Instrumentation (3-5 instruments max)
  • Vocals: yes / no / hummed / choir
  • Mood arc (intro → chorus → outro)

Speech

  • Script (final, not draft — TTS runs verbatim)
  • Voice target + pacing For MiniMax this means a real voice_id, not prose in --voice
  • Pronunciation hints for proper nouns / acronyms

SFX

  • Texture (impact / whoosh / ambience / foley)
  • Duration + envelope (sharp attack vs. gentle swell)
  • Layering note (single hit vs. stacked)

State the plan in 2-3 sentences before dispatching.

Step 2 — Compose the prompt

Use the format the upstream model prefers. Bind audioDuration to the API parameter directly; never put "make it 30 seconds" in prose.

Step 3 — Dispatch via the media contract

Use the unified dispatcher — do not call provider APIs by hand:

"$OD_NODE_BIN" "$OD_BIN" media generate \
  --project "$OD_PROJECT_ID" \
  --surface audio \
  --audio-kind "<music|speech|sfx>" \
  --model "<audioModel from metadata>" \
  --duration <audioDuration seconds> \
  [--voice "<provider voice id (speech only)>"] \
  --output "<short-slug>-<duration>s.mp3" \
  --prompt "<assembled prompt from Step 2 — for speech, the literal script>"

The command prints one line of JSON: {"file": {"name": "...", ...}}. The bytes land in the project; the FileViewer renders the audio transport controls automatically.

Step 4 — Hand off

Reply with: plan summary, the filename returned by the dispatcher, and one sentence on what to try if the user wants a variation (e.g. "swap tempo from 92 to 108 BPM" rather than "make it different").

Hard rules

  • TTS runs your script literally. Proof it before dispatching — even one stray comma changes the cadence.
  • MiniMax TTS rejects free-form voice prose in --voice. Use a real MiniMax voice_id (for example male-qn-qingse) or omit the flag and let the daemon's default voice apply.
  • Music: under 30s = single section; 30–90s = intro + body; 90s+ = full arc. Don't try to fit a 3-act song into 15 seconds.
  • SFX: prefer one well-described layer over a paragraph of "make it cool" — generators reward specific texture words.
  • Save the file every turn. The audio viewer shows transport controls the moment the file lands.

Signals

GitHub stars
98k
Forks
11k
Last commit
Sep 2026

Others that do the same job

Questions

What kinds of audio can it make?
Three kinds, chosen by the project's audioKind setting: music (jingles and beds), speech (voiceover and narration), and sound effects. Each kind routes to different models and has its own planning focus.
Which models does it use?
Music goes to Suno V5 (default), Udio, or Lyria 2; speech to MiniMax TTS (default), Fish, or ElevenLabs V3; sound effects to ElevenLabs SFX (default) or AudioCraft.
Where does the output go?
The finished audio is saved as a single MP3 or WAV file in the project folder.
Advanced
Item type
skill
Key
audio-jingle-nexu-io
Source
github.com/nexu-io/open-design