ONLINEAGENT_OPS 2026.Q3 HOME ARTICLES CRAFT RECORD BLOG HUBS FAQ SEARCH
HOMEARTICLESWHO OWNS WHAT YOU GENERATE
EXPLAINER · RIGHTS

Prompting Music and Voice

Sound happens over time. Almost every audio prompting failure comes from describing it as though it were a picture.

READ6 min
WORDS1,137
SECTIONS4
TYPEGUIDE
CHECKED31 AUG 26

Audio prompting is written like image prompting and it should not be. Sound happens over time, and almost every failure comes from describing it as if it were a picture.

TL;DR — THE SHORT VERSION
  • Sound has a shape in time. A prompt that lists adjectives gets you a texture; a prompt with a structure gets you a piece.
  • Say the arrangement, not the mood. "Sparse, then drums at the halfway point" beats "epic and emotional".
  • Name the instruments and what they do. Which one carries the tune, which one keeps time, what is absent.
  • Voice needs direction, not description. Pace, pauses, and where the emphasis lands — the things a director would say.
  • Silence is an instruction. Most generated audio is too full, because nobody asks for space.
  • Cloning a real voice is a legal question before a technical one. Consent has to cover training, not just use.
IN PLAIN ENGLISH

You are not describing a sound. You are describing a journey a listener takes — where it starts, what changes, where it ends.

An image prompt describes a moment. An audio prompt has to describe a sequence. That single difference explains most of what goes wrong.

Music: structure first, mood last

01

Adjectives produce wallpaper

"Epic emotional cinematic uplifting" is four mood words and no instruction. What comes back is a texture that stays the same for the whole duration — technically matching every word and going nowhere.

Structure is what makes it a piece of music. Say what happens and when: what opens, what enters next, where it lifts, where it stops. Even a rough map beats a pile of adjectives.

A usable shape, in the order that works: genre and tempo → instrumentation → structure over time → production character → what to leave out.

02

Name the jobs, not just the instruments

Listing instruments gives the model a palette. Telling it which one does what gives it an arrangement.

"Piano and strings" is a palette. "Piano carries the melody, strings pad underneath, no percussion until the second half" is an arrangement — and it is the difference between something you can edit to picture and something you cannot.

Say what is absent. "No drums", "no vocals", "no bass" are among the most effective instructions available, because the default is to fill everything.

03

Write for the cut you actually need

Music for video has requirements that generic music does not: a defined length, a clean start, and an ending rather than a fade. Ask for them explicitly.

And ask for space. A track that is busy throughout will fight your voiceover for the whole runtime. "Sparse under dialogue, fuller in the gaps" is a mixing instruction, and it saves an edit later.

TAKEAWAY

If your music prompt has no word describing time — no "opens with", "enters at", "drops out", "builds to" — you are prompting for a texture and you will get one.

Voice: direct it like a performance

04

Adjectives fail here too

"Warm friendly professional" describes a person, not a reading. What you need is what a director says in the booth: where to slow down, where to pause, which word carries the sentence.

The most useful controls, in order of effect:

Pace, and where it changes. Slow on the important line, quicker through the setup.BIGGEST
Pauses, placed deliberately. Punctuation is your main lever; a full stop and a comma do different things.BIG
Emphasis on specific words. Say which. The default stress is usually flat or on the wrong one.BIG
Intent — reassuring, explaining, warning. Actors work from intent, and it moves delivery more than tone words do.USEFUL
Tone adjectives alone. Something, but least of the five.LEAST

Then fix the script, not the settings. Long sentences with stacked clauses read badly aloud whoever is speaking them. If a line will not land, shorten it before you regenerate — the same advice that makes writing better for screen-reader users.

05

The words that get pronounced wrong

Names, places, acronyms, technical terms and anything foreign. Exactly the words a listener cannot reconstruct from context, which is what makes the error expensive.

Two fixes. Spell it phonetically in the input where the tool allows. And check every proper noun in the output rather than listening to the whole thing for vibe — targeted checking finds more in less time.

This is where audio stops being a craft question. A synthetic voice modelled on a real person is governed by the right of publicity, not by copyright — and in some places voice is now protected explicitly, including simulations of it.

The clause people miss: consent to use a recording is not consent to train a model on it. Those are separate permissions and the second one has to be written down.

The full position, with the law and the dates, is on cloning a voice or a face. For commercial work the low-risk default is a voice that belongs to nobody — fully synthetic, not modelled on a specific person.

DISCLOSE SYNTHETIC VOICE WHERE IT MATTERS

Not everywhere — a synthesised read on a product demo needs no announcement. But where a listener would assume a real person is speaking to them, and would feel misled to learn otherwise, say so. Testimonials, anything presented as first-hand experience, anything in a voice that sounds like a specific known person.

The test is the same one this site applies to written work: would the audience's judgement change if they knew?

Before you export

1 — Does the music prompt say what happens over time, not just how it feels?
2 — Have I said what is absent, as well as what is present?
3 — Does it end, rather than fade, if it has to fit a cut?
4 — For voice: have I directed pace, pauses and emphasis rather than adjectives?
5 — Have I checked every proper noun and acronym in the output?
6 — If a real person's voice is involved, do I have written consent covering training?
HONESTY ABOUT THIS PAGE

The prompting advice here is craft, not measurement. It is reasoned from what these tools respond to and from ordinary practice in music and voice direction; this site has not run controlled comparisons, and behaviour differs between tools and versions. Model-specific syntax lives in the cheat sheets, which quote vendor documentation where it exists.

The one part that is not craft is the consent section, which is sourced and dated on its own page. Checked 31 August 2026.

The through-line: sound is a sequence. Describe the shape of it in time, say what is missing, and direct a voice the way you would direct a person.