1,136prompts analyzed
7median words per prompt
21.6%name a tempo
2.4%mention song structure
What the data says
  • Prompts are short. 768 of 1,136 prompts, 68% of the corpus, are ten words or fewer.
  • Genre is doing all the work. Pop appears in 128 prompts, more than any other genre family. Most prompts name a genre and stop.
  • Tempo is mostly missing. Only 245 prompts give a BPM. Of those that do, the median is 100 and almost all sit between 70 and 140.
  • The voice is an afterthought. 69% of prompts say nothing at all about who is singing, or whether anyone is.
  • Nobody asks for a chorus. 20 prompts mention a chorus or hook. 3 mention a verse. 1 mentions a bridge.

What people actually write

The first thing almost every prompt does is name a genre, and for a lot of them it is the only thing they do. Pop leads, which is unsurprising. The interesting part is how flat the tail is: cinematic scoring sits third, ahead of rock and jazz, which says something about what people are actually making. Background music for video, not songs.

Genre terms, by share of promptsHow often each genre family is named across the 1,136 prompts. A prompt can name more than one.
Table view
TermPromptsShare
Pop12811.3%
Hip hop / Rap1049.2%
Cinematic948.3%
Electronic / EDM938.2%
Rock918.0%
Jazz817.1%
Classical817.1%
Folk / Acoustic807.0%
Ambient797.0%
R&B / Soul696.1%
Lo-fi665.8%
Metal585.1%
Reggae / Dub363.2%
Country343.0%

A genre word is a strong instruction and a blunt one. It sets instrumentation, tempo range and production style all at once, which is why it works. It also means two people writing lo fi hip hop get near-identical results, because they have handed over every other decision.

Roughly a third of prompts add an emotional word on top. The vocabulary there is small.

Mood and emotion words, by share of promptsThe second layer after genre, and it clusters hard around a few safe words.
Table view
TermPromptsShare
Chill / Relaxed847.4%
Warm746.5%
Dark655.7%
Romantic615.4%
Nostalgic464.0%
Energetic454.0%
Dreamy403.5%
Aggressive363.2%
Epic353.1%
Uplifting292.6%
Melancholic262.3%
Mysterious242.1%
Melancholic appears in 26 prompts out of 1,136. Sad songs are most of what people want and almost none of what they ask for.

Instrumentation stops at the rhythm section. Drums and bass dominate, then guitar, synth and piano, and after that it drops off a cliff. Saxophone, organ, arpeggios and handclaps together appear in fewer prompts than piano alone. This is the layer with the most room in it.

Instruments named, by share of promptsRhythm section first. Everything else is a long tail.
Table view
TermPromptsShare
Drums16314.3%
808 / bass13712.1%
Guitar1119.8%
Synth1089.5%
Piano1049.2%
Strings474.1%
Trumpet / brass423.7%
Pads393.4%
Vinyl / tape texture242.1%
Saxophone131.1%
Organ121.1%
Arpeggio100.9%

Only one prompt in five names a tempo

This is the finding that surprised us most. Tempo is the single fastest lever on how a track feels, it takes three characters to specify, and 78% of published prompts leave it out. Among the 245 that do name one, the distribution is tight.

Where the tempos landOnly 21.6% of prompts give a tempo at all. Among those that do, the median is 100 BPM and the spread is narrow.

Bins in BPM. Values below 2 prompts omitted for readability.

Table view
BinPromptsShare
50-5962.4%
60-69104.1%
70-792711.0%
80-893413.9%
90-993815.5%
100-1092610.6%
110-1192610.6%
120-1293514.3%
130-139176.9%
140-149176.9%
150-15931.2%
160-16931.2%
170-17920.8%
If nearly every prompt that specifies tempo lands in the same forty-BPM band, then the tracks people generate all move at roughly the same speed. Going deliberately slower or faster than the pack is a free way to sound unlike everything else.

Seven words

Half of all published prompts are seven words or shorter. That is a genre, a mood and maybe one instrument. It is not enough information to specify a song, so the model fills the gap with whatever is most statistically ordinary for that genre, which is exactly why so much AI music sounds the same.

How long a prompt actually isMedian 7 words. Two thirds of every prompt published online is ten words or fewer.
Table view
BinPromptsShare
1-5 words41136.2%
6-10 words35731.4%
11-15 words12911.4%
16-20 words16114.2%
21-30 words645.6%
31+ words141.2%

The voice is the biggest blind spot

356 prompts out of 1,136 say anything about vocals at all. That leaves 69% where the most noticeable element of a song, whether there is a singer and what they sound like, is left entirely to chance.

What prompts say about the voiceThe biggest gap. Most prompts describe the instrumental and leave the singer to chance.
Table view
TermPromptsShare
Mentions vocals at all35631.3%
Instrumental / no vocals12010.6%
Male vocals665.8%
Female vocals655.7%
Layered / harmonies373.3%
Soft / breathy282.5%
Rapped121.1%
Raspy / gritty90.8%

Worth noting that 120 prompts explicitly ask for no vocals, which is a real instruction and a good one. The problem is the silent majority: the prompts that neither ask for a voice nor rule one out.

Almost nobody asks for a song

This is the finding with the biggest practical payoff. Out of 1,136 prompts, 20 mention a chorus or a hook, 3 mention a verse and 1 mentions a bridge.

How often a prompt mentions song structureAlmost nobody asks for a chorus.

Plotted on the same axis as the genre chart, where the leading term reaches 11.3%. Auto-scaling this chart to its own maximum would make these counts look far larger than they are.

Table view
TermPromptsShare
Chorus201.8%
Drop / breakdown111.0%
Intro / outro50.4%
Verse30.3%
Bridge10.1%
People describe a sound, then wonder why they got a loop instead of a song.

A texture prompt produces a texture. If you want a track that goes somewhere, the prompt has to say where. A chorus that lifts an octave. A bridge that drops to near silence. A beat switch under the hook. These are single clauses and they change the output more than any adjective will.

Era references are rare too, under 7% of prompts, and heavily weighted to the eighties. Which means the 2000s, the 2010s and anything genuinely current are wide open as reference points.

Era and decade referencesWhen people do reach for a time period, it is almost always the eighties.

Plotted on the same axis as the genre chart, where the leading term reaches 11.3%. Auto-scaling this chart to its own maximum would make these counts look far larger than they are.

Table view
TermPromptsShare
1980s232.0%
1990s191.7%
Futuristic151.3%
2000s111.0%
1970s100.9%
1960s60.5%
1950s40.4%
2010s30.3%

Test a prompt while you read

Sonx turns a written prompt into a full track with vocals in about a minute, on your phone. Free to start.

The six-layer formula

The data points at a fix that is almost embarrassingly simple. Write six layers instead of two. It takes a prompt from seven words to about twenty-five, which is still one sentence.

  1. Genre, specificallyNot electronic. UK garage, Memphis phonk, liquid drum and bass.
  2. MoodOne or two words, and pick an unusual one. Menacing, wistful, triumphant, numb.
  3. TempoA number. The single highest-leverage word in the whole prompt.
  4. Two or three instrumentsNamed, not implied. Log drum, felt piano, resonator slide guitar.
  5. The voiceWho is singing, how, or that nobody is. Breathy female vocal. Cracked male tenor. Fully instrumental.
  6. One structural instructionThe clause almost nobody writes. A chorus that lifts. A bridge that drops out. A beat switch under the hook.
Before · 7 words, the median promptlo-fi hip hop, chill, rainy vibes
After · 26 words, same ideaLo fi hip hop, rainy melancholy, 72 BPM, dusty Rhodes chords, vinyl crackle, brushed drums, fully instrumental, with a stripped bridge where the drums drop away

We built the library below to that spec. Here is how it compares to what is published elsewhere.

The three layers most prompts skipShare of prompts that include each layer, in the public corpus versus the Sonx library below.
Public corpusSonx library
Names a tempo 21.6% 100%
Directs the voice 31.3% 96%
Asks for a structure 2.4% 92%

Sonx figures cover the 134 song prompts in the library. Lyrics, voice clone, photo and video prompts use different layers and are excluded.

Table view
LayerPublic corpusSonx library
Names a tempo21.6%100%
Directs the voice31.3%96%
Asks for a structure2.4%92%