- Prompts are short. 768 of 1,136 prompts, 68% of the corpus, are ten words or fewer.
- Genre is doing all the work. Pop appears in 128 prompts, more than any other genre family. Most prompts name a genre and stop.
- Tempo is mostly missing. Only 245 prompts give a BPM. Of those that do, the median is 100 and almost all sit between 70 and 140.
- The voice is an afterthought. 69% of prompts say nothing at all about who is singing, or whether anyone is.
- Nobody asks for a chorus. 20 prompts mention a chorus or hook. 3 mention a verse. 1 mentions a bridge.
What people actually write
The first thing almost every prompt does is name a genre, and for a lot of them it is the only thing they do. Pop leads, which is unsurprising. The interesting part is how flat the tail is: cinematic scoring sits third, ahead of rock and jazz, which says something about what people are actually making. Background music for video, not songs.
Table view
| Term | Prompts | Share |
|---|---|---|
| Pop | 128 | 11.3% |
| Hip hop / Rap | 104 | 9.2% |
| Cinematic | 94 | 8.3% |
| Electronic / EDM | 93 | 8.2% |
| Rock | 91 | 8.0% |
| Jazz | 81 | 7.1% |
| Classical | 81 | 7.1% |
| Folk / Acoustic | 80 | 7.0% |
| Ambient | 79 | 7.0% |
| R&B / Soul | 69 | 6.1% |
| Lo-fi | 66 | 5.8% |
| Metal | 58 | 5.1% |
| Reggae / Dub | 36 | 3.2% |
| Country | 34 | 3.0% |
A genre word is a strong instruction and a blunt one. It sets instrumentation, tempo range and production style all at once, which is why it works. It also means two people writing lo fi hip hop get near-identical results, because they have handed over every other decision.
Roughly a third of prompts add an emotional word on top. The vocabulary there is small.
Table view
| Term | Prompts | Share |
|---|---|---|
| Chill / Relaxed | 84 | 7.4% |
| Warm | 74 | 6.5% |
| Dark | 65 | 5.7% |
| Romantic | 61 | 5.4% |
| Nostalgic | 46 | 4.0% |
| Energetic | 45 | 4.0% |
| Dreamy | 40 | 3.5% |
| Aggressive | 36 | 3.2% |
| Epic | 35 | 3.1% |
| Uplifting | 29 | 2.6% |
| Melancholic | 26 | 2.3% |
| Mysterious | 24 | 2.1% |
Melancholic appears in 26 prompts out of 1,136. Sad songs are most of what people want and almost none of what they ask for.
Instrumentation stops at the rhythm section. Drums and bass dominate, then guitar, synth and piano, and after that it drops off a cliff. Saxophone, organ, arpeggios and handclaps together appear in fewer prompts than piano alone. This is the layer with the most room in it.
Table view
| Term | Prompts | Share |
|---|---|---|
| Drums | 163 | 14.3% |
| 808 / bass | 137 | 12.1% |
| Guitar | 111 | 9.8% |
| Synth | 108 | 9.5% |
| Piano | 104 | 9.2% |
| Strings | 47 | 4.1% |
| Trumpet / brass | 42 | 3.7% |
| Pads | 39 | 3.4% |
| Vinyl / tape texture | 24 | 2.1% |
| Saxophone | 13 | 1.1% |
| Organ | 12 | 1.1% |
| Arpeggio | 10 | 0.9% |
Only one prompt in five names a tempo
This is the finding that surprised us most. Tempo is the single fastest lever on how a track feels, it takes three characters to specify, and 78% of published prompts leave it out. Among the 245 that do name one, the distribution is tight.
Bins in BPM. Values below 2 prompts omitted for readability.
Table view
| Bin | Prompts | Share |
|---|---|---|
| 50-59 | 6 | 2.4% |
| 60-69 | 10 | 4.1% |
| 70-79 | 27 | 11.0% |
| 80-89 | 34 | 13.9% |
| 90-99 | 38 | 15.5% |
| 100-109 | 26 | 10.6% |
| 110-119 | 26 | 10.6% |
| 120-129 | 35 | 14.3% |
| 130-139 | 17 | 6.9% |
| 140-149 | 17 | 6.9% |
| 150-159 | 3 | 1.2% |
| 160-169 | 3 | 1.2% |
| 170-179 | 2 | 0.8% |
Seven words
Half of all published prompts are seven words or shorter. That is a genre, a mood and maybe one instrument. It is not enough information to specify a song, so the model fills the gap with whatever is most statistically ordinary for that genre, which is exactly why so much AI music sounds the same.
Table view
| Bin | Prompts | Share |
|---|---|---|
| 1-5 words | 411 | 36.2% |
| 6-10 words | 357 | 31.4% |
| 11-15 words | 129 | 11.4% |
| 16-20 words | 161 | 14.2% |
| 21-30 words | 64 | 5.6% |
| 31+ words | 14 | 1.2% |
The voice is the biggest blind spot
356 prompts out of 1,136 say anything about vocals at all. That leaves 69% where the most noticeable element of a song, whether there is a singer and what they sound like, is left entirely to chance.
Table view
| Term | Prompts | Share |
|---|---|---|
| Mentions vocals at all | 356 | 31.3% |
| Instrumental / no vocals | 120 | 10.6% |
| Male vocals | 66 | 5.8% |
| Female vocals | 65 | 5.7% |
| Layered / harmonies | 37 | 3.3% |
| Soft / breathy | 28 | 2.5% |
| Rapped | 12 | 1.1% |
| Raspy / gritty | 9 | 0.8% |
Worth noting that 120 prompts explicitly ask for no vocals, which is a real instruction and a good one. The problem is the silent majority: the prompts that neither ask for a voice nor rule one out.
Almost nobody asks for a song
This is the finding with the biggest practical payoff. Out of 1,136 prompts, 20 mention a chorus or a hook, 3 mention a verse and 1 mentions a bridge.
Plotted on the same axis as the genre chart, where the leading term reaches 11.3%. Auto-scaling this chart to its own maximum would make these counts look far larger than they are.
Table view
| Term | Prompts | Share |
|---|---|---|
| Chorus | 20 | 1.8% |
| Drop / breakdown | 11 | 1.0% |
| Intro / outro | 5 | 0.4% |
| Verse | 3 | 0.3% |
| Bridge | 1 | 0.1% |
People describe a sound, then wonder why they got a loop instead of a song.
A texture prompt produces a texture. If you want a track that goes somewhere, the prompt has to say where. A chorus that lifts an octave. A bridge that drops to near silence. A beat switch under the hook. These are single clauses and they change the output more than any adjective will.
Era references are rare too, under 7% of prompts, and heavily weighted to the eighties. Which means the 2000s, the 2010s and anything genuinely current are wide open as reference points.
Plotted on the same axis as the genre chart, where the leading term reaches 11.3%. Auto-scaling this chart to its own maximum would make these counts look far larger than they are.
Table view
| Term | Prompts | Share |
|---|---|---|
| 1980s | 23 | 2.0% |
| 1990s | 19 | 1.7% |
| Futuristic | 15 | 1.3% |
| 2000s | 11 | 1.0% |
| 1970s | 10 | 0.9% |
| 1960s | 6 | 0.5% |
| 1950s | 4 | 0.4% |
| 2010s | 3 | 0.3% |
Test a prompt while you read
Sonx turns a written prompt into a full track with vocals in about a minute, on your phone. Free to start.
The six-layer formula
The data points at a fix that is almost embarrassingly simple. Write six layers instead of two. It takes a prompt from seven words to about twenty-five, which is still one sentence.
- Genre, specificallyNot electronic. UK garage, Memphis phonk, liquid drum and bass.
- MoodOne or two words, and pick an unusual one. Menacing, wistful, triumphant, numb.
- TempoA number. The single highest-leverage word in the whole prompt.
- Two or three instrumentsNamed, not implied. Log drum, felt piano, resonator slide guitar.
- The voiceWho is singing, how, or that nobody is. Breathy female vocal. Cracked male tenor. Fully instrumental.
- One structural instructionThe clause almost nobody writes. A chorus that lifts. A bridge that drops out. A beat switch under the hook.
We built the library below to that spec. Here is how it compares to what is published elsewhere.
Sonx figures cover the 134 song prompts in the library. Lyrics, voice clone, photo and video prompts use different layers and are excluded.
Table view
| Layer | Public corpus | Sonx library |
|---|---|---|
| Names a tempo | 21.6% | 100% |
| Directs the voice | 31.3% | 96% |
| Asks for a structure | 2.4% | 92% |