Suno AI Prompts: The Two-Box Rule, and the Part Nobody Agrees On
A Suno AI prompt fails for one reason more than any other. Everything gets written in one place, in one style, as one long line of adjectives. Split it, and the same idea starts landing.
Two boxes. Two jobs. The Style box holds genre, vocal type, tempo, and production. The Lyrics box holds the words that get sung — and, in square brackets, the directions for how those words are performed. Bracketed text is an instruction. Everything outside the brackets gets sung. That is the entire mechanism. Most of what follows is consequence.
Here is the shape that works:
Style
dream pop, ethereal female vocal, lush analog pads, shimmering guitar, 92 BPM, wide stereo, nostalgic
Lyrics
[Intro: soft synth pads, gentle atmospheric build]
[Verse 1: whispered vocal delivery, minimal instrumentation]
City lights are moving under me
[Chorus: explosive, soaring vocals, full production]
We are alive right now
Now the fight.
The separator nobody agrees on
Everyone agrees that piling adjectives behind commas produces mush. Past that, the advice splits three ways, and the three are not compatible.
One camp separates elements inside a single bracket with a pipe: [Chorus | powerful | soaring female vocal | wall of sound | wide reverb | modern pop mix]. The pipe acts as a hard stop. This instruction ended. Read the next one fresh.
Another camp uses line breaks instead, three to five instructions per tag, one per line.
A third keeps commas and simply shortens the list.
The pipe camp has the sharper argument, and it carries two rules the others never state. First, order carries weight, left to right. Section, then mood, then vocal style, then key instruments, then dynamics, then spatial effects, then production style. Lay the foundation before choosing the paint. If the vocal is the point, the vocal instruction goes early. If the instruments lead, they go early. Burying the important element at the end of the stack is the same as deleting it.
Second, there is a ceiling, and it is the number worth remembering: seven. Past seven elements in one tag, the list stops being a checklist and turns back into noise. You are exactly where you started.
Three ways the method breaks:
- More than seven elements in a single tag.
- Contradictory elements in the same stack. Minimalist and wall of sound cannot both win. When two instructions fight, generic is what you get.
- Tags that argue with the Style box. An acoustic folk brief with a heavy distortion tag leaves the model choosing between two directions. Tags reinforce the genre. They do not override it.
Where the disagreement actually matters
Duets
One method says reinforce it in three places: name the duet and both singers in Style, put a header at the top of the lyrics, then label each part inside the lyrics. Another says Style plus labeled parts is enough. Both carry the same warning — alternate the voices line by line and the model starts flipping them. Give each singer full verses. Save the shared label for the chorus.
Live versions
One method puts the crowd direction only in the lyrics: [Intro: stadium crowd ambience, big applause, distant chanting, stage reverb], and notes it works in the outro too. Another insists the lyrics tag alone is not enough and that Style has to open with “live recording at a concert”. A third route skips prompting entirely — take the finished track into Cover, add the live tags there, and let Audio Influence decide how much of the original survives.
Audio Influence
Two recommended numbers for the same slider: 55 and 70. The reasoning behind them matters more than either number. High keeps your song and dilutes the new element. Low pushes the new element hard and starts rewriting the song underneath it. Decide which loss you can live with, then set it.
Exclude Styles
Some treat it as a plain negative prompt — paste in what you do not want. Others say leading with it is the wrong order of operations: fix the Style box first, and only for delicate cases — no instruments at all, an acapella, a track with no guitar — repeat the instruction in brackets at the very top of the lyrics so the model meets it twice. There is also an unresolved argument about whether to prefix a minus sign or write “no guitar”, on the theory that a negation wrapped in a negation reads as its opposite. The field holds 1,000 characters. Length was never the constraint. Reliability is.
The fixes that only surface once
Your taste profile is a cage
Suno keeps a written profile of your musical style and uses it to shape the Style text it generates for you. The field cannot be left empty. When every track starts sounding like the last one, rewrite that profile with something neutral, save it, and switch the style personalization off. The original is still stored — reverting brings it back.
Punctuation is a delivery instruction
A comma is a short breath. A period ends the phrase. A colon sets up what comes next. Three dots buy a longer, more dramatic silence. Change the punctuation and the performance changes, without touching a single tag.
Lock the tempo in three places
Slow tracks expose drift that busy ones hide. Write the BPM in Style. Write it again above the lyrics. Repeat it at the top of every section with an explicit “constant 50 BPM, no build-up, no rhythm increase”. Remind it before each verse, and again before the outro.
Spoken word is for short lines. Monologue is for the dramatic ones
A short spoken phrase sits fine under a spoken-word tag. Cinematic narration — opening, middle, or closing — reads better under monologue.
Clean the voice before you hand it over
Record in a DAW, remove the noise, set the level, then upload. A clean recording still passes the voice check and gives the model far more to work with than a raw phone memo.
Two signals for a foreign-language line
The instruction goes in brackets. The actual words go in parentheses. [female French vocal, soft, perfect French accent] followed by (the French phrase). One says how to perform it. The other says what to say.
Fix the hum after Suno, not inside it
Export the WAV. Open it in Audacity, select a quiet passage, zoom in three or four times, and take a noise profile from a small slice of pure hiss. Then select the whole track and run noise reduction at 6, sensitivity 4, frequency smoothing 3. Gentle wins. Push it harder and you lose the top end, flatten the vocal, and the whole thing turns metallic.
Say where you want the glitch
Broken on purpose is a sound. Ask for corrupted synth noise and a broken robot voice at the top, digital stutters or chopped vocals in the middle, a glitch drop or distorted transition right before the chorus. Place it. Do not hope for it.
Numbers worth keeping
- Style field: 1,000 characters. Exclude Styles: 1,000. Lyrics: 5,000.
- Past roughly 3,000 characters of lyrics, tracks start rushing and skipping sections, even though the box allows more.
- One generation runs up to 8 minutes.
- Since 9 September 2026 the only models are v6, v6-wild and v6-mini. Any guide telling you to pick v5 is describing a menu that no longer exists.
- Weirdness has a documented reference point at 50%. Audio Influence only appears once you upload audio. Style Influence governs how strictly the brief is followed — raise it when one instruction is non-negotiable.
The short version
Front-load. Split the fields. Keep a stack to seven. Never let a tag argue with your Style box. And when the output is close but wrong in one place, fix that place — replace that section, move that block, duplicate that chorus — instead of rolling the dice again.