How to Make AI Music With Text Prompts: A Step-by-Step Guide
Wanderson Jackson
Updated: July 2026 | Author: Wanderson Jackson
TL;DR: Type a description, get a track. Modern AI music generators like Suno, Udio, and ElevenLabs Music turn plain-language prompts into full songs in seconds. The key to good output is a structured prompt: name the genre, mood, tempo, and instruments. This guide walks through the process, from picking a tool to refining prompts to integrating the result into your creative workflow.
Text-to-music AI takes a written description and converts it into audio. You type something like "lo-fi hip hop, warm vinyl texture, 85 BPM, piano and muffled drums" and the model produces a track that matches.
Most tools generate 30 seconds to 4 minutes per pass. You can extend, remix, or regenerate sections until the output fits your project. The quality varies by tool, but the gap between AI music and stock library music has narrowed significantly in 2026.
The main use cases:
Video creators need background tracks without licensing headaches.
Podcasters want intros, outros, and transitions that match their tone.
Advertisers need quick-turnaround tracks for campaigns.
Game developers want ambient and action loops at scale.
Musicians use AI as a starting point, then layer their own production on top.
The key shift: you do not need to read music, play an instrument, or open a DAW. A well-written text prompt gets you 80% of the way there.
Step 1: Pick Your Tool
Not all AI music generators work the same way. Some focus on full songs with vocals. Others specialize in instrumental background tracks. Pick the tool that matches your use case.
Suno
The most popular text-to-music generator as of mid-2026. Suno creates complete songs with vocals, lyrics, and multi-layer arrangements from a single prompt.
Model: v5.5 (latest, behind Pro/Premier paywall; v4.5-all on free tier)
Free tier: 50 credits daily, up to 10 songs per day, no commercial use
Premier: $30/month ($24/month annual), 10,000 credits/month, Suno Studio access
Strengths: Best style adherence. Handles genre blends well. Built-in Studio editor for post-generation editing.
Trade-offs: Commercial rights only apply to songs generated while a paid plan is active (not retroactive from free tier). Copyright ownership is not guaranteed. Some Content ID flags reported.
Best for: Full songs with vocals, demos, theme music, intros
Udio
Suno's main competitor. Known for higher instrumental fidelity and section-level editing.
Free tier: 10 credits/day plus 100 monthly backup credits, up to 3 songs/day
Strengths: Regenerate specific sections (chorus, bridge) without re-rendering the whole track. Better instrumental detail than Suno in many genres.
Trade-offs: Ongoing litigation with Sony Music may affect the platform. Style adherence can be inconsistent. Sometimes ignores lyrics entirely.
Best for: Long-form projects, releases, experimental styles
ElevenLabs Music
An extension of the ElevenLabs voice platform. The music generator emphasizes clean output and copyright-cleared content.
API pricing: $0.15 per minute of generated music
Subscription plans: Starter $6/month, Creator $22/month, Pro $99/month (credit-based for all ElevenLabs features)
Strengths: Copyright-cleared output (marketed as trained on licensed data). Built-in audio editing. Style exclusion controls.
Trade-offs: Decent but not perfect style adherence. Song structure can feel awkward. Higher per-minute cost than credit-based competitors.
Best for: Commercial projects where copyright clearance matters
AIVA
A classical and cinematic music generator focused on instrumental compositions.
Free tier: 3 downloads/month, track durations up to 3 minutes, non-commercial use, credit required
Standard: EUR 11/month (annual), 15 downloads/month, up to 5 minutes, limited monetization (YouTube, Twitch, TikTok, Instagram)
Pro: EUR 33/month (annual), 300 downloads/month, full copyright ownership, all file formats including WAV
Strengths: 250+ style presets. MIDI and audio influence uploads. Full copyright ownership on Pro plan.
Trade-offs: No vocal generation. Focused on cinematic/orchestral styles; less suited for pop, hip-hop, or electronic.
Best for: Film scores, game audio, trailers, corporate videos
Mubert
Sample-based background music generation. Fast output, royalty-free.
Strengths: Generates tracks in under 10 seconds. Lengths from 5 seconds to 25 minutes. Preset-based and prompt-based generation.
Trade-offs: Instrumental only. Less creative variation than Suno or Udio. Output can feel generic.
Best for: Background music for explainers, tutorials, live streams, apps
Avocado AI Music/Audio Studio
Avocado AI includes a Music/Audio Studio as part of its creative workspace. You can generate tracks from text prompts alongside images, video, and sound effects in the same project.
Strengths: Music generation sits inside the same workspace as image and video generation. You can generate a track, create visuals, and build a video ad without switching platforms. Commercial usage rights included on all plans.
Trade-offs: The music generator is a workspace feature, not a dedicated music platform. For advanced music production (stem separation, section-level editing, vocal generation), a dedicated tool like Suno or Udio gives more control.
Best for: Creators who want music, images, and video in one workspace
Step 2: Write a Structured Prompt
The biggest mistake new users make: writing prompts like a search query instead of a creative brief. "Make me a chill song" gives the AI too much room to guess. A structured prompt pins down the output.
The Four-Part Framework
Every good AI music prompt answers four questions:
Component
What It Does
Example
Genre + Subgenre
Sets the primary style
"1980s synthwave" or "boom bap hip hop"
Mood / Emotion
Defines the feeling
"Nostalgic and bittersweet" or "aggressive, dark"
Tempo
Controls the speed
"128 BPM" or "slow ballad, 70 BPM"
Key Instruments
Anchors the sound palette
"Analog synth, gated reverb drums, deep bass"
Format: Structured Tags, Not Prose
AI music models parse keywords better than full sentences. Use comma-separated descriptors:
Bad prompt:
"Create an upbeat synthwave track with female vocals that sounds like something from the 80s."
Most text-to-music tools (Suno and Udio especially) support bracketed meta tags in the lyrics field. These tags tell the model where to place verses, choruses, and instrumental breaks.
Common Meta Tags
Tag
Purpose
[Intro]
Instrumental opening section
[Verse]
Standard verse with lyrics
[Pre-Chorus]
Build-up before the chorus
[Chorus]
Main hook, usually repeated
[Bridge]
Contrasting section for variety
[Instrumental]
No-vocal section for solos or breaks
[Outro]
Closing section
How to Use Them
In tools like Suno, switch to "Custom" mode and enter lyrics with tags:
[Intro] Soft piano, building strings
[Verse 1]
The morning light falls through the glass
A quiet start, the moments pass
[Chorus]
Hold on to the feeling
Hold on to the sound
[Instrumental] Synth solo, building intensity
[Verse 2]
The city wakes, the streets alive
Another day to feel the drive
[Chorus]
Hold on to the feeling
Hold on to the sound
[Outro] Fade out, reverb tail
Meta tags are significantly more powerful than relying on the style prompt alone. They give you structural control that prose descriptions cannot match.
Step 4: Generate and Iterate
AI music generation is not one-shot. Expect to generate 3 to 5 versions before landing on something usable. Budget your credits accordingly.
The Iteration Loop
Generate v1 with your structured prompt. Listen for overall direction, not polish.
Identify what is wrong. Is the genre off? Are the vocals wrong? Is the tempo too fast?
Change one variable at a time. If you change genre, mood, AND tempo between generations, you cannot tell which change fixed (or broke) the output.
Generate v2 with the adjusted prompt.
Repeat until the core direction is right. Fine-tuning happens after.
Credit Budgeting
Each tool uses credits differently:
Tool
Credits Per Song
Approximate Cost Per Song
Suno (Pro)
~5 credits
~$0.02
Udio (Standard)
~10 credits
~$0.04
Avocado AI
4 credits per 30s block
~EUR 0.13 (Intro) to ~EUR 0.05 (Growth)
ElevenLabs Music
Per-minute billing
$0.15/minute via API
At Suno's Pro tier ($10/month for 2,500 credits), you can generate roughly 500 songs per month. That is more than enough for heavy iteration.
When to Stop Iterating
Stop when the track serves its purpose. A background track for a YouTube video does not need to be perfect. A demo for a client pitch does. Match your iteration effort to the stakes.
Step 5: Refine With Prompt Engineering
Once you have the basic direction right, these techniques improve output quality:
The Sandwich Method
Place your most important descriptors first AND last in the prompt. AI models weight early and late tokens more heavily:
Synthwave, analog synths, drum machine, neon aesthetic, 1985 production, synthwave
Negative Prompting
Explicitly exclude unwanted elements. Most tools support some form of negative input:
Style: Dark ambient, atmospheric, cinematic
Avoid: Vocals, bright sounds, major keys, fast tempo
Anchor-and-Shift for Series
When creating multiple tracks for the same project (a podcast series, a video campaign), keep one anchor parameter stable and change only one variable per track:
This maintains consistency across a series without every track sounding identical.
Fixing Common Problems
Problem
Fix
Wrong genre output
Add era and subgenre: "1990s country, Nashville sound" instead of just "country"
AI ignores vocal instructions
Make vocals the first descriptor: "Female vocalist, pop, breathy soprano"
Repetitive intro
Use meta tag: "[Intro] Solo violin, no drums for first 15 seconds"
Lyrics do not match theme
Switch to custom/manual mode and provide complete lyrics with structure tags
Output sounds generic
Add specific production details: "gated reverb snare, sidechained bass, tape saturation"
Step 6: Export and Use
Once you have a track you are happy with, the final steps depend on your workflow.
Standalone Music Tools
If you generated in Suno, Udio, or AIVA:
Download the audio file (MP3 on free tiers, WAV on paid plans).
Check licensing. Suno and Udio grant commercial rights only on paid plans, and only for songs generated while the plan is active. AIVA Pro gives full copyright ownership. ElevenLabs Music is marketed as copyright-cleared.
Import into your project (video editor, DAW, podcast host).
The track lives in your Workspace alongside images, video clips, and other assets.
Drag the track into a video project, pair it with AI-generated visuals, and export without leaving the platform.
Commercial usage rights are included on all plans from EUR 19.99/month.
This is the main advantage of an integrated workspace: you do not need to download, re-upload, and sync assets across three different tools. The music, visuals, and final output all live in the same project.
Tool Comparison
Tool
Vocals
Free Tier
Paid From
Commercial Rights
Copyright Ownership
Best For
Suno
Yes
50 credits/day
$10/month
Paid plans only
Not guaranteed
Full songs with vocals
Udio
Yes
10 credits/day
$10/month
Paid plans only
Not guaranteed
Long-form, experimental
ElevenLabs Music
Yes
Limited
$6/month
Paid plans
Copyright-cleared
Commercial projects
AIVA
No
3 downloads/month
EUR 11/month
Pro plan only
Pro plan only
Cinematic, orchestral
Mubert
No
Yes
$11.69/month
Paid plans
N/A (royalty-free)
Background music
Avocado AI
Via TTS
No
EUR 19.99/month
All plans
Usage rights included
Full creative workspace
Worked Example: Creating a 60-Second Ad Track
Here is a real workflow for creating background music for a product ad.
Goal: Upbeat, modern electronic track for a 60-second skincare product video.
The Prompt
Modern electronic pop, upbeat and fresh, 120 BPM, bright synth pads,
Generated v1 in Suno. Direction was right but the intro was too long. Added meta tag: [Intro] Short, 4 seconds, synth swell into beat.
Generated v2. Better structure, but the bass was too heavy. Changed "subtle bass" to "light sub-bass, minimal."
Generated v3. Good. Downloaded the 2-minute version.
Trimmed to 60 seconds using basic audio editing (crop at the 1-minute mark, add a fade-out).
Exported as WAV for the video editor.
Total credits used: ~15 (5 generations). Total time: 8 minutes.
Adapting for Avocado AI
If you are working in Avocado AI, the same prompt goes into the Music/Audio Studio. Generate the track, then open a video project in the same workspace and drop the audio onto the timeline alongside your product visuals. No export-import cycle needed.
FAQ
Do I need to know music theory to use AI music generators?
No. Text-to-music tools are designed for people with zero musical training. A prompt like "happy acoustic guitar, 100 BPM, warm and sunny" works without any theory knowledge. You do need to describe what you want in concrete terms, but that is creative direction, not music theory.
Can I use AI-generated music commercially?
It depends on the tool and your plan. Suno and Udio grant commercial rights only on paid plans, and only for songs generated while the plan is active. AIVA grants full copyright ownership on its Pro plan (EUR 33/month). ElevenLabs Music is marketed as copyright-cleared. Avocado AI includes commercial usage rights on all plans. Always check the current terms of service before publishing.
How long does it take to generate a song?
Most tools produce a 1 to 4 minute track in 10 to 60 seconds. Suno and Udio typically generate in under 30 seconds. The time investment is in iteration, not generation. Expect to spend 5 to 15 minutes refining prompts and listening to variations.
Can AI music generators create lyrics too?
Yes. Suno and Udio can generate both the instrumental and the lyrics from a single prompt. You can also write your own lyrics and have the AI set them to music. For best results, provide lyrics in custom/manual mode with structure tags ([Verse], [Chorus], etc.).
What is the difference between text-to-music and AI-assisted music production?
Text-to-music generates a complete track from a text prompt (Suno, Udio). AI-assisted production uses AI tools within a traditional DAW to help with specific tasks like mixing, mastering, stem separation, or melody suggestions. The two approaches serve different needs: text-to-music is faster for first drafts; AI-assisted production gives more control for polished final output.
How do I keep a consistent sound across multiple tracks?
Use the anchor-and-shift method. Keep one parameter stable (genre, tempo, or a specific instrument) and change only one variable per track. This maintains a cohesive sound across a series without every track being identical.
Can I upload my own audio to guide the AI?
Some tools support this. AIVA allows MIDI and audio influence uploads. Suno Pro lets you record and upload your own voice. Udio supports audio uploads for remixing. Check your tool's documentation for current upload capabilities.
Is AI music copyright-free?
Not automatically. Copyright status depends on the tool, your plan, and your jurisdiction. AIVA Pro gives you full copyright ownership. Suno and Udio do not guarantee copyright vesting. ElevenLabs Music claims copyright-cleared output. The legal landscape around AI-generated music is still evolving. For high-stakes commercial projects, consult a licensing professional.
If you want one workspace for music, images, and video, start with Avocado AI. Plans start at EUR 19.99/month with commercial usage rights included. Generate a track, build visuals, and export a finished video without switching platforms.
Wanderson Jackson is the founder of Avocado AI, a creative workspace for AI-generated images, video, and music.