How to Write AI Image Prompts That Actually Get You the Image You Want
Wanderson Jackson
Updated July 2026. 9-min read.
TL;DR: Most AI image prompts fail because they are too vague or too cluttered. The fix is a 6-part formula: Subject, Style, Lighting, Composition, Color, and Quality. This guide shows you the formula, explains how different models interpret prompts differently, and gives you ready-to-use templates for product photography, ad creatives, and social media visuals.
Professional photographers and art directors give instructions along six dimensions. AI image generators work the same way. If you leave a dimension blank, the model fills it in with something generic. If you specify all six, you get exactly what you described.
The formula:
Subject + Style and Medium + Lighting and Mood + Composition and Camera + Color and Tone + Quality Modifiers
Every element gives the model a constraint it can act on. The more constraints you provide, the fewer surprises you get.
Here is the difference in practice:
Vague prompt: "A coffee mug on a table"
Structured prompt: "A single matte-black ceramic coffee mug on a light oak dining table, early morning, soft natural window light from camera left casting a long shadow, shot from a 45-degree overhead angle with a 35mm lens, warm neutral tones, shallow depth of field, editorial product photography"
The first prompt gives the model four nouns and no direction. The second gives it a scene, a mood, a camera position, and a style. The second prompt eliminates ambiguity.
Each Element in Detail
1. Subject
The subject is the non-negotiable foundation. Be specific about what is in the image: who or what, where, what they are doing, and how they look.
Vague: "A woman"
Specific: "A 30-year-old woman with shoulder-length dark hair, wearing a cream linen blazer, standing at a kitchen counter pouring oat milk into a ceramic bowl"
The more concrete detail you give, the more control you have. Name materials, colors, actions, and expressions. Avoid abstract nouns like "innovation" or "creativity" since the model cannot render them.
2. Style and Medium
This tells the model how to render the image. Is it a photograph, an illustration, a 3D render, an oil painting? What visual era or movement does it reference?
Examples:
"Editorial product photography, studio lighting, minimalist composition"
"Watercolor illustration with loose brushstrokes, children's book style"
Name a real style or era. The model has learned thousands of visual traditions and will match a named reference faster than a 200-word color description.
3. Lighting and Mood
Lighting is the highest-leverage variable in any prompt. It controls texture, depth, emotion, and whether the image feels flat or cinematic.
Lighting terms that work:
"Soft morning light through linen curtains"
"Dramatic side lighting casting a long shadow"
"Golden hour, warm tones, backlit rim light"
"Single practical lamp casting warm amber light from camera left, deep shadows filling the right side"
Pair a lighting direction with one or two mood words. The model translates emotion into color temperature and contrast.
4. Composition and Camera Angle
This controls spatial organization. Where is the camera? How close? What lens? What is in frame and what is left out?
Effective composition terms:
"Rule of thirds, subject in left third, negative space on right"
"Overhead flat lay, centered, symmetrical"
"Close-up crop from the waist up, shallow depth of field"
"Wide shot, full body, environmental portrait"
"45-degree angle, three-quarter view"
Camera and lens terms that work:
"Shot at eye level with an 85mm portrait lens"
"Low-angle ground perspective captured with a 50mm lens at f/1.4"
"35mm with visible grain"
Naming a focal length gives the model a concrete optical instruction. An 85mm portrait lens compresses background; a 24mm wide lens expands it. The model knows the difference.
5. Color and Tone
Instead of specifying hex codes, describe the palette in natural language. Name the dominant hues and the tonal relationship.
Examples:
"Muted teal and warm amber, low contrast"
"High-contrast black and white with deep blacks"
"Desaturated pastels, soft pink and sage green"
"Rich earth tones, warm brown and burnt orange"
Hex codes produce flat, lifeless output in most models. Natural-language descriptors let the model interpret the mood behind the color choice.
6. Quality Modifiers
These are finishing instructions. They do not fix a weak prompt, but they sharpen a strong one.
"Film grain, analog texture, Kodak Portra 400 look"
"Clean, minimal, no artifacts"
What to avoid: Stacking every modifier you can think of ("masterpiece, best quality, extremely detailed, award-winning, trending on ArtStation"). Modern models respond to specificity, not volume.
How Different Models Handle Prompts
Each AI model has a different prompting style. Writing the same prompt for every model is a common mistake that costs credits and time.
GPT-Image 2 (2 credits/image on Avocado)
GPT-Image 2 excels at text rendering and photorealism. It responds best to natural-language, conversational prompts. Write in full sentences as if you are describing the image to a photographer.
Prompt style: "Create a photorealistic editorial photograph of a minimalist workspace with a white desk, a ceramic mug, and soft natural window light. Shot from a slightly elevated angle."
GPT-Image 2 handles spatial relationships well and produces clean text overlays when requested. It has quality tiers (low, medium, high) that affect both output fidelity and credit cost.
Nano Banana 2 (1 credit/image on Avocado)
The default model on Avocado. Fast, clean detail, and the cheapest option at 1 credit per image. Good for rapid iteration.
Prompt style: Works well with concise, directive prompts. Specify the key elements without overloading the prompt.
Ideogram V3 (2 credits/image on Avocado)
Best for on-image text rendering. If your image needs legible text (logos, posters, signage), Ideogram handles it reliably.
Prompt style: Be explicit about text placement and font style. "A minimalist poster with the headline 'Summer Sale' in bold sans-serif at the top, clean white background."
Recraft V4 (1 credit/image on Avocado)
Design-focused with native style control and SVG output options. Good for logos, icons, and vector-style product renders.
Prompt style: Specify the design style clearly. "Flat vector illustration of a coffee cup, minimal lines, muted palette, isolated on white background."
Seedream V5 Pro (2 credits/image on Avocado)
ByteDance model with native text rendering in 14 languages and strong dense-layout control.
Prompt style: Short, precise prompts outperform long ornate descriptions. Be direct.
Prompt style: Name the cinematic reference. "Cinematic still from a 1970s Wes Anderson film, symmetrical composition, pastel palette."
Prompt Templates by Use Case
Product Photography
"A single [product] on a [surface], shot from a [angle], [lighting direction], [background], [style]. Example: "A single white leather sneaker on a seamless light-gray studio background, shot from a 45-degree angle, dramatic side lighting casting a soft shadow, ultra-sharp product photography, minimalist composition, no text, no logos"
Ad Creative / Social Media Visual
"A [subject] in [setting], [action], [lighting], [mood], [style]. Example: "A young woman in a sunlit cafe laughing while looking at her phone, golden afternoon light through large windows, warm and candid energy, editorial lifestyle photography, shot on 35mm"
UGC-Style Content
"[Subject] [action] in [realistic setting], [natural lighting], [casual framing]. Example: "A person holding a skincare bottle in a bathroom mirror selfie, fluorescent overhead light, casual phone camera angle, slightly grainy, realistic apartment setting"
Brand or Logo Design
"A [style] logo for [brand concept], [color approach], [background]. Example: "A minimal geometric logo for an organic coffee brand, earth tones, clean white background, vector style, no text"
Common Mistakes and How to Fix Them
Mistake 1: Too vague
Problem: "A beautiful landscape"
Fix: Name the specific landscape, time of day, weather, and camera angle. "A misty mountain valley at dawn, golden light breaking through low clouds, shot from a hillside with a 70mm telephoto lens, atmospheric depth"
Mistake 2: Too many subjects
Problem: "A woman, a dog, a car, a city skyline, and a sunset all in one image"
Fix: Limit each prompt to one or two subjects. If you need a complex scene, break it into layers and use an image editor to composite.
Mistake 3: Using hex color codes
Problem: "Background color #BDEE63"
Fix: Describe the color in natural language. "Soft lime green with warm undertones" gives the model room to interpret the color in context rather than locking it to a single flat value.
Fix: Pick two or three that matter. "Ultra-sharp, commercial photography quality, no grain" is more effective than a wall of adjectives.
Mistake 5: Ignoring the negative
Problem: You get text, watermarks, or extra fingers you did not ask for.
Fix: Add explicit exclusions. "No text, no watermarks, no logos, no brand names" at the end of the prompt. On Stable Diffusion models, use a separate negative prompt field.
Mistake 6: One prompt for all models
Problem: You write a long descriptive prompt and use it for GPT-Image 2, Ideogram, and Recraft.
Fix: Adjust the prompt length and style for each model. GPT-Image 2 responds to conversational paragraphs. Ideogram needs explicit text instructions. Recraft needs design-style keywords. Copy-pasting the same prompt across models wastes credits on iterations.
What Actually Matters
The difference between a mediocre AI image and a professional one is not the model or the credits you spend. It is the prompt. A 1-credit image from Nano Banana 2 with a well-structured prompt will outperform a 4-credit image from a premium model with a vague one.
Focus on three things:
Be specific about the subject. Name materials, colors, actions, and expressions. The model cannot guess what you are picturing.
Name the lighting. Lighting direction and quality are the highest-leverage prompt elements. A single phrase like "soft morning window light" can transform a flat image into something editorial.
Iterate in small moves. Change one element per round. If the subject is right but the lighting is wrong, adjust the lighting and leave everything else alone. Rewriting the entire prompt each time introduces new variables and makes it impossible to learn what works.
FAQ
How long should an AI image prompt be?
Between 30 and 80 words is the sweet spot for most models. Shorter prompts leave too much to the model's interpretation. Longer prompts (100+ words) can work for complex scenes but often produce diminishing returns. Start with 40-60 words covering subject, style, lighting, and composition.
Do I need different prompts for different AI models?
Yes. Each model interprets prompts differently. GPT-Image 2 responds well to natural-language paragraphs. Stable Diffusion models prefer structured keyword lists with weights. Ideogram needs explicit text placement instructions. Recraft responds to design-style terminology. Using the same prompt for all models wastes credits.
What is a negative prompt and when should I use one?
A negative prompt tells the model what to exclude from the image: "no text, no watermarks, no extra fingers, no brand logos." Not all models support separate negative prompt fields. Stable Diffusion and Flux models have native negative prompt support. For models without a dedicated field, add exclusions at the end of the main prompt.
How do I get consistent style across multiple images?
Use the same style and lighting description in every prompt. Change only the subject. For example, if you want all your product photos to look the same, fix the lighting ("dramatic side lighting"), the background ("seamless white studio"), and the camera angle ("45-degree overhead") and only swap the product description.
Can AI image generators render text accurately?
It depends on the model. GPT-Image 2, Ideogram V3, and Seedream V5 Pro handle text rendering well. Most other models produce garbled or misspelled text. If your image needs legible on-image text, use a model specifically designed for it.
How many credits does it cost to iterate on a prompt?
Each generation costs the model's per-image rate, regardless of whether it is a first attempt or a tenth iteration. On Avocado AI, Nano Banana 2 costs 1 credit per image, GPT-Image 2 costs 2 credits, and Recraft V4 costs 1 credit. Budget 3-5 generations per image to get the result you want. Use the cheapest model for iteration and the premium model for the final output.
How to Pick the Right Approach in Under 30 Seconds
Need photorealism or on-image text? Use GPT-Image 2 or Ideogram V3.
Iterating fast on a concept? Use Nano Banana 2 at 1 credit per image.
Building brand assets, logos, or vector graphics? Use Recraft V4.
Writing for a global audience with multilingual text? Use Seedream V5 Pro.
Creating cinematic, atmospheric visuals? Use Krea 2 Large.
Working across image, video, and audio in one workspace? Use Avocado AI.
If you want a single workspace where you can generate images, video, music, and sound effects with 15 image models and 7 video models in one credit pool, start with Avocado AI. Plans start at EUR 19.99/month.
Wanderson Jackson is the founder of Avocado AI, a creative workspace for AI-generated images, video, music, and audio. He writes about AI media tools and creative workflows.