How to Write AI Video Prompts That Actually Work: A Practical Guide
Wanderson Jackson
Updated: August 2026 | Reading time: 12 min
TL;DR: The difference between a usable AI video clip and a wasted generation comes down to prompt structure. This guide breaks down the five elements every video prompt needs, shows you templates for common use cases, and explains how to adapt your prompts across different AI video models.
Most AI video generators produce mediocre output not because the model is bad, but because the prompt is vague. "A woman walking in a city" gives the model too many decisions to make: what does she look like? What city? What time of day? What camera angle? What mood?
The fix is not writing longer prompts. It is writing more specific ones. A 40-word prompt with the right structure beats a 200-word prompt that reads like a novel.
This guide teaches you the framework.
The 5 Elements of an Effective Video Prompt
Every high-quality AI video prompt contains some combination of these five elements. You do not always need all five, but the more you specify, the less the model guesses.
Element
What it controls
Example
Subject
Who or what is in the frame
"A female barista in her mid-20s wearing a black apron"
Action
What happens in the clip
"Pours steamed milk into a latte, creating a rosetta pattern"
Setting
Where and when
"Inside a sunlit cafe with exposed brick walls, morning light"
Camera
Angle, movement, lens feel
"Close-up, slow push-in, shallow depth of field"
Mood / Style
Emotional tone and visual aesthetic
"Warm, calm, documentary feel"
A prompt that includes all five reads like a one-sentence film direction. That is the goal: give the AI model a clear scene to render, not a vague concept to interpret.
Quick formula:
[Subject] [Action] in [Setting]. [Camera direction]. [Mood/Style].
Example: "A young chef in a white coat carefully plates a dessert in a dimly lit restaurant kitchen. Medium shot, slight camera drift to the right. Moody, cinematic, warm tungsten lighting."
Step-by-Step: Writing a Video Prompt from Scratch
Here is the process for going from a rough idea to a working prompt.
Step 1: Define the single action
AI video clips are short (typically 5 to 10 seconds). One action is enough. Do not try to pack a narrative arc into a single generation.
Bad: "A woman goes to the store, buys flowers, walks home, and arranges them in a vase."
Good: "A woman arranges fresh tulips in a ceramic vase on a sunlit kitchen counter."
Step 2: Name your subject specifically
Generic subjects produce generic output. The more detail you give about the person or object, the more consistent and interesting the result.
Vague: "A person typing on a laptop."
Specific: "A man in his 30s with short dark hair, wearing a gray crewneck sweater, typing on a silver MacBook at a standing desk."
You do not need to describe every physical feature. Pick two to three distinctive details and the model fills in the rest coherently.
Step 3: Set the scene with sensory detail
Setting is not just location. It is lighting, time of day, weather, and atmosphere. These details dramatically change the output.
Flat: "In an office."
Layered: "In a modern open-plan office with large windows, late afternoon golden light streaming in, slight lens flare."
Step 4: Direct the camera
Camera instructions are the most underused element in video prompts. Even simple directions like "close-up" or "wide shot" change the output significantly.
Useful camera vocabulary for AI video:
Shot type: close-up, medium shot, wide shot, extreme close-up, over-the-shoulder
Lens feel: shallow depth of field, deep focus, wide-angle distortion, telephoto compression
Step 5: Set the mood in one phrase
Mood anchors the visual style. Pick one or two words that describe the emotional tone.
Examples: "warm and nostalgic," "crisp and corporate," "moody and cinematic," "bright and playful," "dark and suspenseful."
Model-Specific Prompt Tips
Different AI video models interpret prompts differently. What works perfectly on one model may produce a different result on another. Here is what to know about the major models available on Avocado AI:
Dreamina Seedance 2.0
Seedance models (Fast, Standard, and Mini variants) respond well to natural-language prompts with strong subject descriptions. They handle motion reliably and produce clean, consistent frames.
Tip: Seedance works best when you describe the action as a continuous verb phrase ("slowly lifts the cup to her lips") rather than a sequence of discrete actions ("lifts cup, then drinks").
Seedance 2.0 Standard (19 credits/5s, Starter+): Best quality. Use for hero content and final deliverables.
Seedance 2.0 Fast (16 credits/5s, all tiers): Faster generation with solid quality. Good for iteration and drafts.
Seedance 2.0 Mini (10 credits/5s, all tiers): Cheapest Seedance option. Ideal for rapid prototyping.
Hailuo Pro
Hailuo Pro (7 credits/6 seconds) produces cinematic output with strong motion dynamics. It handles complex scenes with multiple moving elements better than many competitors.
Tip: Hailuo responds well to cinematic language. Use terms like "tracking shot," "dolly in," and "rack focus" for more filmic results.
Sora 2
Sora 2 Standard (10 credits/8s, Starter+) generates longer clips and handles complex spatial relationships well. The Pro 1080p variant (84 credits/8s, Growth/Pro) adds higher resolution.
Tip: Sora benefits from slightly longer, more descriptive prompts. It interprets narrative context better than shorter models, so you can include brief setup before the action.
Veo 3.1
Veo 3.1 (48 credits/8s, Growth/Pro) supports native audio generation alongside video, making it the only model on Avocado AI that produces video with synchronized sound in a single generation.
Tip: When generating with audio, describe the soundscape in your prompt: "the sound of waves crashing," "quiet office hum with keyboard typing." Veo 3.1 interprets audio cues alongside visual ones.
Kling 3.0
Kling 3.0 Pro (14 credits/5s, Growth/Pro) and Kling 3.0 4K (53 credits/5s, Pro) excel at realistic human motion and facial expressions.
Tip: Kling handles character-driven scenes well. If your prompt focuses on a person performing an action with visible facial expressions, Kling tends to produce the most natural results.
Templates for 6 Common Video Types
Copy these templates and replace the bracketed sections with your specifics.
Product showcase
[Product description] sits on [surface/material] against a [background]. Camera slowly orbits [direction]. [Lighting style]. Clean, minimal, premium feel.
Example: "A matte black wireless headphone sits on a white marble surface against a soft gray gradient background. Camera slowly orbits from left to right. Soft diffused studio lighting with gentle reflections. Clean, minimal, premium feel."
UGC / lifestyle
[Person description] [action] in [everyday setting]. Shot on a phone camera, slightly shaky, natural light. Casual, authentic, relatable.
Example: "A woman in her late 20s with curly hair unboxes a skincare product at her kitchen table. Shot on a phone camera, slightly shaky, natural window light. Casual, authentic, relatable."
Explainer / talking head
[Person description] speaks directly to camera in [setting]. Medium close-up, eye level, static camera. [Lighting]. Professional but approachable.
Example: "A man in a blue button-down shirt speaks directly to camera in a modern home office with bookshelves behind him. Medium close-up, eye level, static camera. Soft ring light fill with natural window key. Professional but approachable."
Cinematic brand ad
[Scene description]. [Camera movement]. [Lighting and color mood]. Cinematic, high production value, [reference style].
Example: "A runner sprints through an empty city street at dawn, breath visible in the cold air. Tracking shot from a low angle, matching her pace. Desaturated blue tones with warm highlights from streetlamps. Cinematic, high production value, Nike-ad energy."
Social media vertical
[Subject] [quick action] in [bright setting]. 9:16 vertical, fast pace, [energy level]. Bold colors, high contrast, attention-grabbing first frame.
Example: "A hand pours matcha powder into a glass, adds milk, and stirs with a bamboo whisk. 9:16 vertical, fast pace, high energy. Bold green and white colors, high contrast, attention-grabbing first frame."
B-roll / establishing shot
[Wide scene description]. [Time of day and weather]. [Camera movement]. Atmospheric, no people [or with people].
Example: "A modern coworking space with floor-to-ceiling windows, warm afternoon light, a few people working at laptops in the background. Slow dolly forward through the space. Atmospheric, editorial feel."
7 Prompt Mistakes That Kill Your Output
1. Packing too many actions into one clip. AI video clips are 5 to 10 seconds. One clear action per clip. If you need a sequence, generate multiple clips and edit them together.
2. Using abstract concepts instead of concrete visuals. "The feeling of success" is not a video prompt. "A woman fist-pumps alone in an empty office after closing her laptop" is.
3. Skipping the camera direction. Without camera instructions, the model picks a default angle that may not serve your vision. Even "medium shot, static camera" is better than nothing.
4. Writing in past tense. Most models respond better to present-tense, continuous-action phrasing. "A dog runs across a field" works better than "A dog ran across a field."
5. Describing what you do NOT want. Most AI video models do not reliably interpret negative prompts the way image models do. Focus on what you want to see, not what to exclude.
6. Ignoring aspect ratio. If you are creating vertical content for Instagram Stories or TikTok, specify 9:16 in your settings. Generating a 16:9 landscape clip and cropping it later wastes detail.
7. Using the same prompt across all models. Each model has strengths and interpretive tendencies. A prompt optimized for Seedance may produce different (and potentially worse) results on Kling or Sora. Test and adapt.
Iterating: How to Refine a Prompt That Almost Works
The first generation is rarely perfect. Here is how to diagnose and fix common issues:
Problem: Wrong framing or composition.
Fix: Add or change the camera direction. Specify "close-up" or "wide shot" explicitly.
Problem: Subject looks wrong or inconsistent.
Fix: Add more specific physical descriptions. Name hair color, clothing, and one distinguishing feature.
Problem: Motion feels unnatural.
Fix: Simplify the action. Break complex movements into a single primary motion. Replace "runs, jumps, and rolls" with "sprints forward."
Problem: Setting is generic or wrong.
Fix: Add sensory details about the environment. Name the time of day, lighting direction, and one prop or architectural detail.
Problem: Mood does not match your vision.
Fix: Add a one-word style reference. "Documentary," "editorial," "cinematic," "handheld," or a visual reference the model recognizes.
The key insight: change one element at a time. If you rewrite the entire prompt after each generation, you cannot tell which change fixed the problem.
How to Write Prompts Inside Avocado AI
Avocado AI gives you access to multiple video models in a single workspace. Here is the efficient workflow:
Start with Seedance 2.0 Mini (10 credits/5s). It is the cheapest video model on Avocado and fast enough for rapid iteration. Use it to test your prompt and get the framing, motion, and composition right.
Refine on Seedance 2.0 Fast (16 credits/5s). Once your prompt produces a solid result on Mini, move to Fast for better quality. This is your working draft stage.
Final render on the right model for the job. Use Seedance 2.0 Standard for polished brand content, Hailuo Pro for cinematic motion, Sora 2 for longer clips, Veo 3.1 when you need audio, or Kling 3.0 for character-driven scenes.
Use Storyboards to plan multi-clip sequences. Place your generated clips in order, preview the flow, and identify gaps before generating additional footage.
Use Flows for repeatable video workflows. If you generate product showcase videos regularly, build a Flow that chains prompt generation, image-to-video, and audio together.
All models are available through Avocado AI's pricing tiers starting at EUR 19.99/month. Credits are shared across all models, so you can mix image and video generations in the same session.
FAQ
What is the ideal length for an AI video prompt?
Between 20 and 60 words. Short enough that every word earns its place, long enough to cover subject, action, setting, and camera. Under 20 words leaves too much to chance. Over 80 words often produces conflicting instructions.
Do I need to specify the video resolution in my prompt?
No. Resolution is controlled by the model and plan settings, not the prompt text. Focus your prompt on the visual content. Check the Avocado AI pricing page for per-model resolution details.
Can I use the same prompt for different AI video models?
You can, but results will vary. Each model interprets prompts differently. A prompt that works well on Seedance may produce a different mood or composition on Sora. The best approach is to write a solid base prompt and then adjust camera and style language for each model.
How do I get consistent characters across multiple clips?
Use the exact same subject description every time. Name the same hair color, clothing, and distinguishing features. For best results, generate a reference image first using a model like Seedream V5 Pro and then use that image as the starting frame for your video generation.
Should I describe camera movement or keep it static?
For product shots and explainers, static or slow movement works best. For brand ads and cinematic content, camera movement adds production value. Start static and add movement only when it serves the story.
How do I write prompts for video with audio?
Currently, Veo 3.1 on Avocado AI supports native audio generation. For other models, generate the video first, then add audio separately using Avocado's Music and Audio Studio. Describe the desired soundscape in your Veo 3.1 prompts: ambient sounds, dialogue cues, or music style.
What is the difference between a video prompt and an image prompt?
Image prompts describe a single frozen moment. Video prompts describe an action happening over time. The key addition is the verb: what moves, changes, or unfolds during the clip. Camera direction also matters more in video because the viewer experiences time and motion.
How to Pick the Right Model in Under 30 Seconds
Need a fast decision? Use this:
Tight budget, testing prompts: Seedance 2.0 Mini (10 credits/5s)
Solid all-rounder for social content: Seedance 2.0 Fast (16 credits/5s)
Cinematic brand ad or hero content: Seedance 2.0 Standard (19 credits/5s) or Hailuo Pro (7 credits/6s)
Need longer clips (8 seconds): Sora 2 Standard (10 credits/8s)
Need video with synchronized audio: Veo 3.1 (48 credits/8s)
Character-driven scenes with facial expressions: Kling 3.0 Pro (14 credits/5s)
Maximum resolution for large displays: Kling 3.0 4K (53 credits/5s) or Sora 2 Pro 1080p (84 credits/8s)
Writing better AI video prompts is a skill that improves with practice. The framework in this guide gives you a starting structure, but the real gains come from generating, reviewing, and refining. Start with the templates above, adapt them to your brand, and iterate.
Written by Wanderson Jackson, founder of Avocado AI. Wanderson built Avocado to give creators and marketing teams access to multiple AI video and image models in a single workspace.