Image to Video AI Prompt: How to Write Prompts That Actually Work
Wanderson Jackson
Updated June 2026
TL;DR: The difference between a 5-second AI clip you can use and one you delete comes down to prompt structure. A strong image-to-video prompt names the subject, describes one clear action, specifies camera movement, and defines the visual style. This guide breaks down the formula model by model so you get usable output on the first or second try, not the tenth.
Image-to-video is different from text-to-video. You already have the visual: the scene, the character, the product, the lighting. Your prompt does not need to describe what the image looks like. It needs to describe what happens next.
Think of it as a director's note, not a scene description. You are telling the model: "Here is frame one. Now move this part, in this direction, at this speed."
Kling AI's own image-to-video guide puts it simply: the formula is Subject + Movement. That is the core. Everything else (camera, style, timing) layers on top of that foundation.
Without a clear subject and movement prompt, most models will produce a static video with a slow pan or subtle drift. Technically animated, but useless for ads, social content, or product demos.
The 6-part prompt framework
Every strong image-to-video prompt has up to six layers. You do not always need all six, but the first three are non-negotiable.
1. Subject
Name what moves. Be specific.
Bad: "The person walks"
Good: "The woman in the red jacket walks toward the camera"
If your image has multiple elements, pick one to animate. Trying to move everything at once produces chaotic output.
2. Action
One clear verb in present tense. One movement per clip.
"picks up the coffee cup"
"turns to face the camera and smiles"
"flies across the frame from left to right"
Multiple actions in a single prompt get muddled. If you need a sequence, generate separate clips and stitch them.
3. Camera
This is the highest-leverage addition most people skip. Camera movement controls pacing, mood, and focus.
Camera term
What it does
Best for
Static shot
No camera movement; stable frame
Product reveals, text overlays
Dolly in / push in
Camera moves toward subject
Intimacy, tension, close-up reveals
Dolly out / pull out
Camera pulls back
Scale reveals, isolation
Pan left / right
Horizontal rotation
Scene reveals, following action
Tilt up / down
Vertical rotation
Power dynamics, establishing scale
Tracking shot
Follows a moving subject
Action sequences, walking shots
Orbit / arc
Circles around subject
Product showcases, dynamic portraits
Handheld
Slight natural shake
UGC, authentic feel
Drone / aerial
High-angle sweeping
Establishing shots, landscapes
4. Style
One or two visual anchors. This tells the model what register to render in.
"cinematic, film grain"
"clean product commercial"
"UGC, handheld iPhone footage"
"anime, Studio Ghibli aesthetic"
"documentary, natural lighting"
Do not stack five adjectives. Two strong references beat a paragraph of vague descriptors.
5. Lighting and mood
Lighting is part of the source image, but you can shift it with the prompt.
"warm golden hour light"
"dramatic rim lighting"
"soft overcast, no harsh shadows"
"neon glow, nighttime"
If your source image already has strong lighting, skip this layer. Adding conflicting lighting instructions confuses the model.
6. Timing and pacing
Some models (Seedance 2.0, Veo 3.1) respond to temporal cues.
"slowly, over 3 seconds"
"quick snap movement"
"gradual acceleration"
For most models, pacing is inferred from the action verb. "Strolls" is slow. "Darts" is fast. Use verb choice over explicit timing when possible.
Model-specific tips
Different models respond to prompts differently. Here is what works on each.
Seedance 2.0 (Dreamina)
Seedance 2.0 is a multimodal directing system, not just a text-to-video engine. Its image-to-video mode accepts up to 9 reference images, 3 video clips, and 3 audio files alongside your text prompt.
What works:
Use @image1 as the first frame syntax to reference uploaded images
Describe camera movement explicitly: "slow dolly in, slight tilt up"
Reference video clips for motion style: "match the pacing from @video1"
Seedance responds well to one-take descriptions: "one continuous shot, no cuts"
What to avoid:
Seedance 2.0 is restricted with people. For character-heavy scenes, consider Kling 3.0 or Veo 3.1
Do not stack multiple complex actions. One movement per generation
Cost on Avocado: 19 credits per 5-second clip. Available on all plans.
Kling 3.0
Kling excels at camera physics and character motion. It handles complex movements like hair blowing, fabric physics, and multi-subject scenes better than most competitors.
What works:
The Subject + Movement formula is Kling's native language
Be physically realistic: "the athlete cycles on the highway with a sense of speed" works; "the athlete flies through the air" does not
Kling handles pan, tilt, dolly, and tracking natively through prompt descriptions
Use Professional mode for higher visual quality
What to avoid:
Complex physical simulations (bouncing balls, objects thrown at high altitude) are still challenging
Descriptions that deviate significantly from the source image cause camera cuts or transitions
Cost on Avocado: 14 credits per 5-second clip (Kling 3.0 Pro, Growth+). 53 credits for 4K (Pro only).
Veo 3.1
Veo 3.1 generates video with native audio support. It follows a five-part prompt structure: Subject, Action, Camera, Setting, and Aesthetic.
What works:
Veo handles cinematic language well: "close-up shot with a slow zoom-in"
Include aesthetic descriptors: "cinematic, photorealistic, brand documentary style"
Veo responds to lighting cues: "warm natural light through windows, soft shadows"
Max prompt length is generous (up to 1,800 words), but shorter is usually better
What to avoid:
More than four main subjects can confuse the model
Abstract or poetic prompts produce unpredictable results. Be concrete
Cost on Avocado: 48 credits per 8-second clip with audio. Growth and Pro only.
Sora 2
Sora 2 is available on Avocado via partner access (OpenAI discontinued the consumer product in April 2026; API access remains available through September 2026).
What works:
Sora handles wide establishing shots and environmental motion well
Use narrative framing: "a slow aerial shot reveals a coastal town at dawn"
Specify shot type explicitly: "wide shot, static camera"
What to avoid:
Character animation is not Sora's strongest suit compared to Kling or Veo
Prompt interpretation can be loose on fast action sequences
Cost on Avocado: 10 credits per 8-second clip (Standard, Starter+). 84 credits for Pro 1080p (Growth+).
Hailuo Pro
Hailuo is the budget-friendly option. At 7 credits per 6-second clip, it is the most cost-efficient model on Avocado for image-to-video.
What works:
Simple, clear prompts with one subject and one action
Good for ambient motion: "leaves rustle in the breeze," "clouds drift across the sky"
Works well for product rotation and slow reveals
What to avoid:
Complex multi-subject scenes
Fast action or intricate character motion
Cost on Avocado: 7 credits per 6-second clip. Available on all plans.
Common mistakes
1. Describing the image instead of the motion. Your source image already shows the scene. The prompt should describe what changes, not what is already there.
2. Stacking multiple actions. "The woman picks up the cup, takes a sip, puts it down, and walks away" is four clips, not one. Generate each action separately.
3. Vague camera instructions. "Make it cinematic" is not a camera instruction. "Slow dolly in from medium shot to close-up, shallow depth of field" is.
4. Ignoring the source image composition. If your image shows a person on the left side of the frame, do not prompt them to walk right (off-screen). Work with the composition you have.
5. Overloading with style adjectives. "Cinematic, dramatic, moody, atmospheric, film noir, high contrast, grainy, 35mm, anamorphic" in one prompt is noise. Pick two.
6. Using text-to-video prompts for image-to-video. Text-to-video prompts describe the entire scene. Image-to-video prompts only describe motion. The source image handles the rest.
Templates you can copy
Product showcase
Subject: [product name] on [surface]
Action: slowly rotates 360 degrees
Camera: medium shot, static camera, slight push in at the end
Style: clean product commercial, studio lighting
Example: "A matte black headphone on a marble surface slowly rotates 360 degrees. Medium shot, static camera with a slight push in at the end. Clean product commercial, soft studio lighting."
UGC-style product demo
Subject: [person] holding [product]
Action: brings product close to camera, smiles
Camera: close-up, handheld
Style: UGC, iPhone footage, natural window light
Cinematic landscape reveal
Subject: [landscape/scene]
Action: fog clears to reveal the full vista
Camera: wide shot, slow dolly out
Style: cinematic, golden hour, shallow depth of field
E-commerce model shot
Subject: model wearing [clothing item]
Action: turns slowly to show the back of the garment
Camera: medium full-body shot, slow orbit
Style: fashion editorial, clean white studio background
Food and beverage
Subject: [drink/food item]
Action: condensation forms on the glass, ice shifts slightly
Camera: close-up, static shot
Style: commercial food photography, dramatic side lighting
What actually matters
The prompt is not the hard part. The hard part is choosing the right source image and knowing which model fits your use case.
A perfectly structured prompt on a bad source image will still produce bad output. Start with a high-quality, well-composed image. Then write a prompt that describes one clear action with one camera movement.
If you are generating at scale for ads or e-commerce, the real advantage is speed. A workspace that lets you generate across multiple models from the same image means you can test Seedance for product shots, Kling for character scenes, and Hailuo for ambient content without switching platforms.
An image-to-video AI prompt is a text instruction that tells an AI model how to animate a still image. Unlike text-to-video prompts that describe an entire scene from scratch, image-to-video prompts focus only on motion: what moves, in what direction, and how the camera responds.
How long should an image-to-video prompt be?
Two to four sentences is the sweet spot. Name the subject, describe one action, specify camera movement, and optionally add a style reference. Longer prompts are not better. Most models ignore filler words and focus on concrete nouns and verbs.
What is the best AI model for image-to-video?
It depends on the use case. Kling 3.0 handles character motion and physics best. Seedance 2.0 supports multi-reference inputs for precise control. Veo 3.1 generates video with native audio. Hailuo Pro is the most cost-efficient at 7 credits per clip. For a workspace that covers all of them, Avocado AI offers every model on one credit pool.
Do I need to describe the image in my prompt?
No. The model can see the image. Your prompt should describe what changes, not what is already there. Focus on subject, action, and camera movement.
How do I write prompts for product video ads?
Start with the product as the subject. Use one clear action (rotate, reveal, zoom in). Specify a clean camera movement (static shot or slow dolly). Add a style anchor like "product commercial" or "e-commerce listing video." Keep it to three sentences.
Can I use the same prompt across different models?
The structure works across models, but results vary. Seedance 2.0 responds to @file references. Kling follows Subject + Movement natively. Veo handles cinematic language well. Test your prompt on two or three models to find the best fit for each content type.
What is the cheapest way to generate image-to-video at scale?
Use Hailuo Pro for ambient and simple motion (7 credits per 6-second clip). Reserve Kling 3.0 and Veo 3.1 for complex scenes that need strong physics or audio. On Avocado AI, all models share one credit pool, so you can allocate credits based on clip complexity.
How to pick in under 30 seconds
Need character motion with realistic physics? Use Kling 3.0.
Want multi-reference control with video and audio inputs? Use Seedance 2.0.
Need video with native audio? Use Veo 3.1.
On a tight budget? Use Hailuo Pro (7 credits per clip).
Want to compare output across models? Use Avocado AI.