The AI Video Prompt Guide
The difference between a bland clip and a scroll-stopping one is rarely the model — it's the prompt. This guide condenses thousands of community generations into the patterns that consistently work with HappyHorse 1.1.
The 6-part prompt formula
Strong prompts answer six questions: Who/what? Doing what? Where? In what light? Seen how? In which style? You can write them as one flowing sentence — order matters less than presence.
Example: 'A weathered lighthouse keeper (subject) lighting an old oil lamp (action) inside a storm-battered lighthouse (setting), warm lamplight against cold blue dusk (lighting), slow push-in shot (camera), cinematic and painterly (style).'
Camera language the model understands
HappyHorse 1.1 responds strongly to film vocabulary. Use these building blocks:
- Movement: static shot, slow dolly-in, tracking shot, aerial orbit, drone flyover, handheld
- Framing: extreme close-up, medium shot, wide establishing shot, low angle, bird's-eye view
- Speed: slow motion, timelapse, real-time
- Focus: shallow depth of field, rack focus, everything in sharp focus
Lighting & mood keywords
Lighting sells realism more than any other keyword group:
- golden hour, blue hour, harsh midday sun, overcast softness
- neon glow, candlelight, moonlight, bioluminescence
- volumetric light rays, rim lighting, silhouette, lens flare
- moody, dreamy, ethereal, gritty, serene
Styles to experiment with
Append a style phrase to completely transform the same scene: 'cinematic film still', 'Studio Ghibli style animation', 'claymation stop motion', 'macro photography', '35mm film grain', 'cyberpunk concept art', 'watercolor painting in motion', 'hyperrealistic 3D render'.
One caution: pick a single dominant style. Stacking five styles produces mud.
Common mistakes (and fixes)
- Too many subjects → focus on one hero subject, let the scene support it
- Multiple conflicting actions → one clear action per clip length
- Vague adjectives ('nice', 'beautiful') → concrete visual language
- Wall-of-text prompts → 1–3 sentences beat a paragraph of keywords
- Forgetting aspect ratio → vertical 9:16 for TikTok/Reels, 16:9 for YouTube
Image-to-video prompts: describe motion, not appearance
With the image-to-video generator your uploaded picture is the exact first frame — subject, colors and composition are already locked in. The prompt's only job is motion: what moves in the scene, how the camera behaves, and how light or weather changes. 'Slow push-in, steam rising from the cup, warm light shifting' beats re-describing the whole scene.
Keep i2v prompts to one or two sentences, pick a single hero motion, and crop your image to the target ratio first (the uploader does it in one click) — the output video inherits your image's aspect ratio.
Reference-to-video prompts: the references handle the who
With the reference images to video generator, up to 9 uploaded images lock your subject's identity — so the prompt should direct everything else: scene, action, camera, lighting and mood. Anchor the subject with phrases like 'the character from the references' and use the extended 2,500-character budget for detailed scene direction.
The more cinematic detail you give (setting, time of day, camera moves, atmosphere), the better r2v re-stages your subject — vague prompts waste the identity lock on generic scenes.
Video-editing prompts: describe the change, protect the rest
AI video editing re-renders your uploaded clip according to an instruction: 'turn the video into a watercolor animation', 'dress the person in the sweater from the reference image', 'change day to rainy night'. State ONE clear edit per run and explicitly protect what must stay: 'keep the camera motion and background unchanged'.
Reference images are the precision tool — attach the exact outfit, product or style and point to it in the prompt. Since editing bills per input second, iterate on a short cut first, then run the full clip.