Clipwave
Media ForgeGenerate images & video, fastMedia StudioNewBuild images & video by chatting with AIWorkflowsAutomate pipelines at scaleClipperNewBreak a video down & cut the highlightsMarketing AgentDone-for-you UGC video ads
Explore NewMCPPricingFAQ
Log inStart free
AI Models/Hailuo 2.3/Prompting

How to prompt Hailuo AI

Hailuo 2.3 animates a still image, so the prompt's only job is to describe what MOVES — the expression, the gesture, the camera, the air in the scene. Everything the image already shows is settled; re-describing it wastes words and invites morphing. The best Hailuo prompts are short (15-30 words), name the subject briefly, give one specific action, and one camera move in plain language. It is the model to reach for when a face has to feel alive.

Anatomy of a Hailuo 2.3 prompt

  1. 1

    Subject anchor

    A brief tag that points at who moves — "the red-haired woman", "the ceramic bottle". Two or three words, not a portrait: the image already carries the appearance, the anchor just prevents subject swapping.

  2. 2

    Action & expression

    The heart of a Hailuo prompt. One specific, physical action with an emotional register: "slowly smiles, eyes narrowing with warmth". Hailuo 2.3's standout skill is facial micro-expression in close-ups — give it expression verbs, not mood adjectives.

  3. 3

    Camera move

    One primary camera movement, written in plain language: "slow push-in", "gentle pan left", "static camera". Do not stack zoom + pan + tilt in a single clip — one move stays stable, three hallucinate.

  4. 4

    Environment motion

    What the world does around the subject: "hair lifts in a light breeze", "steam curls from the cup", "neon signs flicker". Small ambient motion sells realism more than big gestures.

  5. 5

    Pacing word

    One or two motion adjectives that set intensity — "slow, deliberate" or "sudden, explosive". Hailuo maps these words to motion amplitude; more than two starts to conflict.

Template

[brief subject anchor] [specific action + expression], [one camera move in plain language], [ambient motion in the scene], [1-2 pacing adjectives]

10 example prompts that work

Portrait comes alive

i2v · 9:16 · 6s · prompt optimizer on
The young woman slowly smiles, her eyes brightening, a light breeze lifting strands of her hair, she blinks naturally, slow push-in, soft and gentle motion.

The classic Hailuo shot: a generated or real portrait turned into a living moment. Notice what is NOT here — no description of her face, clothes or background; the image owns all of that. One camera move, one expression arc.

Emotional close-up (micro-expressions)

i2v · 16:9 · 10s · prompt optimizer off
The old fisherman's expression shifts from stern to quietly moved — his jaw softens, eyes glisten, a slow blink, the faintest smile forming. Static camera, very slow pacing.

Hailuo 2.3 was tuned for facial micro-expression in close-ups, and this is where it beats most video models. Describe the expression as a journey ("from stern to moved") and lock the camera so nothing competes with the face.

Anime character action

i2v · 16:9 · 6s · prompt optimizer on
The anime swordswoman draws her blade in one fluid motion, cape snapping behind her, cherry blossom petals streaking past, quick pan right following the draw, sharp decisive movement.

Style consistency on anime and illustrated frames is a 2.3 headline improvement — fewer off-style frames and less flicker. Keep the action singular and let the line art stay stable; wild multi-beat choreography reintroduces flicker.

E-commerce product shot

i2v · 1:1 · 6s · prompt optimizer off
The perfume bottle rotates slowly on its base, light sweeping across the glass, fine mist drifting around it, static camera, smooth and controlled motion.

Object stability in ad-style shots is another 2.3 focus — labels and proportions hold while the product moves. Rotate the product, not the camera: a static camera plus subject motion is the most stable e-commerce recipe.

Expressive full-body motion

i2v · 9:16 · 10s · prompt optimizer on
The dancer pushes off her back foot into a slow spin, dress flaring outward, arms unfolding overhead, she lands softly and holds the pose, gentle pan left, fluid continuous motion.

Body motion reads natural when the prompt describes a single continuous phrase of movement with a beginning and an end ("pushes off... lands and holds"). Listing three unrelated moves makes limbs drift.

Physical transformation

i2v · 16:9 · 10s · prompt optimizer on
The bronze statue cracks along its surface, fragments falling away as golden light pours from within, dust drifting in the air, slow push-in, building intensity.

Hailuo treats short physical verbs — crack, melt, ignite, dissolve — as structured events, not vibes. Keep transformations to one material change per clip and give the debris somewhere to go ("fragments falling away").

Camera-led reveal

i2v · 21:9 source · 10s · prompt optimizer off
Slow push-in toward the lit doorway at the end of the corridor, dust motes floating through the light beam, steady cinematic motion.

When the camera IS the story, follow the 2+1 pattern: one or two spatial moves plus one quality word ("steady cinematic"). The subject here is the space itself — no character anchor needed.

Landscape mood shift

i2v · 16:9 · 10s · prompt optimizer on
Fog rolls in across the valley floor, treetops swaying gently, the light warming toward golden hour, very slow pan right, calm and atmospheric.

For establishing shots, prompt the weather and light as the actors. Changes should be gradual verbs ("rolls in", "warming toward") — instant changes ("suddenly night") force the model to abandon the source frame.

Character walks to camera

i2v · 9:16 · 6s · prompt optimizer off
The man in the trench coat turns from the window and walks toward the camera, coat swaying with each step, his expression hardening with resolve, static camera, deliberate pacing.

Movement toward a static camera is more reliable than a tracking shot following the subject — the model handles the natural scale-up well. The expression note gives the approach a reason, which keeps the face coherent as it grows.

Two-clip stitched sequence

i2v ×2 · 16:9 · 6s each · prompt optimizer off
Clip A (from your still): The chef lifts the pan and flames burst upward, embers rising, slow push-in, energetic motion. Clip B (from Clip A's final frame): The flames settle, he plates the dish with a satisfied nod, gentle pan down to the plate, calm motion.

Hailuo generates one shot at a time — sequences are built by stitching: generate Clip A, take its final frame, use it as the source image for Clip B, and change ONE camera axis per clip. Run each prompt as its own generation.

Settings that matter

  • Source image

    The most important "setting". Clean, uncluttered, well-lit frames animate best; blur, overlapping subjects and busy backgrounds degrade motion. Generate the exact frame in Media Forge first, then animate it.

  • Duration

    6s or 10s. Use 6s for a single expression or action beat; 10s only when the motion genuinely has two phases (turn, then walk). Stretching one small action across 10s produces drift.

  • Aspect ratio

    16:9, 9:16 or 1:1 — pick it when you CREATE the source image, since the output follows the frame you feed in. 9:16 for social portraits, 1:1 for product tiles.

  • Prompt Optimizer

    On by default — it expands sparse prompts into fuller ones and is the right choice for short, casual prompts. Turn it off once you are writing precise camera and motion directions you want followed literally.

Do and don't

Do

  • Prompt only the delta: motion, expression, camera — never what the image already shows.
  • Keep it 15-30 words with one specific action; short physical prompts outperform essays.
  • Use a brief subject anchor ("the red-haired woman") before the action to prevent subject swapping.
  • Write camera moves in plain language — "slow push-in", "gentle pan left", "static camera".
  • Give expressions as a journey with verbs: "softens, eyes glisten, a smile forming".
  • Build multi-shot sequences by stitching: last frame of clip A becomes the source of clip B.

Don't

  • Don't combine zoom + pan + tilt in one clip — one primary camera move per generation.
  • Don't re-describe the still ("a beautiful woman with blue eyes in a red dress...") — it confuses the motion and invites morphing.
  • Don't stack quality spam like "8k, masterpiece, ultra-detailed" — it adds nothing and can oversaturate.
  • Don't ask for a wide shot from a close-up source (or vice versa) — the model hallucinates the missing space.
  • Don't use negative-prompt syntax — there is no negative prompt; describe the motion you want instead.
  • Don't write dialogue expecting sound — Hailuo clips are silent; pair with a lip-sync model for speech.

Advanced techniques

Camera direction in plain language

Older Hailuo Director models used bracket tags like [Push in] and [Pan left]; many guides still teach them. Hailuo 2.3 takes camera direction as ordinary language inside the prompt. The stable pattern is "2+1": at most two spatial moves (e.g. a pan plus a subtle push) and one quality word (steady, slow, cinematic). Every axis you add past that multiplies the odds of warped geometry.

  • push-in / pull-back — move toward or away from the subject
  • pan left / pan right — horizontal sweep across the scene
  • tilt up / tilt down — vertical pivot, good for reveals
  • static camera — say it explicitly when only the subject should move
  • slow zoom — gentler alternative to a push-in for portraits
  • quality words: steady, slow, smooth, cinematic — pick one, not four

Sequential stitching for multi-shot stories

Hailuo produces one shot per generation, so sequences are assembled, not prompted. Generate clip A from your still, pull its final frame, and feed that frame back as the source image for clip B with a new prompt — changing only one camera axis per step. Because each clip starts from real pixels of the previous one, identity and lighting carry over far better than regenerating from scratch. Three or four stitched 6-second clips give you a coherent scene that no single 10-second prompt could hold.

Build the frame first, then animate it

Hailuo's ceiling is set by the frame you feed it, which makes it the natural second step of a two-step workflow on Clipwave: compose the exact still in Media Forge with an image model — framing, wardrobe, lighting, aspect ratio all locked — then hand it to Hailuo 2.3 with a motion-only prompt. For talking characters, keep the mouth closed and neutral in the still, animate the performance with Hailuo, then run the clip through a lip-sync model in Media Forge to add the spoken line.

Frequently asked

Does Hailuo 2.3 support negative prompts?

No. There is no negative-prompt field — steering is positive description only. Instead of "no camera shake", write "static camera, smooth motion"; instead of "no distortion", simplify the prompt to one action and one camera move.

Does Hailuo generate audio or dialogue?

No — output is silent video. For a talking character, animate the performance with Hailuo 2.3 first, then run the clip through a lip-sync model in Media Forge with your audio; the expressive face Hailuo produces is an ideal lip-sync base.

Should I pick 6 or 10 seconds?

6s for one action or expression beat — it is the sharper, safer choice. Pick 10s only when the motion has two real phases, like a turn followed by a walk. A single small action stretched over 10 seconds tends to drift and loop.

Prompt Optimizer on or off?

On while you sketch: it expands short prompts into fuller descriptions and generally helps casual use. Off once you are directing precisely — with the optimizer off, your exact camera and motion wording is followed literally instead of being rewritten.

Why does my subject morph into someone else mid-clip?

Usually one of three causes: the prompt re-describes appearance (conflicting with the image), it asks for multiple simultaneous actions, or it lacks a subject anchor. Fix: anchor briefly ("the woman in the chair"), give ONE action, and describe zero appearance.

Can I use Hailuo for text-to-video on Clipwave?

In Media Forge, Hailuo 2.3 is image-to-video — you always start from a frame. For text-to-video with the Hailuo family, use Hailuo 02 Pro inside Workflow Studio, or generate your frame with an image model first and animate it with 2.3.

Try Hailuo 2.3 in Media Forge →About Hailuo 2.3All models
Clipwave

Tell me what you need and grab a coffee. I'll have it ready before you finish the cup.

Product

Marketing AgentWorkflow StudioMedia ForgeTemplatesAI Models

Open in app

Workflow StudioMedia ForgeProjectsAvatars

Company

ContactPrivacy PolicyTerms of ServiceLegal Notice

© 2026 Clipwave

support@clipwave.io