Clipwave
Media ForgeGenerate images & video, fastMedia StudioNewBuild images & video by chatting with AIWorkflowsAutomate pipelines at scaleClipperNewBreak a video down & cut the highlightsMarketing AgentDone-for-you UGC video ads
Explore NewMCPPricingFAQ
Log inStart free
AI Models/Grok Imagine/Prompting

How to prompt Grok Imagine

Grok Imagine is prompted in plain sentences, like describing a photograph to a photographer: subject and moment first, then setting, light, and framing. Its strength is aesthetic photorealism — skin, light, weather, reflections — so the levers that matter are photography words, not style keywords. On Clipwave it works as a loop: generate up to 4 stills, refine the keeper with Grok Imagine Edit using a plain-language change, then animate it into a 5-10 second clip with audio via Grok Imagine Video.

Anatomy of a Grok Imagine prompt

  1. 1

    Subject & moment

    Who or what, caught doing something specific: "two friends laughing at an outdoor cafe table", not "people at a cafe". A concrete moment is what makes the realism read as candid instead of staged.

  2. 2

    Setting & time

    Place plus time of day: "a rainy street at night", "a Scandinavian living room in late afternoon". Time words drive the light, and light is most of this model's magic.

  3. 3

    Light

    The highest-leverage block. Name source, direction and quality: "golden hour sunlight flaring across the frame", "neon signs mixing cyan and magenta on wet skin", "overcast diffused light". One deliberate lighting sentence beats five mood adjectives.

  4. 4

    Camera & framing

    Photography vocabulary lands: close-up, full-length, wide shot, low angle, shallow depth of field, long exposure. Pick one framing and one optical trait per image.

  5. 5

    Mood & finish

    Close with the grade: "cinematic color grade", "warm film-like grain", "muted concrete palette with a single red scarf". Palette accents — one color that pops — are a signature move for this model.

  6. 6

    The edit (Grok Imagine Edit)

    For edits, don't re-describe the image. State the change in natural language plus what must stay: "Change her jacket to a bright yellow raincoat. Keep the pose, face and background unchanged."

Template

[subject in a specific moment], [setting and time of day], [light: source, direction, quality], [framing and optics], [color grade or palette accent]

10 example prompts that work

Cinematic street portrait

9:16 · 2K · 4 images
A close-up portrait of a woman under neon signs on a rainy street at night, rain droplets on the edge of her hood, cyan and magenta light mixing on wet skin, shallow depth of field, candid glance just past the camera, cinematic color grade.

Neon-on-wet-skin is exactly the kind of light interaction this model renders convincingly. "Glance just past the camera" avoids the posed passport-photo look.

Editorial fashion

2:3 · 2K · 4 images
A full-length editorial photo of a man in an oversized wool coat standing in an empty brutalist concrete courtyard, overcast diffused light, muted gray palette with a single red scarf as the only color accent, low camera angle, fashion magazine look.

One accent color against a muted palette is a reliable aesthetic lever. The low angle plus "full-length" fixes both framing and subject scale in one pass.

Golden-hour candid

3:2 · 2K · 4 images
A candid photo of two friends laughing at an outdoor cafe table, golden hour sunlight flaring across the frame, plates and coffee cups mid-conversation, shallow focus on their faces, warm film-like grain, natural skin tones.

"Mid-conversation" props and a caught laugh produce lifestyle realism that stock photos fake. Sun flare is a feature here — name it or you get flat evening light.

UGC-style product realism

4:3 · 1K · 4 images
A handheld phone-camera style photo of a hand holding a glass water bottle in a bright kitchen, slightly imperfect framing, morning window light, soft shadows, realistic reflections and refractions in the glass, authentic everyday feel.

"Slightly imperfect framing" and "phone-camera style" dial realism up by dialing polish down — the honest look that performs in UGC ads. Glass reflections are a strength worth invoking.

Blue-hour cityscape

16:9 · 2K · 2 images
A wide shot of a coastal city skyline at blue hour from a rooftop, long-exposure light trails on the highway below, thousands of glowing windows, deep blue sky fading to orange at the horizon, crisp architectural detail.

"Long exposure" is an optical instruction the model translates into light trails. Blue hour gives the sky gradient that makes skylines feel expensive.

Sunlit interior

4:3 · 2K · 2 images
A sunlit Scandinavian living room in the late afternoon, low sun casting long window shadows across a pale wooden floor, linen sofa, a cat asleep on the armrest, dust motes floating in the light beam, photorealistic detail.

Long shadows and dust motes are the two details that sell "real room, real sun". A single living element (the cat) keeps a still interior from feeling like a render.

Edit: swap the background

Edit · 4:3 · 1K
Keep the bottle exactly as it is and place it on a weathered wooden table at a beach cafe, soft overcast daylight, blurred ocean in the background.

Run on Grok Imagine Edit with your product shot as the source image. The keep-clause leads the prompt — the model preserves the subject and rebuilds the scene around it.

Edit: relight the scene

Edit · 16:9 · 1K
Same scene, but at golden hour: warm low sunlight from the left, long soft shadows, sky tinted peach and violet. Keep every object and the framing unchanged.

Relighting is the highest-value edit — one pass turns a midday shot into evening without regenerating. Describing the new light fully (source, direction, shadow, sky) is what makes it coherent.

Edit: change wardrobe, keep identity

Edit · 2:3 · 1K
Change her jacket to a bright yellow raincoat and add a closed umbrella under her arm. Keep the pose, face, hair and background exactly unchanged.

One change per pass, with an explicit keep-list for identity features. Chaining focused edits beats one prompt asking for five changes — each pass stays verifiable.

Image-to-video: animate the keeper

Video · 9:16 · 5s · 720p
The camera slowly pushes in as steam rises from the coffee cup; she looks up from her book and smiles, hair moving slightly in the breeze from the open window.

Grok Imagine Video starts from your still, so the prompt describes only motion — composition, wardrobe and light are already decided. Audio is generated with the clip, no separate pass.

Settings that matter

  • Aspect ratio

    Seven ratios from 9:16 to 16:9. Pick before writing: 9:16 for stories and reels, 2:3 or 3:4 for portraits, 16:9 for scenes and banners — framing words should match the canvas.

  • Resolution

    Iterate at 1K, rerun the keeper at 2K for delivery. Grok Imagine Edit outputs at 1K, so do heavy edit loops first and the 2K generation pass last when possible.

  • Images per run

    Up to 4 per generation. Use all 4 while exploring — aesthetic variance between takes is real, and picking from four beats re-prompting one.

  • Output format

    JPEG by default; switch to PNG or WebP when your pipeline needs them.

  • Video duration & resolution

    Grok Imagine Video: 5s or 10s, at 480p or 720p. 5s at 720p is the sweet spot for social clips; 10s only when the motion you described genuinely needs the room.

Do and don't

Do

  • Write full natural-language sentences — describe the photo, don't stack keywords.
  • Spend a whole sentence on light: source, direction, quality, color.
  • Use photography vocabulary — shallow depth of field, low angle, long exposure, 35mm-style framing.
  • Generate 4 variations, pick the keeper, then refine it with Edit instead of re-rolling from scratch.
  • In edit prompts, lead with the change and close with an explicit keep-list ("keep the pose, face and background unchanged").
  • Add one imperfection for candid realism — slightly imperfect framing, a caught laugh, dust motes.

Don't

  • Don't write comma-soup keyword lists — this model reads sentences, and prose gives it relationships, not just nouns.
  • Don't ask Edit for five changes at once — chain single-change passes you can verify.
  • Don't re-describe the whole image in an edit prompt — describe the delta and what stays.
  • Don't draft at 2K — explore at 1K with 4 images, then spend 2K on the winner.
  • Don't expect text-to-video on Clipwave — Grok Imagine Video animates an existing image, so generate the still first.
  • Don't bury the subject: put who and what in the first sentence, atmosphere after.

Advanced techniques

The generate → edit → animate pipeline

The three Grok Imagine variants on Clipwave are one workflow. Generate a batch of 4 stills and pick the keeper. Fix what's wrong with focused Edit passes — background, light, wardrobe — one change at a time, so the shot converges instead of mutating. Then hand the final still to Grok Imagine Video with a motion-only prompt for a 5-10 second clip, and the audio arrives generated alongside the picture. You get an ad-ready sequence — hero still plus moving version — without leaving the family, and the identity stays locked because every step derives from the same image.

Edit prompts that behave

Grok Imagine Edit takes your source image plus a natural-language instruction. Two habits make it reliable. First, phrase the target state, not a vague wish: "warm low sunlight from the left, long soft shadows" beats "make the lighting better". Second, say what must survive — models rebuild what you don't anchor, so an explicit keep-list ("keep the pose, face and background unchanged") protects identity and composition. On Clipwave, Edit works from a single source image, so composite ideas belong in the generation prompt, not the edit.

Directing the aesthetic without style presets

There are no style toggles — the look is steered entirely by photographic description. The three levers that consistently move the aesthetic:

  • Light recipe: time of day + source + direction + quality ("blue hour, sodium streetlights, low from the right, hard shadows")
  • Optics: depth of field, exposure length, angle — one optical trait per image reads cleanly
  • Palette accent: a muted base plus one named accent color ("gray concrete, single red scarf")

Which variant to use

  • Grok Imagine

    Default. Text-to-image up to 2K in seven aspect ratios, up to 4 images per run — the exploration and hero-shot engine.

  • Grok Imagine Edit

    Refining a generated or uploaded image with plain-language changes — relight, background swap, wardrobe — while preserving what you anchor. Single source image, 1K output.

  • Grok Imagine Video

    Animating a finished still into a 5s or 10s clip at 480p/720p with audio generated alongside — the finishing move after generate and edit.

Frequently asked

What is Grok Imagine best at?

Aesthetic photorealism: portraits, lifestyle scenes, product realism, cityscapes — anywhere convincing light, skin and reflections decide whether the image reads as a photo. Steer it with photography language; for flat design or heavy in-image typography, use a design-first model instead.

Can I generate video from a text prompt?

On Clipwave, Grok Imagine Video is image-to-video: it animates an existing image. Generate the still with Grok Imagine first (or upload one), then describe only the motion. Clips run 5 or 10 seconds at 480p or 720p, and audio is generated with the video.

How do I write a good edit prompt?

State the change as the target state ("change her jacket to a bright yellow raincoat"), then anchor everything else ("keep the pose, face and background unchanged"). One change per pass — chain edits for multiple changes. Edit works from a single source image.

Should I generate at 1K or 2K?

Explore at 1K with 4 images per run, then rerun the winning prompt at 2K for delivery. Since Edit outputs at 1K, sequence heavy edit work before your final high-res generation pass when resolution matters.

Why generate 4 images per run?

Because take-to-take variance is where this model's aesthetic surprises live. Four takes of one strong prompt routinely beat four rewrites of a weak one — pick the best frame, then converge on it with Edit.

Try Grok Imagine in Media Forge →About Grok ImagineAll models
Clipwave

Tell me what you need and grab a coffee. I'll have it ready before you finish the cup.

Product

Marketing AgentWorkflow StudioMedia ForgeTemplatesAI Models

Open in app

Workflow StudioMedia ForgeProjectsAvatars

Company

ContactPrivacy PolicyTerms of ServiceLegal Notice

© 2026 Clipwave

support@clipwave.io