Text-to-Image-to-Video Pipeline
Build a visual reference before generating video to improve creative control
- Difficulty
- Moderate
- Time to result
- ~weeks to results
- Steps
- 6
- Confidence
- 91%
The Text-to-Image-to-Video Pipeline separates visual design from motion generation. Instead of asking a video model to invent every element directly from text, first use an image model to establish the characters, setting, composition, and style. Refine that still until it represents the intended shot, optionally building neutral character images and environment assets that can be combined into new scenes. Then pass the approved image to the video model as visual scaffolding for animation. This division lets the image model handle detailed scene construction while the video model concentrates on motion. The process is especially useful for recurring characters, advertisements, and narrative projects where uncontrolled changes between shots would undermine continuity.
Origin
Extracted from Marketing Against The Grain, where the host described replacing direct text-to-video generation with a Nano Banana-to-Veo workflow while building a sitcom.
Core principles
- 01A strong still image gives video generation clearer visual scaffolding
- 02Characters, settings, and scenes can be developed as reusable visual assets
- 03Reference images improve continuity and creative control
- 04Image and video models serve different stages of the production process
How to run it
- 1
Define the Shot
Write a concise description of the intended subject, setting, composition, and mood.
Pro tip Separate invariant character traits from details that belong only to the current scene.
Watch out An ambiguous shot description can create visual inconsistencies before video generation even begins.
- 2
Generate the Reference Image
Use an image model to create a still representation of the desired shot and refine it until the key visual choices are correct.
Pro tip Resolve wardrobe, lighting, framing, and text before adding motion.
Watch out Do not advance merely because the first image is aesthetically impressive.
- 3
Build Reusable Assets
When producing multiple scenes, create neutral character references and separate setting images that can be reused.
Pro tip Keep character reference poses and backgrounds simple enough to recombine.
Watch out Changing core character details between references can weaken continuity.
- 4
Compose the Scene
Combine the approved characters and setting into a new image representing the exact scene to animate.
Pro tip Use multiple reference images when the image model supports them.
Watch out Inspect hands, faces, text, and spatial relationships before continuing.
- 5
Animate the Image
Provide the composed image to the video model and describe the desired movement, camera behavior, and timing.
Pro tip Keep the first motion test short so iteration remains inexpensive.
Watch out Motion instructions that contradict the still image can destabilize the result.
- 6
Review Continuity
Check whether identities, environment, motion, and visual style remain consistent before generating adjacent shots.
Pro tip Preserve successful reference assets for subsequent scenes.
Watch out Do not assume continuity will persist automatically across independent generations.
In the wild
A creator generates two recurring characters on neutral backgrounds, creates a separate location, and asks the image model to combine those references into a sitcom scene. The approved scene is then passed to Veo for animation rather than being generated directly from a text prompt.
→ The creator gains greater control over characters and scene composition before spending time on motion generation.
A marketer defines a product shot, uses an image model to refine the product placement and brand styling, and then supplies the accepted image to a video model with camera and motion instructions.
→ The resulting advertisement follows an approved visual direction while still benefiting from rapid AI video production.
Common mistakes
Skipping the Visual Approval Stage
Sending an unreviewed still into video generation compounds errors and makes later corrections more expensive.
Changing References Mid-Sequence
Inconsistent character or environment references can produce visible continuity breaks between generated shots.
Overloading the Motion Prompt
Trying to transform every visual element while also animating the scene can negate the control gained from the reference image.
Is it for you?
Best for
It is best for filmmakers, advertisers, and creators producing AI-generated scenes with recurring characters or controlled settings.
Not ideal for
It is not ideal for spontaneous abstract motion where precise characters, composition, and continuity are unimportant.
From the transcript
“The big unlock for me was actually I used to go text to video and now I go like text to image to video.”
“So I go text to nano banana nana banana to vo.”
“I created neutral backgrounds and then I created a scene a setting and you can like ask it to combine those reference images into another…”
From the episode
Don't Hire a Developer Until You Watch This Gemini 3 Demo