Ingredients-to-Video AI Production Workflow
Build consistent AI videos by storyboarding, creating references, and assembling clips
- Difficulty
- Moderate
- Time to result
- ~days to results
- Steps
- 5
- Confidence
- 98%
The workflow separates creative direction from generation so each stage has a clear purpose. First, the creator owns the central idea and uses AI only to refine it. Next, the idea becomes a timed storyboard containing the visual description, audio, and dialogue for every short scene. The creator then identifies and generates the reference images needed to keep characters, clothing, props, and locations consistent. Those images become ingredients for each video generation rather than relying on text prompts alone. Finally, the resulting clips are reviewed, sequenced, and lightly edited in an accessible tool such as iMovie or CapCut. This staged approach shifts effort toward planning and visual consistency, reducing random generations and making subsequent videos substantially faster to produce.
Origin
Kieran Flanagan developed the workflow after spending roughly 20–25 hours on his first VO3.1 and Nano Banana Pro video, then identifying a process he believed could reduce future production to a couple of hours.
Core principles
- 01Own the core creative idea instead of outsourcing it to AI
- 02Design the complete sequence before generating individual clips
- 03Use reference images to preserve characters, clothing, and environments
- 04Generate video from ingredients rather than text alone
- 05Expect lightweight human editing to remain necessary
How to run it
- 1
Own and Refine the Idea
Start with a clear creative premise, audience, and intended outcome. Use AI as a thought partner to test and refine the concept, but retain responsibility for its originality and taste.
Pro tip Use audience research to refine how an existing idea should be presented rather than asking AI to invent the entire idea.
Watch out A polished generation cannot rescue a generic or poorly chosen concept.
- 2
Build the Timed Storyboard
Break the intended runtime into short scenes that fit the model's clip duration. Specify the visual description, audio, and dialogue for each scene, then iterate until the full sequence works on paper.
Pro tip Tell the model the target runtime and that the production uses 8-second clips so it can propose an appropriately sized sequence.
Watch out Do not begin generating clips while the scene order and content are still unsettled.
- 3
Create Scene Ingredients
For every scene, list the required character, wardrobe, prop, and environment images. Generate those references in Nano Banana Pro, reusing earlier character images whenever a person must remain recognizable.
Pro tip Ask the language model to list the reference images and draft an image prompt for each one.
Watch out Creating every scene independently from text will produce avoidable visual inconsistencies.
- 4
Generate from References
Upload the relevant reference images through Google Flow's ingredients-to-video mode and generate each short clip. Review for dialogue, motion, and visual anomalies before accepting a result.
Pro tip Supply only the references needed for the current scene while preserving established character anchors.
Watch out Multi-character speech and reflected faces can produce incorrect lip movement or speaker assignment.
- 5
Sequence and Lightly Edit
Place the accepted clips in story order and complete transitions, sound adjustments, and other light corrections in a simple editor. Export shorter cuts when the full narrative is longer than the target format.
Pro tip Use iMovie or CapCut if advanced editing features are unnecessary.
Watch out Plan around flaws the generation model cannot reliably correct instead of endlessly regenerating the same impossible shot.
In the wild
Flanagan created a retro sitcom in which Teddy repeatedly trusts an AI companion named Chachi, whose confident errors lead to poisonous berries, a car going over a cliff, and a hurricane mishap. Reused references kept Teddy, his clothing, the computer character, and the studio visually consistent across numerous 8-second clips. The generated clips were then assembled into a roughly two-minute advertisement.
→ A coherent first-version video ad created by a non-video expert, plus a repeatable process expected to reduce future production from approximately 20–25 hours to a few hours.
When Teddy needed to appear in a hospital with different clothing, the creator supplied a previous image of Teddy as a reference and asked Nano Banana Pro to change the wardrobe and setting while retaining his facial identity. This converted an established character into a new scene ingredient instead of recreating him from scratch.
→ The character remained recognizable even though his clothing and environment changed.
Common mistakes
Generating Every Clip from Text
Text-only generation gives the model too little visual grounding, causing characters, clothing, and locations to drift between scenes. Build reusable image references before generating video.
Outsourcing the Core Idea
AI can refine execution, but its initial concepts may lack the taste and differentiation needed for a compelling video. Begin with a human-selected premise.
Forcing Unsupported Scene Behavior
Repeatedly regenerating a scene may not solve limitations such as two characters speaking correctly or a character walking continuously across cuts. Redesign the shot or handle the limitation during editing.
Is it for you?
Best for
It is best for marketers and creators producing short narrative ads, sketches, or branded videos with recurring characters and locations.
Not ideal for
It is not ideal for productions requiring flawless multi-character dialogue, long continuous action, or precise control without manual post-production.
From the transcript
“Everything. But you cannot outsource the idea.”
“a lot of people go from text to video, but you actually want to go from ingredients to video.”
“And that's why you get that scene consistency across each of the 8-second clips.”
From the episode
How to Make the Most Realistic AI Videos (Step-by-Step Tutorial)