Two-by-Two Consistency Storyboarding
Generate four related frames together, then crop and sequence them into a scene.
- Difficulty
- Moderate
- Time to result
- ~days to results
- Steps
- 6
- Confidence
- 99%
The method treats a single two-by-two generated image as a miniature four-shot production unit. The creator asks an image model for four related views of one scene, such as an establishing shot, a closer action shot, a reaction, and an outcome. Because all four panels are generated together, they are more likely to share the same lighting, location, character design, wardrobe, and atmosphere. The creator crops promising panels, upscales each crop into a production-ready frame, and arranges them on a Figma-like board in narrative order. That board makes it easy to test whether the action progresses logically and whether adjacent shots will cut together. Once the scene works as a static sequence, each frame is supplied to a video model for animation, reducing continuity drift and wasted motion-generation attempts.
Origin
PJ Ace shared this as his latest technique for preserving continuity in AI-generated sequences while demonstrating storyboard grids from his Legend of Zelda fan trailer. Extracted from Marketing Against The Grain.
Core principles
- 01Generate related shots together to preserve lighting, characters, and location.
- 02Plan a scene as a sequence rather than as isolated reference images.
- 03Use shot-size variation to create editable visual rhythm.
- 04Upscale selected crops before animation.
- 05Solve continuity on a storyboard board before generating motion.
How to run it
- 1
Divide the script into scenes
Identify short units of action that can be expressed through approximately four connected shots. Give each unit a clear beginning, progression, and endpoint.
Pro tip Choose scenes that keep one location and lighting condition wherever possible.
Watch out A grid containing unrelated moments will lose the consistency advantage.
- 2
Prompt a two-by-two shot grid
Ask the image model to create one two-by-two grid containing four cinematic shots from the same scene. Describe the character, location, time of day, action progression, and desired framing.
Pro tip Request complementary shot sizes rather than four nearly identical compositions.
Watch out Do not assume the model will infer narrative order; specify what changes from frame to frame.
- 3
Select and crop panels
Choose the panels that tell the clearest story and crop each into its own image. Discard panels with identity, anatomy, product, or continuity errors.
Pro tip Generate another grid when fewer than three panels are genuinely usable.
Watch out Do not preserve a weak panel merely to maintain the original four-shot structure.
- 4
Upscale the frames
Use an image model or upscaler to restore detail lost when cropping a grid panel. Confirm that upscaling has not altered the character, costume, setting, or action.
Pro tip Tell the upscaler to preserve composition while enhancing detail.
Watch out Generative upscalers can introduce continuity-breaking details.
- 5
Build the scene board
Arrange the crops in order on a Figma-like board and inspect how they cut together. Move, replace, or regenerate frames until the scene communicates without animation.
Pro tip Alternate shot sizes and include reaction shots to improve editorial flexibility.
Watch out Similar consecutive wide shots can create awkward, visually flat cuts.
- 6
Animate the approved sequence
Send each final frame to the video model and generate motion that advances the planned action. Edit the resulting clips in the same order, adjusting duration and transitions as needed.
Pro tip Use the final frame of one clip as a reference for the next when continuity needs additional reinforcement.
Watch out Do not let an impressive animation override the storyboard's intended action direction.
In the wild
PJ describes prompting four shots of Link jumping over a crumbling bridge at sunset. The shared grid preserves the same sunset lighting and environment across panels; he then crops and upscales the chosen images before animating them as separate shots.
→ The scene gains stronger frame-to-frame consistency than a collection of independently prompted images.
A grid shows the hero jumping, striking the creature, and leaving it on the ground. The creator selects the useful action stages, converts them into individual frames, and lays them out as a readable scene before video generation.
→ The animator receives an explicit visual progression rather than an ambiguous text-only request.
Common mistakes
Generating every shot separately
Independent prompts allow lighting, faces, wardrobe, and locations to drift. Generate related panels together whenever the scene permits it.
Treating references as a storyboard
A few attractive images do not establish action progression or edit points. Arrange distinct shots in narrative order and evaluate the scene as a whole.
Skipping the upscale check
Grid crops may lack resolution, while generative upscaling can alter important details. Inspect every enhanced frame before animation.
Is it for you?
Best for
It is best for creators building multi-shot AI scenes with recurring characters, locations, lighting, or action.
Not ideal for
It is not ideal when every shot requires a radically different environment or when the image model cannot render usable detail in grid panels.
From the transcript
“All you need to do now is create things in like four-shot sequences.”
“And the advantage of these two-by-two frames and and the quality will vary is the advantage is that the lighting and the characters and the…”
“So, I I would do the the two-by-two grid shots which I used in FreePik and then I would bring them onto like a Figma-like…”
From the episode
233M Views in 3 Days: The David Beckham AI Workflow