Controllable Visual Building Blocks
Decompose any creative output into variables you can control and iterate.
- Difficulty
- Easy
- Time to result
- ~days to results
- Steps
- 6
- Confidence
- 98%
The framework begins by asking what must exist in the desired output and which parts can be controlled. For a photograph, Rory Flynn identifies shot type, subject, action, environment, color scheme, camera, lens, mood, emotion, and lighting as visual building blocks. Each becomes a distinct prompt variable rather than being buried inside an unstructured description. The creator first supplies a condensed version of these variables, generates a baseline, and checks which instructions the model followed. When an element is wrong, its corresponding variable can be changed directly without rewriting the entire prompt. More detail is added only as necessary. This decomposition improves troubleshooting, increases repeatability, and can be transferred to video, music, system prompts, and other generated outputs by identifying their own required and controllable components.
Origin
Rory Flynn developed the structure through trial and error while trying to create photorealistic marketing assets. His graphic-design background helped him recognize visual effects, while AI experimentation forced him to learn the photography terms needed to control them.
Core principles
- 01Every output contains implicit variables whether or not you specify them.
- 02Decomposition makes creative generation controllable.
- 03Short structured prompts are easier to troubleshoot.
- 04Isolating variables accelerates iteration.
- 05The method generalizes beyond images.
How to run it
- 1
Define the intended output
Specify what you are trying to create and what function it must serve. Distinguish materially different output types, such as a close-up photograph and a drone shot.
Pro tip Start from the story or marketing purpose the asset must communicate.
Watch out Do not treat superficially similar formats as interchangeable when they create different perspectives.
- 2
Find the non-negotiables
List the categories that will appear in the output whether you specify them or let the model choose. For an image, include composition, subject, setting, color, capture style, mood, and lighting.
Pro tip Ask an LLM to analyze a strong reference and identify its building blocks.
Watch out Unspecified categories are not absent; they are simply decided by the model.
- 3
Turn categories into variables
Represent each building block as a separate, controllable prompt component. Keep the components easy to locate and edit.
Pro tip Use a consistent order so teammates can scan prompts quickly.
Watch out Do not bury key variables inside long prose.
- 4
Generate a condensed baseline
Use the shortest prompt that adequately expresses the core variables. Run it before adding extensive detail.
Pro tip Think of the first prompt as a concentrated base that can be diluted or expanded later.
Watch out Starting with several paragraphs makes individual failures hard to diagnose.
- 5
Audit model compliance
Compare the result with each requested variable and note what was represented correctly. This establishes whether the model understands the structure.
Pro tip Review variables individually rather than judging only the overall aesthetic.
Watch out Do not add new requirements before identifying which existing one failed.
- 6
Iterate one variable at a time
Change the component responsible for the weak part of the output, regenerate, and compare. Add detail gradually until the result is fit for purpose.
Pro tip Preserve successful variables while adjusting the weakest one.
Watch out Changing many variables simultaneously obscures what improved the result.
In the wild
Rory combines motorsport photography, a Red Bull F1 car, a racetrack, warm tones, a 35 mm lens, shallow depth of field, sunset backlighting, center framing, and motion blur. The resulting image reflects the requested colors, movement, perspective, background blur, and lighting direction, demonstrating that seemingly disconnected prompt terms each control a visible property.
→ A structured baseline image whose individual visual properties can be inspected and iterated.
A marketing team needs a realistic lifestyle image for an email. It defines a medium product shot, a customer using the product, a kitchen environment, muted brand colors, an iPhone-photo aesthetic, an optimistic mood, and morning window light, then changes only the lighting variable after the first result looks too dramatic.
→ The team reaches an on-brand image quickly without repeatedly rewriting a long prompt.
Common mistakes
Letting the model choose hidden variables
Omitting a category does not remove it from the output; it delegates the decision to the model and reduces control.
Starting with an oversized prompt
A multi-paragraph prompt makes it difficult to identify the single word or instruction causing an unwanted result.
Changing everything at once
Simultaneous changes prevent the creator from learning which variable caused an improvement or regression.
Is it for you?
Best for
It is best for marketers, designers, and creators who need repeatable control over AI-generated images or other media.
Not ideal for
It is not ideal for purely exploratory creation where surprise matters more than consistency or control.
From the transcript
“If you don't prompt for it, it's going to be provided for you anyway.”
“if we have, you know, lighting represented in our prompt and we don't like the lighting, we know it's in there, we know what to…”
“This sort of thinking right here, this will work for video, this will work for music, this will work for, you know, a system prompt.…”
From the episode
AI Tools to Replace Your $10k+ Creative Agency