MMarketing Against The Grain
← All episodes
08 January 2026

How to Make the Most Realistic AI Videos (Step-by-Step Tutorial)

1Frameworks
8Insights

Frameworks in this episode

Insights & moments

The myth-busts, hot takes, explainers, and tools worth keeping.

Myth Buster· 1

Myth Buster01:00

A Good First AI Video Did Not Take Five Minutes

The hosts reject social-media claims that polished AI videos are effortless one-shot creations. Flanagan estimates that his first serious attempt took 20–25 hours, although the lessons from it could reduce a future version to a couple of hours.

  • The first serious production took approximately 20–25 hours
  • Iteration and troubleshooting accounted for substantial effort
  • A learned process could eliminate roughly 80–90% of that time
  • Published AI-video demos often conceal the work behind the result

Everybody on X or LinkedIn is like, I built this awesome thing. It took five minutes. And you're like grinded away at this thing.

Kieran Flanagan · 01:00

Uh, a couple of hours.

Kieran Flanagan · 01:30
#ai video#production time#iteration#creator workflow

Hot Take· 2

Hot Take07:00

AI Rewards Experts and Unhampered Beginners More Than the Middle

The hosts argue that AI strongly accelerates people who already possess taste and deep domain expertise. They also suggest that complete beginners can experiment freely, while partially knowledgeable practitioners may become constrained by shallow assumptions without having enough expertise to direct the model well.

  • Domain expertise helps users judge and direct AI output
  • AI accelerates an existing ability to script or generate ideas
  • Complete beginners may benefit from fewer assumptions
  • Shallow intermediate knowledge can create hesitation and weak direction

AI is incredible for people with real domain expertise.

Kieran Flanagan · 07:30

you either need to be an expert or completely naive on these things because if like you're in the middle, you just get totally stuck.

Kip Bodnar · 08:00
#domain expertise#ai skills#taste#marketers
Hot Take07:00

AI Rewards Experts and Unhampered Beginners More Than the Middle

The hosts argue that AI strongly accelerates people who already possess taste and deep domain expertise. They also suggest that complete beginners can experiment freely, while partially knowledgeable practitioners may become constrained by shallow assumptions without having enough expertise to direct the model well.

  • Domain expertise helps users judge and direct AI output
  • AI accelerates an existing ability to script or generate ideas
  • Complete beginners may benefit from fewer assumptions
  • Shallow intermediate knowledge can create hesitation and weak direction

AI is incredible for people with real domain expertise.

Kieran Flanagan · 07:30

you either need to be an expert or completely naive on these things because if like you're in the middle, you just get totally stuck.

Kip Bodnar · 08:00
#domain expertise#ai skills#taste#marketers

Explainer· 1

Explainer15:00

The Dialogue, Motion, and Copyright Gotchas Behind the Demo

The production exposed several limitations that were not obvious from the polished final cut. Characters spoke each other's lines, mirrored faces failed to match mouth movement, continuous walking was unreliable, and the model refused to say certain personal or publication names.

  • Speaker assignment can fail when multiple characters share a scene
  • Reflections may not reproduce a character's speech movements
  • Continuous motion can break across generated clips
  • Name-related safeguards can interfere with calls to action
  • Some failures require creative workarounds rather than more prompting

Teddy and the computer could never really speak in the same scene.

Kieran Flanagan · 15:00

It would not allow me to say my name.

Kieran Flanagan · 15:00
#ai limitations#copyright#lip sync#video generation

Story· 2

Story03:00

Why Teddy Had to Die Before the Computer Could Speak

VO3.1 repeatedly assigned all dialogue to Teddy when Teddy and Chachi shared a scene. Flanagan obtained the intended result only after directing the model to kill Teddy during the shot, preventing him from speaking and allowing Chachi to deliver its line.

  • The model struggled to assign dialogue to the correct character
  • Both characters could not reliably speak in one shot
  • Removing Teddy as an active speaker created a workaround
  • Shot design sometimes matters more than prompt repetition

I told it to kill Teddy.

Kieran Flanagan · 03:30

So, that's the only reason the chat says words is cuz I said Teddy is dead.

Kieran Flanagan · 03:30
#vo3.1#dialogue#prompting#ai limitations
Story17:00

The Demo Inspired a Faceless AI Comedy Channel

Working within the model's short clip duration led Flanagan to imagine a sketch channel built around eight- or sixteen-second comedy pieces. He treats the technical limitation as a format constraint that could support an anonymous, faceless creative hobby.

  • Short generation limits can inspire a native content format
  • Comedy sketches naturally fit brief, self-contained clips
  • A faceless channel separates the creator's identity from experimental material
  • Technical constraints can become creative differentiators

I want to create a YouTube channel that's sketches.

Kieran Flanagan · 17:00

I could just have a faceless YouTube channel that's just sketches.

Kieran Flanagan · 17:00
#comedy#youtube#faceless channel#creative constraints

Tool· 1

Tool06:00

Use a Dedicated Voice Tool Instead of Accepting Generated Audio

After reviewing the finished advertisement, Flanagan identifies audio as a clear target for improvement. His next iteration would use ElevenLabs rather than relying entirely on the video model's generated speech.

  • Video generation and voice generation can be separated
  • Dedicated voice tooling may improve dialogue quality
  • Audio should receive its own iteration pass

I think one thing that I will definitely do is use 11 Labs for the audio.

Kieran Flanagan · 06:00
#elevenlabs#audio#voice generation#post-production

Takeaway· 1

Takeaway13:00

The Generations Cost About $2—Time Was the Expensive Part

Despite producing many image and video iterations, Flanagan estimates the direct generation cost at only about two dollars. The episode therefore presents labor, creative judgment, and troubleshooting—not model usage fees—as the dominant costs of this experiment.

  • Image and video generations cost approximately $2 in total
  • Direct tool cost was negligible relative to production time
  • Cheap generation makes extensive experimentation accessible
  • Efficiency depends more on reducing failed iterations than reducing spend

$2 or something.

Kieran Flanagan · 13:00

It was nothing.

Kieran Flanagan · 13:00
#cost#ai economics#video generation#experimentation