MMarketing Against The Grain
← All episodes
04 March 2025

Everything You Need To Know About AI Voice Agents in 2025

3Frameworks
8Insights

Frameworks in this episode

Insights & moments

The myth-busts, hot takes, explainers, and tools worth keeping.

Myth Buster· 1

Myth Buster17:00

The Agent Often Works Before the Enterprise API Does

Large-company voice projects frequently stall for internal technical reasons rather than because customers reject AI conversations. Integration attempts reveal undocumented endpoints, unreliable behavior, unclear errors, and missing data that organizations must repair before deployment.

  • Consumer objections are not always the primary blocker
  • Old organizations often have weak internal APIs
  • Poor documentation slows agent integration
  • Unclear errors and missing data require internal remediation

the big deal ratio of meeting is not even an objection, like most of the time people are in the video issue meeting is technical.

Flo Crivello · 17:00

It's not documented. It errors out half the time. The error codes are not clear

Flo Crivello · 17:30
#enterprise#apis#integration#legacy-systems

Explainer· 4

Explainer01:00

Why Missed Calls Give Voice Agents a Direct ROI

Voice automation creates a particularly visible return because every handled minute replaces a minute of human phone work. For tradespeople and other small businesses, answering even a few otherwise missed calls can recover enough appointments to justify the system.

  • Automated phone conversations are a genuinely new capability
  • Each agent-handled minute frees a human minute
  • Missed calls can represent lost revenue
  • One or two recovered appointments may cover the cost

every minute that an AI voice agent is spending on the phone is a minute that the human is not spending on the phone.

Flo Crivello · 02:00

it needs to book one or two appointments a month for me to realize an ROI on this

Flo Crivello · 02:30
#voice-ai#small-business#roi#missed-calls
Explainer10:00

Restaurants Can Automate Peak-Hour Phone Coverage

Restaurants receive many calls precisely when staff are least able to answer them. A basic receptionist agent can immediately handle opening-hours questions and can later connect to reservation systems for booking-related requests.

  • Peak service hours create simultaneous call demand
  • Many callers ask simple opening-hours questions
  • A basic agent can field routine inbound requests
  • Reservation handling requires integration with the booking system

Full restaurants receive an enormous amount of phone calls at peak hour, right?

Flo Crivello · 11:00

We're open Monday to Friday, from 12pm to 9pm let me know if there's anything else I can help you with.

AI voice agent · 11:30
#restaurants#customer-service#inbound-calls#hospitality
Explainer12:30

Latency and Interruptions Still Expose the Machine

Current voice agents understand context well but remain less natural than people in conversational timing. Delays and awkward interruption handling require users to adjust how they speak, even when the underlying answers are strong.

  • Latency is the leading experience constraint
  • Users may need to pause more deliberately
  • Agents handle interruptions less smoothly than humans
  • Contextual understanding is stronger than conversational timing

Latency right now is the name of the game for these AI voice agents.

Flo Crivello · 12:30

Another thing that these AI agents are still not as good as humans for is interruption

Flo Crivello · 13:30
#latency#interruptions#user-experience#voice-ai
Explainer14:00

Full-Duplex Audio Could Unlock Human-Like Turn-Taking

Many voice agents still convert speech to text, reason over text, and synthesize speech, introducing delay between stages. Flo argues that truly natural timing may require native audio models that can receive and emit audio tokens simultaneously, as humans listen while speaking.

  • Many voice agents are text models wrapped with transcription and speech synthesis
  • Some newer models accept audio tokens directly
  • Current models generally cannot listen and speak simultaneously
  • Full-duplex architecture may be necessary for top-tier latency

I have a hypothesis that is going to require a re architecture of the models.

Flo Crivello · 14:00

they can't receive tokens at the same time as they can send tokens like a human can

Flo Crivello · 15:00
#full-duplex#audio-models#architecture#latency

Story· 1

Story03:00

A Voice Agent Books a Sales Demo in Minutes

Flo builds a voice workflow that receives lead information, checks a Google Calendar, calls the prospect, negotiates a time, asks for meeting context, creates the event, and sends a completion message. The live call demonstrates that current systems can complete a commercially useful workflow rather than merely converse.

  • The agent connects to calendar availability
  • The call offers and confirms a meeting time
  • The agent collects the prospect's discussion topic
  • Post-call actions create the event and report completion

I basically created my AI voice agent right here. It's full steps, it's like two minutes.

Flo Crivello · 05:30

Is there anything specific you'd like to cover during the demo?

AI voice agent · 08:00
#sales#scheduling#demo#voice-agent

Tool· 1

Tool13:00

Speculative Generation Shaves Time from Voice Responses

Lindy reduces perceived delay by generating a possible response before the speaker has definitively finished. If the system detects that the person did stop, it can reuse the work already generated and begin responding sooner.

  • The system repeatedly assumes the speaker may have stopped
  • Response generation begins speculatively
  • Useful speculative work is retained when the turn ends
  • The technique can remove hundreds of milliseconds

behind the scenes, we are continuously pretending as if the person stopped speaking, and we're generating the response, right?

Flo Crivello · 13:00

we had started generating the response 300 milliseconds ago.

Flo Crivello · 13:00
#speculative-generation#latency#optimization#voice-models

Takeaway· 1

Takeaway24:00

AI Agents Trade Instant Perfection for Long-Term Consistency

People often expect an AI agent to work immediately even though they accept weeks of ramp-up from a new employee. Agents may require repeated instruction changes, but once the workflow is calibrated, the same instructions can be followed consistently at very large scale.

  • First-shot performance is an unrealistic expectation
  • Agent onboarding requires patience and iteration
  • Stable prompts produce repeatable behavior
  • Consistency becomes more valuable as volume grows

You onboard an AI agent. You expect it to work like first shot, within the first hour, right?

Flo Crivello · 24:30

be patient and be ready to iterate on your AI agents continuously.

Flo Crivello · 24:30
#consistency#iteration#onboarding#expectations