Voice-to-Text Modality Bridge
Speak ideas quickly, then convert them into text people can process quickly.
- Difficulty
- Easy
- Time to result
- ~days to results
- Steps
- 5
- Confidence
- 98%
This mental model separates information creation from information consumption. Speech is treated as the fastest input channel because people can express connected thoughts without pausing to operate a keyboard. Text remains the preferred output because recipients can scan, search, edit, and revisit it more efficiently than a raw voice memo. Voice AI bridges the two by transcribing speech, interpreting intent, removing verbal noise, and formatting the result as polished text. The practical consequence is that users do not need to choose between an effortless voice message and a readable document: they speak naturally, let the system structure the material, and deliver text suited to the receiving context.
Origin
Alan Gogh formed the model after seeing doctors regain their weekends through AI scribes and observing widespread voice-memo use on WeChat. Extracted from Marketing Against The Grain.
Core principles
- 01Use speech for rapid thought capture.
- 02Use text for rapid information consumption.
- 03Let AI bridge the mismatch between effortless speaking and structured writing.
- 04Choose the strongest modality for each stage of communication.
How to run it
- 1
Identify the thought
Choose a message, idea, or instruction that would take longer to type than to explain aloud.
Pro tip Start with communication you already perform repeatedly, such as email or AI prompts.
Watch out Do not dictate sensitive material where other people can hear it.
- 2
Speak without premature editing
Express the full thought naturally, including the relevant context and desired outcome.
Pro tip Talk as though you are explaining the matter to a capable colleague.
Watch out Stopping repeatedly to perfect each sentence recreates the friction of typing.
- 3
Transform speech into text
Use voice AI to transcribe, remove filler, correct words, and impose useful structure.
Pro tip Use a context-aware tool when the output must match a specific application.
Watch out Raw transcription alone may preserve verbal clutter instead of producing usable writing.
- 4
Review the written output
Scan the result for incorrect names, facts, tone, and formatting before using it.
Pro tip Add recurring names or specialist terms to a custom dictionary.
Watch out Accurate language does not guarantee accurate claims.
- 5
Deliver in the useful modality
Send or store the polished text where recipients can process, search, and reuse it efficiently.
Pro tip Retain the transcript locally when recovery or later reuse matters.
Watch out Do not send a raw voice file when the recipient needs scannable information.
In the wild
A manager needs to send a detailed project update while away from the keyboard. Instead of sending a three-minute audio file that every recipient must play in full, the manager speaks the update into a voice-AI tool. The tool removes filler, separates decisions from action items, and produces a concise written message for the team channel.
→ The manager captures the update quickly while recipients receive searchable, scannable text.
A clinician speaks or records the relevant visit information, and an AI scribe converts it into structured documentation rather than requiring hours of weekend typing.
→ Documentation time falls and the clinician regains time outside clinic hours.
Common mistakes
Sending raw audio as the final product
Speaking may be effortless for the sender while forcing every recipient to spend longer listening. Use speech as the input and readable text as the output.
Treating transcription as interpretation
A literal transcript may still contain filler, poor structure, and inappropriate formatting. The bridge requires an interpretation and cleanup layer.
Is it for you?
Best for
It is best for knowledge workers who generate substantial amounts of email, messages, prompts, or draft content.
Not ideal for
It is not ideal when speaking is socially inappropriate, confidential, or harder than entering a short precise value.
From the transcript
“voice is the fastest way to communicate thoughts to paper, but text is the fastest way to process information.”
“And that's why if we could just find a way to combine the two, which is now possible with voice AI, and a speech code…”
From the episode
Everyone’s Using AI Wrong – This Is the Real Unlock