Think at the speed of speech: why voice input changes everything about AI prompting
Speaking is three times faster than typing, and GPT-Live's full-duplex voice finally keeps up — how to make voice input your default way to prompt AI.
Most of us think faster than we type. Conversational speech runs at somewhere between 130 and 150 words per minute, while the average typist manages around 40 — and that gap never used to matter much, because the hard part of writing was deciding what to say rather than getting it down. Working with AI turns that on its head. The quality of what you get out of these tools depends almost entirely on how much context and intent you give them, which means the keyboard has become the bottleneck between what you know and what the machine gets to work with. Speaking removes it, at roughly three times the speed.
I’ve been convinced about conversational interfaces since I stood in front of a room of executives in 2016 and told them chatbots would change everything (I’ve written about what that prediction got right and wrong). For most of the decade since, the conviction ran ahead of the technology: voice modes were clumsy, turn-based, and forever interrupting you halfway through a thought. That excuse expired on 8 July, when OpenAI launched GPT-Live.
The awkward pause is gone
GPT-Live replaced Advanced Voice Mode as the default voice experience for the 150 million-plus people who use ChatGPT’s voice and dictation features every week, and the change that matters is architectural. Every previous voice mode was turn-based, built around a detector whose job was to guess when you’d finished talking — which is why they’d barge in the moment you paused to think, or sit in silence while you wondered whether they’d heard you. GPT-Live is full duplex: it listens while it speaks, so there’s no guessing at all. Pause mid-sentence and it waits. Interrupt it and it changes course, and if you think aloud for a while it murmurs acknowledgement without taking over.
Since 23 July it’s been in the desktop apps too, where voice can drive agents: you might set a research task running in ChatGPT, get on with your inbox, then cut in with “drop the pricing section and go deeper on the case studies” without breaking stride. OpenAI isn’t alone here either: Gemini Live has offered full-duplex conversation since 2025 and Grok’s voice mode is similarly fluid, while Claude’s voice features are more modest so far. The direction of travel across the major assistants is clear all the same.
Dictation and conversation are different jobs
Before you change anything about how you work, it’s worth separating two things that get lumped together as “voice”.
Dictation uses voice as input: you speak, the words land as text in the prompt box, and you review them before anything happens. It’s the workhorse for precise work — drafting documents, briefing an agent on a complex task, writing anything where you want to check the wording before you commit. Full duplex changes nothing about dictation; it was already good, and it’s still where most of the practical gain lives.
Conversation is the other mode, and the newer one: voice as a control surface, where you’re talking with the model in real time and redirecting it as it responds. This is what GPT-Live actually upgraded, and it suits a different kind of work — thinking out loud through a problem, brainstorming, doing live research where each answer shapes the next question, or steering a long-running agent task without touching the keyboard.
A rough rule: if you’d want to read it before sending, dictate. If the value is in the back-and-forth, converse.
Getting set up
You don’t need anything exotic. Three tiers, in ascending order of commitment:
- Built-in dictation: On a Mac, press the microphone key (or set your own shortcut in System Settings under Keyboard, then Dictation); on Windows, press Win+H. It works in any text field, including every AI chat box. Accuracy is decent rather than brilliant, but it’s already installed and free.
- The AI apps’ own voice features: The ChatGPT desktop and mobile apps offer both dictation and the full GPT-Live conversational mode (GPT-Live-1 on paid plans, with a lighter mini version as the free default) — the waveform icon starts a conversation, and on desktop it can steer agent tasks. Gemini and Claude offer voice in their apps too.
- A dedicated dictation tool: Apps like Wispr Flow or Superwhisper run Whisper-class speech models system-wide, with far better accuracy than the built-in options, automatic punctuation, and the ability to clean up your “ums” as you go. If voice becomes a serious part of your workflow, this tier is worth paying for.
Start with tier one today; move up when the accuracy starts to annoy you.
You ask for better things out loud
The speed turns out to be only half the story, because speaking also changes what you ask for. Typing encourages compression — we trim context to save keystrokes, and the AI gets a terse instruction where a rich briefing would have produced something far more useful. Speech has the opposite economics. Talking for ninety seconds costs nothing, so you naturally include the background, the constraints, the audience, the things you’ve already tried — precisely the context that separates generic output from useful output.
One engineer described his workflow on X as talking “engineering prose” at ChatGPT for up to five minutes, then asking it to restructure the ramble into a formal document. That’s the pattern: speak messily, let the model do the tidying. The messiness is where the context lives.
Making it stick
The early reaction on X follows a recognisable curve: heavy day-one use, then drop-off, because chatting with a charming voice is a novelty and novelties fade. Voice earns a permanent place when you wire it into work you were doing anyway — the briefing you write every week, the agent task you kick off every morning, the decision you’d otherwise pace around the room with. Dictation is the sure bet here; the gain is immediate and it compounds. Whether full-duplex conversation becomes an everyday tool or settles into a narrower role steering long-running work, I’m not sure yet — a month isn’t long enough to know.
The other barrier is cultural, and it’s now the bigger one. Talking to a computer still feels faintly ridiculous, especially in an office, in a way that typing never did. That awkwardness is real, but it’s the same awkwardness that once attached to Bluetooth headsets and video calls — a habit problem, not a technology problem. The way past it is to start where nobody’s listening: dictate with the door closed, in the car, or out on a walk, and let the self-consciousness wear off in private before you ever do it in company.
If you’re spending hours a day prompting AI and still typing everything, you’re leaving the easiest productivity gain available on the table — and most of us use a fraction of what these tools offer as it is. The technology stopped being the excuse in July; the habit is yours to build, and a week of dictating real work is enough to start it.
