ArticlesWriting tools

Why a voice note transcript is not a text message

Plain dictation turns a voice note into text but tends to keep the fillers, restarts and the order you thought of things in. What the research says, and what a rewrite step adds.

By Samet Durgun · Co-founder of Subtext · 19 min read

Speech carries fillers, false starts and mid-sentence corrections that finished writing mostly does not, so a plain transcript of a voice note is not yet a message. Phones will type what you said, and so will Otter, WhatsApp’s transcripts and apps built on open speech models, but a transcript of a rambling voice note is still a rambling message. It keeps the “um”, the correction you made halfway through and the order you happened to think of things in. An app that turns a voice note into a clear text message has to do a second job after transcription. It has to work out what you meant and write that.

The first half is research. Linguists have measured how spoken language differs from written language, speech recognition researchers have measured where transcription fails and for whom, and a few studies have asked why people send voice notes and why recipients put them off. The second half is what the dictation tools people already own do, what a rewrite step adds, and what Subtext does with a voice note.

As a co-founder of Subtext, where voice is one of three ways in next to a typed draft and a screenshot of a conversation, I have an interest in the answer. Every source below is one I opened. Where I could reach only an abstract, I say so and cite no figure from it.

Speech is produced differently from writing

Spontaneous speech is full of material that writing leaves out. Elizabeth Shriberg’s survey of disfluency across several corpora of spontaneous American English speech puts the rate at up to ten per cent of words and more than a third of utterances1. The forms are filled pauses (“uh”, “um”), repetitions (“I I think”), repairs where a word is replaced in the middle of a sentence, and false starts, which she files under deletions, where a clause is abandoned and begun again. In the two human to human corpora she examined, informal telephone conversations and travel planning calls, the per word rate was about six per cent1, in samples of 40,515 and 12,762 words that her SRI paper on the same corpora hand labelled2. In a corpus of people talking to a computer through a push to talk button, the rate fell to under one per cent, and she attributes part of the drop to the button. Speakers could plan the utterance first and then record it1. In a voice note you tap once and talk, and the planning happens out loud.

Spoken fillers and how much speakers differ

Spoken fillers are not noise, which is part of why they are hard to strip. Herbert Clark and Jean Fox Tree argued from several large corpora that “uh” and “um” are words in their own right, planned and produced like any other, used to announce a short or a long delay while the speaker searches for the next word or decides what to say3. They also vary enormously by speaker. In the London-Lund corpus, about 170,000 words from 50 British face to face conversations recorded between 1961 and 1976, mostly among academics, the 65 speakers who produced more than 1,000 words each ranged from 1.2 to 88.5 fillers per 1,000 words, with a median of 17.33. Some transcripts come out nearly clean, others carry a filler every dozen words.

Disfluency also clusters where planning is heaviest. Shriberg notes, citing earlier work, that it is more likely early in a phrase1. In a voice note that is the opening, where you are still deciding what the message is for. In writing you can delete that opening before anyone sees it. In a voice note it usually goes out unless you record the whole message again.

How speech recognition errors differ by speaker

Transcription adds its own errors on top of the ones you spoke. The clearest measurement I found is a 2020 study in PNAS that ran five commercial speech recognisers, from Amazon, Apple, Google, IBM and Microsoft, over 19.8 hours of interview audio from 42 white and 73 Black American speakers4. The average word error rate was 0.19 for white speakers and 0.35 for Black speakers, and every system showed the gap. The best performer, Microsoft, scored 0.15 and 0.27. The worst, Apple, scored 0.23 and 0.45. More than 20 per cent of the Black speakers’ clips came back with at least half the words wrong, against fewer than 2 per cent of the white speakers’ clips4. The authors traced the gap to the acoustic models, the part that maps sound to words, and point to too little audio from Black speakers in the training data.

A second pattern in the same paper matters for voice notes. The systems did somewhat worse on male speakers, which the authors attribute to a more informal style with shorter, more reduced pronunciations and more disfluencies4, the register of a voice note to a friend.

Accents, second languages and words nobody said

Speech recognition also does worse on accents and second language speech. A 2024 paper in JASA Express Letters tested OpenAI’s Whisper model across English accents and reported higher accuracy for native speakers than for speakers with English as a second language, better results for American English than British or Australian, and better results on read speech than on conversation5. I could reach only the abstract, so I am citing the direction and no number. Whisper was trained on 680,000 hours of audio, and its authors describe the models as approaching human “accuracy and robustness”6. Its own documentation adds that “performance varies widely depending on the language”7. An average across speakers says little about a tired voice note in your second language. If that language is English, a correct message can still read as abrupt to a native reader, and a transcript will not tell you.

Whisper can also add words nobody said. A 2024 study ran 13,140 short English clips from 437 participants recorded at US institutions, with and without aphasia, through OpenAI’s hosted Whisper API in April and May 2023. Roughly one per cent of transcriptions contained whole phrases or sentences that were not in the audio, and 38 per cent of those carried at least one explicit harm, such as violence or a false claim of authority. Clips with long pauses were more at risk. On the 187 clips where Whisper had invented text, Google, Amazon, Microsoft, AssemblyAI and RevAI produced no comparable inventions. The clips were interviews, not voice notes, and the study tested the API as it was in 20238.

Why people send voice notes

People send voice notes because they are quick for the sender. The largest study I found is from Ulm University, a 2020 survey of 1,003 people recruited through Amazon Mechanical Turk, mostly in the United States, plus a two week field study of six heavy users9. In the survey, 83 per cent had received a voice message and 73 per cent had sent one. The reasons for sending, taken from 735 open answers by people who had recorded a voice message, sorted into four themes. The theme counts add up to more than the 735 answers, so some answers fell under more than one. Convenience, since talking beats typing on a phone, was named 359 times. The tone and emotion a voice carries and text loses, 229. Situations where typing is awkward or unsafe, such as driving, 203. And the receiver, who had asked for voice or found it easier, 34.

A lab study measured the speed. Its 48 university students, 24 of them native English speakers, copied short phrases on an iPhone, and English speech input ran at 153 words a minute against 52 on the keyboard, 2.93 times faster, although speech left slightly more uncorrected errors in the final text, 0.55 per cent against 0.35 per cent. The task was copying set phrases, which says little about composing a message. Two of the authors worked for Baidu, whose server side recogniser, Deep Speech 2, was tested through an iPhone 6 Plus app, in a paper published in 201710.

Voice notes move the effort from the sender to the receiver

In the Ulm field study, 438 voice messages logged over two weeks ran from 1.5 seconds to just over seven minutes, with a median of 17.5 seconds9. The authors also point out a cost paid before any recipient hears it. Sound is transient, so a recording is cumbersome to review and edit, and a slip of the tongue means recording the whole message again9. Five of the twelve aborted recordings in their field study were exactly that.

The authors write that voice messages “shift the effort necessary for communication towards the receiver”9. Composing is quick, retrieval takes more time and effort than reading a text with the same content, and most of their field participants postponed voice messages at work because a text can be taken in at a glance and a recording cannot. Their proposed fix, in 2020, was automatic transcripts so that an appointment’s date and location buried in a recording could be found by scanning.

Why recipients put them off

Recipients cannot skim, so they go back over what they heard. In the same field study, people played a received voice message 1.37 times on average, and one message was played 13 times9. Reading does not need the replay. A 2019 meta analysis of 190 studies with 18,573 participants puts average silent reading of English nonfiction at 238 words a minute11. Conversational speech runs at about 196 words a minute when you count the whole call and about 164 within a speaker’s own turns, in a large telephone corpus12. The reader also controls the pace and can jump to the question. The listener gets the words at the speaker’s speed, in the speaker’s order.

People are not keen on them, and a little keener on receiving than on sending. A YouGov poll of 2,149 UK adults, fielded on 5 and 6 May 2022 and published that June, asked the smartphone users among them, 1,956 people, about voice notes. Of those, 16 per cent liked sending them and 36 per cent disliked sending them. For receiving them, 22 per cent liked it, 25 per cent disliked it and 29 per cent were ambivalent13. Only among 18 to 24 year olds did more people like receiving them than dislike it, 43 to 28 per cent, and even in that group half disliked sending against 30 per cent who liked it. Of all smartphone users, 63 per cent had never sent one.

How long is too long for a voice note?

In a May 2022 YouGov poll of UK smartphone users, 65 per cent of those who gave an answer said a one minute voice note is too long, and among those who dislike receiving them and gave an answer, 57 per cent were frustrated by 30 seconds. Across all smartphone users, 27 per cent did not know where the limit was. Asked how they preferred to receive a long message, 78 per cent chose text and 14 per cent a voice note, which records a stated preference only13. Some senders seem to know the recording is an imposition. A 2024 analysis of WhatsApp chats in Germany and Spain found people accounting for their choice of audio before, inside and after the message, and the authors describe these accounts as doing social work, including framing the voice note as something worth apologising for14. Across these three sources the sender saves effort and the receiver pays for it, and some senders apologise for that.

What Apple’s Dictation does with a voice note

Dictation gives you the words in the order you said them, and for Apple’s Dictation that is the whole job. Apple’s iPhone guide for iOS 27 describes Dictation as processed on the phone with no internet connection, in many languages, inserting commas, full stops and question marks automatically where the language is supported15. On the newest phones an on device model improves spelling, punctuation and capitalisation, in English. Everything else is a spoken command. You say “new line”, you say “delete” followed by the phrase. The on device claim has a condition. Apple’s privacy page says that if Keyboard Settings do not show that Dictation is processed on the device, what you dictate is sent to Apple’s servers and is not stored unless you opt in to Improve Siri and Dictation. It also says Apple may keep request history and transcripts, linked to a random device identifier and not to your Apple Account, for up to two years16. The page does not say how the two fit together. Rewriting is the job of Apple’s separate Writing Tools, which give no reading of how a text lands.

What Google and Microsoft voice typing do

Google’s basic Gboard voice typing, like Apple’s Dictation, gives you the words as you said them. Tap the microphone, say what you want written and speak the punctuation, although voice typing and punctuation are not available in every language17. Its advanced voice typing, on Pixel 6 and later, adds punctuation as you speak, and on Pixel 9 (excluding 9a) and later a spoken command can proofread, rephrase, shorten or lengthen what you dictated. Google says this runs on device, in English, French, Italian, Japanese and Spanish, with German coming soon18. On Windows, Microsoft’s help page says Fluid dictation, a Copilot+ PC feature, corrects grammar, punctuation and filler words as you speak19. Advanced voice typing rewrites only when you give it a spoken command. Fluid dictation runs without a command, but only on Copilot+ PCs, and its help page does not mention reordering the message. In the Apple, Google and Microsoft help pages cited here I found no feature that turns a whole rambling voice note into a sendable message by default.

What transcription apps and messengers do

Otter, a transcription app, goes a step beyond a phone keyboard. Its speech to text page describes a transcript with speaker labels and timestamps, an automatically generated summary with highlights, and extracted key points and action items20. Otter’s help centre lists English, Spanish, French, German, Japanese and Chinese (Simplified) as its supported languages21, which is more than the speech to text page says. It is built for meetings, so it condenses what was said, and I found no feature on those two pages that drafts what you meant to send. Open speech models sit underneath some apps. Whisper itself produces a transcript and a language label7. The apps built on it vary. MacWhisper advertises filler word removal, summaries and custom prompts on top of the transcript, and offers local or cloud models for that step22.

The messengers have started transcribing the voice notes you receive. WhatsApp added voice message transcripts in November 2024, generated on the device so that WhatsApp itself cannot read them, switched on under Settings and Chats and triggered by a long press on the message23. The post says transcripts started in a few select languages, and I could not confirm the current list. That answers the scanning problem the Ulm authors identified, where the language is covered. What you get is a transcript of someone else’s thinking out loud. You can read it in a meeting, but you still have to write the reply.

What a rewrite step adds

A rewrite step turns the transcript into the message you would have typed if typing were free. It drops the fillers Clark and Fox Tree describe, collapses the repetitions and false starts in Shriberg’s taxonomy, and puts the correction you made halfway through where it belongs, replacing the thing it corrected. It also reorders. Spoken explanations tend to arrive as they occurred to you, context first and the request last, or the request first and the reasons trailing after. A written message can lead with the point.

Two things the step has to leave alone. The first is meaning. A rewrite that smooths your voice note into a different request is worse than the transcript, because the transcript at least said what you said. The second is tone. If you were warm, or blunt, or joking, the written version should be the same person. The failure mode here is the one covered in does AI make your texts sound like a robot, where a rewrite comes back polite and generic.

There is also a judgement a transcript never has to make. Some of what you said in a voice note was for you, not for the recipient. Wondering aloud whether to mention something does not mean you meant to mention it. A rewrite step has to decide what was the message and what was the drafting, and that decision can go wrong in both directions, by keeping a sentence you were talking yourself out of or by dropping one you meant. I found no study that measured how often that happens with any tool, Subtext included.

What Subtext does with a voice note

Subtext transcribes the voice note you record in the app, then writes the message from what you meant and gives back up to three versions, each with a safe-to-send score for how likely it is to land as you meant. The versions are built to keep your meaning and your personality. On a first draft, one stays close to your words with the issues fixed, one is a warmer polish or a different approach, and one is a fuller rewrite. If the wording reads colder or sharper than you intended, the app names the issue and marks the words at fault. Nothing is sent for you.

The model’s instructions cover transcription artefacts, and I read the backend prompts before writing this. When a message is flagged as dictated, the model is told to expect misheard or fused words, odd punctuation and stray filler, to read through them to the intended wording, and to fix them silently without flagging them as something you did wrong. The prompt shows what the model is told to do. I found no independent evaluation of Subtext’s voice transcription or rewriting, so all of this rests on what the app shows and what its prompts and privacy policy say. Since none of the studies above tested Deepgram, the speech model that transcribes Subtext’s audio, I cannot say how far the error patterns for accents, second language speech and casual style carry over to it, so read the message back before you send it.

What happens to a voice recording in Subtext?

Subtext sends a voice recording to Deepgram for transcription, so the audio leaves the phone. Its privacy policy, updated 26 July 2026, names several providers, among them Deepgram for voice recordings, Anthropic for text, images and voice transcripts, Google Firebase for accounts and conversations, and OpenAI only if you switch on read-aloud. The policy adds that when web search is on, search terms drawn from your request go through Anthropic to third-party search providers, and that you can switch web search off in the app’s settings24. It says nothing you submit is used to train an AI model and that none of its AI providers may train on it, with Subtext also setting Deepgram’s model improvement opt out for voice. It says Subtext deletes its own copy of a conversation from its servers five days after you last used it, or 90 days if you pin it. It says Deepgram, with that opt out set, keeps audio only as long as it takes to transcribe it, and that Anthropic deletes inputs and outputs within 30 days, and may keep flagged requests for up to two years and related safety scores for up to seven years. It says Subtext has no zero retention arrangement with any of them24. How other assistants handle your text is compared in which AI writing assistants train on your messages.

Which languages does Subtext support?

Subtext writes the message in the language you spoke, across 17+ languages, and keeps the level of formality you used where the language marks one. A voice note in German comes back as a German message. The summary of a voice note you received is written in the language it arrived in. Google Play’s editors featured this flow in a “Hot Tip” piece that describes tapping the mic icon, speaking, tapping again, and getting a refined message with tags on how the original came across and a copy button25. When I opened the page on 4 October 2026 the listing on the German storefront showed 4.3 stars from 295 reviews and more than 50,000 downloads, a Google Play count only; the star figure differs by country. Using Subtext costs money, and a few analyses are free to start. A recorded voice note comes back as a message you can send.

How do you reply to a long voice message someone sent you?

Subtext also reads the voice messages other people send you, and this is the side where a plain transcript comes closest to enough. WhatsApp’s on device transcript lets you read a two minute recording where your language is covered23. What it does not do is tell you what the person wants. The Ulm study’s own example was a date and a place buried in the audio9. A long voice note from a friend or a manager has the same shape, the question somewhere in the middle and the reasons around it. For people who stall on replies, the reading step is one of five places a reply gets stuck, and most of the Ulm field participants postponed voice messages at work because a recording cannot be taken in at a glance9.

Drop a received voice message into Subtext and it transcribes it, sums up in one line what the person wants from you, and adds two to five action points. It can then draft a reply that picks up the thread, and it does not label the sender’s tone. The backend prompts treat a voice file you attach as most likely sent to you by the other person, and if the model is unsure who is who, it is instructed to ask. The summary is only as good as the transcription underneath it, and one line is hard to read without its context, as what does this text mean shows. For a strong accent or a noisy recording the transcription will carry more errors than for studio English, so play the original before replying to anything that matters. Text and screenshots go through the same summary step, so a message you received gets the same one line summary.

The closest experiments are old and small

The two closest experiments I found to a rewrite step or a summary are from 2003 and 2005, and both are small. In 2003, a DARPA funded experiment had 28 native English speakers read passages of 150 to 250 words of telephone and broadcast speech, as verbatim transcripts and as hand cleaned ones. Readers rated the cleaned text easier to understand, but they answered questions about it no more accurately or quickly, which the authors put down to readers already scoring close to the top. Automatic cleaning of imperfect recogniser output tended to be rated harder than the uncleaned version, though only marginally26. The cleaning removed disfluencies and fixed punctuation but did not rewrite the text into a message.

In 2005, 16 people answered questions about 15 voicemails from automatically extracted summaries. Summaries built from recognised speech got the caller’s name right 57 per cent of the time against 94 per cent for summaries built from human transcripts, while the reason for the call came through equally well, at 78 per cent. People asked for the original audio more often when the summary came from recognised speech, 53 per cent of the time against 30 per cent27. Those were extractive summaries of voicemail made with 2005 speech recognition, so the study sets an expectation only. Neither study measured a message someone would send.

Where the evidence stops

The strong evidence is at the two ends. Spoken language carries disfluency at a measured rate that writing does not, and speech recognition makes more errors for some speakers than for others, with the figures above to show it. The middle is thinner. I found no study that measured how much better a rewritten voice note lands than a raw transcript, or how often a rewrite changes the meaning on the way. The research on voice messaging is a handful of studies, the largest with a survey sample drawn from a US crowdworking platform and a field study of six. The YouGov poll is one country and one year. And the voice note studies predate on device transcripts, so the effort shift they describe is now partly closed by the messengers themselves.

In practice it comes down to how you talk. If you talk in clean sentences, dictation is enough, and your phone already has it. If you talk the way the speakers in those corpora did, with the fillers Clark and Fox Tree counted and the restarts Shriberg catalogued, a transcript hands the recipient your drafting process, and a rewrite step is what turns it into a message. Subtext is one app that does that step, for the voice note you record and the one you received, and if your trouble is a reply you keep editing, why a short reply takes an hour is about that loop. Try Subtext in your browserTry Subtext in your browser

Sources

Numbered in order of first appearance. Pages were checked on 4 and 5 October 2026, and the Subtext privacy policy was opened again on 10 October 2026.

  1. Shriberg, E. (2001). To ‘errrr’ is human: ecology and acoustics of speech disfluencies. Journal of the International Phonetic Association, 31(1), 153 to 169. Corpus study of American English conversation; rates and types of disfluency.
  2. Shriberg, E. Disfluencies in Switchboard. SRI International Speech Technology and Research Laboratory. Conference paper; the PDF carries no year, and its cited references run to 1996. Hand labelled corpora of 40,515 words from 30 speakers (Switchboard) and 12,762 words from 523 speakers (AMEX, calls between SRI employees and travel agents). Funded by DARPA and NSF. . Checked 5 October 2026.
  3. Clark, H. H., and Fox Tree, J. E. (2002). Using uh and um in spontaneous speaking. Cognition, 84(1), 73 to 111. Filler rates per speaker from the London-Lund corpus.
  4. Koenecke, A., Nam, A., Lake, E., Nudell, J., Quartey, M., Mengesha, Z., Toups, C., Rickford, J. R., Jurafsky, D., and Goel, S. (2020). Racial disparities in automated speech recognition. Proceedings of the National Academy of Sciences, 117(14), 7684 to 7689. Five commercial systems, 42 white and 73 Black speakers, 19.8 hours of audio.
  5. Graham, C., and Roll, N. (2024). Evaluating OpenAI’s Whisper ASR: Performance analysis across diverse accents and speaker traits. JASA Express Letters, 4(2), 025206. Abstract only; the full text was not reachable when checked.
  6. Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I. (2022). Robust Speech Recognition via Large-Scale Weak Supervision. arXiv preprint.
  7. OpenAI. Whisper README. GitHub. . Checked 4 October 2026.
  8. Koenecke, A., Choi, A. S. G., Mei, K. X., Schellmann, H., and Sloane, M. (2024). Careless Whisper: Speech-to-Text Hallucination Harms. ACM Conference on Fairness, Accountability, and Transparency (FAccT ’24). 13,140 audio segments from 437 participants recorded at US institutions, with and without aphasia, run through the Whisper API in April and May 2023. Funded by the Pulitzer Center and Cornell’s Center for Social Sciences.
  9. Haas, G., Gugenheimer, J., Rixen, J. O., Schaub, F., and Rukzio, E. (2020). “They Like to Hear My Voice”: Exploring Usage Behavior in Speech-Based Mobile Instant Messaging. MobileHCI 2020. Survey of 1,003 people and a two week field study of 6.
  10. Ruan, S., Wobbrock, J. O., Liou, K., Ng, A., and Landay, J. A. (2017). Comparing Speech and Keyboard Text Entry for Short Messages in Two Languages on Touchscreen Phones. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 1(4), 159. Published December 2017. Lab study of 48 university students, 24 per language, copying phrases on an iPhone. Two authors worked for Baidu USA, whose server side speech system, Deep Speech 2, ran on a Baidu server and was tested through an iPhone 6 Plus app.
  11. Brysbaert, M. (2019). How many words do we read per minute? A review and meta-analysis of reading rate. Journal of Memory and Language, 109, 104047. 190 studies, 18,573 participants.
  12. Yuan, J., Liberman, M., and Cieri, C. (2006). Towards an integrated understanding of speaking rate in conversation. Interspeech 2006. Switchboard telephone conversations.
  13. YouGov. How many Britons like voice notes? 14 June 2022. Poll of 2,149 UK adults, fieldwork 5 and 6 May 2022, 1,956 of them smartphone users; the length figures are shares of respondents who gave an answer. Results tables at . https://yougov.com/en-gb/articles/42817-how-many-britons-voice-notes. Checked 4 and 5 October 2026. The percentages follow YouGov’s article text; the results tables give slightly different totals for disliking, through rounding.
  14. Sampietro, A., and König, K. (2024). The medium is accountable: Metacommunication and media ideologies about voice messages in WhatsApp chats. Discourse & Communication, 18(1), 51 to 71. German and Spanish WhatsApp chats; abstract read through the publisher’s Crossref record, full text paywalled.
  15. Apple. Dictate text on iPhone. iPhone User Guide, iOS 27. . Checked 4 October 2026.
  16. Apple. Siri, Dictation & Privacy. Legal, last updated 14 September 2026. . Checked 5 October 2026.
  17. Google. Type with your voice. Gboard Help. . Checked 4 October 2026.
  18. Google. Use advanced voice typing features. Gboard Help. . Checked 5 October 2026.
  19. Microsoft. Use voice typing to talk instead of type on your PC. Microsoft Support. . Checked 5 October 2026.
  20. Otter.ai. Convert Speech to Text. . Checked 4 October 2026.
  21. Otter.ai. Supported languages. Otter Help Center, article edited 6 May 2026. . Checked 5 October 2026.
  22. MacWhisper. Homepage. Marketing page, undated, so it shows what the product advertises and not how it performs. . Checked 5 October 2026.
  23. WhatsApp. Introducing Voice Message Transcripts. 21 November 2024. . Checked 4 October 2026.
  24. Subtext, Terms, privacy and your account. Privacy policy last updated 26 July 2026. Checked 10 October 2026.
  25. Google Play. Hot Tip: Speak your mind with Subtext. . Checked 4 October 2026; the listing showed 295 reviews and 50K+ downloads on every storefront I opened, and 4.3 stars on the German one.
  26. Jones, D. A., and six co-authors (2003). Measuring the Readability of Automatic Speech-to-Text Transcripts. Eurospeech 2003, 1585 to 1588. 32 native English speakers recruited, 28 analysed. Sponsored by DARPA under Air Force contract F19628-00-C-0002.
  27. Koumpis, K., and Renals, S. (2005). Automatic summarization of voicemail messages using lexical and prosodic features. ACM Transactions on Speech and Language Processing, 2(1). Comprehension test with 16 participants and 15 voicemails; figures read from the author hosted PDF. Funded by an EPSRC ROPA award, GR/R23954.