Speech to text with an accent: how to test your setup

See what studies found about accents and speech-to-text errors, why their averages may not match your voice, and how to check your settings and test your own setup.

Published October 2, 20267 min read

You dictate a short English message for a colleague. Words you said clearly come back as different words, a client's name turns into something else, and fixing the transcript takes longer than you planned. This happens to many people who speak English as a second language, and to people with a regional accent or dialect in their first language.

Here you'll learn what published research says about accents and speech-to-text errors, and what those numbers can't tell you about your own voice. You'll also get a short test for your own setup and see where Yappee fits in.

What studies show about accents and speech-to-text errors

OpenAI documents this for its own speech recognition model. The Whisper model card says the model performs unevenly across languages and differs across accents and dialects of the same language. It strongly recommends testing the model in the specific setting where it will be used. Two studies show how large such differences can be.

A 2020 study comparing Black and white US speakers

In a 2020 study published in PNAS, Stanford researchers tested speech recognition from Amazon, Apple, Google, IBM, and Microsoft on 19.8 hours of matched interview audio from 73 Black and 42 white speakers. Across the five systems, the average Word Error Rate was 0.35 for Black speakers and 0.19 for white speakers. Word Error Rate counts the words a system swaps, drops, or adds and divides that count by the number of words in a correct transcript. That's roughly 35 and 19 errors per 100 words.

The study compared Black and white US speakers. It didn't look at people speaking English as a second language, so 35 percent is not an error rate for accented English in general. The two groups also came from different regions, which the authors name as a limit. The researchers tested the systems in spring 2019. The companies may have updated them since, so the results don't describe today's versions.

A 2025 study of non-native English speakers

A March 2025 preprint tested five systems in the versions available then: AssemblyAI Universal-2, Deepgram Nova-2, RevAI V2, Speechmatics Ursa 2, and Whisper large-v3. The speakers had Arabic, Chinese, Hindi, Korean, Spanish, or Vietnamese as their first language, with four speakers in each group. The study reports Match Error Rate, which counts the same errors as Word Error Rate but divides by a larger total that includes the added words.

On 2,400 short sentences read aloud, Whisper and AssemblyAI averaged 5.4 and 5.6 percent, with no significant difference between them. Averaged across all five systems, the rate varied widely by first language. It was under 1 percent for four US English control speakers and about 14 percent for the Vietnamese first-language group.

Free speech gave a less clear picture. That part used only 22 short narratives, about 26 minutes of audio. The statistical analyses disagreed on whether the systems clearly differed, and the systems handled filler words, repetitions, and self-corrections differently. The study is a preprint, not a peer-reviewed journal article, and its samples were small.

What these numbers mean for your own dictation

The 35 percent from 2020 and the roughly 5 percent from 2025 don't show that speech recognition got seven times better. The studies used different measures, systems, speakers, years, and kinds of speech.

Averages also hide differences between speakers, as the spread by first language in the 2025 study shows. Free speech gave a different picture again, and most work messages are spoken freely.

Taken together, the research shows that results depend on the system, the accent or dialect, and the kind of speech. It isn't a verdict on how well you speak. The number that matters is how your tool handles your voice and your messages, which matches the Whisper model card's advice to test in the real setting. You control the tool and the setup, and none of the steps below ask you to change your accent.

Check your language setting and microphone first

Apple's troubleshooting guide for Mac Dictation suggests these checks:

  • Choose the correct language and region.
  • Keep the microphone clear of objects, such as clothing or your body.
  • Stay in your usual position at your Mac and speak clearly, neither too quietly nor too loudly.
  • Avoid background noise, and try a headset microphone in a noisy room or one with a lot of echo.

Start with the language setting. If the tool expects a different language or regional variant than the one you speak, even clearly spoken words can come out wrong.

Apple wrote these tips for Mac Dictation, and they aren't a measured fix for accent gaps. They're quick to check anyway, and most of them apply to any dictation tool. Keep your normal voice while you do this. Nothing in Apple's guide or the studies above suggests slowing down to single words, over-pronouncing, or copying another accent.

Test two or three setups on your own messages

A small test on your own messages shows what goes wrong for your voice and your work. Adapt the steps to the kind of messages you send.

  1. Write down two or three short messages you would really send. Each one should contain a name, a number, a date, and a negative word such as "not" or "don't".
  2. Dictate each message in two or three setups. For example, compare your laptop microphone at your desk with a headset in a quiet room. You can also compare speaking in whole phrases with pausing every few words, or English with your strongest language plus translation.
  3. Compare each result with what you meant. Count the errors that change the meaning, such as a wrong name, number, or date, or a missing "not". Ignore small wording differences.
  4. Keep the setup with the fewest serious errors. Note the words that come out wrong every time. If your tool has a dictionary, add them there.

Here's a made-up test message:

Hi Anika, the Oyelaran quote for 2,450 euros goes out on Thursday, October 8. Please don't add the setup fee.

It has a colleague's name, a less common client name, an amount, a weekday with a date, and a "don't". A wrong word in any of those spots would change what Anika does next.

Try dictating in your strongest language

If you think and speak most fluently in another language, you can dictate in that language and have the finished text written in English. Add it to your test as one more setup. It isn't a proven accuracy gain, and your own results will show whether it works better for you.

With Yappee, you can record on your iPhone or Mac in any of more than 50 supported languages, such as Portuguese, Polish, or Hindi, and have the text written in English or another supported language. Then you paste the English text into your email or chat.

Translation adds a step where meaning, names, or tone can shift. Check the English text against what you meant, especially names, numbers, and dates. This works best when you read the target language well enough to spot a problem.

Keep names in your dictionary and use Tidy for filler words

If the words that kept coming out wrong in your test are names or terms from your work, the dictionary is the place for them. With Yappee, you can add names and domain terms, such as a client's surname or a product name, so Yappee knows how you want them written. It helps with recurring words and won't close an accent gap for the rest of what you say.

The 2025 study found that systems differ in how they handle filler words and self-corrections in free speech. When you want your own words without them, choose the Tidy format. It removes filler words, false starts, and small errors and keeps your wording. Tidy isn't a verbatim transcript, and it won't catch every misheard word. If recognition turned a name into a different word, you still need to spot it.

Judge Yappee the way you'd judge any other tool, by running the same test on your own messages.

Check names, numbers, dates, and negative words before you send

A low average in a study says little about your next message, and your messages aren't short sentences read in a lab.

Before you send a dictated message, check these details:

  • Names of people, clients, and products
  • Numbers and amounts
  • Dates and weekdays
  • Negative words, because a missing "not" turns a sentence into its opposite

These are the errors that cause real trouble when they slip through. Set the right language, sort out the microphone and the room, and test a few setups on your own messages. Keep the one with the fewest serious errors, and check the details that matter before anyone else reads them.

Put your voice to work

Turn spoken thoughts into text you can use. Available for iPhone and Mac.

Built for iPhone and Mac.