Inclusive Speech

When more synthetic speech labels stop helping

A small amount of careful human annotation can help speech models recognize how different people pronounce the same words.

Interspeech 2026Speech MLPhonetic Transcription
phonetic transcription benchmark chart

Automatically generated phonetic labels helped when human labels were scarce. In our experiments, after about 20 to 30 hours of human annotation, adding synthetic labels could make performance worse across dialects.

Words and sounds

Speech recognition identifies words. Phonetic transcription records the sounds a person makes. Pronunciation feedback, speech therapy, and accent research need those details: two people can say the same word differently.

Having experts listen and label those sounds is expensive. A cheaper option is grapheme-to-phoneme (G2P) labeling, which predicts a word's pronunciation from its spelling. But those labels can miss what the speaker actually said.

What we found

We compared human and G2P labels on an 80-hour benchmark covering native English speakers, non-native speakers, and people with post-stroke aphasia. G2P labels helped when little human annotation was available. With more human annotation, our best approach combined a model pretrained for speech recognition with human phonetic labels.

This reduced weighted phone feature error rate, which measures errors in the sound features being transcribed, by a factor of 2.3 compared with prior systems. On aphasic speech, a model trained with 40 hours of human annotation had under 6% error, compared with over 25% for a model trained with thousands of hours of G2P labels. Only a few of those 40 hours needed to be aphasic speech. Human labels also helped the model generalize to dialects it had not seen during training.

The sweet lesson

Large-scale pretraining gave us a strong starting point. Careful human annotation then taught the model pronunciation details that labels generated from spelling missed. Both contributed to the result.

This is the sweet lesson I took from my time at UW: human expertise can make a system more useful today, even if future methods eventually automate that work. For someone who needs better speech tools now, that improvement matters.

The practical question is where human effort helps most. In our experiments, a few hours of labels for a particular kind of speech made a substantial difference. Better annotation tools can make that work easier, and models can extend its benefit to more speakers.

Read the paper Try the model Code

Interactive Demo

Try pronunciation feedback

Watch a short clip, repeat the phrase, and get pronunciation feedback from our phonetic transcription models. The Koel Labs server processes your audio and does not retain it afterward.

1/12

Want more content, translations, personalization, practice modes, and deeper feedback?
Download my startup's app.