← Back to Typeglot News

Most speech recognizers perform best when a speaker stays inside one language. Real conversations are less tidy. On August 19, AssemblyAI demonstrated Universal-3.5 Pro transcribing mixed-language speech and using a short context prompt to resolve difficult terms. The model launched on July 7, so this week's news is the demonstration, not a new launch.

What AssemblyAI showed this week

In its August 19 demonstration, AssemblyAI showed English mixed with French, Hindi, and Mandarin. The company says its model can follow code-switching across 18 supported languages without a language flag or separate detection pass. It also showed contextual prompting for names and specialist vocabulary.

The goal is not to translate everything into one language. It is to preserve the language of each spoken word. These are provider demonstrations, not benchmarks reproduced by Typeglot. The model itself was introduced in AssemblyAI's July 7 announcement.

Why mixed-language speech is difficult

Code-switching means changing language inside a conversation, sentence, or phrase. Someone may speak English, insert a Polish place name or Italian expression, then continue in English. A system locked to one language can reshape unfamiliar sounds into the words it expects.

Supporting many languages is not the same as following one person as they move among them.

When Typeglot users should choose Auto-Detect

If you regularly switch languages, choose Auto-Detect as Typeglot's input language. Instead of applying one fixed language hint, Typeglot lets its speech-recognition layer identify the language. Selecting one input language remains the best option when you consistently dictate in that language because it gives the recognizer a narrower task.

The tradeoff is small: Typeglot's current internal estimate is about one percentage point of accuracy in the main language. For monolingual dictation, select that language. For regular code-switching, Auto-Detect is usually the better fit. Very short switches, uncommon names, or noisy audio can still cause errors.

Code-switching and cross-language output solve different problems

Preserve mixed-language speech

Auto-Detect helps preserve the languages that were actually spoken.

Choose another output language

Ghost Mode turns speech into finished text in a selected output language.

These workflows are related but distinct. Typeglot's Ghost Mode is for thinking in one language and writing in another. Auto-Detect is the better starting point when the spoken input itself moves among languages. Typeglot does not use AssemblyAI.

What to watch next

The real test is ordinary speech: accents, background noise, uncommon names, and rapid switches. As code-switching improves, products must still explain clearly whether they preserve multiple spoken languages or translate them into one output language.

Frequently asked questions

What is code-switching in speech recognition?

It is changing language within a conversation, sentence, or phrase. The recognizer tries to preserve each word in the language spoken.

Sources

Thinking in one language and writing in another?

See how a direct cross-language output workflow differs from preserving a mixed-language transcript.

Explore Ghost Mode