A studio microphone tagged RU on the left, headphones tagged EN on the right, and a waveform crossing between them — AI dubbing keeping the same voice in another language

Disclosure: this article contains an affiliate link to ElevenLabs. If you sign up through it, Pickurai may earn a commission at no extra cost to you. I paid for and used the tools I write about here.

I spent the last few days listening to Lex Fridman's conversation with Khabib Nurmagomedov. It was recorded in Russian and released dubbed into English by ElevenLabs, and somewhere around the ten-minute mark I stopped noticing that fact. That is the whole review. The dub stopped being a thing I was evaluating and became a thing I was listening through — which is exactly what a translation layer is supposed to do and exactly what none of them managed until very recently.

TL;DR — Key takeaways

  • The seams are gone. I heard the same pipeline on Lex Fridman's Zelenskyy interview in January 2025 and it worked, but I was aware of it the whole time. Nineteen months later, on a longer and more abstract conversation, I wasn't.
  • Recording in the mother tongue was the right call. Khabib's expressive range in Russian is visibly wider than in English. Dubbing from his first language preserves the ceiling of the speaker instead of the ceiling of his second language.
  • Russian → English is a hard pair. No articles, flexible word order, meaning carried by case endings and verbal aspect. Getting this right is a harder problem than Spanish or French to English.
  • Which is why Spanish → English should already be better. Closer word order, more training data, shorter reordering distance. If Russian sounds like this, the easy pairs are further along.
  • The honest caveat: I can judge fluency, not fidelity. I don't speak enough Russian to verify that the translation says what Khabib said.

The Zelenskyy Benchmark

I'd heard this before. In January 2025, Lex Fridman published his conversation with Volodymyr Zelenskyy with AI-dubbed audio tracks, and at the time it felt like a genuine milestone: a multilingual conversation made listenable to an audience that spoke none of the source languages. I listened to the whole thing.

But I listened to it as a technology demo, not as a conversation. The delivery was flat in places where the original clearly wasn't. Sentences arrived intact but stripped — the information survived, the person didn't. Idioms came through as their dictionary meaning rather than their actual force. Every few minutes something would land slightly wrong and pull me back out of the room and into the software.

A good dub isn't one you admire. It's one you forget you're hearing.

Nineteen months. That's the gap between that episode and this one. It is not a long time in any other industry.

What Changed With Khabib

The Khabib conversation should have been harder, not easier. Zelenskyy in wartime is a man delivering positions — grave, deliberate, and fairly close to prepared speech. Khabib talking to Lex Fridman is the opposite: long digressive answers about faith, discipline, his father, grief, wrestling, money, fear, the specific loneliness of retiring at the top. Abstract, personal, unscripted, and full of the kind of half-sentences people produce when they're actually thinking rather than reciting.

None of it flattened. The pauses stayed where he put them. When he slowed down before saying something about his father, the English slowed down too. The dub carried emphasis — the small rise in energy that tells you the speaker cares more about this sentence than the last one — and emphasis is the thing that used to disappear first.

And it's still recognisably him. The English track keeps his timbre, not a stranger's. That is the difference between a translation and a dub, and psychologically it's most of the effect: your brain files what it's hearing under "this man is speaking" instead of "someone is reading me a transcript of this man."

Why Recording in Russian Was the Right Decision

This is the part I keep coming back to, and it's the part that has nothing to do with the technology.

Khabib speaks English. He has done hundreds of English-language interviews and he gets through them fine. But anyone who has heard both knows his English is a narrower instrument than his Russian — a smaller vocabulary, fewer registers, less room for the precise word. That isn't a criticism; it's what a second language is for almost everyone who didn't grow up in it.

The cost of speaking in a second language is invisible because it happens before you say anything. You don't say the thing you mean; you say the nearest thing you can construct. You round off. You drop the qualifier you can't phrase. You pick the anecdote you have the words for over the better one you don't. The interview you get is bounded by the speaker's vocabulary rather than by their thinking.

Recording in Russian removes that ceiling. Khabib says what he actually means, at full resolution, and the machine takes the translation hit afterwards. It's the right order of operations: a dubbing model losing five percent of a rich answer beats a speaker losing thirty percent of it before it ever leaves his mouth.

Better to translate a complete thought imperfectly than to receive a simplified one perfectly.

I notice this in myself, in the other direction. I write and think in Spanish and I work in English every day, and I am a duller person in English. Not less informed — less expressive. The jokes don't fit. If AI dubbing means people can stop being duller versions of themselves in front of an international audience, that's a bigger deal than the audio quality.

A Few Technical Notes on Why This Is Hard

It's worth being precise about what "dubbing" means here, because it isn't one model doing one job. A pipeline like this has to separate the speakers on the track, transcribe each of them, translate the transcript, generate speech in a voice matched to the original speaker, and then fit that speech back into the original timeline. Errors compound down the chain: a diarisation mistake becomes a translation mistake becomes a voice assigned to the wrong person.

Two stages are where the perceived quality actually lives:

Voice identity. The output is generated in a voice derived from the source speaker rather than picked from a stock library. This is the single biggest contributor to "I forgot it was dubbed," and it's why this generation of tools feels categorically different from the dubbed television most of us grew up with, where every foreign actor shared four voices.

Timing. Translated speech is almost never the same length as the original. The system has to compress or stretch each segment to land back on the original timing without pitch-shifting the voice into a cartoon or clipping the ends of clauses. In conversation this is worse than in a monologue, because two people interrupt, overlap and finish each other's sentences, and each of those handoffs is a place where a dub can fall a beat behind and never quite recover.

Then there's the language pair itself. Russian → English is genuinely one of the harder mainstream directions:

  • No articles. Russian has no a / the. Definiteness is inferred from context and word order, so the model has to reconstruct information that simply isn't in the source signal.
  • Flexible word order. Russian marks grammatical roles with case endings rather than position, so speakers reorder freely for emphasis. English marks the same roles with position. The translation has to re-encode emphasis using a completely different mechanism.
  • Verbal aspect. The perfective/imperfective distinction is built into nearly every Russian verb and has no clean English equivalent — it usually has to be paraphrased into tense or an adverb.
  • Reordering distance. Because the decisive word often arrives late in a Russian clause, the system can't translate locally and stay accurate; it has to hold the clause and re-emit it in English order, which is exactly what makes timing hard.

None of this is unsolvable, but it's a longer path than the one from a language that already shares English's clause architecture.

Which Is Why Spanish → English Should Be Even Better

Here's my prediction, and it follows directly from everything above. Russian is not usually first in line. The pattern across this whole field is that English, Spanish, French and German get the training data, the evaluation sets and the product attention first, simply because that's where the corpora and the customers are. Russian arrives later and with less material behind it.

So if Russian → English sounds like this today, then Spanish → English — same clause order, articles in both languages, shorter reordering distance, and an enormous amount of parallel data — should be further along still. I intend to test that properly on my own recordings rather than assume it, but the structural argument points one way.

That matters for anyone producing content outside English. The default advice for a decade was: record in English, reach everyone, accept that you'll sound worse. That trade-off is expiring.

Where It Still Shows

I'm not going to pretend the thing is finished. Three places where I still caught it:

Humour and wordplay. Jokes that depend on the sound or the double meaning of a specific word don't survive any translation, human or otherwise. When one landed on the English track, it landed as a fact rather than as a joke.

Culturally specific vocabulary. Religious and regional terms tend to come through as their closest English approximation, which is correct and slightly hollow at the same time. Some words are load-bearing in one culture and merely descriptive in another.

Fast overlapping speech. The rare moments where both speakers talked at once were the only places the illusion thinned.

And the caveat that matters most: I can judge fluency, not fidelity. My Russian isn't good enough to verify that the English track says what Khabib actually said. What I'm reporting is that it sounded like a real person having a real conversation, which is a claim about the output, not about its accuracy. Someone bilingual should audit that, and until they do, everything above is an impression rather than a measurement.

The tool behind it

ElevenLabs — Dubbing That Keeps Your Own Voice

Upload a video or an audio file, pick a target language, and get back a dub in a voice built from the original speaker's. There's a free tier, so you can run one of your own recordings through it and judge the result on your own ears before paying anything.

Try ElevenLabs free →

The Part I Actually Wanted to Say

I've written before about using ElevenLabs for my own YouTube and TikTok voiceovers, and that was a story about production convenience. This isn't that. This is about access.

For as long as any of us have been alive, the price of being heard internationally has been paid in English. If you didn't speak it well, you were either translated by a stranger's voice or you weren't heard at all — and an enormous amount of what people know, believe and have lived through never crossed a border because of it. The prediction that AI would eventually dissolve that barrier has been made so often it became background noise. Listening to a Dagestani fighter talk about his father in his own language, in his own voice, in English I understood completely, is the first time it stopped feeling like a prediction.

So: thank you to ElevenLabs, and thank you to Lex Fridman for recording it in Russian instead of taking the easy option. That decision is why the conversation was worth listening to, and the technology is why I got to.