Olga Beregovaya — VP of AI at Smartling, who entered natural language processing in 1997 when rule-based machine translation still ruled — interviewed by Angelina on TwoSetAI (73 min). She has spent 25+ years watching the discipline get rebuilt, and thinks the next thing to go is the source text itself.

From hand-built lexicons to transformers

  • Two and a half linguistics degrees, and she never worked as a translator — she studied structure: phonology, syntax, morphology, historical grammar. “Once you understand the structure of the world’s languages, you can extrapolate from there”
  • Her first systems were rule-based MT (Systran, ProMT, AppTek): build a massive lexicon with part-of-speech tagging and metadata, parse the input, then use transfer rules to rebuild the sentence under the target grammar
  • Transfer rules are not word matching — they map grammar to grammar. Going from English into a morphologically rich language like Russian means building whole systems of cases and conjugations
  • The fatal weakness was linguistic coverage: “If you don’t have the word, you don’t have the word” — the engine returns the source word or a random one. Transfer rules are also where most mistakes happen
  • Adding a language was a one-to-two-year endeavor, which is why early engines shipped with a tiny language set
  • Statistical MT (late 2000s) won on quality because n-grams come straight from the corpus with no reconstruction — but the winning systems were hybrids, keeping the parsers and augmenting them

“Translation might die” — the one-to-one mapping does

  • Transformer models were never designed for translation; it is one task among many. “If the same system is supposed to fight your parking ticket, solve your math problem, and then do 10 other things, what are the odds of that system excelling at a peripheral task?”
  • What she expects to die: source→target, text-to-text, the bitext and phrase table — “That paradigm is not going to live for long”
  • What stays: global content. The source does not disappear, it changes form — a marketing brief, a SKU table, a brand book, persona information. You generate from it instead of mapping from it
  • Two reasons to skip the source: pace (you don’t wait for headquarters to produce the assets first) and quality (the source’s English phenomena stop shaping the target, so the output can be more native and more culturally accurate)
  • The trade: translation is predictable because the source is a point of truth. Generation requires trusting the model not to invent
  • Where she will allow it today: blogs and low-level marketing collateral that carries no liability. Legally binding or brand-damaging content still wants the translation path and its guardrails
  • Her timeline for translation-as-a-discipline as we know it: “I’m still giving it two, maybe three years”

Hallucination is a long-tail problem

  • Foundation models are skewed to English in vocabulary and in culture. The less training data a language has, the poorer the lexical coverage and the higher the odds of hallucination — and the failure can sound perfectly fluent while being culturally and factually astray
  • The meme she keeps returning to: someone asks a model whether a mushroom is poisonous, the model confidently says no, then there is a tombstone and an apology
  • Mitigations she lists: semantic similarity between source and target (classic approaches like LaBSE still earn their place next to LLM-as-judge), external search for factual accuracy, named-entity checks for entities that were never in the source, and blunt length sanity checks — five words in the source do not justify twenty in the target
  • Translation has an advantage other AI domains don’t: when you have a source, you have a point of truth to back-translate against
  • For testing, they lean on gold-standard parallel datasets — representation across domains, string lengths, and densities — fed into models to slice and surface failures; referential metrics when human translation exists, non-referential when it does not

Interpretation outlives translation

  • Interpretation is human speech, translation is text — and text-based translation is the part that goes first
  • Old automated interpretation was ASR → machine translation → text-to-speech: three distinct failure points, or ten if you count the sub-models. Now one multimodal model can handle all three steps
  • ASR accuracy and accent handling is the change she feels personally: she used to avoid dictation and yell at previous-generation bank assistants
  • Application space: healthcare and court interpretation (patient support for non-English speakers is “light years” above where it was), and live events — conference interpretation including sign language, what her colleagues call “event globalization”
  • Human interpreters will have to reinvent themselves; the events that needed 200–500 interpreters no longer do
  • Why interpretation keeps its shelf life: “My parents did not speak to me in language model language. My parents spoke to me in Russian. Your parents spoke to you in Mandarin. We still want to hear our mother tongue”

Linguistic colonization is the open question

  • The pivot language (translate into English, then out of it) is dissolving as models go multilingual — and so is the old trick of pivoting through simplified “international English”
  • But if the model is English-centric and speaks to you in an under-resourced language, are you mostly learning the phenomena of an English-centric world? “Is it eventually going to skew my cultural perception of the world? That’s something I don’t quite have an answer to”
  • She is still pro-language-learning: AI should help people learn languages, not only access information in them — she points to Duolingo’s AI work
  • On the long tail she is optimistic: data-collection efforts for under-resourced languages (Data France with the African Language Lab), smaller datasets synthesized into bigger ones. “If a language is under-resourced, make it fully resourced”
  • The philosophical thread she pulls on: a chatbot is suddenly a friend that is tailored to please you. Model and chatbot addiction is now its own discipline — and she admits to four hours of generating poetry styles with a friend, “so who am I to talk about addictions”

Translateese, machine-translateese, modelese

  • Linguists coined translateese for text you can tell was translated — subtle word choices and style that show it was extrapolated from another language
  • Machine-translateese followed: you can spot text post-edited from MT
  • Modelese is the current one: repetitive patterns, overly polished, overly predictable — and the sharper complaint, “the concepts are often very trivial”
  • Her working heuristic against it: ask the model for 10 ideas and go with the 11th, because the model just produced the 10 most obvious ones
  • Google’s old guideline refused to index machine-generated text unless a human curated it. Detection was trivial then and is not now; current search guidance effectively treats human-looking output as human content
  • She expects cross-pollination rather than one-way contamination: exposure to modelese will change how humans write. Her shortest proof: when did you last use proper punctuation in a text message?

Global content delivery is a plumbing problem

  • Why the industry survives: people still want to buy, and make decisions, in their local language. What changes is the operating model, not the demand
  • The mess is real — terminology trapped in DITA XML, content in Git repos, Marketo, Figma, WYSIWYG HTML, plus multimodality, video, and audio. All of it has to be parsed, recognized, and centralized, with linguistic assets governed and curated
  • That is the argument for a platform: the model is the easy part, the piping is not
  • UX is shifting from translating an interface to designing it natively multilingual — engineering constraints of the target locale known at design time, wireframes and prototypes generated per locale (down to local taboos like the infamous white for Japan)
  • Hyper-personalization multiplied by hyper-localization: adapt the language, tone, and visuals to a persona in a specific geography — then multiply that by 7,000+ languages. More opportunity, and more problems

Build versus buy: the line is “when it matters”

  • Plugging in an API is genuinely fine for user-generated content, low-risk content, and quick translation with no brand or legal consequence — one engineer, four hours, done
  • It stops being fine when there is no hallucination mitigation, no quality monitoring, no data governance, no infosec guarantee that your data isn’t feeding a model provider, and no answer for latency as prompts and context grow
  • The scenario that stuck with her, from a Fortune 50 communications SVP: 300,000 employees means 300,000 voices, and no visibility into how good their translated content actually is
  • Look for a platform when accuracy, brand integrity, and governed data are the point, when a mistranslation means something concrete — “pull this lever or not pull this lever, and the next thing you know you’ve dropped a crane on a construction site”
  • Also when inputs come from many sources and need parsing plus centralization, and when a small team has to ship a lot of global content
  • Her own estimate: buying covers ~80% of cases; the remaining 20% are legitimately build-or-plug-in

How do you measure translation quality?

  • Human evaluation first — linguists label and annotate for quality, which delivers both judgment and the data used to train the quality-estimation models
  • The hard part is that annotators disagree: three of them can’t agree with each other, and one can disagree with what she herself said on Wednesday when you ask her on Friday
  • Classic metrics: TER (translation error rate / edit rate, distance from the human gold standard) and BLEU
  • Her team uses 12 metrics — token-based, character-based, word-based, statistical, NLP-based, and semantic — and a matrix for slicing automated quality
  • Favorites: COMET (credit to Alon Lavie — referential and non-referential, semantic-aware), Google’s Metric X (closest to how a human would annotate), and the MQM error taxonomy as a framework for training eval models
  • On model churn: “You wake up in the morning and 10 new models were just released, each of them claiming superiority… it’s almost difficult to stay sane.” Smartling’s answer is to take your content and return a benchmark breakdown across every dimension, then use the stack to phase a model in or out by performance or customer preference

Where to find her

  • Olga Beregovaya on LinkedIn; Smartling posts regularly on language technology
  • She is hiring now: one junior data scientist and one seasoned language-technology / NLP / data-science expert

“The translation process is still much more predictable. When you generate, you really need to trust the model to not dream things up.”