top of page
Undertow_Logo
Undertow · Guide

Can't We Just Use AI?

The move from vibe localization to engineered localization, and what it takes to run a real program.

The web version of the whitepaper Can't We Just Use AI? From ChatGPT chaos to a real localization program in the age of AI, by Nicola Calabrese, founder of Undertow. This page covers the core framework.

The honest answer is yes, and the way you do it decides whether AI becomes the best thing that ever happened to your localization program or an expensive way to damage your brand in eleven languages at once.

Key takeaways
The short version
  • Vibe localization is pasting content into a chatbot, accepting what comes back, and fixing complaints as they arrive. Fine for a Slack message. Fatal for a product launch.
​
  • Engineered localization is AI as a production engine inside a system you designed: language assets as context, quality gates as verification, human judgment aimed where markets are won.
​
  • 95% of enterprise teams already use AI or machine translation, according to Crowdin's 2026 enterprise survey of 152 B2B professionals. Adoption is not the frontier. The gap between using AI and using it well is.
​
  • Most AI quality failures are configuration failures, not model failures.The engine is roughly 10% of the outcome. Everything around it is the rest.​
​
  • Two levels run the program. The localization manager owns the program. A Language Intelligence Lead owns one language. Neither replaces the other.
​
  • AI reliably gets you 80% of the way. The last 20%, brand tone, cultural fit, high stakes content, is the only part competitors cannot copy by subscribing to the same tool.

What is vibe localization?

Vibe localization is describing what you want to an AI model, accepting whatever comes back, and dealing with complaints as they arrive. There is no review against a defined standard, no language assets guiding the output, and no mechanism that stops the same error appearing next week.

​

The term borrows from vibe coding: prompt, accept, paste the error back into the chat when something breaks. No review. No tests. Just vibes.

​

In practice it looks like this. Someone pastes the homepage into a chatbot and asks for it in French. They accept what comes back, paste it into the CMS, and ship it. The campaign goes live. Nobody reviews it. Nobody notices that the product name is rendered differently in paragraph three, or that the tone, warm and conversational in English, has come out stiff and formal in French, because that is what the model defaulted to when nobody told it otherwise.

​

That team is using AI. They would answer yes to the question in the title. And they are accumulating a kind of debt that will not show up in any dashboard they currently look at.

What is engineered localization?

Engineered localization is AI used as a production engine inside a designed system. Language assets tell the model what good looks like before it starts. Terminology is enforced mechanically rather than suggested. Output passes automated checks and sampled human review against a written rubric. Errors are traced upstream and eliminated at the source.

​

The formal name for the framework is the Localization Life Cycle in the Age of AI.

The distinction is not whether you use AI. It is what surrounds it.

Who runs an AI localization program?

The localization manager. The program level is theirs: strategy, meaning which markets and what content to what standard. Technology, meaning the platform, connectors and routing rules. Workflow design, meaning the tiers, SLAs and escalation paths. And stakeholder management, meaning the dashboard and the budget conversation.

​

In a mature setup that is a team. In most companies it is one person carrying every hat, and nothing in this model asks that person to carry more. The model describes functions, not headcount. A solo manager can cover the program level alone, or bring in external capability as an extension of their own role rather than a replacement for it.

​

Whoever holds this level is the architect of the program.

Not everything gets the same treatment

What is a Language Intelligence Lead?

A Language Intelligence Lead (LIL) is the language manager in the age of AI. One per main language: a German LIL, a French LIL, a Japanese LIL. Each owns their language's control levers, runs error analysis, maintains its error library, directs the specialists working in it, and reports quality trends upward.

​​

The German LIL does not review every German segment. They make the system that produces German better every month.

​

The LIL does not run the program. They run a language, inside the program the localization manager runs

Why vibe localization fails quietly

A decade ago, machine translation failed loudly. The output was visibly broken. Today's models fail quietly. The output is grammatically clean, confidently phrased, and plausible to anyone who is not a native speaker.

 

The errors that remain are exactly the kind a non-native reviewer cannot catch. A product term translated generically when the market already knows it by another name. A pun rendered literally into a sentence that means nothing. A Korean speech level that is technically polite and socially wrong for the relationship.

​​

The content looks right. It reads as wrong only to the people it was made for, and they do not file bug reports. They just do not convert.

​

So the damage surfaces late and in the wrong dashboards. Nobody attributes a flat German conversion rate to a campaign vibe localized eight months ago. The CRM says price sensitivity. The post mortem says the market was not ready.​

​

It spread because the barrier to entry collapsed at the moment the pressure peaked. Any marketing manager can build a translation workflow in an afternoon, and every individual choice is locally rational. Nobody designs the aggregate, and the aggregate is where the damage lives: three versions of the brand voice in German, and no translation memory capturing any of it.

​

Tomás · Marketing Operations Lead

His team vibe localized a campaign into Spanish: landing page, three ad variants, the pricing page. He read the output himself. His Spanish was good enough to check that everything was accurate, and it was.

​

​The campaign underperformed from day one. It took a quarter and an external review to find out why. The register drifted from screen to screen. The pricing page used a different term for the core product than the ads leading to it. The call to action was a sentence no native speaker would ever write, correct in every word and wrong as a whole

​

Nothing was mistranslated. Everything was off.

Why it spread

Vibe localization did not spread because localization managers are careless. It spread because the barrier to entry collapsed at the exact moment the pressure peaked.

 

Any marketing manager can now build a translation workflow in an afternoon: a custom chatbot with the brand voice pasted into the prompt, a browser tab that translates anything, a plugin connecting the CMS to a model.

 

Each individual choice is locally rational. Marketing needed campaign copy fast. Support needed help articles in French. Product inherited a machine translation connector nobody remembers configuring. Nobody designed the aggregate, and the aggregate is where the damage lives: three versions of the brand voice in German, no shared terminology, no translation memory capturing any of it, and no way to improve systematically because the work is scattered across tools that do not talk to each other.

​

The spectrum from vibe to engineered localization

The useful frame is not a binary between using AI and not using AI. Everyone is using AI. The real variable is how much structure, verification and human judgment surrounds the output. Most programs sit somewhere in the middle, and position on the spectrum is not a grade. It is a choice that should depend on the stakes.

Spectrum of localization approaches from Vibe Localization to structured AI-assisted and engineered localization, highlighting verification and quality at scale.
A comparison of Vibe Localization, Structured AI-Assisted Localization, and Engineered Localization, across six dimensions including quality, asset usage, error handling, scope, and risk.

The single biggest differentiator is how output gets verified. In engineered localization two mechanisms work together: automated checks for the deterministic parts, and human review against a rubric for register, resonance and cultural fit. Without both, the practice is vibe localization, however sophisticated the prompts are.

​

A team with an expensive model, elaborate prompts and no verification system is vibe localizing. A team with a modest model, enforced assets and a working quality loop is engineering. The model is not the difference. The system around it is.

Language assets are your context

Here is the pattern I see over and over. The AI produces disappointing output, the team blames the engine, evaluates a new model, and the conclusion hardens into "AI is not ready for our content."

​

Examined honestly, most of these are configuration failures, not model failures. A glossary not updated in two years. A translation memory full of pre-rebrand terminology. A style guide written in adjectives a model cannot act on. No context about what the content is or who reads it.

Diagram showing language assets as AI localization context, divided into static context and dynamic context for linguists.

Give the AI what you would give a new linguist on day one.

The encoding work concentrates in three assets. In the research behind my book, Localization in the Age of AI, I started calling them the AI control levers, because that is what they are: the levers that steer output quality before a single word is generated.

​

The glossary is not a terminology spreadsheet. It is explicit instruction: the preferred term, the terms that must not be used and why, and when each applies. "Do not translate X as Y, because Y carries a connotation in German the brand does not want" is far more useful than "X = Z."

​

The style guide governs how the language is written: decimal separators, date and currency formats, guillemets in French, corner brackets in Japanese, spacing in Korean, English loanwords. Written as explicit rules, much of this becomes machine-checkable, which means your automated QA can enforce it.

​

The tone of voice sheet governs how the brand speaks in that market. Du or Sie in German. Which politeness level in Japanese, where politeness is built into the grammar and the wrong level reads as rude or servile. A brand voice defined in English does not map onto other languages by translation.

​

The two get blurred constantly, and the blur is expensive. The test: if the instruction would be true for any brand writing in German, it belongs in the style guide. If it is true only for your brand, it belongs in the tone sheet.

​

One warning that saves real money: stale assets are worse than no assets. A glossary from before the rebrand steers every translation toward abandoned terminology, confidently, in every language at once.

The new localization life cycle

For two decades the loop looked the same: request, scope, hand off, translate, review, QA, deliver. Measured in days, and here is the detail that matters. Almost none of that time is translation. The time lives in the handoffs, the batching, the queues.​

​

Then AI removed the translation bottleneck, and every other bottleneck became visible.

Diagram comparing the traditional localization workflow with the AI-era loop, showing faster production, asset preparation, verification, and continuous localization.

A model translates a batch in minutes, but the batch still waits for intake, routing, review, and for someone to notice it is done. Teams that dropped AI into an unchanged workflow found translation time collapsed and delivery time barely moved.​

​

Two phases carry the weight now. Asset preparation barely existed before and is where quality is actually decided. Verification is the new centre of gravity: automated QA for the deterministic parts, then sampled human review against a rubric for register, resonance and cultural fit.

​

The speed gains are real. In the research behind my book, Localization in the Age of AI, one company went from localization taking days to taking minutes, and another tripled output with the same team. The honest counterweight, from the same research: for many languages, reviewing unguided AI output takes two to three times longer than reviewing an experienced specialist's work.

​

Skip the asset work and AI moves your costs from a phase you were measuring to a phase you were not.

The localization harness

Before the harness, one reset. "AI in localization" is not machine translation, and has not been for about a decade. The assumption leads to a costly pattern: teams buy new AI tooling to do things their existing platform already does, while using a fraction of what the technology covers.

​​

Most localization programs I see are using about ten percent of what is possible. The ten percent is translation.

The other ninety is routing content to the right workflow, extracting terminology before translation begins, scanning source content for problems that will multiply across every language, generating context briefs for linguists, and automating large parts of quality assurance. Most of it needs no new tools. The features sit inside platforms many programs already pay for, switched off or unconfigured. What it needs is workflow design, which is an investment in thinking, not in technology.

​

There is a temptation to treat the model as the system. Output disappoints, the engine takes the blame, and the improvement roadmap becomes a procurement roadmap.

​

In the same 2026 Crowdin survey where 95% of teams reported using AI or machine translation, one in five reported quality incidents or regressions. With adoption that close to universal, the teams having incidents are not running different engines. They are running them with less around them.

Diagram explaining the localization harness, showing how language assets, workflows, platform integrations, and observability support AI localization.

Program = Engine + Harness. The engine belongs to a vendor. The harness belongs to you.

Wired properly, the elements form a loop, not a pipeline. The engine produces output shaped by the assets. Review categorizes the errors. The analysis drives asset updates, which produce better output next batch. A pipeline processes content. A loop gets better every time it runs.​​

​

One detail deserves its own sentence. Enforcement beats reference. A glossary the engine may consult is a suggestion. A glossary the platform enforces before output is generated, blocking forbidden terms and applying preferred ones, is a rule. Every asset that moves from available to enforced removes a class of errors from ever being produced, which is cheaper than catching them.

​

You do not beat rogue workflows with policy. You beat them with a harness good enough that nobody needs to work around it.

Who runs the program, and how

The most effective place to intervene in an AI workflow is upstream, not downstream. Not in the output, but in the instructions. Fixing an error takes minutes. Fixing the instruction that caused it removes that error from every future batch.​​

​

That changes what the program's real product is. The product is not translations. Translations are what the system emits.

Diagram illustrating the localization program model, with localization managers, language intelligence leads, production workflows, QA gates, and verified output.

The system has two levels, and keeping them distinct is what makes the structure work. The program level belongs to localization management. The language level belongs to the Language Intelligence Lead, one per main language. Both can be in-house, fractional, or a mix.​​

​

What the language level cannot be is an afterthought. Handing the title to whichever freelance translator is available does not create a LIL, because the work runs on error categorization, pattern analysis, rewriting assets for machine use, and turning quality data into upstream fixes. Recruit a freelancer per language, each inventing their own method, and you get five languages improving through five incompatible approaches. The language level needs a shared model, so the localization manager can manage the whole rather than five separate crafts.

​

A program with a manager and no LILs has a well-designed factory with nobody calibrating the machines. A program with LILs and no manager has excellent calibration and no factory.

​

There is a second distinction inside the work. In-the-loop mode is segment level: reviewing output, fixing what is wrong. Directing mode is a level up: defining what good looks like before production, maintaining the assets, sampling instead of reading everything. Roughly 45% of a LIL's week goes to high impact review, 25% to the control levers, 20% to error analysis and 10% to reporting.

​

More than half the week is spent making the system better rather than making today's batch acceptable. The post-editor's Monday looks identical every week. The LIL's Monday gets easier every month, because the errors they eliminated upstream stay eliminated.

The 80/20 of going global

AI gets you 80% of the way there. For structured, factual content in well-resourced language pairs, a harnessed engine produces output close to right, and the human role is oversight. Programs that resist automating that tier are spending their scarcest resource on work that no longer needs it.​​

​

Diagram illustrating the 80/20 model of going global, showing AI handling volume, speed, and consistency while humans provide judgment, nuance, and brand expertise.

Here is the strategic point hiding in that split. The 80% is available to everyone. Your competitors have the same engines. Output from an unguided model is, by construction, the average of what everyone else ships. If your German is generic AI German and theirs is generic AI German, language has stopped being a differentiator, and you are competing on price and features alone.

​

The 20% is the only part of localization your competitors cannot copy by subscribing to the same tool.

​

​

It is not a cost. It is the moat.

Yuki · Language Intelligence Lead for Japanese

She reviewed the AI draft of a launch campaign. The English headline leaned on an idiom: "Stop herding cats." The AI translated it accurately, and the Japanese headline became a sentence about cats.

​

Nothing was technically wrong, and the campaign's most important line meant nothing to the audience it was written for. No automated check flags "faithful translation of a joke that no longer exists." Yuki threw the headline away and rebuilt the idea around an image that works for her market.

​

The model got the campaign to 80% in four minutes. Yuki's afternoon gave it a reason to exist in Japan.

That is the answer to the question in the title: yes, for the four minutes. The afternoon is where you win the market.​

​

Quality as a system, not a review step

Ask most programs how they manage quality and the answer is a place: the review step. Ask what the review found last quarter, or whether German is better than six months ago, and the answer is an impression, not a number.

​

What makes review a system rather than an opinion is the rubric: a written definition of what counts as an error, how errors are categorized, how severe each is, and what rate is acceptable per content tier. With a rubric, quality becomes discussable. This batch had two minor terminology errors per thousand words, the threshold for this tier is three, it passes.

Judge the system on the sample, not the demo.

A demo shows the AI at its best, on content chosen to impress, run once. A sample is a representative slice of real content on a fixed cadence, scored against the rubric.​

​

The demo answers "can this work?". The sample answers "does this work on our content?". That distinction matters most when a vendor demo impresses your leadership: the program does not need to argue against the excitement, it needs to put its own numbers on the table.

Diagram illustrating the quality flywheel, showing how localization errors are measured, diagnosed, fixed upstream, verified, and monitored to prevent recurring issues.

One rule keeps the whole protocol honest: change one thing per test window. Update the glossary and the tone sheet in the same week and you will never know which one moved the needle.​

​

Every loop leaves the system permanently better. At one client program we manage, this cadence is a monthly ritual. Two quarters in, German campaign content needed 40% less revision, and the quality report could name the three glossary updates that did most of the work.

The economics, and the CFO conversation

Leadership usually starts and ends with one number: cost per word. That is the wrong number. The right one is the total cost of owning a multilingual presence, including what it breaks.​

​

Vibe localization looks unbeatable in month one. The debt accumulates off the books, in five accounts: the rework tax, the revenue leak, the support surcharge, the emergency premium, and the do-over. On the revenue leak, CSA Research's study of 8,709 consumers across 29 countries found 76% prefer to buy with information in their own language, and 40% will never buy from websites in other languages

Chart comparing the cumulative costs of vibe and engineered localization, showing how upfront investment in engineered localization can reduce costs over time.

Plot the two curves and they cross. Before the crossover, vibe localization is cheaper and leadership is right to ask why you need the structure. After it, the gap widens every month. The strategic error is judging the race at month one, which is exactly where the "AI made this free" narrative judges it.

​

Sofia · Localization Manager

She made the argument twice. The first time she brought a deck about translation quality, coverage and turnaround. The CFO asked about return on investment, got an answer about linguistic consistency, and the meeting ended with a promise to revisit in Q3, which everyone understood to mean never.

​

The second time, French sign-ups were converting at half the English rate. She proposed a six week pilot on the critical path for France, with pre-agreed metrics and a threshold. The pilot cost 12,000 euros. Closing half the conversion gap was worth roughly 43,000 euros in annual recurring revenue per monthly cohort, compounding with every cohort after.

​

​The CFO asked two questions and approved it in the room.

Same program, same person. The first deck asked for a budget. The second offered a bet with a measurable payoff and a bounded downside, and CFOs approve those for a living.

Where to start on Monday

All of this can be built incrementally, one workflow at a time, each step producing the evidence that funds the next. Here is where to start if you run the program. The full paper adds two more playbooks, one for leadership and one for translators and reviewers.

1

Audit your real AI usage, official and unofficial. Every custom GPT, every free MT tab, every forgotten connector. You cannot fix silos you have not found.

2

Build the asset baseline for your priority languages. Glossary as explicit instruction, style guide as machine-checkable conventions, tone sheet per language. Start where the revenue is.

3

Pick one workflow and engineer it end to end. It gives you the before and after numbers everything else will be sold with.

4

Define the rubric before you scale the volume. Error types, severity, thresholds, cadence. It turns "quality feels fine" into a number that survives a budget meeting.

5

Measure, then report in business language. Revision rates, turnaround, and wherever you can get it, conversion and retention by market.

6

​

Deliver the message upward yourself, before a quality incident delivers it for you.

Frequently asked questions

The answer to the question

Can't we just use AI? Yes. Your team is almost certainly using it already, and so is every competitor. Using AI was never the decision. The decision is whether it operates inside a system that makes its output trustworthy, or alongside the wishful thinking that it will be fine.

​

Structure scales. Vibes do not. For the content a company's revenue depends on, the gap between "reads fine" and "right for this market" is where churned customers and quiet brand damage live.

​

AI amplifies the program it lands in. A program with clear strategy, maintained assets and documented process gets amplified in quality and reach. A program without them gets amplified in inconsistency.

​

The human role is evolving, not diminishing. The skills are shifting from producing words to exercising judgment, from translating content to building the system that translates it.

Translation is becoming automated. Judgment is not. Direction, verification, and cultural judgment are the new craft.

Download the full whitepaper

The complete 59 page paper adds the life cycle phase by phase, the quality flywheel, the three playbooks for managers, leadership and linguists, and every figure at print quality. Free, and no form.

Not sure where your program sits on the spectrum?

Undertow works with tech companies as the fractional layer of this exact model: fractional localization management, fractional language intelligence management, and fully managed programs.

​

Talk to us or read more about how we approach AI in localization. For the conversations behind this research, listen to The Multilingual Content Podcast.

bottom of page