A few weeks ago in Dublin, a slide went up in front of a room full of localization leaders: “The localization playbook is dead.” The same week, another session made the opposite case: the localization manager isn’t going anywhere. Both rooms applauded. Two claims that pull in opposite directions and still hold; that’s usually a sign the interesting question is hiding one level down.
“AI won’t take your job, but someone using AI will.” The line is usually traced to a remark by economist Richard Baldwin at the 2023 WEF Growth Summit, and it has been repeated at industry events ever since. Sangeet Paul Choudary calls statements like it “true, but utterly useless”: lines that feel like insight and produce easy agreement, which is exactly what stops people from asking the questions that matter. The slogan version of our industry’s debate (“AI gives localization managers superpowers,” “the human will always be in the loop”) has the same property.
The question hiding one level down is the one I keep offering customers: instead of asking where you want AI, ask where you want humans. The first framing produces a tool list. The second forces you to name the judgments you refuse to delegate, and turns everything else into a design problem. And unlike the slogans, it can be answered with evidence.
Instead of asking where you want AI, ask where you want humans.
The localization job, as presented at LocWorld55
The speakers in Dublin were far past the slogans; they brought deployments, failures, and numbers. What nobody had done was add it all up. So we did: we treated LocWorld55 itself as a dataset. Thirty-eight sessions, covering what practitioners at Dell, Uber, OpenAI, Spotify, Trendyol, Booking.com, Notion and thirty other organizations said and, more usefully, shipped. We worked from the spoken record wherever one exists: transcripts for 34 sessions, about 170,000 words of what was actually said including the Q&A, checked against the published slide decks; most of the recordings are public in the official LocWorld55 Dublin playlist.
We extracted every distinct task and recurring workflow we could attribute to a named enterprise practitioner, tagged each against a written rubric, and counted. Vendor and tool-provider pitches stayed out of the dataset: they describe the job they’d like to sell, not the job being done. The result: 380 tasks and 107 recurring loops, condensed into 43 canonical tasks that describe the localization job as it exists in mid-2026. One thing practitioners described has no slot on that map at all, and what it is says a lot; it arrives at the end of this post.
Intento was there too; my colleague Katya Syromyatnikova moderated the Buy-In Playbook panel with Workday and Indeed. I’ve spent a good part of the weeks since inside these transcripts.
In the final count, one task in six runs on AI end-to-end today, and 29% of the job never touches a model. Both Dublin claims hold at once, and the book Reshuffle explains how.
A job is a bundle of tasks
Reshuffle makes an argument that maps almost perfectly onto what we found. A job, Choudary writes, is “a bundle of tasks… they travel together because, historically, they were cheaper to keep in one person than to split.” AI unbundles the bundle. Drafting becomes a component; so do lookup and classification, and each piece can go to a tool, an agent, or a person. The job then rebundles around the constraints that remain: in his framing, human judgment, AI orchestration, and risk and verification, “because AI can’t audit itself.”
The book’s interactive companion has a simulator where you pick a role and drag an AI-capability slider to watch the bundle come apart. We effectively ran that simulator on field data for one profession, with practitioners as the data source and a conference as the sampling frame.
The method has rough edges, and they’re worth knowing before you trust the numbers. Three sessions remain slides-only and one yielded nothing extractable; everything else now rests on the spoken record. Twelve organizations appear in more than one session, so their repeated numbers are one story told twice, not two confirmations. We used Claude to do the extraction and tagging against our rubric, then reviewed the output; anyone curious about the method can reach out to me. And everything here comes from public talks, published slides, and session recordings, quoted with names attached, because that’s what makes it checkable.
The unbundled job: 43 tasks
Translation execution, the thing outsiders assume the job is, accounts for three of the 43 tasks. The rest is quality standards and shipping decisions, systems engineering, terminology and data work, cultural validation, governance, vendor management, program mechanics, and a large business-facing block: building cases, framing metrics for stakeholders, negotiating trade-offs, earning trust.
Translation execution, the thing outsiders assume the job is, accounts for three of the 43 tasks.
AI didn’t shrink this job. It exposed it. Erik Bremer, who runs content engineering at Dell (3 billion words a year across 28 to 30 languages, 80% of it fully machine-translated back in 2023), told the Dublin audience that AI “exposed every weakness in the systems we’d built around localization.” His executives had no idea the operation existed until the hype cycle made them ask why it did.
Choudary would call the job title itself a historical accident, “an artifact of organizational design, built around coordination problems.” Our map isn’t a job description. It’s an inventory of the coordination problems that currently cluster around one chair.
The second axis: 107 loops
Tasks are one axis. The corpus carries a second: 107 recurring loops, the cycles the work runs in. A task is a responsibility someone owns; a loop is a circuit that repeats, like the one where a quality signal arrives, gets triaged, produces a fix, and changes the next signal. We sorted those into eight families. Quality measurement and diagnosis is the largest at 20 loops, stakeholder and value work accounts for 18, and agent-fleet operations has two, both from a single session, which is roughly where that practice stood in mid-2026.
The two axes cut the same evidence differently rather than nesting inside each other. The stages of a loop are whatever the speaker described, in their own words, while the 43 canonical tasks are the vocabulary we built to compare across sessions. Sorting by family produces a finding of its own. 54 of the loops run the pipeline: production, quality measurement, the data supply that feeds it, and the agent fleet. The other 53 run the organization around it, from winning the mandate to responding to teams who localize without asking. Half the job runs the pipeline. The other half runs the organization around it, and that half is where the human work stays.
Half the job runs the pipeline. The other half runs the organization around it, and that half is where the human work stays.
Mapping every stage of all 107 loops against the canonical tasks puts a number on how the two axes meet: a loop invokes three canonical tasks at its median. About a quarter of loop stages are not work at all, but the events, states and outcomes that make the cycle turn, like a threshold drifting upward or a stakeholder forgetting a favour.
The organization, not the pipeline
Counting which tasks and cycles the sessions actually dwell on gives the map a relief. Coordinating stakeholders is mentioned in 18 of the 38 sessions and measuring quality in 17; running the translation itself is discussed in 8. The most widely shared cycle is terminology and knowledge supply, described in seven separate talks. The conference spent its stage time on the organization around the pipeline, not on the pipeline.
Explore both axes
Start with a family of loops, drill into one to see its stages and the canonical task each stage performs, then into any task for its definition, boundary and evidence. Filters narrow every pane at once.
Open the explorer full width in a new tab →
The headline number and its trap
Across all 380 observed tasks: 16% can run end-to-end on AI today, 51% split between human and AI, 29% stay fully human, and 4% are things teams want from AI but can’t make work yet (chief among them the feedback loop that would turn every human edit into permanent system improvement; more on that below). The verdicts come from the speakers themselves: what they’ve deployed, where humans still sit, what they call unsolved. We only did the tagging.
The mix of sources handed us an accidental experiment worth reporting. Tag the slide decks alone and 21% of the job looks automatable end-to-end; add the spoken record and the share falls to 16%, with the human-AI split growing from 46% to 51%. Slides describe the system as designed; the Q&A keeps adding people. Session after session, the clean pipeline diagram from a deck acquired, in the telling, a calibration task here, an escalation path there, a standing audit behind the dashboard. The pressure that produces the gap even has a sound: Christian Wecke, who runs language technology at DHL, told the room how often his team hears “just AI it, right, leave me alone.” Discount automation claims accordingly whenever the evidence is a slide.
The human share is the number that never moves: 29% in the slides-only tagging and 29% across the full spoken-record corpus. We treat it as the most reliable number in the study: the relational, political, and accountability core of this job is a fixed fraction of it, not a residue waiting its turn to be automated.
Slides describe the system as designed; the Q&A keeps adding people.
The contested middle
The split also isn’t uniform across the map, and the disagreement is data too. Plot every task against where the sessions place it and the corpus reads as a spectrum: at one end Run translation, which seven sessions call AI end-to-end and almost nobody argues about; at the other, coordinating stakeholders and redefining roles, human by consensus. In between sits the contested middle — measuring quality, building tooling, injecting context — where different companies report genuinely different automation levels for the same task, which is exactly where the interesting engineering decisions live.
“Can AI do this task” turned out to be the wrong question, because the system around the tasks is restructuring at the same time; we spent a while stuck on it. The verdicts stayed mushy (“51% partial” tells you nothing) until we switched to a different question: where does the human sit relative to the AI loop? The corpus gives that question five answers.
The five positions of the human
“Partial automation” covers five structurally different arrangements, with different bottlenecks and different economics. We numbered them P1 to P5, from the most human position to the most automated. The vocabulary is half borrowed: “human in the loop” and “human on the loop” are standard terms from machine learning and autonomous-systems research; “under,” “before,” and “outside” are our extensions, added because two prepositions weren’t enough for what the corpus showed. (A caveat for readers who know the autonomy literature: there, taking the human “out of the loop” means full automation; our “outside the loop” marks the opposite pole, where the deliverable itself is human.) The quickest way to tell the five apart is what we may call the holiday test: imagine the humans take a week off, and watch what happens to the work.
The holiday test: imagine the humans take a week off, and watch what happens to the work.
What a week off reveals
P1, the AI-assisted human act. The deliverable is a human act: persuasion, negotiation, a story told to a skeptical CFO. AI takes over much of the prep, but take the week off and the task simply doesn’t happen. Teresa Toronjo of Malt gave a whole talk that lives here, four years of coalition-building to get localization onto the C-suite agenda, re-arguing the case each time a sponsor left. The general manager of global sales handed her the sharpest line of the campaign: “it feels like they told us to go to the moon, and they gave us a boat.” Six months after the global platform finally launched, it was the fourth revenue generator among 17 domains, and she was already re-fighting for the budget.
P2, in-loop judgment. A human gates every item before it moves. A week off here stops everything. This is where classic review lives, and it’s the expensive position, because cost scales with volume.
P3, human-fed. The AI runs, but it consumes a perishable human-made input: terminology, annotation, SME sign-off, in-market knowledge. Output keeps coming through the week off, but it decays as the feed goes stale. Hameed Afssari of Uber named the binding constraint: “data is the gold mine,” and for low-resource languages there isn’t enough of it, which is why his quality-estimation models still need humans and will until that data exists.
P4, design-then-run. Humans define the standards, thresholds, prompts, personas, and routing rules; AI executes them per item, at scale. Take the week off and flow continues fine, right up until the design needs changing. The quality program at Spotify is the cleanest example from Dublin: language tiers, a templatized evaluation harness (the rig that scores every item against the standard), and a stated division of labor straight off their slide: “automation handles volume — humans fine-tune the critical signals.”
P5, exception handling. AI decides by default; humans see only the exceptions, the clusters, the audit samples. The week off barely registers at first, and quality drifts while nobody’s auditing. Trendyol pushes 8 billion words a year through the pipeline, 99.5% machine-produced, and its linguists review only the second-fail exception queue.
These five positions line up with the three constraints Reshuffle expects jobs to rebundle around: AI orchestration is P4; risk and verification are P5 and P3, the audit anchor and the data feed; human judgment is P2, shrinking toward P5, plus the ship decision itself. The data adds a fourth constraint the triad doesn’t name: the business layer, P1, which at more than a quarter of the subtyped partial tasks is too big to leave off the map. Somebody has to sell the change to the people funding it.
The money is in the migration
The migration between positions is where the money is. Of the 195 partial tasks, 115 were specific enough to subtype. P4 leads at 38%, and P4 drifts toward full automation as designs stabilize. P1 is next at 28%. P2, the expensive position, turns out to hold only 10%, which means the industry’s most valuable move, converting P2 into P5, is already well underway. Flo Health showed Dublin the receipts: Indira Lorenzo and her team replaced review-by-default with review-by-exception on push notifications, cut time-to-market from 43 days to 3, removed 70% of manual handovers, and found that low-risk content came out 5% better than under the old regime. The reviewers hadn’t been adding quality. They’d been adding queue time.
The reviewers hadn’t been adding quality. They’d been adding queue time.
Booking.com showed the far end of the same migration, and its newest pathology. The pipeline localizes 79 billion words a year, and “the automation rate is 99.9%,” Mikolaj Szajna told the room: quality estimation scores everything, automated post-editing fixes what it can, and humans see only what still falls below threshold. Except the humans overedit. Their perfectionism flows back into the system, and “the quality estimation model just starts to apply a higher quality threshold than is needed,” so the team now manages threshold drift as a standing calibration task. Even at 99.9%, someone has to watch what the exception handlers teach the machine.
Find the gate, find the human
The loops make the same point from the other side. Measure quality appears inside 44 of the 107 loops, more than any other task and well ahead of running the translation itself (26): whatever the cycle is nominally about, it runs through quality measurement. And across all 107, thirty-nine stages are gates: the moment work passes, stops or forks. A gate on every item is P2; a gate on exceptions only is P5. Gates run more than three times denser in the partial loops than in the human ones, which is another way of saying that the position of the gate is the position of the human.
One caution before you tattoo P1–P5 onto your org chart. These labels describe today’s loops, and the loops themselves get redesigned. Flo didn’t speed up review; it removed review-by-default from the workflow entirely. Position labels are a snapshot, not a destination.
Coordination, not automation
When Malcolm McLean introduced the steel shipping container in 1956, everyone read it as an automation story, faster cranes and fewer dockworkers, and everyone missed the point. The box standardized the interface between ship, train, and truck; what it sold was coordination, and on top of that coordination the entire modern trade economy got built. Reshuffle opens with this story and compresses it into one sentence: “AI is not a tool of automation. It is a mechanism for coordination.”
Every deployed win in our corpus fits that reading, and none of them is “faster translation.”
The builders
Maja Nebes at ServiceNow got tired of translation-memory cleanup being a five-day, $2,000 exercise (that’s what updating two terms across the TM used to cost her team), so she built her own engine with Claude Code. It scans the whole TM, retranslates in context, produces an import-ready file, and did the job in two hours with zero complaints from reviewers. Her summary to the room: “it actually does everything a vendor does.”
Zoetis attacked the other end of the pipeline. The team led by Ian Zhang found terminology drove 66% of their evaluation errors in veterinary diagnostics (“IM” means internal medicine or intramuscular depending on context, and a glossary can’t tell you which), so they built a coaching layer that detects tricky terms and feeds the LLM the concept, not the string. Critical errors dropped 77%.
DHL condensed its build-versus-buy answer into five words, “buy the core, build the brain,” and wired its own LLM API into the CAT tool. Coca-Cola Europacific Partners went furthest. By the end of 2023 its reviewers were asking not to be assigned tasks in the TMS at all, taking the source files elsewhere to finish the job with AI, and the Vice President Digital told Elitza Dublewa-Servatius, who leads translation there: “maybe we just don’t need any technology for translation and localization anymore… This is a dead industry.” Her answer was an internal platform “as simple as ChatGPT” for 41,000 employees; usage tripled against the old solution. That’s the same slide deck that declared the playbook dead. The team that wrote it describes its own trajectory as “from no longer necessary to extremely necessary.”
And Nicolas Jadot of Trendyol, whose team processes around a million pieces of content an hour, gave the corpus its sharpest sentence: “Providers can generate the output. Your system must decide what passes.”
Coordination without consensus
Shadow localization, the conference’s favorite anxiety, reads differently through this lens. AI enables something the container never did: coordination without consensus. The container needed every port to agree on ISO dimensions before anything moved; AI learns each silo’s language and translates between them without asking permission. That’s precisely what marketing teams at Dell, Uber, and KAYAK did, routing around the localization function because suddenly they could. These are not scattered anecdotes: the discovery-and-reinsertion cycle they describe is the same loop, told independently in five different talks, and 87 of the 107 loops in the corpus recur across talks the same way.
Every measured response that worked took the same shape: don’t police the rogue coordination, become the layer it plugs into. CCEP built the tool everyone actually wanted. The KAYAK team inserted itself into the creative workflow it couldn’t block. At Dell, Bremer says his stakeholders now come back to him because his team runs the quality apparatus they lack.
Anna Norek of Cloudflare compressed the whole cycle into one market: one of their sales teams she had spent three years growing with localized collateral stopped coming, because “they had just developed this LLM functionality, and they are just using translations on their own.” Her conclusion describes the standing condition of the job: “keep knocking at those doors. Because people, even grateful, they just forget.” And at Notion, the discovery step itself has been delegated: an agent the team calls the Slack Ghost watches company channels for localization-shaped activity and flags it early, so the team joins conversations it used to find out about after shipping.
Don’t police the rogue coordination. Become the layer it plugs into.
Coordination runs upstream
At KnowBe4, Tyler Balding and Priscila Mottola describe localization as a diagnostic system rather than a delivery phase. A flagship training module drew feedback comments from 127,662 of its 653,236 learners; AI compressed that pile into a single page of signal, and the signal said that complaints arriving as “the German synchronization could be better” and “painfully repetitive” in French were all pointing at one structural gap in the English source. One source fix resolved the friction in every language at once, where the old reflex would have produced thirty per-language patches. The question they now make every team answer first: is this a translation problem or a source problem?
A second KnowBe4 session supplied the measurement underneath the reflex: when the team classified the survey comments by topic, linguistic quality turned out to be under 4% of what users of localized content mention. “Localized content is just content,” Karina Zidan concluded, and outside a short list of non-negotiables each QA specialist now makes the final call alone.
The newest interface: coding agents
The newest coordination interface isn’t for people at all. Most code discussed on the Shifting Left panel is now written by coding agents, which relocates internationalization work from training developers to instructing machines. Hilary Atkisson Normanha described the three-layer setup at Spotify: i18n guidance in the agent prompt, an MCP server the coding agents query for Unicode, CLDR and locale-fallback libraries, and a linter that catches what slipped through and opens automatic fix PRs.
Oleksandr Pysaryuk of GitLab, who was once personally the reviewer blocking bad merges, pronounced that gate dead: “you won’t be able to block anything ever when it’s not humans who are doing things.” At GitLab, a solutions architect who doesn’t write code translated the entire UI into Turkish in about three hours by writing a spec and pointing agents at it. A program manager on the same team, not trained to write code either, shipped an experiment that moved the website language selector from bottom to top, “and the graphs spiked.” Pysaryuk again: “She didn’t ask for permission. She just went for it.” The panel converged on pure container logic: publish your rules in a form machines consume (glossaries, specs, AGENTS.md files) and whatever agent shows up next builds on them instead of around them.
What erodes, and what persists
The Reshuffle chapter with the most to say to localization leaders is the Kodak and Shutterstock comparison. Both faced the same shock, commoditization of the thing they sold. Kodak defended its scarcity (film) and died with it. Shutterstock noticed that the problem its customers had wasn’t access to images but workflow, rights, and brand governance, rebundled around that constraint, and when AI made image generation free, the constraint got more acute and Shutterstock got more valuable. The whole case compresses into one line: Kodak owned a scarcity, Shutterstock owned a constraint.
Nova Patch, director of internationalization at Shutterstock, presented at LocWorld55, and their talk was entirely about the constraint layer: identity, inclusion, and policy in AI output. The company that made the textbook move sent its localization leader to speak about exactly the layer that persists, and hardly anyone clocked the symmetry.
What erodes: the per-word operating model
Now run the same movie on the localization function inside an enterprise. Erik Bremer ran the Dell operation as localization-as-a-service, with what he described as industry-leading cost per word; 128 stakeholder teams used it voluntarily because it was the cheapest path to translated words. That was a scarcity position, and his own verdict in Dublin: AI “has completely broken that operating model.” Production is now abundant (at Trendyol, 99.5% of eight billion words never touch a human; at Booking.com, 99.9% of seventy-nine billion), and the commercial logic follows; on the same stage, Hameed Afssari told the vendors present that with volumes rising and budgets falling, per-word pricing “has got to go.” Renegotiating that model is one of the 43 tasks on the map, and it sits with the manager.
Booking.com measured what the abundant production is worth by switching it off: localized content disabled entirely for six languages, on 10% of traffic, for almost two months, and the extrapolated impact was “a significant chunk of the total revenue of the company.” Szajna added the part that says more about the profession than the number does: “The only reason we were allowed to do that was because, at the time, the company did not believe there was much value. So basically they said: go ahead. So we did. I don’t think we will be able to do that anytime soon.”
What persists: the constraint layer
The constraints that persist are visible all over the corpus, and they are why 29% of the observed tasks stay fully human. Trust and the ship decision: Kathy Mok of OpenAI reframed quality away from error counts to a single question, “would a local user trust this enough to continue?”, with review ending in a decision (fix, ship-and-improve, or ready-to-ship) rather than a score. Governance and compliance: ZEISS and partners walked the room through release gates and decision rights on a worked med-tech launch scenario. Data and terminology: at Zoetis, terminology sits at the base of the whole stack, the “single source of truth” that grounds LLM output, and it’s the same layer that cut their critical errors 77%. And the control layer itself, per Trendyol above.
SAP showed what rebundling around the constraint looks like inside an enterprise, and named what forced it: customer perception. “Why can I ask ChatGPT to translate something for me, but I cannot have SAP documentation translated?” is the question Maite Dupont says made the old static language strategy untenable. The LX2030 strategy that replaced it tiers everything: four language tiers, down to a fourth where customers run AI self-service translation; content-type tiers on top; automatic quality metrics calibrated against human-translated golden standards; and risk assessed as probability times impact, with the deciding question “in the event of failure, what would have the biggest negative impact?” Maite Dupont summarized where her experts now spend their time, and it reads as the constraint thesis in one sentence: language experts “curate datasets and define the quality frameworks that guide AI-only translation.”
Where the reshuffle stalls
Four hard limits showed up in the data.
AI can be confidently wrong at industrial scale. Trendyol ran a model to clean up messy source titles before translation; it looked fine on the benchmark, shipped, and then hallucinated product names in production. The cleanup meant retranslating 3.07 million titles across 11 languages and purging the contaminated translation memories. The team now tracks “AI generation risk” as its own category, next to input risk and output risk.
AI can be confidently wrong at industrial scale.
The low-resource wall is structural, not temporary. Quality-estimation models need training data; long-tail languages don’t generate enough of it; so the confidence scores sag and humans fill the gap. Uber and SAP both landed on the same architecture from opposite directions: tier the languages, automate the top, and staff the tail with people who are getting harder to find, since (per Afssari) the market for Somali linguists with AI skills is not deep.
Some human gates are law, not habit. Valentina Turra of Philips personally reviews every pull request that adds or changes a string in the patient-monitor UI (“nothing can enter without my approval”; the developers call it the Gandalf review), because on hospital units generating 500 to 4,700 alarms a day, a mistranslated abbreviation is a patient-safety event, not a typo. Amanda Thoniati of Evinova, the AstraZeneca digital-health arm, described the harder version: once a clinical-trial app goes live its content is frozen, and any fix, even a typo, means a fresh ethics-committee submission. “It’s not that we find a typo, we fix it, we publish it — that’s it. That doesn’t happen for us.” These review seats don’t convert to P5 with better models, because the binding constraint isn’t model quality. It’s regulatory architecture, and it moves at the speed of regulators.
Almost everyone in the corpus wants the compounding feedback loop, where every human edit makes the system permanently better. Dell, Uber, and Flo Health all called it unsolved. OpenAI, SAP, and Zoetis described theirs as running. Read closely and the contradiction dissolves into scope: the loops that run capture decisions, fix sources, and tune validation rules. The loop that doesn’t exist yet is automatic model retraining that compounds without human curation. Booking.com added a warning from the far end of automation: a feedback channel left unattended can run backwards, as its overediting reviewers proved by silently ratcheting the QE threshold. Verification stays a constraint for as long as the system can’t audit itself, and constraints are where the value sits. Whether that loop closes in the next two years is the single variable that most changes every number in this post.
The rebundled job
Put the pieces together and the localization role that’s forming has five faces, each visible in Dublin with a name attached.
The decision owner, who sets the quality bar and makes the ship call: the Shippability framework at OpenAI, the launch-iterate-block gates at Spotify. The orchestrator-builder, who owns the pipelines, the agents, and the control layer: Nebes, Jadot, the DHL and CCEP teams. The data and terminology steward, who curates what the machines consume: the concept datasets Zhang built at Zoetis, the golden-standard corpora SAP maintains. The cultural compass, the correction layer for what models get wrong about people: the KAYAK screening for stereotype-laden imagery, the identity work Patch presented. And the business translator, who converts all of the above into budget: the four-year campaign Toronjo recounted, Daniel Gonzalez of Workday closing the buy-in panel with “you don’t really learn until you click Publish.”
Two companies already run it
At two companies this job already exists, reached from opposite directions.
Booking.com arrived at it by restructuring. The pipeline localizes 79 billion words a year with a 50-person localization team at that 99.9% automation rate; over two years the in-house translator group went from 250 people to 15, and the team then spent nine months hiring into 14 new roles that do no translation and no QA. Szajna describes the design goal as creating “a big entrepreneurial gap, and kind of recreate the market pressure”: localization leads own a market end to end, backed by data science, and have to build the business case for every intervention they propose. The new job description reads like the business-facing half of our task map with the production half deleted.
Notion arrived at it by delegation. The International Experience team runs a portfolio of named agents (the Terminator for terminology scouting, the Locale Looker for per-market metrics, the Slack Ghost for early detection) that absorbed the operational layer of the job; Mercedes Krimme reports roughly 70 percent average time saved on the tasks the agents took over. The freed hours went upstream: A/B-testing CTAs against the company metric of workspaces created, influencing what features get named in English, “continuing to just move further and further upstream.” She names the skill shift directly, from localization program management “really into AI program management.”
And her team supplied the one task in the whole corpus that our map had no slot for: governing the agents themselves, the rules for what they may do autonomously as other stakeholders start building their own. “It’s probably going to be something that we need to focus on quite soon.”
Execution and governance
Any job divides into two kinds of work: execution, which produces the thing, and governance, which decides what good means, what ships, what gets funded and who answers for it. The corpus splits unevenly along that line. Twelve of the 43 canonical tasks are governance, and not one of them carries an end-to-end verdict. Two of the 99 mentions behind those twelve describe AI doing the work alone, against a fifth of the mentions behind the 31 execution tasks. The ratio inverts on the human side: 58% of governance mentions say the work stays human, against 20% on the execution side.
Governance is not a layer sitting above the pipeline, which is why automating the pipeline leaves it standing. Three of those twelve tasks live inside production itself: defining the quality standard, deciding what ships, and choosing the platform the work runs on. A quality gate is a governance act in an operational uniform, and it is among the most stubbornly human things in the corpus. Booking.com deleted the production half of the job and rebuilt what was left around market ownership and business cases. Notion handed the operational layer to named agents and moved the freed hours upstream. Both teams shed execution and kept governance, and Notion then found itself short a governance task nobody had written down.
AI took the execution half of the bundle. What it handed back was more governance than the old job description had room for.
Where the human work moves
Kathy Mok gave the room a two-word example of where the human work moves. OpenAI runs A/B tests on the “Contact sales” button; instead of asking for a retranslation, her team asks people in-market whether they’d click it and what would make them click, with “would your mother buy this?” as the working test. Changing those two words produced lifts in growth metrics. Her conclusion: humans will always be in the loop, “but their roles will be very different.”
One line from Choudary fits this profession unusually well: “You don’t need an AI strategy. You need a strategy for the conditions AI creates.” Applied to careers, the same logic says a skill is worth only as much as the constraint it resolves. The question for a localization professional in 2026 isn’t which AI course to take. It’s which constraint you intend to own.
The question for a localization professional in 2026 isn’t which AI course to take. It’s which constraint you intend to own.
What to do with this, by role
Back to the question at the top. Once you know where you want humans, the 43 tasks are the menu to design against, and the evidence says where your peers have already put them.
If you run a localization team
Inventory your own bundle against the five positions. Count your P2 hours, then convert them to P5 the way Flo did, with risk gates and audit sampling rather than blanket review; the data says quality holds while cycle time collapses, and the measured P2 share of 10% says your peers are already doing it. Daniel Gonzalez showed how simple the tiering conversation can be: marketing copy has to be pristine, an FAQ can afford an awkward sentence, and the hard line is anything that blocks a user from completing their task.
Guard your P3 feeds (terminology, golden data, SME access) like the assets they are, because they’re what makes your AI better than the ChatGPT tab your stakeholders keep open. And start publishing those assets in a form machines can consume, the way GitLab turned its glossaries and translation rules into specs, because the next consumer of your termbase will be a coding agent somebody else built.
If you fund one
The 16% that automates end-to-end is in the data — take the savings. But the 29% that stays human comes from the same measurement, holds in every subsample of the evidence, and sits in the work that prevents the incidents you’d otherwise read about later: governance, cultural screening, the ship decision. Cut it and the work does not disappear, it relocates. Stakeholders still need other languages, they now have their own models, and with nobody owning the standard they publish first and tell you later. That is shadow localization, and Anna Norek of Cloudflare showed how quietly it starts: a sales team she had supplied for three years stopped asking once it could translate for itself.
If you’re a linguist inside one of these teams
The titles named on stage were specific. Uber now talks about AI orchestrators instead of localization program managers, and quality engineers who prompt-engineer instead of QA-only linguists; at KAYAK, linguists are becoming “both market experts and AI experts,” the people who judge whether AI output is publishable and whether it will convert; at Booking.com the new seats are called localization leads and own markets, not words; at Notion the title on the way in is AI program manager. In P-terms: the seats being added are P3, P4 and P5 (feed the system, set its rules, audit its exceptions), while the pure P2 seat is the one being converted away.
Disclaimer
This analysis isn’t disinterested, and you should weigh it accordingly. The bet we made at Intento, years before I read Choudary, sits on the constraint his triad calls risk and verification. Models supply the abundance; our work is turning what a business means by quality into requirements a harness can enforce and audit on every item, with the ship decision left where it belongs. In the vocabulary of this post, that is P4 and P5 built as a product. Reading the evidence from 38 sessions mostly told me the best practitioners in the industry are building toward the same conclusion, each from their own corner. The playbook isn’t dead. It got unbundled, and the people who spoke at LocWorld55 are the ones rebundling it.
Method, briefly: 38 LocWorld55 Dublin sessions (June 9–11, 2026). Sources: 5 provided transcripts, 29 session recordings from the official LocWorld55 Dublin playlist transcribed from auto-captions and cleaned against the published decks (about 170,000 words of spoken record in total), and 32 slide decks; 3 sessions are covered by slides only, one yielded nothing extractable. Corpus: 380 tasks and 107 recurring loops from enterprise-side speakers, extracted with Claude against a written rubric and reviewed, then consolidated into 43 canonical tasks, with the 107 loops classified into eight families.
Items sourced from vendors, consultants, analysts, and investors (72 tasks, 14 loops) and from academia (13 tasks, 3 loops) are tracked but excluded from the counts as non-representative of the buyer-side job. Twelve organizations appear in more than one session. All quotes are from conference talks, session slides, or session recordings. If you spoke at LocWorld55 and we’ve mischaracterized your work, tell me and I’ll fix it; for questions about the method, reach out to me.
Reading that shaped the frame: Sangeet Paul Choudary, Reshuffle and its interactive companion; his essays “The many fallacies of ‘AI won’t take your job, but someone using AI will'” and “How to reimagine your work in the age of AI”. For the task-level approach to studying work: the Stanford WORKBank audit of automation and augmentation preferences across 844 tasks. LocWorld55 session slides are available to conference attendees through LocWorld; the session recordings are public on the LocWorld YouTube channel.