Over the years we've gotten requests to translate the same document for different customers. Legal documents, mostly. Both parties to the same dispute, each sending us the same contract. This has probably happened to every language service provider at some point.
There's an ethical question there, and a financial one. We tried to make the right call: usually two separate translators, a full wall between them, neither one aware the other exists.
What bothered me most wasn't the ethics. It was the waste. Two skilled translators doing identical work, neither allowed to know about the other. The same document, translated twice, because it had to be.
That memory came back to me recently, because AI translation does something similar to this, at a scale no law firm ever could.

Translation used to be one-to-many. One translator's work served thousands of readers, the same finished translation, reused indefinitely. AI flipped that into one-to-one. Every reader gets a fresh translation, computed just for them, used once, thrown away.
This is worth explaining plainly, because it's the thing that makes the rest of this piece make sense. Translation memory is a database that stores every segment a translator has previously translated, paired with its approved translation. The next time the same or a similar sentence comes up, in a new project, a new client, even years later, the system recognizes it and suggests the existing translation instead of starting from nothing.
We've relied on this for as long as I can remember running a translation business. It's the direct opposite of the two-translator wall from the intro: instead of duplicating work out of necessity, translation memory exists specifically to prevent duplicating work at all.
Generally, no. Most AI translation calls have no built-in memory of a previous identical request, so nearly every session starts at zero. Ask an AI system to translate the same sentence twice, in two separate calls, and it will typically generate the translation fresh both times, unless it's specifically wired into a translation memory system the way human CAT-tool workflows already are.
AI translates the same news article, the same product page, the same regulation, from scratch, over and over, every time someone asks for it. Nothing learned from the last identical request, nothing carried into the next.
Each of those redundant runs burns real compute. GPUs spinning, data centers drawing power, to produce a translation that may already exist from an hour earlier, word for word, for someone else.
A 2026 peer-reviewed study puts optimized, large-scale AI inference at a median of roughly 0.31 watt-hours per query. That sounds small until it's multiplied by billions of daily requests, generative AI systems overall are separately estimated to consume roughly 29.3 terawatt-hours annually, comparable to Ireland's total electricity use.

Neither figure is specific to translation alone, and estimates in this area vary widely depending on model size and setup, worth saying plainly rather than overselling a precise number. But the direction is clear: every avoidable, redundant AI translation call adds a real, measurable energy cost on top of the redundancy itself.
This is where it's worth being precise about what human translation workflows already do differently. Professional translators working with CAT tools use translation memory automatically, across projects, across clients, sometimes across years. It's the direct opposite of starting from zero every time.
| Translation memory (human workflow) | Typical AI translation API call |
|---|---|
| Identical segments are recognized and reused automatically | Most calls have no memory of a previous identical request |
| Knowledge accumulates across projects over time | Each session generally starts fresh |
| Repeat content costs little to no additional work | Repeat content costs the same compute every time |
This is exactly the kind of thing human-in-the-loop translation is positioned to fix, not by replacing AI's speed, but by pairing it with a workflow that already knows how to avoid paying for the same translation twice.
The two-translator wall we built years ago solved an ethics problem at the cost of real, visible waste, two people doing the same job once. AI's version of that waste is invisible, spread across thousands of identical requests a day, but it's the same underlying pattern: work being redone that didn't need to be redone. Accountability was the thing AI alone couldn't replace. Efficiency, it turns out, is the thing nobody's fully rebuilt into it yet either.
Q: What is translation memory?
A: Translation memory is a database that stores previously translated sentences or segments, so identical or similar text encountered again doesn't need to be translated from scratch. It's a core feature of professional CAT tools used in human and hybrid translation workflows.
Q: Does AI translation remember previous translations?
A: Generally, no. Most standalone AI translation calls have no built-in memory of a previous identical or similar request, so the same text translated twice through a typical AI API is generated fresh both times, unless it's specifically integrated with a translation memory system.
Q: How much energy does AI translation actually use?
A: Estimates vary widely depending on model size and setup, but recent research puts optimized large-scale AI inference at a median of roughly 0.31 watt-hours per query. At the scale of billions of daily requests across the AI industry, generative AI systems overall are estimated to consume tens of terawatt-hours annually, comparable to a small country's total electricity use.

Ofer Tirosh is the founder and CEO of Tomedes, a language technology and translation company that supports business growth through a range of innovative localization strategies. He has been helping companies reach their global goals since 2007.
Share:
Post your Comment