The first time a client asked us to run a content pipeline natively in Hindi instead of writing in English and translating afterward, we expected the usual headaches: tone, idiom, getting a human reviewer to sign off on cultural nuance a model can't reliably judge on its own. What we didn't expect was for the same prompt, the same task, the same output length, to suddenly cost meaningfully more and hit context limits we'd never come close to in English. It turns out there's a real, measured reason for that, and it has nothing to do with the model being worse at Hindi. It's the tokenizer. The tax nobody puts in the pricing page A recent paper, The Tokenizer Tax, quantified something a lot of teams building for Indian markets discover the hard way: using OpenAI's cl100k_base tokenizer, the same content in Indian languages takes on average 8 times as many tokens as the English equivalent. Hindi comes in relatively cheap at 4.1x. Malayalam, at the other end, hits 13x. The practical result is that an Indian-language user gets as little as 12% of the effective context window an English-language user gets for the same conversation. That's not a rounding error. If you've designed a RAG pipeline, a long document summarizer, or an agent with multi-turn memory around English-language token budgets, and then pointed the same architecture at Hindi or Tamil content, you should expect it to run out of room, and run up your bill, far sooner than your testing in English predicted. The same paper has a more useful finding buried past the headline number: newer tokenizers close most of this gap. OpenAI's o200k_base cuts the mean tax from 8.0x to 2.1x, a 73% reduction, and the root cause is traceable to vocabulary coverage rather than something inherent to the scripts themselves. Which means the fix, at least for this part of the problem, is mostly a matter of picking the right model family rather than rearchitecting anything. Plenty of teams are still on tokenizers that never got that upgrade, without realizing it's costing them. Tokenization is the easy problem The tokenizer tax is measurable, which makes it the part of this that gets written about. The part that doesn't have a clean benchmark is evaluation, and it's the part that actually slows a project down. In English, you have a reasonable, if imperfect, set of options for judging whether an LLM output is good: other LLMs as judges, established quality heuristics, a large pool of native-speaker reviewers you can hire on short notice. For Hindi, Tamil, or Kannada content, most of that infrastructure is thinner. IBM's IndQA benchmark, built with 261 Indian researchers and linguists across 12 languages, exists specifically because general multilingual benchmarks were mostly testing translation quality, not whether a model could reason through a Tamil proverb or a Hindi literary scenario the way a fluent reader would expect. Few teams outside a handful of labs are actually running anything like it. Most are eyeballing outputs and hoping a native-speaking reviewer catches what's wrong, which turns human review into the bottleneck nobody budgeted time or headcount for. Code-switching breaks the tidy version of this problem Here's the part that a benchmark, however well-built, still won't catch: real Indian users don't write in clean, single-language text. They write in Hinglish, code-switching mid-sentence, transliterating Hindi words into the Latin alphabet, sometimes within the same message. A pipeline tuned and evaluated against pure Hindi or pure English will see this kind of input constantly in production and has usually never been tested against it, because it doesn't show up in either language's benchmark set. We've had to treat code-switched, romanized input as its own test category, not a fallback case, across the client work we do spanning insurance, freight tech, and D2C. It changes what "the model handles Hindi fine" actually means. A model can score well on a clean-Hindi benchmark and still fall over on the messages your actual users send. What this changes about how you'd build the system None of this is a reason to avoid building for Indian-language markets, obviously. It's a reason to stop assuming your English-language architecture and budget transfer over with a translated prompt. A few things we've learned to do differently: Budget context and cost per language, not once. If a RAG pipeline is tuned for a context window that works in English, check what it actually holds once the same documents are in Hindi or Tamil. The tokenizer tax means it might be a fraction of what you assumed. Pick tokenizers deliberately. The gap between cl100k_base and o200k_base is large enough that this single choice matters more than most prompt-engineering effort you'd otherwise spend chasing the same improvement. Build your eval set in the language and script mix your users actually use, including the romanized, code-switched version, not a clean translated version of your English test set. A benchmark score in formal Hindi tells you very little about how the system handles a WhatsApp-style message. Plan for human review as a real cost, not a safety net. Until better automated evaluation exists for Indian languages, native-speaker review is doing work that English-language teams have partly automated away. Budgeting it as an afterthought is how timelines slip. "AI for India" gets talked about as a market-sizing story, a huge population, a fast-growing user base, and it is. What gets left out is that it's also an infrastructure and evaluation problem that's mostly invisible until you've shipped something and watched it behave differently than it did in your English-language testing.
The Hidden Cost of Building AI for Indian Languages
Full Article
Original Source
Read the full article at Hackernoon →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.