AI chatbots need to work in all EU languages, not just English

AI chatbots need to work in all EU languages, not just English

Imagine an EU public-service chatbot gives the eligibility rule in French but drops a condition in Romanian. If both are folded into one aggregate score, the system may appear compliant. For the citizen who receives the incomplete answer, the failure is not cosmetic. It is unequal access to a public service. Earlier this month (2 August), the EU entered a new phase of AI governance. The AI Office and national authorities gained new enforcement powers over provisions now applicable, while high-risk rules will apply later, in December 2027 and August 2028. Europe is also a political community of 24 official languages, whose citizens may address EU institutions in any of them and receive a reply in the same language. Brussels has not yet connected them properly. The high-risk regime will require appropriate levels of accuracy, robustness and cybersecurity. Article 15 directs the EU Commission to encourage benchmarks and measurement methods for assessing those qualities. Article 10 adds that, for high-risk systems, the datasets used for training, validation or testing must take account of the geographical, contextual, behavioural and functional setting in which the system is intended to operate. That does not create a general duty to test every AI system in every European language. But it establishes the correct principle: evidence of performance must reflect the actual context of use. Language and language variety are part of that context whenever they can alter whether a system recognises a legal category, follows an instruction, retrieves the right rule or preserves a safeguard. Benchmark now, regret later? The later dates make this more urgent, not less. Standards, procurement templates and testing practices are being designed now. Once a narrow benchmark is embedded in conformity routines, it becomes difficult to dislodge. The issue is not whether Europe values multilingualism in the abstract, but whether authorities receive comparable evidence when performance varies by language. If the evidence used for enforcement is English-first, supervision will inherit the same blind spot. A model that performs well in English may behave differently when asked an equivalent question in Portuguese, Polish or Greek, especially on local law, recruitment, credit, education or public services. “Multilingual support” is therefore not a compliance result. It is a supplier claim. Recent evidence makes the problem difficult to dismiss. Fluent - but not accurate MuBench evaluated models across 61 languages and found notable gaps between claimed and actual coverage, including persistent disparities between English and lower-resource languages. P3B3, a 2026 benchmark of European and Brazilian Portuguese, found that most tested models favoured the Brazilian variety and showed uneven controllability when prompted for European Portuguese. A system can sound fluent while becoming less accurate, less controllable or less reliable. Europe does not need 24 separate regulatory systems, nor must every product be tested in every official language. The standard should be proportional: systems should be evaluated in the languages, varieties and institutional settings in which they are intended, or reasonably foreseeable, to be used. An EU-wide public chatbot, a cross-border banking service or a recruitment system deployed across several member states should not receive one aggregate score that conceals where performance deteriorates. Providers should disclose which languages were tested, on which tasks, with which model version, and where accuracy or safeguards decline. Europe now needs a multilingual evaluation commons. Four solutions First, the EU should fund native, domain-specific test modules rather than rely mainly on translated English questions. Translation can preserve vocabulary while losing legal concepts, administrative practice, idiom and cultural assumptions. National regulators, universities, language specialists and affected public services should help design these tests. Second, results should be reproducible and versioned. A score without the model version, prompt template, test date, number of repetitions and examples of failure is not durable evidence. Material updates should trigger targeted re-testing in languages and settings where the system has consequential effects. Third, public procurement and conformity documentation should require language-specific performance disclosures. A standard language-performance card could identify the tested language and variety, intended use, data source, sample size, known failure modes, uncertainty and model version. It should report factual accuracy, instruction-following, safety and domain knowledge separately from fluency. Finally, incident reporting should record language as a relevant variable. If a system gives a correct answer in one language and an incomplete or unsafe answer in another, authorities need to see the pattern. Otherwise, failures will be filed as isolated errors rather than evidence of unequal performance. Europe’s multilingualism is often treated as a cultural value. In AI governance, it is also an operational test of equal protection. A regulation available in 24 languages but enforced through evidence gathered mainly in English risks creating two classes of citizens: those whose interactions with AI are measured directly, and those whose protection is inferred from somebody else’s language. If an AI system is expected to serve Europeans in their own languages, the evidence for trusting it must speak those languages too.

Original Source

Read the full article at Euobserver →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.