Stop Calling AI Errors "Hallucinations." They Are Product Failures

Stop Calling AI Errors "Hallucinations." They Are Product Failures

Hallucination’ makes an industrial information failure sound like a charming psychological quirk. If a paid professional invented facts this confidently, we would fire them quickly. My first gripe is this. This is the dictionary definition of a hallucination: A hallucination is a false sensory experience—such as seeing, hearing, smelling, tasting, or feeling something—that feels entirely real but has no external physical source. The most common causes listed are mental health conditions, substances, medical illnesses, or neurological conditions. I know this might get some pushback, but it’s hard to debate, and it’s obviously clear. AI does not hallucinate in the human sense. It does not see, hear, or experience anything. When an AI product presents invented, unsupported, or outdated information as fact, that is a product failure. Dress it up all you like, but it’s just putting lipstick on a pig. Calling it a ‘hallucination’ anthropomorphises the machine, softens accountability, and encourages users to treat false answers as an unavoidable personality trait, which is completely unacceptable, especially from the services you pay for. Google is the most important example, IMO, not because it is the only offender, but because it is attaching generative uncertainty to a search engine that built its reputation on helping us verify reality. If an AI answer cannot be supported, the product should say so, show conventional search results, or simply shut the hell up. Recently, I asked an AI-powered search product a simple question about a current event. It returned a confident answer that was not slightly wrong. It had assembled outdated information, speculation, and what appeared to be a familiar news pattern into a completely false version of reality. Almost as though it was taking whatever the top trending answers were, rather than the most accurate. There were no cautious qualifiers. No admission that the information could not be verified. No small bloke in the corner waving a red flag and yelling, “We may be making this up.” It just delivered fiction in the uniform of research, and the worst part? When challenged, the failure was described using the industry’s favourite escape hatch, “It was a hallucination”. Followed by the obligatory, “You’re completely right to be annoyed and to call me out on that; I provided you with incorrect and outdated information and presented it as fact”. That makes it even worse as far as I am concerned. It wasn’t a hallucination; the machine didn’t see a ghost. The product failed by giving me a false answer presented as fact, and the difference is actually substantial. A human hallucination is not a software bug A human hallucination is a sensory experience that appears real despite there being no corresponding external stimulus. It can involve sight, sound, touch, taste, or even smell. It can be frightening, involuntary, and associated with serious medical conditions. That is the ordinary clinical meaning described by sources such as the Cleveland Clinic. A language model does not have a sensory experience, and it doesn’t become confused about what it saw. It does not wake up distressed because the wallpaper has started negotiating with it, or it’s seen little blue Smurfs running around the house taking up sniper positions with little assault rifles. It processes inputs and generates an output, and that is its one job. When researching, I found this explanation. Depending on the system, it may predict the next token, retrieve information, rank sources, summarise documents, or combine several of these processes. The output can be useful, sophisticated and even astonishingly good. It can also be wrong because the model completed a plausible pattern, the retrieval system found poor material, the ranking system elevated satire, the sources were stale, the query was misread, or the final answer was not properly grounded in the evidence. All of these are technical and product failures, they are not machine psychology. Calling it anything else is a cop-out. We have built these systems up to be more human-like by the language we use with them, and we have also developed a belief that they are always correct because they speak with such authority, confidence, and structure their responses with words that sound very different to the average plonker we talk to all day. Words such as “knows,” “believes,” “thinks,” and “hallucinates” smuggle a mind into a mechanism. That may sound like a philosophical complaint, but in my 25 years in Customer Experience-related industries and roles, it is actually a customer-protection complaint. Once you describe a false output as a hallucination, the error begins to sound like an unfortunate episode the machine suffered and it is clearly designed to immediately invoke a feeling of pity or a sense of forgiveness. The product becomes the patient, and the company becomes the sympathetic carer, yeah right. The customer, apparently, just happened to be standing nearby when the toaster had a difficult afternoon and burnt our crumpets on the number 2 setting. Ladies and gentlemen, accountability has left the building. If a human did this at work, we would not call it a hallucination Imagine I sell you a research service, and you ask for the latest facts about a company, a court case, a medical issue, or a simple sporting result. I fail to check the current information, and I splice together old rumours, likely-sounding names, and fragments from unrelated sources. Then I present the result with absolute confidence and have the audacity to charge you for the privilege. How on God’s green Earth is that a hallucination? Spoiler alert, it’s not, it's gross incompetence, especially considering I had access to the right answers. You would say I failed to do the work, and if I invented the sources as well, you would definitely say I fabricated them. If I kept doing it, you would stop paying me, and depending on the profession and the damage caused, I could be fired, sued, sanctioned, or stripped of the right to continue my profession. You can’t make this up either; we already have a real example. In Mata v. Avianca, lawyers submitted nonexistent cases generated by ChatGPT. The court did not accept ‘hallucination’ as a professional standard. It imposed a US$5,000 penalty. Air Canada provides another useful lesson. Its website chatbot gave a customer incorrect information about bereavement fares. The company was held responsible for the information presented by its own system. The British Columbia tribunal found Air Canada liable rather than treating the bot as an excitable digital employee who had briefly lost contact with reality. That is the correct direction of travel, IMO. A company may automate the answer, but it does not automate away responsibility for the answer. ‘Hallucination’ is excellent product marketing To be clear, I am not arguing that an AI system is lying, because a lie requires intent. The machine does not secretly know the truth, stroke a white cat, and aim to run away with your subscription fees, at least not yet. But ‘hallucination’ is not a neutral alternative. All of these AI systems are better understood as producing language without concern for whether it is true. The word ‘hallucination’ makes a fabricated citation sound like the cost of doing business with a brilliant but eccentric machine. It is the linguistic equivalent of calling an engine fire an unscheduled warmth event. Friendly words reduce pressure on the product owner and designer, because nobody wants to walk into an AGM and say “Fair warning stakeholders: 3.8% of our quarterly report contains fabricated facts and 6.2% of the cited pages go to broken links or are about a different subject, 4.1% turn converted outdated information into a current answer and a heap of them, lets say 2.9% shouldn’t have been answered at all”. It’s much simpler and cheaper to say “The model occasionally hallucinates.” Google is spending the trust it took decades to earn under the guise of an amazing free service This problem is much larger than Google, but Google is the most important case study. For more than two decades, Google trained us to use its name as a verb for checking reality. “Google it” did not mean “ask the internet to improvise.” It meant finding the sources, comparing the results, and working out what was true. Right? Or did I just hallucinate this? Traditional Google Search was never perfect and it contained SEO sludge, spam, stale pages, manipulative headlines, and the occasional forum answer written by a man called PetrolSniffer82. But the basic contract was visible, Google found and ranked pages and the user could inspect them. Simple, and I miss the good old days. Generative search changes that contract, because an AI overview does not merely point toward information. It synthesises information into an answer, and that answer sits above the links and borrows the authority of the Google interface, logo and legacy. The user is no longer only looking at a map of sources, unfortunately for Google, the user is being handed a conclusion. Google knows this trust exists. In its own post following the widely publicised launch failures of AI Overviews, the company wrote that people trust Search for accurate information. It acknowledged that some “odd, inaccurate or unhelpful” overviews had appeared. And the same bloody post argued that AI Overviews generally do not ‘hallucinate’ like other language-model products. Google said errors were more commonly caused by misinterpreting questions, misreading nuances online or not having enough high-quality information. In other words, you can’t explain stuff properly to something that can’t understand what you are saying, who then uses lower-quality information to answer whatever it interpreted your question to be. The defence Google uses is technically interesting and commercially irrelevant and does nothing to restore the quickly eroding confidence in it’s ‘current’ core product. If I ask for a factual answer and your system misreads the query, elevates satire, stitches together unsupported material, and gives me nonsense, I do not care which component should wear the tiny dunce hat; I only care that the service you offered to me, that I have trusted for years, failed. If there is not enough reliable information, the correct answer is not a statistically attractive bedtime story. The correct answer is: I cannot verify this. Here are the best available sources. For a company that still holds an apparent 91% of all search on the planet and 2.5 billion (with a B) monthly users through its AI-powered search feature, I don’t think Google is protecting its search legacy by placing a faster answer above the evidence. It is spending that legacy each time the answer is confidently wrong. The confidence is contagious The damage does not stop with the first false answer. People copy it into emails, presentations, articles, schoolwork, investment memos and group chats. Mario Nawfal and other Merry Men repeat it on X. Another person feeds the claim into a different model, which produces an even cleaner version. By lunchtime, the original fiction has acquired bullet points, a graph and the emotional confidence of a management consultant holding a laser pointer. For those old enough to know, it’s Purple Monkey Dishwasher. The belief that ‘AI is always correct’ is not universal, but the design of these products encourages a dangerous shortcut. The prose is clean, and the answer is faster than Steve McQueen in Bullitt. The interface is calm and slicker than butter, and there are often citations. It looks like the research has already been done, so no need to check it, right? This tendency to accept automated recommendations even when they are wrong is known as automation bias or AI overreliance. We, mere people, can defer to incorrect AI recommendations, even in important decisions. Explanations and citations do not automatically solve the problem, though, and we still need a reason and enough cognitive space to examine them. That becomes harder when the generated answer has removed the friction that once prompted us to compare sources. Manual search is slower, and rather than a bug, that is sometimes a feature. Opening three reputable links lets you see disagreement, dates, who wrote it, context, and the limits of the available evidence. A synthetic answer can compress all of that into one smooth paragraph and accidentally compress uncertainty right out of existence. Many numbers are thrown around on the internet, and some are quoted as saying up to 60% of all cited evidence or posts are incorrect, not worth citing sources, as the numbers do vary wildly, but one thing is for certain: the first failure is machine-generated, and the second is human-amplified. Stop asking users to carry the entire burden The standard disclaimer tells us that AI can make mistakes and we should check important information. OK. Fair enough. Users should check important information. But let us translate the commercial arrangement: The company generates the answer. The company presents it in an authoritative interface. The company receives the subscription, advertising value, and/or strategic benefit. The user is assigned responsibility for discovering whether the answer was fabricated. That is a wonderful business model if you can get it. I order a pizza and pay for it, and if it’s ham and pineapple instead of pepperoni, be damn sure I am asking for my money back. Then again, my pizza doesn’t carry some “The Pizza artist can make mistakes; you should check your order if it’s important to you, but checking it stops you from eating it, not paying for it”. Imagine an Indian restaurant serving a mystery curry with a note saying, “Ingredients may be fictional. Please conduct your own laboratory analysis before swallowing.” At some point the disclaimer stops being caution and becomes an admission that the product is not ready for the job it has been given and that you are paying good Deniros for. Users have responsibilities and so do the vendors. If an AI system is sold as creative brainstorming, invention is part of the deal. If it is sold as search, research, legal assistance, medical guidance, financial analysis, or factual support, invention is a defect, and the product isn’t ready to be paid for yet. IMO, the acceptable error rate depends on the task, and the obligation to describe the error honestly does not. Call the failure what it is The industry does not need one replacement buzzword, it needs a vocabulary specific enough to force action. When a system fails, say how it failed: Fabricated fact: The answer asserted something for which no reliable evidence existed. Fabricated citation: The source, case, paper, quotation, or URL did not exist. Citation mismatch: The source existed but did not support the claim attached to it. Retrieval failure: Relevant and available information was not found or was ranked below poor material. Freshness failure: Old information was presented as current. Source-quality failure: Satire, spam, propaganda or low-quality content was treated as authoritative. Query-interpretation failure: The system answered a different question from the one asked. Unsupported synthesis: Individual sources were real, but the conclusion assembled from them was not. Abstention failure: The system should have said it could not verify the answer, but generated one anyway. This language is less magical and more useful. I had pushed back recently on Google AI and asked why it gave me incorrect information, and it simply stated, “I thought you were in a hurry; therefore, I didn’t research long enough to source the correct information”. WHAT? Engineers can test it, and the product teams can measure it, then the executives can report it, and customers can understand it. Regulators can ask who knew about it and what was done. That’s true accountability. Most importantly, nobody gets to hide behind the idea that the machine had a conniption. What an accountable AI search product should do The solution is not to abandon generative AI, I mean, these systems are too useful for that, and improving too quickly. The solution is to stop granting them a special exemption from ordinary product standards. An accountable factual AI service should: Prefer no answer to an unsupported answer. “I cannot verify that” is a feature, not a defeat. Attach evidence to individual claims. A decorative list of links at the bottom is not enough. Each factual claim should be traceable to material that actually supports it. Separate retrieval from generation. Users should be able to see what the sources say before or alongside the model’s synthesis. Show freshness clearly. Current questions require current sources, not a linguistic smoothie made from last year’s rumours. Publish error rates by task and failure type. One blended “accuracy” score can hide the exact situations in which a product becomes dangerous. Create a visible correction trail. If a widely distributed answer was false, the correction should not vanish into a silent model update. Return conventional search results when synthesis is unreliable. If the AI cannot improve on manual search, get it out of the way. Own the output. The company chose the model, retrieval system, interface, launch date, and marketing claim. Responsibility should follow those choices. Minimum SLA’s. If there are above a certain number of incorrect answers given, you get your next month free. This one I love. This is not anti-AI. It is what taking AI seriously looks like. The machine did not hallucinate. The company shipped the answer. Words shape tolerance. Calling fabricated output a hallucination teaches the public to treat falsehood as an unavoidable side effect of intelligence. It is neither, it is a known failure mode of systems that companies have chosen to put in front of paying customers, patients, students, workers, and billions of search users. And fast becoming a staple in every school in the civilised world. If we are not careful, we could head towards Idiocracy (the movie) Idiocracy: IQ test - YouTubeWe would never accept this language from a human professional. Your accountant did not hallucinate your tax return, and your lawyer did not hallucinate six court cases; your executives at the office didn’t hallucinate your quarterly revenue, well, outside of Bernie Ebbers from Worldcom, and your researchers don’t hallucinate the sources.If they failed to verify it, misrepresented it or made it up and got fired as a result, you cease to pay them. AI companies should not receive a softer standard because their products can produce the apology in beautifully formatted prose. Google and the rest of the industry have built remarkable systems. Now they need to stop describing serious information failures as though the software briefly saw a dragon in the server room. Call the failure what it is, measure it, disclose it, correct it, and if the system doesn’t know, teach it to shut up.

Original Source

Read the full article at Hackernoon →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.