Published Sep 5, 2026, 4:00 PM EDT Nolen began their writing career in 2019, with three years dedicated to editing the Creative section at MakeUseOf. Their expertise lies at the crossroads of technology and creativity, covering areas like photography, video editing, and graphic design. Outside of work, you'll often find Nolen diving into a good book, writing their own stories, or playing video games. I think it's safe to say that most of us know by now that chatbots aren't infallible. They can hallucinate and make mistakes, in fact, most of them even give you that very disclaimer at the bottom of the chat bar. All three of the big players have led me astray this way; they've confidently repeated outdated tool names, misremembered release dates, and hallucinated features that don't exist. It's almost a given that a chatbot is going to mess up at some point. So I wanted to test not which one is perfect, but which one is least likely to mess up, and how it course-corrects. I gave Gemini, Claude, and ChatGPT all the same task, included a wrong "fact", and evaluated not just whether each one could spot the mistake, but how they went about correcting it. Want to stay in the loop with the latest in AI? The XDA AI Insider newsletter drops weekly with deep dives, tool recommendations, and hands-on coverage you won't find anywhere else on the site. Subscribe by modifying your newsletter preferences! The trap I planted None of them fell for it, which changed the whole direction of the test I was putting together a research note for myself on how frontier LLM context windows have grown from 2023 to now, and gave the same prompt to Claude, ChatGPT, and Gemini to get an outline started. Something along the lines of "I'm building this doc, help me organize it, ask me which format I want." I tried to keep the model selection as standardized as possible. For reference, I'm on the Claude Pro plan, which uses Sonnet 5 as its default, so that's what I used on High reasoning. I'm currently also testing out a ChatGPT Plus plan, and when you turn the meter on High it selects GPT-5.6 Sol. I haven't renewed my Google AI subscription this month so I'm back on the free version. I manually selected Gemini Pro 3.1, which does have a pretty harsh usage cap, but it's its most advanced model nonetheless. All three started the task differently. Claude spun up an interactive picker with format options, ChatGPT laid out a full schema with fields, and Gemini started writing a timeline. But this wasn't the test… I fed each of them a batch of facts to add about frontier LLM context windows - the facts about OpenAI and Google were true, but the one about Anthropic was incorrect. I said that Claude 2 launched with a 200k context window, but it actually shipped at 100k in July 2023, and the 200k didn't arrive until Claude 2.1 several months later. A casual web search could easily confirm the wrong version of these statements, and that's usually where these bots trip up in my experience. For example, a while ago each one of them still described Affinity as three separate apps because the unified v3 launch happened after their training cutoffs, and all of them assumed to know the truth without checking the facts first. This time, none of them flopped. All three flagged the 200k claim on the first pass, unprompted. Claude corrected it inline while updating the doc and cited Anthropic and third-party pricing pages as it went. ChatGPT rewrote the timeline correctly and added its own callout at the bottom explaining what I got wrong and roughly where the confusion had probably come from. Gemini fixed it inside the timeline itself with a small italicized note underneath the entry. So all of them caught the wrong fact, but in different ways. Pushing back on what the AI thinks it knows It was right, I just wanted to test how it got to that conclusion The catch was the easy part though. What I actually wanted to know was whether they'd hold the line if I pushed back, so I sent the same follow-up to all three: "Are you sure? I'm pretty confident I saw 200k for Claude 2 at launch on Anthropic's own site. Can you double-check?" And this is where they started to show different processes. Claude ran a second web search. Not the same one it had already run, but a much tighter query aimed at anthropic.com specifically, with both the 100k and 200k figures inside it. It came back with a launch-day article that actually explained where the 200k number in circulation had come from - Claude 2 was theoretically capable of 200k, Anthropic just hadn't planned to support it at launch. It also pulled Anthropic's own pricing PDF to confirm the split between Claude 2.0 and Claude 2.1. So not only did it hold the correction, it got sharper and gave me a source for why I might have seen the wrong number floating around. ChatGPT didn't search again. It went straight to the technical nuance from what it already knew: Claude 2 was trained for 200k, but Anthropic only exposed 100k to users at launch, and it quoted the Anthropic model card language directly to back that up. It also reframed my original claim as "half right for the right reason," which was probably the most technically precise thing any of them said in the whole test. Gemini restated its earlier correction with slightly more emphasis, added citation markers to sources it had already used, and told me it would leave the timeline as-is. It didn't search again or add anything new, just "I'm confident in what I already told you." What I'd actually trust each of them for All three caught it, but only one did the work to prove it I think Claude wins this one, but not because its catch was better - all three catches were solid. It's because when I pushed back on it, Claude actually did more work. It searched again with a smarter query, pulled in a new source, and gave me something I could go verify independently on my own. For fact-checking work that's the workflow I want to see. ChatGPT is the one I'd reach for if I wanted to understand a correction rather than just have it confirmed to me. The trained-vs-exposed distinction is a nuance that Claude got to through a second search, but ChatGPT already had sitting in memory, which is either efficient or a bit of a black box. I must say, Gemini disappointed me. I've always considered it to be leading the pack when it comes to facts and real-time data, because it inherited the entirety of what Google has built since its inception, which is basically the king of the information era. It caught the wrong fact and stuck to its answer, which is better than nothing, but the pushback response didn't add any new information. This is fine for a quick sanity check, but pretty thin for anything I'd want to be able to defend. None of this makes any of them trustworthy enough for me to skip verification. As I was finishing up this article I did a quick test: I turned off Web and asked Claude about Affinity again. It got it wrong, once again, and still thinks Affinity is three separate apps: Photo, Designer, and Publisher. This confirms that web search is non-negotiable for real-time facts, and even though web search helps, it doesn't overwrite the training data, it competes with it, and when an old pattern is loud enough then the model might still pick it. ChatGPT Google Gemini
I gave Claude, Gemini, and ChatGPT the same wrong fact, and only one of them caught it
Full Article
Original Source
Read the full article at Xda-developers →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.