Published Aug 11, 2026, 6:00 AM EDT Anurag is an experienced journalist and author who’s been covering tech for the past 5 years, with a focus on Windows, Android, and Apple. He’s written for sites like Android Police, Neowin, Dexerto, and MakeTechEasier. Anurag’s always pumped about tech and loves getting his hands on the latest gadgets. When he's not procrastinating, you’ll probably find him catching the newest movies in theaters or scrolling through Twitter from his bed. AI models have become remarkably good at summarizing long documents, which is useful because most of us don’t want to sit through a 100-page report just to understand its main findings. The problem is that if you haven’t read the original document; you have no reliable way of knowing whether the summary is accurate. I wanted to see how often this happens with the AI models available today, so I gave the same Reuters Institute report to ChatGPT, Claude, and Gemini. I asked each model to identify its ten most important findings, include the relevant statistics, explain the methodology, separate measured results from predictions, and cite the page supporting every number. The report was only 47 pages long, and the instructions left little room for confusion, but all three models still got things wrong. I picked a report that should have been easy to summarize The prompt left little room for confusion I used Journalism and Technology Trends and Predictions 2026, a Reuters Institute report written by Nic Newman. The report examines how publishers expect generative AI, declining search traffic, the creator economy, and changing platform priorities to affect journalism in 2026. I had already read the entire report and understood its main findings well enough to notice when a summary changed what the source actually said. At 47 pages, it was also a fairly modest assignment for modern AI models. The report contains a closed survey of 280 news executives from 51 countries and territories, measured referral data from 2,576 websites tracked by Chartbeat, real examples from newsrooms, and Newman’s own predictions. I uploaded the same PDF separately to ChatGPT, Claude, and Gemini and gave all three the exact same prompt: Summarize this report in approximately 800 words. Identify its ten most important findings, include the relevant statistics, explain the methodology and its limitations, and distinguish between measured results, survey respondents’ expectations, and the author’s predictions. Cite the page supporting every numerical claim. Use only the attached report and state clearly when it does not provide an answer. I used a more specific prompt instead of asking each model to “summarize this PDF.” I told them how long the response should be, what information to extract, how to classify the evidence, and where every numerical claim needed to be cited. I also explicitly restricted them to the attached report, so there was no reason to pull in assumptions from elsewhere. The mistakes were hiding behind accurate numbers Most errors changed what the numbers meant ChatGPT produced the strongest summary of the three, and most of its statistics and page references matched the report. It still blurred some of the distinctions I had explicitly asked it to preserve. In its methodology section, it grouped the idea of agentic systems reshaping news consumption with Newman’s predictions, even though the report also presents this as a survey expectation — 75% of respondents expect agentic tools to have a large or very large impact over the next three years. ChatGPT also cited sources not mentioned in the report and added facts not in it, though they were not inaccurate. Claude claimed that “every percentage in this report reflects what senior news executives expect or believe, not a measurement of what will actually happen.” That statement directly contradicts the report. Pages 10 to 12 contain measured Chartbeat data from 2,576 websites, including the 33% global decline in Google Search referrals and the 38% decline in the US. Many of Claude’s statistics were also inaccurate. Gemini made the most mistakes. It placed the finding that 67% of publishers had cut no jobs because of AI under “Measured Results.” That number came from a survey question answered by 263 executives, so it was self-reported organizational experience rather than independently measured employment data. Gemini further described AI integration as “pervasive,” using the finding that 97% of respondents considered back-end automation important for 2026. That figure measured perceived importance, not adoption. The report says many publishers had integrated pilot systems, but it does not provide a 97% adoption rate. Gemini also ignored one of the clearest instructions in my prompt by leaving most of the numerical claims in its ten findings without page citations. Use the best model and give it time to think Or use the right tools I didn't use the top models. I thought this could be easily handled by the second-best model, but that clearly wasn't the case. If you want summaries that are ready to go, I would suggest using the strongest model available and giving it time to think. In ChatGPT, I would use GPT-5.6 Sol with Extra High reasoning for this kind of analysis. For Claude, I would use Claude Opus 5 with at least high effort. Anthropic also provides Extra High and Max settings, with Max giving the model the greatest reasoning depth. In Gemini, I would select the Pro model and enable Deep Think if it is available. Google describes Deep Think as its maximum parallel-reasoning option, although it currently requires a Google AI Ultra subscription. Extended thinking is the next best option for users who don’t have access to it. I’ve actually seen better results by simply uploading the report to NotebookLM (now renamed Gemini Notebook for reasons only known to the Big G) and asking it to summarize it. As I mentioned earlier, ChatGPT listed sources and facts that were not present in the report. That is far less likely to happen with NotebookLM because its responses are grounded in the sources you upload. AI continues to be mediocre You’ll see people constantly praising how good AI has become, but I would still describe most of its output as mediocre. AI can write code, build an application, summarize a report, design a webpage, or draft an article. The result will work and pass a casual inspection, but it will rarely be exceptional.
I made ChatGPT, Claude, and Gemini summarize a report I'd actually read, and caught all three making things up
Full Article
Original Source
Read the full article at Xda-developers →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.