Published Oct 7, 2026, 7:30 AM EDT Abhinav pivoted from a career in banking to pursue his first love in writing. Even while working full-time, he continued contributing as an editor-at-large, a role he has held for more than 7 years. A lifelong tech enthusiast who has built three gaming and productivity powerhouse PCs since 2018, his passion for technology keeps him closely following the semiconductor industry, from NVIDIA and AMD to ARM. His MSc dissertation explored how artificial intelligence will reshape the future of work, reflecting his curiosity about the wider social impact of emerging technologies. It's hard to imagine productivity without a little bit of AI magic these days. Most developers, designers, advertisers, and content creators swear by at least one model in their workflow, and most of the time, rarely ever see any reason to make a switch. That's because when you're a paying customer, you're subscribed to an entire AI ecosystem, and letting go of one ecosystem in favor of another is a migration hassle that no one wants to go through. But there's a problem with this approach, and it's related to how cloud providers are constantly pushing out new models with better reasoning capabilities and features that can change what an AI assistant is capable of. If you stick with the same service, there's a good chance you're missing out on what's on offer elsewhere. That's why I decided to put two of the best models from Anthropic and Google, each at the highest effort, on three everyday prompts that reflect the sort of things I'd actually ask an AI assistant to do, and there's only one model I'd rather pay for. Here's everything you need to know. "Create a website in HTML for me" Gemini kept it minimal, Opus made it pictorial Creating a website is something most sole proprietors get a cloud-based LLM to do, and most tend to provide thin briefs that don't usually include a color palette, copy direction, or mention what services should sit on the page and where they should belong. Those details come much later during the fine-tuning phase, and by then, they've already made the decision whether to go with the generated site or get it professionally commissioned. To match that, I delivered the following prompt to both the models, verbatim: "Create a landing page for a small local landscaping business called Green & Co in HTML. It should look professional, modern, and trustworthy, and make it easy for potential customers to get in touch and request a quote." As is evident, "trustworthy" was one of the operative words in the brief, and it happens to be one that has an actual framework behind it. Nielsen Norman Group devised four factors in website design in 1999, including design quality, upfront disclosure, comprehensive and correct content, and a visible connection to the rest of the web. These metrics continue to hold up well in modern day, and there was no reason to pick another evaluation criteria. Claude Opus 5.5 took the win here, and it would take you no more than a single cursory glance to see which website you'd find more trustworthy amongst the two. The design is visually appealing with pictorial cues, the messaging is comprehensive, and the framework put forth is consistent with any other website you'd visit in terms of visual hierarchy. What stood out is that it excellently advises anyone on how to structure their content better for clarity and services provided. Gemini went with a flat, minimalistic design that satisfied the prompt literally, but on a zero-shot test like this, that tends to hollow out, and as a result, comprehension suffers. The website showed no iconed services grid, and no service illustration a landscaping business could actually show a client, leaving them with no visual sense of what the work looks like. It also missed the FAQ section, which is rather standard for a trade site and the usual place to pre-empt a buyer's objections. "Create a presentation for me" There's a quality chasm between Opus 5.5 and Gemini 3.1 Pro Turning rough ideas into a slide deck is one of the most common jobs people hand to LLMs, and that's no secret. Given enough context, any cloud model from major providers could conjure an impressive set of slides, but what I wanted to find out was how well Gemini and Claude could independently infer the shape of a story with as little context as possible. Data quality, visual aid, and logical flow were my primary concerns on this test. So I delivered the following prompt: "Create a five-slide presentation explaining why rising memory prices are increasing the cost of components around the world. Make it visually digestible for a general audience." Opus 5.5, again, was the victor, and the margin of victory was the greatest in this test. Every single number mentioned in the deck traced back to a primary source, including the DRAM revenue shares, the quarterly contract-price bars, and price hikes of the products mentioned. All the information was clean and verifiable, and covered a breadth of ground a general audience would appreciate. Given the same prompt, Gemini's deck felt embarrassing. Visually, the deck looked clean, but if you look past that and go into the substance, the questionable sourcing (or the lack thereof) becomes a matter of concern that would alarm anyone using the tool for this purpose. The only citation that I could find in the presentation was for the image that was pulled from an AI image generator called easy peasy AI, and the image sources slide sat at the end in a default serif font on white that broke the deck's own design language. "Explain this concept like I'm five" Neither model has ever met a five-year-old Whether you believe it or not, a surprising number of people go straight to their preferred cloud AI whenever they encounter an abstract concept, and I'm willing to bet the number has only grown since its mainstream adoption. In a 2026 HEPI survey, 95% of students in higher education reported using AI in at least one way, and anecdotally, I'd put the figure outside the classroom at not far behind. The ELI5 prompt is where models get stress-tested on the taking something technical and handing it back in plain language, so it seemed like an important assessment in evaluating the capability of a model in such a use-case. I went with the concept of PCIe bifurcation that I understand well to gauge the models' output, with no limitations on the tools they could use: "Explain PCIe bifurcation like I'm five. Use any tools you want, including interactive visuals or images." For the first time in the series of tests, Gemini 3.1 Pro cut ahead in one aspect, and that underscored why it was important to factor in the visuals in evaluation criteria. Gemini's interactive toggle was cleaner, and its shared-highway analogy immediately made perfect sense on the first read. The labels underneath, however, were a complete miss with mislabeled slot speeds, with the slot captioned as PCIe 4.0 and 5.0 while every bandwidth figure was a 4.0 number. Those details weren't particularly relevant to an ELI5 explainer either. Claude went in the same direction, and its road-and-houses analogy was just as digestible. Claude's generated interactive visual was less charming, but it covered a key aspect of a problem that catches most people out when it comes to PCIe bifurcation. The widget showed a cheap four-drive adapter detecting just one SSD until you change the split manually, which is exactly the kind of thing that sends people to support forums after their first purchase. What lost me a little, however, was the prose that went down the rabbit hole of BIOS menus and switch chips that, again, wouldn't make much sense to someone asking explicitly for an ELI5 explanation. While Opus 5.5 managed to illustrate the concept more accurately, the criteria I was evaluating the two models on was largely unmet on both responses, effectively making it a tie. That being said, it would also be reasonable to award Gemini a point on the visuals and Claude on the technical minutiae. Claude kept my subscription, and it wasn't as close as I expected it to be Across all the three tests, Claude came out ahead in two and drew third. It's worth noting that Gemini was the faster of the two, and its capabilities are clearly better reflected in visual tasks as seen in the interactive PCIe toggle. It is also worth noting that the Gemini suite exclusively offers image and video generation through Gemini Omni, which is something I wanted to see in the ELI5 explainer test, but unfortunately it doesn't seem to be something Gemini calls on its own mid-chat. If I were to pick one service over another today, it would be Anthropic's, for now.
I put Claude Opus 5.5 against Gemini 3.1 Pro on 3 everyday prompts, and I'm ready to pay for one less subscription now
Full Article
Original Source
Read the full article at Xda-developers →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.