Published Jul 31, 2026, 6:30 PM EDT Mahnoor Faisal is a tech journalist covering AI and productivity tools with bylines at XDA, SlashGear, MakeUseOf, Laptop Mag, and Android Police. She's been writing professionally since she was sixteen, and has since penned hundreds of articles. This includes in-depth coverage of AI tools like NotebookLM to breaking news across the AI space. Her passion for technology started when she received her first iPod Touch (4th generation) on her 8th birthday, and she's been deep in the tech world ever since. Currently pursuing a degree in computer science, Mahnoor brings both a journalist's eye and a technical foundation to her coverage of how AI is reshaping the way we work and learn. When AI chatbots were still in their honeymoon phase, most of us never gave a second thought to the model actually answering our questions. There were no usage limits to speak of, no real cost to being wasteful, and no spread of models to choose from even if you wanted to. You got one model, it gave you an answer, and there was nothing to weigh it against. We never developed a sense for matching a task to a model, because the situation never asked us to. As time passed and new AI labs came into the picture, that completely changed. Suddenly there wasn't one model but a whole lineup of them. There were fast ones, cheap ones, heavyweight ones, and each one came with its own price and its own daily ceiling. The choice we never had to make became one we made constantly, whether we realized it or not. And I, apparently, was making it wrong every single time...until I decided to switch to Sonnet and Haiku for most of my work. I stuck to Opus for almost everything You know how when you're about to buy a new device, you always instinctively lean toward the highest-end model on the shelf? The one with the most storage, the fastest chip, the specs you'll almost certainly never max out — just in case? That was me with the model picker in AI tools. In terms of Claude, it was me with Opus. Given that I've been on the Max plan since I started using the tool, I had access to Opus, Anthropic's highest-end (excluding Fable) model. So naturally, that's the one I reached for. Every stupid question I had, every summary, every quick edit, every explanation, every reworded sentence, every single thing went to Opus. It didn't matter that most of these tasks wouldn't have known the difference between a heavyweight model and the very first GPT model we had seen. Opus was there, and Opus was the best, so Opus got everything. Why would I reach for anything less? And while there seemed to be no downside to using Opus for every single thing for a fair bit, the answer to that question began showing up more and more as I began to lean on Claude for more and more tasks. I'd be deep into a session and would then receive a polite "you've reached the limit, come back later" message. This began happening more and more, and I reacted the way I suspect most people do. I got annoyed, I complained, I blamed the plan for being stingy. And to be fair, a lot of the blame fell on Anthropic too. The company admitted to tightening usage limits during peak hours as demand surged. So, I wasn't entirely imagining it. However, what I never once did was stop and ask whether I was the problem as well. So I moved most of my work to Sonnet and Haiku Matching the model to the moment Now, more heavyweight models consume your usage allowance far faster than lighter ones. Every Opus message draws down your limit at a steeper rate than the same request sent to Sonnet or Haiku would. Just to give you a better idea, Fable 5, a Mythos-tier model that sits above Opus, is priced at roughly double Opus 5's rate. Essentially, this means that on subscription plans, it draws down your usage allowance at around twice the speed. So, the heavier the model you point at a task, the less of your week you get before the wall comes up. Now think about it: if I point a heavy model at a task that doesn't really require it, it still costs me the same steep drawdown as a task that genuinely needed the horsepower. The model doesn't charge less just because the job was easy. Sure, if the task is marginal and wrapped up quickly, it won't drain a huge chunk on its own. One throwaway question to Opus isn't going to end your week. However, these tasks aren't rare and make up most of what I do each day. Dozens of them, and each one used up the same large share of my Opus allowance as a difficult task would. Added together, they were the reason I ran out of usage so often. For instance, AI models have gotten very good at visual tasks. I have a skill set up within Claude Code that looks at every image in a folder and writes alt text for each one based on rules I've specified. I usually run it on batches of around 50 images, sometimes three times a day. And until recently, I was sending all of that to Opus too when it never really needed it. Haiku, admittedly, didn't do a great job. But Sonnet 5 handles it faster than Opus did, at a fraction of the usage cost, and I genuinely can't tell the difference in the output. To give you a better idea, Opus 5 runs at $5 per million input tokens and $25 per million output. Sonnet 5 sits at $3/$15, which is around 40% less on both. Haiku 4.5 goes even lower, at $1/$5 — a flat fifth of what Opus costs on both input and output. While these are API rates and not what I pay directly on my plan, they're exactly what determines how quickly each model eats into my usage allowance. It makes no sense to spend Opus-level drawdown on a task that Sonnet handles just as well! For months, that's exactly what I'd been doing, over and over, without once stopping to question it. The alt-text workflow is also just one example! I rely on Claude a lot for studying too, like turning lecture notes into summaries, drilling myself with practice questions, and breaking down concepts I didn't catch the first time. None of that needs Opus either. A model summarizing my own notes back to me, or generating a set of recall questions, is doing exactly the kind of well-defined, low-ambiguity work that Sonnet handles without breaking stride. Yet every one of those had been going to Opus too, purely out of the same reflex. I've now moved all of these tasks selectively to Sonnet and, in the lighter cases, Haiku, and the truth is that my day-to-day hasn't changed at all, except that I stopped running into limits! I'm not saying don't use Opus There's a reason it sits at the top of the lineup, and there's a category of work where you'll feel its absence immediately: genuinely hard reasoning, long and tangled problems, the tasks where a lighter model quietly goes shallow, and you don't notice until the output is wrong. When I hit something like that, Opus is still the first thing I reach for, and I don't think twice about it. However, for a lot of tasks, Opus is simply overkill, and you're likely wasting a good chunk of your usage allowance without getting anything back for it. The output wouldn't have been any better, and you just paid more to produce it. The key is matching the model to what the task actually needs, and you'll begin to notice an immediate difference in the value you get out of your subscription.
I moved half my work to Haiku + Sonnet and couldn't tell the difference for most of it
Full Article
Original Source
Read the full article at Xda-developers →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.