I'm not paying $20 for ChatGPT or Claude because a free local LLM does everything I need

I'm not paying $20 for ChatGPT or Claude because a free local LLM does everything I need

Published Oct 6, 2026, 7:30 AM EDT Abhinav pivoted from a career in banking to pursue his first love in writing. Even while working full-time, he continued contributing as an editor-at-large, a role he has held for more than 7 years. A lifelong tech enthusiast who has built three gaming and productivity powerhouse PCs since 2018, his passion for technology keeps him closely following the semiconductor industry, from NVIDIA and AMD to ARM. His MSc dissertation explored how artificial intelligence will reshape the future of work, reflecting his curiosity about the wider social impact of emerging technologies. I've spent some time reflecting on my LLM usage habits, and, inevitably, three names from the biggest cloud AI providers were on my list. ChatGPT and Claude, which together covered a broad chunk of my daily workflow and the menial things that I used to spend a lot of my time on. Each brought something to the table that others don't, which had me convinced for the longest time that there simply wasn't one model capable of replacing all three. There was a price to this convenience, of course. Both the services offered $20 Pro or Plus plans, and having subscriptions to both of them was eating into my spare change. Naturally, I tried looking for a local model that could handle everything I was throwing at the cloud models, and to my surprise, I found one that could handle most, if not all of it. Qwen 3.8-27B changed how I saw local AI It does the job of two subscriptions and costs me nothing Before using Qwen 3.8-27B, I'd assumed it meant settling for something smaller and possibly less intelligent than a cloud model that I was replacing. Perhaps that's why running Qwen 3.8-27B felt like it defied all expectations I had surrounding local AI. The UD-IQ4_XS build from Unsloth is a 13.3GB download that fits on my 4070 Ti Super (with 16GB of GDDR6X VRAM) with context room to spare at 16K context window, and llama.cpp serves it over localhost at speeds that stop being a complaint after a few minutes of use. The model holds about 33.7 tokens per second across runs, and in a Pygame test where I asked it to build a working Snake game in a single file, it did exactly that in one shot. Plain text queries give no room for complaint either, as the model is excellent at the sort of ordinary summarization and analysis work that I usually go to a cloud model for. ChatGPT was the easiest to replace Most of what I use it for doesn't need a model running in a data centre One of the things I use ChatGPT for most often is grammar checking and tidying up my emails, and on the off chance something verbose lands in my inbox, pulling the key details out of it. On rare days, I might even use it to trim down bloated sentences when I get too carried away, or turn a rough thought filled with jargon into something I can communicate more effectively. I also rely heavily on sentiment analysis when I'm researching products and services. I'll often go through dozens of user forums and discussion threads to understand how consumers are reacting to a particular product, service, or pricing change and then use an LLM to codify that information. Those are all tasks that a local model can handle without breaking a sweat, and that's something I learned only after a few days of using the Qwen and Gemma family of models. To put this to test, I fed Qwen 3.8-27B a deliberately rambling email about PC component prices with a dozen figures scattered through it. The model was able to pull every number and categorize it into a clean table, and that was enough for me to be convinced that it's a model that can take down a subscription from my bank statement. Claude is brilliant, until it tells me to go away Five-hour limits are counterproductive for something iterative like coding Anthropic's frontier models are undoubtedly better at coding, and if you don't believe me, you'd only have to go through a dozen tests I've put them through against Google and OpenAI's models. One might even argue that the rise of Claude changed developer workflows permanently, and it'd be a reasonable argument. I use it much the same way, particularly when I'm working on Python projects that require a lot of back-and-forth. The problem is that Anthropic's models come with a lot of limits and guardrails, even for customers like myself who would happily pay $20 a month for some extra usage and performance from its top models. It's no surprise that tweaks to get past Claude's rate limits or to mitigate them as far as possible have their own dedicated tutorials these days. The biggest problem that anyone who relies on Claude faces is that when the five-hour limit runs out, there isn't much you can do besides use your credit card to make it work again, which is a charge on top of a subscription. Waiting for the limit to reset completely destroys the tempo of an iterative coding project, because by the time I'm back (five hours later), I have to go through the thread again and try to pick up where I left off, potentially losing ideas for refining whatever app I was working on. That's a problem that's almost entirely fixed by running a local model. Qwen running locally has no five-hour wall, no weekly cap, and no overage bill that piles up and stares at me at the end of the month. I can leave it iterating on a project overnight if I want, and the only cost will be the 285W the GPU draws from the wall. As an added advantage, nothing I send to Qwen leaves my PC. That goes for both my half-written apps and the benchmarking data, none of which touch a server that another organization controls the retention policy for. llama.cpp Llama.cpp is an open-source framework that runs large language models locally on your computer. Qwen models are free, but they are priceless in the right workflow Qwen 3.8-27B doesn't exactly top the benchmark charts, but for the simple things that I need an LLM for, it absolutely does not have to. Although I found the model that fits the needs of my workflow (and my GPU's VRAM) perfectly, there's a perfect model out there on Hugging Face for every card, from the Gemma family that runs comfortably on my 16GB laptop to smaller Qwen builds that would run on even older hardware. All of them cost nothing besides the hardware they run on, so there's no reason not to tinker and give them a go.

Original Source

Read the full article at Xda-developers →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.