A used Tesla V100 has quietly become the cheapest way to run local LLMs at home

A used Tesla V100 has quietly become the cheapest way to run local LLMs at home

Published Jul 28, 2026, 12:00 PM EDT His love of PCs and their components was born out of trying to squeeze every ounce of performance out of the family computer. Tinkering with his own build at age 10 turned into building PCs for friends and family, fostering a passion that would ultimately take shape as a career path. Besides being the first call for tech support for those close to him, Ty is a computer science student, with his focus being cloud computing and networking. He also competed in semi-pro Counter-Strike for 8 years, making him intimately familiar with everything to do with peripherals. In terms of running local LLMs, the RTX 3090 is the firm king of consumer GPUs used to run local AI on a budget with its massive 24 GB of VRAM. Above that, it's either the RTX 5090, or you start to venture into the enterprise realm, which can sound scary and proprietary. While it definitely has its warts, it's a lot friendlier than it sounds. The Nvidia Tesla V100 has slowly been falling out of fashion in data centers for some time, and as a result, loads of them get shoveled onto eBay for discount prices. Both the 16 GB and 32 GB models are worth picking up, as long as you understand that they're slowly being deprecated in ways that aren't circumventable for some tools. 32GB of HBM2 for less than a used RTX 3090 Larger models at home have never been more affordable The PCIe Tesla V100 comes in 16GB and 32GB configurations built on the same GV100 silicon: 5,120 CUDA cores, 640 tensor cores, and 900GB/s of memory bandwidth across a 4096-bit HBM2 interface, all of it on a dual-slot PCIe 3.0 x16 card rated at 250W. The card it's competing with in regard to local AI is a used RTX 3090, which gives you 24GB of GDDR6X at 936GB/s, so on bandwidth the two are effectively tied, with the 3090 slightly ahead. On capacity the V100 is significantly larger at the high end, and on price the comparison currently favors the V100, though the spread is wide enough that it pays to shop carefully. As of late July 2026, 32GB PCIe cards start around $400 to $500 from overseas sellers, while a refurbished unit from a US seller with a return window runs closer to $880, and fan-equipped versions on AliExpress sit around $900. Used RTX 3090s, meanwhile, have been drifting the wrong way for buyers, with eBay averages hovering near the $1,000 mark and price trackers putting the realistic range somewhere between $700 and $1,050 depending on condition and seller. At the low end, then, you are looking at roughly half the price of a 3090 for a third more memory, and at the high end, you are looking at price parity for that same extra capacity. Neither of those is a bad trade if capacity is what is blocking you, and that extra eight gigabytes is a very significant jump. It is the difference between running a 30B-class model at a quantization you actually like with room left for a serious context window, versus trimming the KV cache until the thing fits. The 16GB PCIe card is a different proposition at a different price, with active listings starting near $281 and running well past $2,000 depending on the seller. It makes sense mainly as a dedicated inference card that leaves your gaming GPU alone, not as a capacity play, primarily. It's CUDA, and that counts for a lot It's not an exotic card The reason this card deserves consideration over the cheaper AMD accelerators floating around the same price bracket is that Volta is not an exotic target in the slightest, at least in the context of local AI workloads. It is CUDA compute capability 7.0 and has tensor cores, Volta being the architecture that introduced them. Ollama, llama.cpp, and the rest of the local inference stack are all built against CUDA first and everything else second, which helps your setup be seamless at home. Mostly. Nvidia has already started walking away from Volta The door is closing There's nearly always a catch with enterprise gear, and this is the big one with these V100 cards: CUDA Toolkit 13.0 removed offline compilation and library support for Maxwell, Pascal, and Volta. Applications built with the 12.x series continue to work, but newer toolkits cannot target these architectures at all, and the 13.3 release notes say that support for pre-Turing architectures has been dropped. What this means for you, a potential Tesla V100 buyer, is that tools will potentially cease to function with this card in the near future. Everything that works today will continue to work, but for example, PyTorch moved compute capability 7.0 users onto the CUDA 12.6 wheel and has an open proposal to deprecate Volta outright. LM Studio's CUDA runtime straight up doesn't support it, and prebuilt llama.cpp binaries have dropped Volta, so you'll have to compile yourself if you want to use it. There's also no BF16 support and no FP8, which is a serious omission. Compare all of this to the RTX 3090, which dodges all of these issues besides FP8 support, and the cost discrepancy starts to make a bit more sense. Running one at home isn't straightforward, but it can be done This isn't a gaming card The PCIe V100 is a server part and behaves like one. It uses a passive bidirectional heatsink that requires system airflow to stay inside its thermal limits, which means it will cook itself in an ATX case with two intake fans. There are no display outputs, and it also requires a separate eight pin adapter to power it, but that's usually included by many eBay sellers. The thermal issues can be solved with a 3D-printed shroud and a dedicated fan, which is what almost everyone running these at home has started to do. The V100 is a genuinely good deal, but it comes with some significant asterisks The V100 32GB is a genuinely good deal wearing a genuinely alarming label, but if you can get past that, it is the cheapest way through that 24 GB wall of the RTX 3090. If you want your inference stack to just work, and to keep working when llama.cpp ships its next release, buy the 3090 and do not look back, but otherwise, a V100 could be the silver bullet in your local AI inference setup.

Original Source

Read the full article at Xda-developers →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.