4 Sources
[1]
AI enthusiast adds Nvidia Tesla V100 as loud as a lawnmower to gaming PC for $266 -- 32GB of VRAM rig can run 27 billion parameter model at 32 tokens per second
Now they have a total 32GB VRAM system on a budget, for local LLM inference. A computing enthusiast has repurposed a very noisy and largely obsolete enterprise GPU (with lots of VRAM) for local LLM inference purposes. They are now enjoying a system that has doubled its total VRAM quota to 32GB for just a $266 (£200) outlay. That's a good result, especially in the midst of a RAMpocalypse. Oscar Molnar explains that a cheap Tesla V100 SXM2 with 16GB HBM2 was sourced, as was an SXM2-to-PCIe adapter, and a PWM mod for the loud-as-a-lawnmower cooler, to complete this VRAM expansion for the hefty local LLMs project. Indeed, these GPUs do look cheap right now, as I can see them listed on eBay US for under $140 each, if you don't mind buying from China. As mentioned above, you can't just get one of these Tesla V100 SXM2 cards with abundant VRAM and plug it into your PC. Molnar says they spent about $66 on an SXM2-to-PCIe adapter, also on eBay. You might think that was enough. However, the PC and local LLMs enthusiast baulked at the noise of "the fan from hell," which came as standard with the Tesla V100 SXM2. That shrieking cooler was measured outputting 82dB of noise. Molnar described it as "somewhere between a garbage disposal and a lawnmower." This may be the most complicated tweak yet, but basically the existing fan wires just needed rerouting and plugging into the motherboard PWM fan header. You could also simply purchase a "2.54mm male to PH2.0 female jumper cable" for the task. Apparently, the fan only needs to run at 10% to keep the Tesla V100 under 50C at full load. 27 billion parameter LLM runs at 32 tokens per second With the hardware all now fitted and finessed, Molnar had a 32GB VRAM system at their disposal - that's a PC with RTX 4080: 16GB VRAM, Ada architecture and Tesla V100: 16GB VRAM, Volta architecture. They note you can get Tesla V100s with 32GB of VRAM, but they are double the price. Getting the system to make use of this 32GB of total VRAM for LLMs wasn't tricky, says the DIYer. They used NixOS with a legacy Nvidia driver that overlapped support for both Volta and Ada architectures. Testing a local LLM, they got a 27 billion parameter model running at 32 tokens per second, which they say is "fast enough for interactive use" and faster than most cloud API alternatives. Follow Tom's Hardware on Google News, or add us as a preferred source, to get our latest news, analysis, & reviews in your feeds.
[2]
Old Nvidia GPUs with 24GB VRAM are crushing new cards at local AI inference, and here's why
Richard is the PC Hardware Lead at XDA and has been covering the technology industry for almost two decades. He's been building PCs since young, and when not creating content, you can often find him inside a chassis somewhere. You've seen us cover the Nvidia GeForce RTX 3090 and how this now six-year-old GPU refuses to give way to successors for running local AI inference. It's all down to the VRAM and how much is available on the card, making this quite the compelling purchase, even if it happens to be from several generations prior. The RTX 3090 may not be as good for gaming today, but it remains an absolute monster of a local large language model (LLM) powerhouse. It's all in the RTX 3090's VRAM. There's plenty of it to go around One highlight of the RTX 3090 is the 24 GB of GDDR6X VRAM. That sounds like a lot because it is, even by today's standard. The RTX 5080 only comes with 16 GB, so the RTX 3090 is still quite the GPU for cramming as much data into memory. This so happens to be just what running local AI inference demands. Plenty of high-speed memory, and even though it's GDDR6X, it's still vastly speedier than DDR5 memory, making it ideal for local agents. What good is a more capable model if your more powerful GPU doesn't have enough memory to run it? It's what has led to the RTX 3090 becoming one of the most desirable GPUs for a home lab setting. Originally launched as the flagship GPU for gaming with a whopping 10,496 CUDA cores, third-gen Tensor Cores, and 24 GB of GDDR6X VRAM on a 384-bit memory bus, the RTX 3090 is a beast that can clock out at 936 GB/s for memory bandwidth. It was excessive for the time, but these specifications have allowed the card to age like fine wine. Local AI applications are notably more accessible these days, and more people are looking to cut subscriptions with cloud counterparts and run some AI at home. That's where a dedicated discrete GPU comes into play, and you'd struggle to find one better than the RTX 3090 for value. If you were to purchase a brand-new Nvidia GPU, you'd have to go with the obscenely expensive RTX 5090 to get more than 24 GB of VRAM. This difference in memory can matter more than raw compute performance. Because what good is a more capable model if your more powerful GPU doesn't have enough memory to run it? The RTX 5080 alone is better equipped than the RTX 3090 yet has two-thirds of the RAM, resulting in smaller models being the only option to cram all that data into memory. It's faster, sure, but that doesn't matter if slower system RAM has to come in to pick up the slack with the data overflow. VRAM determines what models you can use And running out of memory can lead to terrible performance An AI model is only as good as the underlying hardware. It's why you'll experience different performance with the same model but on two very different systems. Local inference requires memory to store the parameters, but it also needs to fill the RAM up with temporary data, runtime overhead, and cache. So that parameter number doesn't always translate well to how much memory is required to run that model well. Even though a 14 GB model will fit on a 16 GB GPU, it won't run terribly well without optimization. And all that supplementary data is vital for getting the most out of these AI agents. The cache, especially so, which saves the model from having to repeat the same work whenever a new token is generated. If you plan to run some longer context workloads, this is where considerable memory overhead can mean the difference between getting better results sooner or having to start a fresh session because you've hit the limit of what the GPU can physically offer. The RTX 3090 can run models up to around 20 GB or so without issue. The ability to automatically offload some of the model into normal system RAM is great for allowing people to try larger models, but it does result in severe penalties to performance that can sometimes lead to sub-par results with responses. If the model and all its data cannot fit on the same GPU, the system now has to work between the GPU, CPU, and RAM over PCIe. That may not sound like a substantial stepdown, but it's huge compared to the speeds GPU memory can run at. So while newer GPUs may be outright faster than the RTX 3090, they may have to offload some of the work to slower memory, which is where the RTX 3090 can shine by keeping it all locally and offering better sustained performance. The difference between 16 GB and 24 GB is massive Not just in VRAM capacity but also model support Close A smaller, faster 8B model may be perfect for answering simple questions and providing lightweight text generation, but for anything heavier, you'll need substantially more parameters to work with. Coding along needs some hefty models to really handle difficult problems, analysis, and lengthy instructions. This is where 24 GB of VRAM can make a world of difference compared to 16 GB. I've run countless models on my old RTX 4060 Ti with 16 GB of VRAM. It's not the fastest GPU on the block, but the RAM is pretty good for the price. Bump that up to 24 GB with beefier internals, and now you're looking at having the ability to load up 27B parameter models with quantization. And this is where the RTX 3090 can really hold its own against newer cards. The choice of model has a heavier impact on output quality than the hardware used to increase token-generation speed. And just because it's a few generations old, that doesn't mean the RTX 3090 cannot perform once everything is loaded into VRAM. 936 GB/s of memory bandwidth is great compared to modern consumer-grade GPUs. LLM inference has multiple phases, and they all require different parts of the system. Prefill is far more compute-intensive, while decoding will rely more on memory bandwidth as the GPU will need to read the model weights with each processed token. So, it's not that the RTX 3090 offers better performance for local inference; it's why people are flocking to the GPU for AI. It's because it has enough RAM for more breathing space with larger models. It's an unusual addition to the used market Sourcing a used GPU for AI inferencing is more challenging if you don't have the funds to cover the cost of more expensive workstation cards or an RTX 5090. The RTX 3090 is in an interesting position with listings that aren't too out of reach thanks to those offloading their older GPU for an upgrade. It still demands quite the chunk of change to be parted with, but it's a fantastic local AI-crunching GPU compared to other options. NVIDIA RTX 3090 A superb GPU available at reduced prices for running local AI workloads. See at eBay Expand Collapse
[3]
A used Tesla V100 has quietly become the cheapest way to run local LLMs at home
His love of PCs and their components was born out of trying to squeeze every ounce of performance out of the family computer. Tinkering with his own build at age 10 turned into building PCs for friends and family, fostering a passion that would ultimately take shape as a career path. Besides being the first call for tech support for those close to him, Ty is a computer science student, with his focus being cloud computing and networking. He also competed in semi-pro Counter-Strike for 8 years, making him intimately familiar with everything to do with peripherals. In terms of running local LLMs, the RTX 3090 is the firm king of consumer GPUs used to run local AI on a budget with its massive 24 GB of VRAM. Above that, it's either the RTX 5090, or you start to venture into the enterprise realm, which can sound scary and proprietary. While it definitely has its warts, it's a lot friendlier than it sounds. The Nvidia Tesla V100 has slowly been falling out of fashion in data centers for some time, and as a result, loads of them get shoveled onto eBay for discount prices. Both the 16 GB and 32 GB models are worth picking up, as long as you understand that they're slowly being deprecated in ways that aren't circumventable for some tools. 32GB of HBM2 for less than a used RTX 3090 Larger models at home have never been more affordable The PCIe Tesla V100 comes in 16GB and 32GB configurations built on the same GV100 silicon: 5,120 CUDA cores, 640 tensor cores, and 900GB/s of memory bandwidth across a 4096-bit HBM2 interface, all of it on a dual-slot PCIe 3.0 x16 card rated at 250W. The card it's competing with in regard to local AI is a used RTX 3090, which gives you 24GB of GDDR6X at 936GB/s, so on bandwidth the two are effectively tied, with the 3090 slightly ahead. On capacity the V100 is significantly larger at the high end, and on price the comparison currently favors the V100, though the spread is wide enough that it pays to shop carefully. As of late July 2026, 32GB PCIe cards start around $400 to $500 from overseas sellers, while a refurbished unit from a US seller with a return window runs closer to $880, and fan-equipped versions on AliExpress sit around $900. Used RTX 3090s, meanwhile, have been drifting the wrong way for buyers, with eBay averages hovering near the $1,000 mark and price trackers putting the realistic range somewhere between $700 and $1,050 depending on condition and seller. At the low end, then, you are looking at roughly half the price of a 3090 for a third more memory, and at the high end, you are looking at price parity for that same extra capacity. Neither of those is a bad trade if capacity is what is blocking you, and that extra eight gigabytes is a very significant jump. It is the difference between running a 30B-class model at a quantization you actually like with room left for a serious context window, versus trimming the KV cache until the thing fits. The 16GB PCIe card is a different proposition at a different price, with active listings starting near $281 and running well past $2,000 depending on the seller. It makes sense mainly as a dedicated inference card that leaves your gaming GPU alone, not as a capacity play, primarily. It's CUDA, and that counts for a lot It's not an exotic card The reason this card deserves consideration over the cheaper AMD accelerators floating around the same price bracket is that Volta is not an exotic target in the slightest, at least in the context of local AI workloads. It is CUDA compute capability 7.0 and has tensor cores, Volta being the architecture that introduced them. Ollama, llama.cpp, and the rest of the local inference stack are all built against CUDA first and everything else second, which helps your setup be seamless at home. Mostly. Nvidia has already started walking away from Volta The door is closing There's nearly always a catch with enterprise gear, and this is the big one with these V100 cards: CUDA Toolkit 13.0 removed offline compilation and library support for Maxwell, Pascal, and Volta. Applications built with the 12.x series continue to work, but newer toolkits cannot target these architectures at all, and the 13.3 release notes say that support for pre-Turing architectures has been dropped. What this means for you, a potential Tesla V100 buyer, is that tools will potentially cease to function with this card in the near future. Everything that works today will continue to work, but for example, PyTorch moved compute capability 7.0 users onto the CUDA 12.6 wheel and has an open proposal to deprecate Volta outright. LM Studio's CUDA runtime straight up doesn't support it, and prebuilt llama.cpp binaries have dropped Volta, so you'll have to compile yourself if you want to use it. There's also no BF16 support and no FP8, which is a serious omission. Compare all of this to the RTX 3090, which dodges all of these issues besides FP8 support, and the cost discrepancy starts to make a bit more sense. Running one at home isn't straightforward, but it can be done This isn't a gaming card The PCIe V100 is a server part and behaves like one. It uses a passive bidirectional heatsink that requires system airflow to stay inside its thermal limits, which means it will cook itself in an ATX case with two intake fans. There are no display outputs, and it also requires a separate eight pin adapter to power it, but that's usually included by many eBay sellers. The thermal issues can be solved with a 3D-printed shroud and a dedicated fan, which is what almost everyone running these at home has started to do. The V100 is a genuinely good deal, but it comes with some significant asterisks The V100 32GB is a genuinely good deal wearing a genuinely alarming label, but if you can get past that, it is the cheapest way through that 24 GB wall of the RTX 3090. If you want your inference stack to just work, and to keep working when llama.cpp ships its next release, buy the 3090 and do not look back, but otherwise, a V100 could be the silver bullet in your local AI inference setup.
[4]
A Modder's RTX 4080 Was Enough To Play AAA Games, But Not For Running LLMs, So He Integrated NVIDIA's Tesla V100 At A Throwaway Price To Run 27B AI Models
Running higher-quality AI models means your existing GPU will have to be equipped with a ton of VRAM to make the experience enjoyable. For one RTX 4080 owner, playing the most visually taxing and graphically demanding games might be a walk in the park, but for running LLMs, it's a Herculean task. Wanting to accomplish both feats in a single gaming PC, a modder successfully ran NVIDIA's Tesla V100 in his system, but encountered a few challenges along the way. Adding the Tesla V100 grants the modder 32GB of usable VRAM, sufficient for running significantly improved AI models like Qwen 3.6 at 32 tokens/second Before you ask, it's impossible to attach the Tesla V100 to a desktop motherboard like a "plug and play" job, requiring Tymscar to purchase an SXM2-to-PCIe adapter. Sourcing the GPUs with 16GB of HBM2 memory and the accessory cost him £200, which translates into roughly $266. On eBay, you can grab these for only $100 apiece. With the Tesla V100's 5,120 CUDA cores and 4,096-bit bus width that delivers 900GB/s of bandwidth, the GPU still has some computing juice remaining. Successfully sourcing an SXM2-to-PCIe adapter wasn't the most difficult of challenges, but it isn't simple either, especially when you find out later that the Tesla V100 doesn't have a PCIe slot, display inputs, or PCIe power connectors. As you can see in the image below, attaching the GPU to the vapor chamber-style cooler might be perfectly fine for those who don't mind the excessive noise, but at 82dB, it'll make anyone uncomfortable. With a little tweaking here and there to reduce the fan noise by using a 9V battery and a PWM jumper, the V100's modded cooler was now operating at 10 percent of the original maximum RPM. With this problem out of the way, the modder had successfully found a way to get 32GB of usable VRAM into his system. Now, no game will ever require a whopping 32GB of video memory, unless you decide to run a newer title at 16K resolution, so the best use case would be to fire up Qwen3.6 27B. The modder was running Qwen3.6-27B-MTP quantized at Q5_K_M, which comes in at 19GB, and with a context size of 128K tokens, there was sufficient VRAM to run the LLM at 32 tokens per second. Prompt processing was between 133 and 160 tokens per second, making it decent performance if you happen to stumble across previous-generation AI GPUs for home-computing purposes. Best of all, for less than $300, you can have your very own small to medium-sized AI models running at home, free of cost and without any internet connection. News Source: Tymscar Follow Wccftech on Google to get more of our news coverage in your feeds.
Share
Copy Link
Computing enthusiasts are discovering that older Nvidia GPUs like the Tesla V100 and RTX 3090 outperform newer cards for local AI inference, thanks to their abundant VRAM. One modder added a $266 Tesla V100 to their gaming PC, achieving 32GB of total VRAM to run a 27 billion parameter model at 32 tokens per second. The trend highlights how memory capacity matters more than raw compute power for running LLMs at home.
A computing enthusiast has successfully integrated an Nvidia Tesla V100 enterprise GPU into their gaming PC for just $266, doubling their system's total VRAM to 32GB and enabling robust local AI inference capabilities
1
. Oscar Molnar sourced a Tesla V100 SXM2 with 16GB HBM2 memory and paired it with an RTX 4080, creating a system capable of running a 27 billion parameter model at 32 tokens per second. The setup required an SXM2-to-PCIe adapter costing around $66 and a PWM modification to tame the card's notoriously loud cooler, which originally measured 82dB—described as "somewhere between a garbage disposal and a lawnmower"1
. After rerouting the fan to a motherboard PWM header, the cooler now runs at just 10% capacity while keeping the Tesla V100 under 50°C at full load4
.
Source: Wccftech
The Nvidia Tesla V100 has quietly emerged as one of the most cost-effective AI hardware options for home lab AI setups. As data centers phase out these cards, they're flooding the secondhand market at prices starting around $140 on eBay from international sellers
1
. The PCIe Tesla V100 comes in 16GB and 32GB configurations featuring 5,120 CUDA cores, 640 tensor cores, and 900GB/s of memory bandwidth across a 4096-bit HBM2 interface. For those seeking to run local LLMs at home, the 32GB version offers roughly half the price of a used RTX 3090 while providing a third more memory—a significant advantage when VRAM capacity for AI determines which models you can actually run3
.The surge in interest around older Nvidia GPUs stems from a fundamental reality: VRAM capacity matters far more than compute performance when running large language models locally. The RTX 3090, a six-year-old GPU, continues to dominate as the king of consumer cards for local AI inference thanks to its 24GB of GDDR6X VRAM
2
. This capacity advantage allows it to outperform newer cards like the RTX 5080, which despite having superior architecture only packs 16GB of memory. What good is a more capable model if your more powerful Nvidia GPU doesn't have enough memory to run it?2

Source: XDA-Developers
When running local AI inference, memory needs extend beyond just storing model parameters. The system must also accommodate temporary data, runtime overhead, and cache—all critical for optimal performance
2
. Even though a 14GB model technically fits on a 16GB GPU, it won't run well without optimization. The cache proves especially vital for longer context workloads, saving the model from repeating work when generating new tokens per second. When models exceed available VRAM, systems automatically offload data to slower system RAM over PCIe, creating severe performance penalties that can render cloud-based AI inference alternatives faster than your local setup2
.The difference between 16GB and 24GB VRAM isn't just numerical—it's transformational for model support. Smaller 8B models handle simple questions adequately, but coding assistance, analysis, and complex instructions demand substantially larger parameter counts. This is where gaming PC for running LLMs configurations with 24GB or more VRAM shine, enabling users to run 30B-class models at preferred quantization levels with room for serious context windows
3
.Related Stories
While the Tesla V100 offers compelling value for budget-friendly local LLM inference, potential buyers should understand the hardware compatibility limitations ahead. The Volta architecture represents CUDA compute capability 7.0, and Nvidia has begun walking away from this generation
3
. CUDA Toolkit 13.0 removed offline compilation and library support for Maxwell, Pascal, and Volta architectures, meaning applications built with newer toolkits cannot target these cards at all. PyTorch has moved compute capability 7.0 users onto the CUDA 12.6 wheel with an open proposal to deprecate Volta entirely, while LM Studio's CUDA runtime doesn't support it and prebuilt llama.cpp binaries have dropped Volta support3
.
Source: XDA-Developers
The Tesla V100 also lacks BF16 and FP8 support—serious omissions for AI hardware optimization
3
. Compare this to the RTX 3090, which dodges most of these issues besides FP8 support, and the cost discrepancy between cards starts making sense. Tools that work today will continue functioning, but new software may cease supporting these cards in the near future. Despite these limitations, the V100's CUDA compatibility with Ollama, llama.cpp, and the broader local inference stack—all built against CUDA first—makes setup relatively seamless compared to cheaper AMD accelerators3
.For those monitoring the market, 32GB PCIe Tesla V100 cards currently range from $400-$500 from overseas sellers to around $880 for refurbished units from US sellers with return windows, while used RTX 3090s hover near $1,000 with realistic ranges between $700-$1,050
3
. The 16GB PCIe variant starts near $281, positioning it primarily as a dedicated inference card rather than a capacity play3
. As more people seek to cut subscriptions and run AI at home, these older high-VRAM cards represent a pragmatic path forward, provided users accept the eventual software support sunset that accompanies aging enterprise hardware.Summarized by
Navi
[1]
[2]
11 May 2026•Technology

08 Sept 2025•Technology

01 Jun 2026•Technology

1
Technology

2
Technology

3
Policy and Regulation
