2 Sources
[1]
AI enthusiast adds Nvidia Tesla V100 as loud as a lawnmower to gaming PC for $266 -- 32GB of VRAM rig can run 27 billion parameter model at 32 tokens per second
Now they have a total 32GB VRAM system on a budget, for local LLM inference. A computing enthusiast has repurposed a very noisy and largely obsolete enterprise GPU (with lots of VRAM) for local LLM inference purposes. They are now enjoying a system that has doubled its total VRAM quota to 32GB for just a $266 (£200) outlay. That's a good result, especially in the midst of a RAMpocalypse. Oscar Molnar explains that a cheap Tesla V100 SXM2 with 16GB HBM2 was sourced, as was an SXM2-to-PCIe adapter, and a PWM mod for the loud-as-a-lawnmower cooler, to complete this VRAM expansion for the hefty local LLMs project. Indeed, these GPUs do look cheap right now, as I can see them listed on eBay US for under $140 each, if you don't mind buying from China. As mentioned above, you can't just get one of these Tesla V100 SXM2 cards with abundant VRAM and plug it into your PC. Molnar says they spent about $66 on an SXM2-to-PCIe adapter, also on eBay. You might think that was enough. However, the PC and local LLMs enthusiast baulked at the noise of "the fan from hell," which came as standard with the Tesla V100 SXM2. That shrieking cooler was measured outputting 82dB of noise. Molnar described it as "somewhere between a garbage disposal and a lawnmower." This may be the most complicated tweak yet, but basically the existing fan wires just needed rerouting and plugging into the motherboard PWM fan header. You could also simply purchase a "2.54mm male to PH2.0 female jumper cable" for the task. Apparently, the fan only needs to run at 10% to keep the Tesla V100 under 50C at full load. 27 billion parameter LLM runs at 32 tokens per second With the hardware all now fitted and finessed, Molnar had a 32GB VRAM system at their disposal - that's a PC with RTX 4080: 16GB VRAM, Ada architecture and Tesla V100: 16GB VRAM, Volta architecture. They note you can get Tesla V100s with 32GB of VRAM, but they are double the price. Getting the system to make use of this 32GB of total VRAM for LLMs wasn't tricky, says the DIYer. They used NixOS with a legacy Nvidia driver that overlapped support for both Volta and Ada architectures. Testing a local LLM, they got a 27 billion parameter model running at 32 tokens per second, which they say is "fast enough for interactive use" and faster than most cloud API alternatives. Follow Tom's Hardware on Google News, or add us as a preferred source, to get our latest news, analysis, & reviews in your feeds.
[2]
A Modder's RTX 4080 Was Enough To Play AAA Games, But Not For Running LLMs, So He Integrated NVIDIA's Tesla V100 At A Throwaway Price To Run 27B AI Models
Running higher-quality AI models means your existing GPU will have to be equipped with a ton of VRAM to make the experience enjoyable. For one RTX 4080 owner, playing the most visually taxing and graphically demanding games might be a walk in the park, but for running LLMs, it's a Herculean task. Wanting to accomplish both feats in a single gaming PC, a modder successfully ran NVIDIA's Tesla V100 in his system, but encountered a few challenges along the way. Adding the Tesla V100 grants the modder 32GB of usable VRAM, sufficient for running significantly improved AI models like Qwen 3.6 at 32 tokens/second Before you ask, it's impossible to attach the Tesla V100 to a desktop motherboard like a "plug and play" job, requiring Tymscar to purchase an SXM2-to-PCIe adapter. Sourcing the GPUs with 16GB of HBM2 memory and the accessory cost him £200, which translates into roughly $266. On eBay, you can grab these for only $100 apiece. With the Tesla V100's 5,120 CUDA cores and 4,096-bit bus width that delivers 900GB/s of bandwidth, the GPU still has some computing juice remaining. Successfully sourcing an SXM2-to-PCIe adapter wasn't the most difficult of challenges, but it isn't simple either, especially when you find out later that the Tesla V100 doesn't have a PCIe slot, display inputs, or PCIe power connectors. As you can see in the image below, attaching the GPU to the vapor chamber-style cooler might be perfectly fine for those who don't mind the excessive noise, but at 82dB, it'll make anyone uncomfortable. With a little tweaking here and there to reduce the fan noise by using a 9V battery and a PWM jumper, the V100's modded cooler was now operating at 10 percent of the original maximum RPM. With this problem out of the way, the modder had successfully found a way to get 32GB of usable VRAM into his system. Now, no game will ever require a whopping 32GB of video memory, unless you decide to run a newer title at 16K resolution, so the best use case would be to fire up Qwen3.6 27B. The modder was running Qwen3.6-27B-MTP quantized at Q5_K_M, which comes in at 19GB, and with a context size of 128K tokens, there was sufficient VRAM to run the LLM at 32 tokens per second. Prompt processing was between 133 and 160 tokens per second, making it decent performance if you happen to stumble across previous-generation AI GPUs for home-computing purposes. Best of all, for less than $300, you can have your very own small to medium-sized AI models running at home, free of cost and without any internet connection. News Source: Tymscar Follow Wccftech on Google to get more of our news coverage in your feeds.
Share
Copy Link
Oscar Molnar transformed his gaming PC into a local LLM powerhouse by adding a Nvidia Tesla V100 GPU for just $266. Combined with his RTX 4080, the system now boasts 32GB of VRAM and can run a 27 billion parameter model at 32 tokens per second. The project required an SXM2-to-PCIe adapter and creative noise reduction solutions to tame the enterprise GPU's lawnmower-loud cooling system.
Oscar Molnar, an AI enthusiast and modder, has successfully transformed his gaming PC into a powerful local LLM inference machine by integrating a Nvidia Tesla V100 enterprise GPU. The project, completed for approximately $266, demonstrates how budget-conscious AI hobbyists can achieve 32GB of VRAM without breaking the bank during ongoing memory price fluctuations
1
.The Tesla V100 SXM2 variant, sourced with 16GB of HBM2 memory, now pairs with Molnar's existing RTX 4080 to create a dual-GPU configuration capable of running mid-sized AI models locally. This cost-efficient alternative to cloud-based AI inference offers users complete control over their AI workloads without subscription fees or internet dependency
2
.Integrating the Nvidia Tesla V100 into a standard gaming PC presented immediate hardware compatibility obstacles. Unlike consumer graphics cards, the Tesla V100 SXM2 format lacks standard PCIe slots, display outputs, or conventional power connectors. Molnar addressed this by purchasing an SXM2-to-PCIe adapter for roughly $66 on eBay, bringing the total investment to around $266
1
.The Tesla V100 units themselves are available on eBay for under $140 each from international sellers, with some listings as low as $100. These enterprise GPUs feature 5,120 CUDA cores and a 4,096-bit memory bus delivering 900GB/s bandwidth, providing substantial computing power despite their age
2
.
Source: Tom's Hardware
The most significant challenge emerged from the Tesla V100's industrial cooling system, which Molnar described as sounding "somewhere between a garbage disposal and a lawnmower." The stock fan measured a deafening 82dB, making it unsuitable for home computing environments. Through creative AI hardware optimization, Molnar rerouted the fan wiring to connect with a motherboard PWM header, enabling software-controlled fan speeds. Running at just 10 percent capacity, the modified cooling solution keeps the Tesla V100 under 50°C at full load while drastically reducing noise
1
.Alternatively, enthusiasts can purchase a 2.54mm male to PH2.0 female jumper cable to achieve similar noise reduction results without extensive modifications
1
.Related Stories
With the hardware fully operational, Molnar configured NixOS with a legacy Nvidia driver supporting both Volta and Ada architectures, enabling simultaneous use of the Tesla V100 and RTX 4080. The system now successfully runs Qwen 3.6 27B-MTP quantized at Q5_K_M, which requires 19GB of memory. With a context size of 128K tokens, the setup achieves 32 tokens per second during generation—performance Molnar describes as "fast enough for interactive use" and superior to many cloud API alternatives
1
.Prompt processing speeds range between 133 and 160 tokens per second, demonstrating respectable throughput for a budget-friendly local LLM inference configuration
2
. While 32GB Tesla V100 variants exist, they command double the price, making the 16GB version a more practical choice for hobbyists.This project highlights an emerging trend where AI enthusiasts repurpose enterprise hardware to run 27 billion parameter model workloads at home. As local AI inference gains traction, expect more creative solutions leveraging older datacenter GPUs. For those monitoring AI hardware markets, Tesla V100 pricing and availability may shift as demand from the hobbyist community increases.
Summarized by
Navi
[1]
11 May 2026•Technology

08 Sept 2025•Technology

28 Dec 2025•Technology

1
Technology

2
Technology

3
Technology
