A Modder Swapped out an RTX 4080 for Server Hardware to Run LLMs!

TECH NEWS – Although the Nvidia GeForce RTX 4080 is adequate for gaming, one modder encountered issues while performing AI-related tasks.

 

In order to run higher-quality AI models, a GPU must have ample VRAM to ensure a satisfactory experience. An RTX 4080 can easily handle the most visually demanding and graphically intensive games, but running large language models (LLMs) is a different story. A modder who wanted to accomplish both tasks with a single gaming PC successfully installed the Nvidia Tesla V100, though he faced a few challenges along the way. The Tesla V100 cannot be connected to a desktop motherboard as a plug-and-play device, so the modder, known as Tymscar, had to purchase an SXM2-PCIe adapter. The GPU, equipped with 16 GB of HBM2 memory, and the adapter cost him 200 pounds. Despite having 5,120 CUDA cores and a 4,096-bit bus width, which provides 900 GB/s of bandwidth, the GPU still has some computing power to spare.

Obtaining an SXM2–PCIe adapter wasn’t the most difficult task, but it wasn’t easy, especially when it was discovered later that the Tesla V100 lacks a PCIe slot, display inputs, and PCIe power connectors. As shown in the image below, connecting the GPU to the vapor chamber cooler is fine for those not bothered by excessive noise. However, at 82 dB, it will be unpleasant for most people. After making a few minor modifications (using a 9V battery and a PWM jumper to reduce fan noise), the modified V100 cooler now operates at 10% of its original maximum RPM.

After resolving this issue, the modder found a way to integrate 32 GB of usable VRAM into his system. He ran Qwen3.6-27B-MTP, quantized to Q5_K_M, which takes up 19 GB. With a context size of 128K tokens, enough VRAM was available to run the LLM at a rate of 32 tokens per second. Prompt processing ranged from 133 to 160 tokens per second, considered decent performance.

Best of all, you can run your own small- and medium-sized AI models at home for less than $300, with no internet connection required.

Source: WCCFTech, Tymscar

Avatar photo
theGeek is here since 2019.

No comments

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

theGeek Live