A frontier-class DeepSeek model now runs on a home gaming PC, slowly
A frontier-class AI model is now running on a home gaming PC, and the owner says it is slow as porridge.
The post went up on r/LocalLLaMA on August 3, 2026. Someone got DeepSeek-V4-Flash-0731 running locally at a Q3 quantization on an Intel Windows machine with 24GB of VRAM. Not a cloud endpoint, not an API key. The model, on a desk, at home.
Twenty months of collapse in one quote
Their framing is the part worth quoting: “In less than 20 months we’ve gone from super expensive cloud models only, to being able to run a Q3 quant of DeepSeek on an Intel Windows PC with a very average 24GB of VRAM.”
The same subreddit is full of variations on this right now. Another user documented the full official V4-Flash checkpoint, a 284B mixture of experts, hitting 33 tokens per second on two used RTX 3090s plus a secondhand quad-Xeon server. Secondhand hardware, frontier weights, usable speed.
The caveats that stay attached
Two things temper the excitement. A Q3 quant is a compressed model, not the full one, and quantization degrades knowledge unevenly rather than smoothly, so you cannot assume the local copy behaves like the cloud version on every task. And “it runs” is doing a lot of work when the speed is that low; porridge-speed inference is a demo, not a daily driver.
Why the floor matters anyway
Still, the direction is unmistakable: the floor for running serious models keeps dropping, and it is dropping into equipment people already own. Every step down in hardware requirements moves frontier-adjacent capability from institutions to individuals, from API terms of service to a local process nobody meters.
The gap between what a lab can run and what a gaming PC can run is still large. It is just measurably smaller than it was 20 months ago, and the measurements are being posted from living rooms.
Sources
ANOTHER News is published by ANOTHER, an AI-native content agency. Daily coverage also runs on Instagram.