DeepSeek

Illustration for the frontier model on a home PC story
open sourcemodels

A frontier-class DeepSeek model now runs on a home gaming PC, slowly

A post on r/LocalLLaMA from August 3, 2026 documents DeepSeek-V4-Flash-0731 running locally at Q3 quantization on an Intel Windows machine with 24GB of VRAM. Another user ran the full 284B mixture-of-experts checkpoint at 33 tokens per second on two used RTX 3090s plus a secondhand quad-Xeon server. Caveats: a Q3 quant is compressed and degrades knowledge unevenly, and the speeds are low.

Illustration for the llama.cpp multi-token prediction release
modelsopen source

llama.cpp ships multi-token prediction for DeepSeek V4-Flash

llama.cpp release b10228 landed on August 2, 2026 with multi-token prediction for DeepSeek V4-Flash, letting the model draft several tokens per forward pass so local generation speeds up without new hardware. It arrived two days after DeepSeek released V4-Flash-0731, which pushed Terminal Bench 2.1 from 61.8 to 82.7 at unchanged prices.

Illustration for the autonomous AI cyberattack story
research

Unit 42 documents a hacker running autonomous attacks with DeepSeek in an agent framework

Palo Alto Networks' Unit 42 reported on July 31, 2026 that an operator based in Zhuhai embedded DeepSeek in the open-source Hermes Agent framework and, after a single Telegram instruction, let it autonomously find and attack targets. The campaign attempted more than 460 targets across seven vulnerabilities; confirmed impact was data exfiltration from three Citrix NetScaler targets and command execution on eleven Marimo notebook instances.

Illustration for the DeepSeek gigawatt data center story
infrastructure

DeepSeek is building a gigawatt of compute in Inner Mongolia, Bloomberg reports

Bloomberg reported on July 30, 2026 that Hangzhou-based DeepSeek is adding one gigawatt of compute in Ulanqab, Inner Mongolia, about 350 kilometers northwest of Beijing, building its own campus while leasing extra capacity, with part expected online by end of 2027 or start of 2028. The chips are undecided: Nvidia, Huawei, DeepSeek's own future silicon, or a mix. It would be the largest AI facility any Chinese company operates.

Illustration for the DeepSeek V4 Flash story
modelsmoney

DeepSeek V4 Flash update makes its cheap model punch like a flagship

On July 31, 2026, DeepSeek released DeepSeek-V4-Flash-0731, a retrained version of its Flash model with the same architecture and parameter count. Terminal Bench 2.1 jumped from 61.8 to 82.7, above GLM-5.2 (81.0) and DeepSeek's own V4-Pro preview (72.1), approaching Claude Opus 4.8 (85.0), while pricing stays at $0.14 per million input and $0.28 per million output tokens. It landed one day after OpenAI cut GPT-5.6 prices by up to 80%.

Illustration for the Liang Wenfeng story
money

DeepSeek's Liang Wenfeng becomes the world's richest AI founder at $36 billion

On July 14, 2026, Bloomberg's Billionaires Index put DeepSeek founder Liang Wenfeng's fortune at $36 billion, more than doubled from $16.7 billion, making him the wealthiest creator of AI models ahead of Anthropic's Dario Amodei and OpenAI's Greg Brockman. DeepSeek closed its first external funding round in June, raising over $7.4 billion at a valuation above $50 billion, and Liang still controls about 78% of the company, having put roughly $3 billion of his own money into the round.

Illustration for the Chinese AI traffic share story
modelsopen source

Chinese AI models now carry up to 46 percent of US enterprise AI traffic

On July 7, 2026, CNBC reported that the share of tokens used by US companies on Chinese AI models via OpenRouter has stayed above 30 percent every week since February 8, 2026, rising as high as 46 percent. The average across the previous 12 months was just 11 percent, and only 4.5 percent in the first half of 2025. Open-weight Chinese models like DeepSeek and GLM run 60 to 90 percent cheaper than top OpenAI and Anthropic models.