Twitter/X

DeepSeek V4 Flash 0731 can be run locally

Brief

DeepSeek V4 Flash 0731 can be run locally (lossless 4-bit on 168 GB RAM, 3-bit on 110 GB RAM) and is claimed to outperform V4 Pro. GGUF models are available on Hugging Face and the model runs via Unsloth or llama.cpp. DeepSeek’s official V4‑Flash API is now public beta with enhanced Agent features, Responses API and Codex support.

Why it matters

DeepSeek V4 Flash 0731 can be run locally: lossless 4-bit on 168 GB RAM and 3-bit on 110 GB RAM; GGUF artifacts posted on Hugging Face (huggingface.co/unsloth/DeepS…) and runnable via Unsloth or llama.cpp (post dated 2026-07-31).

Key details

  • DeepSeek claims V4-Flash 0731 outperforms V4 Pro; the DeepSeek‑V4‑Flash official API entered public beta with upgraded Agent capabilities, benchmark scores surpassing V4‑Pro‑Preview, native Responses API format, and Codex adaptation.
Source evidence

DeepSeek V4 Flash 0731 can now be run locally! 🐳

Run DeepSeek V4 Flash lossless 4-bit on 168GB RAM and 3-bit on 110GB RAM.

V4 Flash 0731 outperforms V4 Pro. Run via Unsloth or llama.cpp. Smaller quants coming today.

Guide: unsloth.ai/docs/models/deeps…
GGUF: huggingface.co/unsloth/DeepS…

DeepSeek (@deepseek_ai)

🚀 DeepSeek-V4-Flash Official API is now LIVE in public beta!

🔷 We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive performance leap below! 👇
🔷 The official V4-Flash now natively supports the Responses API format and is fully adapted for Codex!

Check out the configuration details in our official API docs: api-docs.deepseek.com/quick_…

— https://nitter.net/deepseek_ai/status/2083084415157022911#m