Disclosure: As an Amazon Associate I earn from qualifying purchases. This site contains affiliate links.
DeepSeek-V4-Flash-Vision-Exp Is Real Weights. Not a GX10 Sit.
deepseek-ai/DeepSeek-V4-Flash-Vision-Exp created 2026-08-31T06:16 UTC, lastModified 2026-08-31T12:23 UTC. License MIT. pipeline_tag image-text-to-text. 48 safetensor shards. Hugging Face usedStorage 167,819,616,863 bytes.
Key takeaways
- deepseek-ai/DeepSeek-V4-Flash-Vision-Exp created 2026-08-31T06:16 UTC, lastModified 2026-08-31T12:23 UTC. License MIT. pipeline_tag image-text-to-text. 48 safetensor shards. Hugging Face usedStorage 167,819,616,863 bytes.
- First official multimodal V4-Flash weights. Card: visual modules on the DeepSeek-V4-Flash architecture. Text sibling deepseek-ai/DeepSeek-V4-Flash-0731 (created Jul 31; usedStorage 166,888,735,421). Architecture DeepseekV4ForCausalLM. config: 43 layers, 256 routed experts, top-6, 1 shared, 1M context.
- Unsloth GGUFs same day (unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF). Only UD-Q4_K_XL and UD-Q8_K_XL. I summed the shards: Q4_K_XL 155,095,241,184 bytes (155.10 GB), Q8_K_XL 161,869,615,584 (161.87 GB). No IQ1. Every quant still carries a ~49–50 GB shard. Both over 128GB unified. Not a sit.
Local LLMs on NVIDIA Spark / ASUS GX10
DeepSeek posted Vision-Exp this morning. The instrument is a Hugging Face card, created Aug 31 at 06:16 UTC and last updated 12:23 UTC. deepseek-ai/DeepSeek-V4-Flash-Vision-Exp. License: MIT. pipeline_tag: image-text-to-text. 48 safetensor shards. I pulled the API: usedStorage is 167,819,616,863 bytes.
The card's own sentence: first experimental multimodal model in the DeepSeek-V4 family. Visual modules on the DeepSeek-V4-Flash architecture, then continued training. Text sibling is deepseek-ai/DeepSeek-V4-Flash-0731 (created Jul 31; usedStorage 166,888,735,421). Architecture DeepseekV4ForCausalLM. From config.json: 43 layers, 256 routed experts, 6 experts per token, 1 shared expert, 1,048,576 context. The Vision card does not print a total-parameter count. I am not inventing one.
I am not quoting their eval table. I am not inventing tokens per second.
The fit test is still the file size. Official weights are about 168 GB. That already misses 128GB unified.
Unsloth posted GGUFs the same morning. The card is unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF, last updated Aug 31 at 12:42 UTC. Two quants, no IQ1. I summed the .gguf shards on the tree:
- UD-Q4_K_XL: 155,095,241,184 bytes (155.10 GB)
- UD-Q8_K_XL: 161,869,615,584 (161.87 GB)
Every quant still carries a ~49–50 GB shard. Both over 128GB unified even before KV. A 155 GB file list is not a sit.
I run Grok when it earns it, and a GX10 when I want the weights in the room. This family is still on the far side of that split. If a later quant actually sits in 128GB and I load it, that note comes next. Not this card, and not a download I did not start.
Frequently asked questions
Yes. The Hugging Face card deepseek-ai/DeepSeek-V4-Flash-Vision-Exp was created Aug 31, 2026 at 06:16 UTC and last updated 12:23 UTC. License is MIT. pipeline_tag is image-text-to-text.
No. Official usedStorage is 167,819,616,863 bytes. Unsloth's smallest GGUF, UD-Q4_K_XL, sums to 155,095,241,184 bytes. There is no IQ1 on that tree. File size is not a sit. I did not load this on the GX10.
No. The text sibling is deepseek-ai/DeepSeek-V4-Flash-0731, created Jul 31. This card is the first official multimodal V4-Flash weights. I am not rewriting that text note, and I am not inventing a parameter count the Vision card does not print.
Affiliate Disclosure: As an Amazon Associate I earn from qualifying purchases. This site contains affiliate links.
