Disclosure: As an Amazon Associate I earn from qualifying purchases. This site contains affiliate links.
llama.cpp 0.4.0 Is a Tag, Not a GX10 Bench
ggml-org tagged llama.cpp v0.4.0 on Sep 4, 2026 (published 19:56:47Z). Nightly build on that release is b10809. ggml is bumped to v0.23.0.
Key takeaways
- ggml-org tagged llama.cpp v0.4.0 on Sep 4, 2026 (published 19:56:47Z). Nightly build on that release is b10809. ggml is bumped to v0.23.0.
- Headline items from the notes, not from my bench: initial Qwen3.8-Flash-Next (qwen4exp, #27742, optimization still pending); Nemotron-3-Puzzle (#25444) plus per-layer expert routing (#28323); DeepSeek-V4-Flash-Vision-Exp mtmd (#28133) and input vision fix (#28154); prevent RAM peaking at load (#27483); lazy tensor reading / --lazy-mode (#27794/#27969); sparse FA for DSV4/GLM/Qwen4exp (#27970).
- DFlash2 (#27816) is in the tag; that sit stays parked until I actually sit it. Apple RDMA (#26421) stays off the calendar. This is not a rewrite of the Flash Next, Nemotron, Vision, K2, or NVIDIA-HF notes. Not a GX10 sit.
Local LLMs on NVIDIA Spark / ASUS GX10
ggml-org tagged llama.cpp v0.4.0 today. The instrument is a GitHub release, published Sep 4, 2026 at 19:56:47Z. Nightly on that page: b10809. ggml is v0.23.0.
Headline items from the notes, not from my bench:
- Initial Qwen3.8-Flash-Next (
qwen4exp) support (#27742); optimization still pending - NVIDIA Nemotron-3-Puzzle-75B-A9B (#25444) plus per-layer expert routing (#28323)
- DeepSeek-V4-Flash-Vision-Exp mtmd (#28133) and input vision fix (#28154)
- Prevent RAM peaking at load (#27483)
- Lazy tensor reading /
--lazy-mode(#27794, #27969) - Sparse FA for DeepSeek-V4 / GLM / Qwen4exp (#27970)
That list is the release. It is not leftover unified memory, and it is not tokens per second from this box.
I am not rewriting the Flash Next, Nemotron Puzzle, Vision, K2 Horizon, or NVIDIA Hugging Face notes here. Those pages already carry the size and sit claims. This page is the stack tag that pulled several of those PRs into one version.
Two lines that stay parked even though they show up in the changelog: DFlash2 support (#27816) stays off the calendar until I sit it. Apple RDMA as an RPC transport (#26421) stays off the calendar.
I run Grok when it earns it, and a GX10 when I want the weights in the room. This tag is the software side of that split. If I upgrade and time something, that note comes next. Not this changelog.
Frequently asked questions
Yes. ggml-org tagged v0.4.0 on Sep 4, 2026 at 19:56:47 UTC. The nightly build on that release page is b10809. ggml is v0.23.0.
No. The tag lists initial architecture support and a loader fix. Those are software receipts. I already wrote separate not-a-sit notes for the models. This page is the stack tag only.
No. This note is the GitHub release. If I upgrade and measure something on the GX10, that is a different article.
Affiliate Disclosure: As an Amazon Associate I earn from qualifying purchases. This site contains affiliate links.
