llama.cpp 0.4.0 Is a Tag, Not a GX10 Bench
ggml-org tagged llama.cpp v0.4.0 on Sep 4, 2026. Nightly on that release is b10809. ggml is bumped to v0.23.0. This is a changelog note, not a GX10 sit.

> I run open-weight models on boxes I own — NVIDIA Spark, ASUS GX10, Grok, local agents — then write what worked. No brochure. Games included.
Local AI leads: open-weight models on hardware I actually own, plus Grok, vibe coding, and the games that come out of that stack. Marketing archives stay in the hubs — they are not the front door.
Open-weight models, local inference stacks, VRAM planning, and homelab setups for running AI on your own hardware.
Open hub From the labSmall browser games shipped from the same stack — including Moon Trail, built with a local agent.
Open hubIn-depth reviews, comparisons, and guides for the latest AI tools reshaping how we work and create.
Open hubSEO strategies, content marketing insights, and growth tactics for the AI-powered marketing landscape.
Open hubBreaking developments in AI, cloud computing, data centers, and the future of technology.
Open hubNex AGI posted Nex-N2.5-mini under Apache 2.0. Official BF16 usedStorage is 70.24 GB of weights. Sep 9 fold: mradermacher GGUFs are real, Q2_K through Q8_0, 12.94 to 36.90 GB. All fit a GX10 as a file-size claim. Official serve is patched SGLang on 2×H100. I have not loaded it.
Read full articleggml-org tagged llama.cpp v0.4.0 on Sep 4, 2026. Nightly on that release is b10809. ggml is bumped to v0.23.0. This is a changelog note, not a GX10 sit.
NVIDIA's Hugging Face headline says $12.9303 billion. The 8-K describes approximately $11.9 billion for stockholders plus up to approximately $1.0 billion for employee retention—not $11.9 billion cash.
Official FP8 and NVFP4 fit 128GB on disk. llama.cpp PR 28323 merged; first nightly with the expert-layer fix is b10796. I have not loaded this on the GX10. Not a sit.
IFM posted K2 Horizon under Apache 2.0: six sizes from 0.9B to 375B-A23B. Official under-~110 GB GGUF/FP8 siblings fit on paper. 375B FP8 does not. Not a bench. Not a GX10 load.
Local AI, vibe coding, real hardware
Open-weight models on hardware I own. Grok and local agents when they earn it. Skeptical of the brochure.
Read Full BioHardware unboxings, local-model runs, Grok experiments, and the occasional cookie bet.
Follow @wikiwayne