Disclosure: As an Amazon Associate I earn from qualifying purchases. This site contains affiliate links.
Spark-X2.5 Fits 128GB. That Is Not a Bench.
XHToken/Spark-X2.5-4B lastModified 2026-08-31T15:59 UTC. Apache 2.0. ungated. Hugging Face usedStorage 8,234,897,398 bytes. 1.7B usedStorage 3,426,045,486. Collection XHToken/spark-x25.
Key takeaways
- XHToken/Spark-X2.5-4B lastModified 2026-08-31T15:59 UTC. Apache 2.0. ungated. Hugging Face usedStorage 8,234,897,398 bytes. 1.7B usedStorage 3,426,045,486. Collection XHToken/spark-x25.
- Card: hybrid attention (1 full + 3 sliding-window), native 1M claimed. config max_position_embeddings 1048576, sliding_window 512, model_type spark2_5. Serve stacks are not stock: pinned SGLang image, XHToken/Spark-plugin for vLLM, XHToken/llama.cpp fork, XHToken/Spark-MLX-LLM.
- 8.23 GB of weights fit 128GB unified easily. The open question is 1M KV on that box, not disk. I have not loaded this. Vendor benches stay theirs. Not NVIDIA Spark / GX10 branding.
Local LLMs on NVIDIA Spark / ASUS GX10
XHToken posted Spark-X2.5. The instrument is a Hugging Face card, last updated Aug 31 at 15:59 UTC. XHToken/Spark-X2.5-4B. License: Apache 2.0. Ungated. I pulled the API: usedStorage is 8,234,897,398 bytes. The 1.7B sibling (XHToken/Spark-X2.5-1.7B) is 3,426,045,486. Collection: XHToken/spark-x25. GitHub: XHToken/Spark-X2.5.
This is not NVIDIA Spark. The GX10 hardware notes stay the hardware notes.
The card's own architecture sentence: hybrid attention, one full-attention layer with three sliding-window layers, native context up to 1M tokens. From config.json: model_type spark2_5, architecture Spark2_5ForCausalLM, max_position_embeddings 1,048,576, sliding_window 512, 36 layers. The layer list matches that 1-full-plus-3-sliding pattern.
I am not quoting their vendor bench table against Qwen3.5 or Gemma4. I am not inventing tokens per second.
Serve is not stock. The card pins SGLang to lmsysorg/sglang:nightly-dev-cu13-20260827-20621aa1 with --tool-call-parser spark25 and --context-length 1048576. vLLM wants XHToken/Spark-plugin, not stock vLLM. llama.cpp / Ollama / LM Studio point at the XHToken/llama.cpp fork. MLX is XHToken/Spark-MLX-LLM. Official GGUFs exist under the same org; I am not opening a second write for those.
8.23 GB of official BF16 weights fit 128GB unified easily. The untested question is 1M KV on that box, not disk. I have not loaded Spark-X2.5 on the GX10. Spark X2.5-293B is a Sep 7 teaser, not live.
I run Grok when it earns it, and a GX10 when I want the weights in the room. This family sits on the local side of that split, as a download. If I sit 4B and actually measure 1M, that note comes next. Not this card.
Frequently asked questions
Yes. Hugging Face cards XHToken/Spark-X2.5-4B and XHToken/Spark-X2.5-1.7B lastModified Aug 31. License Apache 2.0. Press dated Sep 1, 2026 (iFLYTEK / Ciyuan Xinghuo via XHToken).
The official BF16 weights do, as a file-size claim. 4B usedStorage is 8,234,897,398 bytes. 1.7B is 3,426,045,486. That is not a 1M-context sit. I have not loaded it, and I am not quoting leftover VRAM or tokens per second.
No. This is XHToken Spark-X2.5. The GX10 / NVIDIA Spark hardware notes stay those notes. Do not mix the names.
Affiliate Disclosure: As an Amazon Associate I earn from qualifying purchases. This site contains affiliate links.
