WikiWayne
Local AIGamesAI ToolsTech NewsAboutBlogContact

As an Amazon Associate I earn from qualifying purchases. This site contains affiliate links.

WikiWayne

Independent notes on local AI, real hardware, Grok, and the games that come out of that stack.

Categories

  • Local AI Hub
  • Local AI
  • AI Tools
  • Digital Marketing
  • Tech News

Quick Links

  • About Wayne
  • Contact
  • Links
  • Methodology
  • Editorial Standards
  • Disclosures
  • Privacy Policy
  • Sitemap

Follow on X

Hardware, local models, Grok, and what I actually ship.

Follow @wikiwayne
WikiWayne© 2026
PrivacyMethodologyEditorialDisclosuresTermsSitemap

Disclosure: As an Amazon Associate I earn from qualifying purchases. This site contains affiliate links.

Home/Local AI/Nex-N2.5-mini Fits 128GB. That Is Not a Sit.
Back to Blog
Nex-N2.5-mini Fits 128GB. That Is Not a Sit.
Local AI

Nex-N2.5-mini Fits 128GB. That Is Not a Sit.

Published: September 9, 2026

nex-agi/Nex-N2.5-mini created 2026-09-08T05:08 UTC, lastModified 16:17 UTC. Apache 2.0. Ungated. Hugging Face usedStorage 70,237,676,471 bytes. 16 safetensor shards. No GGUF on the official tree. Card says 35B params; architecture Qwen3_5MoeForConditionalGeneration, model_type qwen3_5_moe.

Key takeaways

  • nex-agi/Nex-N2.5-mini created 2026-09-08T05:08 UTC, lastModified 16:17 UTC. Apache 2.0. Ungated. Hugging Face usedStorage 70,237,676,471 bytes. 16 safetensor shards. No GGUF on the official tree. Card says 35B params; architecture Qwen3_5MoeForConditionalGeneration, model_type qwen3_5_moe.
  • config.json text side: 40 layers, num_experts 256, num_experts_per_tok 8, max_position_embeddings 262144, hybrid layer_types mixing linear_attention and full_attention. vision_config present, so multimodal. Official serve is the patched image nexagi/sglang:v0.5.18-nex-patch with --tp 2 on 2×H100. Not stock llama.cpp, not stock SGLang.
  • 70.24 GB of official BF16 weights fit 128GB unified on paper. I have not loaded Nex-N2.5-mini on the GX10. Vendor benches on the card stay theirs. Pro weights are still empty on HF; Max is real but ~1.65 TB. mradermacher GGUF tree has mmproj files only as of this morning. Not a sit. Not a bench.
2 min read
local-ai, hardware, open-weights
Wayne Lowry, WikiWayne author
Wayne Lowry

Local LLMs on NVIDIA Spark / ASUS GX10

Nex AGI posted Nex-N2.5-mini yesterday. The instrument is a Hugging Face card, created Sep 8 at 05:08 UTC and last touched at 16:17 UTC. nex-agi/Nex-N2.5-mini. License: Apache 2.0. Ungated. I pulled the API: usedStorage is 70,237,676,471 bytes. 16 safetensor shards. No GGUF on the official tree. Collection: nex-agi/nex-n25. GitHub: nex-agi/Nex-N2.5.

The card's model line says 35B params. Architecture Qwen3_5MoeForConditionalGeneration, model_type qwen3_5_moe. From config.json, text side: 40 layers, num_experts 256, num_experts_per_tok 8, max_position_embeddings 262,144. layer_types is a hybrid list mixing linear_attention and full_attention. There is a vision_config, so this is a multimodal card, not a text-only one.

The card publishes vendor bench tables. Terminal-Bench, SWE, OSWorld, the usual columns. I am not quoting them as a sit. I am not inventing tokens per second.

Serve is not stock. The card ships a patched image, nexagi/sglang:v0.5.18-nex-patch, and the mini recipe is a single node with 2×H100 at --tp 2. That is a customized SGLang fork, not stock SGLang, not stock llama.cpp, not Ollama. Official path is patched SGLang on 2×H100, not my box.

Community GGUF, said plainly: mradermacher/Nex-N2.5-mini-GGUF appeared this morning. As of my check it held two files, Nex-N2.5-mini.mmproj-f16.gguf and Nex-N2.5-mini.mmproj-Q8_0.gguf. Those are the vision projector. There is no full runnable GGUF weight set in that tree yet. I am not claiming stock llama.cpp GGUFs are ready.

Family beat, one paragraph. nex-agi/Nex-N2.5-Pro has a card under Apache 2.0, usedStorage about 3,193,112 bytes, and zero safetensor shards. Pro weights are still empty on HF; Max is real but ~1.65 TB — not this note's sit. nex-agi/Nex-N2.5-Max usedStorage is 1,654,577,675,225 bytes, tagged deepseek_v4, text-only, with a 16×H200 multi-node recipe on the card. Neither one is a GX10 sit.

70.24 GB of official BF16 weights fit 128GB unified on paper. The untested question is 256-expert routing plus 262K context on that box, not disk. I have not loaded Nex-N2.5-mini on the GX10. Not a sit. Not a bench.

I run Grok when it earns it, and a GX10 when I want the weights in the room. This card sits on the local side of that split, as a download. If a stock loader lands and I actually sit mini, that note comes next. Not this card.

  • nex-agi/Nex-N2.5-mini
  • nex-agi/Nex-N2.5
  • Nex-N2.5 collection

Frequently asked questions

Yes. Hugging Face card nex-agi/Nex-N2.5-mini was created 2026-09-08T05:08:43Z and last modified 2026-09-08T16:17:10Z. License Apache 2.0. Ungated. It sits in the nex-agi/nex-n25 collection alongside Max and Pro.

The official BF16 weights do, as a file-size claim. Hugging Face usedStorage is 70,237,676,471 bytes, about 70.24 GB. 128GB unified has room for that on paper. I have not loaded it on the GX10, and I am not quoting leftover VRAM or tokens per second.

Not from what I checked. The official tree has no GGUF. The card's serve path is a patched SGLang image, nexagi/sglang:v0.5.18-nex-patch, with a 2×H100 reference deploy for mini. mradermacher/Nex-N2.5-mini-GGUF existed on the morning of Sep 9 but held only mmproj files, not a runnable weight set.

Pro has a card under Apache 2.0 but zero safetensor shards and about 3.19 MB of usedStorage, so the weights are still pending. Max has weights: usedStorage 1,654,577,675,225 bytes, about 1.65 TB, tagged deepseek_v4 and text-only. Neither is a GX10 sit, and neither is this note.

Affiliate Disclosure: As an Amazon Associate I earn from qualifying purchases. This site contains affiliate links.

Related Articles

local ai

Nemotron Puzzle Fits 128GB on Paper. Loader Fix Landed. Not a GX10 Sit.

3 min read

local ai

K2 Horizon Fits on Paper. Not a GX10 Sit.

3 min read

local ai

VibeVoice ASR-Streaming Fits 128GB. That Is Not a Sit.

2 min read