WikiWayne
Local AIGamesAI ToolsTech NewsAboutBlogContact

As an Amazon Associate I earn from qualifying purchases. This site contains affiliate links.

WikiWayne

Independent notes on local AI, real hardware, Grok, and the games that come out of that stack.

Categories

  • Local AI Hub
  • Local AI
  • AI Tools
  • Digital Marketing
  • Tech News

Quick Links

  • About Wayne
  • Contact
  • Links
  • Methodology
  • Editorial Standards
  • Disclosures
  • Privacy Policy
  • Sitemap

Follow on X

Hardware, local models, Grok, and what I actually ship.

Follow @wikiwayne
WikiWayne© 2026
PrivacyMethodologyEditorialDisclosuresTermsSitemap
Home/Local AI

Local AI

Open-weight models, local inference stacks, VRAM planning, and homelab setups for running AI on your own hardware.

View all local AI articles

Cornerstone guides

Fourteen pillar articles at /blog/<slug> — always on the blog URL, never moved under /local-ai/.

P01

Run Open-Weight Models Locally (2026)

P02

LM Studio vs Ollama vs llama.cpp: Which Local AI Tool?

P03

Ollama vs LM Studio (2026)

P04

Best GPU for Local AI (2026)

P05

VRAM Requirements for Local LLMs

P06

Quantization Explained for Local AI

P07

llama.cpp Complete Guide

P08

ComfyUI Local Stable Diffusion Guide

P09

KoboldCpp Local LLM Guide

P10

Open WebUI for Local AI

P11

Local AI Model Tracker (2026)

P12

MLX on Apple Silicon for Local AI

P13

Raspberry Pi Local AI: Limits and Use Cases

P14

Self-Hosting for Beginners (2026)

Latest in local AI

Nex-N2.5-mini Fits 128GB. That Is Not a Sit.
local ai

Nex-N2.5-mini Fits 128GB. That Is Not a Sit.

Nex AGI posted Nex-N2.5-mini under Apache 2.0. Official BF16 usedStorage is 70.24 GB of weights. Sep 9 fold: mradermacher GGUFs are real, Q2_K through Q8_0, 12.94 to 36.90 GB. All fit a GX10 as a file-size claim. Official serve is patched SGLang on 2×H100. I have not loaded it.

4 min read Sep 9, 2026
llama.cpp 0.4.0 Is a Tag, Not a GX10 Bench
local ai

llama.cpp 0.4.0 Is a Tag, Not a GX10 Bench

ggml-org tagged llama.cpp v0.4.0 on Sep 4, 2026. Nightly on that release is b10809. ggml is bumped to v0.23.0. This is a changelog note, not a GX10 sit.

2 min read Sep 4, 2026
Nemotron Puzzle Fits 128GB on Paper. Loader Fix Landed. Not a GX10 Sit.
local ai

Nemotron Puzzle Fits 128GB on Paper. Loader Fix Landed. Not a GX10 Sit.

Official FP8 and NVFP4 fit 128GB on disk. llama.cpp PR 28323 merged; first nightly with the expert-layer fix is b10796. I have not loaded this on the GX10. Not a sit.

3 min read Sep 4, 2026
K2 Horizon Fits on Paper. Not a GX10 Sit.
local ai

K2 Horizon Fits on Paper. Not a GX10 Sit.

IFM posted K2 Horizon under Apache 2.0: six sizes from 0.9B to 375B-A23B. Official under-~110 GB GGUF/FP8 siblings fit on paper. 375B FP8 does not. Not a bench. Not a GX10 load.

3 min read Sep 3, 2026
A small cube computer on a wooden lab bench next to a printed model card with a modest file size circled in red, a microphone
local ai

VibeVoice ASR-Streaming Fits 128GB. That Is Not a Sit.

Microsoft posted streaming ASR weights today: 1.5B usedStorage 5.64 GB, 7B 17.35 GB. MIT. Both fit a GX10 as a file-size claim. This is not TTS, and I have not loaded it.

2 min read Sep 2, 2026
A small black cube computer on a wooden bench inside a red dashed capacity outline, next to a stack of translucent weight blo
local ai

AngelSlim Hy4 GGUFs Are Real Weights. Not a GX10 Sit.

AngelSlim posted Hy4-preview GGUFs. Smallest card size is 213.66 GiB. Still not a GX10 sit. Stock llama.cpp does not run them.

2 min read Sep 2, 2026
A small cube computer on a wooden lab bench next to a printed model card with a modest file size circled in red
local ai

Spark-X2.5 Fits 128GB. That Is Not a Bench.

XHToken posted Spark-X2.5 4B and 1.7B under Apache 2.0. Official 4B BF16 is 8.23 GB of weights. That fits a GX10. Native 1M is claimed. I have not timed it, and this is not NVIDIA Spark.

2 min read Sep 1, 2026
A small cube computer on a wooden lab bench next to a printed model card with a photo clip and a large file size circled in r
local ai

DeepSeek-V4-Flash-Vision-Exp Is Real Weights. Not a GX10 Sit.

Unsloth IQ1_S is 82.44 GB. ggml-org Q2_K_S is 98.59 GB. llama.cpp b10762 has PR 28133. Still not a GX10 sit.

3 min read Aug 31, 2026
A small cube computer on a wooden lab bench next to a printed model card with a large file size circled in red
local ai

Tencent Hy4 Preview Is Real Weights. Not a GX10 Sit.

Tencent posted Hy4-preview on Aug 28: 770B / 49B active, Apache 2.0. Official usedStorage is 1.56 TB. Official FP8 is 814 GB. No public GGUF. Not a GX10 sit.

2 min read Aug 31, 2026
A small cube computer on a wooden lab bench next to a printed model card with a large file size circled in red
local ai

Qwen3.8-Flash-Next Is Real Weights. Not a GX10 Sit.

Qwen posted Flash-Next on Aug 26: 125B / 6B active, plus 51B n-gram and 4B MTP. Official BF16 is 360 GB. Unsloth GGUFs landed; the smallest is 72.5 GB and still not a GX10 sit.

3 min read Aug 26, 2026
A small cube computer on a wooden lab bench next to a printed changelog stamped APPROVE in red
local ai

Open WebUI 0.11.1 Puts a Human Gate in Front of Tool Calls

Open WebUI tagged v0.11.1 on Aug 25. Admins can pause each tool call for allow or deny. Automations and temp chats stay out. Their 1000x stream claim is theirs.

2 min read Aug 25, 2026
A small cube computer on a wooden lab bench next to a printed model card with a file size circled in red
local ai

IBM Granite 4.2 Fits 128GB. That Is Not a Bench.

IBM posted Granite 4.2 on Aug 25: Apache 2.0, dense 3B/8B/30B, thinking switch, 128K native. Official 30B BF16 is 58.55 GB of weights. That fits a GX10. I have not timed it.

2 min read Aug 25, 2026