WikiWayne
Local AIGamesAI ToolsTech NewsAboutBlogContact

As an Amazon Associate I earn from qualifying purchases. This site contains affiliate links.

WikiWayne

Independent notes on local AI, real hardware, Grok, and the games that come out of that stack.

Categories

  • Local AI Hub
  • Local AI
  • AI Tools
  • Digital Marketing
  • Tech News

Quick Links

  • About Wayne
  • Contact
  • Links
  • Methodology
  • Editorial Standards
  • Disclosures
  • Privacy Policy
  • Sitemap

Follow on X

Hardware, local models, Grok, and what I actually ship.

Follow @wikiwayne
WikiWayne© 2026
PrivacyMethodologyEditorialDisclosuresTermsSitemap
Home/Local AI/Local AI

Local AI

Open-weight models, local inference stacks, VRAM planning, and homelab setups for running AI on your own hardware.

Local AI topical hub
Nex-N2.5-mini Fits 128GB. That Is Not a Sit.
local ai

Nex-N2.5-mini Fits 128GB. That Is Not a Sit.

Nex AGI posted Nex-N2.5-mini under Apache 2.0. Official BF16 usedStorage is 70.24 GB of weights. Sep 9 fold: mradermacher GGUFs are real, Q2_K through Q8_0, 12.94 to 36.90 GB. All fit a GX10 as a file-size claim. Official serve is patched SGLang on 2×H100. I have not loaded it.

4 min read Sep 9, 2026
llama.cpp 0.4.0 Is a Tag, Not a GX10 Bench
local ai

llama.cpp 0.4.0 Is a Tag, Not a GX10 Bench

ggml-org tagged llama.cpp v0.4.0 on Sep 4, 2026. Nightly on that release is b10809. ggml is bumped to v0.23.0. This is a changelog note, not a GX10 sit.

2 min read Sep 4, 2026
Nemotron Puzzle Fits 128GB on Paper. Loader Fix Landed. Not a GX10 Sit.
local ai

Nemotron Puzzle Fits 128GB on Paper. Loader Fix Landed. Not a GX10 Sit.

Official FP8 and NVFP4 fit 128GB on disk. llama.cpp PR 28323 merged; first nightly with the expert-layer fix is b10796. I have not loaded this on the GX10. Not a sit.

3 min read Sep 4, 2026
K2 Horizon Fits on Paper. Not a GX10 Sit.
local ai

K2 Horizon Fits on Paper. Not a GX10 Sit.

IFM posted K2 Horizon under Apache 2.0: six sizes from 0.9B to 375B-A23B. Official under-~110 GB GGUF/FP8 siblings fit on paper. 375B FP8 does not. Not a bench. Not a GX10 load.

3 min read Sep 3, 2026
A small cube computer on a wooden lab bench next to a printed model card with a modest file size circled in red, a microphone
local ai

VibeVoice ASR-Streaming Fits 128GB. That Is Not a Sit.

Microsoft posted streaming ASR weights today: 1.5B usedStorage 5.64 GB, 7B 17.35 GB. MIT. Both fit a GX10 as a file-size claim. This is not TTS, and I have not loaded it.

2 min read Sep 2, 2026
A small black cube computer on a wooden bench inside a red dashed capacity outline, next to a stack of translucent weight blo
local ai

AngelSlim Hy4 GGUFs Are Real Weights. Not a GX10 Sit.

AngelSlim posted Hy4-preview GGUFs. Smallest card size is 213.66 GiB. Still not a GX10 sit. Stock llama.cpp does not run them.

2 min read Sep 2, 2026
A small cube computer on a wooden lab bench next to a printed model card with a modest file size circled in red
local ai

Spark-X2.5 Fits 128GB. That Is Not a Bench.

XHToken posted Spark-X2.5 4B and 1.7B under Apache 2.0. Official 4B BF16 is 8.23 GB of weights. That fits a GX10. Native 1M is claimed. I have not timed it, and this is not NVIDIA Spark.

2 min read Sep 1, 2026
A small cube computer on a wooden lab bench next to a printed model card with a photo clip and a large file size circled in r
local ai

DeepSeek-V4-Flash-Vision-Exp Is Real Weights. Not a GX10 Sit.

Unsloth IQ1_S is 82.44 GB. ggml-org Q2_K_S is 98.59 GB. llama.cpp b10762 has PR 28133. Still not a GX10 sit.

3 min read Aug 31, 2026
A small cube computer on a wooden lab bench next to a printed model card with a large file size circled in red
local ai

Tencent Hy4 Preview Is Real Weights. Not a GX10 Sit.

Tencent posted Hy4-preview on Aug 28: 770B / 49B active, Apache 2.0. Official usedStorage is 1.56 TB. Official FP8 is 814 GB. No public GGUF. Not a GX10 sit.

2 min read Aug 31, 2026
A small cube computer on a wooden lab bench next to a printed model card with a large file size circled in red
local ai

Qwen3.8-Flash-Next Is Real Weights. Not a GX10 Sit.

Qwen posted Flash-Next on Aug 26: 125B / 6B active, plus 51B n-gram and 4B MTP. Official BF16 is 360 GB. Unsloth GGUFs landed; the smallest is 72.5 GB and still not a GX10 sit.

3 min read Aug 26, 2026
A small cube computer on a wooden lab bench next to a printed changelog stamped APPROVE in red
local ai

Open WebUI 0.11.1 Puts a Human Gate in Front of Tool Calls

Open WebUI tagged v0.11.1 on Aug 25. Admins can pause each tool call for allow or deny. Automations and temp chats stay out. Their 1000x stream claim is theirs.

2 min read Aug 25, 2026
A small cube computer on a wooden lab bench next to a printed model card with a file size circled in red
local ai

IBM Granite 4.2 Fits 128GB. That Is Not a Bench.

IBM posted Granite 4.2 on Aug 25: Apache 2.0, dense 3B/8B/30B, thinking switch, 128K native. Official 30B BF16 is 58.55 GB of weights. That fits a GX10. I have not timed it.

2 min read Aug 25, 2026
A small cube computer on a wooden lab bench next to a printed GitHub release circled in red
local ai

llama.cpp 0.3.0 Is a Tag, Not a GX10 Bench

ggml-org tagged llama.cpp v0.3.0 on Aug 25. It adds dots3-note, GLM-4.5-Air MTP, DeepSeek 4 -sm tensor, and a moov-at-EOF video fix. A GitHub tag is not a GX10 number.

2 min read Aug 25, 2026
Timed Qwen 3.8 27B decode receipts on a single ASUS Ascent GX10
local ai

Qwen 3.8 27B on One GX10: The Receipts

I cheered Qwen 3.8 27B before the weights existed. Then I loaded NVFP4 on the ASUS Ascent GX10 and timed it. Threads, comments, and the numbers from that week.

6 min read Aug 16, 2026
A 27 billion parameter open-weight model running on a compact local AI workstation
local ai

Qwen 3.8 27B Locally: We Believe in You

I posted Qwen 3.8 27B three times and meant it. Why a 27B open-weight model still matters on a 128GB Spark-class box, and how I would actually load it.

4 min read Aug 11, 2026
Local AI agent coding a lunar trail game on a compact desktop supercomputer
local ai

I Built Moon Trail with Hermes + DeepSeek V4 Flash on the GX10

Hermes agent, DeepSeek V4 Flash 0731, ASUS Ascent GX10 — local only. How I shipped the Moon Trail browser game without cheating on a cloud model.

4 min read Aug 10, 2026
ASUS Ascent GX10-class 150mm square AI slab on a homelab desk
local ai

The ASUS Ascent GX10 Showed Up. Here's What I'm Actually Running on It.

I unboxed an ASUS Ascent GX10 — NVIDIA DGX Spark in a 150mm cube — and started running local models on 128GB of unified memory. No brochure. First notes from the bench.

5 min read Aug 8, 2026
Best GPU for Local AI (2026) — WikiWayne local-AI hero
local ai

Best GPU for Local AI (2026)

Cornerstone WikiWayne guide: Best GPU for Local AI (2026). Open-weight, practitioner-tested local AI.

8 min read Jun 13, 2026
Best Used GPUs for Local AI on a Budget (2026) — WikiWayne local-AI hero
local ai

Best Used GPUs for Local AI on a Budget (2026)

Shopping the secondary market without overspending on VRAM you cannot use.

9 min read Jun 13, 2026
Your First ComfyUI Workflow for Local SDXL — WikiWayne local-AI hero
local ai

Your First ComfyUI Workflow for Local SDXL

Load checkpoints, wire KSampler, export PNGs locally.

8 min read Jun 13, 2026
ComfyUI Local Stable Diffusion Guide — WikiWayne local-AI hero
local ai

ComfyUI Local Stable Diffusion Guide

Cornerstone WikiWayne guide: ComfyUI Local Stable Diffusion Guide. Open-weight, practitioner-tested local AI.

9 min read Jun 13, 2026
CPU-Only Local LLM Privacy Tradeoffs — WikiWayne local-AI hero
local ai

CPU-Only Local LLM Privacy Tradeoffs

Slower tokens, stronger air-gap story.

8 min read Jun 13, 2026
GPU Offload Layers Explained for Local LLMs — WikiWayne local-AI hero
local ai

GPU Offload Layers Explained for Local LLMs

What `-ngl` / GPU layer sliders actually do.

8 min read Jun 13, 2026
Homelab Docker Stack: Ollama + Open WebUI — WikiWayne local-AI hero
local ai

Homelab Docker Stack: Ollama + Open WebUI

Compose services for local chat without cloud relay.

8 min read Jun 13, 2026
How Much VRAM for Llama 3 8B? — WikiWayne local-AI hero
local ai

How Much VRAM for Llama 3 8B?

Quant-specific VRAM bands for Meta Llama 3 8B class models.

8 min read Jun 13, 2026
Install Ollama on Windows, Mac, and Linux (2026) — WikiWayne local-AI hero
local ai

Install Ollama on Windows, Mac, and Linux (2026)

Step-by-step Ollama install paths for the three major desktop OS families.

8 min read Jun 13, 2026
KoboldCpp Creative Writing Setup — WikiWayne local-AI hero
local ai

KoboldCpp Creative Writing Setup

Import GGUF and tune narrative sampling locally.

8 min read Jun 13, 2026
KoboldCpp Local LLM Guide — WikiWayne local-AI hero
local ai

KoboldCpp Local LLM Guide

Cornerstone WikiWayne guide: KoboldCpp Local LLM Guide. Open-weight, practitioner-tested local AI.

8 min read Jun 13, 2026
llama.cpp CUDA Build Quickstart on Linux — WikiWayne local-AI hero
local ai

llama.cpp CUDA Build Quickstart on Linux

Compile with GPU backends for NVIDIA cards.

8 min read Jun 13, 2026
llama.cpp Complete Guide — WikiWayne local-AI hero
local ai

llama.cpp Complete Guide

Cornerstone WikiWayne guide: llama.cpp Complete Guide. Open-weight, practitioner-tested local AI.

8 min read Jun 13, 2026
llama.cpp vs Ollama: When to Switch — WikiWayne local-AI hero
local ai

llama.cpp vs Ollama: When to Switch

Leave the managed service when you need custom builds or flags.

7 min read Jun 13, 2026
LM Studio: Download Models Step by Step — WikiWayne local-AI hero
local ai

LM Studio: Download Models Step by Step

Use the LM Studio catalog without guessing quant labels.

8 min read Jun 13, 2026
LM Studio vs Ollama vs llama.cpp: Which Local AI Tool? — WikiWayne local-AI hero
local ai

LM Studio vs Ollama vs llama.cpp: Which Local AI Tool?

Cornerstone WikiWayne guide: LM Studio vs Ollama vs llama.cpp: Which Local AI Tool?. Open-weight, practitioner-tested local AI.

8 min read Jun 13, 2026
Local AI Model Tracker (2026) — WikiWayne local-AI hero
local ai

Local AI Model Tracker (2026)

Cornerstone WikiWayne guide: Local AI Model Tracker (2026). Open-weight, practitioner-tested local AI.

8 min read Jun 13, 2026
Local LLM Checklist: Keep Data Off the Cloud — WikiWayne local-AI hero
local ai

Local LLM Checklist: Keep Data Off the Cloud

Network egress, logging, and backup habits for homelabs.

9 min read Jun 13, 2026
MLX on Apple Silicon for Local AI — WikiWayne local-AI hero
local ai

MLX on Apple Silicon for Local AI

Cornerstone WikiWayne guide: MLX on Apple Silicon for Local AI. Open-weight, practitioner-tested local AI.

8 min read Jun 13, 2026
Install MLX on Apple Silicon for Local Llama — WikiWayne local-AI hero
local ai

Install MLX on Apple Silicon for Local Llama

Python venv, mlx-lm, and a tiny model smoke test.

8 min read Jun 13, 2026
NVIDIA vs AMD GPU for Local LLMs (2026) — WikiWayne local-AI hero
local ai

NVIDIA vs AMD GPU for Local LLMs (2026)

CUDA maturity vs ROCm tradeoffs for GGUF stacks.

7 min read Jun 13, 2026
Ollama OpenAI-Compatible API for Local Apps — WikiWayne local-AI hero
local ai

Ollama OpenAI-Compatible API for Local Apps

Point agents and UIs at `http://localhost:11434/v1`.

7 min read Jun 13, 2026
Open WebUI for Local AI — WikiWayne local-AI hero
local ai

Open WebUI for Local AI

Cornerstone WikiWayne guide: Open WebUI for Local AI. Open-weight, practitioner-tested local AI.

8 min read Jun 13, 2026
Open WebUI + Ollama Connection Guide — WikiWayne local-AI hero
local ai

Open WebUI + Ollama Connection Guide

Docker and bare-metal pairing for household chat.

7 min read Jun 13, 2026
Pull Your First Open-Weight Model in Five Minutes — WikiWayne local-AI hero
local ai

Pull Your First Open-Weight Model in Five Minutes

From zero to a working chat with a small Ollama tag.

8 min read Jun 13, 2026
Q4 vs Q8 Quant Quality Tradeoffs — WikiWayne local-AI hero
local ai

Q4 vs Q8 Quant Quality Tradeoffs

When to spend extra gigabytes on higher precision.

8 min read Jun 13, 2026
Quantization Explained for Local AI — WikiWayne local-AI hero
local ai

Quantization Explained for Local AI

Cornerstone WikiWayne guide: Quantization Explained for Local AI. Open-weight, practitioner-tested local AI.

9 min read Jun 13, 2026
Raspberry Pi 5 and Small LLM Limits — WikiWayne local-AI hero
local ai

Raspberry Pi 5 and Small LLM Limits

What runs at usable speed on 8 GB Pi hardware.

8 min read Jun 13, 2026
Raspberry Pi Local AI: Limits and Use Cases — WikiWayne local-AI hero
local ai

Raspberry Pi Local AI: Limits and Use Cases

Cornerstone WikiWayne guide: Raspberry Pi Local AI: Limits and Use Cases. Open-weight, practitioner-tested local AI.

8 min read Jun 13, 2026
Run Open-Weight Models Locally (2026) — WikiWayne local-AI hero
local ai

Run Open-Weight Models Locally (2026)

Cornerstone WikiWayne guide: Run Open-Weight Models Locally (2026). Open-weight, practitioner-tested local AI.

8 min read Jun 13, 2026
VRAM Requirements for Local LLMs — WikiWayne local-AI hero
local ai

VRAM Requirements for Local LLMs

Cornerstone WikiWayne guide: VRAM Requirements for Local LLMs. Open-weight, practitioner-tested local AI.

8 min read Jun 13, 2026
What Is GGUF? The Local LLM File Format Explained — WikiWayne local-AI hero
local ai

What Is GGUF? The Local LLM File Format Explained

GGUF packs tensors and metadata for llama.cpp-compatible runners.

8 min read Jun 13, 2026
Ollama vs LM Studio (2026): Which Local AI Runner Fits Your Workflow? — WikiWayne local-AI hero
local ai

Ollama vs LM Studio (2026): Which Local AI Runner Fits Your Workflow?

Ollama favors CLI and API automation; LM Studio favors GUI model browsing. Compare setup, VRAM use, and GGUF workflows on real hardware.

8 min read Jun 13, 2026
A Raspberry Pi 5 with cables connected sitting on a desk next to a terminal window showing Docker containers running
local ai

Self-Hosting for Beginners: Run Your Own Services in 2026

A complete guide to self-hosting your own services in 2026 — from hardware and Docker to Nextcloud, Vaultwarden, and OpenClaw on a Raspberry Pi or VPS.

13 min read Feb 25, 2026
All articles