Disclosure: As an Amazon Associate I earn from qualifying purchases. This site contains affiliate links.
AngelSlim Hy4 GGUFs Are Real Weights. Not a GX10 Sit.
AngelSlim/Hy4-preview-GGUF lastModified 2026-09-01T05:11:08 UTC. Hugging Face usedStorage 932,151,934,379 bytes. Base model is tencent/Hy4-preview (770B MoE / ~49B active, Apache 2.0), already covered.
Key takeaways
- AngelSlim/Hy4-preview-GGUF lastModified 2026-09-01T05:11:08 UTC. Hugging Face usedStorage 932,151,934,379 bytes. Base model is tencent/Hy4-preview (770B MoE / ~49B active, Apache 2.0), already covered.
- Card table: Hy4-preview-STQ1_0.gguf 213.66 GiB, Hy4-preview-UD-IQ1_M.gguf 219.83 GiB, Hy4-preview-Q4_K_M.gguf 435.20 GiB. LFS bytes: 229,412,839,872 / 235,351,974,336 / 467,292,398,016. None fit 128GB.
- hyv4 is not upstream in llama.cpp. The card ships patches in hy4-preview-patch/ against pinned commit 0cea36222. STQ1_0 is MIX-STQ1_0 (~2.38 bpw) and cites llama.cpp PR #22836 for the format. Their 8xH20 benches stay theirs. I did not download or load this. Not a sit.
Local LLMs on NVIDIA Spark / ASUS GX10
On Aug 31 I wrote that Hy4-preview had no public GGUF. That line is stale. AngelSlim posted one. The instrument is a Hugging Face card, last updated Sep 1 at 05:11 UTC. AngelSlim/Hy4-preview-GGUF. I pulled the API: usedStorage is 932,151,934,379 bytes.
The base is the same tencent/Hy4-preview I already covered: 770B MoE, about 49B active, Apache 2.0. I am not re-benching the base. This note is about whether a quant changes the fit answer. It does not.
The card's own size table, three files:
Hy4-preview-STQ1_0.gguf: 213.66 GiB. LFS bytes 229,412,839,872.Hy4-preview-UD-IQ1_M.gguf: 219.83 GiB. LFS bytes 235,351,974,336.Hy4-preview-Q4_K_M.gguf: 435.20 GiB. LFS bytes 467,292,398,016.
The card's own full-residency VRAM line: about 214 GiB for STQ1_0, about 435 GiB for Q4_K_M. Smallest file on the card is 213.66 GiB. 128GB unified does not hold it. The fit test is still the file size, and the file size still says no.
Stock llama.cpp does not run these. hyv4 is not upstream. The card ships its own patches in hy4-preview-patch/ against pinned llama.cpp commit 0cea36222. STQ1_0 is MIX-STQ1_0, roughly 2.38 bits per weight, and the card cites llama.cpp PR #22836 for the format. So the smallest Hy4 file also needs a patched build to open. That is a second gate on top of the 214 GiB.
The card carries 8xH20 numbers from AngelSlim. Those are theirs. I did not time this on a GX10. I did not download the weights. I did not build the patched llama.cpp. I am not inventing tokens per second.
I run Grok when it earns it, and a GX10 when I want the weights in the room. A real GGUF moves Hy4 one step closer to the local side of that split, and the smallest one is still 213.66 GiB against a 128GB box. If a later quant actually sits in 128GB on a stock build and I load it, that note comes next. Not this card.
Frequently asked questions
Yes. AngelSlim/Hy4-preview-GGUF was last updated Sep 1, 2026 at 05:11 UTC. Hugging Face usedStorage is 932,151,934,379 bytes. The Aug 31 note said no public GGUF. That line is now stale, and this is the follow-up.
No. The smallest file on the card is Hy4-preview-STQ1_0.gguf at 213.66 GiB. The card's own full-residency VRAM figure is about 214 GiB for STQ1_0 and about 435 GiB for Q4_K_M. 128GB unified does not hold any of the three. I did not load this on the GX10.
No. The hyv4 architecture is not upstream. The card points at patches in hy4-preview-patch/ applied to llama.cpp commit 0cea36222, and the STQ1_0 file uses the MIX-STQ1_0 format from llama.cpp PR #22836. I did not build it, and I am not quoting their 8xH20 tokens per second as mine.
Affiliate Disclosure: As an Amazon Associate I earn from qualifying purchases. This site contains affiliate links.
