Disclosure: As an Amazon Associate I earn from qualifying purchases. This site contains affiliate links.
VibeVoice ASR-Streaming Fits 128GB. That Is Not a Sit.
microsoft/VibeVoice-ASR-Streaming-1.5B created 2026-09-02T15:51 UTC. MIT. pipeline_tag automatic-speech-recognition. 3 safetensor shards, sum 5,628,388,290 bytes. Hugging Face usedStorage 5,640,625,810. Ungated. No GGUF.
Key takeaways
- microsoft/VibeVoice-ASR-Streaming-1.5B created 2026-09-02T15:51 UTC. MIT. pipeline_tag automatic-speech-recognition. 3 safetensor shards, sum 5,628,388,290 bytes. Hugging Face usedStorage 5,640,625,810. Ungated. No GGUF.
- 7B sibling microsoft/VibeVoice-ASR-Streaming-7B usedStorage 17,349,013,521. 8 shards, sum 17,348,198,658. Architecture VibeVoiceForASRStreamingTraining. Card: who-said-what as speech arrives. Languages en/zh/es/pt/de/ja/ko/fr/ru/it.
- Serve is GitHub, not llama.cpp: clone microsoft/VibeVoice. Playground aka.ms/vibeasr is cloud. TTS siblings stay TTS. Size-alone both fit 128GB unified. I did not load this.
Local LLMs on NVIDIA Spark / ASUS GX10
Microsoft posted streaming ASR weights this morning. The instrument is a Hugging Face card, created Sep 2 at 15:51 UTC. microsoft/VibeVoice-ASR-Streaming-1.5B. License: MIT. pipeline_tag: automatic-speech-recognition. Ungated. 3 safetensor shards. I summed them: 5,628,388,290 bytes. The API usedStorage is 5,640,625,810. The 7B sibling (microsoft/VibeVoice-ASR-Streaming-7B, created 15:46 UTC) usedStorage is 17,349,013,521. Eight shards sum to 17,348,198,658. No GGUF on either tree.
This is not TTS. microsoft/VibeVoice-1.5B stays the TTS card. Offline diarized ASR is microsoft/VibeVoice-ASR. Streaming is who-said-what as speech arrives. Architecture VibeVoiceForASRStreamingTraining. model_type vibevoice. The card lists English, Chinese, Spanish, Portuguese, German, Japanese, Korean, French, Russian, and Italian. Customized hotwords. I am not quoting their eval figure.
The card points at GitHub for install, not llama.cpp, not Ollama, not Comfy. Clone microsoft/VibeVoice. The playground at aka.ms/vibeasr is cloud. I am not standing FastAPI or vLLM up on the GX10.
5.64 GB and 17.35 GB of official weights fit 128GB unified as a file-size claim. That is not a sit. I have not loaded VibeVoice-ASR-Streaming on the GX10. I am not inventing tokens per second.
I run Grok when it earns it, and a GX10 when I want the weights in the room. This family sits on the local side of that split, as a download. If I actually run the file demo, that note comes next. Not this card.
Frequently asked questions
Yes. Hugging Face cards microsoft/VibeVoice-ASR-Streaming-1.5B and microsoft/VibeVoice-ASR-Streaming-7B were created Sep 2, 2026. License MIT. pipeline_tag is automatic-speech-recognition.
The official weights do, as a file-size claim. 1.5B usedStorage is 5,640,625,810 bytes. 7B is 17,349,013,521. That is not a sit. I have not loaded this on the GX10, and I am not inventing leftover VRAM or tokens per second.
No. TTS is still microsoft/VibeVoice-1.5B. Offline 60-minute who/when/what is microsoft/VibeVoice-ASR. This card is streaming ASR. I am not rewriting those notes.
Affiliate Disclosure: As an Amazon Associate I earn from qualifying purchases. This site contains affiliate links.
