Instructions to use mlx-community/NemotronLabs-VoiceChat-11B-mlx-bf16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-community/NemotronLabs-VoiceChat-11B-mlx-bf16 with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir NemotronLabs-VoiceChat-11B-mlx-bf16 mlx-community/NemotronLabs-VoiceChat-11B-mlx-bf16
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
NemotronLabs-VoiceChat-11B — MLX (bf16)
MLX conversion of nvidia/NVIDIA-NemotronLabs-VoiceChat-11B,
a duplex speech-to-speech model that listens and speaks at the same time. 11.1 B parameters,
converted from the original float32 release (44.4 GB) for Apple Silicon.
| tier | size | quantization |
|---|---|---|
| bf16 (this repo) | 22.2 GB | none |
| 8bit | 13.9 GB | 8-bit, LLM + vocab heads only |
| 4bit | 9.2 GB | 4-bit, LLM + vocab heads only |
What is inside
The conversion keeps all four components of the model, complete and in the original layout.
| component | parameters | description |
|---|---|---|
| language backbone | 7.71 B | Nemotron-H hybrid — 56 layers: 27 Mamba-2, 25 MLP, 4 grouped-query attention. Hidden size 4480, vocabulary 131072 |
| vocabulary heads | 1.76 B | token embedding, LM head, and a function-calling head |
| speech encoder | 0.61 B | Conformer encoder with a mel front end |
| speech generation | 1.00 B | Gemma-3 TTS backbone, mixture-of-Gaussians head, neural audio codec, 31-stage residual VQ |
Duplex operation runs at 80 ms frames (12.5 frames/s), 16 kHz input and 22.05 kHz output.
What runs in MLX today
Two parts of the model are directly usable, with runnable examples below:
- the language backbone, via
mlx-lm'snemotron_himplementation - the audio codec — waveform to residual-VQ codes and back
The Conformer speech encoder and the full duplex loop are not yet implemented in MLX. The weights for them are present and complete in this repository.
Example 1 — text generation with the language backbone
pip install mlx mlx-lm
python mlx/extract_llm.py --src . --dst ./voicechat-llm
python mlx/chat_example.py --model ./voicechat-llm --prompt "The capital of France is"
The source model ships no text tokenizer, so extract_llm.py pairs the backbone with a Nemotron-H
base tokenizer (vocabulary 131072) and verifies the size matches before writing.
Example 2 — audio codec round trip
pip install mlx numpy
python mlx/codec_example.py --repo . --wav input.wav --out output.wav
Encodes a waveform to 512-dimensional latents and 31-stage codes, then decodes back to audio. Reconstruction is about 8 dB SNR at 12.5 frames per second, and runs roughly 6x faster than real time on an M4 Max.
Quantization
Weights are stored in bfloat16 without quantization.
Measured performance
Apple M4 Max, 68.7 GB unified memory, language backbone only, greedy decoding:
| load | peak memory | first token | throughput | |
|---|---|---|---|---|
| bf16 | 3.1 s | 17.96 GB | 646 ms | 23.4 tok/s |
| 8-bit | 1.9 s | 10.22 GB | 803 ms | 33.2 tok/s |
Audio codec: about 6x real time.
Notes
- Text-only prompting is outside the model's training distribution. It was trained to emit interleaved
text and audio-codec tokens, so generations may end with audio-side tokens such as
<SPECIAL_12>. - The codec expects 22.05 kHz mono audio. Other rates decode at the wrong speed.
- Every tensor from the source checkpoint is present; the conversion is verified for completeness and for finite values in all tiers.
License and attribution
Released under OpenMDW-1.1, following the original model. Base model and architecture by NVIDIA;
see nvidia/NVIDIA-NemotronLabs-VoiceChat-11B.
The MLX audio-codec implementation follows
NVIDIA NeMo Speech.
- Downloads last month
- 103
Quantized
Model tree for mlx-community/NemotronLabs-VoiceChat-11B-mlx-bf16
Base model
nvidia/NVIDIA-Nemotron-Nano-12B-v2-Base