NemotronLabs-VoiceChat-11B — MLX (bf16)

MLX conversion of nvidia/NVIDIA-NemotronLabs-VoiceChat-11B, a duplex speech-to-speech model that listens and speaks at the same time. 11.1 B parameters, converted from the original float32 release (44.4 GB) for Apple Silicon.

tier size quantization
bf16 (this repo) 22.2 GB none
8bit 13.9 GB 8-bit, LLM + vocab heads only
4bit 9.2 GB 4-bit, LLM + vocab heads only

What is inside

The conversion keeps all four components of the model, complete and in the original layout.

component parameters description
language backbone 7.71 B Nemotron-H hybrid — 56 layers: 27 Mamba-2, 25 MLP, 4 grouped-query attention. Hidden size 4480, vocabulary 131072
vocabulary heads 1.76 B token embedding, LM head, and a function-calling head
speech encoder 0.61 B Conformer encoder with a mel front end
speech generation 1.00 B Gemma-3 TTS backbone, mixture-of-Gaussians head, neural audio codec, 31-stage residual VQ

Duplex operation runs at 80 ms frames (12.5 frames/s), 16 kHz input and 22.05 kHz output.

What runs in MLX today

Two parts of the model are directly usable, with runnable examples below:

  • the language backbone, via mlx-lm's nemotron_h implementation
  • the audio codec — waveform to residual-VQ codes and back

The Conformer speech encoder and the full duplex loop are not yet implemented in MLX. The weights for them are present and complete in this repository.

Example 1 — text generation with the language backbone

pip install mlx mlx-lm
python mlx/extract_llm.py --src . --dst ./voicechat-llm
python mlx/chat_example.py --model ./voicechat-llm --prompt "The capital of France is"

The source model ships no text tokenizer, so extract_llm.py pairs the backbone with a Nemotron-H base tokenizer (vocabulary 131072) and verifies the size matches before writing.

Example 2 — audio codec round trip

pip install mlx numpy
python mlx/codec_example.py --repo . --wav input.wav --out output.wav

Encodes a waveform to 512-dimensional latents and 31-stage codes, then decodes back to audio. Reconstruction is about 8 dB SNR at 12.5 frames per second, and runs roughly 6x faster than real time on an M4 Max.

Quantization

Weights are stored in bfloat16 without quantization.

Measured performance

Apple M4 Max, 68.7 GB unified memory, language backbone only, greedy decoding:

load peak memory first token throughput
bf16 3.1 s 17.96 GB 646 ms 23.4 tok/s
8-bit 1.9 s 10.22 GB 803 ms 33.2 tok/s

Audio codec: about 6x real time.

Notes

  • Text-only prompting is outside the model's training distribution. It was trained to emit interleaved text and audio-codec tokens, so generations may end with audio-side tokens such as <SPECIAL_12>.
  • The codec expects 22.05 kHz mono audio. Other rates decode at the wrong speed.
  • Every tensor from the source checkpoint is present; the conversion is verified for completeness and for finite values in all tiers.

License and attribution

Released under OpenMDW-1.1, following the original model. Base model and architecture by NVIDIA; see nvidia/NVIDIA-NemotronLabs-VoiceChat-11B. The MLX audio-codec implementation follows NVIDIA NeMo Speech.

Downloads last month
103
Safetensors
Model size
11B params
Tensor type
BF16
·
I64
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/NemotronLabs-VoiceChat-11B-mlx-bf16

Collection including mlx-community/NemotronLabs-VoiceChat-11B-mlx-bf16

Free AI Image Generator No sign-up. Instant results. Open Now