Instructions to use WaveCut/Nanbeige4.2-3B-heretic-MLX-DWQ-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use WaveCut/Nanbeige4.2-3B-heretic-MLX-DWQ-4bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("WaveCut/Nanbeige4.2-3B-heretic-MLX-DWQ-4bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use WaveCut/Nanbeige4.2-3B-heretic-MLX-DWQ-4bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "WaveCut/Nanbeige4.2-3B-heretic-MLX-DWQ-4bit"
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "WaveCut/Nanbeige4.2-3B-heretic-MLX-DWQ-4bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use WaveCut/Nanbeige4.2-3B-heretic-MLX-DWQ-4bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "WaveCut/Nanbeige4.2-3B-heretic-MLX-DWQ-4bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default WaveCut/Nanbeige4.2-3B-heretic-MLX-DWQ-4bit
Run Hermes
hermes
- OpenClaw new
How to use WaveCut/Nanbeige4.2-3B-heretic-MLX-DWQ-4bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "WaveCut/Nanbeige4.2-3B-heretic-MLX-DWQ-4bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "WaveCut/Nanbeige4.2-3B-heretic-MLX-DWQ-4bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- MLX LM
How to use WaveCut/Nanbeige4.2-3B-heretic-MLX-DWQ-4bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "WaveCut/Nanbeige4.2-3B-heretic-MLX-DWQ-4bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "WaveCut/Nanbeige4.2-3B-heretic-MLX-DWQ-4bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WaveCut/Nanbeige4.2-3B-heretic-MLX-DWQ-4bit", "messages": [ {"role": "user", "content": "Hello"} ] }'
Nanbeige4.2-3B Heretic MLX DWQ 4-bit
A 4-bit, group-size-32 Distilled Weight Quantization (DWQ) release of
WaveCut/Nanbeige4.2-3B-heretic.
This repository includes a small trusted-code MLX-LM adapter because Nanbeige reuses 22 physical decoder layers over two loops and is not a standard Llama layout at runtime. The adapter preserves shared weights while allocating 44 independent KV caches, one for each loop/layer execution. Unsupported optional Nanbeige architectures are rejected explicitly.
DWQ calibration
- 4 bits, group size 32.
- 1,024 training samples and 32 validation samples.
- Maximum sequence length: 1,025 tokens.
- Seed: 20260722.
- Corpus: 528 agentic trajectories plus 528 coding-reasoning examples, deterministically shuffled.
- Corpus SHA-256:
a7cfdbe02c124304bf1282bbd5ed7162bfa72dec6750b60ed2d3a68000c7a554. - Agentic source:
TIGER-Lab/SWE-QA-Pro-SFT-Trajectoriesatb8f5b8a8dcf90bca8b6d70adedac0d20dca02b86. - Coding source:
nvidia/OpenCodeReasoningat20a1ca19c0d050fe9057fc08339d6b370ec1c67a.
| Validation loss | Value |
|---|---|
| Initial RTN | 0.284 |
| Final DWQ | 0.043 |
MLX-LM revision: cf10f962b7a20e63a6df43dbf0faf06070153d40.
Usage
The model file is repository code, so load it only after reviewing
nanbeige_mlx.py and pass --trust-remote-code.
mlx_lm.generate \
--model WaveCut/Nanbeige4.2-3B-heretic-MLX-DWQ-4bit \
--trust-remote-code \
--prompt "Implement a bounded async worker pool in Python." \
--max-tokens 256
Exact artifact hashes and clean-load smoke-test results are recorded in
release-manifest.json.
- Downloads last month
- 480
4-bit
Model tree for WaveCut/Nanbeige4.2-3B-heretic-MLX-DWQ-4bit
Base model
Nanbeige/Nanbeige4.2-3B-Base