Instructions to use gghfez/DeepSeek-V3.1-IQ2_KS with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- llama-cpp-python
How to use gghfez/DeepSeek-V3.1-IQ2_KS with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="gghfez/DeepSeek-V3.1-IQ2_KS", filename="DeepSeek-V3.1-IQ2_KS-00001-of-00005.gguf", )
llm.create_chat_completion( messages = [ { "role": "user", "content": "What is the capital of France?" } ] ) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use gghfez/DeepSeek-V3.1-IQ2_KS with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf gghfez/DeepSeek-V3.1-IQ2_KS:Q2_K # Run inference directly in the terminal: llama cli -hf gghfez/DeepSeek-V3.1-IQ2_KS:Q2_K
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf gghfez/DeepSeek-V3.1-IQ2_KS:Q2_K # Run inference directly in the terminal: llama cli -hf gghfez/DeepSeek-V3.1-IQ2_KS:Q2_K
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf gghfez/DeepSeek-V3.1-IQ2_KS:Q2_K # Run inference directly in the terminal: ./llama-cli -hf gghfez/DeepSeek-V3.1-IQ2_KS:Q2_K
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf gghfez/DeepSeek-V3.1-IQ2_KS:Q2_K # Run inference directly in the terminal: ./build/bin/llama-cli -hf gghfez/DeepSeek-V3.1-IQ2_KS:Q2_K
Use Docker
docker model run hf.co/gghfez/DeepSeek-V3.1-IQ2_KS:Q2_K
- LM Studio
- Jan
- vLLM
How to use gghfez/DeepSeek-V3.1-IQ2_KS with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "gghfez/DeepSeek-V3.1-IQ2_KS" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "gghfez/DeepSeek-V3.1-IQ2_KS", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/gghfez/DeepSeek-V3.1-IQ2_KS:Q2_K
- Ollama
How to use gghfez/DeepSeek-V3.1-IQ2_KS with Ollama:
ollama run hf.co/gghfez/DeepSeek-V3.1-IQ2_KS:Q2_K
- Unsloth Studio
How to use gghfez/DeepSeek-V3.1-IQ2_KS with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for gghfez/DeepSeek-V3.1-IQ2_KS to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for gghfez/DeepSeek-V3.1-IQ2_KS to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for gghfez/DeepSeek-V3.1-IQ2_KS to start chatting
- Atomic Chat new
- Docker Model Runner
How to use gghfez/DeepSeek-V3.1-IQ2_KS with Docker Model Runner:
docker model run hf.co/gghfez/DeepSeek-V3.1-IQ2_KS:Q2_K
- Lemonade
How to use gghfez/DeepSeek-V3.1-IQ2_KS with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull gghfez/DeepSeek-V3.1-IQ2_KS:Q2_K
Run and chat with the model
lemonade run user.DeepSeek-V3.1-IQ2_KS-Q2_K
List all available models
lemonade list
ik_llama.cpp imatrix Quantizations of deepseek-ai/DeepSeek-V3.1
This quant REQUIRES ik_llama.cpp fork to support the ik's latest SOTA quants and optimizations! Do not download these big files and expect them to run on mainline vanilla llama.cpp, ollama, LM Studio, KoboldCpp, etc!
NOTE ik_llama.cpp can also run your existing GGUFs from bartowski, unsloth, mradermacher, etc if you want to try it out before downloading my quants.
I made this for myself and my RAM+VRAM setup. For more ik_llama quants of this model, discussions, perplexity measurements, see @ubergarm's DeepSeek-V3.1 Collection
👈 Quant details
#!/usr/bin/env bash
custom="
# First 3 dense layers (0-3) (GPU)
# Using q8_0 for attn_k_b since imatrix might not have these tensors
blk\.[0-2]\.attn_k_b.*=q8_0
blk\.[0-2]\.attn_.*=iq5_ks
blk\.[0-2]\.ffn_down.*=iq5_ks
blk\.[0-2]\.ffn_(gate|up).*=iq4_ks
blk\.[0-2]\..*=iq5_ks
# All attention, norm weights, and bias tensors for MoE layers (3-60) (GPU)
# Using q8_0 for attn_k_b since imatrix might not have these tensors
blk\.[3-9]\.attn_k_b.*=q8_0
blk\.[1-5][0-9]\.attn_k_b.*=q8_0
blk\.60\.attn_k_b.*=q8_0
blk\.[3-9]\.attn_.*=iq5_ks
blk\.[1-5][0-9]\.attn_.*=iq5_ks
blk\.60\.attn_.*=iq5_ks
# Shared Expert (3-60) (GPU)
blk\.[3-9]\.ffn_down_shexp\.weight=iq5_ks
blk\.[1-5][0-9]\.ffn_down_shexp\.weight=iq5_ks
blk\.60\.ffn_down_shexp\.weight=iq5_ks
blk\.[3-9]\.ffn_(gate|up)_shexp\.weight=iq4_ks
blk\.[1-5][0-9]\.ffn_(gate|up)_shexp\.weight=iq4_ks
blk\.60\.ffn_(gate|up)_shexp\.weight=iq4_ks
# Routed Experts (3-60) (CPU)
blk\.[3-9]\.ffn_down_exps\.weight=iq3_ks
blk\.[1-5][0-9]\.ffn_down_exps\.weight=iq3_ks
blk\.60\.ffn_down_exps\.weight=iq3_ks
blk\.[3-9]\.ffn_(gate|up)_exps\.weight=iq2_ks
blk\.[1-5][0-9]\.ffn_(gate|up)_exps\.weight=iq2_ks
blk\.60\.ffn_(gate|up)_exps\.weight=iq2_ks
# Token embedding and output tensors (GPU)
token_embd\.weight=iq5_k
output\.weight=q8_0 # Changed to q8_0
"
custom=$(
echo "$custom" | grep -v '^#' | \
sed -Ez 's:\n+:,:g;s:,$::;s:^,::'
)
./build/bin/llama-quantize \
--custom-q "$custom" \
--imatrix /fast/DeepSeek-V3.1.imatrix \
/fast/bf16/DeepSeek-V3-00001-of-00030.gguf
/fast2/quants/DeepSeek-V3.1-IQ2_KS.gguf \
IQ2_KS \
- Downloads last month
- 18
2-bit
Model tree for gghfez/DeepSeek-V3.1-IQ2_KS
Base model
deepseek-ai/DeepSeek-V3.1-Base