Is this working on your fork?

#1
by gghfez - opened

Latest from git:

/content/llama.cpp# git log -n 1
commit 02e112f46b11d611ed9a576ac2c15993b230f116 (HEAD -> cohere2-moe-support, origin/cohere2-moe-support)
Merge: 7b08a0eca 4f0e43da6
Author: Csaba Kecskemeti <csaba.kecskemeti@gmail.com>
Date:   Thu May 21 21:07:53 2026 -0700

    Merge branch 'ggml-org:master' into cohere2-moe-support

I just tried to launch it

./build/bin/llama-server -m /content/CohereLabs.command-a-plus-05-2026-bf16-GGUF/CohereLabs.command-a-plus-05-2026-bf16.Q3_K_M-00001-of-00007.gguf --port 6006 -ngl 99 -c 8192
0.00.103.199 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
0.00.103.204 I device_info:
0.00.173.545 I   - CUDA0   : NVIDIA RTX PRO 6000 Blackwell Server Edition (97250 MiB, 96691 MiB free)
0.00.173.560 I   - CPU     : AMD EPYC 9B45 (181130 MiB, 181130 MiB free)
0.00.173.661 I system_info: n_threads = 24 (n_threads_batch = 24) / 48 | CUDA : ARCHS = 1200 | USE_GRAPHS = 1 | PEER_MAX_BATCH_SIZE = 128 | BLACKWELL_NATIVE_FP4 = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | AVX512 = 1 | AVX512_VBMI = 1 | AVX512_VNNI = 1 | AVX512_BF16 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 | 
0.00.173.692 I srv  llama_server: n_parallel is set to auto, using n_parallel = 4 and kv_unified = true
0.00.173.731 I srv          init: running without SSL
0.00.173.771 I srv          init: using 47 threads for HTTP server
0.00.173.866 I srv         start: binding port with default address family
0.00.175.087 I srv  llama_server: loading model
0.00.175.096 I srv    load_model: loading model '/content/CohereLabs.command-a-plus-05-2026-bf16-GGUF/CohereLabs.command-a-plus-05-2026-bf16.Q3_K_M-00001-of-00007.gguf'
0.00.175.141 I common_init_result: fitting params to device memory ...
0.00.175.142 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
0.00.518.064 E llama_model_load: error loading model: check_tensor_dims: tensor 'blk.0.attn_q.weight' has wrong shape; expected   4096,   4096, got   4096,  16384,      1,      1
0.00.518.080 E llama_model_load_from_file_impl: failed to load model
0.00.518.121 E common_fit_params: encountered an error while trying to fit params to free device memory: failed to load model
0.00.743.969 W load: special_eos_id is not in special_eog_ids - the tokenizer config may be incorrect
0.00.843.973 E llama_model_load: error loading model: check_tensor_dims: tensor 'blk.0.attn_q.weight' has wrong shape; expected   4096,   4096, got   4096,  16384,      1,      1
0.00.843.981 E llama_model_load_from_file_impl: failed to load model
0.00.843.988 E common_init_from_params: failed to load model '/content/CohereLabs.command-a-plus-05-2026-bf16-GGUF/CohereLabs.command-a-plus-05-2026-bf16.Q3_K_M-00001-of-00007.gguf'
0.00.843.994 E srv    load_model: failed to load model, '/content/CohereLabs.command-a-plus-05-2026-bf16-GGUF/CohereLabs.command-a-plus-05-2026-bf16.Q3_K_M-00001-of-00007.gguf'
0.00.843.996 I srv    operator(): operator(): cleaning up before exit...
0.00.845.456 E srv  llama_server: exiting due to model loading error
DevQuasar org

Hmm trying to recall is I’ve tested with the llama-server. I think I’ve not. Try the simple-chat . I can took at the state of this later. I have t spent much time to make to fully ready, working on other things.

Sign up or log in to comment

Free AI Image Generator No sign-up. Instant results. Open Now