WeSpeaker ResNet34-LM — GGUF (ggml conversion)
GGUF conversion of
Wespeaker/wespeaker-voxceleb-resnet34-LM,
a 256-dimensional speaker-embedding model, for the --diarize-method foxnose
diarizer in CrispStrobe/CrispASR.
⚠ Licence — attribution is required
These weights are CC-BY-4.0, inherited from the upstream model. Several downstream projects describe them as Apache-2.0; that is incorrect — the wenet-e2e/wespeaker code is Apache-2.0, the published weights are CC-BY-4.0 and carry an attribution requirement. If you redistribute these files, keep the attribution.
Source: Wespeaker/wespeaker-voxceleb-resnet34-LM
Upstream: https://github.com/wenet-e2e/wespeaker
Licence: https://creativecommons.org/licenses/by/4.0/
Files
| File | Size | Notes |
|---|---|---|
wespeaker-resnet34-lm-f32.gguf |
26.5 MB | reference precision |
wespeaker-resnet34-lm.gguf |
23.9 MB | conv kernels F32, linear F16 — recommended |
Conv kernels stay F32 in both, and that is a speed choice. An F16 conv kernel is numerically fine (cosine 0.99999724 against the PyTorch oracle) and would shrink the file to 13.3 MB, but ggml's CPU conv path is 2.2× slower on it — 297 ms vs 133 ms per 1.2 s window, measured back to back over the same 352 windows on an M1. ResNet34 is ~94% of embedding time, so 10 MB of disk is not worth it. F16 on the 2-D linear is free and yields an identical embedding (cosine 0.99999744 vs 0.99999747).
Architecture
ResNet34 [3,4,6,3] over 80-bin Kaldi fbank, TSTP pooling, Linear(5120→256).
BatchNorm is folded into every convolution at conversion time (219 → 74
tensors) and the ArcMargin projection head is training-only and dropped
(11.25 M → 6.6 M params).
Three details that decide correctness, traced to wespeaker/cli/speaker.py
rather than assumed:
- the waveform is int16-scale (
torchaudio.load(normalize=False), sincewavform_normdefaults to False), window is hamming, then per-utterance CMN; - the 2-D map is height=freq, width=time — TSTP reduces over time and
flattens (channel, freq) with freq fastest, which is the order
seg_1's 5120 columns are in; - TSTP's std uses torch's unbiased (n−1) variance,
+1e-7inside the sqrt; - the output is
seg_1(stats)raw — no ReLU, no BatchNorm, no L2 normalisation.
Verification
Per-stage against the upstream PyTorch model run as an oracle
(crispasr-diff wespeaker), on an 11 s clip:
| stage | cos_mean |
|---|---|
| fbank | 0.999999 |
| stem / layer1–4 | 0.99997 – 0.999995 |
| stats | 0.999999 |
| embedding | 0.999997, cosine(emb, ref) 0.99999747 |
Discriminative check on real audio: two windows of the same speaker score cosine 0.595, against 0.100 for a different speaker.
End-to-end, the CrispASR diarizer built on this model scores DER 3.93% against the upstream Python pipeline's own output (same pinned speaker count, 0.25 s collar) with zero speaker confusion — the residual is entirely false alarm from a different speech-segmentation source.
⚠ That is a PARITY number, not an accuracy one. It says this port
reproduces the reference implementation; it does not say either is 96% right.
Measured against HUMAN labels on VoxConverse dev (40 files, --diarize-max- speakers 8, whisper-tiny segments, 0.25 s collar):
| value | |
|---|---|
| DER | 33.1% |
| speaker count exactly right | 18/40 (45%) |
| within ±1 speaker | 34/40 |
Estimating the number of speakers is the weak link, not the embeddings — the
embedding matches the PyTorch oracle to cosine 0.99999747. Pass
--diarize-num-speakers N when you know the count and the picture improves
sharply. Numbers near 3–7% DER quoted elsewhere in this project's history came
from an 8-file subset that turned out to be unrepresentative: identical code
scores 7.3% there and 33.1% on the fuller corpus.
Usage
crispasr -m <asr-model.gguf> -f audio.wav \
--diarize --diarize-method foxnose \
--diarize-embedder wespeaker-resnet34-lm.gguf
Conversion
python models/convert-wespeaker-to-gguf.py \
--model Wespeaker/wespeaker-voxceleb-resnet34-LM \
--output wespeaker-resnet34-lm.gguf
The CrispASR runtime is an independent implementation written from the published architecture; no upstream source is incorporated.
Provenance and EU AI Act Art. 53 note
- Upstream model: Wespeaker/wespeaker-voxceleb-resnet34-LM — published by
Wespeaker. - Upstream licence:
cc-by-4.0. This repository redistributes under the same terms; it grants no rights the upstream licence does not. - What was done here: format conversion and/or quantisation only (GGUF/GGML). No training, no fine-tuning, no merging, no distillation, no change to architecture, vocabulary or capability. Only the numeric representation of the upstream weights differs.
- Training data: documented — where it is documented at all — by the upstream provider; see the upstream model card. No training data was used, added or selected by this repository.
- Provider status: under Regulation (EU) 2024/1689 the upstream authors remain the provider of this model. Converting the serialisation format does not make this repository the provider of a new general-purpose AI model, and no such claim is made. Questions about training content, copyright policy or model capability belong upstream.
- Downloads last month
- 582
32-bit
Model tree for cstr/wespeaker-resnet34-lm-GGUF
Base model
Wespeaker/wespeaker-voxceleb-resnet34-LM