File size: 16,554 Bytes
f80a379
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2a58e13
 
 
f80a379
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c6fe9df
 
267dd9a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f80a379
 
 
 
 
267dd9a
 
 
f80a379
267dd9a
 
 
f80a379
267dd9a
f80a379
267dd9a
 
 
 
 
 
 
 
f80a379
 
267dd9a
 
f80a379
 
267dd9a
f80a379
 
267dd9a
f80a379
267dd9a
 
f80a379
 
 
 
267dd9a
 
 
 
f80a379
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2ad38c9
 
 
 
 
f80a379
 
 
 
 
267dd9a
 
 
f80a379
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
---
language:
- kab
license: apache-2.0
base_model: ylacombe/omniASR_W2V_300M_SSL
tags:
- kabyle
- taqbaylit
- berber
- amazigh
- automatic-speech-recognition
- speech-to-text
- wav2vec2
- ctc
- low-resource
pipeline_tag: automatic-speech-recognition
metrics:
- wer
- cer
model-index:
- name: Fadhma-300M
  results:
  - task:
      type: automatic-speech-recognition
      name: Kabyle speech recognition
    dataset:
      type: fsicoli/common_voice_22_0
      name: Common Voice 22.0 Kabyle test split (15,003 clips, 888 unseen speakers)
    metrics:
    - type: wer
      value: 25.65
      name: Word error rate (%, beam search with 5-gram shallow fusion)
    - type: cer
      value: 8.01
      name: Character error rate (%, beam search with 5-gram shallow fusion)
    - type: wer
      value: 30.12
      name: Word error rate (%, greedy CTC)
    - type: cer
      value: 8.53
      name: Character error rate (%, greedy CTC)
---

# Fadhma-300M

A 300M-parameter CTC acoustic model for **Kabyle** (Taqbaylit, `kab`, Latin script), trained
on 144.31 hours of read speech and decoded against a 5-gram language model built from three
million normalised Kabyle sentences.

It transcribes Kabyle at **8.01% character error rate** and **25.65% word error rate** on
15,003 held-out utterances from **888 speakers it has never heard**. It is, as far as we can
establish, the first Kabyle speech recognition system published with an error rate measured
on a speaker-disjoint split — and the first CTC result for the language at any size.

Its output is written in canonical Kabyle orthography, because the transcripts it learned
from were repaired before it saw them.

## Results

Common Voice 22.0 Kabyle test split: 15,003 utterances, 16.00 hours, 888 speakers present in
no other split. Hypothesis and reference are reduced by the same normalisation policy before
either is scored.

| decoding | character error rate | word error rate |
|---|---|---|
| **beam search, 5-gram shallow fusion** | **8.01%** | **25.65%** |
| greedy CTC, no language model | 8.53% | 30.12% |

Three things worth reading carefully.

**The language model buys words, not phonemes.** It takes 4.5 points off the word error rate
and 0.5 off the character error rate, which is the shape you should expect: an n-gram cannot
fix a misheard sound, it can only prefer a spelling that is a word. Both rows are reported
because a single number here would hide which half of the system produced it.

**There is no comparison table, deliberately.** Every number in this card was produced by one
harness on this split. No competing system has been run through it, so none appears. Two
figures circulate for Kabyle and neither is comparable: Meta's Omnilingual ASR reports CER
6.2, but that figure lives in `per_language_results_table_7B_llm_asr.csv` — a 7B model with an
LLM decoder, 23× the parameters and the architecture this model deliberately does not use.
No per-language CTC result was published for any language. The honest comparison is
`facebook/omniASR-CTC-300M`: public, Apache-2.0, exactly this size, and unscored on Kabyle by
anyone. Running it costs no training and has not been done. Until it is, treat this model's
error rate as a first measurement rather than a ranking.

**Character error rate and word error rate diverge here for a structural reason.** Kabyle
attaches object pronouns and directional particles with hyphens — `-d`, `-n`, `-awen`, `-iw`.
A hyphen the model writes as a space is one character wrong and two words wrong. At 92%
character accuracy that accounts for a large share of the 25.65%, and it is a property of the
orthography rather than of the acoustics.

## Intended use

Transcribing Kabyle speech: voice notes, oral archives, interviews, broadcast, and speech
input to downstream NLP. The output is canonical Kabyle Latin, so it can be fed to
[Amrouche-1.3B](https://huggingface.co/agbalu/Amrouche-1.3B) or
[Masinissa-31M](https://huggingface.co/agbalu/Masinissa-31M) without a repair step.

**Not suitable for**: any language other than Kabyle; audio with heavy background music or
noise without voice-activity filtering first; or medical, legal or safety-critical
transcription without human review. No safety evaluation of any kind has been performed.

## Why CTC and not an encoder-decoder

This model exists to convert Kabyle's one abundant resource into its scarcest one: the
project has 571 validated hours of speech and roughly 70 million unique tokens of text, and
transcription is what turns the first into the second.

That purpose forbids a decoder. An encoder-decoder carries an internal language model, which
is why Whisper hallucinates fluent text nobody said — and a hallucination here would be
fabricated Kabyle entering a corpus under a provenance record claiming a human spoke it. CTC
emits characters aligned to frames. It can mishear. It cannot invent a sentence.

Whisper is also refuted on measurement, not preference: 6 of the 10 Kabyle-specific letters
have no token in its vocabulary and decode through UTF-8 byte fallback, its fertility on
Kabyle is 3.3600 tokens per word against this project's tokenizer at 1.8837, and at a 3.463 s
mean clip length against a fixed 30 s window, 11.5% of its compute is signal.

## Architecture

Fine-tuned from Meta's Omnilingual ASR SSL encoder
([`ylacombe/omniASR_W2V_300M_SSL`](https://huggingface.co/ylacombe/omniASR_W2V_300M_SSL)),
which lists `kab_Latn` among the 1,668 languages in its pretraining.

| | |
|---|---|
| Total parameters | **315,482,152** |
| Trainable | 311,287,848 |
| Frozen | 4,194,304 — the 7-layer convolutional feature extractor |
| Layers / hidden / heads | 24 / 1,024 / 16 (head size 64) |
| Feed-forward | 4,096, GELU |
| Front end | 7 convolutions, kernel 10,3,3,3,3,2,2, stride 5,2,2,2,2,2,2 — one frame per 20 ms |
| CTC classes | **40**: `[PAD]`, `[UNK]`, `\|`, 36 Kabyle letters and `-` |
| Loss | CTC, `ctc_zero_infinity=True` |

The convolutional front end is frozen, so the acoustic representation learned across Meta's
multilingual pretraining survives contact with 144 hours of one language.

**The 40 classes are derived from the transcripts' own characters — after normalisation, not
before, and the ordering is the whole point.** Common Voice's Kabyle transcripts are raw
contributor text and carry the same homoglyph substitution as every other Kabyle source:
Greek `ε` accounts for **10.6%** of that letter's occurrences. Take the inventory first and
the homoglyph becomes a 41st class, so the model learns to reproduce the corruption at corpus
scale. Gate on the alphabet without normalising and the character is *deleted* instead, which
splits the word in two and makes every word error rate computed over it wrong. Normalising
first avoids both: 40 classes from a 128-character raw inventory, with 1.73% of rows repaired.

Nobody else fine-tuning Kabyle ASR has a reference orthography to check against. That is what
the first four phases of this project bought.

## Training data

**Common Voice 22.0 Kabyle**, CC-0, normalised at `normaliser` **1.3.0+rules1.0.0**.

| split | clips | hours | speakers |
|---|---|---|---|
| train | 152,478 | 144.31 | 162 |
| dev | 15,002 | 15.21 | 135 |
| test | 15,003 | 16.00 | **888** |

Speaker overlap between splits is **zero**, verified on the built splits rather than trusted
from the upstream partition. Every clip passed a CTC feasibility check first — a target longer
than the available acoustic frames makes the loss infinite.

The weights are Apache-2.0 and every clip and transcript is CC-0, so unlike this project's
text models there is no licence composition to disclose: all 182,483 clips are public domain.

## Training recipe

| | |
|---|---|
| Objective | CTC |
| Optimiser | AdamW, lr 1e-4, β (0.9, 0.98), weight decay 0.01 |
| Batch | dynamic micro-batches of **160 audio-seconds**, gradient accumulation 4 |
| Schedule | 250 warmup steps, cosine decay to 1e-5 |
| Steps | **6,568** — 8 epochs |
| Precision | bfloat16 autocast |
| Masking | `mask_time_prob` 0.05, `layerdrop` 0.0 |
| Hardware | one NVIDIA A10G, 24 GiB |
| Seed | 20260813 |

Validation improved monotonically across every evaluation:

| step | 2,500 | 4,500 | 5,250 | 5,750 | 6,000 | **6,568** |
|---|---|---|---|---|---|---|
| dev loss | 0.4275 | 0.3599 | 0.3572 | 0.3506 | 0.3425 | **0.3340** |
| greedy CER | 10.83% | 9.49% | 9.24% | 8.97% | 8.67% | **8.53%** |
| greedy WER | 36.67% | 33.76% | 32.69% | 31.62% | 30.78% | **30.12%** |

**160 audio-seconds is a measured ceiling, not a round number.** At 320 the run died asking
the allocator for 1.95 GiB in a single tensor: wav2vec2's first convolution turns one 16 kHz
channel into 512 channels at 3.2 kHz, and `layer_norm` sits on autocast's fp32 list, so a
budget stated in seconds of audio reaches the allocator expanded 102×. The first step
survived it and the second did not, because AdamW does not allocate its moment buffers until
the first optimizer step — a training loop that survives one step has not proved it fits.

## Language model

`5gram.klm`, 186.65 MB, built from **AƔBALU-Text v1**, the largest clean Kabyle text
corpus assembled — **3,246,174 lines from its 3,041,989 records**, because 204,185 of
them carry an internal newline and `lmplz` reads a sentence per line:

```
lmplz --order 5 --prune 0 1 2 3 --skip_symbols --discount_fallback
build_binary trie 5gram.arpa 5gram.klm
```

`--skip_symbols` is required rather than defensive: the corpus contains lines carrying `<s>`
and `</s>`, and `lmplz` aborts on them. `trie` rather than `probing` because the model is
memory-mapped once per container and array compression is what makes 186 MB affordable beside
the acoustic model.

Those arguments live in `agbalu.speech.lm` as values the test suite asserts on, so what
produced the published binary is checkable rather than a shell line in a document.

Shallow fusion runs at α 0.5, β 1.5 — `pyctcdecode`'s documented defaults, **unswept**. A
sweep needs the dev split decoded once per setting, which is GPU time this has not had, so
the 4.5-point word error reduction above is a floor on what the fusion is worth rather than a
tuned result.

## Usage

Greedy, with `transformers`, `torch` and `librosa` — the architecture is `Wav2Vec2ForCTC`,
one of the library's own, so no `trust_remote_code` is needed:

```python
import librosa
import torch
from transformers import AutoProcessor, Wav2Vec2ForCTC

REPO = "agbalu/Fadhma-300M"
processor = AutoProcessor.from_pretrained(REPO)
model = Wav2Vec2ForCTC.from_pretrained(REPO).eval()

waveform, _ = librosa.load("sample.wav", sr=16_000, mono=True)
inputs = processor(waveform, sampling_rate=16_000, return_tensors="pt")
with torch.inference_mode():
    logits = model(inputs.input_values).logits

print(processor.batch_decode(logits.argmax(-1))[0])
```

That is **CER 8.53 / WER 30.12**. The 5-gram takes it to 8.01 / 25.65, and needs
`pyctcdecode` and `kenlm`:

```python
import json

import librosa
import torch
from huggingface_hub import hf_hub_download
from pyctcdecode import build_ctcdecoder
from transformers import AutoProcessor, Wav2Vec2ForCTC

REPO = "agbalu/Fadhma-300M"
processor = AutoProcessor.from_pretrained(REPO)
model = Wav2Vec2ForCTC.from_pretrained(REPO).eval()

with open(hf_hub_download(REPO, "vocab.json"), encoding="utf-8") as handle:
    vocabulary = json.load(handle)

# `pyctcdecode` reads label 0 as the CTC blank and a literal space as the word delimiter,
# where this vocabulary writes `[PAD]` and `|`.
labels = [token for token, _ in sorted(vocabulary.items(), key=lambda item: item[1])]
labels = ["" if token == "[PAD]" else " " if token in ("|", "[UNK]") else token
          for token in labels]

decoder = build_ctcdecoder(labels, kenlm_model_path=hf_hub_download(REPO, "5gram.klm"))

waveform, _ = librosa.load("sample.wav", sr=16_000, mono=True)
inputs = processor(waveform, sampling_rate=16_000, return_tensors="pt")
with torch.inference_mode():
    logits = model(inputs.input_values).logits[0].float().numpy()

print(decoder.decode(logits, beam_width=100, alpha=0.5, beta=1.5))
```

Three things above are not stylistic and will produce wrong output if changed.

**The label substitution.** Handed this vocabulary unchanged, `pyctcdecode` emits `[PAD]`
and `|` verbatim in every transcript, and every error rate computed from one is wrong.

**`.float()` before `.numpy()`.** The model runs in bfloat16, which numpy cannot represent:
without it, `numpy()` raises `TypeError: Got unsupported ScalarType BFloat16`.

**`AutoProcessor`, not a bare `Wav2Vec2FeatureExtractor()`.** The published extractor sets
`do_normalize: true`; a default one feeds the model a distribution it never saw and returns
a plausible, worse transcript with nothing raising.

## Limitations

**162 training speakers against 888 in the test split. This is the binding constraint.**
It is a property of the corpus, not of the recipe: Common Voice Kabyle is 571 validated hours
contributed by a small number of people, and no amount of additional training reaches speakers
who are not in it. Wherever this model fails, speaker generalisation is the first place to
look, and the roughly 395 validated hours outside these three splits are what would move it.

**Hyphenated clitics are where character accuracy and word accuracy part company**, as above.

**Read speech only.** Common Voice is read aloud from written prompts. Spontaneous
conversation, overlapping speakers, telephone bandwidth and regional variation are entirely
unbenchmarked, and the register mismatch cuts both ways: the language model was built from
*written* Kabyle, which is not what a prompt reader sounds like either.

**Not evaluated beyond this split.** One corpus, one domain, 15,003 test utterances. Treat
everything else as unmeasured.

**No safety evaluation of any kind** has been performed.

**The output carries no punctuation and no capitals.** The CTC vocabulary is 40 characters and
none of them is a mark, so that is a property of the architecture rather than a shortfall of
the training. [`agbalu/Belaid-31M`](https://huggingface.co/agbalu/Belaid-31M) puts both back
— one utterance at a time, which is the shape this model emits.

## Files

| file | description |
|---|---|
| `model.safetensors` | 315,482,152 acoustic parameters |
| `config.json` | the architecture, `Wav2Vec2ForCTC` with a 40-class head |
| `preprocessor_config.json` | the feature extractor, with `do_normalize: true` |
| `vocab.json`, `tokenizer_config.json`, `added_tokens.json` | the 40 CTC classes as a `Wav2Vec2CTCTokenizer`, so `AutoProcessor` resolves |
| `5gram.klm` | the KenLM binary, 186.65 MB |

The training checkpoint, with AdamW's moments and the resume state, is not published. Ask if
you need it.

## Reproduction

```bash
make modal-asr-repack                  # decode the audio once, CPU only
make modal-asr TASK=lm                 # build the 5-gram
make modal-asr-train EPOCHS=8          # train
make modal-asr TASK=evaluate           # score the test split, both decodings
```

The evaluation writes its results to the run's volume rather than to the repository, so the
figures in this card come from that run's own output. Re-running the last command regenerates
them.

## The name

**Fadhma Aït Mansour Amrouche** (1882–1967) began, in 1930, writing down the songs and tales
she had inherited from her ancestors — hearing Kabyle and putting it on paper, which is
precisely what this model does.

Her daughter **Taos** names the translation model
[Amrouche-1.3B](https://huggingface.co/agbalu/Amrouche-1.3B), and the pair split the
oral-tradition work the way the two models do: the mother wrote the tradition down, the
daughter carried it into French while singing only in Kabyle.

The naming is homage; it implies no endorsement by anyone.

## Citation

```bibtex
@software{agbalu_fadhma_2026,
  title  = {Fadhma-300M: CTC speech recognition for Kabyle},
  author = {AƔBALU},
  year   = {2026},
  url    = {https://huggingface.co/agbalu/Fadhma-300M},
  note   = {Fine-tuned from ylacombe/omniASR_W2V_300M_SSL on Common Voice 22.0 Kabyle;
            normaliser 1.3.0+rules1.0.0}
}
```

## Licence

**Apache-2.0** for the weights and code. The audio and transcripts are CC-0 from Mozilla
Common Voice. The language model derives from AƔBALU-Text v1, a third of which has no
resolvable licence — see [Masinissa-31M](https://huggingface.co/agbalu/Masinissa-31M) for that
corpus's composition before redistributing derivatives of the `.klm` binary.