CrisperWhisper2 Korean Verbatim (base, 74M)

Korean verbatim ASR (74M), the base of a three-size family. Transcribes spontaneous Korean as spoken, keeping the fillers and repeats (μ–΄, 음, κ·Έ, 막 ...) that standard ASR drops. Built by transferring CrisperWhisper2.0's verbatim strength (an English-only model) into Korean.

CER 0.167 at 74M, beating whisper-small (244M, 0.171) at one third the size.

Non-commercial / research use only. Derived from a non-commercial teacher (CrisperWhisper2.0) and restricted data (KsponSpeech). See Licensing below.

Model family

Model Params Korean CER
small 244M 0.103
base (this model) 74M 0.167
tiny 39M 0.216

Method

This base model is soft-KD distilled from our Korean small model (there is no CrisperWhisper checkpoint at this size to initialize from, so the advantage is transferred by distillation).

Results: vs Whisper and CrisperWhisper

Held-out KsponSpeech valid, 300 spontaneous Korean clips. CER lower is better; disfluency = fillers kept; tag-leak = leaked markup tokens (0 is clean).

Model Size CER Disfluency Tag-leak
small (CrisperWhisper-init, ours) 244M 0.103 92% 0
base (KD from small, ours) 74M 0.167 92% 0
tiny (KD from small, ours) 39M 0.216 86% 0
small-B (whisper-init, ablation) 244M 0.265 92% 0
whisper-small (baseline) 244M 0.171 60% 0
CrisperWhisper2.0_small (baseline, markup-suppressed) 244M 0.682 60% 0

CrisperWhisper2.0_small is run with its 31 markup tokens suppressed, so it emits clean text instead of looping on [verbatim_N]. This barely changes its CER (0.68 vs 0.71): it is English-centric and does Korean poorly either way.

What it shows:

  • Our base (74M) beats whisper-small (244M) on Korean (0.167 vs 0.171) at one third the size.
  • CrisperWhisper-init (0.103) vs whisper-init (0.265), same data: the CrisperWhisper prior transfers to Korean and cuts CER 2.6x. Both reached the same training loss, so the gain is generalization.
  • base-KD (74M) beats the whisper-init 244M ablation (0.265): the advantage propagates down via distillation.
  • Our models keep 86 to 92% of disfluencies. The original CrisperWhisper, even markup-suppressed, sits at CER 0.682 on Korean: it drops clauses, hallucinates English, or repeats characters.

Transcription examples

Complete outputs on held-out clips (nothing trimmed). REF is the human verbatim transcript; CrisperWhisper is the fair markup-suppressed run.

Example 1: a long turn opening with the filler γ€ŒμŒγ€

REF             음 그리고 인디 κ²Œμž„ μ•„λ‹ˆ 인디 κ²Œμž„μ΄λž˜ 인디 μŒμ•…λ„ μ’‹μ•„ν•œλ‹€κ³  ν•˜λ”λΌκ³  κ·Έλž˜μ„œ 듀어봐야 λ˜λŠ”λ° κ·Έκ±° μ•„μ§κΉŒμ§€ λͺ» λ“€μ–΄λ΄€μ–΄ 였늘 집에 κ°€λ©΄μ„œ κΌ­ 듀어봐야지. 응.
small (ours)    음 그리고 인디 κ²Œμž„ μ•„λ‹ˆ 인디 κ²Œμž„μ΄λž˜. 인디 μŒμ•…λ„ μ’‹μ•„ν•œλ‹€κ³  ν•˜λ”λΌκ³ . κ·Έλž˜μ„œ 듀어봐야 λ˜λŠ”λ° κ·Έκ±° μ•„μ§κΉŒμ§€ λͺ» λ“€μ–΄λ΄€μ–΄. 였늘 집에 κ°€λ©΄μ„œ κΌ­ 듀어봐야지. 응.
base  (ours)    응. 그리고 인디 κ²Œμž„ μ•„λ‹ˆ 인디 κ²Œμž„μ΄λΌ 인디 μŒμ•…λ„ μ’‹μ•„ν•œλ‹€κ³  ν•˜λ”λΌκ³ . κ·Έλž˜μ„œ 듀어봐야 λ˜λŠ”λ° κ·Έ μ•„μ§κΉŒμ§€ λͺ» λ“€μ–΄λ΄€μ–΄. 였늘 집에 κ°€λ©΄μ„œ κΌ­ 듀어봐야지. 응.
tiny  (ours)    음 그리고 인디 κ²Œμž„ μ•„λ‹ˆ 인디 κ²Œμž„μ΄λΌ 인디 μŒμ•…λ„ μ’‹μ•„ν•œλ‹€κ³  ν•˜λ”λΌκ³  κ·Έλž˜μ„œ 듀어봐야 λ˜λŠ”λ° κ·Έκ±° μ•„μ§κΉŒμ§€ λͺ» λ“€μ–΄λ΄€μ–΄. 였늘 집에 κ°€λ©΄μ„œ κΌ­ 듀어봐야지. 응.
whisper-small   그리고 μΈλ””κ²Œμž„.. μ•„λ‹ˆ μΈλ””κ²Œμž„μ΄λž˜ 인디 μŒμ•…λ„ μ’‹μ•„ν•œλ‹€κ³  ν•˜λ”λΌκ³  κ·Έλž˜μ„œ 듀어봐야 λ˜λŠ”λ° μ•„μ§κΉŒμ§€ λͺ» λ“€μ–΄λ΄€μ–΄ 였늘 집에 κ°€λ©΄μ„œ κΌ­ 듀어봐야지
CrisperWhisper  그리고 인디 μŒμ•…λ„ μ’‹μ•„ν•œλ‹€ ν•˜λ”λΌκ³ ? κ·Έλž˜μ„œ 듀어봐야 ν•˜λŠ”λ° μ•„μ§κΉŒμ§€ λͺ» λ“€μ–΄λ΄€μ–΄. 였늘 집에 κ°€λ©΄μ„œ κΌ­ 듀어봐야지.

What differs: our models reproduce the full turn, opening filler γ€ŒμŒγ€ and closing γ€Œμ‘γ€ included. whisper-small drops the γ€ŒμŒγ€, γ€Œκ·Έκ±°γ€, and the final γ€Œμ‘γ€. CrisperWhisper drops the whole γ€ŒμΈλ”” κ²Œμž„ μ•„λ‹ˆ 인디 κ²Œμž„μ΄λž˜γ€ clause.

Example 2: filler γ€Œμ–΄γ€ plus a self-repeat γ€Œλ‚˜ λ‹Œν…λ„ ... λ‚˜ λ‹Œν…λ„γ€

REF             μ–΄ λ‹Œν…λ„ 맨 μ²˜μŒμ— λ‚˜ λ‹Œν…λ„ 맨 μ²˜μŒμ— λ‚˜μ˜¬ λ•Œ 그런 건데 μ•„ 근데 μ§„μ§œ μ’€ μ—¬
small (ours)    μ–΄. λ‹Œν…λ„ 맨 μ²˜μŒμ— λ‚˜ λ‹Œν…λ„ 맨 μ²˜μŒμ— λ‚˜μ˜¬ λ•Œ 그런 건데 μ•„ 근데 μ§„μ§œ μ’€ μ—¬
base  (ours)    μ–΄ λ‹Œν…λ„ 맨 μ²˜μŒμ— λ‚˜ λ‹Œν…λ„ 맨 μ²˜μŒμ— λ‚˜μ˜¬ λ•Œ 그런 건데 μ•„ 근데 μ§„μ§œ μ’€ μ—¬ν–‰
tiny  (ours)    μ–΄. λ‹Œν…λ„ 맨 μ²˜μŒμ—λŠ” λ‹Œν…λ„ 맨 μ²˜μŒμ—λŠ” ν•  λ•Œ κ·ΈλŸ¬λŠ”λ° μ•„ 근데 μ§„μ§œ μ’€ μ—¬
whisper-small   λ‹Œν…λ„ 맨 μ²˜μŒμ— λ‚˜μ˜¬ λ•Œ 그런 건데 μ•„ 근데 μ§„μ§œ μ’€ μ—¬..
CrisperWhisper  'Nintendo Manchera' 맨날 'Nintendo Manchera' λ‚˜μ˜¬ λ•Œ 그런 건데, μ•„, 근데 μ§„μ§œ μ’€,

What differs: our small/base keep the filler γ€Œμ–΄γ€ and the self-repeat γ€Œλ‚˜ λ‹Œν…λ„ ... λ‚˜ λ‹Œν…λ„γ€. whisper-small drops the γ€Œμ–΄γ€ and collapses the repeat. CrisperWhisper hallucinates the Korean word γ€Œλ‹Œν…λ„γ€ into English "Nintendo Manchera".

Example 3: a short turn that is mostly the filler γ€Œμ•„γ€

REF             μ•„ μ•„ μ•„ 톀이 λ†’μ•„μš”? μ•„
small (ours)    μ•„ μ•„ μ•„ 톀이 λ†’μ•„μš”? μ•„
base  (ours)    μ•„ μ•„ μ•„ 토읡 λ†’μ•„μš”? μ–΄.
tiny  (ours)    μ•„ μ•„ 토읡도 λ²Œμ–΄?
whisper-small   μ•„, μ•„, μ•„, μ•„, 톀이 λ†’μ•„μš”?
CrisperWhisper  -γ…Žγ…Ž. -γ…Žγ…Ž.

What differs: our small reproduces it exactly, all four γ€Œμ•„γ€ and γ€Œν†€μ΄ λ†’μ•„μš”γ€. whisper-small adds a fourth γ€Œμ•„γ€ and drops the trailing one. CrisperWhisper degenerates into "-γ…Žγ…Ž.". (On this very short clip our base and tiny mishear γ€Œν†€μ΄γ€ as γ€Œν† μ΅γ€, an honest small-model slip.)

Rare-word slips are real (γ€Œν†€μ΄γ€ to γ€Œν† μ΅γ€) and we do not hide them (our CER is 0.10 to 0.22). The point is which model preserves how the sentence was actually spoken, fillers and repeats included.

Usage

from transformers import WhisperForConditionalGeneration, WhisperProcessor
import librosa, torch

repo = "Sky-Kim/crisper-whisper2-base-finetuned-ko"
model = WhisperForConditionalGeneration.from_pretrained(repo).eval()
proc  = WhisperProcessor.from_pretrained(repo)

audio, _ = librosa.load("your_korean.wav", sr=16000, mono=True)
feat = proc.feature_extractor(audio, sampling_rate=16000, return_tensors="pt").input_features
ids  = model.generate(feat, language="ko", task="transcribe", max_new_tokens=200)
print(proc.tokenizer.decode(ids[0], skip_special_tokens=True))

Architecture: Whisper 512 dim, 6+6 layers with the CrisperWhisper 51,896-token vocabulary. Emits plain verbatim text (no markup tags). On-device target: Unity Sentis 2.6 (ONNX / Float16 export).

Training data

Dataset Role License
KsponSpeech (AI Hub / NIA) Korean verbatim ground-truth labels (about 236,600 clips) AI Hub terms (restricted)
Zeroth-Korean Korean read speech (about 6,000 clips) CC BY 4.0
CrisperWhisper2.0_small (nyralabs) small init weights, KD teacher for base/tiny, tokenizer Non-Commercial Research
Whisper (openai) tiny / base skeletons MIT

Labels are KsponSpeech ground truth (markup parsed to verbatim text, fillers kept). Korean only, encoder frozen. small: 15k steps. base and tiny: soft-KD from small, 12k steps. Eval is the held-out KsponSpeech valid split.

Licensing and redistribution

Effective license is the most restrictive of the components, so this is non-commercial / research use only.

  • CrisperWhisper2.0 (nyralabs), used as the small init, the KD teacher for base and tiny, and the tokenizer, is under the Nyra Health Non-Commercial Research License (commercial use requires a commercial license). These weights inherit that restriction.
  • KsponSpeech (AI Hub / NIA), the Korean training labels, is not an open license. Redistribution and commercial use are restricted by the AI Hub Terms of Use.
  • Zeroth-Korean is CC BY 4.0 (attribution). Whisper skeletons are MIT.

Commercial use requires a commercial license from Nyra Labs plus clearing the AI Hub KsponSpeech terms (or retraining with permissive components). License chain recorded to our knowledge (2026-08); not legal advice. Verify current terms before distributing. Full attribution in NOTICE.md.

Limitations

Korean only (no English). Verbatim output keeps disfluencies, which raises CER against cleaned references but is the point for verbatim use (meeting notes, speech analysis, subtitles as said). tiny is less accurate than base and small. Numbers are in-domain (KsponSpeech).

Downloads last month
-
Safetensors
Model size
72.6M params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Sky-Kim/crisper-whisper2-base-finetuned-ko

Finetuned
(2)
this model
Finetunes
1 model