canary-1b-v2 β€” timestamps_asr_model β€” MLX q8

Quantized MLX port of the timestamps_asr_model bundled inside nvidia/canary-1b-v2. It is the dedicated CTC forced-alignment head paired with canary; vocabulary is byte-identical to canary's sentencepiece (16 384 pieces).

Architecture 24-layer Conformer encoder + 1Γ—1 CTC head
Params 626 M
Mel features 128
d_model 1024
CTC vocab 16 384 + blank
Subsampling 8Γ— (dw_striding)
Self-attention full rel-pos
Quantization bits=8, group_size=64 (MLX default)
Source extracted from canary-1b-v2.nemo

Files

File Size
config.json preprocessor + encoder + CTC config (with inline vocab)
model.safetensors q8 weights + bf16 scales/biases
tokenizer.model canary's sentencepiece (1:1 copy)

License

CC-BY-4.0 β€” inherited from nvidia/canary-1b-v2.

Downloads last month
14
Safetensors
Model size
0.2B params
Tensor type
BF16
Β·
U32
Β·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for TechHara/timestamps-asr-canary-ctc-mlx-q8

Finetuned
(10)
this model