CAMPPlus zh-cn-common-200k (ONNX)

Speaker embedding model for diarization, exported from ModelScope damo/speech_campplus_sv_zh-cn_16k-common.

Specs

  • Architecture: CAM++ (Context-Aware Masking)
  • Training data: CN-Celeb + CN Common (~200k speakers)
  • Embedding dim: 192
  • Input: feats โ€” [batch, 80, time] (80-dim FBANK, 16 kHz, 10ms hop)
  • Output: embs โ€” [batch, 192]

Compatible with

pyannote-rs EmbeddingExtractor (same feats/embs tensor names as wespeaker_en_voxceleb_CAM++.onnx).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support