vconnx · kNN-VC (pure ONNX)

ONNX export of kNN-VC (Baas et al., Interspeech 2023, MIT) for the vconnx voice-conversion library: WavLM-Large layer-6 feature extractor + prematched HiFi-GAN vocoder; the kNN regression step is pure numpy at inference time. Any-to-any voice conversion at 16 kHz, CPU-only.

file role size
wavlm_layer6.onnx feature extractor (fp32) 387 MB
wavlm_layer6_q8.onnx feature extractor (int8) 98 MB
hifigan_knnvc.onnx vocoder (fp32) 63 MB
hifigan_knnvc_q8.onnx vocoder (int8) 25 MB

Export parity vs the torch reference (max abs): WavLM see report · HiFi-GAN see report (full reports in the repo).

Usage

from vconnx import VoiceCloner
out = VoiceCloner(engine="knnvc").clone_voice("source.wav", "reference.wav")

Provenance and upstream license: see PROVENANCE.md. Weights derive from bshall/knn-vc (MIT) and microsoft WavLM (MIT).

Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including TigreGotico/voiceclonnx-knn-vc