Request to add ClipCannon Voice Clone Pipeline to TTS Arena

#114
by cabdru - opened

Hi TTS Arena team,

I'd like to request adding ClipCannon to the Arena for evaluation.

What it is:
A personalized voice cloning pipeline built on Qwen3-TTS-12Hz-1.7B-Base. Zero model modification -- all gains from pipeline engineering (Full ICL mode, best-of-N candidate selection, Resemble Enhance post-processing).

Benchmark results (independently verified):

  • 0.961 mean cross-encoder SECS (WavLMForXVector) across 10 novel sentences
  • 0.975 max -- inside the same-session human verification band (0.95-0.99)
  • +0.080 above Microsoft's "human parity" threshold (VALL-E 2: 0.881)
  • 0.000 WER (perfect word accuracy)

Links:

How to integrate:
ClipCannon runs locally on a single GPU (RTX 5090, ~4GB VRAM for Qwen3-TTS). I can provide an API endpoint or generate samples on request for Arena evaluation. Happy to work with whatever integration method you prefer.

The system is optimized for personalized voice cloning (known speaker with reference data), not zero-shot cloning of strangers. For Arena evaluation, I'd provide a reference recording and the system generates any text in that voice.

Demo video:
https://youtu.be/9re6jYR6GZg
Shows the original recording, audio stripped, then AI clone saying words I never said.

Let me know what you need from my side.

Chris Royse

Sign up or log in to comment