91 MB
12 files
Updated about 1 month ago
Name
Size
benchmarks
.gitattributes1.52 kB
xet
README.md1.25 kB
xet
smart-turn-v3.0.onnx8.76 MB
xet
smart-turn-v3.1-cpu.onnx8.68 MB
xet
smart-turn-v3.1-gpu.onnx32.4 MB
xet
smart-turn-v3.2-cpu.onnx8.68 MB
xet
smart-turn-v3.2-gpu.onnx32.4 MB
xet
README.md

Smart Turn v3.x

Smart Turn is an open‑source semantic Voice Activity Detection (VAD) model that tells you whether a speaker has finished their turn by analysing the raw waveform, not the transcript.

Links

Model architecture

  • Backbone: Whisper Tiny encoder
  • Head: shallow linear classifier
  • Params: 8M
  • Checkpoint: 8 MB ONNX (int8 quantized), 32MB ONNX (unquantized)

How to use

Please see the blog post and GitHub repo for more information on using the model, either standalone or with Pipecat.

Thanks

Thank you to the following organisations for contributing audio datasets:

Total size
91 MB
Files
12
Last updated
Jul 13
Pre-warmed CDN
US EU US EU

Contributors