AI Tools.

Search

voice activity detection models

3 models · ranked by HuggingFace downloads

segmentation-3.0

Pyannote segmentation-3.0 is a speaker segmentation model for detecting speaker changes, overlapping speech, and voice activity in audio. It produces frame-level predictions used as input to the full speaker diarization pipeline. The model can also run standalone for voice activity detection or overlapped speech detection without the full diarization stack.

5,859,290 ↓ · 1,577 ♡

segmentation

Pyannote segmentation (v1.x) is the earlier version of pyannote's speaker segmentation model for voice activity detection and speaker change detection, preceding the current segmentation-3.0. It is used within older pyannote speaker diarization pipelines. MIT licensed.

4,670,849 ↓ · 693 ♡

Namo-Turn-Detector-v1-Korean

Namo-Turn-Detector-v1-Korean is a Korean-language end-of-utterance detection model from VideoSDK Live, built on a DistilBERT backbone and exported to ONNX for low-latency inference. Turn detection determines when a speaker has finished speaking, enabling real-time voice assistant and transcription systems to trigger correctly. Korean-specific tuning addresses Korean prosody and speech patterns.

413,343 ↓ · 1 ♡