How it works
- Speaker embeddings extracted per voiced window
- Online clustering with minimum-cluster-duration heuristic
- Label stability via Viterbi smoothing
- Optional enrollment for known clinicians
Documentation Index
Fetch the complete documentation index at: /llms.txt
Use this file to discover all available pages before exploring further.
Speaker separation for conversational STT
POST /api/providers/{id}/voice-enroll
Content-Type: audio/wav