Speaker turns
Each stretch of speech is labeled and grouped so the dialogue structure survives export.
Speaker labels
Diarization answers one question well: who spoke when. VoxTextor turns raw audio into a speaker-separated transcript so you can follow the conversation instead of guessing at paragraph breaks.
Each stretch of speech is labeled and grouped so the dialogue structure survives export.
See when a new voice enters without scrubbing the waveform or second-guessing blank lines.
Keep who-spoke markers through TXT, SRT, VTT, and Markdown so editors can drop the file in as-is.
Separate host and guest for show notes, pull quotes, and social clips.
Track who committed to what without replaying the whole call.
Compare responses across participants without manual tagging marathons.
Diarization is not voice recognition. It separates speakers by acoustic pattern, not by name. Overlapping speech, heavy crosstalk, and identical voices can still confuse a model — yours included. We label the turns; you attach the identities.
For one-on-one interviews this is almost always enough. For chaotic town halls, plan a quick review pass.
Speaker diarization is included in Pro. Start with 30 free minutes every month to try the pipeline. MP3, M4A, WAV, MP4, MOV, WEBM.
Audio & video · any common format
Free without sign-in · up to 15 min · no credit card