Overview
Transcribe audio or video and label who said what. Send a media_url (optionally language and num_speakers) and it runs Whisper v3 speech-to-text with speaker diarization, returning an utterances array grouped by speaker with start/end timestamps and text, plus per-speaker stats (utterance count, seconds spoken, word count) and total duration_seconds. Files up to 60 minutes are supported. Use it as a speaker diarization API, who-said-what transcription tool, multi-speaker transcript generator, or
Health
x402 Payment Validation
Recent Health Checks
| Time | Status | HTTP | Latency | Error |
|---|---|---|---|---|
| 2026-08-29 22:34:46 | healthy | 402 | 433ms | |
| 2026-08-29 12:28:14 | healthy | 402 | 629ms | |
| 2026-08-29 03:35:02 | healthy | 402 | 16ms |