In-browser automatic speech recognition with word-level timestamps and speaker segmentation
Audio or video files supported