Transcription with Speaker Identification

Transcribe audio with speaker labels and timestamps. Upload or record below, review each voice, and export here. Labels group speech; they do not verify identity.

Speaker detection starts on with auto count. Open settings before you submit.
My recordsCloud processing · up to 1 GB / 60 minutes. Final usage is verified by the server.
Upload audio or video

Available with transcription credits; sign in to submit. Only upload or record voices you have permission to process. Speaker names are labels you review, not verified identities.

Estimated cost: 0 min

Free minutes are included. Upload a file or record audio to start.

Choose your file and settings before signing in; sign in to submit. Processing uploads audio to cloud services. You start with 5 free minutes after verification, then account credits. This is not unlimited free local transcription. Review the estimate and your balance before starting. See plans and usage limits and our privacy policy.

Choose your review workflow · checked September 10, 2026

WhisperWeb vs Otter vs Happy Scribe

Choose by what happens after voices are separated. For a recording you already have, WhisperWeb keeps upload, speaker review, and download on this page. A recurring meeting workflow or a different editing system may suit another service better.

Speaker labeling workflows, not accuracy rankings
ProductSpeaker workflowConsider it when
WhisperWebGroups voices within a recording. You review and save label names; supported exports use those revisions.You want a file-based review with account credits and one saved task.
OtterIts documentation describes learning from speaker tags to recognize speakers in future conversations.Recurring conversations and cross-conversation speaker tagging matter to you.
Happy ScribeIts editor supports adding, renaming, and unassigning speakers. Export choices depend on the plan.You prefer its transcript editor and export settings.

Sources: Otter speaker identification and Happy Scribe transcript editor. This compares documented workflows, not measured accuracy. Check each provider’s current plan before buying.

Who spoke when

Speaker Labels Are Not Identity Verification

Speaker diarization transcription divides speech into turns and groups turns that appear to come from the same voice. “Speaker 1” means a group inside this recording. It does not establish a legal name, authenticate a caller, or compare a voiceprint against a verified person.

Listen before changing a label to “Interviewer” or a person’s name. Renaming a group changes its displayed name; it does not repair a passage assigned to the wrong group. Check short replies and interruptions individually before quoting anyone.

Reviewed transcript printout with three timestamped speaker turns beside a recorder
Illustrative commercial transcript layout, AI-generated. Not a customer file or measured model output.

How It Works

  1. 01 / INPUT

    Upload Your Recording

    Choose a file or record here with microphone permission. Keep speaker detection enabled. Open settings to use automatic speaker count, or supply a count if you know it. A count is a hint, not a promise that every voice will be separated. Review the credit estimate and sign in to submit.

  2. 02 / REVIEW

    Check Each Change of Voice

    Processing and results stay on this URL. Replay timestamped segments, compare repeated appearances of each speaker, and save label changes in the speaker panel. Refreshing a task URL restores its server state. If labels are missing, the task fails visibly instead of accepting plain text.

  3. 03 / DELIVER

    Export the Saved Version

    Download DOCX, TXT, JSON, SRT, or VTT. Saved speaker names carry into supported transcript and caption exports. Review subtitle cues separately after changing prose: editing a paragraph does not automatically retime captions. Recordings keeps the task and its source page together.

See an Example: Three Voices, One Decision

Use our existing synthetic practice recording to check speaker boundaries. The reference below is the authored script, not a transcription-provider output or an accuracy claim. Upload the clip in the tool above, then compare the returned groups against the audio. Names in the script are roles to verify manually.

00:00Speaker 1 / Alex

The Harbor Line launch moves to Thursday. Does anyone object?

00:03Speaker 2 / Priya

Engineering can ship the checklist by Wednesday. I did not name a reviewer.

00:08Speaker 3 / Sam

Ops will reprint the floor signs. No date was given.

00:11Speaker 1 / Alex

Then Thursday is the decision. We still need a reviewer for Priya’s checklist.

For a two-voice interruption check, download the podcast practice clip. Its final beat overlaps voices. A useful review asks whether a short “yes” was lost, whether one voice became two groups, and whether a shared timestamp needs manual correction. Neither clip proves performance on your recording.

Supported Files and Plan Limits

The upload control lists accepted formats. The shared service currently limits each file to 1 GB and 60 minutes. Audio is uploaded for cloud processing; an account and sufficient transcription credits are required. Speaker detection uses the existing transcription route and credit rules. Check plans and usage limits before submitting. Failed tasks follow the existing refund rules; a retry starts only when you choose it.

Use this for research interviews, panel discussions, and recorded editorial conversations where attribution needs a human review. A closer microphone and turn-taking help more than renaming a noisy file. Results and downloads require the task owner’s signed-in session.

Frequently Asked Questions

Does the tool know each speaker’s real name?

No. Automatic labels distinguish voices within the file. You can add names after listening, but that is your annotation, not identity verification or voice authentication.

What happens when people talk at the same time?

Words may be omitted or assigned to the wrong group. Very short turns can also merge into a neighboring speaker. Replay the passage and preserve uncertainty instead of treating the label as proof.

What if speaker detection is unavailable?

Read the displayed error. This page requires detection to submit and rejects a provider result with no speaker-labeled speech. An account, credit, permission, or provider failure does not silently switch your request to plain transcription.

Can I keep editing and download later?

Yes. Save changes before exporting and reopen the task from Recordings or its task URL. Selected files alone are not saved server tasks; if the browser cannot restore an unsigned draft, it will ask you to select the file again.