How it works
When you transcribe audio or video with multiple people talking, Speak AI automatically detects and separates different speakers. Each speaker is labeled (Speaker 1, Speaker 2, etc.) and their dialogue is organized by paragraph.
Accuracy
Speaker identification works best when:
- Speakers have distinct voices
- People talk one at a time (minimal overlap)
- Audio quality is good with clear separation
- Each speaker uses a dedicated microphone
In noisy environments or recordings with lots of crosstalk, speaker detection may occasionally merge or split speakers incorrectly.
Renaming speakers
After transcription, you can rename speakers to their real names:
- Click on any speaker label in the transcript
- Type the person’s name
- Press Enter
The name applies to every instance of that speaker throughout the transcript. You can also use AI Chat: “Change Speaker 1 to John Smith”.
For more details on managing speakers, see our speaker editing guide.
Speaker analytics
Once speakers are identified, Speak AI tracks:
- Speaking time per person
- Word count per speaker
- Words per minute (speaking pace)
- Percentage of conversation
These analytics are visible on the media detail page and can be analyzed across multiple files on the Explore page.
Overview
When Speak AI transcribes your audio or video, it automatically identifies different speakers and labels them (Speaker 1, Speaker 2, and so on). You can rename these to real names so your transcript is clear and useful.
Renaming a speaker
- Open your transcribed file
- Click any speaker label in the transcript (for example, “Speaker 1”)
- Type the person’s name
- Press Enter to confirm
The name applies to every instance of that speaker throughout the transcript. You can rename speakers both in the read-only view and while editing the transcript.
Renaming multiple speakers
- Click any speaker label to open the speaker editor
- Update each speaker’s name
- Press Enter after each name
Merging two speakers into one
If the transcription split one person into two speakers, rename one of them to exactly match the other’s name. Speak combines them into a single speaker and confirms with “Speakers merged.”
Resetting all speakers
While editing a transcript, use the Reset Speakers button to set every speaker back to the default labels. This is the quickest way to start over when the labels are mixed up, and then relabel them with the correct names.
Using AI Chat to rename speakers
You can also use AI Chat to rename speakers with natural language:
- “Change Speaker 1 to John Smith”
- “Rename Speaker A to Sarah and Speaker B to Mike”
- “The interviewer is Jane Doe”
Tips for better speaker identification
- Good audio quality helps: Clear audio with minimal crosstalk makes it easier to separate speakers
- Dedicated microphones: When each person has their own microphone, speaker detection is most accurate
- Name speakers early: Renaming speakers right after transcription makes AI Chat more useful (“What did John say about the budget?”)
Troubleshooting
- Speakers labeled incorrectly? Rename them, merge two labels by giving them the same name, or use Reset Speakers while editing to start fresh.
- Too many speakers detected? Background noise or audio quality issues can add extra speaker labels. Rename or reset the extras.
How it works
Speak AI automatically detects different speakers in your recordings, but the level of detail depends on how the recording was captured.
Virtual meeting recordings (Meeting Assistant)
When the Speak AI Meeting Assistant joins your Zoom, Google Meet, Teams, or Webex call:
- Speakers are automatically identified by their meeting participant names
- If calendar sync is enabled, names from the calendar invite are used
- Each speaker gets their own label from the start
- Speaker analytics (word count, speaking time, pace) are calculated per person
Uploaded in-person recordings
When you upload a recording from a phone, handheld recorder, or other device:
- Speakers are detected by voice patterns and labeled as Speaker 0, Speaker 1, Speaker 2, etc.
- The system separates speakers based on voice differences, but cannot identify names automatically
- You need to rename speakers manually after transcription
Renaming speakers
For uploaded recordings, rename speakers right after transcription:
- Open the transcribed file
- Click on any speaker label (e.g., “Speaker 0”)
- Type the person’s name and press Enter
The name applies throughout the entire transcript. You can also use AI Chat: “Change Speaker 0 to Dr. Smith”.
Tips for better speaker detection in uploaded recordings
- Clear audio helps: Minimize background noise and crosstalk
- Separate microphones: If possible, use individual microphones for each speaker
- Central placement: Place the recording device in the center of the table
- Speak one at a time: Overlapping speech is the hardest scenario for speaker detection
Once speakers are named, AI Chat becomes much more powerful. You can ask “What did Dr. Smith say about the treatment plan?” and get speaker-specific answers.
Evaluating Speak AI for a team? Book a demo and see your own recordings analyzed.
Related: Transcription · Accuracy