A single file can run up to 10 hours. Nothing special is asked of you between one minute and ten hours: a recording that lands a few minutes over what you planned uploads exactly like any other.
Past 4 hours, Speak AI splits the file into segments and transcribes them in order, so a very long recording takes longer to come back than its length alone suggests.
Set the language yourself on anything over 4 hours instead of leaving it on automatic detection. The segmented path needs to know the language up front, and most of the languages work with it, though a number of the regional variants do not. When Speak AI cannot segment a long file it fails it with a message asking you to split it into shorter clips and upload each separately.
Work through a long recording
A conference session or a deposition takes hours to listen back to. Transcribe it once, then read the summary instead:
- Upload the audio or video, for example an MP3 or MP4.
- Wait for the transcript. Even a multi-hour recording is typically ready in a few minutes.
- Ask AI Chat for what you need, such as “Summarize this entire transcript into 5 key bullet points”. Themes break the same content down by topic.
- Select a point in the answer to jump straight to that moment in the audio.
Run speaker identification before you summarize a recording with several voices. Once Speak AI knows who said what, you can ask targeted questions like “What did the judge say?”.
If the file is too large
Compress the audio rather than splitting the recording. Dropping the bitrate from 128kbps to 64kbps cuts the file size sharply with no noticeable loss in speech clarity, and speech transcribes just as well. For a recording longer than 10 hours, split it into shorter files and upload them separately.
See File formats for size limits by plan, and Upload errors if an upload fails outright.
Evaluating Speak AI for a team? Book a demo and see your own recordings analyzed.
Related: Uploads · CSV import