---
title: "File formats"
description: "MP3, WAV, M4A, FLAC, AAC, MP4, MOV, AVI and more upload directly, alongside TXT, DOCX and PDF, plus PNG and JPEG images read by OCR."
---

> Documentation Index
> Fetch the complete documentation index at: https://docs.speakai.co/llms.txt
> Use this file to discover all available pages before exploring further.

# File formats

Speak AI reads most common audio and video formats, plus documents and images. If your file is on this list, drag it in and it transcribes and analyzes like any other upload.

## Supported formats

### Audio

- **MP3**, the most common audio format
- **WAV**, uncompressed audio and the highest quality
- **M4A**, Apple audio
- **M4P**, protected Apple audio
- **AAC**, Advanced Audio Coding
- **FLAC**, lossless compressed audio
- **OGG**, open-source audio
- **WEBM**, web audio
- **AMR**, the format most phone voice recorders produce

### Video

- **MP4**, the most common video format
- **MOV**, Apple QuickTime video
- **M4V**, Apple video
- **AVI**, Windows video
- **WMV**, Windows Media Video
- **FLV**, Flash video

### Documents, images, and links

- **Documents**: TXT, DOCX, and PDF for text-based analysis
- **Images**: PNG and JPEG, read with OCR
- **URLs**: YouTube and Vimeo links, and direct media links

Spreadsheets take a different route. To analyze a CSV, use [CSV import](/help/uploads/csv-import/), which turns each row into its own file or note.

## File limits

- **Duration**: up to 10 hours per file. See [Duration limits](/help/uploads/duration/).
- **File size**: on the free plan, each file can be up to 2GB. Paid plans upload larger files. Compressed formats such as MP3 and M4A fit longer recordings inside the same size limit.
- **PDFs**: 50MB per file, whatever your plan.
- **Links**: 200MB per import, so a long video usually uploads faster as a file than as a URL.

## Get the best transcription

- **MP3 at 128kbps** suits most recordings: a small file with good speech clarity.
- **WAV** gives the highest accuracy but the files are much larger.
- If a file is too large, compress the audio bitrate. 64kbps still transcribes speech well.
- For video, Speak AI extracts the audio track automatically. Video quality does not affect transcription accuracy.

## Convert an unsupported file

MKV and WMA are the two formats people most often expect to find on the list above. Neither one gets past the file picker, so convert those before you upload.

Convert to MP3 for audio or MP4 for video with a free tool:

- [HandBrake](https://handbrake.fr/) for video
- [Audacity](https://www.audacityteam.org/) for audio
- An online converter such as CloudConvert or Zamzar

Having trouble with a specific format? Send us a message and we can help.

Evaluating Speak AI for a team? [Book a demo and see your own recordings analyzed](https://calendly.com/speak-ai/demo?utm_source=docs&utm_campaign=book-demo).

---
Related: [Uploads](/help/uploads/) · [CSV import](/help/uploads/csv-import/)

Source: https://docs.speakai.co/help/uploads/formats/index.mdx
