90+ languages · speaker labels · SRT / VTT · results in ~1 minute

Turn audio & video into accurate text

Drop in a podcast, meeting, or lecture — audio or video. We detect the language, label who said what, and hand back searchable, click-to-replay text with subtitle exports.

Sign in to transcribe your first file

Free accounts get 3 transcriptions every day. No credit card, no trial countdown.

Free: 3 files a day, up to 30 minutes each, resets midnight UTC · a 30-minute recording takes about a minute · no credit card

Works withMP3WAVM4AAACFLACOGGMP4MOVWEBMMKV

Free every day.

3 transcriptions a day on the house — and every file gets the full toolkit, not a stripped-down teaser.

Speaker labels

One checkbox, and every line comes back tagged by voice and time-coded.

Click-to-replay editor

Click any sentence to hear that exact moment — checking a quote takes seconds.

Subtitles in one click

Export SRT or VTT and drop it straight into YouTube, Premiere, or Final Cut.

Six export formats

TXT, DOCX, PDF, CSV, SRT, VTT — download several in one go.

A searchable archive

Find the minute a topic came up instead of re-listening to hours of tape.

Video straight in

Upload MP4, MOV, WEBM, or MKV as-is — no audio extraction first.

Used byPodcastersJournalistsStudentsResearchersProduct teamsLawyersDoctorsVideo editorsCourse creators

Powered by Whisper.

OpenAI's speech model runs your files — here's when it's accurate, and when it isn't.

Two engines, routed automatically

Whisper handles most files; a diarization engine takes over for multi-speaker or long audio. The file decides — you never pick a model.

Honest accuracy

Near-perfect on a decent mic. Expect slips on names, jargon, and crosstalk — no 100% promises. Test your own audio free.

Timestamps you can check

Model-native timings, typically within a second — click any line and verify it against the audio.

90+ languages, auto-detected

English, Spanish, Japanese, Chinese, and more. Mixed recordings resolve to the dominant language.

Private by default

Stored encrypted, never used to train models, no human ever listens — gone when you delete.

Flat pricing, not a taxi meter.

Per-minute services turn a single hour of audio into a $10–15 invoice. Here, one flat price covers everything you transcribe in a month — and the free plan never expires.

Free

$0/ forever

For trying it out — or light, everyday use.

  • 3 transcriptions per day, every day
  • Files up to 30 minutes / 50MB
  • Timestamps & speaker labels
  • TXT, PDF, DOCX, CSV, SRT, VTT export
  • 90+ languages
  • Standard processing queue
Most popular

Unlimited

$10/ month, billed yearly ($120)

For people who transcribe as part of their work.

  • Unlimited transcriptions
  • Speaker identification on every file
  • All AI tools — summaries, cleanup, translation
  • Priority processing
  • No per-file length limit — files up to 100MB

Questions people actually ask.

About accuracy, privacy, limits, and what “free” really means. The skeptical ones included.

How accurate is it, really?

On clean audio — a decent mic, one or two speakers, minimal background noise — expect near-perfect text with occasional slips on names and jargon. Accuracy drops with heavy crosstalk, strong accents, or noisy rooms. The best test is your own audio: the first transcriptions are free.

Why do I need to sign in with Google?

Sign-in is how we give every person a real daily free quota instead of a one-time teaser. One click with Google — no password to invent, no email verification dance, and we only receive your name, email, and avatar.

What does free actually include?

3 transcriptions every day, forever — each up to 30 minutes long. Not a 7-day trial, not a one-time credit. The counter resets at midnight UTC. If you need more or longer, the Unlimited plan removes both ceilings.

Which languages are supported?

Over 90, including English, Spanish, Portuguese, French, German, Japanese, Korean, and Chinese. Language is detected automatically — mixed-language recordings generally resolve to the dominant one.

Can it handle video files?

Yes. Upload MP4, MOV, WEBM, or MKV and we transcribe the audio track directly — no need to extract audio first.

How does speaker identification work?

Tick “Identify speakers” before uploading. The AI separates voices and labels each segment (Speaker 1, Speaker 2, …). It's most reliable with clear turn-taking — interviews, meetings, podcasts — and less so when people talk over each other.

Can I trust the timestamps?

They come straight from the speech model and are typically within a second of the audio. Click any line in the editor to hear the original recording at that moment and check for yourself. Subtitle exports keep the raw segment timings, so captions line up in your video editor.

Is my audio private?

Your files are private to your account, stored encrypted, never shared, and never used to train AI models. Transcription runs on automated systems only — no human listens to your audio. Delete a file and it's gone from our storage.

What can I export?

PDF and DOCX documents, plain text (TXT), spreadsheets (CSV), and subtitle files (SRT, VTT) — with optional section timestamps. You can select several formats and download them in one go.

Exactly which file formats can I upload?

Audio: MP3, WAV, M4A, AAC, FLAC, OGG, OPUS, WMA, AIFF. Video: MP4, MOV, WEBM, MKV, MPEG, AVI. If your recorder or editor produced it, it almost certainly works — and if a format ever fails, tell us and we'll add it.

Why is this so much cheaper than human transcription?

Because no human does the work. A professional transcriber charges $1–2 per audio minute and needs days; a speech model costs us cents per hour and needs about a minute. You trade a small accuracy margin for a hundred-fold price difference — and for anything you plan to publish, you can click each line to verify it against the audio.

How much can I actually transcribe?

Free: 3 files a day, every day — that's up to 90 recordings a month without paying a cent. Unlimited: no cap on the number of transcriptions; the only rule is one account per person.

Who is behind this?

Hi 👋 — Transcript.so is built by a tiny independent team, not a VC-fueled giant. We answer our own support email, ship improvements weekly, and keep the service fast and affordable because our costs are engineered to stay low.

The next recording doesn't have to become an evening of typing.

Your first transcript is about a minute away — free, no card.

Transcribe a file — free