Convert WAV to Text

Upload an uncompressed WAV from a field recorder, interface, or DAW, review the text against playback, and export TXT or DOCX without leaving this page. Each file may run to 1GB and 60 minutes, and because PCM is uncompressed the size cap usually arrives first. Runs use your account credits.

My recordsCloud processing · up to 1 GB / 60 minutes. Final usage is verified by the server.

Sign in to create a saved task. If your file is unavailable after signing in, select it again.

Dashboard

New Transcription
0 min

How do you want to transcribe?

Upload audio or video
Estimated cost: 0 min

Free minutes are included. Upload a file or record audio to start.

Choose your file and settings before signing in; sign in to submit. Processing uploads audio to cloud services. Review the estimated credits and your current balance before starting. See plans and usage limits and our privacy policy.

How It Works

1. Upload Your WAV

Pick the .wav your recorder, audio interface, or DAW wrote. Set the spoken language or leave detection on, and turn on speaker labels when more than one person talks. The duration and credit estimate appear before anything is submitted, so check them against your own player first. A renamed file is still whatever it was: the bytes have to open as a real RIFF/WAVE stream, not merely carry a .wav extension.

2. Review the Result

Progress, processing state, and the finished transcript all stay on this page. Scrub the audio in the player while you fix names, figures, and punctuation — recorder and DAW takes carry room tone and handling noise that push recognition off far more often than a studio file does. Treat speaker labels as a first pass to verify, not a finished attribution. When a run fails, read the message before resubmitting; a rejected file usually needs a different export rather than another attempt.

3. Export Your Text

Save first, then take TXT for a plain file or DOCX to carry on editing in a word processor. Both exports read the saved transcript, so unsaved changes do not travel with them. Each run is filed under Recordings alongside work started elsewhere in the account, and its task link reopens the same result without re-uploading a WAV that may run to hundreds of megabytes.

WAV Is a Container, Not a Codec

A .wav file is a RIFF container holding a WAVE form. Recorders and DAWs almost always write linear PCM into it, and that is the case this page is written for. The same container can also carry ADPCM, µ-law, GSM, or even an MP3 payload. This page does not decode every WAV variant in your browser before upload: it accepts the .wav extension and leaves the final codec decision to the transcription server and its provider, so an unusual payload can still be refused after you submit.

Two channels are not two speakers. A stereo WAV is one mixed take spread across a left and a right channel, and the speaker labels here are inferred from analysing that mix, never from the channels themselves. A dual-mono field recording with one lavalier per side is treated the same way, because the channels are combined before recognition runs.

Uncompressed PCM preserves whatever reached the microphone, which is not the same as capturing it well. It cannot lift a voice out of a noisy room, tighten a distant mic, or undo clipping, and exporting a compressed file back out to WAV re-inflates the byte count without returning one discarded detail. What PCM does cost is size: minute for minute a WAV runs roughly ten times a typical MP3, which is why the byte cap, not the clock, is usually what ends a long session.

See an Example

The clip below is an original synthetic practice input, never a customer recording. It carries the same authored script as the MP3 practice file on our MP3 page, re-encoded to PCM, so it demonstrates the upload, review, and export path on a small file rather than the fidelity of an untouched recorder take. The text further down is that authored reference script, not a claimed model result; punctuation and recognition may differ.

Codec
Linear PCM, 16-bit (pcm_s16le)
Sample rate
48 kHz
Channels
Mono
Duration
14.05 seconds
File size
1,348,882 bytes (about 1.3 MiB)
Credits
0.5 (under 30 seconds)

Download practice WAV · Download reference text

Reference script and review exercise

Welcome to the Field Notes podcast. Today we are testing a small change: writing down one useful idea after every interview. Before publishing, replay the recording, check each guest's name, and keep the original audio with your notes.

Once the run finishes, compare “Field Notes” and the possessive in “guest's name” against the script above. Insert a paragraph break ahead of “Before publishing,” save, and download both TXT and DOCX; each file should carry that saved break. The point of the exercise is spotting differences and confirming exports, not scoring accuracy — a short, clean, synthetic clip says nothing about how a noisy field recording will transcribe.

Supported Files and Plan Limits

Uploads here are .wav only; other formats belong in the general workspace or on the format-specific pages. The cloud limits are 1GB per file and 60 minutes of audio, whichever arrives first. Do the arithmetic before a long session: an hour of 48 kHz 16-bit stereo is about 660 MiB and clears the cap comfortably, the same hour at 24-bit lands near 990 MiB and only just clears it, and 96 kHz stereo passes 1GB well before the hour is up. Split a long recording in your editor first — every part you submit becomes its own task.

Credits follow duration rather than file size: up to 30 seconds costs 0.5 credit, and past that one credit per started minute, so a 660 MiB WAV and a small MP3 of the same hour cost exactly the same. On submission the server re-checks eligibility, duration, size, and your balance. Optional AI actions bill separately. Check the current plans rather than assuming unlimited free processing.

Frequently Asked Questions

Can you transcribe a stereo WAV as two separate speakers?

No. Two channels are two halves of one mix, not two microphones this page can route independently. Labels are inferred from the combined audio, so people talking over each other still need a human pass. Where each person really was recorded to a separate channel, split them in an editor and submit the takes as separate tasks.

Does uncompressed WAV give a more accurate transcript than MP3?

Only when the WAV is the original capture. Microphone placement, room acoustics, and overlapping speech shape the result far more than the container does, and re-exporting a compressed file to WAV enlarges it without recovering anything the encoder discarded.

Why was my .wav file rejected?

Usually the bytes are not a readable WAVE stream: a renamed file, a truncated transfer, an unusual payload inside the RIFF container, or a recording past the 1GB or 60-minute limit. Re-export it as linear PCM from your recorder or editor and submit that file instead.

Is WAV transcription free, and does it run on my own machine?

Neither. The audio is uploaded and processed in the cloud against account credits, and whatever free allowance you have shows up through your account and plan. Choosing a file on its own starts nothing; you sign in and submit to begin a run.

Where can I find my transcript after leaving?

Under Recordings in the same account, or by reopening this page with the task link. A saved result stays with the task owner. Refreshing before you submit clears the local selection, because a browser cannot hand back file contents from a filename alone, so choose the WAV again.