Browser-local audio

Join Audio Files Online

Arrange 2–10 MP3, WAV, or M4A files in the order you want, preview the sequence, and join them into one continuous MP3. Clips are normalized to 44.1 kHz stereo. Joining runs in this browser and does not use transcription credits.

Sequential join → one MP3
2–10 clips

Join 2–10 MP3, WAV, or M4A clips. Up to 100 MB per file, 250 MB and 30 minutes combined.

No clips yet. Choose two or more audio files to begin.

Clips are resampled to 44.1 kHz stereo, then placed one after another. There is no overlay and no crossfade.

Join audio in this browser without transcription credits. Sign in only if you want the arrangement in My records.

Joining is not mixing: clips play one after another with no overlay or crossfade. Review the privacy policy for how browser-local tools are described, and plans and usage limits for the billed transcription tools this page does not use.

How to join audio files online

Stay on this page from the first clip to one downloadable MP3. Nothing here sends your audio through speech recognition.

  1. 1

    Add 2–10 audio clips

    Choose MP3, WAV, or M4A files. The browser reads each duration as you add it, so a mixed-format set is normal rather than an error.

  2. 2

    Arrange the sequence

    Move any clip up or down, or remove one. The numbered list is the exact playback order; there is no automatic sorting and no hidden shuffle.

  3. 3

    Join and download one MP3

    Every clip is resampled to 44.1 kHz stereo and placed one after another, then encoded to a single MP3 you can preview before saving.

Public sample

A real three-clip join you can reproduce

These labeled files are original practice clips, not a customer upload and not an accuracy score.

The sample set mixes all three accepted formats on purpose: an MP3 recorded at 44.1 kHz mono, a WAV at 48 kHz mono, and an M4A at 22.05 kHz stereo. Joining them proves the tool is normalizing sample rate and channels instead of refusing a mixed set.

The failure sample is a real decoder failure, not just a renamed text file: corrupt-outro.m4a keeps a readable header and duration but has damaged audio frames. The join is configured to abort on the first decode error, so it stops instead of skipping the damaged frames and shipping a silently short MP3.

Download the three clips, add them in order, and join. The expected result is one 44.1 kHz stereo MP3 that plays the first, second, and third sentence in sequence. The committed expected file lets you compare duration and channel layout yourself.

Sample inputs

Expected output

  • expected-joined.mp3 — 44.1 kHz stereo MP3
  • Three sentences in source order, no overlap and no crossfade.
  • Duration within about 100 ms of the sum of the three inputs, allowing for MP3 frame padding.

Generated with the same FFmpeg graph this page runs locally. It is a reproducibility aid, not a benchmark.

Arrange clips in sequence

A joiner is an order editor first. The list you see is the file you get.

Joining is positional. The first clip plays first and the last clip plays last, with nothing inserted between them. Use the up and down controls on each row to change that order before you join. The controls are ordinary buttons, so they work with a keyboard and a screen reader, not only with a mouse drag.

The running total updates as you add, remove, and reorder clips. Durations come from the browser's own media metadata reader. If a file cannot be read, its duration shows as unavailable and the join step re-probes it before encoding, so a corrupt file is caught rather than silently skipped.

The order is not remembered from a previous session unless you save the arrangement while signed in. A saved task restores the clip names, order, and output settings; it cannot restore the original File objects after the browser tab is closed, so you choose the same files again to rebuild the MP3.

Combine mixed audio formats

MP3, WAV, and M4A can sit in the same sequence because every clip is converted to one common shape first.

A typical project mixes a phone voice memo, a recorder WAV, and an exported MP3. Those files rarely share a sample rate or channel count. Before concatenating, the tool forces every clip to 44.1 kHz stereo with one sample format, then appends them in order. That is what makes an MP3 and a 22.05 kHz M4A line up without a pitch shift.

Accepted input is MP3, WAV, and M4A. Video files, AC3, OGG, FLAC, and documents are rejected. The per-file limit is 100 MB, the combined limit is 250 MB, and the combined duration limit is 30 minutes. The browser checks these limits when you add clips, and the local execution step rechecks size, decodability, and duration before it encodes anything.

The output is a new MP3 at 96, 128, 192, 256, or 320 kbps. Re-encoding an already-compressed MP3 adds one more generation of loss; it will not recover detail the source already lost. If fidelity matters more than a single portable file, keep the originals.

Joining versus mixing

These are different jobs, and this page only does one of them.

Joining, also called concatenation, plays clips one after another on a single timeline. It is the right operation for an interview cut into segments, a set of voice memos in date order, or a lesson split across several recordings. The result is longer than the longest input and every source is audible in full.

Mixing, also called overlay or multi-track, plays clips at the same time and blends them: background music under narration, two speakers in a duet, or layered sound design. That requires per-clip gain, pan, and timing decisions, and it usually needs a real multi-track editor.

This tool does not mix, and it does not crossfade. There is no overlap between clips and no volume balancing between them. If a clip is much louder than the next, that difference carries into the output. Use a desktop audio editor or a dedicated mixing tool when the deliverable is a blend rather than a sequence.

WhisperWeb vs desktop editors vs cloud audio joiners

Choose the workflow that matches the job. Public product pages checked 19 September 2026. Plan details change.

DecisionWhisperWebDesktop audio editorCloud audio joiner
Where the audio goesWhisperWebAvailable: Stays in your browserClips are processed locally with a verified WebAssembly runtime. The source audio is not uploaded.Desktop audio editorAvailable: Stays on your machineInstall an application and manage your own files, storage, and updates.Cloud audio joinerLimited: Uploaded to a hostHosted joiners usually require an upload first and may keep files on their servers.
Joining vs mixingWhisperWebAvailable: Sequential join onlyOrdered concatenation with sample-rate and channel normalization. No overlay and no crossfade.Desktop audio editorAvailable: Multi-track mixing availableOverlay, gain, pan, crossfade, and effects are built into professional editors.Cloud audio joinerLimited: Often sequencing plus paid upsellsMany hosted tools mix conversion, editing, and transcription behind separate plans.
Mixed formats and limitsWhisperWebAvailable: MP3, WAV, and M4A together2–10 clips, 100 MB per file, 250 MB and 30 minutes combined, output capped at 64 MB.Desktop audio editorAvailable: Broad codec supportLarge projects are limited mainly by disk space and machine performance.Cloud audio joinerLimited: Varies by vendorUpload caps, retention windows, and format support differ per product and plan.
Credits and transcriptionWhisperWebAvailable: No transcription creditsThis page never calls a speech model. It does not deduct minutes and does not create a transcript.Desktop audio editorAvailable: No cloud minutesOffline editing does not consume a hosted service's minutes.Cloud audio joinerNot available: Sometimes bundled with paid ASRSome hosted suites route audio through paid speech-to-text even when you only want one file.

WhisperWeb details match the joiner on this page. Desktop and cloud columns describe common approaches seen on public editor and joiner pages on 19 September 2026, not a paid test of every vendor. See our plans and usage limits for billed speech tools, which this page does not use.

Use WhisperWeb when you have two to ten ordinary audio clips and want one MP3 without an upload, an account, or a transcription charge. Use a desktop editor when you need to mix tracks, cut precisely, or repair audio. Use a cloud joiner only when you cannot run something locally and you accept the host's upload, retention, and pricing terms.

One continuous file for the next step

The download is an ordinary MP3. Open it in any player, editor, or notes app.

Several audio waveforms being placed end to end into one continuous track on a studio timeline

From several takes to one timeline

Podcasters, students, and interviewers often record in short segments and need a single file for a player, a transcript, or an archive. Reordering the clips before joining is the whole edit. Nothing is uploaded, and the joined MP3 is ready to hand off in one download.

Generated use-case illustration, not an actual customer result. The runnable sample above is separate and downloadable.

Limits, privacy, and cost

Accepted input is 2–10 MP3, WAV, or M4A files. Each file is capped at 100 MB, the combined input at 250 MB, and the combined duration at 30 minutes. The generated MP3 is capped at 64 MB for browser memory safety. The front end shows these limits and the local execution step enforces them again before encoding.

Joining runs on your device. The selected audio is not sent to a speech model and does not deduct transcription minutes. The page downloads the verified WebAssembly encoder after you start a join; that engine download is separate from your media, and your files are not part of it.

A signed-in save stores the ordered clip names, sizes, durations, and output settings under your account so the arrangement can be reopened from My records. It does not store the audio bytes, because they never leave the browser. After a tab closes, choose the same files again to rebuild the MP3. Tasks are owner-only and are not placed in a public index.

This page does not log audio content or sign-in secrets. Only process audio you have the rights to use. A join does not improve a bad recording, and re-encoding to MP3 adds one lossy generation.

Frequently asked questions

Can I mix tracks at the same time?

No. This page joins clips in sequence; it does not overlay them. Each clip plays entirely, one after another, with no crossfade and no per-clip volume control. For background music under a voice, or two speakers at once, use a multi-track audio editor instead.

Can I reorder clips?

Yes. Every row has move-up and move-down buttons, and the numbered list is the exact output order. The controls work with a keyboard and a screen reader. Removing a clip also removes it from the sequence before encoding.

What happens when one file is invalid?

The whole join stops and no partial file is produced. A wrong format is rejected when you add it; a corrupt or silent file is caught when the local runtime probes each clip. Replace or remove the failing clip, then join again. You never get an MP3 that quietly skipped a clip.

Are my audio files uploaded?

No. The join runs in your browser with a verified WebAssembly encoder. The audio bytes stay on your device; only an optional signed-in save stores the clip arrangement, not the audio itself.

Does joining use transcription credits?

No. This page does not call a speech model and does not deduct transcription minutes. It starts no transcript. If you also need written text, use the Audio Transcript Studio with the joined file.