Back to all tools
Free · runs in your browser

Audio joiner

Put two or more recordings in order and join them into one file, with a silence gap or a crossfade at each join. Every clip is decoded and joined inside your browser tab and saved as an MP3 or a lossless WAV, so nothing is uploaded.

Drag two or more audio files here, or choose them below. They are decoded in this tab and never uploaded. MP3, WAV, FLAC and M4A open in every current browser; Ogg, Opus and WebM open in Chrome, Edge and Firefox, and older versions of Safari cannot open them.

Your file is decoded by your own browser. It is never uploaded, and nothing is written to disk unless you click download.

Merge audio files without handing them to a server

Many online audio joiners upload your files, stitch them together on a server and hand back a download link. For a playlist of sound effects that hardly matters. For a client call, a doctor’s voicemail or an interview recorded in three parts, it leaves a copy of the recording on a disk you know nothing about.

This page does the whole job in the tab. Each file is decoded by your browser’s own audio decoder, the clips are laid end to end in memory, and the result is written out as a file you save yourself. The MP3 encoder is LAME compiled to WebAssembly, running on your machine. You can check: open the network panel, join some files and watch. The only requests are for the page’s own code — including, the first time you export an MP3, the encoder — and none of them carries your audio.

How to join audio files here

  • Add the files. Drop them on the box or use the picker, and add more whenever you like. The list shows each clip’s length, channels, sample rate and level.
  • Put them in order. The arrow buttons move a clip up or down, from the keyboard too: focus stays with the clip you moved, so you can keep pressing.
  • Choose how the clips meet. A silence gap of 0 to 10 seconds, or a crossfade of 0 to 5 seconds — one or the other, since a join cannot both overlap and leave space.
  • Listen, then export. The joined length shows before anything is built. Preview plays the join in the page; export saves it as an MP3 at the bitrate you pick, or as a 16-bit WAV.

Joining voice notes, podcast segments and interview parts

Voice notes arrive in bursts: half an idea recorded while walking, the rest an hour later, a correction after that. The iPhone’s Voice Memos app saves AAC audio in M4A files by default, which every current browser decodes. WhatsApp voice notes are Opus audio in an Ogg container, usually with an .opus or .ogg extension, which Chrome, Edge and Firefox decode. Join them with a half-second gap and you have one file to send or archive instead of seven.

Podcast segments are often recorded separately: a cold open, the interview, a sponsor read, an outro over music. Join with a short gap or none, and crossfade only where music meets music. The start time the list shows for each clip is that segment’s timestamp in the finished episode — what chapter markers and show notes need.

Interviews in parts happen when a phone call interrupts the recorder or a battery dies. Join the parts with a gap of a second or so: the break really happened, and a small pause is more honest than pretending it did not. If you want the words rather than the audio, join first and transcribe once — the audio to text tool runs a speech recognition model on your own device, and one file is less work than three.

Gap or crossfade: choosing the join

A gap inserts digital silence — samples that are exactly zero — between each pair of clips. For speech it is almost always the right choice: a few tenths of a second reads as a breath, a second or two as a change of subject. At zero the clips butt straight together.

A butt join has one trap. A recording rarely ends at exactly zero; its last sample might sit at a fifth of full scale while the next clip starts somewhere else entirely, and that instantaneous jump is heard as a click. The “fade 5 ms at each join” option ramps each edge over five milliseconds — enough to remove the jump, far too short to hear as a fade. Untick it to leave the clip edges untouched.

A crossfade overlaps the end of one clip with the start of the next, fading one out as the other fades in. It suits music, ambience and room tone. On speech it lays the last words of one clip over the first words of the next, so keep it short there, or use a gap.

Why the crossfade uses an equal-power curve

In a linear crossfade each clip is at half amplitude halfway through, 6 dB down. When the two clips are different recordings their power adds rather than their amplitude, so the pair sums to half power: an audible 3 dB dip in the middle of every join.

An equal-power crossfade uses a quarter of a sine wave for the fade-in and a quarter of a cosine for the fade-out. At the midpoint each sits at about 0.707, 3 dB down, and because sine squared plus cosine squared is always one, the combined power stays level throughout. The trade-off appears only when both sides are the same signal, such as a recording crossfaded into a copy of itself, where equal power bumps up by 3 dB. Separate recordings are the uncorrelated case, which is why this joiner uses it.

The arithmetic of the joined length

Gaps add time and crossfades remove it. Three clips joined with one-second gaps come out two seconds longer than the clips laid end to end; with two-second crossfades, four seconds shorter. A crossfade can never be longer than the clip it runs over, and a clip with a join at both ends gives each at most half of itself, so a short clip in the middle can shorten its crossfades. When that happens the page says which join and by how much.

Mismatched sample rates and channel counts

Sample rate is how many times per second the waveform was measured: 44,100 for CD audio and most music, 48,000 for video and most recording hardware. Joining clips at different rates without converting would play one of them at the wrong speed and pitch.

In practice the browser settles most of this first. The Web Audio API decodes every file at the sample rate of the audio context, by default the rate of your sound output, so a 44.1 kHz MP3 and a 48 kHz WAV usually arrive at the same rate. Where they do not, each clip is resampled to the highest rate in the list: going up keeps everything the lower-rate clip had, while going down would discard the top of the higher-rate ones.

If any clip is stereo, the output is stereo and each mono clip is copied identically to both channels, which places it in the centre, where a mono recording belongs. No width is invented. Tick “mix down to mono” to average everything to one channel instead, which loses nothing for a single voice. It halves a WAV’s size; an MP3’s size is set by the bitrate you choose, so a mono MP3 is no smaller, but all of its bits go to the one channel. A clip with more than two channels contributes its front left and right.

Mismatched loudness, and why this page leaves it to you

The commonest fault in a joined file is not a click or a sample rate. It is one clip far quieter than its neighbours: the voice note recorded with the phone on the table, the interview part picked up by a laptop across the room. The listener reaches for the volume at every join.

This joiner does not change levels, because every automatic loudness fix makes judgement calls about peaks, noise and dynamics that are better made where you can hear the result. It does measure. The level beside each clip is the RMS of its non-silent parts: the audio is cut into 50-millisecond blocks and the quiet ones dropped using the gating approach of the ITU-R BS.1770 loudness standard, without its frequency weighting. That is enough to spot a problem, not a loudness meter. Any clip 6 dB or more below the loudest is flagged — half the amplitude, and clearly audible when one clip follows another. To fix it, raise that clip’s level in an editor first: Audacity is free and has both a Normalize and a Loudness Normalization effect.

Lossless WAV or MP3: which export to choose

WAV here is 16-bit PCM behind a standard 44-byte header, with no lossy codec in the way: if your sources were MP3 or M4A, the WAV holds their decoded samples at 16-bit precision, changed only by the gaps, fades and channel mixing you asked for. The cost is size, about 11 MB a minute at 48 kHz stereo. Choose it when the file is going into an editor, a transcription service or an archive.

MP3 is encoded once, at the bitrate you choose. If your sources were already MP3 or AAC, that is a second generation of lossy compression, applied to audio that has already lost material once. At 128 kbps and above that is rarely noticeable on speech; for music, 192 or 256 kbps leaves more headroom. At 128 kbps a minute is about 1 MB. Choose MP3 when the file is going to a person, a phone or a podcast host.

Why not just glue the MP3 files together?

You can concatenate MP3 files byte for byte — copy /b on Windows, cat on a Mac — and because an MP3 is a sequence of short frames, the result often plays. It also keeps every flaw of its parts. A LAME-encoded or VBR first file opens with a header stating its own length, which some players trust, so they show the wrong duration and seek to the wrong place. The second file’s tags land mid-stream, files at different sample rates do not belong in one stream at all, and each file’s encoder padding stays at the seam as a short gap. Decoding and encoding once avoids all of that for the cost of one generation, and WAV avoids even that.

What this joiner deliberately does not do

  • It does not trim. Clips are joined whole. To cut dead air or a false start, run the clip through the audio trimmer first.
  • It does not change volume or copy tags. No normalisation, compression or EQ, and titles, artwork and chapters from the source files are not carried over, because decoding keeps only the audio.
  • It is bounded by your browser’s memory. Every clip is decoded to 32-bit floating-point samples, so an hour of 48 kHz stereo holds about 1.3 GB whatever it weighed on disk, and the joined copy needs as much again. Single files over 200 MB are refused, and the page warns when a job would hold more than about 1 GB.
  • It is not a single-file converter. For one recording in the wrong format, the M4A to MP3, WAV to MP3 and WEBM to MP3 converters do that directly.

Where VoiceSnap Pro fits

This page is a free tool from the people building VoiceSnap Pro, a voice-to-text dictation app for macOS and Windows. Hold one keyboard shortcut, speak, and clean punctuated text appears in whatever field your cursor is already in — an email, a Slack message, a document. Every dictation is also saved to a searchable notes library. It is a one-time $39 purchase, not a subscription.

The connection to an audio joiner is the voice note. Plenty of people think out loud in fragments and end up here stitching them together; dictation turns each fragment into text as you say it, and the notes library keeps the pieces in one searchable place. VoiceSnap Pro has not shipped yet, so there is nothing to download today — join the waitlist for one email on release day. Until then, the audio to text tool will transcribe the file you just joined, and the microphone test will tell you whether your next recording is going to be one of the quiet ones.

Questions people ask

No. Each file is read from your disk and decoded by your browser's own audio decoder, the clips are joined in the tab's memory, and the MP3 or WAV is written in the tab as well. None of the page's network requests carries your audio: joining fetches only the page's own code, plus the MP3 encoder the first time you export an MP3.

Yes. Every clip is decoded to plain samples first, so the source formats do not need to match. MP3, WAV, FLAC and M4A open in every current browser; Ogg, Opus and WebM open in Chrome, Edge and Firefox. The joined file is written in one format, MP3 or WAV, whatever the inputs were.

Use a gap for speech. Half a second to a second of silence reads as a natural pause, and a gap of zero butts the clips together, with a 5 millisecond fade at each edge to stop clicks. Use a crossfade for music, ambience or room tone, where one sound should blend into the next. On speech a crossfade lays the end of one sentence over the start of the next, so keep it short.

Gaps add time and crossfades remove it. With three clips, two one-second gaps make the result two seconds longer than the clips laid end to end, and two two-second crossfades make it four seconds shorter. The joined length is shown before you export. A crossfade that would be longer than a short clip allows is shortened automatically, with a note saying which join was affected.

The browser decodes every file at the sample rate of your sound output, so clips usually arrive at one rate already; if they do not, each is resampled to the highest rate in the list. If any clip is stereo, the result is stereo and each mono clip is copied to both channels, which places it in the centre. Tick the mix down to mono option to make everything one channel instead.

Exporting as MP3 means the decoded audio is encoded again, which is a second generation of lossy compression. At 128 kbps and above that is rarely noticeable on speech; for music, 192 kbps or more leaves more headroom. Exporting as WAV avoids the second generation entirely, at several times the file size.

No. It joins exactly what you give it and does not change levels. It does measure each clip and flags any that is 6 dB or more quieter than the loudest, so you know before you export. To fix it, raise that clip's level in an audio editor such as Audacity, then add the corrected file here.