Back to all tools
Free · runs in your browser

Audio trimmer

Drop in a recording, drag the handles across a real waveform and export just the part you want. This audio trimmer decodes the file inside your browser tab, so nothing is uploaded to anybody.

Your file is never uploaded.

It is read straight off your disk with the browser’s own file API and decoded by the browser’s own audio decoder, inside this tab. There is no server call in this page at any point — you can watch the network panel and see it. Export writes a WAV file, because a lossless container needs no encoder and adds no second generation of compression.

Drag an audio or video file here, or choose one below. MP3, WAV, FLAC and M4A open everywhere; Ogg and Opus need Chrome or Firefox.

Your file is decoded by your own browser. It is never uploaded, and nothing is written to disk unless you click download.

What an online audio trimmer usually does with your file

It uploads it. Of the sites ranking for this search, around four in ten POST your recording to a server, cut it there and hand back a download link. The page rarely says so — you infer it from the progress bar that appears before anything is visible. Your audio then sits on someone else’s disk for as long as their retention policy says, which for a free tool is usually a number nobody there can tell you.

For a podcast intro that does not matter. For a client call, a medical note, an unreleased track or an interview with a source it matters, and it is the reason this page exists. Everything here happens in the tab: the file is read off your disk, decoded by your browser, drawn to a canvas, cut in memory and written back out as a WAV you save yourself. You do not have to take that on trust — open the network panel, load a file, export a clip, and watch the request list stay empty.

How this audio cutter works online without uploading anything

Worth stating plainly, because “private” is a claim and this is a description.

  • Reading. File.arrayBuffer() pulls the bytes off your disk into the tab’s memory. No part of that involves the network.
  • Decoding. AudioContext.decodeAudioData hands those bytes to the same decoder your browser uses to play a video, and gets raw samples back. That is why MP3, WAV, FLAC and M4A/AAC open everywhere while Ogg and Opus need Chrome or Firefox — you are borrowing the browser’s codec list, not ours.
  • Drawing. Three minutes of audio is roughly eight million samples and your screen is eight hundred pixels wide, so the waveform is one minimum and maximum per column. That scan runs in a Web Worker, and the result is cached so a resize costs nothing.
  • Cutting. A trim is a slice of a typed array between two sample indices. What comes out is what went in, unless you asked for a fade, a downmix or a rate change.
  • Writing. The clip becomes 16-bit integers, interleaved, behind a hand-written 44-byte RIFF header. That is the whole WAV format, which is why this page carries no audio library.

Using it as an audio file cutter online, step by step

  • Drag the file onto the drop zone, or use the file picker.
  • Drag across the waveform to rough in a selection, then drag either handle. Or tab to a handle: an arrow key moves 100 milliseconds, shift and an arrow moves a second.
  • Press space to hear exactly what you selected before committing to it.
  • Type into the start and end fields when you know the timecode. They read and write mm:ss.mmm and accept plain seconds too.
  • Choose keep or delete. Delete is the one people forget exists, and it is the right tool for taking a cough out of the middle of a take.
  • Export. The WAV lands in your downloads folder.

Why you cannot trim an MP3 online without losing a generation

An MP3 is not a recording; it is a description of a recording that was already thrown away. To cut it, something must decode it back to samples. If the tool then writes another MP3, that second encoder makes its own decisions about what to discard, on audio that has already lost material once. One round trip is usually inaudible on speech and audible on cymbals; four are audible on everything.

There is a way to trim an MP3 with no loss — rewriting the container on frame boundaries without touching the audio data — but it only lands cuts every 26 milliseconds or so, and doing it correctly means parsing frame headers, LAME gapless tags and ID3 chunks. That is a real piece of software, not a page feature.

So this tool decodes once, cuts, and writes WAV. Nothing is compounded, and if you need an MP3 at the end you make it once, from the WAV, at a bitrate you chose rather than one a web page never disclosed.

Clicks at the cut, and what the fade control actually fixes

Cut a waveform at an arbitrary point and the sample value there is almost never zero — it might be at 0.6 of full scale. The audio before the join ends at 0.6, the audio after starts at 0, and that instantaneous step is, in frequency terms, energy smeared across the whole spectrum. You hear it as a click.

Editors normally solve this by snapping cuts to zero crossings. This tool uses a short linear fade, the same fix from the other direction: ramp the first and last few milliseconds from silence and there is no step left to reproduce. Ten milliseconds is the default because it kills the click and is far too short to perceive as a fade. Set it to zero when you cut into silence and want the samples untouched; raise it to 500 or 1000 for a fade you can hear.

Sample rate, bit depth, and what this audio trimmer’s WAV export contains

Sample rate is how many times per second the waveform was measured — 44,100 for CD audio, 48,000 for video — and it sets the highest frequency the file can hold, at half the rate. Bit depth is how finely each measurement is stored. This page writes 16-bit: about 96 dB of dynamic range, more than any speech recording needs, and half the size of 24-bit.

One wrinkle most tools hide: the Web Audio API decodes at the sample rate of your sound hardware, not the file’s. Load a 44.1 kHz MP3 on a machine whose output runs at 48 kHz and the browser resamples during decoding, before this page sees a single sample. So the tool reports the rate it actually got rather than the one written in the file. That is browser behaviour no in-page trimmer can avoid, and you should know about it rather than wonder later.

Why 16 kHz mono is the right preset for speech-to-text

Speech recognition models are overwhelmingly trained on 16 kHz mono. That rate holds everything up to 8 kHz, which covers the intelligible range of the human voice, consonants included, and discards the octave above it, where speech carries almost nothing. A second channel is redundancy when one person talks into one microphone. Hand an engine a 48 kHz stereo file and it downmixes and resamples anyway, usually less carefully than you would.

The payoff is size: ten minutes at 48 kHz stereo is about 110 MB, and about 19 MB at 16 kHz mono. When you downsample here, every input sample inside an output sample is averaged rather than discarded — a crude low-pass, but a real one, because naive decimation folds high frequencies back into the speech band as a metallic buzz. To clean up the text that comes back, use the transcript cleaner.

What this tool deliberately does not do

  • It exports WAV only. No MP3, no M4A, no Ogg. An encoder means a megabyte of JavaScript and a second lossy generation.
  • It is single-track. One file, one selection, one cut — no timeline, no crossfades, no joining two recordings.
  • There are no effects. No noise reduction, normalisation, compression or EQ. Fade, mono downmix and resample are the whole list, because each is arithmetic simple enough not to go wrong.
  • Large files are bounded by your browser, not by us. Decoding expands audio to 32-bit floating-point samples, so an hour of 44.1 kHz stereo is roughly 1.2 GB in memory whatever it weighs on disk. This page estimates that before decoding, warns above 50 MB and refuses above 200 MB — a refusal you can act on beats a tab that dies silently.
  • It does not transcribe. There is no speech recognition here at all; it cuts audio. For dictation in the browser, that is the voice typing tool.
  • It is not a video editor. A video file loads and its audio track decodes, but the export is audio only.

Where VoiceSnap Pro fits

This page is a free tool from the people building VoiceSnap Pro, a voice-to-text dictation app for macOS and Windows. Hold one keyboard shortcut, speak, and clean punctuated text appears in whatever field your cursor is already in. It punctuates and paragraphs as it transcribes, strips filler words, and saves every dictation to a searchable notes library. It is a one-time $39 purchase rather than a subscription — the same instinct as this page: pay once, or pay nothing, and do not rent the basics.

The connection to an audio trimmer is narrow but honest: if you are here because you recorded something and want the useful minute out of it, dictation is the thing that means you never make that recording at all. VoiceSnap Pro has not shipped yet — there is no download, only a waitlist. Join the waitlist for one email on release day. Meanwhile the microphone test tells you whether the input you are about to record with is the one you think it is.

Questions people ask

No. The file is read from your disk into the tab’s memory, decoded by your browser’s own decoder and cut in memory. This page makes no network request with your audio, and nothing touches your disk until you click export and the browser saves the WAV. Close the tab and every sample is gone. The network panel will confirm it.

Free, no sign-up, no watermark, no daily limit and no length cap beyond what your browser’s memory allows. The catch: the work happens on your computer, so a long file spends your RAM and CPU instead of a server’s.

Whatever your browser decodes: MP3, WAV, FLAC and M4A/AAC everywhere, plus Ogg Vorbis and Opus in Chrome and Firefox. Safari has neither, and no browser decodes WMA or AMR. If a decode fails, the message names the format and your browser rather than shrugging.

Because WAV is raw samples plus a 44-byte header, so it needs no encoder, and because re-encoding would add a second lossy generation to audio that has already been through one. Convert once at the end if you need a smaller file. The subtitle converter is the sibling tool for the caption files that travel with the same recording.

Yes, on current mobile Safari and Chrome, and the handles respond to touch. But phones have much less memory, so the practical size limit is lower, and precise cuts on a small screen are awkward — use the numeric fields there.

Select the part you want gone and switch to “delete the selection”. The two remaining pieces are joined and the fade is applied at the outer edges of the result. For a clean join at the splice itself, cut during a pause rather than mid-word.