Send Audio on WhatsApp
16 MB is the cap. This works out how to get under it.
any size. It is read on your device, not uploaded
How to use it
- Drop the recording onto the page. Its length and true size appear immediately, along with whether it is already under WhatsApp's 16 MB.
- Check the line about speech or music. Everything underneath depends on which it is, and if the guess is wrong, one tap corrects it.
- Pick a route. The pre-selected one is whichever gets you the fewest files without sounding bad; the others are listed because sometimes a different trade is the right one.
- Press "Hear it first" to encode fifteen seconds at exactly those settings and play it back before you commit to the whole file.
- Download, then attach it in WhatsApp using Audio rather than Document. Otherwise it arrives as a file card instead of something you can tap and play.
Where the 16 megabytes comes from
WhatsApp has two entirely separate ceilings, and almost every argument about its limits comes from people standing on different ones. Anything sent asmedia, meaning a photo, a video, or an audio file attached with the audio button, is capped at 16 MB. Anything sent as a document can be up to 2 GB. Same app, same chat, same file, two different answers depending on which attachment menu you used.
That cap is old, and it has survived because it is not really about storage. WhatsApp delivers to phones on networks that are still metered and slow across much of the world, and 16 MB is roughly what can be pushed to a handset without the transfer becoming its own event. It is a delivery decision rather than a technical one, which is also why it has not risen the way the document limit did.
In practice the cap translates to about seventeen minutes of ordinary 128 kbps stereo audio. That is enough for a song and not remotely enough for a lecture, a sermon, an interview, or any of the other things people actually try to send. The arithmetic is unforgiving because a lossy file's size is simply its bitrate multiplied by its duration: halve the bitrate and you halve the bytes, with no cleverness available in between.
Why speech and music get treated so differently
The single biggest lever is what kind of recording it is. Speech occupies a narrow band of frequencies, contains long pauses, and carries almost nothing useful in the stereo image, which means an encoder can describe it with remarkably few bits. Music does none of those things. It is wide, dense, continuous, and often depends on stereo placement for instruments to stay distinguishable.
This is why folding a recording to mono is worth doing for a voice memo and damaging for a song. Mono does not make a lossy file any smaller, since bitrate times duration is still the whole story, but it lets the encoder spend every one of those bits on a single channel instead of splitting them across two. At low bitrates that is close to doubling the quality for free. On music, the same operation can make instruments vanish where the two channels cancel.
The practical effect is a large gap in what fits. Speech encoded as mono Opus stays clean down to around 32 kbps, which puts roughly an hour inside WhatsApp's 16 MB. Music needs something closer to 96 kbps before it stops sounding compressed, which puts the same limit at about twenty minutes. Splitting only becomes necessary past those points, and it is the honest answer when it does: three files that sound right are better than one that does not.