ihateaudio

Voice Changer

Deeper, higher, robot, phone. Pick one and hear it.

Drop an audio file here

a voice note or phone recording works fine

MP3 · WAV · M4A · AAC · OGG · FLAC and moreYour file stays on your device. Always.

How to use it

  1. Drop a recording onto the page. A voice note, a take from your phone, anything your browser can decode. It stays on your device.
  2. Choose a preset. The line underneath says what it does to the sound and whether it changes the length of the recording.
  3. Add a manual pitch shift if the preset alone is not enough. It runs after the preset, so you can push Telephone lower or lift Radio slightly.
  4. Press Preview to render the chain and listen to it in the player. Pitch shifting is heavy, so it runs when you ask rather than on every change.
  5. Download the result. The filename records the preset you used, so a robot version saves as recording-robot.mp3.

What a voice is made of

Speech has two independent parts, and every effect on this page works on one or the other. The first is the source: the vocal folds open and close at some rate, producing a buzz. That rate is the pitch, typically around 85 to 180 Hz for adult men and 165 to 255 Hz for adult women. The second is the filter: the throat, mouth, tongue and nose form a set of resonant cavities that emphasise certain bands and suppress others. Those emphasised bands are the formants, and they are what turn a buzz into a recognisable vowel.

Formant positions depend mostly on the size and shape of the cavities, not on the note being sung. That is why you can recognise the same person speaking high or low, and it is why simply moving the pitch never quite produces a different person. It produces the same person transposed. Which is a perfectly good effect, but a different thing from a disguise.

Two families of effect

Deeper, Higher, Chipmunk and Giant all move pitch. Deeper and Higher use time-stretching, so the timing of the speech is untouched and only the register changes. Chipmunk and Giant change the playback rate instead, so pitch and speed move together and the recording gets shorter or longer. That coupling is what makes Chipmunk sound cartoonish rather than merely high: the delivery is rushed as well as raised, exactly as it would be on a tape running fast.

Telephone, Radio and Robot leave the pitch alone and change the filter instead. Telephone throws away everything outside roughly 300 Hz to 3.4 kHz, the bandwidth a phone line actually carried, which strips the chest resonance out of the voice and leaves the thin, close sound everyone recognises. Radio narrows the band less aggressively, lifts the presence region and then compresses hard, which is close to what a broadcast chain does to keep a voice at a constant level. Robot adds a very short delay with heavy feedback, producing a fixed comb of resonances that your ear attributes to a machine rather than a mouth.

Stacking them without making mud

The manual pitch control runs after the preset, which lets you combine families: Telephone plus four semitones down gives a voice on a bad line from someone considerably larger, and Radio plus two up sits a narrator slightly forward. Keep the manual shift small when you have already used a pitch-based preset, because two stretching passes compound each other's artefacts.

Beyond about seven semitones in either direction the result stops being a person and becomes an effect. That is often what you want, but decide it deliberately rather than arriving there by stacking. And do the ordinary cleanup first: trim the silence at the start, get the level right, and remove obvious background noise before you process. Every effect on this page treats the noise in a recording exactly as carefully as it treats the voice.

Questions

Will this make me unrecognisable?
Not reliably. A pitch shift changes the register of a voice but leaves the things people actually identify you by: rhythm, accent, phrasing, breathing, word choice, and the room you recorded in. Simple shifts can also be undone by shifting back. Treat these as effects for content, not as a way to stay anonymous when it matters.
Why do Chipmunk and Giant change the length of the recording?
Because those two work by changing the playback rate, the way a tape machine does. Chipmunk plays the file at 1.5x, so it finishes a third sooner and every frequency rises by about seven semitones. Giant plays it at 0.7x, so it runs longer and drops about six. Deeper and Higher use time-stretching instead and leave the length exactly as it was.
What is the difference between Deeper and Giant, then?
Deeper moves only the pitch: the words come out at the same speed, four semitones down, so it still sounds like someone talking normally in a lower register. Giant slows everything, so the delivery drags as well as descending. The difference between a low voice and a very large slow one. Pick Deeper for a natural result and Giant when you want the character.
How does the Telephone preset get that sound?
It band-limits the recording to roughly 300 Hz to 3.4 kHz, which is the actual bandwidth a traditional phone line carried, and adds a small lift around 1.8 kHz where speech intelligibility lives. Everything below and above is simply gone. That missing low end is why a phone voice sounds thin, and why the effect is so recognisable from a fraction of a second.
Why does the Robot preset sound metallic rather than just delayed?
It feeds the signal through a delay of about eighteen milliseconds with heavy feedback. At that spacing the repeats are far too close to hear as echoes; instead they interfere with the original and produce a comb filter with peaks every fifty-odd hertz. Your ear reads a fixed pattern of resonances as a mechanical resonator rather than a throat.
Can I use this on a live call or in a game?
No. This page processes a file you already have and gives you a new file back. Changing your voice inside a call needs a virtual audio device installed on your machine, which a web page cannot provide. Record first, process here, then send the result.
Does the source recording quality matter?
A great deal, particularly for Radio. That preset compresses hard, and compression raises quiet material. Which means room noise, air conditioning and hiss come up with the voice. Record close to the microphone in the quietest space you have. A clean recording survives heavy processing; a noisy one gets worse with every stage.