Voice Changer
Deeper, higher, robot, phone. Pick one and hear it.
a voice note or phone recording works fine
How to use it
- Drop a recording onto the page. A voice note, a take from your phone, anything your browser can decode. It stays on your device.
- Choose a preset. The line underneath says what it does to the sound and whether it changes the length of the recording.
- Add a manual pitch shift if the preset alone is not enough. It runs after the preset, so you can push Telephone lower or lift Radio slightly.
- Press Preview to render the chain and listen to it in the player. Pitch shifting is heavy, so it runs when you ask rather than on every change.
- Download the result. The filename records the preset you used, so a robot version saves as recording-robot.mp3.
What a voice is made of
Speech has two independent parts, and every effect on this page works on one or the other. The first is the source: the vocal folds open and close at some rate, producing a buzz. That rate is the pitch, typically around 85 to 180 Hz for adult men and 165 to 255 Hz for adult women. The second is the filter: the throat, mouth, tongue and nose form a set of resonant cavities that emphasise certain bands and suppress others. Those emphasised bands are the formants, and they are what turn a buzz into a recognisable vowel.
Formant positions depend mostly on the size and shape of the cavities, not on the note being sung. That is why you can recognise the same person speaking high or low, and it is why simply moving the pitch never quite produces a different person. It produces the same person transposed. Which is a perfectly good effect, but a different thing from a disguise.
Two families of effect
Deeper, Higher, Chipmunk and Giant all move pitch. Deeper and Higher use time-stretching, so the timing of the speech is untouched and only the register changes. Chipmunk and Giant change the playback rate instead, so pitch and speed move together and the recording gets shorter or longer. That coupling is what makes Chipmunk sound cartoonish rather than merely high: the delivery is rushed as well as raised, exactly as it would be on a tape running fast.
Telephone, Radio and Robot leave the pitch alone and change the filter instead. Telephone throws away everything outside roughly 300 Hz to 3.4 kHz, the bandwidth a phone line actually carried, which strips the chest resonance out of the voice and leaves the thin, close sound everyone recognises. Radio narrows the band less aggressively, lifts the presence region and then compresses hard, which is close to what a broadcast chain does to keep a voice at a constant level. Robot adds a very short delay with heavy feedback, producing a fixed comb of resonances that your ear attributes to a machine rather than a mouth.
Stacking them without making mud
The manual pitch control runs after the preset, which lets you combine families: Telephone plus four semitones down gives a voice on a bad line from someone considerably larger, and Radio plus two up sits a narrator slightly forward. Keep the manual shift small when you have already used a pitch-based preset, because two stretching passes compound each other's artefacts.
Beyond about seven semitones in either direction the result stops being a person and becomes an effect. That is often what you want, but decide it deliberately rather than arriving there by stacking. And do the ordinary cleanup first: trim the silence at the start, get the level right, and remove obvious background noise before you process. Every effect on this page treats the noise in a recording exactly as carefully as it treats the voice.