What the readout shows. The big letter is the nearest musical note, and the number is your fundamental frequency — how many times per second your vocal folds are closing. Higher number, higher pitch.
The cents value shows how far off the note you are. A hundred cents is one semitone; zero means you're exactly on the note.
The thin bar underneath is confidence — how sure the detector is. It drops during breaths, whispers, and background noise, and that's normal.
Pitch over time
Standby
-- Hz
Elapsed 0.0sTarget 180–220 Hz-- Hz
What the graph reads. Each point is one pitch measurement, about twenty-one per second, scrolling right to left — the right edge is now.
The orange line is your pitch. Gaps are unvoiced sounds: breaths, pauses, and consonants like s, f, t.
The shaded band with dashed edges is your target range. Time spent inside it is what "in target" counts.
Flat lines read as robotic. A line that moves inside the band is what natural speech looks like.
The second icon swaps to a waveform, which shows loudness rather than pitch — useful for checking you aren't clipping or too quiet.
The half-filled circle turns on greyscale shading behind the trace.
100%
SpaceRecordNNextRResetSSave
Spectrogram
F1 --F2 --F3 --F4 --0–5 kHz · time →
What the waterfall shows. Frequency runs bottom to top (0–5 kHz), time scrolls right to left, and brightness is how much energy sits at that frequency. It is the closest thing to seeing your voice.
The horizontal bands are formants — resonances of your vocal tract, not your pitch. They are what makes a voice sound bright and forward or dark and chesty.
F1 and F2 do most of the perceptual work. Raising them generally means a smaller, more forward oral space. F3 and F4 shift far less and are shown for completeness.
Formants move as you change vowels, so read the trend across a whole sentence, not any single instant.
The dotted line is your pitch (F0). Pitch and formants are independent — you can move one without the other, and training both is the point.
These are estimates from linear predictive coding. A quiet room and a decent mic change the numbers noticeably.
Progress
Sessions0
Latest--
Change--
Per session--
Best--
Save a couple of sessions and this will tell you, in words, what the numbers are doing.
The only chart that shows whether this is working. Each point is one saved session; the newest is highlighted. Switch metric with the buttons above.
Median pitch with your target band behind it. In target is the share of voiced time inside that band. Variation is pitch movement in semitones — a high pitch delivered flatly still reads as robotic. Resonance is the composite tract-size score.
The dashed line is a least-squares trend, and “per session” is its slope — the honest measure of movement, since it uses every point rather than just the first and last.
Faint columns show how long each session ran. A big jump off a thirty-second session is mostly noise.
Voice change is measured in months, and it wobbles with illness, tiredness and time of day. Read the drift, not the day-to-day.
CSV exports every stored number for your own analysis, or to hand to an SLP. Only summary values are kept — never audio.
Target range
to
Notes or hertz — D3, F#3, Bb2 and 147 all work.
Notes or hertz. If a coach has given you a range in notes — "sit between D3 and F3" — type it straight into the custom boxes. Sharps and flats both work, and you can mix units: D3 to 180 is a valid band. The label shows the band both ways so you can hand the numbers back in whichever form is wanted.
What these presets do. They move the shaded band on the graph and set what counts as "in target". They change nothing about your voice — they're just where you're aiming today. They are labelled in hertz rather than by gender, because the same numbers serve anyone working in any direction, and a label telling you which band is "yours" would be wrong more often than right.
Typical speaking ranges overlap far more than these labels suggest, in every direction. Treat them as starting points, not categories.
Pick the band you can sit in comfortably for a whole sentence. If you can only hit it by pushing, drop one preset.
Pitch alone rarely decides how a voice is heard — resonance does a lot of the work.
Resonance
Darker · largerBrighter · smaller
Ticks = adult population meansSession avg --
No signal
F1 · openness-- Hz
F2 · brightness-- Hz
F3-- Hz
F4 no reference-- Hz
Vowel check
Everything here rests on the formant tracker being right about your setup. This holds it to two vowels whose formants are known, and tells you whether the numbers can be trusted.
The two ticks are the adult population reference means, lower and higher. Formants move a lot with each vowel — “ee” sits far from “ah” no matter how you speak — so read the drift across a whole sentence, not any instant.
Read the average, not the needle. The geometric mean swings hugely with the vowel you happen to be on — for one speaker "had" and "who’d" differ by more than the gap between two typical speakers. So a single instant tells you almost nothing, and the session average across a whole sentence tells you a great deal. The scale is anchored so that the two adult population references sit at 30% and 70%, which puts ordinary speech mid-scale with room to move either way.
How much is enough? There is no single right number, because formants shift with every vowel. The gauge shows where your overall resonance sits between the two adult population reference means, so you can read direction of travel rather than guess at raw hertz.
Pitch carrying it alone: the needle parked at one end across a whole sentence while pitch does all the work. That usually sounds strained.
A useful place to be: drifting steadily toward whichever end you are working on, while speech still feels effortless. For most people this shifts over weeks, not days.
Too much: two signals. If vowels start colliding — “bead” and “bid” blurring — you have pushed past intelligibility. If your throat feels tight or high, you are lifting the larynx rather than reshaping the mouth, and that is the road to strain, not progress.
F1 tracks jaw and throat openness; F2 tracks tongue position and the space behind the lips. F2 is usually the one people can move most, and moving it is what most listeners notice.
The single resonance figure is the geometric mean of F1, F2 and F3, which weights the three equally: a 10% shift in any one of them moves the number by the same amount. That is a deliberate simplification — perceptually F1 and F2 matter more than F3 — so treat the figure as a direction of travel rather than a perceptual score.
The bars show a slow average over roughly the last second and a half, not the instant reading. Formants swing hard between vowels, so a live needle is unreadable noise; the faint grey mark is the instantaneous value if you want to see it move.
The teal tick is your average for this session — the number worth comparing week to week.
These are estimates from a browser mic. Watch the direction across sessions, not the exact hertz.
Vowel space
Centroid F1--Centroid F2--Points0vs baseline--
Reading the vowel space. Every voiced moment plots as a dot: F2 across (bright to dark, left to right) and F1 down (closed to open). This is the standard clinical view of resonance, and it is the clearest single picture of vocal tract shape.
Different vowels naturally sit in different places — the grey letters are roughly where typical adult vowels land. The cloud you make while talking is your vowel space.
Resonance work generally moves the whole cloud up and to the left: smaller, more forward oral space. What matters is your cloud moving over weeks, not where it starts.
The circle marks your current centroid. Press the ◎ button to pin a baseline, then compare later sessions against it.
Vowel identity has to stay intelligible — if you push so far that vowels collide, speech gets hard to understand. Movement of the whole cloud beats chasing one corner.
Landmarks are rough population averages, not targets. Yours will differ, and that is fine.
Clips stay in memory only — they are never uploaded, and they disappear when you close the page. Download to keep one. The shelf holds your last six takes so you can hear them back to back.
Why recording matters more than any meter. You cannot hear your own voice accurately while speaking. Bone conduction adds low frequency, so your voice sounds deeper and fuller to you than it does to anyone else. Playback is the only way to hear what others hear.
Record a sentence, then listen without watching the numbers. Your ear is the instrument that ultimately matters.
The summary covers only the recorded stretch, so you can compare takes directly instead of whole-session averages.
Variation is pitch movement in semitones. Flat delivery reads as robotic regardless of pitch height, and intonation is a strong perceptual cue in its own right.
If a take sounds strained to you on playback, it was strained. Trust that over any number on this page.
Session
Average--Hz
Median--Hz
Range--Hz
In target--%
Std dev--Hz
Variation--st
Voiced0:00
What these numbers read. All of them count only voiced time — silence and breaths are excluded, so a long pause won't drag your average down.
Average and median are your habitual pitch. Median ignores stray octave jumps, so when the two disagree, trust the median.
Range is lowest to highest measured. A very wide range usually means detection errors, not singing.
In target is the share of voiced time inside your band. Anything above roughly sixty per cent while reading is solid.
Std dev is how much your pitch moves. Near zero reads as monotone; twenty to forty Hz is lively, expressive speech.
Greyscale scale
LH
LowerMiddleHigher
Now--Session avg--Range--Time in upper zone--
Pitch--
Resonance--
Speak to analyse
What the scale shows. The dark marker is right now. The teal line is your session average, the shaded band is the range you have covered, and the two ticks are where the adult population reference means fall. Pitch and resonance are shown separately, because a single blended number moves for either reason and cannot tell you which. The combined figure is still at the foot if you want one, but the pair above it is the useful part: when pitch runs well ahead of resonance you are reaching with the larynx, which is the road to strain rather than progress.
Pitch alone is a weak predictor. A high pitch with unchanged resonance sits near the middle here, which matches how such a voice is usually heard.
The reference points are adult population averages, and the two overlap heavily. Individuals vary enormously, and plenty of people are read exactly as they intend well outside either mark.
It is a rough visual cue built from two noisy estimates — not a verdict. No measurement on this page decides how anyone is perceived, and it cannot know anything about you.
The shading behind the pitch trace is still pitch-only: a texture for spotting high and low passages at a glance.
Practice
How to use a session. Short and frequent beats long and occasional — ten minutes most days does more than an hour once a week.
Warm up first, always. Cold voices strain.
Work one thing at a time: pitch today, resonance tomorrow. Chasing both at once usually produces a strained version of neither.
Save a session at the end and compare across days rather than within one. Voices vary hour to hour.
Stop if anything aches. Soreness is the one signal worth obeying immediately.
Saved sessions
What gets kept. Pressing save stores only summary numbers — average pitch, voiced duration, and a thumbnail of the pitch line. No audio is ever recorded, kept, or sent anywhere.
These live in this browser's local storage on this device. Clearing your browser data removes them.
The bars are a miniature of the pitch trace, so you can spot a steadier day at a glance.
Voice quality
--Jitterpitch steadiness
—
--Shimmerloudness steadiness
—
--HNRtone vs breath
—
Press the button, then hold one steady vowel
Reference ranges: jitter under 1%, shimmer under 3.8%, HNR above 20 dB. These vary with microphone, room and person — compare against your own earlier results, not anyone else's.
Not about how you sound to anyone. These measure strain and hoarseness — the things that can actually injure you. They need a held vowel, so press the button and sustain one steady “ah”; on running speech the numbers are meaningless.
Jitter is how much the period between glottal pulses wobbles. Healthy voices sit under 1.2%. Higher values suggest tension, dehydration, or fatigue.
Shimmer is how much the loudness of each pulse wobbles. Under 5% is typical for clean phonation. Breathiness or strain push it up.
HNR (Harmonic-to-Noise Ratio) is how much tonal signal sits above the noise floor. Higher is cleaner; below 10 dB usually sounds rough or breathy.
These are estimates from a browser mic. Use them to watch trends across sessions, not to diagnose pathology.
Guided sequencer
Warm-up · Hum glide
Next: Glides on "ee"
00:00
Press Start to begin a structured warm-up sequence.
Difficulty: Standard
Structured practice beats random drills. The sequencer walks you through a complete session — warm-up, glides, reading, cool-down — with auto-advance and audio cues so you never have to guess what comes next.
Each stage has a set duration. A gentle beep marks transitions.
Difficulty adjusts target range width and glide speed based on your detected vocal range.
Standard mode uses the target band you have already set. Gentle widens it by 20 Hz. Challenge narrows it by 15 Hz and speeds up transitions.
Range detector
Tap the arrow to sustain "ah" from your lowest to highest comfortable note
Lowest--Highest--Usable range--Suggested target--
Your voice is not a preset. This maps your comfortable speaking and singing range by tracking pitch while you glide from low to high. The suggested target band sits inside the densest part of your natural cloud.
Sustain a smooth "ah" for about 10–15 seconds, starting at your lowest comfortable pitch and sliding up to your highest.
The plot shows pitch over time. Flat spots are your comfortable zones; spikes are usually strain or breaks.
The suggestion ignores the top 15% (often falsetto) and bottom 10% (unreliable low end) to find a sustainable speaking target.
Breath support
Shallow · chestDiaphragmatic · supported
Stability--Trend--
Speak to analyse breath support
The hidden engine. Breath support is the single biggest factor in a stable, effortless voice. This module tracks the stability of your amplitude envelope — a proxy for subglottal pressure consistency.
Diaphragmatic breathing produces a smooth, steady envelope with gentle swells. The marker sits in the green zone.
Shallow chest breathing creates a ragged envelope with sudden drops and spikes. The marker drifts left or right of centre.
The histogram shows the last 30 seconds of stability scores. A consistent mid-line means good support.
This is an estimate from mic amplitude, not a pressure transducer. Use it to watch trends, not to diagnose.
Prosody
Your contourExpected statementExpected questionStress
Contour--Stress count--
Pitch patterns, not just pitch points. Prosody is how intonation conveys meaning — questions rise, statements fall, stress highlights key words. Flat prosody reads as robotic regardless of absolute pitch.
The orange line is your pitch contour across the last sentence. The teal dashed line is a typical statement fall; the red dashed line is a typical question rise.
Cyan dots mark detected stress — simultaneous pitch and amplitude accents.
Use this to check you are not flattening out when you focus on pitch height. Intonation is a strong perceptual cue in its own right.
Transcription
Web Speech API — tap play to start
Trap words are highlighted in orange. Words that often drag pitch down are marked in red.
Bridging pitch and content. Real-time transcription lets you see which words pull your pitch down or tighten your resonance — the "trap words" that undo training in spontaneous speech.
Requires a browser that supports the Web Speech API (Chrome, Edge, Safari). All processing stays on-device.
Orange highlights are common trap words: no, hello, okay, actually, just, think, problem.
Red marks are words where the model detected a pitch drop below your target band.
The transcript is never stored or sent anywhere beyond the browser's own speech engine.
Session goals
0%
Daily progress
0Min practised today
0Sessions this week
0Day streak
--Best session
Goals turn diagnostics into habits. Set a daily target for minutes in the target band and a weekly session count, then watch the ring fill as you practice.
The ring shows today's progress toward your minute goal. The numbers update live during recording.
A "session" is any recording you save. The week resets every Monday.
Streak counts consecutive days with at least one saved session.
JSON export includes every stored session in structured format for backup or import into other tools.
Profiles
Each profile keeps its own sessions, baseline, goals and detected range. Data stays in this browser.
One device, multiple people. Profiles let a speech therapist switch clients, or let you track different goals — speaking voice versus singing, for example.
Switching profiles reloads that profile's saved sessions, vowel baseline, and goals instantly.
Nothing is uploaded; each profile is a separate key in localStorage.
Delete a profile to wipe its data permanently.
Pitch histogram
Peak--Spread--Bars0
Where your time actually goes. The graph shows how many voiced frames fell into each pitch bucket. The peak is your habitual pitch; the spread tells you how much you move around.
A sharp peak with narrow spread reads as monotone. A broad hump with a clear centre is natural, expressive speech.
If the peak sits well below your target band, resonance work may do more than pushing pitch.
The histogram only counts voiced frames — silence and breaths are ignored.
Note matcher
-----Hz
flatin tunesharp
Pick a note and speak
How to use the note matcher. Pick a target note, play the reference tone, then sing or speak that pitch while watching the needle.
The needle shows cents deviation: how far you are from the exact pitch. Zero is dead centre. ±50 cents is a quarter-tone — still recognisably the same note.
Use this for pitch anchoring exercises: hold a steady note, then slide it up or down by exact semitones.
The reference tone uses a gentle sawtooth with a soft attack so it doesn't fatigue your ear.
Match by feel first, then check the needle. Internalising the sensation matters more than hitting the number.
Voice range
Lowest--Highest--Dynamic--Points0
Reading the range profile. Each dot is one voiced frame: pitch up the side, loudness across. This is the standard clinical picture of how you use your voice.
A comfortable speaking voice clusters in the middle of your range at moderate loudness. The cloud should feel centred, not crammed at the top or bottom.
If your speaking pitch sits at the very top of the scatter, you're working hard. If it's at the bottom, you have headroom to explore.
Dynamic range is the difference between your loudest and quietest voiced frames. Expressive speech needs at least 10 dB of movement.
Record a passage, then look at the shape. A good target band sits inside a dense part of your natural cloud.
How syllable rate is estimated. The detector looks for peaks in the loudness envelope — each burst of energy roughly corresponds to a spoken syllable.
Typical conversational English runs 4–6 syllables per second (240–360 / min). Reading aloud is often slower.
Raising pitch while keeping rate natural is the goal. Rushing usually tightens the throat and drags resonance back.
The count resets with the session. Use it to check you aren't speeding up when you reach for a new pitch.
Accuracy depends on mic quality and room noise. Treat it as a trend indicator, not a transcript.
Spectral features
Centroid
--
Rolloff
--
Flatness
--
ZCR
--
RMS
--
Press Load to enable (downloads ~30 kB)
What Meyda measures. These are standard spectral descriptors used in speech and music analysis, computed by the Meyda library (loaded on demand).
Spectral centroid is the "brightness" of the sound — where the spectral mass is centred. Higher centroid = brighter, more forward resonance.
Spectral flatness tells you how noise-like vs tonal the signal is. Pure tones score near 0; white noise near 1. Breathiness raises flatness.
Rolloff is the frequency below which 85% of the energy sits. A low rolloff against a high pitch indicates a darker timbre, and the reverse a brighter one.
ZCR (zero-crossing rate) counts how often the waveform crosses zero. High ZCR with low pitch often means frication or breathiness.
Safety and how to use
Pitch is only half of it.Whether a voice reads the way you intend depends at least as much on resonance as on raw pitch.
The zones are guidelines, not rules. The band you pick is where you are aiming today, not a category you have to land in — people are read across a wide range once resonance and intonation come together.
Do not push into strain. If it hurts, feels tight, or leaves you hoarse, stop and rest.
A gender-affirming speech-language pathologist is the gold standard. This is for between-session practice, not a substitute for professional guidance.
Every measurement here is approximate. Browser audio on a laptop mic is noisy — watch trends across sessions rather than any single number.
All processing happens in your browser. No audio is recorded or uploaded — saved sessions are summary numbers kept in this browser only.
Let’s find your starting point
This takes about a minute, and nothing is recorded or sent anywhere.
First, which direction are you working in? You can change this whenever you like, and it only sets a starting band — it does not restrict anything.
Read this out loud in your ordinary speaking voice. Do not reach for anything — this is a measurement, not an exercise. If you push here, everything downstream is calibrated to a voice you cannot hold.
The line changes every few seconds. Read whatever is showing — a paragraph of varied speech measures a speaking voice far better than one sentence said six times, which drifts into a chant.
100%
Start the measurement and watch the bar. If it barely moves, raise the gain until it does — feeding this from monitor output usually needs a good deal more than 100%.
You can change the band any time from the Target range module, and the guide (H) explains the rest of the screen. This window opens each time the page loads — close it with the × or Skip.
How to use this
Start here
If you skipped the guided start, you can still do it by hand: record thirty seconds of ordinary speech, look at the median, and set a band one preset away from it.
Press Record, then read the sentence at the top out loud in your ordinary speaking voice. Do not try to hit anything yet — the first minute is a measurement, not an exercise. Stop, and look at where your median landed.
Then set a voice goal and a target band you can sit in for a whole sentence without pushing. If you can only reach it by straining, drop one band. Straining is the one thing here that can actually hurt you.
What the screen is telling you
The command bar stays at the top: current pitch in hertz, the nearest note, whether you are inside the band, and how much of your voiced time has been in it. The sentence sits directly beneath and stays there while you scroll.
Everything below is optional detail. You do not need any of it to practise.
Workspaces
The rail on the left switches which modules are on screen. Speak is live feedback, Resonance is formants, Drills is guided exercises, Check is measurements that need a held vowel, and Progress is the across-session view.
Each workspace remembers its own arrangement. Open Modules & goal to add or close any module in the workspace you are currently in; other workspaces are unaffected. Reset layout restores just that one.
Pitch is only part of it
Pitch is the easiest thing to measure, which is why it dominates apps like this one. Resonance — the size the vocal tract sounds like — carries at least as much weight perceptually, and intonation carries more than most people expect. A voice held flat reads as flat whatever its pitch.
Practice mode
Press M or the Practice button for pitch and sentence only, at a size you can read from arm's length. Escape leaves it.
Looking after your voice
Stop if anything aches, tickles or feels tight. Warm up before drills. Short frequent sessions beat long ones. Voice change is measured in months, and it moves with illness, tiredness and time of day — read the drift across sessions, not any single day.
A gender-affirming speech-language pathologist remains the gold standard. This is for practice between sessions, not a substitute for one.
Your data
Everything stays in this browser. Sessions store summary numbers only, never audio. Clips live in memory and disappear when you close the page unless you download them. CSV and JSON export everything stored.
Keyboard shortcuts
SpaceRecord / StopNNext sentenceRReset sessionSSave sessionPPitch trace modeWWaveform modeGToggle greyscaleFToggle formantsOToggle pitch overlayZZoom spectrogramLLog frequency scaleMPractice modeHHow to use this?Show this reference1–4Switch colourwayEscClose any modal