What a minimal pair is and why it matters
A minimal pair is two words identical except for one sound. They matter because they prove which differences the language treats as meaningful. Japanese has relatively few consonants and vowels, so it leans heavily on length and pitch to keep words apart, and those are exactly the features most learners' ears ignore.
The reason to drill pairs rather than single words is that a contrast is the thing being learned. Saying a long vowel correctly ten times in isolation does not teach you to distinguish it from the short one; alternating the two does.
Perception comes before production. Learners who cannot hear a contrast produce something in between, and are then surprised to be misunderstood, because to them the word sounded fine.
Long and short vowels
This is the highest-value contrast in Japanese, because it is frequent, it changes common words, and almost no learner arrives with it. A long vowel is held for two beats rather than one, which means it is a timing distinction rather than a quality one: the vowel does not change colour, only duration.
Exaggerate at first. A long vowel that feels absurdly long to you is usually close to correct, and the exaggeration can be reduced once the contrast is stable.
| Japanese | Romaji | Meaning |
|---|---|---|
| おばさん / おばあさん | obasan / obāsan | Aunt / grandmother |
| おじさん / おじいさん | ojisan / ojīsan | Uncle / grandfather |
| ここ / こうこう | koko / kōkō | Here / high school |
| とる / とおる | toru / tōru | To take / to pass through |
| ゆき / ゆうき | yuki / yūki | Snow / courage |
| びよういん / びょういん | biyōin / byōin | Salon / hospital |
The last pair is the one to be careful with in real life. Asking for the salon when you meant the hospital is a well-known learner story for a reason.
Single and double consonants
The small tsu marks a doubled consonant, which in speech is a beat of silence before the consonant is released. It occupies a full mora, so the word takes the same time whether or not you produce the pause, and skipping it makes the word shorter and unrecognisable.
The trick that works for most learners is to think of it as a held stop rather than a repeated consonant. Close the mouth, wait one beat, then release.
| Japanese | Romaji | Meaning |
|---|---|---|
| きて / きって | kite / kitte | Come / stamp |
| いた / いった | ita / itta | Was there / said |
| さか / さっか | saka / sakka | Slope / author |
| ぶか / ぶっか | buka / bukka | Subordinate / prices |
| おと / おっと | oto / otto | Sound / husband |
| まて / まって | mate / matte | Wait / waiting |
Pitch accent pairs
Pitch accent is where the pitch drops within a word. It distinguishes fewer pairs than length does, and context usually rescues the meaning, so it is a lower priority. It is worth knowing about early anyway, because pitch is a large part of what makes speech sound native, and because a handful of common pairs really do get confused.
The contrast is a drop, not a stress. Nothing gets louder or longer; the pitch simply falls after a particular mora.
| Japanese | Romaji | Meaning |
|---|---|---|
| はし | HA-shi, pitch drops first | Chopsticks |
| はし | ha-SHI, pitch drops last | Bridge |
| あめ | A-me, pitch drops first | Rain |
| あめ | a-me, flat | Sweets |
| かき | KA-ki, pitch drops first | Oyster |
| かき | ka-ki, flat | Persimmon |
Attach the particle when you drill these. The drop is often only audible on the particle that follows, so practising the bare noun hides the contrast.
A ten-minute daily routine
The routine has three parts and one rule: only one contrast per day. Mixing long vowels and small tsu in the same session slows both, because the ear is being asked to retune twice.
Keep the same six pairs for a week. Novelty feels productive and is not; the gain comes from repetition of the same discrimination until it is automatic.
- Two minutes of listening only: hear each pair spoken, no production.
- Four minutes alternating aloud: long, short, long, short, exaggerated.
- Two minutes in a sentence, because the contrast has to survive context.
- Two minutes recording and checking, or scoring the lines in a lab.
Checking that your two versions are actually different
The hardest part of minimal pair practice alone is verification. Learners routinely produce two versions that sound different to them and identical to everyone else, and no amount of repetition fixes a contrast that is not actually being made.
Unihongo's pronunciation lab is built for this check: it plays a native model line, records your attempt, and scores it, with a focus mode for long vowels, one for the small tsu and one for pitch accent, so you can point the scoring at the contrast you are drilling instead of at general accuracy. Two words from the same pair that score the same are a clear signal that the difference is not there yet.
Once a pair scores apart reliably, take it into speech. Saying the word in a conversation with your AI Sensei in the immersive 3D classroom, under the load of building a sentence, is a much harder test than saying it alone, and the session report will show whether it held.
A contrast is finished when it survives conversation, not when it is correct in a drill. That gap is normal and takes a few weeks to close.

