Five vowels that do not move
The Japanese vowel system is small and stable. Five vowels, each produced at a single position, held steady from the start of the sound to the end. There are no diphthongs in the English sense within a single mora: where two vowels meet, they are two separate beats rather than one gliding sound.
This is the opposite of English, where many vowels move during production. The vowel in the English word say begins in one place and ends in another, and speakers are entirely unaware of the movement.
Importing that movement into Japanese produces vowels that sound approximately right and unmistakably foreign, because the listener hears a sound that changes when it should not.
What each vowel does
The positions themselves are not difficult, and most learners can produce all five within minutes. What follows is a description of the target rather than a set of exercises, because the work is in stability rather than in placement.
The u is the only one that is genuinely unfamiliar to speakers of most European languages, because it is made without lip rounding.
| Vowel | Mouth | Watch for |
|---|---|---|
| あ | Open, central | Do not push it forward |
| い | Front, close, spread | Do not let it glide |
| う | Back, close, unrounded | Do not round the lips |
| え | Front, mid | Do not glide towards the i |
| お | Back, mid, lightly rounded | Do not glide towards the u |
Three of the five warnings are about gliding. That is the single most common vowel problem for learners from English and several other languages.
Reduction is the bigger problem
English weakens vowels in unstressed syllables, collapsing them towards a neutral central sound. This is so automatic that speakers cannot easily hear themselves doing it, and it is disastrous in Japanese.
Reducing a Japanese vowel removes two things at once: the quality of the vowel and the beat it was carrying. A reduced vowel is shorter, so the mora timing collapses along with the sound, and the word ends up both mispronounced and mistimed.
This is why English speakers often find that their pronunciation is worse in long words than in short ones. Longer words have more syllables to reduce.
| Japanese | Romaji | Meaning |
|---|---|---|
| ありがとうございます | a-ri-ga-to-u-go-za-i-ma-su | Every beat gets full value |
| おはようございます | o-ha-yo-u-go-za-i-ma-su | No syllable is weakened |
| よろしくおねがいします | yo-ro-shi-ku-o-ne-ga-i-shi-ma-su | Long, and evenly timed |
| しつれいします | shi-tsu-re-i-shi-ma-su | Watch the middle beats |
| だいじょうぶです | da-i-jo-u-bu-de-su | The o-u is two beats |
| おつかれさまでした | o-tsu-ka-re-sa-ma-de-shi-ta | Nine beats, all equal |
The devoiced vowels that are not reductions
There is one apparent exception worth understanding, because learners often mistake it for permission to reduce. In certain positions, the vowels i and u are devoiced: the mouth still makes the vowel shape and the beat is still there, but the vocal cords do not vibrate.
This is not the same as English reduction. The timing is preserved and the mouth position is unchanged; only the voicing stops. Learners who treat it as deletion produce a word that is a beat short.
- It happens mainly to i and u between two voiceless consonants.
- It also happens at the end of a word after a voiceless consonant.
- The beat remains; only the voicing is dropped.
- It is regional and speaker-dependent, so it is not compulsory.
- Producing the vowel fully is never wrong, only slightly careful-sounding.
Where two vowels meet
When two vowels come together in Japanese they are two separate beats, not a single gliding sound. This is the point where English habits do the most damage, because English would naturally merge them.
The sequence in the word for teacher is a good test: the last two vowels are two distinct beats, and merging them into one gliding sound is one of the most recognisable learner pronunciations.
| Japanese | Romaji | Meaning |
|---|---|---|
| せんせい | se-n-se-i | Two beats at the end, not one |
| とけい | to-ke-i | Three beats |
| たいへん | ta-i-he-n | Four beats |
| あおい | a-o-i | Three separate vowels |
| いいえ | i-i-e | Three beats |
| ゆうめい | yu-u-me-i | Four beats |
In fast casual speech some of these do merge, and that comes later. Producing them as separate beats is correct and is what you should be able to do first.
Checking for drift you cannot hear
The difficulty with vowel quality is that gliding and reduction are invisible from the inside. They are automatic habits from your first language, which means the sound you produce matches your intention exactly while still being wrong.
A recording helps and a comparison helps more. Unihongo's pronunciation lab plays a native model line and scores your attempt against it, which surfaces the drift and reduction that self-listening misses, and it lets you repeat one line as many times as needed without any of the awkwardness of asking a person.
Then check it under load. Vowels that are steady in a drill and reduced in conversation are the normal pattern, and that gap closes with speaking volume rather than with more drilling, which is what regular sessions in the immersive 3D classroom provide.
Fix length first if you have to choose. Length changes words; quality changes how foreign you sound, which matters less than being understood.

