Why it feels like nothing is happening
Vocabulary gives daily feedback: you knew nine words this morning and ten this evening. Pronunciation gives almost none, because the change is slow, and because you cannot hear your own output accurately enough to notice a small improvement.
The result is that learners conclude pronunciation is not improving, stop working on it, and thereby make that true. The work was probably fine; the measurement was the problem.
The fix is to measure on the timescale the change actually happens on. Monthly, with a recording of the same passage, is about right.
Realistic timelines by component
The components change at very different rates, and knowing which is which prevents both impatience and premature conclusions about your own ability.
Individual sounds are fastest, because each is one item with a clear target. Rhythm is slowest, because it is a global habit applied to everything.
| Component | To correct in a drill | To hold in conversation |
|---|---|---|
| One consonant | 2-6 weeks | 1-3 months |
| Vowel length | 2-4 weeks | 3-6 months |
| The small tsu | 2-4 weeks | 3-6 months |
| The moraic n | 3-6 weeks | 3-6 months |
| Overall rhythm | 2-3 months | 6-12 months |
| Word pitch accent | Ongoing | 1-2 years |
The gap between the two columns is the important part. Correct in a drill arrives early and means much less than it feels like it does.
The gap between drilled and automatic
Every entry in that table has two stages, and learners routinely mistake the first for completion. Producing a sound correctly when you are thinking about it is a different achievement from producing it while thinking about what to say.
This is why pronunciation appears to regress. It has not; the attention that was maintaining it has been redirected to the content of the conversation, which is exactly what should happen.
The second stage is closed by volume of speaking rather than by more drilling. Once a sound is correct in isolation, additional isolated repetitions do very little, and what it needs is to occur repeatedly in real sentences.
- Stage one: correct when you are attending to it. Weeks.
- Stage two: correct in a rehearsed sentence. Weeks to a couple of months.
- Stage three: correct in unrehearsed conversation. Months.
- Stage four: correct when tired, in noise, or under stress. Longer still.
What actually determines the rate
The strongest single factor is feedback frequency. A learner practising twenty minutes a day with no way to check is largely rehearsing whatever they already do, because the error is inaudible to them; the same learner with a check per attempt improves several times faster.
The second factor is whether you can hear the target. If a contrast is not yet perceptible to you, production practice does very little, and the perception has to come first.
| Factor | Effect | What to do |
|---|---|---|
| Feedback per attempt | Very large | Score or compare every rep |
| Can you hear the target | Very large | Do perception work first |
| Daily consistency | Large | Ten minutes beats an hour weekly |
| Speaking volume | Large | Needed for stage three |
| Total practice hours | Moderate | Less than you would think |
| Age of starting | Small for intelligibility | Not the limiting factor |
The bottom row is worth stating plainly. Adults improve intelligibility substantially, and age is not the reason a given learner's pronunciation is not moving.
Why progress is uneven
Pronunciation improves in steps rather than smoothly. Several flat weeks followed by a sudden change is the normal shape, and it catches learners out because the flat period feels like failure.
There is also interference between components. Working hard on one sound often makes another temporarily worse, because attention is finite, and starting to use contractions frequently costs rhythm for a few weeks.
- Expect flat periods; they are consolidation rather than failure.
- Expect one thing to worsen while another improves.
- Expect regression when tired, in noise, or in a difficult conversation.
- Do not change your target every week; a fortnight is the minimum.
- Judge on a monthly recording, not on how a session felt.
Measuring instead of guessing
Everything above depends on measurement, because unmeasured pronunciation work is a learner practising in the dark and concluding from feel, which is exactly the faculty that is unreliable here.
Two measurements are enough. A monthly recording of the same passage shows the slow trend, and per-attempt scoring during practice supplies the frequent feedback that the table above identifies as the largest factor. Unihongo's pronunciation lab covers the second, with focus modes so the score is about the feature you are working on rather than a general impression.
The third stage, holding it in conversation, needs its own measurement. A session in the immersive 3D classroom puts the sound under load, and the report shows the words that lost it, which is the signal that you have moved from drilling to transfer.
Give any pronunciation target a month before judging it. Most abandoned pronunciation work is abandoned inside three weeks, which is before the change would have been visible.

