Why you cannot hear yourself in real time
While you speak, most of your attention is on the sentence you are building. What is left is not enough to evaluate the sounds you have just made, so you hear what you intended rather than what came out. This is not carelessness; it is a hard limit on attention, and it applies to confident speakers as much as nervous ones.
A recording moves the evaluation to a moment when you have nothing else to do. The same ear that missed a shortened vowel while speaking catches it immediately on playback, because it is now the only job.
This is also why a listener's correction sometimes feels wrong. You remember producing the long vowel you intended. The recording settles it.
A setup that takes two minutes
The equipment does not matter. A phone voice memo held thirty centimetres away in a room without a television produces a recording good enough to hear vowel length, pitch and pausing. Do not spend a week choosing a microphone instead of speaking.
What does matter is having a prompt ready, because deciding what to talk about while recording produces a recording of you deciding what to talk about. Keep a list of ten prompts and pick one at random.
- One prompt, chosen before you press record.
- Sixty to ninety seconds, timed, no restarts.
- A quiet room, phone at arm's length, no headphones.
- Say the date and the prompt at the start so the file identifies itself.
- Stop when the timer ends, even mid-sentence.
First pass: listen for flow
On the first playback, ignore pronunciation completely. You are counting events: how many times you stopped mid-sentence, how many sentences you abandoned, how long the longest silence was, how many times you fell back into English.
Flow problems have flow fixes, and they are mostly about strategy rather than knowledge. A four-second silence before a common answer means the answer is not yet automatic. An abandoned sentence usually means you started with a structure you could not finish, and the fix is to start smaller.
| Japanese | Romaji | Meaning |
|---|---|---|
| ちょっと考えさせてください。 | Chotto kangaesasete kudasai. | Let me think for a moment. |
| なんと言えばいいかな。 | Nan to ieba ii ka na. | How should I put it? |
| 言い方を変えますね。 | Iikata o kaemasu ne. | Let me put that another way. |
| まず結論から言うと…… | Mazu ketsuron kara iu to... | To give the conclusion first... |
| 例を挙げると…… | Rei o ageru to... | To give an example... |
| うまく説明できないんですが…… | Umaku setsumei dekinai n desu ga... | I cannot explain it well, but... |
These are not filler. Each one buys thinking time while keeping the floor and telling the listener what is coming, which is what a fluent speaker does with a pause.
Second pass: listen for sound
On the second playback, ignore content entirely. You are listening for four things: vowel length, double consonants, the flatness or shape of pitch across a word, and whether the rhythm gives every mora roughly equal time.
Do it with a pen. Write down the exact words that sounded wrong, not the categories. A note that says long vowels is useless next week; a note that says the second syllable of the word for teacher was short is a drill you can run.
- Any word where a long vowel came out short, or the reverse.
- Any small tsu that vanished, so the word ran together.
- Any word where every mora got a different length, English style.
- Any sentence where the final particle rose when it should have fallen.
Scoring instead of cringing
An unstructured reaction to your own voice is discouraging and produces nothing. A fixed scorecard produces a comparable number and a fix list, and after three months the numbers show movement that your feelings will not.
Keep it to five items and score each out of five. The absolute value does not matter. The trend does.
| Item | What a 5 sounds like | What a 2 sounds like |
|---|---|---|
| Pausing | Pauses fall between sentences | Pauses fall mid-clause |
| Vowel length | Long vowels held fully | Long and short merged |
| Rhythm | Even mora timing | Stressed like English |
| Sentence completion | Every sentence finished | Several abandoned |
| Recovery | Repairs without stopping | Restarts from the beginning |
Score the same prompt each month. Different prompts produce different scores for reasons that have nothing to do with your Japanese.
Letting a pronunciation lab do the scoring
Self-scoring has a ceiling: you cannot reliably hear a contrast you have not yet learned to produce, so the errors you are least aware of are the ones your own review will keep missing. This is where an objective score earns its place.
Unihongo's pronunciation lab gives you a native model line, records your attempt, and scores it, with focus modes for pitch accent, long vowels, the small tsu, the Japanese r and particles. Because the score is per line, you can take the three words your recording review flagged and drill exactly those rather than repeating whole paragraphs.
Pair it with conversation practice rather than replacing it. The lab tells you whether a sound is right in isolation; a spoken session with an AI Sensei in the immersive 3D classroom tells you whether it survives when you are also thinking about what to say, and the report afterwards shows which words slipped under that load.
Words that score well alone and badly in conversation are not pronunciation problems. They are attention problems, and they resolve as the surrounding grammar becomes automatic.

