Why it works as a test
People are generous listeners. They use context, they guess, and they will not tell you that a word was unclear, which means conversation gives you very little information about your own intelligibility.
A transcription engine does none of that. It maps sound to text and reports what it heard, and its mistakes are therefore informative in a way that a polite listener's nods are not.
That bluntness is the whole value. It is not a better judge than a person; it is a more honest one.
| Feedback source | Honesty | Detail |
|---|---|---|
| A friendly listener | Low | Almost none |
| A teacher | High | High |
| Voice input | Blunt | Sound level only |
| A recording of yourself | High | Whatever you can hear |
Voice input fills a specific gap: frequent, free and immediate, at the level of whether the sounds landed.
Setting it up
The setup is short and the most common mistake is skipping it: dictating Japanese while the keyboard language is still set to English produces nonsense that has nothing to do with your pronunciation.
Add Japanese as a keyboard or dictation language, select it before you speak, and use a quiet room with the phone at a consistent distance.
- Add Japanese as a dictation or keyboard language.
- Switch to it deliberately before each attempt.
- Use a quiet room; noise produces misleading errors.
- Hold the device at a consistent distance.
- Speak at normal conversational volume, not louder.
What its errors usually mean
Certain mistakes recur across learners, and they map onto the features of Japanese that non-native speakers most often flatten: vowel length, doubled consonants and the syllabic nasal.
These are timing features rather than sound-quality features, which is why they survive long after a learner's individual sounds have become accurate.
| Japanese | Romaji | Meaning |
|---|---|---|
| おばさん | obasan | Aunt, or middle-aged woman |
| おばあさん | obāsan | Grandmother |
| きて | kite | Come |
| きって | kitte | Stamp |
| かた | kata | Shoulder, or way of doing |
| かった | katta | Bought |
If the transcript keeps giving you the short version, you are not holding the long one long enough. The fix is duration rather than effort.
Running the check
The method is simple and should be quick, or you will not do it often. Choose a few sentences you actually say, dictate each three times, and look only at repeated errors.
One-off errors are usually noise or a slip. Errors that appear all three times point at something systematic in how you produce that sound.
- Pick five sentences from your real speaking.
- Dictate each three times.
- Ignore errors that appear once.
- Note the repeated ones and what they have in common.
- Work on one feature at a time, not five words.
What it cannot tell you
The limits matter, because learners who rely on this alone can plateau at intelligible-but-foreign without realising it.
Transcription engines are trained to be robust and will happily understand speech with wrong pitch accent, unnatural rhythm and a strong foreign accent. It also has no view on whether your sentence was appropriate, polite or natural.
So it answers one question well and everything else not at all.
| Question | Voice input answers it |
|---|---|
| Were my sounds clear enough? | Yes |
| Was my vowel length right? | Usually |
| Was my pitch accent right? | No |
| Did I sound natural? | No |
| Was the grammar correct? | No |
| Was the register appropriate? | No |
Where it fits
Treat it as a quick self-check between sessions rather than a source of instruction. It tells you that something is wrong and rarely what to do about it.
For the what-to-do part you need feedback that can hear the difference between a clear sound and a native-like one. Unihongo's pronunciation lab works at that level, isolating the sounds and the timing features that transcription only hints at, and the report after a session in the 3D classroom separates what was unclear from what was merely accented.
Used together the division is clean: the phone tells you whether you cleared the bar, and the lab tells you how to move it.
If your phone understands you reliably, you have passed the intelligibility floor. That is a real milestone and the point at which accent work starts to pay off.

