What a general-purpose chatbot does well
Be fair to the chatbot first. It explains grammar at any length, in your own language, and will rephrase until the explanation lands. It translates in either direction instantly, so the sentence you could not parse in a novel is clear in seconds. It produces example sentences on demand, compares two patterns you keep confusing, and never tires of the same question asked four ways. For understanding Japanese, it is an extraordinary tool.
- Ask why a native speaker chose one pattern over another in a sentence you met
- Generate ten example sentences for a grammar point, then ask for ten more at a lower level
- Translate a paragraph you read and ask which words were essential
- Check a sentence you wrote and ask what a native speaker would say instead
Where text fails a speaking learner
Now look at what happens when the same tool is used for speaking practice. You type, so every sentence is composed, checked and edited before it counts, and your mouth is never involved. You read the reply, so you never train your ear, and words arrive at the speed you choose rather than the speed a speaker chooses. You never hear yourself, so you cannot notice that your tsu sounds like su or that your pitch flattens every question. Nothing is listening, so pronunciation feedback is impossible in principle, not merely absent.
There is also no scene. A chatbot will play a waiter if asked, but the waiter is a paragraph, and you can scroll up, reread the menu and take a minute over your reply. Real scenes have faces, timing and a goal you can fail. Finally, the exit to English is always one keystroke away, and because the tool answers in whatever language you use, most learners drift there by the second exchange without noticing.
The drift into English, and the lines that stop it
The drift deserves its own section because it is the failure learners least notice. In a text window, asking for help in English costs nothing, so help gets asked for in English, and the practice quietly becomes a conversation about Japanese rather than in it. The cure is a small set of Japanese lines for asking for help, said aloud, so the meta-conversation stays in the target language. They work with any teacher, human or AI, and they are the first things to make automatic.
| Japanese | Romaji | Meaning |
|---|---|---|
| 日本語で何と言いますか? | Nihongo de nan to iimasu ka? | How do you say it in Japanese? |
| 予約ってどういう意味ですか? | Yoyaku tte dō iu imi desu ka? | What does yoyaku mean? |
| もう少し簡単な言葉でお願いします。 | Mō sukoshi kantan na kotoba de onegai shimasu. | In simpler words, please. |
| 例文をください。 | Reibun o kudasai. | Please give me an example sentence. |
| 例えば、明日は雨が降りそうです。 | Tatoeba, ashita wa ame ga furisō desu. | For example: it looks like rain tomorrow. (the teacher's reply) |
| 分かりました。もう一度やってみます。 | Wakarimashita. Mō ichido yatte mimasu. | Understood. I will try once more. |
What a 3D classroom with voice adds
A purpose-built AI Sensei starts from the opposite assumption: the session is spoken unless you choose otherwise. In Unihongo's immersive 3D classroom you face a teacher who speaks first, and your reply is your voice. That single default reverses every failure above at once. You must retrieve at speaking pace, you listen because there is nothing to read, you hear yourself, and something is listening back. The classroom supplies the mild stakes of a lesson that a text box cannot, because someone is visibly waiting for you.
Around that default sit four things a chat prompt cannot easily provide. You choose the teacher's personality, gentle, strict, funny or relaxed, once, and the tone of every correction follows from it. Roleplay missions stage a cafe, a hotel, a job interview or a workplace call with a goal you can fail. Pronunciation coaching reacts to the sounds you actually made. And a session report afterwards lists the corrections, so the review is a short list rather than a scrollback you have to reread.
Personality choice and the tone of correction
Teacher style is not cosmetic. The same learner sentence gets a different reply from each personality, and that changes what you practise next. A gentle teacher follows the content and keeps you talking; a strict one stops the flow to fix a sound; a funny one throws in a sideways question that forces an unplanned answer; a relaxed one keeps the register casual so you hear plain forms. Rotate them over a week and you get four kinds of pressure from one sentence.
| Japanese | Romaji | Meaning |
|---|---|---|
| 昨日は友達と映画を見ました。 | Kinō wa tomodachi to eiga o mimashita. | Yesterday I watched a film with a friend. (the learner) |
| いいですね。どんな映画でしたか? | Ii desu ne. Donna eiga deshita ka? | Nice. What kind of film was it? (gentle) |
| もう一度。「みました」の「し」をはっきり言ってください。 | Mō ichido. "Mimashita" no "shi" o hakkiri itte kudasai. | Once more. Say the shi in mimashita clearly. (strict) |
| 映画?ポップコーンは食べましたか? | Eiga? Poppukōn wa tabemashita ka? | A film? Did you eat popcorn? (funny) |
| へえ、面白かった? | Hē, omoshirokatta? | Oh, was it good? (relaxed, casual form) |
Beginners usually do best starting gentle and moving to strict once the basic lines are automatic, because correcting every sound from day one stalls the conversation before there is anything to correct. Use the relaxed teacher when you want to hear casual Japanese, and label those forms in your notes, since the polite forms in the other rows are still the safer default with a stranger.
Lines you say in a voice session that you would never type
The clearest way to see the difference between the tools is to list what actually comes out of a learner's mouth in a spoken session. None of the lines below would ever be typed, because typing removes the hesitation, the self-repair and the requests for repetition that make up half of real conversation. They are not mistakes to eliminate. They are the connective tissue that lets you hold a turn while you think, and a text window never lets you practise them.
| Japanese | Romaji | Meaning |
|---|---|---|
| えっと、昨日は… | Etto, kinō wa... | Um, yesterday... |
| あの、すみません。 | Ano, sumimasen. | Er, excuse me. |
| そうですね… | Sō desu ne... | Well, let me think... |
| えー、何だっけ。 | Ē, nan dakke. | Uh, what was it again. (casual) |
| あ、違います。三時じゃなくて、四時です。 | A, chigaimasu. Sanji ja nakute, yoji desu. | Oh, no. Not three, four o'clock. |
| えっと、二泊…二泊です。 | Etto, nihaku... nihaku desu. | Um, two nights... two nights. |
| 言い直します。 | Iinaoshimasu. | Let me say that again. |
| 聞き取れませんでした。もう一度お願いします。 | Kikitoremasen deshita. Mō ichido onegai shimasu. | I could not catch that. Once more, please. |
| ちょっと待ってください。 | Chotto matte kudasai. | Just a moment, please. |
| はい、はい。 | Hai, hai. | Yes, yes. (showing you are following) |
| なるほど。 | Naruhodo. | I see. |
| あ、そうか。分かりました。 | A, sō ka. Wakarimashita. | Ah, right. Got it. |
| 先生、今の発音は合っていますか? | Sensei, ima no hatsuon wa atte imasu ka? | Sensei, was that pronunciation right? |
| 「つ」が「す」に聞こえました。もう一度。 | "Tsu" ga "su" ni kikoemashita. Mō ichido. | Your tsu sounded like su. Once more. (the teacher's reply) |
Two of these deserve daily practice: the self-correction, sanji ja nakute yoji desu, because it lets you fix a number without restarting the sentence, and the repeat request, because it keeps you in Japanese when you have lost the thread. The guides to Japanese filler words and aizuchi cover both in depth; here the point is only that a spoken tool is the only place you will ever rehearse them.
A mission compared with a chat prompt
Ask a chatbot to interview you and you get a list of questions you can answer at leisure, in writing, with a dictionary open. Start the job interview mission and the teacher asks the first question aloud, waits, and reacts to what you said rather than to what you meant. The lines below are the spine of that mission. The learner's replies are short by design; the skill being trained is answering promptly and asking for a repeat without leaving Japanese, not producing an elegant paragraph.
| Japanese | Romaji | Meaning |
|---|---|---|
| 自己紹介をお願いします。 | Jikoshōkai o onegai shimasu. | Please introduce yourself. (interviewer) |
| 山田と申します。大学でデザインを勉強しました。 | Yamada to mōshimasu. Daigaku de dezain o benkyō shimashita. | My name is Yamada. I studied design at university. |
| なぜこの会社を選びましたか? | Naze kono kaisha o erabimashita ka? | Why did you choose this company? (interviewer) |
| 御社の製品をよく使っているからです。 | Onsha no seihin o yoku tsukatte iru kara desu. | Because I often use your company's products. |
| すみません、もう一度質問をお願いできますか? | Sumimasen, mō ichido shitsumon o onegai dekimasu ka? | Sorry, could you repeat the question? |
| 本日はありがとうございました。 | Honjitsu wa arigatō gozaimashita. | Thank you for today. (interviewer) |
| こちらこそ、ありがとうございました。 | Kochira koso, arigatō gozaimashita. | Thank you too. |
Notice that the mission uses onsha for your company and to mōshimasu for the name, both of which a beginner will have read about and never said. The guide to honorifics explains the forms; the mission is where you find out whether you can produce them while someone waits. Expect the second attempt to go much better than the first, and the fourth to feel almost boring, which is the moment to switch the personality to strict.
Side by side
| Question | General-purpose chatbot | AI Sensei in a 3D classroom |
|---|---|---|
| How you answer | You type | You speak aloud by default |
| What you hear | Nothing, unless you add a voice tool | The teacher's Japanese, spoken to you |
| Who you face | A text box | A teacher standing in a classroom |
| Pronunciation | Cannot hear you | Listens and corrects the sounds it hears |
| Scene | Whatever you describe in the prompt | Staged missions: cafe, hotel, job interview, workplace call |
| Teacher style | Set by your prompt each time | Gentle, strict, funny or relaxed, chosen once |
| After the session | A scrollback you must reread | A session report with corrections |
| Explanations | Long, detailed, on demand | Short, spoken, then back to practice |
| Translation | Instant, any direction, any length | Brief help inside the lesson |
| Cost of drifting to English | None, so you drift | Noticeable, so you stay in Japanese |
| Best for | Understanding | Producing |
The table is deliberately blunt, so add the caveat: a chatbot with a voice add-on closes some of the gaps in the middle column, and a learner with strong discipline can force any tool to work. The earlier comparison of an AI tutor with a general chatbot covers that workflow question; this table is about defaults, because defaults are what you actually do when you are tired.
Pronunciation coaching and the session report
Pronunciation is the sharpest dividing line. A text tool cannot hear you, so its advice on sounds is generic, however well written. Voice coaching listens to the sounds you produced and points at the one that drifted, which is the only kind of feedback that changes a habit. It is not perfect: a strong accent can be mis-heard, and a correct sound is occasionally flagged. Treat a repeated flag on the same sound as a real signal and a one-off as noise.
The session report is the other half. After a spoken session you have nothing to reread. The report lists the corrections, the lines you stalled on and the words you asked about, so the review takes five minutes and can be done aloud. Learners who skip it lose most of the value of the session; learners who read it and say each corrected line twice tend to make the same mistake only once.
The sensible way to use both
- Speak first: open the classroom before you open a chat window, so the day's Japanese starts as speech.
- Take the session report to the chatbot: paste a correction you did not understand and ask why, then ask for three more examples.
- Bring the examples back: say those three examples aloud in the next session, ideally inside a mission where they fit.
- Use the chatbot for reading and translation between sessions, never as the place you rehearse a scene.
- Keep the languages separate: the chatbot in your own language for understanding, the classroom in Japanese for production.
Honest limits of both
Neither tool knows your life. The chatbot cannot tell whether an explanation matched the sentence you actually met, and the AI Sensei does not know your colleagues, your street or what you said last week unless you bring it into the scene. Neither will hold you accountable when you stop showing up. Neither reproduces the nerves of a real interview panel or the speed of a busy izakaya, and both can be confidently wrong on occasion, so check a surprising claim from either against a reliable reference or a human teacher.
Within those limits the split is clear. Text is for understanding. A voice classroom is for production, and it is the only one of the two that will ever ask you to open your mouth. Real people are for everything else, and the point of both tools is to arrive in front of them with fewer beginner mistakes left to make.

