Why the limits are worth knowing
The case for practising with an AI partner is straightforward and strong: it is available at the hour you are free, it does not switch to English, it never tires of your pace, and nothing is at stake when a sentence comes out badly. Those properties are what produce the speaking volume most learners cannot otherwise arrange.
None of that is undermined by being clear about the limits. A learner who knows where to verify gets the benefit and avoids the failure modes; a learner who assumes uniform reliability does not.
The limits below are specific rather than general. This is not an argument that AI practice does not work, which would be contradicted by the fact that it obviously does for the parts it covers.
Verify explanations, not practice
The clearest failure mode is confident wrong explanation. Asking why one form is more natural than another can produce a fluent, well-structured answer that is mistaken, and there is nothing in the delivery to signal it.
Questions about nuance, regional usage, etymology and why something sounds unnatural are the highest-risk category, because they are where confident-sounding answers are hardest to check by feel.
| Question type | Reliability | What to do |
|---|---|---|
| Is this sentence natural | Good | Usually trust it |
| Correcting a grammar error | Good | Usually trust it |
| Why is this more natural | Mixed | Verify |
| Regional or generational usage | Mixed | Verify with a person |
| Etymology or history | Weak | Verify with a reference |
| Would this offend someone | Weak | Ask a person |
The pattern is that identifying an error is more reliable than explaining it. Take the correction and be sceptical of the reasoning behind it.
Comfortable practice and false confidence
Practice with an AI partner is comfortable by design: the level adapts, nothing is at stake, and you are never switched to English. That comfort is why it produces volume, and it also means the practice is easier than the situation it is preparing you for.
The risk is a learner who is fluent in practice and finds a real conversation much harder, then concludes the practice was worthless. It was not; it simply did not include the variables that make real conversation hard.
The remedy is a periodic reality check rather than a change of method. One real conversation a week tells you whether the gains are transferring.
- Raise the level deliberately; comfortable practice teaches less.
- Book one real conversation a week as a calibration, however short.
- Expect real conversation to feel harder; that is the missing variables, not a verdict.
- Do not judge your level from practice alone.
What needs a person
Some things are not weaknesses of a particular tool but of the situation. Practice cannot supply the variety of real accents, the noise of a real room, being interrupted, or the consequence of being misunderstood by someone who matters.
Cultural and social judgement is the other category. Whether a phrasing would land badly with a specific person in a specific relationship is a question about lived social knowledge, and it is worth asking a person.
- Accent variety, which practice cannot reproduce at real breadth.
- Noise, interruption and overlapping speech.
- Social consequence, which changes how you perform.
- Whether something would offend a particular person.
- Current slang and generational usage, which move quickly.
Where it is genuinely strong
It is worth being equally specific about the other direction, because the strengths are real and are exactly the things learners struggle to arrange.
Anything that needs many repetitions, no audience and no scheduling is where this is strongest, and that covers a large part of what building fluency actually requires.
| Need | Why it fits | Example |
|---|---|---|
| Daily volume | No scheduling | A session at any hour |
| Repetition | No social cost | The same scene five times |
| Pronunciation scoring | Objective, repeatable | One line, drilled |
| Rehearsal | Safe to fail | A call before you make it |
| Record keeping | Automatic | Session reports |
| Awkward questions | No embarrassment | Did that sound rude |
The last row is underrated. Asking whether something sounded rude, repeatedly, is genuinely difficult to do with a person and is exactly the feedback learners most lack.
A shape that uses both
The practical conclusion is an allocation rather than a choice. Put the volume where volume is cheap and the judgement where judgement lives.
Concretely: daily spoken sessions with an AI Sensei in Unihongo's immersive 3D classroom for the repetitions, the session report for the drill list, flashcards for the words and the pronunciation lab for the sounds; then one real conversation a week to calibrate, and a person for anything about register, offence or current usage.
That shape gets the availability advantage without building on unverified ground, which is the whole point of knowing where the limits are.
Being specific about limits is not a reason to avoid a tool. It is what lets you rely on it for the things it is reliable for.

