Two registers, not one language with two speeds
It is tempting to treat speech as writing delivered aloud, with a few casual words swapped in. In Japanese that model breaks down quickly, because the differences are structural: which particles appear, how clauses are joined, how long a sentence runs, and which sentence endings are available at all.
This matters for learners because most study material is written, so the input is systematically skewed towards the register you use less often. Learners who read a great deal and speak little develop a large gap between what they can produce and what sounds natural out loud.
The correction is not to abandon written forms. It is to learn the spoken register as a separate thing, with its own habits, and to know which one you are in.
Particles that disappear in speech
Casual spoken Japanese drops particles that written Japanese requires, and the drops are systematic rather than random. Wa and o go most often; ga and ni go sometimes; de and to rarely go, because they carry information that cannot be recovered from context.
Dropping happens when the relationship is obvious. Nobody drops a particle that would leave the sentence ambiguous, which is why the rule is easier to acquire by listening than by memorising.
| Japanese | Romaji | Meaning |
|---|---|---|
| ご飯食べた? | Gohan tabeta? | Have you eaten? (o dropped) |
| これ、いくら? | Kore, ikura? | How much is this? (wa dropped) |
| 時間ある? | Jikan aru? | Do you have time? (ga dropped) |
| 明日、行く? | Ashita, iku? | Are you going tomorrow? |
| それ、どこで買ったの? | Sore, doko de katta no? | Where did you buy that? (de stays) |
| 田中さんと話した? | Tanaka-san to hanashita? | Did you talk with Tanaka? (to stays) |
Notice that de and to survive in the last two. Those particles carry meaning that context cannot supply, so they are not dropped.
Contractions that are standard, not sloppy
Spoken Japanese contracts heavily, and these contractions are neutral rather than casual: they appear in ordinary polite conversation and in speech by people being careful. Failing to use them is a stronger signal of learner Japanese than almost anything else.
Learn to hear them first. A large proportion of the listening difficulty learners report with natural speech is contraction rather than speed.
| Japanese | Romaji | Meaning |
|---|---|---|
| 食べてる | tabeteru | From tabete iru, eating |
| 言っちゃった | itchatta | From itte shimatta, said it (regrettably) |
| 見とく | mitoku | From mite oku, will look in advance |
| なくちゃ | nakucha | From nakute wa, have to |
| やっぱ | yappa | From yappari, as expected |
| すみません | sumimasen | Often heard as suimasen in speech |
Sentence length and how clauses are joined
Written Japanese tolerates long sentences with several subordinate clauses, because the reader can go back. Speech does not, because the listener cannot. Spoken Japanese uses shorter units chained with connectors, and the chaining does the work that subordination does on the page.
For learners this is good news: the spoken register is structurally easier. Three short clauses joined with kedo and sorede are both more natural and easier to produce than one carefully nested sentence.
| Feature | Written | Spoken |
|---|---|---|
| Sentence length | Long, nested | Short, chained |
| Joining | Subordination | Connectors between units |
| Particles | All present | Often dropped |
| Forms | Full | Contracted |
| Endings | Plain or polite | Plus ne, yo, no, kana |
| Repetition | Avoided | Normal and useful |
The last row surprises people. Repeating yourself in speech is helpful to the listener, whereas in writing it looks like poor editing.
Endings that only exist in speech
Spoken Japanese has a set of sentence endings that carry attitude and manage the conversation: seeking agreement, softening, expressing realisation, checking. They barely appear in formal writing, so learners who study from text often produce sentences with no ending at all, which sounds abrupt.
These are worth adopting early, because they do a lot of social work for very little grammatical cost.
| Japanese | Romaji | Meaning |
|---|---|---|
| そうですね。 | Sō desu ne. | That is so, is it not. |
| これでいいよね? | Kore de ii yo ne? | This is fine, right? |
| 明日だっけ? | Ashita dakke? | It was tomorrow, was it not? |
| 行けるかな。 | Ikeru ka na. | I wonder if I can go. |
| やっぱりそうだったんだ。 | Yappari sō datta n da. | So it was that after all. |
| ちょっと違うかも。 | Chotto chigau kamo. | It might be a little different. |
Keeping the two apart
The risk of learning the spoken register well is bleed in the other direction: contractions in an email, dropped particles in a report, a casual ending in a meeting. That is a real problem and it is noticed, because writing is where formality is expected to be exact.
The way to keep them separate is to practise them separately and to know which one you are in. Speaking practice should be spoken, and writing practice should be written, which sounds obvious and is exactly what a text-based chatbot blurs: you type a conversation, so the register you rehearse is a hybrid that belongs to neither.
Practising conversation aloud with an AI Sensei in Unihongo's immersive 3D classroom keeps the spoken register where it belongs, and the session report will flag written forms appearing in speech, which is the mistake this guide is about.
A good test: read your last email aloud. If it sounds like something you would say, it is probably too casual for the page.

