Affiliate disclosure: some links in my articles are affiliate links, which means I may earn a commission if you buy through one, at no extra cost to you. As of 4 September 2026 I have four approved affiliate partnerships. One of them is JapanesePod101, which is one of the products I assess on this page, and it pays 25 percent on a cookie that does not expire. Two of the others are Japanese knife retailers and have nothing to do with language learning. The fourth is The Japan Shop, a bookshop approved on 4 September 2026, and that one does sell Japanese learning material of the kind this page is about, so I am naming it rather than calling it unrelated. There is no link to it on this page. Other pages here do carry links to that shop, the largest being in my guide to starting to read Japanese, and my affiliate disclosure page keeps the full list current. There are no tracked links on this page, so it currently earns me nothing today, and every judgement below was written before any of those approvals and has not been changed since. You should also know before you start that the single most effective fix I describe on this page is free, the second most effective one runs a ten dollar programme I am not in, and the three products that would pay me the most all fail the main test I am applying. I have marked which is which next to every product.
You have probably been told that the gap between textbook Japanese and real Japanese is slang, contractions, and speed. That the textbook gives you 私は学生です and the street gives you 学生っす, and the work is to learn the second column.
That is not wrong. It is also not the gap. You can learn every contraction in the language and still have conversations that feel like pulling a cart uphill, and the reason is something almost nobody names, because it is not in the part of the conversation you were taught to think of as yours.
Here is the short version. In Japanese, the listener has a job. It is a continuous, audible, high-frequency job, and your textbook cast you as the speaker in every single dialogue you have ever studied. You have spent years practising one half of a two-person task. The half you never practised is the half that makes a conversation feel normal to the person across from you.
I am going to give you the numbers, explain what actually goes wrong in real time, and then go through the products to see which ones do anything about it. I will tell you now that the answer to the last part is disappointing, and that it is the third time in a row this has happened on this site.
The numbers, before anything else
Japanese listeners make noise. Constantly.
The figures I trust here come from research summarised in a University of Hawai’i NFLRC conference paper on teaching 相槌 (aizuchi, the listener’s responses) in Japanese-as-a-foreign-language classrooms. I have not read the original studies, only this summary, so treat the specific numbers as reported rather than as something I verified at source:
- Maynard (1993): Japanese speakers produce roughly twice as many backchannels as Americans.
- Mizutani (1988): native speakers average fifteen to twenty aizuchi per minute.
- Komiya (1986): one every 9.6 seconds in a televised interview, and one every 6.1 seconds on the telephone.
Sit with the last one. On the phone, a Japanese listener makes a supportive noise about every six seconds. Not a reply. A noise: うん, ええ, はい, へえ, そうなんだ. If you are on the phone with me and you stop doing it for fifteen seconds, I will stop talking and say もしもし, because as far as my ears are concerned the line has dropped or you have walked away.
This is the thing your textbook did not tell you. Not because it was a bad textbook, but because textbooks are made of written dialogue, and written dialogue prints the words that carry information. Aizuchi carry almost no information. They are pure channel maintenance, and they are omitted from the page for the same reason nobody transcribes breathing.
Why silence is not neutral
In English, a quiet listener is often a good listener. Attention can be silent. You nod occasionally, you keep eye contact, you let the speaker finish, and interrupting is the rude thing.
Japanese does not have a neutral silent setting for the listener, in the same way it has no neutral setting for politeness. If you say nothing while I talk, that silence is not read as an absence of signal. It is read as a signal. Depending on the situation, it lands as: I disagree, I am annoyed, I am waiting for you to finish so I can leave, or I am someone very senior to you and am withholding approval on purpose.
That last one matters more than people expect. In a Japanese meeting, the person who backchannels least is usually the most senior person in the room. Silence from a listener is a status move. When a learner sits quietly out of politeness, they are unwittingly performing seniority at their own boss, and nobody will ever explain this to them, because explaining it would require saying out loud that they were behaving as though they outranked someone.
So the first correction to make: you are not being polite by listening quietly. You are producing a marked, slightly cold, slightly presumptuous version of listening, and getting warm behaviour back anyway because everyone has decided in advance that the foreigner is doing their best.
The mechanism that makes real conversation collapse
Here is the part I think is genuinely useful, because it explains a specific failure you have probably experienced and blamed on your listening comprehension.
Japanese sentences are built with slots in them where the listener is supposed to come in.
Real spoken Japanese does not arrive as clean units. It arrives in chunks joined by hooks: 〜んですけど, 〜でね, 〜てさ, 〜じゃないですか. Every one of those is a place where the speaker deliberately eases off and waits about half a second for you to make a noise. The speaker is not pausing because they finished. They are pausing because they need a receipt before they spend the next clause.
Now watch what happens when the listener does not know this.
The speaker leaves the slot. The learner, trained on written sentences, hears an incomplete clause and waits politely for the sentence to end. The speaker, receiving nothing, assumes the listener did not understand, and either repeats the chunk more slowly or abandons the sentence and asks 大丈夫? The learner, who understood perfectly, now believes their comprehension failed. Both people quietly downgrade their estimate of how well this is going.
Or the other version: the learner correctly identifies the pause as a gap and starts a full reply into it. The speaker was mid-sentence and continues. They collide. The learner apologises, the speaker apologises, and the learner concludes they are bad at conversation. They are not. They filled a six-second slot with a six-second response, when the slot was built for a syllable.
This is the actual textbook-versus-real gap, and it has nothing to do with vocabulary. It is a difference in how the two languages divide the labour of keeping a conversation alive, and you cannot fix it by learning more words, because the words in question are うん and へえ and you already know them.
はい does not mean yes
The single most consequential thing textbooks get wrong here is a translation you learned in week one.
はい is taught as yes. As an aizuchi, it very often means nothing of the sort. It means: I am here, I am following, keep going. Research on aizuchi makes the same distinction — utterance-internal backchannels signal continuation and understanding, not agreement.
The business consequence is famous and entirely real. A Japanese counterpart listens to your proposal saying はい, はい, はい throughout, and then declines. The foreign side feels misled, and describes Japanese people as indirect or dishonest. Nobody lied. Every はい meant received, and the learner’s textbook had glossed all of them as agreed.
Once you know this, the receiving side of it becomes obvious too. When you say はい while I am talking, I do not hear you agreeing with me. I hear you keeping the line open. If you actually want to agree with me you have to do something else, which brings me to the part that gets learners in trouble in the opposite direction.
Doing it wrong is worse than doing it too little, in two specific ways.
First, はいはい — the doubled form, delivered quickly — is dismissive. It means yeah, yeah, I’ve got it, move on. A learner drilling はい as a listening noise and speeding up under pressure produces this by accident on a regular basis, and it reads as impatience with the person speaking.
Second, なるほど. Learners love it, because it feels like the sophisticated one and it is all over the phrasebooks. But なるほど is an evaluation: it means I see, that checks out. Passing judgment on the content of what someone said is a move that flows downward, so plenty of Japanese workplaces treat なるほど toward a superior as subtly out of line. Your manager does not need you to certify that their point was sound. Nobody will correct this either. It is exactly the failure shape I described for politeness levels: the form is correct, the social condition is not, and the cost is invisible and cumulative.
So the inventory is not the hard part, and the inventory is the only part anyone teaches. What matters is placement, frequency, and which one you are allowed to use on the person in front of you.
The evidence that teaching the words does not work
I want to be careful here, because this cluster has already burned me once on a claim of this shape. In my first article I wrote that no product taught pitch accent systematically, discovered a JapanesePod101 lesson that did, and had to correct a published article. So this time I went looking specifically for evidence against my thesis.
I found it, and it turns out to make the thesis stronger.
The Hawai’i paper I cited above is not about apps. It studied university Japanese courses that taught aizuchi explicitly: two hours of direct instruction at the start of the intermediate course, an hour and a half in the advanced course, reinforced through scripts, model dialogues, and handouts. This is far more coverage than any commercial product gives it.
The finding was that this instruction had limited transfer to spontaneous conversation.
That is the whole problem in one sentence. Aizuchi is not knowledge. It is a real-time motor habit running on a six-second clock while you are simultaneously parsing a language you are not fluent in. Explaining it for two hours produces students who can define it and cannot do it, in the same way that explaining a tennis serve produces nobody who can serve.
Which means the correct test for a learning product is not does it cover aizuchi. It is: does it ever put you in the listener’s chair and require you to make the right noise at the right moment, under time pressure, with a consequence for missing it?
That is the test I applied below. One product came close. None passed.
How I judged these
Three questions, in order.
1. Does it expose you to real conversational Japanese at real speed? Not slowed studio dialogue with the hooks edited out. Speech with 〜んですけど and 〜でね in it, including the other person’s noises, so that you at least hear the rhythm you are supposed to join.
2. Does it teach the listener’s inventory? The words themselves, what they mean, and ideally which ones are safe with whom.
3. Does it make you produce the listener’s half in real time? The actual test. Not typing an answer, not repeating a line after a beep. Coming in during someone else’s sentence, in the gap, without being told the gap is there.
The comparison
| Product | Real conversational speed? | Teaches the aizuchi inventory? | Makes you produce it in real time? | Can it pay me? |
|---|---|---|---|---|
| A real human (italki, HelloTalk) | Yes, by definition | Only if you ask | Yes. The only thing here that does | italki ~$10, HelloTalk no program |
| Native TV and video | Yes, and in volume | No, but it shows every pattern | No. You are a viewer | Free on YouTube and Netflix |
| JapanesePod101 | Yes, plus slowed versions | Best coverage here. A lesson teaching そうそう / だよねー / うんうん / へえー | No. Definitional only | Yes |
| Rocket Japanese | Yes, native-speed conversations | Not that I could find on their pages | Closest attempt. You voice both sides, but the lines are scripted | Yes, and the highest rate — I am not in it |
| Lingopie | Yes. Real shows with subtitles | No | No. Watching, with a click-to-look-up layer | Yes, 30% recurring |
| Pimsleur | Partly. Clean studio dialogue | No | No. You are always the responding speaker | Yes, and the biggest per sale |
| Migaku | Yes, it runs on real video | No | No | Unverified — no program on their own site |
| Bunpro | No. It is written grammar | No. No grammar points exist for these | No | No — points, not cash |
| WaniKani | Not applicable | No, and it never claimed to | No | No — no affiliate program exists |
| Rosetta Stone | No | No | No | Yes, and I still say skip it |
Product by product
A real human — the answer, and the one that barely pays
I am putting this first because it is the honest ranking and because it costs me money to do so.
The listener’s half is a real-time skill with a live partner. There is exactly one category of product that provides a live partner: lessons with a tutor (italki), or language exchange with someone learning your language (HelloTalk). Nothing else on this list can generate the thing you need, which is a person producing a hook and then waiting half a second.
An AI voice partner does not fill this gap, for a reason I set out in my comparison of AI Japanese conversation apps: it reads your pause as the end of your turn.
Two practical notes, because “get a tutor” is useless advice on its own. Tell the tutor what you are working on. A Japanese tutor will not spontaneously drill your backchannels, because from the inside it does not look like a skill; it looks like being a normal person. Ask them explicitly to talk at you for two minutes about anything while you do nothing but respond, and to tell you when it felt off. And do it on audio, not video, at least sometimes. Video lets you get away with nodding, which is a real aizuchi channel but not the one that is failing you on the phone.
italki pays somewhere around ten to fifteen dollars for a new student’s first purchase — their own help pages say at least ten. That is below the threshold I use for deciding whether a product is worth building a business around, and it is the top recommendation on this page anyway. HelloTalk has no conventional affiliate program at all, just an ambassador scheme for creators. Can it pay me? About ten dollars, and nothing, respectively.
Native TV and video — free, and better than the paid version of itself
The second most effective thing is to watch Japanese people talk to each other and pay attention to a channel you have been ignoring.
Not for the dialogue. For the listener. Pick any Japanese talk show, variety programme, or two-person YouTube video, watch the person who is not talking, and count. You will hear the six-second rhythm inside a minute, and you will start to notice that the noises are not random: へえ for new information, そうなんだ for mild surprise, うんうん for keep going, ええ for polite keep going, なるほど for evaluation, and a completely different intonation on each depending on whether it is sincere.
This costs nothing. It is on YouTube right now. I want to be clear about that before I mention the products that package it.
Lingopie — the packaged version of the free thing
Lingopie is real Japanese television with dual subtitles and click-to-look-up. It is genuinely pleasant, and the content is authentic in a way studio dialogue is not, which means the aizuchi are all in there.
But it does not teach the listener’s role, and the fundamental activity is watching. Its affiliate program is the most generous of any product here — the official page states 30% recurring commission, paid Net 30 by PayPal — and I still have to tell you that a Netflix subscription with Japanese audio and Japanese subtitles does most of the same job. Pay for Lingopie if the friction of the lookup layer is what stops you watching. Do not pay for it expecting it to fix this problem. Can it pay me? Yes, more than most.
JapanesePod101 — the best coverage of the inventory, and it stops there
Of everything I checked, this is the product that actually teaches aizuchi as a topic. I confirmed a lesson in the Absolute Beginner series, 5 Phrases Your Teacher Will Never Teach You, that covers そうそう, だよねー, うんうん and へえー, and the hosts explain the intonation difference on へえー between polite interest and genuine surprise, which is a distinction I would not have expected a beginner lesson to make. Their blog covers it as well.
The ceiling is the format. The lesson tells you what the words mean. It does not ask you to produce one, and there is no mechanism anywhere in the product that puts you in the listener’s chair with a clock running. It gets you the inventory, which is roughly the two hours of explicit instruction the Hawai’i study found did not transfer.
There is a second, real strength: the dialogues are played at natural speed and then broken down, so you are at least hearing the hooks. Can it pay me? Yes.
Rocket Japanese — the closest anyone gets, and it is still not it
Rocket’s Play the Part series is the only feature I found that is even structurally adjacent to the right exercise. Their own page describes it as practising “both sides of practical conversations”, with 60 practice conversations and 1,382 phrases run through their voice recognition scoring.
Both sides. That is the right instinct, and no other product on this list has it.
But what is scored is whether you said a scripted line accurately, and the line you are given is a speaker’s line, not the listener’s half. Reading the other character’s dialogue is not the same as reacting to a person, because the script tells you when your turn is. The entire difficulty is that in real conversation nothing tells you when your turn is. I found no mention of aizuchi or backchannelling anywhere in Rocket’s material for this feature.
I should be straight about the incentive here: Rocket’s content-partner programme pays the highest rate of anything on this page, and this is the second consecutive article in which I have marked Rocket down on the thing the article is about. Their product is good. It is a complete, well-built course. It just does not do this. Can it pay me? Yes, and at the highest rate here — I am not in that programme.
Pimsleur — a speaking machine, permanently pointed the wrong way
I keep rating Pimsleur’s method highly and I will again: nothing else reliably forces you to open your mouth on schedule. If your problem is freezing, this is the fix.
But look at the structure of a Pimsleur lesson. A prompt is given, a silence follows, you speak into it. Their own description of the flow is listen and learn, then recall through roleplay transcripts and speed rounds, then confidence through an AI voice coach. In every one of those, you are the person producing the content and the machine is the one waiting. The role you occupy is the exact opposite of the role you need to practise, and it is that way by design, for good reasons.
So it is a strong product with a structural blind spot on this specific axis. It also has the largest single commission of anything I have looked at in this cluster, and it is going in the bottom half of this table. Can it pay me? Yes, and the most per sale.
Migaku — the same category as Lingopie, with less transparency
Migaku turns real video into study material, which puts it in the right family for this problem. The reason I cannot rank it properly is that when I went to check its affiliate terms, the affiliate page on its own site returned a 404, and the rate is not published anywhere I can verify. Third-party networks list a program; no figures.
That does not make it a bad product. It does mean I am not going to tell you what it pays, and you should notice that a great many of the “best Japanese app” articles you have read are published on Migaku’s own blog. That was the finding that started this whole cluster: the people writing the comparisons are usually selling one of the entries. Can it pay me? Unverified.
Bunpro — I checked, and there is nothing here
Bunpro is my usual answer for grammar production, and it deserves credit for being the one product in this cluster that makes you type the form yourself. So I searched its grammar library specifically for the interjections.
There is no grammar point for なるほど. None for そうなんだ, none for へえ. There is a そう entry, but it is the 〜そう construction meaning seems / appears, at N4, which is a different thing entirely.
That is not a criticism of Bunpro so much as an observation about what a grammar SRS can be. Aizuchi has no conjugation, no rule, and no correct written answer. It is a timing skill, and there is no way to put it in a text box. Can it pay me? No — points, not cash.
WaniKani — still not in this race, still my most-recommended product
Kanji and vocabulary. No conversation, no listener role, nothing to evaluate here. I include it in every article in this cluster so its absence is not mistaken for a mark against it, and because it remains the thing I recommend most often overall while having no affiliate program in existence. Can it pay me? No.
Rosetta Stone — three articles, three times skipped
Picture matching cannot teach a timing behaviour whose entire meaning is social. It does not teach the inventory, it does not create the situation, and its core method — inferring meaning from images without explanation — is structurally unable to convey do not say なるほど to your boss.
It would pay me a commission. Can it pay me? Yes, and no.
What I went looking for and could not find
The obvious product does not appear to exist yet, and I want to describe it precisely, because if you are reading this and you build language software, this is a real gap.
What is needed is an AI conversation partner that talks while listening to you — one that produces natural hooks, holds a half-second slot, and reacts when you miss it or fill it wrongly. Voice models can already do the hard part of this. The academic side has been at it for a while; there is a 2024 CHI paper on an Aizuchi-bot built to explore exactly this kind of co-adaptive interaction.
I searched the consumer AI Japanese conversation apps and could not find one that drills the learner’s own backchannel timing. I am stating that as I could not find one, not as none exists — this cluster has already taught me the difference the hard way, and the products in this space ship features weekly. If you know of one, I would like to hear about it, and I will update this page.
The pattern I have now hit three times
This is the fourth article on this site about Japanese learning products, and something has come up in three of them in a row.
In the pitch accent article, two products display pitch information and neither one ever requires you to produce a correct pitch pattern. In the casual speech article, every product teaches both registers and not one asks you to choose between them for a specific person. Here, several products teach the aizuchi vocabulary and none of them puts you in the listener’s chair.
The shape is identical every time: the material is present, the production requirement is absent. And I think the reason is structural rather than lazy. All three are things a grader cannot check. There is no right answer to mark, only a right moment, and a right moment requires the software to be a participant rather than an examiner. Multiple-choice pedagogy has quietly determined the boundary of what Japanese courses teach, and everything outside that boundary — pitch, register selection, backchannel timing — is exactly the set of things that make you sound like a person rather than a well-drilled textbook.
That is not a reason to avoid these products. It is a reason to know what you are buying: an excellent supply of material, and no supply at all of the thing that turns material into behaviour.
So what should you actually do
Today, for free: watch one Japanese conversation video and ignore the speaker. Watch the listener for five minutes. That single change in attention is most of the insight on this page, and it will make you notice something you have heard ten thousand times and never registered.
This week: pick three. うん for casual, ええ or はい for polite, and へえ for anything new to you. Three is enough. Learners who try to deploy the full inventory sound like they are performing Japanese, and the doubled はいはい problem is waiting for anyone who drills too hard.
In your next conversation: aim for the hooks. When you hear 〜んですけど or 〜でね or 〜てさ, that is your cue, and it is not the end of the sentence. Do not wait for the end of the sentence. In Japanese the end of the sentence is often the least interesting place for you to speak.
When you are ready to practise properly: shadow for ten minutes a day, then book one tutor session, on audio, and spend it listening. Tell them in advance that you want them to talk and you want feedback on your reactions. Nobody uses a lesson this way and it is the highest-value hour available to you on this specific problem.
What not to do: do not buy another course to fix this. Every course on this page will sell you more input, and input is not what is failing. You already understand more Japanese than you can comfortably participate in, and that gap is the whole complaint you started with.
One last thing, from my side of it
I want to describe what it is actually like on the other end, because I think it reframes the effort.
When a learner backchannels well, I do not think their Japanese is good. I do not think about their Japanese at all. That is the entire point. The conversation just runs — I get my receipts, I spend the next clause, we get somewhere. Afterwards I would probably describe them as easy to talk to, which is a comment about them as a person and not about their grammar.
When a learner does not, I work harder than I realise. I over-explain, I check whether they followed, I shorten what I was going to say. I do not experience this as their Japanese is weak; I experience it as the conversation being slightly heavy. And I have no idea that I am doing it, which is why you will never be told.
So the reason this is worth your attention is not that it makes you sound native. It is that it is the cheapest possible improvement to how people feel after talking with you, and it is made of four syllables you already know. Everything else in this cluster — the pitch patterns, the register calls — takes months. This one you can start on in the next conversation you have, and the person you are talking to will never know why it went well.