Affiliate disclosure: some links in my articles are affiliate links, which means I may earn a commission if you buy through one, at no extra cost to you. As of 4 September 2026 I have four approved affiliate partnerships. One of them is JapanesePod101, which is one of the products I assess on this page, and it pays 25 percent on a cookie that does not expire. Two of the others are Japanese knife retailers and have nothing to do with language learning. The fourth is The Japan Shop, a bookshop approved on 4 September 2026, and that one does sell Japanese learning material of the kind this page is about, so I am naming it rather than calling it unrelated. There is no link to it on this page. Other pages here do carry links to that shop, the largest being in my guide to starting to read Japanese, and my affiliate disclosure page keeps the full list current. There are no tracked links on this page, so it currently earns me nothing today, and every judgement below was written before any of those approvals and has not been changed since. Two of the three things I recommend hardest below are free, one of them is run by a registered non-profit, and neither pays anybody a commission.
You can follow a conversation. You read news articles. You know several thousand words. And when it is your turn to speak, what comes out is slow, short, and nothing like what is in your head.
Everyone tells you the same thing: you need more speaking practice. That is true and it is not useful, because it does not tell you which of the three separate problems you have.
Here they are, because they need different fixes.
One, retrieval. You know the word and cannot get it out in time. This is a speed problem in a system that works.
Two, the motor problem. Your mouth is running on the wrong clock. This is the one almost nobody names, it is the most Japanese-specific of the three, and it is the reason a learner with excellent grammar can still be hard for me to follow.
Three, fear. Perfectionism, waiting to be ready, the readiness that never arrives.
Most advice treats all three as the same thing. This article is mostly about the second one, because it is invisible from the inside, and because there is a technique that fixes it that was invented in Japan, for Japanese people, and has been quietly working here for thirty years.
First, why you are in this position
I want to defend you for a paragraph, because this gap is not a personal failing and it is not evidence that you studied badly.
Recognition always outruns production. In every language, for every learner, you can understand far more than you can say, and the gap is normal at every level. But in Japanese the gap is unusually wide, and the reason is structural rather than personal.
I have spent this whole site documenting it. No mainstream product makes you produce a pitch pattern. None puts you in the listener half of a conversation and requires a response on time. And the exam that shapes the entire industry contains no speaking component at all, so nothing in the market has an incentive to build production practice.
You did not fail to practise output. You were sold a stack that has almost no output in it.
The part nobody names: your mouth is running the wrong clock
Here is the thing I most want you to take away.
Japanese is a mora-timed language. English is a stress-timed language. These are two different rhythmic systems and your mouth has spent your whole life in one of them.
In English, stressed syllables land at roughly even intervals and everything between them gets squeezed. Comfortable becomes two and a half beats. That compression is automatic and invisible to you.
Japanese does not compress. Every mora takes one beat, and it holds that beat regardless of what is around it:
- 東京 is not two syllables. It is to-o-kyo-o, four beats.
- The small っ is a beat of silence. Not a pause for effect, a beat you must actually wait through.
- ん is a full beat by itself.
- A long vowel is two beats, and the second one is not optional.
Now the consequence, and this is the native-side observation I would most want a learner to hear. When a learner speaks Japanese with English rhythm, the first thing that breaks is not pitch. It is length, and length changes words.
The small っ inside いっぽん is one of these beats, which is why counters break so early for learners. I cover that in my guide to Japanese counter apps.
- おばさん is a middle-aged woman. おばあさん is a grandmother. One beat apart.
- ビル is a building. ビール is beer.
- きた is came. きいた is heard or asked.
If you compress the long vowel because English compresses unstressed vowels, you have not produced an accented version of the word. You have produced a different word, and I have to reconstruct which one you meant from context. That reconstruction is the tax your listener is paying, and it is much heavier than the tax from wrong pitch, which I described in the pitch accent article as a small continuous cost.
And you cannot hear yourself doing it. That is the definition of a motor problem: the error is in the timing of your own production, and your own production is the thing you have the least objective access to.
Which is exactly what shadowing is for
Shadowing means listening to speech through headphones and repeating it out loud at the same time, about a second behind, without waiting for a gap and without reading along.
It is not repetition after the speaker. That is a different and much easier exercise. The whole point is the overlap, because the overlap removes your ability to set your own tempo.
And here is the part I enjoy telling learners of Japanese. Shadowing did not come out of Japanese-as-a-foreign-language teaching. It came out of simultaneous interpreter training, and it entered mainstream language education in Japan in 1992, when a teacher named Ken Tamai encountered it at an interpreter training school in Osaka and brought it into English teaching. It was later formalised in the research literature by Kadota.
So the technique you are being told to use on Japanese was invented in Japan, refined here, and used by an entire generation of Japanese people to attack their own version of exactly your problem: high comprehension, frozen output. My own school English education included it. When I tell you this works, I am not repeating something I read. I am describing the standard equipment.
Why it fixes the clock specifically. When you speak at your own pace, you unconsciously apply your native rhythm, because nothing is stopping you. When you speak on top of somebody else, you cannot. You either match the mora timing or you fall behind and lose the line. The exercise makes the correct rhythm a requirement rather than a suggestion, and after enough repetitions your mouth stops needing the model.
It also happens to train the other two problems as a side effect. Retrieval speeds up because you are producing whole chunks at speed rather than assembling sentences word by word. And fear declines, because you have spent twenty hours making Japanese noises without a single human being judging any of them.
How to shadow without wasting your time
Most people who try this do it wrong in one of four specific ways.
Pick material you already understand. Shadowing is not comprehension practice. If you are decoding meaning you cannot spare the attention for timing. The clip should be one you could follow easily on first listen.
Keep it short and loop it. Twenty to forty seconds, repeated many times in one session, is the format. Working through a long audio once is the most common mistake and it produces almost nothing.
Go through the stages. Listen without speaking. Then mumble along quietly, matching only rhythm. Then shadow at full voice. Then shadow without looking at the text. The last stage is where the benefit is, and skipping to it too early is why people conclude they cannot do this.
Record yourself once a week and listen back. This is unpleasant and it is the only way to hear the thing you cannot hear while producing it. Compare your recording to the model on one specific axis: length. Are your long vowels actually long. Did you hold the small っ. Did you give ん its own beat.
Ten minutes a day of this beats an hour of anything else on the specific problem it solves.
What shadowing does not fix
I want to be honest about the boundary, because this technique gets oversold as a general fluency hack.
Shadowing trains how you say things. It does not train what to say, to whom. It will not teach you to pick the right politeness level for the person in front of you, it will not give you the timing of the listener responses a Japanese conversation requires, and it will not stop you from using an honorific in the wrong direction.
Those are decision problems and shadowing is a motor drill. You will still need a human being for the decisions. What shadowing does is remove the mechanical layer of the difficulty, so that when you do get in front of a human, your attention is free for the decisions instead of being consumed by your own mouth.
Which of the three problems do you actually have
Before you buy anything, spend fifteen minutes finding out what is broken. The three problems look identical from the inside and they have different fixes, so guessing wastes months.
Test one, for retrieval. Take a sentence you would want to say in a real conversation. Alone, with no listener and no time limit, say it out loud in Japanese. If it comes out fine given thirty seconds and falls apart in a conversation, your knowledge is intact and your retrieval speed is the bottleneck. The fix is drilling whole chunks until they come out as single units, which is what Pimsleur and Anki sentence decks are for.
Test two, for the motor problem. Find a short audio clip with a transcript. Read the transcript out loud on your own, recording yourself. Then listen to the model and your recording back to back. You are not listening for accent. You are listening for length: long vowels, small っ, ん. If your version is consistently shorter than the model, or if your recording finishes noticeably earlier, that is the clock, and shadowing is the fix.
Test three, for fear. Can you write the sentence in a chat window without difficulty and not say it to a person. If your written Japanese is dramatically better than your spoken Japanese with the same time available, the bottleneck is not linguistic. The fix is volume of low-stakes speaking, and the lowest stakes available are talking to yourself, then an AI, then a paid tutor whose job is to be patient, in that order.
The middle step has since become a category of its own. I have written a separate comparison of AI apps for practising Japanese conversation, and the short version is that they are excellent for fear and useless for register.
Most people have two of the three. The common combination is the clock plus fear, and they feed each other: you sound wrong to yourself, so you speak less, so the motor skill never develops. That loop is why shadowing works better than its mechanics suggest. Twenty hours of making Japanese noises with nobody listening breaks the loop from both ends at once.
And one thing that is not on the list. If the honest answer is that you cannot say the sentence even with unlimited time, then this article is not your problem yet. That is a knowledge gap and you need vocabulary and grammar, not speaking drills. Speaking practice does not install the words. It gets the installed words out faster.
How I judged these
Three questions.
1. Can you loop a short clip easily? The single most important feature. A tool that makes you scrub a timeline manually will not survive contact with your motivation.
2. Is there real audio at natural speed with a transcript? Both halves matter. Audio without a transcript blocks the early stages, and a transcript without natural-speed audio trains you on speech nobody uses.
3. Does anything give you feedback? Almost nothing does. This is the same hole this site keeps finding, and here there is one genuine exception.
The comparison
| Tool | Loop a clip easily | Natural audio plus transcript | Feedback on your voice | Can it pay me? |
|---|---|---|---|---|
| Language Reactor with Netflix or YouTube | Yes. Line by line, auto pause | Yes, anything you watch | No | No program I could find |
| Speechling | Yes, sentence by sentence | Yes, with native models | Yes. A human coach, ten free recordings a month | No. It is a non-profit |
| JapanesePod101 | Yes, and dialogue is played twice | Yes. Full transcripts, ideal source material | No | Yes |
| Dedicated shadowing apps | Yes, built for exactly this | Varies | Usually not | No |
| Ganbatte Shadowing | Yes. Seventy graded audio lessons from a Japanese university | Yes, built for learners | No | No |
| Rocket Japanese | Partly. Record and compare against a model | Yes | Automated scoring only | Yes, at the highest rate here — I am not in it |
| Pimsleur | Not shadowing. Prompt and respond in gaps | Studio audio | Automated | Yes, and the most per sale |
| A tutor on italki | Not applicable | Not applicable | Yes, the real thing | About ten dollars |
Tool by tool
Language Reactor with Netflix or YouTube — the free setup, again
The same tool I recommended for reading Japanese turns out to be the best free shadowing rig as well, because it gives you exactly the two controls the exercise needs: repeat one subtitle line, and auto pause at the end of it. Pick a forty second stretch of a show you have already watched and loop it.
Using material you enjoy matters more here than in any other exercise, because shadowing is repetitive by design and the failure mode is quitting. Can it pay me? Nothing, and there is no program to join.
Speechling — the only human feedback at this price, and it is a non-profit
You record yourself reading or repeating a sentence, and a human coach listens and sends back a correction. The free tier allows ten recordings a month, and unlimited coaching is about twenty dollars a month. It is run by a registered non-profit that states it charges to cover costs and offers scholarships to people who cannot pay.
I have been writing for months that no product gives feedback on your production. This is the exception, and I want to give it full credit. Ten corrections a month is not much, and it is infinitely more than zero, which is what every other product on this page offers. Send the sentences you find hardest to say, not the ones you already say well.
There is no affiliate program and I would recommend it anyway. Can it pay me? Nothing.
JapanesePod101 — the best source material, if you already subscribe
Not a shadowing tool, but the thing shadowing needs most: a large library of dialogue at natural speed, played twice, with a full transcript and vocabulary attached, graded by level so you can find clips you actually understand.
The second playback and the transcript are what make this workable for the early stages. If you already pay for it, you are sitting on several hundred hours of ideal shadowing material and probably using it only for listening. Can it pay me? Yes.
The dedicated shadowing apps — purpose built, and I have not tested them
The app stores contain a cluster of apps built specifically for this, with modes corresponding to the stages I described, looping, speed control, and progress tracking. There is also Ganbatte Shadowing, a set of seventy graded audio lessons produced by a Japanese university, which is the kind of source I trust for level-appropriate material.
I have not installed these and I am not reviewing their quality, only pointing at the category. If the friction of setting up loops in a video player is what stops you, a purpose-built app removes that friction and that is worth something. Can it pay me? Nothing.
Rocket Japanese — the closest a mainstream course gets
Rocket records your voice and compares it to a native model with automated scoring, which is the nearest thing to feedback in any of the large courses. It is not shadowing, because you speak after the model rather than over it, and the scoring judges whether the recogniser understood you rather than whether your rhythm was right.
I have marked Rocket down in every article where the topic touched production, and this is the one where I can give it real credit: among the big paid courses, it is the only one that listens to you at all. Its programme carries the highest rate on this page, which is why I want the qualification stated as clearly as the praise, and I should be exact about the rest of it: I am not in that programme, so the rate is one I could earn rather than one I am earning. Can it pay me? Nothing today.
Pimsleur — forced output, wrong exercise for this problem
Pimsleur makes you speak on schedule, which is genuinely valuable and is why I keep recommending it for the fear problem. But the structure is prompt and respond in a gap, so you are always setting your own tempo, which is precisely the condition under which your native rhythm reasserts itself.
It is a good product aimed at a different one of the three problems. Can it pay me? Yes, the most per sale, and it is not the tool for the clock.
A tutor — for the decisions, not the drill
Do not spend tutor hours on shadowing. You can shadow alone for free and you cannot get corrected alone at all.
Spend the hour on the things a recording cannot judge: whether the register was right for the person, whether the sentence was natural rather than merely grammatical, whether you sounded like yourself. Can it pay me? About ten dollars.
A four week plan that actually fits in a life
Week one. Ten minutes a day. One forty second clip, the same clip all week, through the four stages. Choose something easy and slightly boring rather than exciting and hard.
Week two. Same routine, new clip. Record yourself on day seven and listen to it once, checking only vowel length and small っ.
Week three. Add Speechling and send your ten free recordings. Use the hardest sentences you have, not the ones you have already mastered.
Week four. Shadow without the transcript from the start. Then book one tutor session and talk. The purpose is not the tutor, it is to find out what changed.
Then keep the ten minutes. This is a motor skill and motor skills decay. Ten minutes forever is a better investment than an hour for a month.
One last thing, from my side of it
I want to describe the moment this pays off, because it is not the one you are expecting.
You are not going to suddenly sound native. What happens is smaller and much more useful: the person you are speaking to stops working. They stop doing the half-second reconstruction of whether you said ビル or ビール. They stop slowing down. The conversation moves at conversation speed, and afterwards they do not think about your Japanese at all.
That is the whole prize, and it is the same prize I described in the article about the listener half of a conversation. Fluency, from the other side of the table, is not brilliance. It is the absence of effort in the person listening to you.
And the technique that gets you there was built in this country, for interpreters, and then handed to schoolchildren. It works. It is free. It is ten minutes. The reason nobody sells it to you is that there is nothing to sell.