What this tool does
The hard part of reading Japanese is not turning 漢字 into かんじ. It is knowing which reading to use, and that answer is not in the kanji. 行 has three on-readings and three kun-readings and none of them is the right one for 行く, where the kana that follows decides it. 今日 is きょう and not こんにち, 一人 is ひとり and not いちじん, 大人 is おとな and not だいにん, and a tool that reads the kanji one at a time gets every one of these confidently wrong. This one reads words, then uses the okurigana to pick between the readings, and tells you when it had to guess.
How it works
A word is tried first, and it is the step that does the real work. There is no rule that turns 今日 into きょう or 一本 into いっぽん, so a list of the common words whose reading is not the sum of their kanji has to exist, and it is consulted before anything else. That list is drawn from JMdict and the words in it are the ones a reader actually meets, so 行く is いく rather than い.く and 学校 is がっこう rather than がくこう. Surnames come before it, because 田中 is たなか and not でんちゅう and no rule on earth derives that — it is a fact about a family, not about a character.
Whatever is left is a run of kanji, and the kana directly after that run is what settles the reading. With no kana following it, the run is a compound and every kanji takes its on-reading, which is why 東京 is とうきょう and 大切 is たいせつ. With kana following it, the last kanji is a verb or an adjective stem and takes a kun-reading instead, and the kana after it is left as written rather than rebuilt. That is why 読んで comes out as よんで without anything needing to know about rendaku, and why 生まれる picks うまれる out of fourteen readings of 生: the dictionary okurigana まれる lines up with the まれる that was written, and い.きる does not.
Kana after a kanji are sometimes not okurigana at all. が, を, に, は and the rest are particles, and です and だ and ので are copulas and connectors, and in front of those the kanji is a noun and wants its bare kun-reading. Without that check 猫です comes out as ビョウです, since 猫 has no dotted kun-reading to match です against, and 勉強でした comes out as ベンつよでした. One verb is handled by hand because no dictionary marks it: 来ます is きます and not くます, and that change happens to the verb rather than to the kanji.
Every reading is then labelled with how it was arrived at, which is the part that matters more than it sounds. A reading from a list is certain. A reading the okurigata chose is likely. A reading a rule had to take with nothing to confirm it is marked as inferred and underlined on the page, and a kanji the dictionary has no reading for is marked with a question mark and counted rather than given an invented answer. A study tool that is quietly wrong teaches the wrong thing, and the student has no way to tell. One that says it is unsure can be checked.
Worked example
A paragraph from a lesson has to be read tonight, and reading it one character at a time produces a page of plausible kana that is mostly not what the text says.
- Paste the paragraph in and the reading appears above each kanji rather than beside it
- 今日は comes out as きょう and not こんにち, because the word is looked up before its kanji are
- 行きます is いきます, with く taken from the word and きます left as the okurigana it is
- 田中さん is たなかさん, from the surname list, rather than でんちゅうさん from reading the kanji
- A given name the tool has no entry for is underlined and counted as unconfirmed, not silently guessed
The paragraph is readable, the kana can be copied as text, the romaji is there to check a reading against, and the one place the tool is unsure is visible rather than buried.
Accuracy and limitations
- Kanji outside the standard list are reported as unknown rather than read. That covers old forms like 學 and 圓, rare characters, and most given names, and it is a deliberate choice: a question mark is more useful to a student than a confident invention.
- A given name is a harder problem than a surname. Names are read, not composed, and there is no rule for them at all, so a name outside the list is underlined as unconfirmed and the reading shown is the ordinary compound reading, which is often wrong.
- A long vowel is not marked in the romaji, which is what Hepburn does and what dictionaries print, so ビール comes out as biru and ビール's short form ビ is also bi. It cannot be helped without a macron, and doubling the vowel instead produces biiru, which no dictionary prints either.
- The romaji of a word read from kanji contains the long vowel written out, because hiragana writes a long vowel as a second vowel. 東京 is therefore toukyou and not Tokyo, and no amount of post-processing on the kana can recover the length.
Frequently asked questions
- What is furigana and why do I need it?
- Furigana is the small phonetic gloss written above kanji in Japanese children's books, manga and textbooks, and it exists because a kanji on its own does not tell you how to read it. This tool puts that gloss over any text you paste, using the reading of the word the kanji is actually in rather than the reading of the character on its own, which is the difference between きょう and こんにち for 今日.
- Why does it say a reading is unconfirmed?
- Because some readings cannot be worked out from rules, and a tool that prints all of them with equal confidence teaches you the wrong ones along with the right ones. Readings taken from the word and surname lists are marked as certain, readings the following kana chose are marked as likely, and a reading a rule had to take with nothing to confirm it is underlined and counted. Check those against a dictionary, and the rest you can rely on.
- Does it work offline and is my text uploaded?
- Yes to both. The kanji readings, the word list, the surname list and the romaji tables are all bundled into the page, so the reading happens in your browser tab with no network call at all. Nothing is sent, stored or logged, which is the only way a tool like this can be trusted with a draft you have not published yet.
- What romaji style does it use?
- Hepburn by default, which is what English-language dictionaries print: しゃ as sha, ふ as fu, じ as ji. Nihon-shiki is there as a second option, the classical spelling that writes し as si, ち as ti, つ as tu and ふ as hu, which is what linguistics papers and older dictionaries use. A long vowel is not marked under either, so ビール is biru rather than biiru.
- Why is my name or an old kanji not read?
- Because there is no reading in the dictionary for it. The kanji table covers the standard school list, and the surname list covers the hundred commonest Japanese family names, which between them cover most text. A given name or an old-form kanji like 學 falls outside both, and rather than invent something it is marked with a question mark and counted, so you can look it up yourself and know which part of the page to distrust.
- Can I use it to find the kanji I have not learned yet?
- Yes. The grade filter marks every kanji above the school grade you pick, so on a grade 2 setting the harder characters in a passage are highlighted in the reading. That turns the tool from a converter into a way of seeing what a text is leaning on, which is usually the actual thing a learner wants to know about it.