Jump to content

User:TJones (WMF)/Notes/Japan Eset OK en I Zat Ion—Kuromoji, ICU, and Sudachi

From mediawiki.org

January/February 2025 — See TJones_(WMF)/Notes for other projects. See also T318269. For help with the technical jargon used in the Analysis Chain Analysis, check out the Language Analysis section of the Search Glossary.

Fortunately, Japanese tokenization (and other analysis) in practice is generally better than Japan eset ok en I zat ion! Nonetheless, there are... quirks.

Background

[edit]

I looked at the Japanese analysis package Kuromoji, back in 2017, but eventually decided not to deploy it. Now, more than 7 years later, I was hoping that Kuromoji might have improved, as well as my ability to understand its nuances, and my confidence to make bold(ish) changes to improve performance (working with languages you can't even kinda read can be scary!).

I don't think Kuromoji has changed much—I have read complaints about a lack of updates to the dictionary—but I started this re-review because I hoped that we could still get more out of it by customizing and configuring it.

Also, having gained more experience with the ICU tokenizer over the years, and having made it the default tokenizer everywhere it was practical to do so, I decided to include ICU tokenization as a sort of baseline in the eventual analysis by Japanese speakers.

During the initial tokenization review phase, one of our reviewers suggested investigation Sudachi, which is a different library that supports tokenization and other Japanese NLP. Compared to Kuromoji, its tokenization dictionary has been upgrade more often and more recently. Since I was able to crib some analysis from my past experience with Kuromoji, and because Sudachi is even quirkier than Kuromoji, adding Sudachi to the mix more than doubled the effort needed here—though it looks like it will have been worth it in the end!

I've also changed my approach to evaluation. The large-scale automated tokenization evaluation used last time is nice in theory, but in practice some answers are ambiguous, and sometimes there are better answers and answers that are less good but still very acceptable. (As an analogous example in English, you could argue that high school or football are better as one token each, or better split into two tokens, but both answers are reasonable—unlike Japan eset ok en I zat ion, which is clearly just wrong all over the place.) Smaller-than-ideal tokens (at the extreme, like CJK unigrams from the standard analyzer) are better than longer-than-ideal tokens (like splitting CJK tokens only on spaces and punctuation). Consistency—which is much harder to test for, alas—is the best.

Mini-Glossary of Writing Terms

[edit]

Japanese writing is made up of several scripts.

Hiragana and katakana (collectively kana) are native syllabaries. To my non–Japanese-reading eye, hiragana are often collectively rounder (あ, ち, よ) and katakana are more angular (ク, ヒ, コ), though individual symbols, especially hiragana, may not follow the general trend, and with fancy fonts all bets are off. (And へ and ヘ are often identical anyway—though it depends on the font.)

Kanji are the logographic symbols adapted from Chinese. Chinese characters are called hanzi in Chinese, hanja in Korean, and sometimes Han characters (sometimes further shortened to just Han) in English. Many Unicode characters use "CJK Compatibility Ideograph" in their descriptions because they are used by Chinese, Japanese, and Korean, but these are usually characters originally from Chinese, so CJK is also common. N.B.: "Han" and "CJK" are the terms I use most when talking about Chinese characters in the (Unicodian) abstract, so if I slip up, and mention Han or CJK (or hanzi), I generally mean kanji.

Latin and katakana characters come in fullwidth and halfwidth varieties. In monospaced fonts, typical Japanese characters are twice as wide as typical Latin characters. However, there are also fullwidth Latin characters and halfwidth カタカナ (katakana) characters (and various varieties of punctuation in either) which are used for historical technical reasons or general typographic/design reasons.

Data

[edit]

I sampled 5K Japanese Wikipedia articles (which are longer), plus 10K each from the shorter texts of Japanese Wikipedia queries, Japanese Wiktionary entries, and Japanese Wiktionary queries. These samples are for generally testing the analysis chains, not limited to tokenization.

I also created an intentionally heavily multilingual, multi-script sample, including one medium-long line of text from each of almost 120 wikis, plus a number of words (mostly names of languages taken from the SiteMatrix) in non-kana, non-kanji scripts, sampling a few extra with less common diacritics. In the 2017 analysis of Kuromoji, the default config simply deleted text in many scripts, so this small corpus helps us monitor that.

For tokenization, I intended to create corpora with 50 random sentences/query snippets each from the four Wikipedia/Wiktionary article/query corpora. Somehow, I only got 25 sentences from the Wiktionary entry corpus, which I didn't notice until after getting reviews. I also sampled 15 of the longest sentences from the Wikipedia article sample (since they seem more likely to be complex and pose tokenization/parsing problems). I also included a mere 8 samples that happened to generate particularly long tokens, to see how my "optimized" config handled them, and to compare between tokenizers.

Later, when Sudachi was added to the mix, I found another 7 sentences/snippets that had "interesting" results from Sudachi—particularly long tokens using the default config, and mixed-script tokens, to see if they seemed generally right or wrong, and whether the input text was treated the same by the other tokenizers. (i.e., seeing if the others get them wrong, to figure out whether Sudachi is bad at these snippets, or they are just really hard snippets and all the tokenizers have trouble. This is similar to the long sentences and long token snippets above.)

CJK Analyzer Analysis

[edit]

The CJK analyzer—which is the default from Elasticsearch/Lucene for Chinese, Japanese, and Korean—creates overlapping bigram tokens, normalizes fullwidth and halfwidth characters, lowercases non-CJK tokens, and uses some English stop words.

We currently use the CJK analyzer only for Japanese at this point, having upgraded to more language-specific analyzers for Chinese and Korean.

Our version has our usual upgrades (including ICU normalization, processing acronyms, homoglyphs, and camelCase, etc.). We do see acronym, homoglyph, and camelCase tokens in the samples, as well as very long Thai tokens.

The CJK analyzer provides a baseline for processing non-Japanese (i.e., not kana or kanji) text. For Japanese tokens, the bigrams are not directly comparable to any decent attempt at proper tokenization.

Kuromoji

[edit]

Note: We're using the Kuromoji analyzer for Elastic 7.10.

Monolithic Kuromoji

[edit]

Observations on the monolithic Kuromoji analyzer (i.e., configured simply as analyzer type kuromoji):

  • It filters non-Japanese scripts, other than Latin, Cyrillic, and Greek. Specifically, Ahom, Alchemical Symbols, Arabic, Armenian, Balinese, Bamum, Batak, Bengali, Bhaiksuki, Bopomofo, Buginese, Canadian Syllabics, Cham, Cherokee, Coptic, Cuneiform, Cypriot, Deseret, Devanagari, Duployan, Egyptian Hieroglyphs, Ethiopic, Georgian, Glagolitic, Gothic, Gujarati, Gurmukhi, Hangul, Hanunoo, Hebrew, IPA extensions, Javanese, Kaithi, Kannada, Kayah Li, Khmer, Lao, Lepcha, Limbu, Linear A, Linear B, Malayalam, Math Letters (Greek & Latin), Meetei Mayek, Mende Kikakui, Meroitic Cursive, Meroitic Hieroglyphs, Mongolian, Myanmar, New Tai Lue, Nko, Ol Chikim, Old Persian, Old Turkic, Oriya, Runic, Siddham, Sinhala, Sora Sompeng, Sundanese, Syloti Nagri, Syriac, Tagalog, Tai Le, Tai Tham, Tai Viet, Tamil, Telugu, Thaana, Thai, Tibetan, Tifinagh, Tirhuta, Ugaritic, Vai, and Yi are all stripped.
    • That's a long and fairly nerdy list of scripts (including some ancient and extinct), but very few come solely from my multilingual sample. A random sample of Japanese Wiktionary has a lot of non-Japanese scripts in it! In our case, the plain field would pick up the slack and index these, but the text field should be pulling its weight, too!
    • Greek words are treated very weirdly and somewhat inconsistently. I've figured out more of it than in 2017, but some mysteries still remain. If a Greek word is 2-5 characters long, it gets split into individual letters (στη → σ + τ + η). If it is 6-8 letters long, it is left whole. If it is 11 letters or longer, it gets overlapping tokens, including the whole word, each letter up to the last 7 as single letters, and the last 7 as one token (βικιπαίδεια → β + βικιπαίδεια + ι + κ + ι + παίδεια). Words that are 9 or 10 letters long can either be left whole, or split up like 11+-letter words; I think it is related to the first letter of the word, or the specific letters in the word, but the exact pattern is unclear and not currently worth puzzling out.
  • The Kuromoji tokenizer splits on ...
    • ... apostrophes, periods, and colons. (All reasonable.)
    • ... combining characters: Ма́харов → ма + ́ + харов. (Not reasonable!)
    • ... the function application character (U+2061), which is probably a good thing, but is definitely a hardcore nerdy concern.
    • ... halfwidth katakana middle dots (・ U+FF65), but not fullwidth katakana middle dots (・ U+30FB). (Update: sometimes. It's not entirely predictable in all contexts.)
    • ... script changes, so ABCDАБГД is tokenized as ABCD + АБГД, and words with homoglyphs, like chocоlate, get split up, too: choc + о + late
  • Numbers are split off from text around them, in both CJK and non-CJK text.
    • Fullwidth or mixed fullwidth and halfwidth numbers, especially starting with fullwidth 1 can be broken up in weird ways. 1933 → 1 + 9 + 3 + 3, 1933 → 1933, 1933 → 1 + 933.

The monolithic Kuromoji analyzer provides no folding, but the docs suggest you use the ICU normalizer character filter to normalize before tokenization! That's something I missed last time.

I've noticed that in certain contexts, things that shouldn't matter, like leading or trailing spaces, can change how text is parsed. Here's one example where adding an initial space dramatically changes the tokenization:

  • "なかむらつね" → な + かむら + つ + ね
  • " なかむらつね" → なか + むら + つね

This is frustrating because it occasionally leads to seemingly random changes.

I found the example above because I had two versions of the analyzer that should have had identical output, but they differed on a very small number of tokens. Because of some Sudachi-related issues (foreshadowing!) I updated the script that feeds input to the analyzer, which changed the precise chunks that text was bundled up into (i.e., multiple lines joined together with spaces to save on >90% of API calls, which can be really slow). As best as I can figure, in one batch, the "なかむらつね" was at the beginning of batch (no space before it), and in another it was bundled with the text before it, separated by a space.

Unpacked Kuromoji

[edit]

I unpacked Kuromoji according to the docs. I must've been a little less careful last time, or something has changed, because there were some differences between monolithic and unpacked Kuromoji last time. (Probably from me not quite disabling all of our automatic analysis upgrades.) This time the results were identical on all samples.

However, with the unpacked analyzer you can use the "explain" feature and get more details on what each stage of the analyzer does, and see some internal information. It turns out that kuromoji_part_of_speech eats tokens labelled "symbol-misc", which includes emoji, combining accents (which are split off by the tokenizer), IPA symbols (like ʰ ː ɪ ̯) and all the foreign scripts that disappeared.

Automatic Upgrades

[edit]

Re-enabling the automatic upgrades gave the expected changes over baseline Kuromoji, mostly affecting non-CJK text: acronyms, camelCase, ICU normalization. There are a few changes to CJK tokens where the parsing is context dependent, and the context changed; for example, replacing parens with spaces (by word_break_helper).

Default Kuromoji plus Upgrades vs CJK
[edit]

Comparing Kuromoji with automatic upgrades to prod CJK (with the same upgrades), highlights the differences in tokenization and some of the edge cases in how Kuromoji deals with non-Japanese text.

The biggest change is in how many two-character tokens had their counts decrease. For example, the CJK tokenizer found 2421 instances of "かし" (anywhere those two characters appeared next to each other), but Kuromoji only found 14 instances that it thought were actually the word かし. The total number of tokens parsed by CJK was 6.6M, while Kuromoji only had 2.8M—those overlapping bigrams really add up!

Tokens in most foreign scripts disappeared. Non-Japanese mixed-script tokens (e.g., Latin/Cyrillic/Greek homoglyphs, stylistic mixtures like Justiφ's, etc.) and letter+number tokens all disappeared.

Latin, Greek, and Cyrillic tokens have all of their various diacritics because I have not yet enabled ASCII folding. The reasoning behind adding ASCII folding to CJK is the same as adding it to Kuromoji—and then allowing it to upgrade to Japanese-specific ICU folding when possible.

Fullwidth Number Fix

[edit]

Adding a character filter to map fullwidth numbers (0123) to halfwidth numbers (0123) solves all the problems of inconsistent number parsing—e.g., 1933 → 1 + 9 + 3 + 3.

However, it also sometimes loses parsing of months, e.g. 1月 ("first month", i.e., January) gets broken up to 1 + 月. On the other hand, that's what already happens with halfwidth-numbered months: 1月 → 1 + 月. The halfwidth variants are much more common anyway, so that's fine.

ICU Normalizer Char Filter

[edit]

We normally upgrade our lowercase token filter (i.e., after tokenization) to the ICU normalization token filter. However, there is also an ICU normalization character filter (i.e., before tokenization).

Elastic recommends using it with Kuromoji, specifically to prevent problems with tokenizing fullwidth Latin. (of can be tokenized as o + f in some contexts, for example.) Somehow I missed that recommendation last time, too. (Maybe they upgraded the docs when I wasn't looking!)

ICU normalizer char filter observations:

  • It definitely changes the parsing of some strings of Japanese text.
    • Zero-width spaces are (or at least can be?) treated as a hard word breaks by the kuromoji_tokenizer, but the icu_normalizer removes them. Not sure if this bad (many are put there on purpose) or good (many end up in place by accident, some people won't use them, inconsistencies in parsing based on invisibles is always a little sketchy).
  • Some characters that get regularized (e.g., ℃, ℆, ㉅, ㊎, ㊩, ㎡, ㏆, ㏢, ㏫, ㌆, ㌳, ㌶, ㍃) would have been filtered by the tokenizer otherwise. Some would have been eaten by the kuromoji_part_of_speech filter, like 🄩.
  • We get arguably better processing of characters that get normalized into multiple characters with punctuation (e.g., ⑾, ⒓ → "(11)", "12."). Normalizing before tokenizing means plain numbers can match.
  • Fixes some tokenization inconsistencies, e.g. ⼎ (U+2F0E) is dropped by the kuromoji_tokenizer, but 冫 (U+51AB) is not. Technically, one is a radical (which is—or historically was—an element incorporated into other characters) and the other is a full-fledged character. They have different items in Japanese Wiktionary (but a combined one on English Wiktionary). This is a consistent patterns across basic radicals that are also full characters (e.g., 冖 vs ⼍, ⼀ vs 一, ⼁ vs 丨, ⼂ vs 丶, ⼃ vs 丿, ⼄ vs 乙, ⼅ vs 亅).
  • Fixes problem with fullwidth Latin words (like "of") sometimes getting split up letter-by-letter.
  • Everything gets lowercased. This seems ok as long as icu_normalizer is the last character filter (or at least after relevant other char filters), so that, for example, camelCase_splitter can still do its thing.
  • Offset ranges can change to include invisibles the tokenizer would break on, which probably doesn't matter most of the time.
  • Halfwidth katakana middle dots are converted to fullwidth katakana middle dots, which will at least treat them consistently on the halfwidth/fullwidth dimension.

I did find an odd bug in the char filter version of icu_normalizer: the reported token in "explain" output is truncated at the first changed letter (e.g., input στις → output στι (not στισ); even when just lowercasing: abCde → ab (not abcde)). The internal token passed on to next filter is actually fine, so this seems to be a bug in bookkeeping rather than in processing.

Unpacked and Customized POS Filter

[edit]

I unpacked the part-of-speech (POS) filter so that I could modify it. Filtering particles usually makes sense, and that's most of the filters, so we probably want it to continue using the those filters in Japanese.

I tried turning off some of the individual non-particle POS filters to see what they apply to. (They have names in both Japanese and English—which is much appreciated!)

  • If we stop filtering "記号-一般 / symbol-misc":
    • All the "foreign" scripts other than Latin, Cyrillic, and Greek reappear!
      • Some less common Latin/IPA (ɗ, ʔ, ɪ, ꝡ, Ə, ɢ), Cyrillic (ԥ, ꙗ), and Greek (ἅ, ἐ, ἑ, ῖ, 𝈁, 𝈆, 𝈦, 𝈰, 𝈻) characters reappear, too.
    • Individual tokens of naked combining diacritics reappear: ́ ̄ ̤ ̯ .
    • Some actually misc symbols (plus emoticons, flags) reappear: , 😌, 👀, 🌘, 🍳, 𝌞, 𝄞, 🇵🇼.
    • Some (high-number and very high-number) CJK symbols reappear: 鿃, 𘭂, 𘭏, 𘰮, 𘰼, 𘲄, 𘲅, 𘳂, 𱍔, 𱏗, 𱗎, 𱚣, 𱚥, 𱝧, 𱝨, 𱝩, 𱟝, 𱤏, 𱥺, 𱦸, 𱫃, 𱭐, 𱰿, 𱲊, 𱴌, 𲂅, 𲂇, 𲅜, 𲉊, 𬚩, 𱁬.
    • 〇 (ideographic number zero) is treated as "symbol-misc" when it appears by itself (i.e., surrounded by spaces), but as a "noun-numeric" when around other numeric symbols (e.g., 一 or 二)... though if it is the last character parsed in a string, it becomes a "symbol-misc" again. The symbol-misc version comes back.
    • Katakana, hiragana, and kana iteration marks reappear: ヽ, ゝ, ゞ, 々.
    • Some katakana symbols reappear: ノ.
      • This is the syllable "no", its hiragana analog, の, was never filtered; two of them (ノノ) is tagged as "noun-common", and three (ノノノ) is "noun-proper-organization". (Actually, it's more complex and context-dependent than that, but we'll stop there.)
    • Some ancient hentaigana symbols reappear: 𛁒𛁳𛁹𛁂𛀽𛁣𛁑𛁺𛁜𛁪, 𛀙, 𛀚, 𛀻, 𛀾, 𛁓, 𛁔, 𛁿, 𛂧, 𛃵.
  • If we stop filtering "その他-間投 / other-interjection":
    • Only よ ("yo") and ァ (small "a") are tagged as interjections in my sample.
      • ァ (halfwidth small "a") is not tagged as an interjection, but it gets regularized by icu_normalizer to ァ (small "a"), which is. ア ("a") is not tagged as an interjection.
      • Neither よ ("yo") nor ょ (small "yo") are tagged as an interjections alone, so よ ("yo") at least requires some other context to be tagged that way.
  • If we stop filtering "フィラー / filler":
    • One new token, なんか, is allowed through. Several other tokens (あ, あの, あー, うん, え, えと, えー, ま) have from a slightly increased (2 to 5) to a rather increased (46 to 308) number of instances. This reiterates that the POS tagging is context dependent.
  • Stopping filtering "記号 / symbol" or "非言語音 / non-verbal" had no effect on my samples.

I disabled symbol-misc, other-interjection, and filler filtering before proceeding.

Unneeded Filters

[edit]

The ICU normalization char filter seems helpful/necessary. When we have it, do we need the ICU token filter, too? Apparently not! Removing the lowercase filter it upgrades from created no diffs.

According to the docs, the cjk_width token filter seems to only normalize fullwidth and halfwidth variants. ICU normalization should do that, too. Do we need both? Apparently not! Removing it created no diffs.

More Tweaks

[edit]
  • Combo Filter: From the Nori (Korean) config, I copied the idea of a character filter that filters combining characters (a.k.a., a "combining character character filter filter", but that's confusing... maybe it's better in German?). The new kuromoji_combo_filter is a bit more expansive than its Nori counterpart and strips U+300 through U+362, which is most of the Unicode block of basic combining diacritical marks.
    • It's particularly helpful on Cyrillic/Russian names, which tend to have an acute accent to show stress. So instead of Ма́харов → ма + ́ + харов, we get Ма́харов → махаров.
    • There are also Greek, IPA, and other uncommonly accented Latin characters that no longer cause splits—particularly in detailed phonetic information from Wiktionary.
    • It also gets rid of all those naked combining character tokens.
  • ICU Folding: ICU folding (and if ICU is unavailable, ASCII folding) makes sense with the same Japanese-specific exceptions used by the Japanese version of the CJK analyzer.
    • The biggest impact is accented Latin text, just because there is more of it in Japanese wikis than other scripts, and thus more chance of a new collision in my samples (e.g., abōminōsam and abominosam). Still, there are plenty of regularizations in other non-CJK tokens, and matches in Greek, Cyrillic, Hebrew, Arabic, and a few Indic scripts.
    • Even more of the naked modifier and combining character tokens disappear.

Final Custom Kuromoji Config

[edit]

The final wiki-optimized config for Kuromoji (version ES 7.10) is (not necessarily listed in analyzer order):

  • A char filter to remove the most common combining diacritics for alphabetic scripts, which otherwise get split by the Kuromoji tokenizer.
  • The icu_normalizer char filter (rather than the usual token filter) if ICU is available; otherwise, a simple char filter to remap fullwidth numbers to their halfwidth versions, and cjk_width to clean up any other fullwidth/halfwidth variants after tokenization.
  • The baseline kuromoji_tokenizer, and kuromoji_baseform and kuromoji_stemmer token filters, Japanese stop words, and lowercasing.
    • We can remove lowercasing if ICU is available as the icu_normalizer char filter takes care of it.
  • A customized Kuromoji part-of-speech filter to allow foreign scripts and a few other terms through.
  • ASCII folding for non-CJK diacritics, upgraded to Japanese-specific ICU folding if ICU is available.

Sudachi

[edit]

Note: We're using the Sudachi analyzer v3.0.0 for Elastic 7.10.2.

I Can't Explain

[edit]

I discovered that using the explain flag in the Analyzer API causes an exception when the Sudachi tokenizer is enabled. This is a known issue with the "morpheme" attribute.

It's pretty annoying because internal info like the part-of-speech tags would be nice to see, so you can more easily filter on the right ones.

Fortunately we can still get at the non-Sudachi attributes by specifying an explicit list of attributes we want to see in the explain output. That breaks up the big "Sudachi analyzer" black box into a smaller "Sudachi filter" black boxes—this will be a recurring theme.

(I've learned that Sudachi v3.2.0 may fix the explain problem. If only I had found out sooner!)

Monolithic and/or Unpacked Sudachi

[edit]

Rather than initially report on the monolithic Sudachi analyzer's behavior, I want to be able to ascribe the behavior to specific filters when possible, in case we want to address it or try to modify it.

After running a monolithic baseline, I tried to unpack Sudachi according to the docs. However, it gave different results!

The unpacked analyzer did not filter a lot of tokens that the monolithic analyzer did. It seems that the default configuration for the sudachi_part_of_speech filter wasn't doing anything. I don't know if the unpacking info was wrong, I did something wrong, the default configuration file was missing, misplaced, had bad permissions, or what.

I was able to reconstruct the monolithic behavior with no diffs by explicitly listing all of the default part-of-speech tag filters, which are listed in Sudachi's stoptags.txt file. That's a fine solution for an interim step since I will probably want to customize that list anyway.

However, not having the explain feature was still pretty annoying.

Observations
[edit]

Below are some observations on the Sudachi analyzer (i.e., either monolithic, configured simply as analyzer type sudachi, or unpacked to be equivalent as far as testing on my corpora can tell). These are issues mostly other than straightforward tokenizing of Japanese text.

The Good:

  • Does some normalization in sudachi_baseform: fullwidth (1 → 1, g → g), halfwidth (カ → カ), encircled (① → 1, ⓓ → d, ㋭ → ホ), superscript (¹ → 1, ʲ → j), subscript (₁ → 1), characters with punctuation (⒓ → 12), "telegraph" symbols (㏫ → 12), script variants (ℓ → l), math variants (𝒏 → n, 𝓉 → t), "square" units (㌆ → ウォン), multi-char symbols (℆ → c + u, ㎡ → m2), other misc variants (〸 → 十)
    • It doesn't matter in practice, but I find it funny that it lowercases Cherokee letters. As I've mentioned elsewhere, the lowercase letters are rarely used, and uppercase is the default. The icu_normalizer, which we will add eventually, uppercases them back to where they started from.

The Bad:

  • Makes some really long tokens:
    • ラボグロウンダイヤモンドオリジナルジュエリーオンラインショップ
    • ヤマハミュージックエンタテインメントホールディングス
    • サイバネティックスパワードスーツクォンタムテクター
    • MS&ADインシュアランスグループホールディングス
    • These tend to be kana, and there are other config settings that may be able to help with this.
    • I learned that = (fullwidth equals sign) can used as a dash, as in ブレスト=リトフスク ("Brest-Litovsk", the former name of a city in Belarus). That's not wrong, but in the same way I want "Brest-Litovsk" to be two tokens, ブレスト=リトフスク is better as two tokens... especially in this case where the current name (and likely search query) is just ブレスト ("Brest").
  • Filters a lot of things (generally done by sudachi_part_of_speech, though the tokenizer makes some choices, too):
    • High-numbered kanji (from CJK Unified Ideographs Ext B–G, but not the "CJK Unified Ideographs" block, or the Ext A block) are filtered, at least as individual characters.
    • A few specific individual hiragana are filtered: ゃ • ゅ • ゆ • ゑ
      • ゑ is an obsolete kana, which is probably different from the others, which can be used as particles.
    • Scripts other than Chinese/Japanese, Latin, Greek, and Cyrillic—pretty much the same list as Kuromoji.
    • Some less common Latin and Cyrillic letters are filtered, like Latin ə, ɔ, and ʋ, and Cyrillic ԕ and ԛ, and words are split on these letters.
      • Others letters are only disallowed as single letters, like the more common Latin å and Cyrillic з. (This is pretty weird!)
    • Greek single letters are generally filtered. Of the standard modern alphabet, only ς and Ω/ω are not filtered as single letters. (Also pretty weird!) Longer words of unaccented Greek are left alone.
    • Names of Greek letters in English and katakana and names of English letters in katakana are also filtered, like nu, rho, オメガ ("omega"), アルファ ("alpha"), イオタ ("iota"), イプシロン (alternate form of "epsilon"), エックス (ekkusu / "X"), ゼット (zetto / "Z"), ダブリュー (daburyū / "W").
    • Some emoji (🌸, 🍎) are ignored by the sudachi_tokenizer, others (🥹, 🫥) are stripped by sudachi_part_of_speech.
    • Some less common punctuation symbols (‼, ⁉) are ignored by the tokenizer. Single dashes are ignored, but multiple dashes (--, ---, ----, etc.) are tokenized, normalized to an em dash (—) by sudachi_baseform (even 45 dashes in a row!) and then filtered by sudachi_posfilter. (What a ride!)
  • Splits on character set changes (like the ICU tokenizer).. like our old friend chocоlate → choc + о + late, or gγг → g + γ + г.
  • Splits on "character set changes" that aren't really different character sets:
    • Combining accents: Ма́харов → ма + ́ + харов
    • Uncommon Latin/IPA letters: dɔg → d + ɔ + g
    • Accented Greek letters: δικεῖν → δικε + ῖ + ν
  • Capitalization of output tokens is all over the place: this is a man → This + is + a + MAN; love story → LOVE STORY

The Ugly:

  • It can be really arbitrarily inconsistent in the way it parses and normalizes some characters.
    • is sometimes normalized as m2 and sometimes m + 2, depending on the context of the numerals and spaces around it.
      • 0㎡ → 0 + m2 if not followed by a space, 0 + m + 2 if followed by a space. So, "0㎡ 0㎡" → 0 + m2 + 0 + m + 2.
      • 8㎡ → 8 + m2 (space or not), but 08㎡ → 08 + m + 2 if followed by a space.
      • 3㎡ → 3M + 2.
      • The internals are kind of wild. "0㎡ " (with a space) → 0 + ㎡ + "" (empty string) by the tokenizer, but then sudachi_baseform converts that to 0 + m + 2. The lack of explain detail is really annoying.
    • There are tokens normalized as An & AN, and cm & CM. Furthermore, both An & AN tokens could be in the original text as AN, An, or an, and both cm and CM tokens could be in the text as cm or CM. I was able to find examples and 9CM/9cm → cm, 5CM/5cm → CM, though any CM at the very end of the string is capitalized, regardless of the number before it. This is easily fixed with a lowercase/icu_normalizer filter after the Sudachi components, but it's weird......
    • rock 'n' roll → "rock 'n' roll", but rock'n'roll → rock'n' + Roll.
    • A couple of notably long tokens include アゼルバイジャン・ソビエト社会主義共和国 and NBCユニバーサル・エンターテイメントジャパン (both with a middle dot, ・). If we remove the middle dot, the first remains a single token, though it doesn't match the one with the dot. Without the dot, the second gets broken up into three tokens.
      • Fortunately, there are some settings we can change to break up these longer tokens.
  • Some tokens have stray final spaces: "coma ", "optical ", and "engel ". This seems to require a space after the word (as opposed to punctuation or it being the end of the string), but it is somewhat context dependent in complex ways. "engel angel" → "engel " + "angel", but "engel engel" → "engel" + "engel".
  • Lots of Latin multi-word tokens that are generally English, though there are also some French phrases that I would consider as having been borrowed into English—though some, like prêt-à-porter, are pretty niche.
    • Understandable examples: e-mail, ping-pong (but not ping pong), café au lait and cafe au lait, follow-up (but not follow up), low profile (but not low-profile), SBI Holdings, Inc. (but not without that exact punctuation), Organization for Economic Cooperation and Development (that's a long one!), Platform as a Service.
    • Less understandable examples: excuse me, over time, the floor, want to, would you, once more.
    • Obvious problem examples: Mark Twain will not match Twain, Mark; neither United States nor America will match United States of America; guitar will not match steel guitar.
  • Lots of tokens with punctuation in them: cf., vol., vs., /m, #9, @Tokyo, MS&AD, P&G, GOOD LUCK!!
    • These are distinct from cf, vol, and vs, which is pretty unexpected.
    • GOOD LUCK!! is a single token. Zero or one exclamation point gets tokenized as GOOD LUCK, while two or more would be indexed as GOOD LUCK!!
      • Stand Up, STAND UP!, and STAND UP!! are all distinct—and capitalized as shown here.
      • "(財)", with fullwidth parens, and "(財)" with halfwidth parens are both filtered by the default configuration for sudachi_part_of_speech, but "財" without parens is kept. It means "wealth" or "treasure" so the meaning doesn't somehow make it make more sense.
      • This category includes some long, mixed-script tokens, too: Yahoo!オークション, M&Aキャピタルパートナーズ, MS&ADインシュアランスグループホールディングス.
    • The lack of explain was pretty annoying.
  • There's a common pattern where kanji followed without spaces by kana in parens (regular/halfwidth or fullwidth parens) results in a single token: 世(ごせ) → 世. The sudachi_tokenizer parses it as a single token, and the sudachi_baseform filter removes the parenthesized part.
    • Not everything that fits the basic pattern as I described it gets tokenized in one piece. There might be a limit of four kana in the parens. But there is definitely no matching of kanji to kana, so non-standardly formatted non-pronunciation information could get ignored, too.
    • It's not unmotivated because the parenthesized part can be there to provide pronunciation information. But what if an editor wants to search for that bit specifically?
    • One of my samples had three source tokens for the indexed token "金": 金(おかね), 金(かね), and 金(キン). The first one is an error, as the context is "お 金(おかね)", which probably has an extra space (should probably be "お金(おかね)", which is indexed as "お金").
    • Another interesting example is "金(かね)持(も)ち" which, despite the spacing, has no spaces in it. It gets tokenized as one token and normalized to 金持ち, which is impressive.
    • I don't love the fact that word_break_helper converts regular/halfwidth parens "()" to spaces, but not fullwidth parens "()".
    • The plain field can probably pick up the slack on this one.
  • Generates some zero-width non-empty tokens; these are essentially empty input tokens. For example, ㏆ ("C/kg") is normalized as two tokens, "c" from offset 0 to 1, and "kg" from offset 1 to 1 (i.e., zero length). The "kg" can be matched, but not highlighted, for example.
    • This is in some ways the opposite of the usual empty output tokens we see (when, for example, a token made up of nothing but a diacritic, like "`", has the diacritic deleted, leaving an empty token).
    • Other tokenizers aren't much better. The ICU tokenizer skips over it, so nothing matches it. The SmartCN tokenizer (used for Chinese) creates a single token "c/kg" which can't match anything else. This is the least bad of the options, though I think I'd prefer if both "c" and "kg" had offset from 0 to 1, but that can cause problems with some filters that could come later.
    • LATE-BREAKING UPDATE: Some (maybe most? defnitely not all...) single-character numbers and letters in parens (e.g., ⑶, ⒇, ⒳, 🄩, ㈂, ㈒, ㈣, ㉃), also get tokenized with start and final offsets that are equal. The tokenizer spits out empty tokens, but sudachi_baseform restores them. This also means these can't be highlighted.

According to the docs, "the sudachi_baseform token filter replaces terms with their Sudachi dictionary form. This acts as a lemmatizer for verbs and adjectives." Other options are sudachi_normalizedform and sudachi_readingform, which perhaps I should also investigate.

Customized POS Filter

[edit]

Sudachi's default part-of-speech filter has many expected categories—conjunctions, auxiliary verbs, and seven kinds of particles. They also have three kinds of symbols and nine kinds of punctuation—including separate categories for left parens, right parens, commas, periods, and three kinds of ASCII art.

Oddly, it seems that foreign scripts (I tested samples from more than 60 scripts) are tagged as both "補助記号" ("punctuation") and "補助記号,一般" ("punctuation-misc"), so disabling either by itself does nothing. That was pretty annoying to figure out. Not being able to use the explain API feature did not help. Disabling both allows foreign scripts through, along with the Japanese prolonged sound mark (ー) and a few individual katakana small letters.

Disabling the three "symbol" filters ("記号" ("symbol"), "記号,一般" ("symbol-misc"), and "記号,文字" ("symbol-letters")) allows through some individual Cyrillic, Greek, and hiragana characters, along with the names of the Greek letters (in katakana or English) and names of English letters (in katakana). It also let through URL parts (.com, .jp, www., http:// — but not https://).

Disabling the "感動詞,フィラー" ("interjection-filler") filter allows through some... wait for it!... filler interjections... though they seem validly searchable. Most are single kana with either a prolonged sound mark (ー) or wavy dash (〜) after them.

ICU Normalizer

[edit]

Since Kuromoji recommended the icu_normalizer char filter (rather than our usual token filter), I decided to try it with Sudachi, too. It did not go particularly well. At least we have the token filter to fall back on.

Char Filter Technical Difficulties
[edit]

For some reason—the lack of explain details is pretty annoying!—when both the ICU normalizer character filter and the Sudachi tokenizer are enabled, API calls through curl sometimes throw errors if the input text is more than 4096 characters.

Because of this, I had to change my analysis program to send smaller chunks of text to the API. The chunking joins lines of text with spaces. The chunking changes caused a few analysis changes in Japanese text because words that had previously been in different chunks were now adjacent, or vice versa. The addition of spaces also affected Latin tokens.. this was how I discovered the An / AN and cm / CM variation above.

The smaller chunk sizes also meant more API calls, which slowed down the processing by ~3.5x.... or so I originally thought!

Now I'm not sure, because I also ran some direct load time tests (more details below under Load Timings), but without the icu_normalizer char filter added to the Sudachi analyzer, the load ran in ~35 seconds. With the filter, it took 4 minutes and 34 seconds, or 7.8x longer. (It seems that even with that slow down, the API call overhead dominated.)

Oddly, this seems to be an interaction between the icu_normalizer character filter and the sudachi_tokenizer. The Sudachi analysis chain is slower, but not that slow, and the icu_normalizer character filter has no real effect when used with the kuromoji_tokenizer, icu_tokenizer, or standard tokenizer.

The icu_normalizer character filter seems unusable with Sudachi at this point, but we might as well see what it did, if anything.

Char Filter Observations
[edit]
  • Zero-width non-joiners (zwnj) in Arabic and Telugu, zero-width joiners (zwj) in Sinhala, and zero-width space (zwsp) in Khmer and Japanese used to break words, now they don't. There are some long Khmer tokens.
  • Turkish İnternet getting normalized as i + combining dot (i̇) causes a word break, and that makes me sad. It also causes some weird offset errors. "İnternet" ⇒ I + ̇ + nterne (no final t). "İnternet x" ⇒ I + ̇ + nternet (no x). "İnternet x y z" ⇒ I + ̇ + nternet + x + y (no z). However, if the final t, x, or z is capitalized, it works. Whaaat?
    • I can't tell if this is an icu_normalizer char filter bug or sudachi_tokenizer bug because icu_normalizer char filter's "filtered_text" attribute is cut off after the double-dotted i! However, it's not a problem with icu_normalizer char filter and the icu_tokenizer or standard tokenizer.
  • Overall, the icu_normalizer char filter fixes some non-CJK tokenizing issues like καλλίστῃ ⇒ καλλίστη (instead of καλλίστ + ῃ), but causes others like İnternet above. Not much impact on CJK text (unlike with Kuromoji), other than removing the few zero-width spaces in Japanese Wiktionary that cause word beaks.
Back to the Token Filter
[edit]
  • Various zero-width characters are not removed before tokenization, so they can still cause word breaks.
  • İnternet is parsed okay by the sudachi_tokenizer (İ is not in a different character class from plain Latin, it seems) and sudachi_baseform normalizes it correctly.
  • Greek καλλίστῃ is still split into καλλίστ + ῃ by the sudachi_tokenizer, and ῃ is normalized to ηι... so close!
  • Overall, it does the same good things icu_normalizer token filter always does.
Lowercasing
[edit]

Given that the icu_normalizer token filter does good things, and the fact that case-variant tokens are possible (e.g., An / AN and cm / CM above), tacking on lowercase to the Sudachi analysis chain and letting it upgrade to icu_normalizer seems like a good idea in general.

Automatic Upgrades

[edit]

Enabling the automatic upgrades generally gave the expected changes over baseline Sudachi, mostly affecting non-CJK text, like normalizing acronyms and splitting camelCase. (No homograph fixes because of the way the tokenizer breaks things up, alas.)

However, word_break_helper screws up some tokenization, like "SBI Holdings, Inc.", which requires the final period, and "http://", ".com", and "www.", which need the punctuation.

It seems that dotted_i_fix isn't strictly required when sudachi_baseform is in the analysis chain, but I don't feel like that's very future-proof.

Improved Splitting

[edit]

The sudachi_tokenizer makes some long tokens, but it has two other mode options to break longer items into shorter ones. For example:

  • In mode C (default) 選挙管理委員会 ("Election management committee") is one token.
  • In mode B, 選挙管理委員会 → 選挙 + 管理 + 委員会 ("election" + "manage" + "committee")
  • In mode A, 選挙管理委員会 → 選挙 + 管理 + 委員 + 会 ("election" + "manage" + "committee member" + "group")

There is also a token filter, sudachi_split, that breaks up some of the longer tokens into smaller tokens, but keeps both the full token and the new sub-token. It has both a "search" mode and a more aggressive "extended" mode. The "search" mode seems to basically add in the tokens from tokenizer A mode if they aren't already present:

input 選挙管理委員会 / アバラカダブラ / 手荷物
tokenizer A 選挙 + 管理 + 委員 + 会 + アバラカダブラ + 手 + 荷物
tokenizer A + split search 選挙 + 管理 + 委員 + 会 + アバラカダブラ + 手 + 荷物
tokenizer B 選挙 + 管理 + 委員会 + アバラカダブラ + 手荷物
tokenizer B + split search 選挙 + 管理 + 委員会 + 委員 + 会 + アバラカダブラ + 手荷物 + 手 + 荷物
tokenizer C 選挙管理委員会 + アバラカダブラ + 手荷物
tokenizer C + split search 選挙管理委員会 + 選挙 + 管理 + 委員 + 会 + アバラカダブラ + 手荷物 + 手 + 荷物

With tokenizer mode B and split mode "search", some long tokens, like MS&ADインシュアランスグループホールディングス ("MS&AD Insurance Group Holdings") get split up (ms&ad + インシュアランス + グループ + ホールディングス).

Some tokens, like サイバネティックスパワードスーツクォンタムテクター ("Cybernetic Powered Suit Quantum Techter") do not get split up, even with tokenizer mode A. Setting sudachi_split to "extended" mode will add additional tokens, one for each kana character. However, sudachi_baseform turns around and filters all of the new one-kana tokens... not sudachi_part_of_speech, but sudachi_baseform. Weird.

The lack of explain detail is again pretty annoying here. I'd like to see what info Sudachi has on サイバネティックスパワードスーツクォンタムテクター, and on the individual kana tokens sudachi_split generated.

More Tweaks

[edit]
  • Combo Filter: As with Kuromoji (and Nori), it makes sense to filter out the most common combining diacritics that cause problems for the tokenizer. The new sudachi_combo_filter strips U+300 through U+362, which is most of the Unicode block of basic combining diacritical marks. The results are pretty much the same as those for Kuromoji:
    • It's particularly helpful on Cyrillic/Russian names, which tend to have an acute accent to show stress. So instead of Ма́харов → ма + ́ + харов, we get Ма́харов → махаров.
    • There are also Greek, IPA, and other uncommonly accented Latin characters that no longer cause splits—particularly in detailed phonetic information from Wiktionary.
    • It also gets rid of all those naked combining character tokens.
  • ICU Folding: ICU folding (and if ICU is unavailable, ASCII folding) makes sense with the same Japanese-specific exceptions used by the Japanese version of the CJK analyzer. Again, since this affects non-CJK tokens the most, the results are the same as for Kuromoji:
    • The biggest impact is accented Latin text, just because there is more of it in Japanese wikis than other scripts, and thus more chance of a new collision in my samples. Still, there are plenty of regularizations in other non-CJK tokens, and matches in Greek, Cyrillic, Hebrew, Arabic, and a few Indic scripts.
    • Even more of the naked modifier and combining character tokens disappear.
  • Variation Selector Filter: I noticed this time that there are a few empty tokens that come from variation selectors being tokenized alone.
    • Looking at some contrived examples, this could cause unnecessary splitting of tokens, since the tokenizer splits on them. However, that doesn't come up in our actuall data.
    • The empty tokens also go away when ICU folding is enabled because we automatically add remove_empty filter after it.
  • Bidi Mark Filter: Since the tokenizer splits on bidi marks, filtering them could prevent token splits as well. However, out of all my samples, I found one example of an Arabic word with bidi marks in the middle, which did get split as a result. The rest—Arabic, Hebrew, and many other scripts—all had bidi marks at the end of the token, where splitting didn't change anything. I'm tempted to ignore them.
    • Sudachi doesn't make tokens out of bidi marks (at least with our current config).

The variation selector and bidi mark issues are the same for Kuromoji, but are not super common, and the empty tokens are handled by the remove_empty filter, so we can ignore them unless there is more evidence of a problem.

Custom Sudachi Config

[edit]

At this point, I must say that I did not have high hopes for Sudachi, so I didn't work out all the details of the custom config. However, the issues that would be important enough to deal with include:

  • Readily Fixable:
    • A char filter to remove the most common combining diacritics for alphabetic scripts, which otherwise get split by the Sudachi tokenizer.
    • The sudachi_tokenizer using mode B, and sudachi_split using mode search to get the best balance between recall and precision out of longer Japanese tokens.
    • A customized sudachi_part_of_speech filter to allow foreign scripts and a few other terms through.
    • The baseline sudachi_baseform and sudachi_ja_stop token filters.
    • Add lowercase, upgradable to icu_normalizer when ICU is available.
    • ASCII folding for non-CJK diacritics, upgraded to Japanese-specific ICU folding if ICU is available.
  • Need Work:
    • Prevent punctuation-only differences in tokens, like vs / vs. or "(財)" / 財 or Stand Up, STAND UP!, and STAND UP!!
    • Clean up tokens with final spaces.
    • Split on fullwidth equals sign (as in ブレスト=リトフスク)—and check whether Kuromoji needs that, too.
  • Need Thinking:
    • Is it possible to break up multi-word Latin tokens (café au lait; excuse me; ping-pong; SBI Holdings, Inc.) that isn't super expensive—regex filter bad!
    • Is it worth it to deal with the kana-in-parens issue, like 世(ごせ) → 世, and is it worth saving 金(かね)持(も)ち → 金持ち. Should we add fullwidth parens "()" to word_break_helper?
    • Is there any way to improve the inconsistent processing of ㎡, and are there other characters that misbehave like that?
    • Doing something about mid-dots (as in アゼルバイジャン・ソビエト社会主義共和国), probably either removing it or replacing it with a space.
  • Not Right Now (maybe later, maybe never):
    • Repair multi-script (and multi-"script") tokens split by the tokenizer (like chocоlate, dɔg, and δικεῖν).
    • Filtering variation selectors and bidi marks in a char filter.

I also opened a ticket on the Sudachi GitHub repo with a lot of these issues.

Load Timings

[edit]

Online discussions I saw commented on how Sudachi is fairly slow. I ran some load tests on my laptop to see how it compares to the Kuromoji and the CJK analyzer.

For these tests I do a direct index of a json file of Japanese Wikipedia articles. It has 2000 articles, and it is 76MB of data. (Loads have to be broken up into pieces around 90MB, so this is a round number below that threshold.) There is also less content than there ideally would be because the CJK characters are \u encoded in the json (e.g., "生物分類表" is encoded as "\u751f\u7269\u5206\u985e\u8868").

The tests are not very precise because of the random load on my laptop, and how awake Docker is feeling at any given moment. But I'm not looking for 1% changes, I'm really worried about 300% changes—though 25% changes are worth paying attention to. I run about 3 (usually 2–4) loads each time to get a more representative average, and I rerun configs I want to compare so that the things being compared were run within a few minutes of each other. As a result, comparing A to B one day, B to C another day, and A to C yet another day may result in the relative proportions of A:B and B:C not jiving with A:C. But things are usually within a few percentage points, and always seem to be within 10%. With those caveats out of the way....

The current production version of the CJK analyzer (with all of our various generic upgrades) is about 20% slower than the plainest version of the CJK analyzer, which is pretty fast.

Note that the measures below are still using fairly generic unpacked versions of Kuromoji and Sudachi to get a rough idea of their relative speed. Adding in additional filters to handle issues can slow things down a little more (or in the case of Sudachi and the ICU normalizer char filter, a lot!—see above).

  • Plain Kuromoji is about 8% slower than than the current production version of CJK, and adding in all of our generic upgrades, it's about 16% slower. Quite reasonable for (presumably) good tokenizing.
  • Plain Sudachi is 36% slower than plain Kuromoji, and 47% slower than production CJK. Sudachi with our upgrades is 34% slower than upgraded Kuromoji, and 56% slower than production CJK.

So, very roughly, Kuromoji is about 15% slower than our current production CJK, and Sudachi is about 35% slower than Kuromoji (and thus about 55% slower than prod CJK).

A 15% increase in analysis time (Kuromoji) would not be too bad. A 55% increase in analysis time (Sudachi) would be a very serious concern if it were applying to all of our indexes. (And that's why I wrote a plugin to handle the generic issues of acronyms and camelCase.. the regex versions I cobbled together work, but they are sloooooooooooow.)

But for one language even ~1.5x analysis time is not so bad, especially not in exchange for (presumably) good tokenization. Only eight wikis out of ~1000 are Japanese.. though Japanese Wikipedia and Wiktionary are pretty big and Commons and Wikidata have Japanese content.

Speaker Review of Tokenization

[edit]

Two Japanese speakers volunteered to review—Thanks, Jeena and Thomas!!

Data, Process, Etc.

[edit]

I originally tokenized Japanese snippets (article sentences and queries) using Kuromoji and the ICU tokenizer, using the best-looking options for tokenization mode and part-of-speech filtering. Since we use the ICU tokenizer by default in other languages, it seemed like both a good baseline, and a chance to get a better sense of the ICU tokenizer's performance on Japanese.

See the Data section above for more details, but, in brief, I planned to pull 50 examples each from Japanese Wikipedia and Wiktionary, and from user queries to those two wikis. Somehow, I only got 25 sentences from the Wiktionary entry corpus. I also included a double handful of "interesting" snippets, including some very long sentences, and snippets that generated particularly long tokens, to compare edge cases between the various tokenizers.

I grouped the samples in chunks and ordered them so that if not all items were reviewed (it's a tedious task, for sure), I'd be more likely to have a decent sample across the various types. I highlighted the differences between tokenizers, or collapsed them to one tokenization when they agreed, and highlighted stop words, particles, and other text that was omitted from the final tokenization. I included multiple output tokens when generated (e.g., including 手荷物 and 手 + 荷物 output for input 手荷物), and showed the final normalized tokens in the tokenization.

Both Jeena and Thomas reviewed more than 100 snippets, and Thomas reviewed all ~200. Based on the rate of agreement between them and the general subjectivity of ratings and the squishiness of the rating system, I'm very comfortable extrapolating from the items that only Thomas reviewed. I really, really appreciate all the time both reviewers took, and I mostly forgive Thomas for inflicting Sudachi on me. ;)

While reviewing, Thomas suggested Sudachi as a viable alternative, so I explored it as well. I added a small number of snippets where Sudachi had "interesting" behavior. I re-tokenized all of the previous samples with Sudachi, again using the best-looking options for tokenization mode and part-of-speech filtering. I removed any that were the same as Kuromoji or ICU. (When Kuromoji and ICU differed, Sudachi was much more likely to agree with Kuromoji.)

Thomas then reviewed all of the tokenizations where Sudachi differed from both Kuromoji and the ICU tokenizers (I copied over his and Jeena's evals from the ones that were tokenized the same).

Reviews & Scores

[edit]

Since reviewing is a tedious task, I only asked the reviewers to give a single grade of good, ok, or bad to the whole tokenization, though I encouraged them to include any comments they wanted to make.

For some of my analysis, I converted scores to numerical values from approximately 0 ("bad") to 1 ("good"). "ok" was 0.5. Sometimes the reviewers' ratings had modifiers, like "good, but barely" or "bad, but it got this hard thing right". I assigned + and - modifiers worth ±0.15. Modified values include "good-", "good+", "bad+", and "ok-". Nothing got a rating of "bad-", but there were a couple of "good+" ratings ("there's a typo, but it got it right anyway") so the scores got from 0 to 1.15. Where Jeena and Thomas both gave scores, I averaged them.

I also compared this fine-grained scoring to a more coarse bad/ok/good == 0/½/1 score, and it didn't make much difference.

A few snippets for review were dropped because they had objectionable material in them and I didn't want to subject the reviewers to that. A couple were missed here and there because of my poor formatting or other issues. And a couple of snippets were dropped because they turned out to be mostly gibberish (the equivalent of asking how to properly break "dshkflgdlusvds" into words in English).

The tables below show the summarized ratings for the various samples.

Japanese Wikipedia Japanese Wikipedia user queries
# good ok bad good% ok+% best best% # good ok bad good% ok+% best best%
kuromoji 50 43 6 1 86% 98% 35 70% 48 40 3 5 83% 90% 38 76%
icu 50 26 10 14 52% 72% 21 42% 48 33 3 12 69% 75% 31 62%
sudachi 49 49 0 0 100% 100% 49 98% 48 42 3 3 88% 94% 43 86%
sud-kurΔ 14% 2% 28% 4% 4% 10%
Japanese Wiktionary Japanese Wiktionary user queries
# good ok bad good% ok+% best best% # good ok bad good% ok+% best best%
kuromoji 25 13 6 6 52% 76% 8 32% 45 35 4 6 78% 87% 36 80%
icu 25 12 10 3 48% 88% 9 36% 45 36 5 4 80% 91% 36 80%
sudachi 23 21 1 1 91% 96% 21 84% 44 38 4 2 86% 95% 42 93%
sud-kurΔ 39% 20% 52% 9% 9% 13%

Key

  • #: number of samples reviewed; see above for why some are less than 50, and why Sudachi sometimes has fewer than the others.
  • good: number of samples with an average rating of 0.6 or better (i.e., better than "ok").
  • bad: number with an average rating of 0.25 or worse (i.e., threshold halfway between "ok" and "bad").
  • ok: number with a rating between "good" and "bad".
  • good%: percentage rated "good".
  • ok+%: percentage rated "ok" or better (i.e., "ok" or "good").
  • best: number of times this tokenizer had the best score; the best score is not necessarily "good"; there were lots of ties (Kuromoji and Sudachi agree fairly often), so the sum is more than 100%.
  • best%: percentage (of the max) that are "best"; Sudachi loses out for having snippets skipped; in the wiki sample it was best in 49 of the 49 that were rated, but it only scores 98% because there were 50 samples. Snoozing equals losing.
  • sud-kurΔ: The difference in percentages between Sudachi (always more) and Kuromoji (always less); I added this to force myself to admit Sudachi is better and I'll have to deal with its idiosyncrasies somehow.

The table below, "All-Wiki Weighted", shows the weighted counts when the four buckets above are each scaled up to 50 samples and combined. I also calculated the Wilson score interval for the combined good% and ok+%. (It's the preferred confidence interval for small numbers of values and values close to 0% or 100%. It's also asymmetric and won't give ranges over 100%!) I would not be shocked if I didn't quite calculate the combined confidence interval correctly, but it still demonstrates who the winner is.

All-Wiki Weighted
# good ok bad good% wilson ok+% wilson
kuromoji 200 150 26 25 74.8% 68.4%–80.3% 87.6% 82.3%–91.5%
icu 200 124 39 37 62.2% 55.3%–68.9% 81.5% 75.5%–86.3%
sudachi 200 183 10 8 91.3% 86.6%–94.5% 96.2% 92.6%–98.1%
sud-kurΔ 16.5% 8.7%

And for those who don't like looking at tables of numbers—what is wrong with you?!—here's a nice graph of the good% and ok+%:

Analysis (a.k.a. Sudachi Fails to Suck)

[edit]

Keep in mind that these are sentence-level metrics, so at the token level the accuracy would be even higher. Token-level accuracy is much harder to count, though: the reviewers have to annotate their reviews in much more detail, I'd have to figure out if one token being split in two or two tokens being combined counts as one error or two, and I'd have to wrestle with the question of whether one- or two-token sentences should be weighted equally or much less than 15-token sentences. Life is too short. (i.e., "I love deadlines... they make a cool whoooshing sound as they fly by.")

So, given the caveats about sentence-level metrics, and that an "ok" rating is ok enough, Kuromoji seems good enough to deploy—though its performance on the Wiktionary sample isn't the greatest.

Sudachi, however, is clearly better. Even if the error bars are not 100% copacetic, they are 94.3%–98.2% copacetic, and Sudachi, despite its quirks and inconsistencies, is clearly better at its core job—parsing Japanese.

What Next?

[edit]

Did I mention that Sudachi is quirkier than a Manic Pixie Dream Girl in an early 2000s Hollywood romcom? We've got a list of issues to fix or think about (see Custom Sudachi Config above) that I don't want to go unaddressed.

Sudachi is currently available from its developers for various versions of Elasticsearch (including our current 7.10), and OpenSearch versions 2.6—2.18. Unfortunately, our current mid-term upgrade/migration plans have us stopping at OpenSearch 1.x—not forever, but long enough that the lack of Sudachi plugin would have to be addressed.

We've ported (and backported, and sideported, and diagonallyported) plugins before, so we could if we had to, but it'd be great if we didn't have to.

Also of note: the Sudachi dictionary repo is available.

So, my next step is going to be investigating the possibilities related to mixing and matching (and hacking) dictionaries.

In particular:

  • How hard is it to use a user-specified dictionary with Sudachi or Kuromoji?
  • Do the newer Sudachi dictionaries fix some of its quirks?
  • How hard is it to recompile an edited dictionary for Sudachi or Kuromoji?
  • Can we get the Kuromoji plugin running with a provided or edited Sudachi dictionary, and how well does it do?

Once we have a handle on these issues, we'll decide on our approach for actual deployment. Options I see include:

  • Settling for Kuromoji until we migrate/upgrade to OpenSearch 2.x+ because it's the simplest option from a technical/development/deployment standpoint.
  • Using Kuromoji with a better/newer/Sudachier dictionary.
  • Porting Sudachi to OpenSearch 1.x, and using either the default dictionary or an edited one.

Depending on which option we go with, there will be a few accompanying analysis chain customizations to handle the remaining quirks.

Dictionaries, Dictionaries, Dictionaries Change of Plans!

[edit]

I decided to work out a decent custom Sudachi config using the default components available. I was a little unsure how best to have our config code know whether or not we had a custom dictionary installed, but I figured that a good config for the default components would be a useful starting point.

As I worked through the config, I managed to find decent accommodations for all of the biggest issues I have with Sudachi, so the plan has changed. Rather than introduce the complexity of custom dictionaries, I'm going to customize our config to handle the shortcomings of the default dictionary (and probably try to suggest some changes to the dictionary upstream). This will make it easier to upgrade Sudachi in the future, and to pick up any more important core improvements to the dictionary without having to maintain a custom dictionary in parallel.

In addition to being easier to implement and maintain, having a good config for both Kuromoji and Sudachi (with "better" Sudachi overriding "good enough" Kuromoji when both are available), will allow us to gracefully fall back to Kuromoji if porting Sudachi to OpenSearch 1.x turns out not to be possible or worth the effort, and upgrade back to Sudachi when we reach OpenSearch 2.x.

Enter word_delimiter_graph
[edit]

I had originally configured a number of custom character-by-character mappings (=, =, and ゠, which are used like hyphens; and lots of misc halfwidth and fullwidth punctuation that show up in tokens—all mapped to spaces by a character filter), a token filter to clean up tokens with word-final spaces, and a janky char filter hack to replace spaces with a non-space character (half-fill space 〿 was one of the better options) to block multi-word Latin tokens. To keep the last one from also mucking up the parsing of Japanese text, I had to limit it to spaces between Latin characters. That turned out to be too limiting, so I had to allow some optional punctuation in there. It was still a char-for-char replacement, which kept the char filter bookkeeping from getting too complex..

And then I remembered the word_delimiter token filter—and was reminded by the docs of its undeprecated cousin, word_delimiter_graph. Seems more efficient because it is one filter instead of three (char filter for punctuation, char filter for converting spaces between words, and a token filter for trimming word-final spaces), it won't apply to most tokens and should bail quickly, it does its expensive munging on short strings (tokens instead of whole documents), and the internals should have been reasonably optimized for this task in a way that configuring other kinds of generic filters can't be.

I set up a slightly customized version that:

  • does not split on case changes (that should have already been handled elsewhere, and I don't entirely trust sudachi_baseform not to introduce spurious camelCase).
  • does split on numerics. I don't love this, but it solves all the ㎡ problems by breaking things up the same everywhere.
  • does stem English possessives. This isn't foolproof by a long shot, because the tokenizer splits on apostrophes by default (generating plenty of possessive s tokens), but for words with 's in the Sudachi dictionary, like Master's, the possessive 's does get stripped off.
    • So it doesn't prevent all extra instances of English possessive s in the index, but it does allow possessives in the Sudachi dictionary to match their non-possessive counterparts. (Possessives that aren't in the Sudachi dictionary will already match their non-possessive counterparts.)

Even though word_delimiter_graph is configured to adjust offsets so they point to the sub-word, it doesn't always work. In particular, when camelCase processing splits a word into two or more words (AprilFool → April Fool) that the tokenizer recognizes as a multi-word token, the bookkeeping is thrown off, and the offsets for the sub-words are the same as the original token. Thus if you search for april, you could get match highlighting like this: AprilFool April Fool. Not great, but not the worst thing ever.

My word_delimiter_graph filter has to come after all the Sudachi-specific filters, though, or they all return nothing.. even when word_delimiter_graph does nothing. I guess it eats the Sudachi-specific attributes.

Because it comes late, it can't block the parenthetical-eating (e.g., 蜜蜂(みつばち) → 蜜蜂) in sudachi_baseform, so I restored the mapping of those to spaces, which seems to do the expected thing.

word_delimiter_graph also eats—and splits on!—tildes (~/~), but not the similar-to-identical looking wave dash (〜). Tildes seem to give somewhat better parses for the various use cases (see link), so I mapped the wave dash to the fullwidth tilde. (The tilde-like characters are only common in my Wikipedia sample, where wave dash is about twice as common as the fullwidth tilde, and ~25x more common than the regular tilde.)

It turns out that word_delimiter_graph has even better coverage for some uncommon edge cases I hadn't noticed, including script-specific punctuation, like Arabic ۔ and ؛, Armenian ։, Devanagari ।, Ethiopic ።, Hebrew ־, Meetei Mayek ꯫, Ol Chiki ᱾, and Tibetan ་ and ༝.

More Observations:

  • The default treatment of = as a hyphen is inconsistent (it's probably based on what's in the dictionary). word_delimiter_graph makes them more consistent.
    • シュモラー=シューレ ("Schmoller-Schule") is split into シュモラー and シューレ. With a space instead of =, it is unchanged.
    • ブレスト=リトフスク ("Brest-Litovsk") is kept as one token; with a space instead of =, it is two tokens.
    • A weird one is that ジュリアンヌ=カトリーヌ ("Juliane-Catherine", with hyphen) is split into two tokens (ジュリアンヌ ("Juliane") and カトリーヌ ("Catherine")), while ジュリアンヌ カトリーヌ ("Juliane Catherine", with space) is split in three (ジュリ ("Julie") + アンヌ ("Anne") + カトリーヌ ("Catherine")). Using word_delimiter_graph doesn't really address this particular inconsistency.
    • All three of = (fullwidth equals sign), = (equals sign), and ゠ (katakana-hiragana double hyphen) get used as hyphens, so letting word_delimiter_graph split tokens on them does make sure they are treated more consistently, as long as there aren't any Juliane/Julie Anne issues.
  • word_delimiter_graph does break up numbers with commas in them. (word_break_helper is already splitting numbers on periods, so at least any inconsistent number punctuation (1.000.000 vs 1,000,000 and 3,14159 vs 3.14159) will be treated the same, even if precision is sub-optimal.
  • Looking more closely at ㏆ and the other square characters with slashes (㎧, ㎨, ㎮, ㎯, ㏞, ㏟), they are actually pretty okay.
  • ㎡ is so random and not super common, and word_delimiter_graph breaks up the weird number parsing to again be consistent, even if precision is sub-optimal.
  • word_delimiter_graph also takes care of the remaining tokens with middle dots. (word_break_helper handles actual middle dots (·) while word_delimiter_graph handles katakana middle dots (・) and halfwidth katakana middle dots (・).
Final Sudachi Custom Config
[edit]

The final wiki-optimized config for Sudachi (version ES 7.10) is (not necessarily listed in analyzer order):

  • A char filter to remove the most common combining diacritics for alphabetic scripts, which otherwise get split by the Sudachi tokenizer.
  • A char filter to normalize wave dash (〜) to fullwidth tilde (~), and fullwidth parens ( & ) to spaces.
  • The default sudachi_baseform and sudachi_ja_stop filters.
  • A modified sudachi_tokenizer set to split_mode B, plus a new sudachi_split token filter and a lightly customized word_delimiter_graph token filter, all to minimize overly long Japanese & Latin tokens, and prevent multi-word Latin tokens.
  • A customized Sudachi part-of-speech filter to allow foreign scripts and a few other terms through.
  • Lowercasing (to normalize sudachi_baseform output) and ASCII folding for non-Japanese diacritics, both upgraded—to ICU normalization and Japanese-specific ICU folding—when ICU is available.

Load Timings, Addendum

[edit]

Adding word_delimiter_graph to Sudaci actually seems to speed up load times by a few percent, even over the baseline unpacked Sudachi config. Presumably because adding an occurence of sbi, holdings, and inc to existing index entries is faster than creating a new index entry for sbi holdings, inc.

What Nexter?

[edit]

We still need to decide on our approach for actual deployment. Options I see include:

  • Enabling Sudachi now, and then, either:
    • ✓ Porting Sudachi to OpenSearch 1.x.
    • Reverting to Kuromoji until we get to OpenSearch 2.x if the port doesn't take (whether for technical or non-technical reasons).
  • Settling for Kuromoji until we migrate/upgrade to OpenSearch 2.x+.

I prefer Sudachi now, though it may frustrate and confuse our users if we have to downgrade to Kuromoji for a little while.. but we can burn that bridge when we get to it!

Sudachi Updates

[edit]

March 2025—It looks like we're going to port Sudachi to OpenSearch 1.3 (the porting is done, just need some analysis chain analysis and testing). We're standardizing on Sudachi 3.3.0 and the 20250129 core dictionary with OpenSearch. I've been using Sudachi 3.0.0 and the 20241023 core dictionary with Elasticsearch, so I need to upgrade my Elasticsearch version to match as a baseline for general Elastic vs OpenSearch testing. A few observations:

  • Upgrading the dictionary caused only a few small changes in my regression test sets. A few longer tokens got broken up into shorter tokens. Pretty much what you'd expect for a periodic dictionary update.
  • Upgrading Sudachi itself had a nice effect that the morpheme attribute no longer causes problems, and you can see what's going on with that attribute. I'm a bit surprised that partOfSpeech is an array with a fixed 6 slots; most tokens have 2–3 non-default elements ("*") in the array, with some having 4. Their arrangement is odd.. they are often at the beginning of the array, but sometimes some are at the end of the array (with 2–3 "*" elements in the middle). Weird.
  • Sudachi 3.3.0 handles some multi-character codepoints better. Things like ℆ and ㏆ no longer generate zero-length tokens. Both tokens generated for each (c and u or C and KG) have the same offsets (i.e., they overlap). Similar for date elements like ㏢ (which becomes 3 and with the same offsets). Enclosed Alphanumerics (⑾, ⒇, ⒳) no longer generate empty tokens, either. Nice!
  • There are some mildly different shenanigans with "3㎡"—2163㎡ and 2413㎡ are parsed differently, for example—but overall very low impact.

sudachi_word_delim Update

[edit]

July 2025 (T396529)—During reindexing Erik discovered that we have some tokens that cause indexing errors because their offsets are out of order.

A simple example is 1級 (offsets 0-2). sudachi_split (which breaks up longer tokens into sub-parts to improve recall and keeps the original to improve precision) breaks it up into 1(0-1) and 級(1-2). Later, sudachi_word_delim (which breaks up some Sudachi cruft like multi-word tokens ("good luck"), tokens with numbers ("2人乗り"), tokens with other symbols ("@tokyo", "ヤフオク!"), tokens with hyphen-like characters ("Hewlett-Packard", "CD-ROM", "クレルモン=フェラン"), etc.) re-splits the original 1級, giving tokens 1(0-1), 級(1-2), 1(0-1), 級(1-2).

The duplication isn't great, but the real problem is that the tokens aren't in offset order.

I looked at a few possible solutions:

  • Configure sudachi_word_delim to not split on numbers. At the moment, all of the problem cases come from tokens with numbers, but that isn't guaranteed to always be the case.
  • Configure sudachi_word_delim to not adjust offsets. This is kind of okay in some case, but in others, searching for "クレルモン" would highlight all of "クレルモン=フェラン", or searching for "Packard" would highlight all of "Hewlett-Packard". Kind of a sledgehammer approach.
  • Having two word_delimiter_graph filters. The first wouldn't break on numbers, but would adjust offsets. The second would break on numbers, but not adjust offsets. This is brittle in that if some aspect of the Sudachi dict that causes the problem changes and it's not just numbers that cause the problem, we'd break again.
  • Use flatten_graph (as recommended by the docs). The new offsets can create zero-width non-empty tokens. In our example above, the new token stream is 1(0-1), 級(1-1), 1(1-1), 級(1-2). It is a bit overly aggressive in that it flattens offset rangeds that don't cause problems. For example, あたたかさ is tokenized as あたたかさ(0-5), あたたか(0-4), さ(4-5), which indexes fine. But flatten_graph changes the first token to あたたかさ(0-4) so that it doesn't extend beyond the second token. (Perhaps this is useful preprocessing for other token filters one could use?)

Using flatten_graph seems to cause the fewest problems. Highlighting actually works fine in the examples I tested, presumably from either the duplicate tokens or from the plain field providing different tokens for highlighting.

It also has the lowest impact, affecting fewer than 0.045% of tokens in all my samples (articles and queries from Japanese Wikipedia and Wiktionary).

The impact on indexing speed is about +14% for Japanese text (Japanese-language wikis, and Japanese text on Commons or Wikidata). Other options were about +9%, so using flatten_graph isn't much worse, and the impact is limited to a relatively small subset of all of our data.

Very Long Token Update

[edit]

August 2025 (T402220)—During reindexing Erik discovered another problem! Very long strings of unparsable characters cause Java IndexOutOfBoundsException errors during reindexing. The original problem was found on a Japanese Wikipedia page with tens of thousands of emoji (🤓) with no spaces. I isolated the problem to the sudachi_tokenizer, and it seems to be the same regardless of split_mode (A, B, or C) or other configuration settings.

For the emoji character, the problem emerges between 10K and 11K characters, though before I figured out these tighter bounds, I was testing strings of various characters that were 8K, 16K, 32K, and 64K characters long.

Generally, I tried to find a single character of various types that I can repeat which gets tokenized as one long token. I used ワ for katakana, 四 for Han, 🤓 for emoji, 𐌳 for Gothic, X for Latin, digits (0–9), and ร for Thai. I couldn't find a single character that worked for hiragana, but the two characters ゛け (i.e., ゛け゛け゛け゛け゛け...) or just one ゛ followed by multiple け (゛けけけけけけけ....) work; I prefer repeating both because when it gets internally chunked (foreshadowing!), all the chunks start with ゛け, which keeps them from being parsed into smaller pieces.

8k

  • In all cases, a token of 8k repeated characters (or 4k copies of ゛け) is handled fine, and returned as one long token. (Though strings of 🤓 generate no tokens, because they are filtered by the tokenizer, which is fine. I added Gothic 𐌳 to the list of test cases to have a multibyte character that passes through the tokenizer.)

16k

  • For 16k-character tokens, 四, ゛け, X, and digits all generate one long token.
  • 16k-long tokens of ワ, 🤓, 𐌳, and ร cause IndexOutOfBoundsException errors.

32k

  • For 32k-character tokens, 四, X, and digits all generate one long token. 🤓 generates no tokens, which is correct.
  • 32k-long tokens of ゛け and ร cause IndexOutOfBoundsException errors.
  • 32k-long tokens of 𐌳 gets broken up into blocks of 4096 𐌳 (internally represented as 24,576-char tokens of repeated \uD800\uDF33).
  • 32k-long tokens of ワ crash my OpenSearch instance, and I have to tear down and restart my Docker image!

64k

  • For all 64k-long tokens, the original token gets broken up into chunks of 4096 characters. As before, 🤓 generates no tokens. The 4096-char chunks of 𐌳 are 24,576 characters of \uD800\uDF33 internally.

Note that ร causes errors with both 16k and 32k strings, while 🤓 and 𐌳 only cause errors at 16k. I think that this has something to do with Thai characters being 3 bytes, while this emoji and Gothic characters are four bytes.

Because everything recovers (no errors) at 64k strings, my guess is that there is something going on with a mismatch between counting bytes and counting characters to trigger chunking. When the string is short, it doesn't matter because it is too short to cause problems. When it is really long, it doesn't matter because it triggers the chunking no matter which way you count. I don't think there's any problem with the internal \u representation of 𐌳 being 12 bytes per character, but it's not impossible.

(It would be nice if stupidly long tokens where handled more consistently. A 10k-char token being indexed as one token while an 11k-char token gets chunked into 3 pieces (4096, 4096, and 2808 characters long) is unexpected. Breaking all tokens over 4096 into 4096-char chunks would be better.)

There's a Filter for That

[edit]

To prevent the error, I added a char filter to chunk 8k stretches of text with no spaces and no token-breaking punctuation (i.e., matching [^、。\s「」・]{8000}). We do have stretches of natural text (including possibly older texts on Japanese Wikisource) with over 500 CJK characters with no spaces or token-breaking punctuation, but I couldn't find any examples in the thousands—so this seems like a generous threshold to keep from mangling natural Japanese text.

In my local tests, the new filter adds 2-5% to reload times, which seems to be the ballpark for the minimum from any additional filter—though the Sudachi analysis chain is already a little slow, so this is admittedly 2-5% of a bigger number.

This filter seems to work, and I'm not too upset about breaking up long weird strings because I don't think there are many valuable 8K tokens out there (at least in our use case). But, in general, depending on a char_filter to prevent rare indexing failures is uncomfortably brittle, so I opened an upstream ticket (with pretty much the same content as this update) to fix it at the source.