fix: a one-character word ending on a boundary must not block sentence splitting - #369
Merged
hayashi-mas-wap merged 1 commit intoOct 2, 2026
Conversation
NonBreakChecker::has_non_break_word counted `input[i..]` (the rest of the input) instead of the matched word `input[i..eos_byte]`, so any word that ends exactly on the candidate boundary blocked it whenever more text followed, including the terminator itself (`。`, `!`, `?`, `♪` are dictionary entries). `今日は晴れ。明日は雨。` came back as one sentence. Use the word length, as the comment says and as Java's hasNonBreakWord does, and keep scanning when the word is a single character so a longer entry crossing the boundary at the same position is still found. Add regression tests. Fixes WorksApplications#367 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
hayashi-mas-wap
self-requested a review
October 1, 2026 08:19
hayashi-mas-wap
approved these changes
Oct 2, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes #367.
With the lexicon checker (
NonBreakChecker),has_non_break_wordblocks a boundary whenever a word ends exactly on it and more text follows. This includes the terminator itself:。,!,?and♪are SudachiDict entries. For example,今日は晴れ。明日は雨。comes back as one sentence fromSentenceSplitter::with_checkerand from the CLI's default--split-sentences yes. Java Sudachi splits it into two.The cause is that the
Ordering::Equalbranch countsinput[i..], the rest of the input, instead of the matched wordinput[i..eos_byte]. Its comment says "check that there are more than one character in the matched word".Changes
sudachi/src/sentence_detector.rs: use the word length, as Java'shasNonBreakWorddoes. When the word is a single character, returnNoneinstead ofSome(false)so the scan continues. A longer entry at the same position that crosses the boundary (e.g.。なinあ。なに。) still blocks the split, as in Java.sudachi/tests/sentence_detector.rs: add regression tests with a small lexicon (。,娘。,な。な,。な):今日は晴れ。明日は雨。→ 18 (currently 33)モーニング娘。です。→ 30ばな。なです。→ 21あ。なに。→ 15Tests
cargo test -p sudachiondevelop(3b68eb9) with this change: 361 passed, 0 failed.cargo fmt --all -- --check: cleancargo clippy -p sudachi: no warnings in the touched filesleft: 33, right: 18.今日は晴れ。明日は雨。and the!,?,♪variants now split.本当!?うそ。still does not split, the same as Java, because the multi-character entry!?ends on the boundary.This is independent of #368 (the ASCII
"fix for #366), and the two merge without conflicts.AI assistance: Claude (Claude Code) helped reproduce the issue, write the patch and tests, and draft this description. I reviewed the change.
🤖 Generated with Claude Code