Skip to content

fix: a one-character word ending on a boundary must not block sentence splitting - #369

Merged
hayashi-mas-wap merged 1 commit into
WorksApplications:developfrom
h1431532403240:fix/non-break-checker-single-char-word
Oct 2, 2026
Merged

hayashi-mas-wap merged 1 commit into
WorksApplications:developfrom
h1431532403240:fix/non-break-checker-single-char-word

Conversation

@h1431532403240

Copy link
Copy Markdown
Contributor

Summary

Fixes #367.

With the lexicon checker (NonBreakChecker), has_non_break_word blocks a boundary whenever a word ends exactly on it and more text follows. This includes the terminator itself: 。, !, ? and ♪ are SudachiDict entries. For example, 今日は晴れ。明日は雨。 comes back as one sentence from SentenceSplitter::with_checker and from the CLI's default --split-sentences yes. Java Sudachi splits it into two.

The cause is that the Ordering::Equal branch counts input[i..], the rest of the input, instead of the matched word input[i..eos_byte]. Its comment says "check that there are more than one character in the matched word".

Changes

  • sudachi/src/sentence_detector.rs: use the word length, as Java's hasNonBreakWord does. When the word is a single character, return None instead of Some(false) so the scan continues. A longer entry at the same position that crosses the boundary (e.g. 。な in あ。なに。) still blocks the split, as in Java.
  • sudachi/tests/sentence_detector.rs: add regression tests with a small lexicon (。, 娘。, な。な, 。な):
    • 今日は晴れ。明日は雨。 → 18 (currently 33)
    • モーニング娘。です。 → 30
    • ばな。なです。 → 21
    • あ。なに。 → 15
    • a splitter-level check

Tests

  • cargo test -p sudachi on develop (3b68eb9) with this change: 361 passed, 0 failed.
  • cargo fmt --all -- --check: clean
  • cargo clippy -p sudachi: no warnings in the touched files
  • Without the fix, the new test fails with left: 33, right: 18.
  • CLI with SudachiDict small 20260723: 今日は晴れ。明日は雨。 and the !, ?, ♪ variants now split. 本当!?うそ。 still does not split, the same as Java, because the multi-character entry !? ends on the boundary.

This is independent of #368 (the ASCII " fix for #366), and the two merge without conflicts.

AI assistance: Claude (Claude Code) helped reproduce the issue, write the patch and tests, and draft this description. I reviewed the change.

🤖 Generated with Claude Code

NonBreakChecker::has_non_break_word counted `input[i..]` (the rest of the
input) instead of the matched word `input[i..eos_byte]`, so any word that
ends exactly on the candidate boundary blocked it whenever more text
followed, including the terminator itself (`。`, `!`, `?`, `♪` are
dictionary entries). `今日は晴れ。明日は雨。` came back as one sentence.

Use the word length, as the comment says and as Java's hasNonBreakWord
does, and keep scanning when the word is a single character so a longer
entry crossing the boundary at the same position is still found. Add
regression tests.

Fixes WorksApplications#367

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Copilot AI lite review requested due to automatic review settings September 24, 2026 19:14

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@hayashi-mas-wap
hayashi-mas-wap self-requested a review October 1, 2026 08:19
@hayashi-mas-wap
hayashi-mas-wap merged commit 1c21940 into WorksApplications:develop Oct 2, 2026
16 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

3 participants