Skip to content

fix(java): escape text in the HTML output - #1391

Merged
tushuhei merged 1 commit into
google:mainfrom
kwy404:fix-java-html-escape
Sep 25, 2026
Merged

tushuhei merged 1 commit into
google:mainfrom
kwy404:fix-java-html-escape

Conversation

@kwy404

@kwy404 kwy404 commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

The Java HTMLProcessor has the same escaping problem that #1381 fixed in the Python processor. jsoup unescapes character references in text nodes, and PhraseResolvingNodeVisitor appended the characters back as is. So HTMLProcessor.resolve(phrases, "a&lt;b&gt;c") returned a<b>c, and Parser.translateHTMLString("今日は&lt;b&gt;天気です。") returned a real <b> element instead of the escaped text.

This re-escapes &, < and > when writing text back. Attribute values were already fine, since jsoup escapes them in Attribute.toString(), and <script> and <style> contents are DataNodes, so they are not affected. Added a test that fails before the change and passes after.

Tested with the Java unit tests (HTMLProcessorTest and ParserTest) and google-java-format.

jsoup unescapes character references in text nodes, and HTMLProcessor
wrote the characters back as is, so `a&lt;b&gt;c` came out as `a<b>c`.
This re-escapes `&`, `<` and `>` in text, like the Python processor does
since google#1381. Attribute values are already escaped by jsoup.
@tushuhei
tushuhei merged commit 28c6e13 into google:main Sep 25, 2026
18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

2 participants