Skip to content

perf(java): copy plain runs after the first escape when writing JSON strings - #4122

Merged
chaokunyang merged 2 commits into
apache:mainfrom
pavel-ptashyts:perf/json-escaped-string-write
Oct 6, 2026
Merged

chaokunyang merged 2 commits into
apache:mainfrom
pavel-ptashyts:perf/json-escaped-string-write

Conversation

@pavel-ptashyts

@pavel-ptashyts pavel-ptashyts commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Why?

This is the writing side of #4055 / #4058. Once a string being written contains its first character that needs escaping (a quote, a backslash, a control character, or non-ASCII text with escapeNonAscii), both JSON writers write the rest of it one character at a time through writeEscapedChar - a switch, a capacity check and one store per character - and never return to a bulk copy. The cost of a string therefore follows the distance from that character to its end, not the number of characters to escape: one quote near the front of a 20 KB string makes it 13 times slower to write than with no quote at all. In Utf8JsonWriter a compact latin1 string longer than the fixed-size fast paths is also rescanned with charAt up to that character, so even a quote at the very end costs 5 times. Details and measurements in #4121.

What does this PR do?

  • Utf8JsonWriter: a compact latin1 String that the fast paths reject now goes from writeStringChars to writeLatin1StringEscaped, which copies each plain run a word at a time with the existing isJsonAsciiWord predicate, writes only the byte that stopped it through writeEscapedChar, and resumes. The character tails (writeStringSlow for String and CharSequence) copy the plain ASCII run up to the next stop with one capacity check per run, then handle the stop character exactly as before.
  • StringJsonWriter: the same in all four tails. writeLatin1StringSlow and writeLatin1StringUtf16Slow copy plain words and resume the word copy after a byte above 0x7F, running the word predicate at most once per eight bytes. In writeStringSlow the run follows the current output coder, because a character above 0xFF upgrades the output to UTF16 in the middle of a string; the UTF16 runs (writeStringSlow, writeStringUtf16Slow) clear latin1Output when they copy such a character.
  • Every run copy reserves capacity for at most RUN_CHUNK (8192) characters at a time, the way writeEscapedUtf16 already bounds its reservation, so (length - i) << 1 or pos + (length - i) cannot overflow for a huge string and skip grow before the word stores. A run that ends at a chunk boundary rather than at a stop just continues. The copy is inlined into each tail: a helper call per stop measurably slowed strings with many escapes.
  • The hot writeString(String) methods and their fixed-size fast paths are untouched; all new code is in the cold paths, following the C2-layout notes in both classes. The output is byte-identical.
  • Tests in JsonStringTest:
    • writeEscapeFollowedByOtherText compares both writers with JsonStringEscaper (a character-at-a-time reference) for a stop character (quote, backslash, newline, control character, latin1 above 0x7F) followed by latin1 text, 0x7F/0x80, text outside latin1, a surrogate pair and further escapes, with 0-20 characters before it and 0-8 after it so stops land in every lane of a word, with and without trailing text, for String and CharSequence input (each on a fresh writer), utf8 output and LATIN1 and UTF16 string output, with and without escapeNonAscii. Accepting a quote inside the latin1 run, or ignoring escapeNonAscii in the UTF16 run, makes it fail.
    • writeRunsAcrossReservationChunks: a quote, latin1 above 0x7F, text outside latin1 and a surrogate pair at 8192 +/- 9 characters after an escape, so they land on and around a chunk boundary.
    • writeRejectsUnpairedSurrogateAfterAnEscape: an unpaired high or low surrogate after an escape, inside a long plain run and at the end, is still rejected by both writers, with and without escapeNonAscii.
    • writeEscapeThenWideTextAfterUtf16Reset: a writer reset after a UTF16 result starts in UTF16 while assuming its output fits latin1; the run copy must withdraw that assumption. Removing the latin1Output clear makes it fail.
    • The old code was correct but slow, so these tests pin behaviour and also pass on it.

Related issues

AI Contribution Checklist

  • Substantial AI assistance was used in this PR: yes
  • If yes, I included a completed AI Contribution Checklist in this PR description and the required AI Usage Disclosure.
  • If yes, my PR description includes the required ai_review summary and screenshot evidence or equivalent persisted links of the final clean AI review results from both fresh reviewers described in AI_POLICY.md, the Fory-guided reviewer and the independent general reviewer, on the current PR diff or current HEAD after the latest code changes.

Completed checklist from AI_POLICY.md section 9:

  • Substantial AI assistance was used in this PR: yes
  • I included the standardized AI Usage Disclosure block below.
  • I can explain and defend all important changes without AI help.
  • I reviewed AI-assisted code changes line by line before submission.
  • I completed line-by-line self-review first and fixed issues before requesting AI review.
  • I ran two fresh AI review agents on the current PR diff or current HEAD after the latest code changes: one Fory-guided reviewer using AGENTS.md and .agents/ci-and-pr.md, and one independent general reviewer in a separate clean-context review session that was not pointed to .agents/ci-and-pr.md or any copied Fory-specific review checklist.
  • I addressed all AI review comments and repeated the review loop until both ai reviewers reported no further actionable comments.
  • I attached equivalent persisted links of the final clean AI review results from both fresh reviewers on the current PR head in this PR body.
  • I ran adequate human verification and recorded evidence (checks run locally, pass/fail summary, and confirmation I reviewed results).
  • I added/updated tests and specs where required.
  • I validated protocol/performance impacts with evidence when applicable.
  • I verified licensing and provenance compliance.
AI Usage Disclosure
- substantial_ai_assistance: yes
- scope: investigation, code drafting, tests, benchmark harness, PR text
- affected_files_or_subsystems: java/fory-json writers (Utf8JsonWriter, StringJsonWriter) and JsonStringTest
- ai_review: line-by-line self-review of the AI-drafted change is pending by the contributor (see unchecked items above). Two-reviewer loop on fresh sessions for every round, six rounds: Fory-guided reviewer Codex gpt-6-astra (xhigh), independent general reviewer Claude Fable 5.1. Fixed from review: each input type written by a fresh writer in the test so a CharSequence reaches the mid-string LATIN1-to-UTF16 upgrade itself; 0x7F/0x80 tails and strings ending right after a stop; surrogate test with and without escapeNonAscii; the StringJsonWriter latin1 tails resume the word copy after a byte above 0x7F, run the word predicate at most once per eight bytes and document that they are only reached without non-ASCII escaping; a test for the latin1Output clear after a UTF16 reset; every run copy reserves at most RUN_CHUNK characters at a time so the capacity arithmetic cannot overflow for huge strings (found by the Fory-guided reviewer), with a chunk-boundary test; the run copies are inlined into each tail after a helper call per stop measurably slowed strings with many escapes. Kept by decision: hot writeString fast paths untouched; String and CharSequence overloads stay duplicated like the existing code; tests check behaviour against JsonStringEscaper rather than speed. Final round on the current head: both reviewers report no further actionable comments.
- ai_review_artifacts: Fory-guided reviewer: https://github.com/apache/fory/pull/4122#issuecomment-5997450696 ; independent general reviewer: https://github.com/apache/fory/pull/4122#issuecomment-5997451342 (both on commit 72008b9e9a15f2f9fe1de2fdd3cf529a849056dd, the current head, final round, no further actionable comments)
- human_verification: see "Verification" below
- performance_verification: see "Benchmark" below
- provenance_license_confirmation: all code written for this PR against the Apache Fory sources; no third-party code introduced

Verification

JDK 25.0.2, Windows 11 x86_64, from java/:

  • mvn -pl fory-json -am install -DskipTests: success
  • mvn -pl fory-json test: 1276 tests, 0 failures, 0 errors, 0 skipped
  • JsonStringTest#writeEscapeFollowedByOtherText with a quote accepted inside the latin1 run of Utf8JsonWriter, and with escapeNonAscii ignored in the StringJsonWriter UTF16 run copy: fails, as intended
  • JsonStringTest#writeEscapeThenWideTextAfterUtf16Reset with the latin1Output clear removed from the UTF16 run copy: fails (中 comes back as -), as intended
  • mvn -pl fory-json spotless:check checkstyle:check: success, 0 Checkstyle violations

Does this PR introduce any user-facing change?

  • Does this PR introduce any public API change? No.
  • Does this PR introduce any binary protocol compatibility change? No. The JSON output is byte-identical.

Benchmark

Standalone harness from #4121 (not JMH): a bid whose adm field is a 19998-character VAST XML string, pure ASCII, 612 quotes and 345 newlines, the first quote at character 15. Variants change only the characters that need escaping inside adm, keeping its length. ForyJson.builder().build(), best of 7 rounds of 5000 writes after 30000 warm-up writes, one JVM per build, builds alternating three times on an otherwise idle machine; JDK 25.0.2. Range over the three runs, microseconds per write:

variant toJsonBytes main f85877f toJsonBytes this PR toJson main toJson this PR
nothing to escape 3.6-3.8 3.6-3.7 3.4-3.5 3.4-3.5
one quote at the very end 18.7-19.0 6.0-6.4 3.3-3.4 3.3-3.4
one quote at character 1 44.5-45.0 4.2-4.6 33.8-36.8 3.7-4.0
the real string 44.3-45.0 9.9-10.4 36.5-40.7 14.6-15.8

A bean of eight short strings, four of them with one escape each: 286-291 ns -> 181-191 ns (toJsonBytes), 255-272 ns -> 231-235 ns (toJson). The output digest is identical between the builds in every case. For reference, Jackson 3.1 writeValueAsString takes 12.5 us on the string without escapes and 25.0 us on the real one.

🤖 Generated with Claude Code

…strings

Once a string being written contained its first character that needs
escaping, both JSON writers wrote the rest of it one character at a time
through writeEscapedChar and never returned to a bulk copy, so the cost of a
string followed the distance from that character to its end rather than the
number of characters to escape. In Utf8JsonWriter a compact latin1 string
longer than the fixed-size fast paths was also rescanned with charAt up to
that character before the tail started.

Utf8JsonWriter now sends a compact latin1 string that the fast paths reject
to a loop that copies each plain run a word at a time and writes only the
stop byte through writeEscapedChar. Every character tail of both writers
copies the plain run up to the next stop with one capacity check per run and
handles the stop character as before; in StringJsonWriter the run follows the
current output coder, since a character above 0xFF upgrades the output to
UTF16 in the middle of a string. The fixed-size fast paths in writeString are
unchanged, and the output is byte-identical.

Fixes apache#4121

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@pavel-ptashyts

Copy link
Copy Markdown
Contributor Author

AI review artifact - Fory-guided reviewer (Codex gpt-6-astra, reasoning effort xhigh; fresh session, prompted to follow AGENTS.md, .agents/ci-and-pr.md and .agents/languages/java.md; read-only, static review) on commit 72008b9, final round (6).

Reviewed commit `72008b9e9a15f2f9fe1de2fdd3cf529a849056dd` as the Fory-guided AI reviewer.

No further actionable comments.

Residual risks/testing gaps:

- Static review only; contributor-reported tests, mutation checks, and benchmarks were not independently rerun.
- Other JDK versions, disabled compact strings, and big-endian execution are not covered by the supplied validation evidence.

@pavel-ptashyts

Copy link
Copy Markdown
Contributor Author

AI review artifact - independent general reviewer (Claude Fable 5.1; separate clean-context session, not pointed to .agents/ci-and-pr.md or any Fory review checklist, following the independent-review carve-out in AGENTS.md; read-only) on commit 72008b9, final round (6).

Reviewed commit 72008b9e9a15f2f9fe1de2fdd3cf529a849056dd.

No further actionable comments.

What I checked to reach that conclusion:

  • Equivalence of the run predicates with the per-character path. In every tail the "plain" test is exactly the set of characters the old writeEscapedChar would have written raw for that coder and escape setting: isJsonAscii for UTF-8 output, isJsonLatin1 with the 0x7F/0xFF cap for LATIN1 output, isJsonUtf16 with the 0x7F/0xFFFF cap for UTF16 output, and isJsonAsciiWord/isJsonAsciiByte/isJsonLatin1Byte for compact Latin1 bytes. I traced the packed-word predicate including inter-byte carries; it rejects every byte at or above 0x80, so word copies never swallow a byte that UTF-8 output must encode as two bytes.
  • Invariant that the Latin1 byte tails run only without non-ASCII escaping. All callers of writeLatin1StringSlow, writeLatin1StringUtf16Slow and writeLatin1StringEscaped sit after the early return on escapeNonAscii in writeString, so copying bytes above 0x7F raw there matches the old behaviour.
  • Surrogates. The run copy stops at any surrogate in all variants, the index rewrite from i + 1 >= length to post-increment i >= length is equivalent, and the bound uses length rather than the chunk limit, so a pair straddling a chunk boundary is handled and an unpaired surrogate at a chunk end is still rejected.
  • Capacity. Each chunk reserves one byte per input unit for LATIN1/UTF-8 output and two for UTF16, word stores only happen when a full word fits before the chunk limit, and the position and buffer references are re-read after every escape call and every grow. The mid-tail LATIN1 to UTF16 upgrade through writeCharRaw is followed by re-reading coder and position on the next chunk iteration.
  • Output coder. The widest accumulator clears latin1Output exactly when the old per-character writeCharRawUtf16 would have; the upgrade path still clears it inside writeCharRawUtf16.
  • Unescaped and short strings. All fast paths are untouched. The only new branch on a hot route is the coder check at the top of the UTF-8 writeStringChars(String), which is reached only after a fast path has rejected the string on content, and it folds away when strings are not byte-backed.
  • Tests. The matrix covers every tail (String and CharSequence, UTF-8, LATIN1, UTF16 and the mid-string upgrade), stops at chunk boundaries, 0x7F/0x80 tails, strings ending at a stop, unpaired surrogates after an escape, and the reset-after-UTF16 latin1Output case. One non-actionable note: in the reset test the String input goes through writeUtf16StringBytes, which clears latin1Output up front on byte-backed little-endian JVMs, so only the StringBuilder half exercises the run copy's clearing there. That is fine since the CharSequence tail is the only one where that clearing can matter on such JVMs.

Compare each bounded run reservation with the remaining buffer capacity so a
large cursor cannot overflow the check before packed stores. Keep the existing
copy paths and allocation failure behavior, and document the capacity proof.

Validated the before/after behavior with real 2147483645-byte storage on JDK 25.
The fixed UTF8 and UTF16 paths reject insufficient capacity before copying runs.

@chaokunyang chaokunyang left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@chaokunyang
chaokunyang merged commit f53dd7f into apache:main Oct 6, 2026
82 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Java] Strings are written one character at a time from the first character that needs escaping to the end (java, fory-json)

2 participants