Repository navigation
perf(java): copy plain runs after the first escape when writing JSON strings - #4122
Merged
chaokunyang merged 2 commits intoOct 6, 2026
Merged
Conversation
…strings Once a string being written contained its first character that needs escaping, both JSON writers wrote the rest of it one character at a time through writeEscapedChar and never returned to a bulk copy, so the cost of a string followed the distance from that character to its end rather than the number of characters to escape. In Utf8JsonWriter a compact latin1 string longer than the fixed-size fast paths was also rescanned with charAt up to that character before the tail started. Utf8JsonWriter now sends a compact latin1 string that the fast paths reject to a loop that copies each plain run a word at a time and writes only the stop byte through writeEscapedChar. Every character tail of both writers copies the plain run up to the next stop with one capacity check per run and handles the stop character as before; in StringJsonWriter the run follows the current output coder, since a character above 0xFF upgrades the output to UTF16 in the middle of a string. The fixed-size fast paths in writeString are unchanged, and the output is byte-identical. Fixes apache#4121 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
Author
|
AI review artifact - Fory-guided reviewer (Codex gpt-6-astra, reasoning effort xhigh; fresh session, prompted to follow |
Contributor
Author
|
AI review artifact - independent general reviewer (Claude Fable 5.1; separate clean-context session, not pointed to Reviewed commit No further actionable comments. What I checked to reach that conclusion:
|
Compare each bounded run reservation with the remaining buffer capacity so a large cursor cannot overflow the check before packed stores. Keep the existing copy paths and allocation failure behavior, and document the capacity proof. Validated the before/after behavior with real 2147483645-byte storage on JDK 25. The fixed UTF8 and UTF16 paths reject insufficient capacity before copying runs.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why?
This is the writing side of #4055 / #4058. Once a string being written contains its first character that needs escaping (a quote, a backslash, a control character, or non-ASCII text with
escapeNonAscii), both JSON writers write the rest of it one character at a time throughwriteEscapedChar- a switch, a capacity check and one store per character - and never return to a bulk copy. The cost of a string therefore follows the distance from that character to its end, not the number of characters to escape: one quote near the front of a 20 KB string makes it 13 times slower to write than with no quote at all. InUtf8JsonWritera compact latin1 string longer than the fixed-size fast paths is also rescanned withcharAtup to that character, so even a quote at the very end costs 5 times. Details and measurements in #4121.What does this PR do?
Utf8JsonWriter: a compact latin1Stringthat the fast paths reject now goes fromwriteStringCharstowriteLatin1StringEscaped, which copies each plain run a word at a time with the existingisJsonAsciiWordpredicate, writes only the byte that stopped it throughwriteEscapedChar, and resumes. The character tails (writeStringSlowforStringandCharSequence) copy the plain ASCII run up to the next stop with one capacity check per run, then handle the stop character exactly as before.StringJsonWriter: the same in all four tails.writeLatin1StringSlowandwriteLatin1StringUtf16Slowcopy plain words and resume the word copy after a byte above 0x7F, running the word predicate at most once per eight bytes. InwriteStringSlowthe run follows the current output coder, because a character above 0xFF upgrades the output to UTF16 in the middle of a string; the UTF16 runs (writeStringSlow,writeStringUtf16Slow) clearlatin1Outputwhen they copy such a character.RUN_CHUNK(8192) characters at a time, the waywriteEscapedUtf16already bounds its reservation, so(length - i) << 1orpos + (length - i)cannot overflow for a huge string and skipgrowbefore the word stores. A run that ends at a chunk boundary rather than at a stop just continues. The copy is inlined into each tail: a helper call per stop measurably slowed strings with many escapes.writeString(String)methods and their fixed-size fast paths are untouched; all new code is in the cold paths, following the C2-layout notes in both classes. The output is byte-identical.JsonStringTest:writeEscapeFollowedByOtherTextcompares both writers withJsonStringEscaper(a character-at-a-time reference) for a stop character (quote, backslash, newline, control character, latin1 above 0x7F) followed by latin1 text, 0x7F/0x80, text outside latin1, a surrogate pair and further escapes, with 0-20 characters before it and 0-8 after it so stops land in every lane of a word, with and without trailing text, forStringandCharSequenceinput (each on a fresh writer), utf8 output and LATIN1 and UTF16 string output, with and withoutescapeNonAscii. Accepting a quote inside the latin1 run, or ignoringescapeNonAsciiin the UTF16 run, makes it fail.writeRunsAcrossReservationChunks: a quote, latin1 above 0x7F, text outside latin1 and a surrogate pair at 8192 +/- 9 characters after an escape, so they land on and around a chunk boundary.writeRejectsUnpairedSurrogateAfterAnEscape: an unpaired high or low surrogate after an escape, inside a long plain run and at the end, is still rejected by both writers, with and withoutescapeNonAscii.writeEscapeThenWideTextAfterUtf16Reset: a writer reset after a UTF16 result starts in UTF16 while assuming its output fits latin1; the run copy must withdraw that assumption. Removing thelatin1Outputclear makes it fail.Related issues
AI Contribution Checklist
yesyes, I included a completed AI Contribution Checklist in this PR description and the requiredAI Usage Disclosure.yes, my PR description includes the requiredai_reviewsummary and screenshot evidence or equivalent persisted links of the final clean AI review results from both fresh reviewers described inAI_POLICY.md, the Fory-guided reviewer and the independent general reviewer, on the current PR diff or current HEAD after the latest code changes.Completed checklist from
AI_POLICY.mdsection 9:yesAI Usage Disclosureblock below.AGENTS.mdand.agents/ci-and-pr.md, and one independent general reviewer in a separate clean-context review session that was not pointed to.agents/ci-and-pr.mdor any copied Fory-specific review checklist.Verification
JDK 25.0.2, Windows 11 x86_64, from
java/:mvn -pl fory-json -am install -DskipTests: successmvn -pl fory-json test: 1276 tests, 0 failures, 0 errors, 0 skippedJsonStringTest#writeEscapeFollowedByOtherTextwith a quote accepted inside the latin1 run ofUtf8JsonWriter, and withescapeNonAsciiignored in theStringJsonWriterUTF16 run copy: fails, as intendedJsonStringTest#writeEscapeThenWideTextAfterUtf16Resetwith thelatin1Outputclear removed from the UTF16 run copy: fails (中comes back as-), as intendedmvn -pl fory-json spotless:check checkstyle:check: success, 0 Checkstyle violationsDoes this PR introduce any user-facing change?
Benchmark
Standalone harness from #4121 (not JMH): a bid whose
admfield is a 19998-character VAST XML string, pure ASCII, 612 quotes and 345 newlines, the first quote at character 15. Variants change only the characters that need escaping insideadm, keeping its length.ForyJson.builder().build(), best of 7 rounds of 5000 writes after 30000 warm-up writes, one JVM per build, builds alternating three times on an otherwise idle machine; JDK 25.0.2. Range over the three runs, microseconds per write:toJsonBytesmain f85877ftoJsonBytesthis PRtoJsonmaintoJsonthis PRA bean of eight short strings, four of them with one escape each: 286-291 ns -> 181-191 ns (
toJsonBytes), 255-272 ns -> 231-235 ns (toJson). The output digest is identical between the builds in every case. For reference, Jackson 3.1writeValueAsStringtakes 12.5 us on the string without escapes and 25.0 us on the real one.🤖 Generated with Claude Code