Repository navigation
Check character codes before regexes in liquid-html-parser tokenizers - #1333
Open
PhilippeCollin wants to merge 1 commit into
Open
PhilippeCollin wants to merge 1 commit into
PhilippeCollin wants to merge 1 commit into
Conversation
- `scanLiquidOpen` returns early unless the character is `{`, instead
of four failed `startsWith` calls at every `<`, `-`, quote, `/`, `>`
and `=` the text fast path stops at.
- The Liquid expression tokenizer tests digits with char codes instead
of evaluating `/\d/` literals per token, and only runs the whitespace
regex when the character could be whitespace.
Each check is a pre-filter in front of the existing `startsWith` or
regex, so letting extra characters through can't change tokens; only
rejecting a real `{` or digit could, and the suite catches that.
PhilippeCollin
added this pull request to stack #1335
October 9, 2026 13:52
PhilippeCollin
marked this pull request as ready for review
October 9, 2026 13:52
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #1329. Review that first; this PR's diff is only the last commit.
Why
After #1329, the profile shows two avoidable costs on characters that can't match:
scanLiquidOpenruns at every position where the text fast path stops (<,-, quotes,/,>,=) and makes fourstartsWithcalls ({{-,{{,{%-,{%), even though all four need a{.tokenizeMarkup) runs the sticky whitespace regex before every token, and evaluates/\d/.test(ch)twice per token. A regex literal inside a loop creates a newRegExpobject each time it's evaluated.What
scanLiquidOpenreturnsfalseimmediately unless the current character is{.tokenizeMarkuptests digits by char code (\dwithout theuflag is ASCII 0–9 only). It only runs the whitespace regex when the character could be whitespace, since every\scharacter is at most U+0020 or at least U+00A0.Each change adds a cheap pre-check in front of the existing
startsWithor regex, which still makes the decision. Letting extra characters through therefore can't change tokens. Only rejecting a real{or ASCII digit could, and doing that on purpose fails between 389 and 1,268 existing tests. New tests pin non-ASCII digits and a trailing-.Benchmarks
Apple M3 Pro, Node 24.15. The input is the parser's 429 fixture files (Dawn, Horizon, base theme; 3.4 MB). #1329 and this PR ran in alternating processes, 5 rounds × 30 passes. Each value is the median.
tokenizetoLiquidHtmlASTtoLiquidASTTokens and ASTs are unchanged. A differential run against #1329 over about 25,700 inputs (every
.liquidfile in the repo, plus random and mutated templates) found 0 differences, positions included.liquid-html-parser,prettier-plugin-liquidandtheme-check-commontests pass.Related: #1330 and #1331 are independent follow-ups from the same profiling.