Skip to content

Check character codes before regexes in liquid-html-parser tokenizers - #1333

Open
PhilippeCollin wants to merge 1 commit into
liquid-html-parser-tokenizer-fast-pathfrom
liquid-html-parser-markup-fast-checks
Open

PhilippeCollin wants to merge 1 commit into
liquid-html-parser-tokenizer-fast-pathfrom
liquid-html-parser-markup-fast-checks

Conversation

@PhilippeCollin

Copy link
Copy Markdown

Stacked on #1329. Review that first; this PR's diff is only the last commit.

Why

After #1329, the profile shows two avoidable costs on characters that can't match:

  • scanLiquidOpen runs at every position where the text fast path stops (<, -, quotes, /, >, =) and makes four startsWith calls ({{-, {{, {%-, {%), even though all four need a {.
  • The Liquid expression tokenizer (tokenizeMarkup) runs the sticky whitespace regex before every token, and evaluates /\d/.test(ch) twice per token. A regex literal inside a loop creates a new RegExp object each time it's evaluated.

What

  • scanLiquidOpen returns false immediately unless the current character is {.
  • tokenizeMarkup tests digits by char code (\d without the u flag is ASCII 0–9 only). It only runs the whitespace regex when the character could be whitespace, since every \s character is at most U+0020 or at least U+00A0.

Each change adds a cheap pre-check in front of the existing startsWith or regex, which still makes the decision. Letting extra characters through therefore can't change tokens. Only rejecting a real { or ASCII digit could, and doing that on purpose fails between 389 and 1,268 existing tests. New tests pin non-ASCII digits and a trailing -.

Benchmarks

Apple M3 Pro, Node 24.15. The input is the parser's 429 fixture files (Dawn, Horizon, base theme; 3.4 MB). #1329 and this PR ran in alternating processes, 5 rounds × 30 passes. Each value is the median.

All 429 files #1329 this PR Time saved
tokenize 10.0 ms 9.2 ms 8%
toLiquidHtmlAST 42.1 ms 39.9 ms 5%
toLiquidAST 44.8 ms 41.7 ms 7%

Tokens and ASTs are unchanged. A differential run against #1329 over about 25,700 inputs (every .liquid file in the repo, plus random and mutated templates) found 0 differences, positions included. liquid-html-parser, prettier-plugin-liquid and theme-check-common tests pass.

Related: #1330 and #1331 are independent follow-ups from the same profiling.

- `scanLiquidOpen` returns early unless the character is `{`, instead
  of four failed `startsWith` calls at every `<`, `-`, quote, `/`, `>`
  and `=` the text fast path stops at.
- The Liquid expression tokenizer tests digits with char codes instead
  of evaluating `/\d/` literals per token, and only runs the whitespace
  regex when the character could be whitespace.

Each check is a pre-filter in front of the existing `startsWith` or
regex, so letting extra characters through can't change tokens; only
rejecting a real `{` or digit could, and the suite catches that.
@PhilippeCollin
PhilippeCollin added this pull request to stack #1335 October 9, 2026 13:52
@PhilippeCollin
PhilippeCollin marked this pull request as ready for review October 9, 2026 13:52
@PhilippeCollin
PhilippeCollin requested a review from a team as a code owner October 9, 2026 13:52

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant