Repository navigation
Speed up liquid-html-parser by tokenizing plain text in one step - #1329
Open
PhilippeCollin wants to merge 2 commits into
Open
PhilippeCollin wants to merge 2 commits into
PhilippeCollin wants to merge 2 commits into
Conversation
The document tokenizer advanced one character at a time and ran every mode's startsWith checks on each one, even inside plain text. Most characters cannot start a token in the current mode, so jump straight to the next one that can and let the main loop re-check it there. Liquid tag and output bodies use indexOf for their closing delimiter. The `<` tag-start check uses char codes instead of a regex. Tokens and ASTs are unchanged.
The existing tests only had quotes directly after `=`, where the main loop checks them without the fast path, so a helper missing a stop character could pass CI. The oracle suites that might catch it are skipped there because the theme fixtures are gitignored. Route the five text branches through one `nextTextCandidate` dispatcher, so a test-only `tokenizeWithoutFastPath` can run the same loop one character at a time. Tests now: - put each token start in the middle of plain text in every mode - compare `tokenize` with `tokenizeWithoutFastPath` on every string of up to 4 token-relevant characters, in every entry state
PhilippeCollin
marked this pull request as ready for review
October 8, 2026 22:28
PhilippeCollin
added this pull request to stack #1335
October 9, 2026 13:52
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Profiling
toLiquidHtmlASTon real themes shows the document tokenizer takes about two thirds of parse time. It moves forward one character at a time and runs up to eightstartsWithchecks on every character, including plain text, attribute values and Liquid expressions. Most of those characters can't start a token: in normal HTML content, only{,<or-can.What
When the tokenizer reaches a character that is plain text, it now skips to the next character that could start a token in the current mode, instead of moving forward by one:
{,<,-{,/,>,=, straight or curly quotes{, either quote of the open pair%}/}}(viaindexOf), or the-just before itAll of the existing
match()checks stay as they are. If a stop position turns out not to start a token, the main loop treats it as text, exactly as before. The</</tag-name check also uses char codes now instead of a regex.The risky direction is a helper that misses a stop character: the tokenizer would then read a token as text. To make that testable, the five text branches go through one dispatcher, and a test-only
tokenizeWithoutFastPath(not exported from the package) runs the same loop one character at a time, likemain.Tokens and ASTs are unchanged. The tokenizer's design is the same; only how it moves through text changed.
Benchmarks
Apple M3 Pro, Node 24.15. The input is every
.liquidfile from the parser's bench fixtures (scripts/download-themes.ts): Dawn, Horizon and the base theme, 429 files, 3.4 MB. Baseline and branch builds ran in alternating processes, 5 rounds × 30 passes. Each value is the median.maintokenizetoLiquidHtmlASTtoLiquidASTtoLiquidHtmlASTtoLiquidHtmlASTtoLiquidHtmlASThorizon/snippets/icon.liquid, 130 KB)The repo's own
parser.bench.ts(vitest bench) shows the same result.full-theme-parsewent from 108.7 ms to 55.3 ms (2.0×). Underper-file-parse, all 429 files got faster: the geometric mean is 2.0×, the minimum 1.1× and the maximum 5.0×.Prettier: the plugin's parse step (
toTolerantLiquidHtmlAST) drops from 122 ms to 64 ms per pass over the same files. Formatting time is mostly printing, so end-to-endprettier.formatimproves by roughly 2–7% when the two parsers alternate in one process. That is close to the noise between separate runs. The formatted output is byte-identical.Worst case: the change stays linear. The one input I found that gets slower is 200K characters that are all
-(5.4 ms → 5.8 ms), because every character is a candidate. Inputs that are all lone<or{get about 2× faster.Checking that output is identical
Enforced in CI (tokenizer unit tests, no fixtures needed, about 0.7 s):
<div class= "a b">puts a quote after a space instead of right after=.tokenizematchestokenizeWithoutFastPathon every string of up to 4 characters drawn from{ } % - < > / = ! " ' “ ” ‘ ’ ‚ a, space and newline. That is 137,561 strings, each checked in 11 entry states (eachTokenizeOptionsvariant, plus inside{%and{{).-lookback, or the open/close quote of the pair, and pointing a mode at the wrong helper. Each one fails these tests without the theme fixtures, with 5 to 278 tests failing.Local only:
liquid-html-parsertests: 2771 passed, 4 skipped. That includes the parser-oracle and tolerant-corpus suites, with golden ASTs generated frommainfor all 429 theme files. CI skips those suites because the fixtures are gitignored and never downloaded, and they compare ASTs without positions.prettier-plugin-liquidandtheme-check-commontests also pass.mainand this branch on about 25,700 inputs and found 0 differences. It checkedtokenizeoutput under every option and full ASTs, positions included, fromtoLiquidHtmlAST,toLiquidASTand both tolerant entry points. The inputs were every.liquidfile in this repo, 20K random strings built from Liquid and HTML fragments, and 5K truncated or spliced theme files.