Skip to content

Speed up liquid-html-parser by tokenizing plain text in one step - #1329

Open
PhilippeCollin wants to merge 2 commits into
mainfrom
liquid-html-parser-tokenizer-fast-path
Open

PhilippeCollin wants to merge 2 commits into
mainfrom
liquid-html-parser-tokenizer-fast-path

Conversation

@PhilippeCollin

@PhilippeCollin PhilippeCollin commented Oct 8, 2026 •

Copy link
Copy Markdown

Why

Profiling toLiquidHtmlAST on real themes shows the document tokenizer takes about two thirds of parse time. It moves forward one character at a time and runs up to eight startsWith checks on every character, including plain text, attribute values and Liquid expressions. Most of those characters can't start a token: in normal HTML content, only {, < or - can.

What

When the tokenizer reaches a character that is plain text, it now skips to the next character that could start a token in the current mode, instead of moving forward by one:

Mode Stops at
HTML content {, <, -
Inside an HTML tag {, /, >, =, straight or curly quotes
Quoted attribute value {, either quote of the open pair
Liquid tag / output body first %} / }} (via indexOf), or the - just before it

All of the existing match() checks stay as they are. If a stop position turns out not to start a token, the main loop treats it as text, exactly as before. The < / </ tag-name check also uses char codes now instead of a regex.

The risky direction is a helper that misses a stop character: the tokenizer would then read a token as text. To make that testable, the five text branches go through one dispatcher, and a test-only tokenizeWithoutFastPath (not exported from the package) runs the same loop one character at a time, like main.

Tokens and ASTs are unchanged. The tokenizer's design is the same; only how it moves through text changed.

Benchmarks

Apple M3 Pro, Node 24.15. The input is every .liquid file from the parser's bench fixtures (scripts/download-themes.ts): Dawn, Horizon and the base theme, 429 files, 3.4 MB. Baseline and branch builds ran in alternating processes, 5 rounds × 30 passes. Each value is the median.

Workload main this PR Time saved Speedup
All themes, tokenize 63.0 ms 10.1 ms 84% 6.3×
All themes, toLiquidHtmlAST 96.7 ms 42.6 ms 56% 2.3×
All themes, toLiquidAST 98.8 ms 44.7 ms 55% 2.2×
Dawn (88 files), toLiquidHtmlAST 26.1 ms 12.4 ms 52% 2.1×
Horizon (283 files), toLiquidHtmlAST 61.1 ms 23.8 ms 61% 2.6×
Base theme (58 files), toLiquidHtmlAST 9.1 ms 5.0 ms 45% 1.8×
Largest file (horizon/snippets/icon.liquid, 130 KB) 2.3 ms 0.3 ms 87% 7.5×

The repo's own parser.bench.ts (vitest bench) shows the same result. full-theme-parse went from 108.7 ms to 55.3 ms (2.0×). Under per-file-parse, all 429 files got faster: the geometric mean is 2.0×, the minimum 1.1× and the maximum 5.0×.

Prettier: the plugin's parse step (toTolerantLiquidHtmlAST) drops from 122 ms to 64 ms per pass over the same files. Formatting time is mostly printing, so end-to-end prettier.format improves by roughly 2–7% when the two parsers alternate in one process. That is close to the noise between separate runs. The formatted output is byte-identical.

Worst case: the change stays linear. The one input I found that gets slower is 200K characters that are all - (5.4 ms → 5.8 ms), because every character is a candidate. Inputs that are all lone < or { get about 2× faster.

Checking that output is identical

Enforced in CI (tokenizer unit tests, no fixtures needed, about 0.7 s):

  • Each token start sits in the middle of plain text, for every mode and entry state: Default, inside a tag, each of the six opening quotes, and Liquid tag and output bodies. For example, <div class= "a b"> puts a quote after a space instead of right after =.
  • tokenize matches tokenizeWithoutFastPath on every string of up to 4 characters drawn from { } % - < > / = ! " ' “ ” ‘ ’ ‚ a, space and newline. That is 137,561 strings, each checked in 11 entry states (each TokenizeOptions variant, plus inside {% and {{ ).
  • I planted 17 bugs in the helpers and the dispatcher: removing each stop character, the - lookback, or the open/close quote of the pair, and pointing a mode at the wrong helper. Each one fails these tests without the theme fixtures, with 5 to 278 tests failing.

Local only:

  • liquid-html-parser tests: 2771 passed, 4 skipped. That includes the parser-oracle and tolerant-corpus suites, with golden ASTs generated from main for all 429 theme files. CI skips those suites because the fixtures are gitignored and never downloaded, and they compare ASTs without positions. prettier-plugin-liquid and theme-check-common tests also pass.
  • A differential harness (not committed) compared main and this branch on about 25,700 inputs and found 0 differences. It checked tokenize output under every option and full ASTs, positions included, from toLiquidHtmlAST, toLiquidAST and both tolerant entry points. The inputs were every .liquid file in this repo, 20K random strings built from Liquid and HTML fragments, and 5K truncated or spliced theme files.

The document tokenizer advanced one character at a time and ran every
mode's startsWith checks on each one, even inside plain text. Most
characters cannot start a token in the current mode, so jump straight to
the next one that can and let the main loop re-check it there.

Liquid tag and output bodies use indexOf for their closing delimiter.
The `<` tag-start check uses char codes instead of a regex.

Tokens and ASTs are unchanged.
The existing tests only had quotes directly after `=`, where the main
loop checks them without the fast path, so a helper missing a stop
character could pass CI. The oracle suites that might catch it are
skipped there because the theme fixtures are gitignored.

Route the five text branches through one `nextTextCandidate` dispatcher,
so a test-only `tokenizeWithoutFastPath` can run the same loop one
character at a time. Tests now:
- put each token start in the middle of plain text in every mode
- compare `tokenize` with `tokenizeWithoutFastPath` on every string of
  up to 4 token-relevant characters, in every entry state

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant