Repository navigation
Trim text around Liquid tags in linear time in liquid-html-parser - #1330
Open
PhilippeCollin wants to merge 1 commit into
Open
PhilippeCollin wants to merge 1 commit into
PhilippeCollin wants to merge 1 commit into
Conversation
`value.replace(/\s+$/, '')` retries the match at every position of each whitespace run, so its cost grows with the square of the run length. It was the most expensive single step when parsing Dawn and Horizon. `trimStart()`/`trimEnd()` strip exactly the characters `\s` matches (both use the spec's WhiteSpace and LineTerminator sets), in linear time. A test pins that set and the long-run case.
This was referenced Oct 8, 2026
PhilippeCollin
marked this pull request as ready for review
October 9, 2026 12:58
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Profiling
toLiquidHtmlASTon Dawn, Horizon and the base theme showed one regular expression taking about 15% of all parse time. It isvalue.replace(/\s+$/, ''), which trims trailing whitespace from text next to Liquid tags.That pattern isn't anchored at the start, so the regex engine tries a match at every position inside each whitespace run, and each attempt reads to the end of the run before failing at
$. A run of k spaces costs about k² steps. Raw-tag bodies such as indented{% schema %}JSON or{% style %}CSS contain many such runs.What
The trimming now uses
trimEnd(), andtrimStart()for the leading side. These strip exactly the characters\smatches, because the spec defines both with the same WhiteSpace and LineTerminator sets, and they run in linear time. A new test pins that character set, including a zero-width space that must be kept. A second test parses text containing a 100,000-space run within 1 s; onmainit takes about 5 s.Benchmarks
Apple M3 Pro, Node 24.15. The input is the parser's 429 fixture files (Dawn, Horizon, base theme; 3.4 MB). Builds ran in alternating processes, 5 rounds × 30 passes. Each value is the median.
main:toLiquidHtmlASTmain:toLiquidASTtoLiquidHtmlASTtoLiquidASTThe relative gain is larger with #1329, because that PR removes most of the tokenizer time that otherwise dominates. A text node with a long whitespace run also stops scaling quadratically: an 80,000-space run takes 3.5 s on
mainand 3 ms with this change.ASTs are unchanged.
liquid-html-parser(including the local fixture oracle suites),prettier-plugin-liquidandtheme-check-commontests pass. A differential run againstmainover about 25,700 inputs (every.liquidfile in the repo, plus random and mutated templates) found 0 differences in tokens or ASTs, positions included.Related: #1329 (tokenizer fast path). Found while profiling after it.