Repository navigation
Accept whitespace around = in HTML attributes - #1332
Merged
Merged
Conversation
The parser read `data-ratio = '{{ r }}'` and `srcset= "…"` as a valueless
attribute followed by junk attributes built from the value. Browsers and
the old ohm grammar accept whitespace on either side of `=`.
The tokenizer now emits each whitespace run inside a tag as its own Text
token, and the parser skips those tokens between attributes, before `=`,
and before the value. Both attribute loops share `parseAttributeAfterName`.
`scanForHtmlCloseTag` now compares every Text token between `</` and `>`,
so `</script marker>` with a non-breaking space stays body text.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
charlespwd
approved these changes
Oct 9, 2026
EvilGenius13
approved these changes
Oct 9, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What are you adding in this PR?
@shopify/liquid-html-parsernow accepts whitespace around=in HTML attributes, as browsers and the old ohm grammar do.Before, the recursive-descent parser ended the attribute name at the space and found no
=. It then turned each part of the value into an attribute of its own:<img data-ratio = '{{ r }}'>data-ratio(empty),{{ r }}data-ratio='{{ r }}'<img srcset= "{{ a }}, {{ b }} 2x">srcset=(empty),{{ a }},,,{{ b }},2xsrcset="{{ a }}, {{ b }} 2x"Dawn 5's footer has the
srcset= "…"form, so this shows up in real themes.Changes:
Texttoken.=, and before the value. This matches the old grammar's syntactic attribute rules (AttrDoubleQuoted = attrName "=" doubleQuote …), which skippedspaceimplicitly. Both attribute loops (element and Liquid-branch) share a newparseAttributeAfterName, so curly-quoted values inside an{% if %}in attribute position now get the same double/single mapping as top-level ones. The parser no longer modifiestoken.start.scanForHtmlCloseTag: compares everyTexttoken between</and>, not just the first one. Without this, the new tokenization would make</script marker>(with a non-breaking space) inside a<script>or<style>body close the tag early and throw. That text is body text, both before this PR and in browsers.What's next? Any followup issues?
These cases still differ from the HTML tokenizer. None of them changed in this PR:
title=Don't)a==banda=b=cshould give the values=bandb=ca/bshould split intoaandbTophatting
html.test.ts: whitespace on either side of=, newlines around=, unquoted values, attributes inside a Liquid branch, and NBSP close-tag candidates inscriptandstyle. Each fails without its fix.tokenizer.test.tsupdated for the split whitespace tokens.liquid-html-parser,prettier-plugin-liquid,theme-check-commonandtheme-language-server-commonpass locally (7,152 tests).data-ratio = 'x'asdata-ratio='x'. It used to break the attribute apart.Before you deploy
changeset🤖 Generated with Claude Code