Skip to content

Flatten WeakTopologicalOrdering into a single entry vector - #9220

Merged
tlively merged 56 commits into
mainfrom
wto-flat
Oct 9, 2026
Merged

tlively merged 56 commits into
mainfrom
wto-flat

Conversation

@tlively

@tlively tlively commented Oct 6, 2026

Copy link
Copy Markdown
Member

Replace the recursive std::variant/Cycle representation of
WeakTopologicalOrdering with a single contiguous std::vector
where cycle ends store the header block and jump target index. This
avoids per-cycle vector allocations and recursive evaluation in
WTOWorklist::run.

Benchmark results across 16 WebAssembly modules (3 iterations,
interleaved):

  • --constraint-analysis:
    • Geomean: 1.606s -> 1.592s (-0.9%)
    • Total time: 64.05s -> 63.67s (-0.6%; faster on 11/16 modules)
  • --rse:
    • Geomean: 0.871s -> 0.852s (-2.2%)
    • Total time: 26.03s -> 25.39s (-2.5%; faster on 12/16 modules)

Add `src/cfg/wto.h` with `WeakTopologicalOrdering` (`WTO`) and `WTOWorklist` built on top of `DomTree`. In a reducible CFG ordered in reverse postorder, every cycle is a natural loop headed by a block that dominates all blocks in the cycle, allowing a Bourdoncle Weak Topological Ordering to be constructed directly from the dominator tree and natural loops of the CFG.

Include unit tests in `test/gtest/wto.cpp` and `TODO` comments noting follow-on optimizations.
Replace RPOQueue with WTOWorklist in ConstraintAnalysis and
RedundantSetElimination so that loops stabilize before flow values
propagate to downstream blocks. This avoids quadratic/cubic blowups on
functions with sequential loops while also speeding up general workloads.

Benchmark results across 16 WebAssembly modules (3 iterations):
- --constraint-analysis:
  - esbuild.wasm: 381.60s -> 7.59s (-98.0%, 50.3x speedup)
  - 15 non-esbuild modules geomean: 1.564s -> 1.487s (-4.9%)
  - 15 non-esbuild modules total: 72.00s -> 57.24s (-20.5%)
  - All 16 modules geomean: 2.205s -> 1.646s (-25.4%)
  - All 16 modules total: 453.60s -> 64.83s (-85.7%)
- --rse:
  - esbuild.wasm: >600s (timeout) -> 1.62s (>370x speedup)
  - 15 non-esbuild modules total: 26.40s -> 25.47s (-3.5%)
  - All 16 modules geomean: N/A -> 0.922s (total: 27.09s)
Read each basic block's reverse-postorder index from contents.index in
DomTree instead of allocating and populating an
unordered_map<BasicBlock*, Index>, and skip self-loop backedges
immediately with predIndex >= index. Update OnceReduction and
test/example/domtree.cpp to initialize contents.index, and remove the
redundant index initialization loop in WeakTopologicalOrdering.

Benchmark results across 16 WebAssembly modules (3 iterations,
interleaved):
- --constraint-analysis:
  - Geomean: 1.646s -> 1.576s (-4.3%)
  - Total time: 64.83s -> 62.82s (-3.1%)
- --rse:
  - Geomean: 0.922s -> 0.859s (-6.8%)
  - Total time: 27.09s -> 25.40s (-6.3%)
During reverse-RPO natural loop discovery in WeakTopologicalOrdering,
collapse each discovered loop body into its header using union-find with
path compression, and skip over already-collapsed inner loops when
walking immediate dominators in dominates(). This prevents outer loops
from re-traversing inner loop bodies, bounding natural loop discovery to
O(E alpha(N)) instead of O(N * depth) on deeply nested loops.

Benchmark results across 16 WebAssembly modules (3 iterations,
interleaved):
- --constraint-analysis:
  - Geomean: 1.576s -> 1.564s (-0.7%)
  - Total time: 62.82s -> 62.40s (-0.7%)
- --rse:
  - Geomean: 0.859s -> 0.853s (-0.7%)
  - Total time: 25.40s -> 25.06s (-1.3%)
When the CFG has no backedges (checked via CFGWalker::loopTops),
evaluate queued blocks in a single reverse-postorder pass in
WTOWorklist::run without constructing DomTree or
WeakTopologicalOrdering.

Benchmark results across 16 WebAssembly modules (3 iterations,
interleaved):
- --constraint-analysis:
  - Geomean: 1.610s -> 1.606s (-0.2%)
  - Total time: 64.36s -> 64.05s (-0.5%; dart_essentials: 2.92s -> 2.59s, -11.6%)
- --rse:
  - Geomean: 0.870s -> 0.871s (+0.1%)
  - Total time: 26.10s -> 26.03s (-0.3%; dart_essentials: 2.00s -> 1.71s, -14.4%)
Replace the recursive std::variant/Cycle representation of
WeakTopologicalOrdering with a single contiguous std::vector<Entry>
where cycle ends store the header block and jump target index. This
avoids per-cycle vector allocations and recursive evaluation in
WTOWorklist::run.

Benchmark results across 16 WebAssembly modules (3 iterations,
interleaved):
- --constraint-analysis:
  - Geomean: 1.606s -> 1.592s (-0.9%)
  - Total time: 64.05s -> 63.67s (-0.6%; faster on 11/16 modules)
- --rse:
  - Geomean: 0.871s -> 0.852s (-2.2%)
  - Total time: 26.03s -> 25.39s (-2.5%; faster on 12/16 modules)
@tlively
tlively requested a review from kripken October 6, 2026 07:10
@tlively
tlively requested a review from a team as a code owner October 6, 2026 07:10
Comment thread src/cfg/wto.h Outdated
// end of the cycle headed by `block` (`cycleTarget` is the entry index of the
// cycle header).
struct Entry {
BasicBlock* block = nullptr;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
BasicBlock* block = nullptr;
BasicBlock* block;

This made me think it was optional, and I don't see a benefit to this default?

Base automatically changed from wto-fast-paths to main October 9, 2026 18:17

@kripken kripken left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, comments look very clear now!

@tlively
tlively enabled auto-merge (squash) October 9, 2026 23:05
@tlively
tlively merged commit 6de3a0b into main Oct 9, 2026
16 checks passed
@tlively
tlively deleted the wto-flat branch October 9, 2026 23:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants