Repository navigation
perf(vm): defer full bytecode generation until payload reuse - #137
Merged
Merged
Conversation
steipete
added a commit
to openclaw/openclaw
that referenced
this pull request
Oct 8, 2026
Advance the shared Bun pin from fc53 to 42bd for deferred VM bytecode generation, integral heap-sampling sizes, and inspector snapshot cleanup (openclaw/bun#137, openclaw/bun#140, openclaw/bun#141). Keep WebKit cb8d6f202b and update the four verified Darwin/Linux artifact pins and CI documentation. Gate A passed: 35 paired selections, 270 targeted cells, no candidate regressions, ten Node-hidden smoke steps with zero Node attempts, and supplemental JSC checks. Both Darwin architecture gates passed. Preserve the baseline-only UI palette intermittent in #167000; #166780 is not claimed fixed. No new admissions qualified. Pin staging tests, changed checks, independent P2 review, exact-head CI, and ClawSweeper passed without a CI rerun.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Repeated
node:vmcompilations currently generate a separate, recursively compiled tree on the first private-cache hit, even when the original block is still alive. Snapshot the function bodies already generated by execution on that live hit, and defer full generation until an incomplete payload is actually reused after GC.Record snapshot completeness from the encoder's existing leaf map while GC is deferred, then discard the raw pointer map as before. This covers nested functions and preserves still-lazy decoded bytes. Each entry attempts full promotion at most once. The 256 MiB budget, exact-source/key validation, immutable payload ownership and weak decoded shortcuts stay in place. Public
createCachedData/produceCachedDatakeeps source-based full generation.How did you verify your code works?
Matched release/LTO builds on one Linux x64 c7a.16xlarge, pinned to eight CPUs with four non-isolated Vitest workers. OpenClaw is frozen at
497de2e869f6a329fac4acc571ed46d6b4454521, baseline Bun atfc53bf8c0fccd0dcbcccfcf75e2c306ab8953cae, paired WebKit atcb8d6f202b5a396caa204ee1bb75d78175aa841a, and Node at 24.21.0. Every process uses fresh private state/cache/tmp directories. Full outcomes, worker policy, zero steal and descendant settlement gate qualified observations.Discord has three qualified Node/baseline processes and four candidate processes; providers has three per arm. Wall ratios are 0.9414 (95% process-bootstrap interval 0.9402–0.9524) and 0.9683 (0.9627–0.9710). Two Discord controls with CPU-steal ticks remain recorded and excluded; a predeclared balanced round supplies the required qualified counts. RSS is the process maximum; its median passes the +5% gate, but these small samples do not establish a universal memory bound. Bun remains approximately 22% and 12% behind Node on these configs.
Separate native profiles show parse/bytecode CPU falling from 81.3 to 58.6 seconds in Discord (−27.9%) and from 63.5 to 50.3 seconds in providers (−20.7%). Inclusive full-generation cost falls 43.0→21.2 and 25.9→13.4 seconds. Samples are normalized to diagnostic GNU CPU, with roughly 96–97% coverage and no lost events; these profiles are not pooled with ordinary timings. One providers diagnostic with a steal tick was retained and excluded, and its separately named retry qualified.
Upstream research found no direct patch for this fork-private cache.
oven-sh/bun#40174concerns the module disk cache; this change requires no engine API or artifact update.Hosted Linux x64 and macOS arm64 build/test CI passed on exact head
68d3c8b1b0c0fda75778b14d56995c946ad58f65. Formatting, JavaScript lint and source-lints also passed. The ancillary issue-finder could not start analysis because its provider credentials are not configured.