Skip to content

fix: bound benchmark scenario wall-clock time to fit CI's job timeout - #1707

Merged
johnnyreilly merged 1 commit into
mainfrom
benchmark-scenario-time-budget
Sep 2, 2026
Merged

johnnyreilly merged 1 commit into
mainfrom
benchmark-scenario-time-budget

Conversation

@johnnyreilly

Copy link
Copy Markdown
Member

Summary

Ports the wall-clock safety cap from the copilot/implement-new-tsgo-api-support branch's benchmark harness (commit b9efbc4) onto main, so both branches' benchmark mechanisms match.

  • Caps each benchmark scenario's wall-clock time in run-side.mts (60s, after a minimum sample floor) instead of relying solely on a fixed iteration count.
  • Cheap scenarios are unaffected (they already finish under the cap); scenarios with an unexpectedly high per-iteration cost stop once they've collected enough measured samples, rather than risking the whole job blowing past the 20-minute CI timeout.
  • On main today (classic ts-loader, cheap in-process instantiation) this cap essentially never triggers - verified with a local run against this branch's own checkout, where every scenario completed at its full iteration count with no early cutoff. It's here as a safety net: if a future change (here or in a dependency) makes some scenario meaningfully more expensive per iteration, the benchmark degrades gracefully (fewer samples, still reports) instead of timing out with no results at all - exactly what happened on the tsgo branch before this cap was added there.

Test plan

  • yarn lint passes
  • yarn build passes (TypeScript 6.0.2)
  • npx tsc --project test/benchmark-tests/tsconfig.json --noEmit passes
  • Ran node test/benchmark-tests/run-benchmark.mts locally (self-comparison) - all 6 scenarios completed at full iteration count, no early cutoffs, deltas near zero as expected

🤖 Generated with Claude Code

The tsgo sync API spawns a native child process per ts-loader instance,
making a "cold build" iteration ~10-20x more expensive than the classic
API's cheap in-process instantiation; touching a widely-imported (hub)
file is similarly ~35x more expensive per incremental rebuild due to
per-dependant recheck round trips over the sync RPC channel. The fixed
iteration counts (tuned for the classic API's cost profile) made the
cold-typeCheck and hub-touch scenarios alone take ~28 minutes combined,
blowing the 20-minute CI job timeout before a single scenario finished.

Cap each scenario's wall-clock time in run-side.mts instead of guessing
a smaller fixed iteration count, so cheap scenarios keep their full
sample size while expensive ones stop once they've collected enough
measured samples. Confirmed locally: full default run (300 files) went
from never finishing to completing in ~4 minutes.
@github-actions

github-actions Bot commented Sep 2, 2026

Copy link
Copy Markdown

Benchmark (Ubuntu)

Scenario transpileOnly PR branch median (ms) base branch median (ms) Δ vs base branch
Cold build false 901.0 901.0 -0.0%
Incremental rebuild (leaf touch) false 62.8 61.6 +1.9%
Incremental rebuild (hub touch) false 358.2 355.8 +0.7%
Cold build true 545.2 546.0 -0.2%
Incremental rebuild (leaf touch) true 32.1 31.7 +1.4%
Incremental rebuild (hub touch) true 32.8 32.2 +1.9%

PR branch = /home/runner/work/ts-loader/ts-loader, base branch = /home/runner/work/ts-loader/ts-loader-main. 2 warmup + 10 measured iterations per scenario, median reported. Report-only - no threshold fails this check.

@github-actions

github-actions Bot commented Sep 2, 2026

Copy link
Copy Markdown

Benchmark (Windows)

Scenario transpileOnly PR branch median (ms) base branch median (ms) Δ vs base branch
Cold build false 1390.0 1401.8 -0.8%
Incremental rebuild (leaf touch) false 101.3 98.8 +2.6%
Incremental rebuild (hub touch) false 643.4 664.1 -3.1%
Cold build true 796.5 787.9 +1.1%
Incremental rebuild (leaf touch) true 40.6 39.6 +2.3%
Incremental rebuild (hub touch) true 39.5 40.2 -2.0%

PR branch = C:\source\ts-loader, base branch = C:\source\ts-loader-main. 2 warmup + 10 measured iterations per scenario, median reported. Report-only - no threshold fails this check.

@johnnyreilly
johnnyreilly merged commit 7748b9f into main Sep 2, 2026
190 of 191 checks passed
@johnnyreilly
johnnyreilly deleted the benchmark-scenario-time-budget branch September 2, 2026 11:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant