Add opt-in Ractor-local GC metrics to the Ractor harness - #537
eightbitraptor wants to merge 7 commits into
Conversation
Add gc_total_time_warmup/bench in milliseconds and gc_global_count_warmup/bench (on Ruby versions with support for global_gc_count) to the JSON output, plus a global column in the per-iteration stdout table. Ignore this if the benchmarked Ruby doesn't support global gc counts (introduced during Ruby 4.1 dev cycle).
enable with RUBY_BENCH_RACTOR_GC=1 The mode requires Ruby 4.1 or newer (because of the Ractor-local GC.stat and per-Ractor GC.measure_total_time);
Passing --ractor-gc to run_benchmarks.rb now sets RUBY_BENCH_RACTOR_GC=1
|
I think having a miscount of |
|
@luke-gruber ruby/ruby#19147 I went a different way, by tracking the number of globals initiated by each Ractor. We can then subtract that from the major count to get the true local major count. It's difficult and I guess we could go one way or the other but I opted not to just discount globals as majors, because they do run a major GC. I'm not wedded to this decision though. |
|
I don't know, that seems kind of hacky. The fact that global GC does a pseudo-"major" GC on each objspace is an implementation detail that imo shouldn't leak out to Edit: to be honest, I don't really care how they're measured so if it's easier to do it that way I'm fine with it. What matters is getting good benchmarks and then fixing the GC issues. |
GC.stat(:global_gc_count) now counts the global cycles the calling Ractor initiated, and global cycles no longer count as majors: count == minor + major + global in both scopes. This removes the overlap that forced the controller-observed global column. gc_global_count joins the worker-sum whitelist, the tables show a plain additive global/iter (worker sum) column and an unstarred global/iter ratio, and minor GC % divides by the total count. Only controller compacts/iter keeps its star: compact_count increments in every object space, so it still cannot be summed across workers. The preflight gains a fork-based probe (the benchmark process never creates a Ractor, so multi-ractor mode stays unarmed): a child arms a Ractor, runs one explicit global GC.start, and the child's main Ractor must see no scoped delta while the process total moves. Builds without the attribution change fail before warmup with NotImplementedError.
|
This is very useful for our work, so I'm +1 on merging it. It's a lot of code added to ruby-bench, but it's self-contained enough that I think it's okay. |
./run_benchmarks.rb --category ractor --ractor-gcsamplesGC.statand GCtotal time in every worker Ractor's own object space on each measured
iteration. The JSON output records raw per-worker samples, worker-sum series,
a controller-observed compaction count, and the measurement scope:
gc_scope: "ractor-local-workload",gc_stat_scope: "ractor-local",gc_measure_total_time_scope: "ractor-local", plus the target'sgc_config.The target needs Ruby 4.1 or newer with per-Ractor global GC attribution
(ruby/ruby#19147). Older targets fail before warmup with
NotImplementedError; a fork-based probe checks the semantics without thebenchmark process ever creating a Ractor.
What it looks like
Per-iteration harness output:
Single-executable summary (absolute columns):
Comparison summary (ratios, then base → comparison counts):
How the numbers compose
With ruby/ruby#19147,
GC.stat(:global_gc_count)counts the global(stop-the-world) cycles the calling Ractor initiated, and global cycles no
longer count as majors. Every worker-sum column is additive:
global/iteris the global cycles the sampled workers initiated. No overlap,no caveat paragraph.
One exception keeps its star:
controller compacts/iter*.compact_countincrements in every object space during a global compacting cycle, so it can
never be worker-summed. We record the main Ractor's delta once per iteration.
The entire
GC metric notes:block is one sentence:Failure handling and alignment
settings are restored. Covered by subprocess tests against a real build.
Indexes stay aligned, and the table renders N/A. Absence is never zeroed.
Feedback wanted
This is a draft mainly to get eyes on the presentation:
(worker sum)suffixes,GC ms/worker, and the remainingstarred
controller compacts/iter*. Do these read well?GC ms/workerdivides each iteration's worker-sum GC time by its sampledworker count, then takes the mean. Is the mean enough, or do we want
max/median?
GitHub comment box, and these columns break that. Layout suggestions are
very welcome.