Repository navigation
Gate agentic inline accuracy on the entry the client writes - #115
Merged
Merged
Conversation
The inline gate looked up an accuracy entry named agentic_combined, which is the performance dataset's name in the Kimi K3 and DSV4 example configs. The client never writes that name: it records the inline score under dataset_name "performance", tagged dataset_type "performance", whatever the performance dataset is called. No real submission ever reached the gate, so inline accuracy was never checked for Kimi K3, Qwen3.6-35B-A3B or DeepSeek-V4.1-Flash, and a failing score only raised a warning. Find the entry by dataset_type, as the client asks of consumers. An untyped entry named performance, from clients that predate dataset_type, still counts. An accuracy file with more than one dataset_type performance entry is now invalid; the client scores the performance run once and refuses more than one performance dataset. Tests use the client's own output shape. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅ |
The README's agentic gate table said inline accuracy is "per point — every point must clear", which reads as an accuracy run being required at every Pareto point. The checker requires one per mandatory region (accuracy-coverage) and gates whatever results are present. Say so in the README, agentic_targets and the inline gate's docstring. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The agentic inline accuracy gate never runs on a real submission. It looks up an accuracy entry named
agentic_combined, which is the performance dataset's name in the Kimi K3 and DSV4 example configs (Qwen's isagentic_coding). The reference client never writes that name: it records the inline score underdataset_name: "performance", taggeddataset_type: "performance", whatever the performance dataset is called (execute.py, accuracy.py). For Kimi K3, Qwen3.6-35B-A3B and DeepSeek-V4.1-Flash the gate therefore only warned "No point reports aagentic_combinedaccuracy score", and a submission below the inline threshold passed without--strict.Fix. The inline entry is found by
dataset_type: "performance", which is how the client's ownaverage_accuracyidentifies it and how its comments ask consumers to filter. An untyped entry namedperformance, from clients that predatedataset_type, still counts. An entry namedperformancebut typedaccuracydoes not.New validation. An accuracy file with more than one
dataset_type: "performance"entry is now anaccuracy-validerror. The client scores the performance run once and rejects configs with more than one performance dataset, so no real output trips it.Scope. The lookup is used only by the agentic inline gate. The §15 single-turn gate, SWE-bench mean-of-N and OSL range are unchanged.
Tests. The fixture now writes the client's shape. New cases cover a native
accuracy_scoreslist, lookup by type rather than name, the untyped legacy entry, and the two entries that must not count as inline. A duplicate-performance case is added to the malformed-file tests. Full suite passes (1535), and ruff and mypy are clean.Run against a real Qwen3.6-35B-A3B GB300 submission (9 points), the gate now reports inline accuracy at every point (56.43–56.74 ≥ 55.86) where it previously only warned.
Not in this PR.
swe_bench. The client names that entry after the configured dataset, so a submission that names it differently gets a warning instead of a gate.dataset_type: performanceone. No current single-turn example writes one.🤖 Generated with Claude Code