Skip to content

Gate agentic inline accuracy on the entry the client writes - #115

Merged
arav-agarwal2 merged 2 commits into
mainfrom
fix/agentic-inline-performance-entry
Oct 9, 2026
Merged

arav-agarwal2 merged 2 commits into
mainfrom
fix/agentic-inline-performance-entry

Conversation

@anandhu-eng

Copy link
Copy Markdown
Collaborator

The agentic inline accuracy gate never runs on a real submission. It looks up an accuracy entry named agentic_combined, which is the performance dataset's name in the Kimi K3 and DSV4 example configs (Qwen's is agentic_coding). The reference client never writes that name: it records the inline score under dataset_name: "performance", tagged dataset_type: "performance", whatever the performance dataset is called (execute.py, accuracy.py). For Kimi K3, Qwen3.6-35B-A3B and DeepSeek-V4.1-Flash the gate therefore only warned "No point reports a agentic_combined accuracy score", and a submission below the inline threshold passed without --strict.

Fix. The inline entry is found by dataset_type: "performance", which is how the client's own average_accuracy identifies it and how its comments ask consumers to filter. An untyped entry named performance, from clients that predate dataset_type, still counts. An entry named performance but typed accuracy does not.

New validation. An accuracy file with more than one dataset_type: "performance" entry is now an accuracy-valid error. The client scores the performance run once and rejects configs with more than one performance dataset, so no real output trips it.

Scope. The lookup is used only by the agentic inline gate. The §15 single-turn gate, SWE-bench mean-of-N and OSL range are unchanged.

Tests. The fixture now writes the client's shape. New cases cover a native accuracy_scores list, lookup by type rather than name, the untyped legacy entry, and the two entries that must not count as inline. A duplicate-performance case is added to the malformed-file tests. Full suite passes (1535), and ruff and mypy are clean.

Run against a real Qwen3.6-35B-A3B GB300 submission (9 points), the gate now reports inline accuracy at every point (56.43–56.74 ≥ 55.86) where it previously only warned.

Not in this PR.

  • SWE-bench is still looked up by the name swe_bench. The client names that entry after the configured dataset, so a submission that names it differently gets a warning instead of a gate.
  • The §15 gate averages every entry and does not skip a dataset_type: performance one. No current single-turn example writes one.

🤖 Generated with Claude Code

The inline gate looked up an accuracy entry named agentic_combined, which is
the performance dataset's name in the Kimi K3 and DSV4 example configs. The
client never writes that name: it records the inline score under
dataset_name "performance", tagged dataset_type "performance", whatever the
performance dataset is called. No real submission ever reached the gate, so
inline accuracy was never checked for Kimi K3, Qwen3.6-35B-A3B or
DeepSeek-V4.1-Flash, and a failing score only raised a warning.

Find the entry by dataset_type, as the client asks of consumers. An untyped
entry named performance, from clients that predate dataset_type, still
counts. An accuracy file with more than one dataset_type performance entry
is now invalid; the client scores the performance run once and refuses more
than one performance dataset. Tests use the client's own output shape.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 8, 2026

Copy link
Copy Markdown

MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅

The README's agentic gate table said inline accuracy is "per point — every
point must clear", which reads as an accuracy run being required at every
Pareto point. The checker requires one per mandatory region
(accuracy-coverage) and gates whatever results are present. Say so in the
README, agentic_targets and the inline gate's docstring.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@arav-agarwal2
arav-agarwal2 merged commit 064aa17 into main Oct 9, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants