feat(adapters): systematic-debugging scenario pack - #254
Conversation
Add a systematic-debugging skill scenario pack to the Superpowers adapters.SuperpowersEvaluator, alongside verification-before-completion. Scenarios judge mechanically-detectable process discipline (all reuse the existing rule-based judge ops; no change to the evidence machinery): - investigate-before-fix: reproduce a failing test before fixing, then re-run and verify (the Iron Law). - failing-test-before-fix: establish a failing signal before the fix, then reach green (Phase 4). - single-fix-not-test-gamed: fix the source so the *unmodified* test passes, rather than gaming the test. Deliberately NOT judged: whether the agent truly understood the root cause — that is beyond a rule judge (the OSS project uses an LLM verifier for skill compliance). Documented as an opt-in real-harness smoke; the change was built /validated offline (16 unit tests) without a live Claude/Codex CLI. Refs microsoft#132.
Per independent review (no P1; P3-nits): - Rename scenario ids for honesty: reproduce-and-verify-before-done and fix-source-not-test-gamed (they check reproduce->fix->verify and fix-source-not-test-game, not semantic root-cause or a strict single-edit). - Keep the declared protected_files_unchanged check so offline unit tests can assert fail-closed on a test-game (the runner also auto-appends it; the duplicate is idempotent/harmless).
|
Thanks for adding the scenario pack. The current evidence does not actually prove reproduce-before-fix ordering: it records aggregate pytest failure/success counts, while |
Make the existing opt-in real-harness caveat explicit and current: the --compare-baseline baseline-versus-skill run and the ordered reproduce-before-fix live evidence were validated with offline fixtures + adversarial-order unit tests only; the real-harness runs require a POSIX host with an authenticated claude CLI and were not executed here.
6c7e135 to
9316a17
Compare
|
Thanks for the re-review. Addressed the ordered-sequence concern (commit 9316a17):
Note: I added an opt-in --compare-baseline real-harness run (score delta of the candidate skill vs the same scenario without it), but I could not execute it here - this PR was developed without an authenticated Claude/CLI on a POSIX host. That's documented in the module docstring; the live harness run remains to be executed on such a host. |
Summary
Adds a
systematic-debuggingscenario pack to the existing Superpowers evaluation adapter (skillopt_sleep/adapters/superpowers.py), alongsideverification-before-completion. Refines issue #132 by extending the adapter to a second checkable skill.What it does
The three scenarios judge mechanically-detectable process discipline, reusing the existing rule-based judge ops (no change to the evidence machinery):
reproduce-and-verify-before-done— observe a failing run, then re-run and verify after editing (guards against fix-without-reproduce / no-verify).failing-test-before-fix— establish a failing signal before the fix, then reach green (Phase 4).fix-source-not-test-gamed— fix the source so the unmodified test passes, rather than gaming the test (fail-closed viaprotected_files_unchanged).Honest boundaries
python -m skillopt_sleep.adapters.superpowers --skill systematic-debuggingA reviewer/CI with a working
claudeCLI can run it to get real empirical evidence.Scope
skillopt_sleep/adapters/superpowers.py+tests/test_systematic_debugging_scenarios.py.verification-before-completionscenarios or the evidence machinery.Refs #132.