Skip to content

Latest commit

 

History

History
198 lines (172 loc) · 11.8 KB

File metadata and controls

198 lines (172 loc) · 11.8 KB

Regression test selection (csharp:test-pass)

The csharp:test-pass audit can narrow dotnet test to the tests a change may affect. New selectors first ship advisory/shadow only (SHADOW-BEFORE-ENFORCE): the selector computes the subset it would run, the full suite still runs, and a shadow record captures whether any deselected test failed. Real skipping is gated on accumulated shadow data showing zero unsafe skips over the calibration window (the soundness gate, readyForEnforcement). The project-graph and coverage modes below are the ENFORCING modes, each enabled by operators only after that gate reports ready for the matching selector.

Modes (Audit:TestSelection:Mode)

Mode Behaviour
all (default) Full suite; the emitted command is byte-identical to the legacy path. Instant kill-switch: hot-reloading back to all disables all selection.
coverage-shadow Coverage selector computes the would-be subset; full suite still runs; one structured test-selection shadow log line is emitted per run.
project-graph ENFORCING: the project-graph selector's subset is executed via --filter; any selector error, missing/stale data, global-target touch, or ambiguous result falls back to the full suite (fail-safe). Opt in only after the soundness gate reports zero unsafe skips.
coverage ENFORCING: the coverage selector's subset — nested inside the project-graph superset (coverage can only shrink it, never grow beyond it) — is executed via --filter. Fallback ladder: coverage rung → project-graph rung (the superset) → full suite, on missing/stale coverage data, selector error, or global-target touch. Opt in only after the soundness gate reports zero unsafe skips for the coverage selector.

The value is case-insensitive (coverage_shadow also parses) and hot-reloads via IOptionsMonitor. An unrecognised value fails fast at load.

Selectors

  • Project-graph (ProjectGraphTestSelector): maps each changed file to its owning MSBuild project (via the baseline's file → project map, built by walking the project-reference graph) and selects the baseline's precomputed affected tests, plus tests defined in the changed files. A change owned by an ALWAYS-FULL project forces the full suite (see below).
  • Coverage (CoverageTestSelector): selects tests whose recorded per-test coverage intersects the changed lines, NESTED INSIDE the project-graph superset — every candidate (coverage hit, test defined in a changed file, test with NO coverage record) is kept only when the superset already contains it. Coverage only shrinks that superset, never grows beyond it (defense in depth against a poisoned/stale coverage map). A change no recorded coverage intersects carries no signal, so the ladder descends to the project-graph rung instead of narrowing.

Both fall back to the full suite on ANY uncertainty: no/unknown changeset, no baseline, stale baseline, global targets (Directory.Build.*, Directory.Solution.*, Directory.Packages.props, global.json, NuGet.Config, appsettings*.json, CodeyBox.slnx, .github/workflows/ — exact matches, plus Audit:TestSelection:Coverage knobs), changes owned by an ALWAYS-FULL project (Audit:TestSelection:Coverage:AlwaysFullProjects — defaults: src/CodeyBox.Core/CodeyBox.Core.csproj, tests/CodeyBox.Tests/CodeyBox.Tests.csproj; extend with source-generator projects or other shared roots), whole-file changes, changed test files (which may define unrecorded tests), or files no record references. Running more tests is always safe.

Baseline artifact (per-test coverage map)

Audit:TestSelection:Coverage:BaselineSandboxPath (default /opt/codeybox/test-selection/baseline.json) points at the sandbox-side JSON artifact codeybox-test-selection-baseline/1:

{
  "format": "codeybox-test-selection-baseline/1",
  "commit": "<main HEAD sha>",
  "producedAtUtc": "<timestamp>",
  "fileProject": { "src/Foo/Bar.cs": "src/Foo/Foo.csproj" },
  "projects": { "src/Foo/Foo.csproj": ["Ns.Foo.BarTests"] },
  "tests": {
    "Ns.Foo.BarTests": {
      "file": "tests/Foo.Tests/BarTests.cs",
      "covers": { "src/Foo/Bar.cs": [10, 11, 12] }
    }
  }
}

Producer (operational): tools/CodeyBox.TestSelectionBaseline is the CLI an orchestrator job (and a human) runs against a checkout after every merge to main. It builds the tree, enumerates tests with dotnet test --list-tests, walks the solution (*.slnx parsed directly, *.sln via dotnet sln list) / ProjectReference edges, maps each source file to its project via dotnet msbuild -getItem:Compile (the evaluated truth, so globs and Compile Remove are honoured), resolves each test's defining file from the portable PDB, then collects per-test coverage by running each listed test in isolation through the same coverlet collector the coverage gate uses:

dotnet test <project> --no-build --filter FullyQualifiedName=<test> \
  --collect "XPlat Code Coverage" --results-directory <per-test-dir>

Each Cobertura document is parsed with CoberturaParser (executable-line semantics) and keyed with CoberturaParser.ToRepositoryRelative — the producer does not re-implement those. Only lines with a positive hit count enter that test's covers map; a listed test with no coverage record is still emitted, with an empty map, so the selector must-include it. Size caps (MaxBaselineBytes, MaxBaselineTests, MaxBaselineCoveredLines) are enforced on the fully assembled document; a cap miss fails loudly and never writes a truncated file.

dotnet run --project tools/CodeyBox.TestSelectionBaseline -- produce \
  --repo /path/to/checkout \
  --output /opt/codeybox/test-selection/baseline.json

Mechanism, licences, measured cost, and rejected alternatives: tools/CodeyBox.TestSelectionBaseline/README.md.

Distribution: bake the file into the audit baseline image at the path above, OR fetch the CI artifact to that sandbox path at sandbox setup, OR (let the orchestrator do it) enable automated post-merge production:

  • Per-project opt-in: "TestSelectionBaselineEnabled": true on the project.
  • Global kill-switch and knobs: Audit:TestSelection:BaselineProduction (Enabled, Timeout per job, MaxRetainedPerProject host-side retention, ProducerBinary guest binary — default codeybox-test-selection-baseline, resolved on the job sandbox's PATH, so bake the producer into the audit baseline image; SaturatedPoolRecheckDelay).
  • After each successful merge to the project's base branch the orchestrator schedules one sandboxed production job for that commit (bounded concurrency 1 per project — a second merge before the job starts supersedes it, newest commit wins; the job measures the live base tip at execution so squash-merge divergence converges on the freshest main). The job consumes the global sandbox budget like any other phase and parks while the pipeline is saturated, so work-item phases always win.
  • The artifact is stored host-side keyed by project + commit. At audit sandbox setup the orchestrator stages the newest stored baseline whose commit is an ancestor of the item's base tip at BaselineSandboxPath (read-only); when no stored baseline is fresh within MaxBaselineAge, nothing is staged and the selector runs the full suite.
  • Progress is observable via GET /audit/test-selection/baseline (latest commit, age, size, last production duration or error, pending / running state) and the test-selection.baseline_produced / test-selection.baseline_failed events. A production failure is recorded there but never blocks merges or audits.

Staleness bound: regeneration on every merge to main means the map is at most one merge stale. MaxBaselineAge (default 7 days) is an additional age backstop for a quiet main; a stale/missing/unparseable baseline (or a commit mismatch, when the current commit is known) falls back to the full suite. Size caps (MaxBaselineBytes, MaxBaselineTests, MaxBaselineCoveredLines) bound the untrusted artifact before buffering.

Structural invariants

  • SHADOW-BEFORE-ENFORCE — DotnetTestAuditor executes BuildInvocation(TestSelection.All, …) on every shadow run; the narrowed --filter argv is computed for the shadow record only, never executed. The enforcing project-graph and coverage modes execute the narrowed argv only after the soundness gate reported zero unsafe skips over the calibration window (for the matching selector).
  • FULL-SUITE-ON-MAIN — the merge/release path (IRequiredBuildVerifier / process:required-build) takes no ITestSelector dependency and always runs everything, enforced in code (see TestSelectorTests), not config. The selector is consulted ONLY for per-item audit runs.
  • FAIL-SAFE — selector errors, missing/stale data, global-target touches, and ambiguous results all resolve to the full run; Mode=all (hot-reload) is an instant kill-switch.

Shadow records (the shared validation harness)

Each shadow run emits one TestSelectionShadowRecord through ITestSelectionShadowSink (production: structured log; tests: in-memory) with the verdict safe-for-this-run | unsafe-skips-observed | full-suite | unverifiable, the deselected set, the full run's failures, and their intersection (the unsafe skips). This is the shared harness every selector reports through; the enforcement gate for real skipping consumes the persisted per-run telemetry derived from these records — see "Test-selection soundness" in docs/quality/audit.md (report: GET /audit/test-selection/soundness; CLI: codeybox audit test-selection-soundness). No layer may switch from shadow to enforcing until that gate reports readyForEnforcement: true (zero unsafe skips across the calibration window).

Per-run telemetry (audit report + dashboard)

Every csharp:test-pass invocation — shadow AND full-suite — records a testSelection block on its audit report (audit_reports.test_selection_json, served as testSelection on each auditor in GET /workitems/{id}/audit-reports, rendered on the Audit Reports and Timeline dashboard pages):

Field Meaning
mode Live selection mode (All, CoverageShadow, ProjectGraph, Coverage).
selector Selector that decided (coverage, project-graph, or none when neither shadow nor enforcement was active).
layers Layers consulted, innermost first (["project-graph","coverage"] for the coverage selector, which refines the project-graph superset; ["project-graph"] for enforcing project-graph runs).
selectedCount / totalCount WOULD-BE subset / known universe size. 0/0 means the universe was unknown (no baseline) — the dashboard shows "full suite". For a full-suite fallback with a known universe, both equal the universe size. For enforcing runs, the EXECUTED subset / universe size.
estimatedSavedFraction Proportional estimate: deselected / total in [0,1]. The dashboard multiplies it by the run's durationMs (est. saved 62.5% (~75s)). Zero for full-suite runs. For shadow runs this is an estimate, not a measurement — the full suite always ran.
assessment Shadow verdict (safe-for-this-run | unsafe-skips-observed | full-suite | unverifiable), plus enforced-subset for enforcing runs that executed a narrowed subset (deselected tests were skipped, so no safe/unsafe claim is made; the soundness gate ignores these runs).
fallbacks Which fallback-ladder rungs fired (e.g. no per-test coverage baseline is available, or the coverage → project-graph descent marker project-graph rung:). Empty when the selector narrowed without falling back.
detail Operator-facing selection detail (capped at 4000 chars).

See tests/CodeyBox.Tests/CoverageTestSelectionTests.cs and docs/coverage.md (aggregate gate) / docs/quality/audit.md (DotnetTestAuditor).