The csharp:test-pass audit can narrow dotnet test to the tests a change
may affect. New selectors first ship advisory/shadow only
(SHADOW-BEFORE-ENFORCE): the selector computes the subset it would run, the
full suite still runs, and a shadow record captures whether any deselected
test failed. Real skipping is gated on accumulated shadow data showing zero
unsafe skips over the calibration window (the soundness gate,
readyForEnforcement). The project-graph and coverage modes below are the
ENFORCING modes, each enabled by operators only after that gate reports ready
for the matching selector.
| Mode | Behaviour |
|---|---|
all (default) |
Full suite; the emitted command is byte-identical to the legacy path. Instant kill-switch: hot-reloading back to all disables all selection. |
coverage-shadow |
Coverage selector computes the would-be subset; full suite still runs; one structured test-selection shadow log line is emitted per run. |
project-graph |
ENFORCING: the project-graph selector's subset is executed via --filter; any selector error, missing/stale data, global-target touch, or ambiguous result falls back to the full suite (fail-safe). Opt in only after the soundness gate reports zero unsafe skips. |
coverage |
ENFORCING: the coverage selector's subset — nested inside the project-graph superset (coverage can only shrink it, never grow beyond it) — is executed via --filter. Fallback ladder: coverage rung → project-graph rung (the superset) → full suite, on missing/stale coverage data, selector error, or global-target touch. Opt in only after the soundness gate reports zero unsafe skips for the coverage selector. |
The value is case-insensitive (coverage_shadow also parses) and hot-reloads
via IOptionsMonitor. An unrecognised value fails fast at load.
- Project-graph (
ProjectGraphTestSelector): maps each changed file to its owning MSBuild project (via the baseline'sfile → projectmap, built by walking the project-reference graph) and selects the baseline's precomputed affected tests, plus tests defined in the changed files. A change owned by an ALWAYS-FULL project forces the full suite (see below). - Coverage (
CoverageTestSelector): selects tests whose recorded per-test coverage intersects the changed lines, NESTED INSIDE the project-graph superset — every candidate (coverage hit, test defined in a changed file, test with NO coverage record) is kept only when the superset already contains it. Coverage only shrinks that superset, never grows beyond it (defense in depth against a poisoned/stale coverage map). A change no recorded coverage intersects carries no signal, so the ladder descends to the project-graph rung instead of narrowing.
Both fall back to the full suite on ANY uncertainty: no/unknown changeset, no
baseline, stale baseline, global targets (Directory.Build.*,
Directory.Solution.*, Directory.Packages.props, global.json,
NuGet.Config, appsettings*.json, CodeyBox.slnx,
.github/workflows/ — exact matches, plus Audit:TestSelection:Coverage
knobs), changes owned by an ALWAYS-FULL project
(Audit:TestSelection:Coverage:AlwaysFullProjects — defaults:
src/CodeyBox.Core/CodeyBox.Core.csproj, tests/CodeyBox.Tests/CodeyBox.Tests.csproj;
extend with source-generator projects or other shared roots), whole-file
changes, changed test files (which may define unrecorded
tests), or files no record references. Running more tests is always safe.
Audit:TestSelection:Coverage:BaselineSandboxPath (default
/opt/codeybox/test-selection/baseline.json) points at the sandbox-side JSON
artifact codeybox-test-selection-baseline/1:
{
"format": "codeybox-test-selection-baseline/1",
"commit": "<main HEAD sha>",
"producedAtUtc": "<timestamp>",
"fileProject": { "src/Foo/Bar.cs": "src/Foo/Foo.csproj" },
"projects": { "src/Foo/Foo.csproj": ["Ns.Foo.BarTests"] },
"tests": {
"Ns.Foo.BarTests": {
"file": "tests/Foo.Tests/BarTests.cs",
"covers": { "src/Foo/Bar.cs": [10, 11, 12] }
}
}
}Producer (operational): tools/CodeyBox.TestSelectionBaseline is the
CLI an orchestrator job (and a human) runs against a checkout after every
merge to main. It builds the tree, enumerates tests with dotnet test --list-tests, walks the solution (*.slnx parsed directly, *.sln via
dotnet sln list) / ProjectReference edges, maps each
source file to its project via dotnet msbuild -getItem:Compile (the
evaluated truth, so globs and Compile Remove are honoured), resolves each
test's defining file from the portable PDB, then collects per-test
coverage by running each listed test in isolation through the same
coverlet collector the coverage gate uses:
dotnet test <project> --no-build --filter FullyQualifiedName=<test> \
--collect "XPlat Code Coverage" --results-directory <per-test-dir>
Each Cobertura document is parsed with CoberturaParser (executable-line
semantics) and keyed with CoberturaParser.ToRepositoryRelative — the
producer does not re-implement those. Only lines with a positive hit count
enter that test's covers map; a listed test with no coverage record is
still emitted, with an empty map, so the selector must-include it. Size
caps (MaxBaselineBytes, MaxBaselineTests, MaxBaselineCoveredLines)
are enforced on the fully assembled document; a cap miss fails loudly and
never writes a truncated file.
dotnet run --project tools/CodeyBox.TestSelectionBaseline -- produce \
--repo /path/to/checkout \
--output /opt/codeybox/test-selection/baseline.jsonMechanism, licences, measured cost, and rejected alternatives:
tools/CodeyBox.TestSelectionBaseline/README.md.
Distribution: bake the file into the audit baseline image at the path above, OR fetch the CI artifact to that sandbox path at sandbox setup, OR (let the orchestrator do it) enable automated post-merge production:
- Per-project opt-in:
"TestSelectionBaselineEnabled": trueon the project. - Global kill-switch and knobs:
Audit:TestSelection:BaselineProduction(Enabled,Timeoutper job,MaxRetainedPerProjecthost-side retention,ProducerBinaryguest binary — defaultcodeybox-test-selection-baseline, resolved on the job sandbox's PATH, so bake the producer into the audit baseline image;SaturatedPoolRecheckDelay). - After each successful merge to the project's base branch the orchestrator
schedules one sandboxed production job for that commit (bounded
concurrency 1 per project — a second merge before the job starts
supersedes it, newest commit wins; the job measures the live base tip at
execution so squash-merge divergence converges on the freshest
main). The job consumes the global sandbox budget like any other phase and parks while the pipeline is saturated, so work-item phases always win. - The artifact is stored host-side keyed by project + commit. At audit
sandbox setup the orchestrator stages the newest stored baseline whose
commit is an ancestor of the item's base tip at
BaselineSandboxPath(read-only); when no stored baseline is fresh withinMaxBaselineAge, nothing is staged and the selector runs the full suite. - Progress is observable via
GET /audit/test-selection/baseline(latest commit, age, size, last production duration or error, pending / running state) and thetest-selection.baseline_produced/test-selection.baseline_failedevents. A production failure is recorded there but never blocks merges or audits.
Staleness bound: regeneration on every merge to main means the map is at
most one merge stale. MaxBaselineAge (default 7 days) is an additional
age backstop for a quiet main; a stale/missing/unparseable baseline (or a
commit mismatch, when the current commit is known) falls back to the full
suite. Size caps (MaxBaselineBytes, MaxBaselineTests,
MaxBaselineCoveredLines) bound the untrusted artifact before buffering.
- SHADOW-BEFORE-ENFORCE —
DotnetTestAuditorexecutesBuildInvocation(TestSelection.All, …)on every shadow run; the narrowed--filterargv is computed for the shadow record only, never executed. The enforcingproject-graphandcoveragemodes execute the narrowed argv only after the soundness gate reported zero unsafe skips over the calibration window (for the matching selector). - FULL-SUITE-ON-MAIN — the merge/release path (
IRequiredBuildVerifier/process:required-build) takes noITestSelectordependency and always runs everything, enforced in code (seeTestSelectorTests), not config. The selector is consulted ONLY for per-item audit runs. - FAIL-SAFE — selector errors, missing/stale data, global-target
touches, and ambiguous results all resolve to the full run;
Mode=all(hot-reload) is an instant kill-switch.
Each shadow run emits one TestSelectionShadowRecord through
ITestSelectionShadowSink (production: structured log; tests: in-memory)
with the verdict safe-for-this-run | unsafe-skips-observed |
full-suite | unverifiable, the deselected set, the full run's failures,
and their intersection (the unsafe skips). This is the shared harness every
selector reports through; the enforcement gate for real skipping consumes
the persisted per-run telemetry derived from these records — see
"Test-selection soundness" in docs/quality/audit.md (report: GET /audit/test-selection/soundness; CLI: codeybox audit test-selection-soundness). No layer may switch from shadow to enforcing
until that gate reports readyForEnforcement: true (zero unsafe skips
across the calibration window).
Every csharp:test-pass invocation — shadow AND full-suite — records a
testSelection block on its audit report (audit_reports.test_selection_json,
served as testSelection on each auditor in
GET /workitems/{id}/audit-reports, rendered on the Audit Reports and
Timeline dashboard pages):
| Field | Meaning |
|---|---|
mode |
Live selection mode (All, CoverageShadow, ProjectGraph, Coverage). |
selector |
Selector that decided (coverage, project-graph, or none when neither shadow nor enforcement was active). |
layers |
Layers consulted, innermost first (["project-graph","coverage"] for the coverage selector, which refines the project-graph superset; ["project-graph"] for enforcing project-graph runs). |
selectedCount / totalCount |
WOULD-BE subset / known universe size. 0/0 means the universe was unknown (no baseline) — the dashboard shows "full suite". For a full-suite fallback with a known universe, both equal the universe size. For enforcing runs, the EXECUTED subset / universe size. |
estimatedSavedFraction |
Proportional estimate: deselected / total in [0,1]. The dashboard multiplies it by the run's durationMs (est. saved 62.5% (~75s)). Zero for full-suite runs. For shadow runs this is an estimate, not a measurement — the full suite always ran. |
assessment |
Shadow verdict (safe-for-this-run | unsafe-skips-observed | full-suite | unverifiable), plus enforced-subset for enforcing runs that executed a narrowed subset (deselected tests were skipped, so no safe/unsafe claim is made; the soundness gate ignores these runs). |
fallbacks |
Which fallback-ladder rungs fired (e.g. no per-test coverage baseline is available, or the coverage → project-graph descent marker project-graph rung:). Empty when the selector narrowed without falling back. |
detail |
Operator-facing selection detail (capped at 4000 chars). |
See tests/CodeyBox.Tests/CoverageTestSelectionTests.cs and
docs/coverage.md (aggregate gate) / docs/quality/audit.md
(DotnetTestAuditor).