Repository navigation
L3/L4 populate 0.7% of callables on multi-program repos — dataflow uses one root tsconfig #111
Description
Activity
Attempted fix, and why it does not close this
Branch
fix/issue-111-per-program-dataflow(c17288c) — not merged, do not merge as-is.What it does
Threads
BuiltProgram[]throughstartExtractionso a callable is looked up in the AST index of
the program that owns its file, matching whatbuildSymbolTableand the L2 call graph already
do. Worker tasks are partitioned within a program so every task's files share one tsconfig — as a
side effect this also fixes a latent-j Nvs-j 1divergence on multi-program repos, since all
tasks previously received a singletsConfigFilePath. Adds a coverage line
(dataflow: extracted N of M callables (P%), warning under 50%) so a near-empty L3 can no longer
pass for a successful one.Correct on the fixture:
test/multi-tsconfig.test.tsnow asserts at-a 3that callables owned by
the nestedweb/program get acfg, and it is break-checked — reverting to root-only lookup fails
it. Full suite 227 pass / 6 skip / 0 fail.Why it is incomplete: it OOMs at vscode scale
run peak RSS outcome sequential, all indexes built up front 26.9 GB killed, exit 133 sequential, one index resident at a time 28.6 GB killed, exit 133 -j 4(worker path)29.2 GB killed, exit 133 Exit 133 is a JSC heap OOM abort, and it is silent — no error, no output file. Bounding the index
changed nothing, which is the diagnostic: the memory is in the Projects, not the index.
indexCallableDeclswalks a project, which forces tsc to parse and bind every file in it.
Previously only the root program was ever walked, so the other 91 stayed lazy; now all 92
materialize.Bun
Workers are threads in one process, so-j Ngives each worker its own JSC heap but does not
bound total process memory — which is why the worker path did not rescue it either.What it actually needs
Nothing can be freed today because
src/core.tsstarts extraction (:42) concurrently with the
call-graph solve (:54-68), joining at:97— both hold every program live for the whole run, so
no program is ever done being used. Two candidate designs:- Serialize, then dispose. Run the call graph first, then extraction, releasing each program's
Project once its callables are extracted. Costs the documented concurrency win. - Bounded project pool. Materialize at most N programs at a time, rebuilding on demand.
Either is a design change, not a patch — which is why this issue stays open.
Do not do
Falling back to root-only indexing when the program count is high would restore today's behaviour:
it completes and is quietly wrong. The silence is the actual defect; a crash is at least honest.Status of the shipped release
v1.1.0 still has the original behaviour:
-a 4on vscode completes in 12m22s and populates 1,204 of
174,767 callables (0.7%). L1 and L2 are unaffected, and single-root repositories are unaffected.- Serialize, then dispose. Run the call graph first, then extraction, releasing each program's
Scaling context and the sequenced path for making this affordable: #112. Step 2 there (bounding the resident program set) is what unblocks this issue.
Problem
src/core.tswarns that the L3 dataflow stage uses the root tsconfig for every program, calling it"nested-program files may under-resolve (see #56)". Measured on vscode, the effect is far larger
than under-resolution: levels 3 and 4 populate almost nothing.
-a 4 --graphs cfg,dfg,pdg,sdgon microsoft/vscode (18,391 files, 9,351 modules, 174,767callables), analyzer v1.1.0, exit 0 in 12m22s / 24.7 GB:
cfgcdg/ddg/summarycfgparam_in/param_outstatementbody nodescallnodes)The 44 modules that do get flow are all top-level scripts outside
src/(.github/skills/...,scripts/chat-simulation/...). vscode has no roottsconfig.json— its TypeScript lives undersrc/tsconfig.json— so the single root program the dataflow workers build covers essentially noneof the repository.
The L4 output is only ~5 MB larger than the L2 output (1,030 MB vs 1,026 MB), which is the tell:
the run succeeds and produces nearly no flow.
Scope boundary
The level-2 call graph is already fully per-program and is NOT affected — 1,149,984 edges on the
same run. Levels 1 and 2 are correct. Single-root-tsconfig repositories are unaffected: the same
binary on a repo with a root tsconfig populates
cfgfor a majority of its callables.This is pre-existing and shipped in v1.1.0; it is not a regression from the checker-guard or
allowJsfixes in #103.Cause
src/core.tspassesmat.tsConfigFilePath(one root config) tostartExtraction, andsrc/dataflow/worker.tsbuilds its Project from that single config. EachBuiltProgramalreadycarries its owning
configPath— the file-to-program map exists at the call site — but it is notthreaded into the workers. The code comment at
src/core.tsdocuments this as deferred.Goals
buildSymbolTablealready does for the symbol table and the call graph
collected, so a near-empty L3 cannot pass as success
callables in the nested program
Caveats and known risks
only when
programs.length > 1; a repo with no root tsconfig and one nested program is silent.precisely because it does almost no work.
projects may need pooling rather than one project per worker per program.
Definition of done
-a 4on a multi-program repository populatescfg/cdg/ddgfor the large majority of collectedcallables, verified by re-running the vscode measurement above — not by fixture tests alone.