diff --git a/PROGRESS.md b/PROGRESS.md index d52ca59..b24a168 100644 --- a/PROGRESS.md +++ b/PROGRESS.md @@ -153,3 +153,5 @@ ADR-0019 PARKED, 2026-10-01 — durable meta write via an O_DSYNC fd (2 → 1 fd ADR-0020 SPIKE, 2026-10-01 — 24-byte page-header spike approved for measurement only; the format change is not approved and awaits the spike's numbers. Branch `qdequele/zerodb-header24-spike`. ADR-0021 PRODUCTION PASS, 2026-10-02 — true in-place WRITE_MAP hardened on the spike branch (not merged): B1 unsafe-fn broker with the sole sanctioned call in zerodb-core::dirty; B2 superseded by a stronger M1 miri finding (under Stacked Borrows a cached whole-map reference is UB after any in-place write — the in-map writer now borrows its map view lazily per access, like readers; spill-time re-derivation kept for the heap-staged paths); B3 map-aware fault backend (brokered regions journaled, sealed per sync, image cuts open the real writable map; vacuousness tripwire + smoke pin); B4 loom justified as no-new-model (ADR-0021 §hardening), 180 s stress writemap-in-place variant, M1.13 erased-cursor audit + miri battery in heed-zerodb; M1 zerodb_io::testmap::TestWriteMap (test-backing feature) + writemap_in_place_miri battery; M2 release-mode TXN-62 typed guard in allocate; M3 abort-after-spill / nested-fanout / put_reserved batteries parameterized over WRITE_MAP. SPEC 04 TXN-71/C5a/TXN-45b and SPEC 06 REC-20 amended in the same change. + +ADR-0022 SPIKE, 2026-10-03 — meta free-list annex (format v2): the per-commit freed PIL rides inside the CRC-covered meta page; the GC tree is the spill/cold path (over-cap sets, reader-gated carries). Steady-state freelist_save drops from ~2 GC-tree puts + 1 delete + 1 leaf COW per commit to zero tree ops (~1.2–1.5 µs → ~0.1 µs local 4 KiB census; one fewer dirty page per commit), and churn steady-state files shrink ~12% vs v1. First-cut lesson: folding the carried annex into `freed` up front starved the save's own allocations into per-commit extends (unbounded ratchet, caught by churn_parity_general); the carried remainder must stay a gated in-save pool (GC-12 source 4) merged at placement. Pinned crash seed re-pinned (198) per its documented protocol. Implemented on coordinator-relayed approval; ratified directly 2026-10-05 (format v2 + seed 198), re-gated on main after ADR-0021 merged (test 653/0, miri 206/0, loom 9/0, stress 3/0, crash-test-quick clean, fuzz-quick clean). diff --git a/README.md b/README.md index 73cbdb7..18d92be 100644 --- a/README.md +++ b/README.md @@ -60,11 +60,33 @@ bytes. Full list: [`docs/COMPATIBILITY.md`](docs/COMPATIBILITY.md). ## Performance -On its real consumers, ZeroDB runs at **LMDB-level performance** (Meilisearch -indexing 1.00×, search 1.03×; hannoy search 0.95×). Against the field — the +On its real consumers, ZeroDB runs at **LMDB-level performance**: Meilisearch +indexing 1.01× (movies) and 0.99× (incremental hackernews additions), search +0.94× — time relative to LMDB, lower is better; hannoy search 0.95×. + +**ZeroDB vs LMDB, current tree** — YCSB on a Graviton4 with local NVMe, +10 M × 128 B, kops/s (higher is better). `WRITE_MAP` is opt-in on both engines +(LMDB's `MDB_WRITEMAP`, ZeroDB's in-place `WRITE_MAP`): + +| Workload | LMDB | zerodb | LMDB `WRITE_MAP` | zerodb `WRITE_MAP` | +|---|---:|---:|---:|---:| +| YCSB A (no-sync, 2 GiB cap) | 256 | 211 | 545 | 408 | +| YCSB B (no-sync, 2 GiB cap) | 490 | 355 | 1,199 | 847 | +| YCSB B (fsync) | 238 | **258** | — | — | + +Durable writes are ahead of LMDB (1.08×); no-sync writes trail it by 18–27%, and +`WRITE_MAP` makes ZeroDB 1.6–1.7× faster than default LMDB, still behind LMDB's +own `WRITE_MAP`. Details: +[`benches/results/2026-10-05-meta-annex-real-case.md`](benches/results/2026-10-05-meta-annex-real-case.md). +_(2026-10-05, 2 reps.)_ + +**Against the field** — the [rust-storage-bench](https://github.com/marvin-j97/rust-storage-bench) suite behind the *fjall 3* article, on a Graviton4 with local NVMe — throughput -in kops/s (higher is better; per-row winner in **bold**): +in kops/s (higher is better; per-row winner in **bold**). This run predates +in-place `WRITE_MAP` and the meta free-list annex, and uses a different setup +(8 GiB cap, 2 GiB cache, 100 B values), so its numbers don't compare with the +table above: | Workload | LMDB | zerodb | fjall 3 | rocksdb | redb | sqlite | |---|---:|---:|---:|---:|---:|---:| @@ -74,7 +96,7 @@ in kops/s (higher is better; per-row winner in **bold**): | feed | **55** | 42 | 37 | 31 | 16 | 38 | | 100 M keys | **93** | 79 | 76 | 65 | 15 | 36 | -ZeroDB tracks LMDB closely — at parity on durable writes (YCSB B), a little ahead +In that run ZeroDB tracked LMDB closely — at parity on durable writes (YCSB B), a little ahead on 4 KB values, behind on the write-heavy no-sync mix (its weakest path). The LSM engines (fjall, rocksdb) take the raw write-throughput rows; the B-trees (LMDB and ZeroDB) keep read p99 in microseconds where the LSMs run to hundreds. Per-engine diff --git a/benches/results/2026-10-05-meta-annex-real-case.md b/benches/results/2026-10-05-meta-annex-real-case.md new file mode 100644 index 0000000..a65fcf9 --- /dev/null +++ b/benches/results/2026-10-05-meta-annex-real-case.md @@ -0,0 +1,109 @@ +# Meta free-list annex (ADR-0022) — real-case three-column check — 2026-10-05 + +The commit-ladder A/B for ADR-0022 only showed a win on tiny no-sync commits +(`commit/batch/n1` −22%; `n10k` and the fsync rungs flat). This run checks it on +real workloads: YCSB through rust-storage-bench, and Meilisearch's own +`cargo xtask bench` workloads. + +Three columns throughout: + +| column | tree | +|---|---| +| **LMDB** | heed 0.22.1 (the Meilisearch LMDB fork) — drift control | +| **ZeroDB before** | `8b066a6` — main with ADR-0021 (in-place `WRITE_MAP`), no annex | +| **ZeroDB after** | `b01b75b` — `8b066a6` + ADR-0022 (the annex), nothing else | + +`*-wm` rows open the environment with `WRITE_MAP` (LMDB `MDB_WRITEMAP`, ZeroDB +in-place `WRITE_MAP`). + +## YCSB — Graviton4 NVMe (the target hardware) + +AWS `m8gd.xlarge` (Graviton4, 4 vCPU, 16 GiB), Ubuntu 24.04 aarch64, local NVMe +instance store, ext4 `noatime`. rust-storage-bench (with the `zerodb` backend), +10 M items × 128 B, json corpus, 8 MiB engine cache, 60 s per run; no-sync runs in +a `MemoryMax=2G` scope, durable runs uncapped. Fresh data dir and dropped page +cache before every run. 2 reps, the second in reverse order. + +### YCSB A (50/50 read/update), no-sync + +| config | ops/s (reps) | median | ÷ LMDB | write p50 | write p99 | +|---|---|---:|---:|---:|---:| +| LMDB | 255k / 257k | 256k | 1.00 | 6.4 µs | 13.9 µs | +| ZeroDB before | 207k / 207k | 207k | 0.81 | 14.0 µs | 27.7 µs | +| **ZeroDB after** | 213k / 209k | 211k | 0.82 | **12.0 µs** | **24.8 µs** | +| LMDB-wm | 545k / 545k | 545k | 2.13 | 1.3 µs | 4.1 µs | +| ZeroDB-wm before | 421k / 404k | 412k | 1.61 | 4.4 µs | 9.3 µs | +| **ZeroDB-wm after** | 412k / 405k | 408k | 1.60 | **3.3 µs** | **7.8 µs** | + +after ÷ before: default **1.018** (rep range 1.008–1.028); wm 0.991 (0.963–1.020). + +### YCSB B (95/5 read/update), no-sync + +| config | ops/s (reps) | median | ÷ LMDB | write p50 | write p99 | +|---|---|---:|---:|---:|---:| +| LMDB | 492k / 487k | 490k | 1.00 | 7.1 µs | 15.4 µs | +| ZeroDB before | 368k / 358k | 363k | 0.74 | 15.1 µs | 31.9 µs | +| **ZeroDB after** | 372k / 339k | 355k | 0.73 | **13.1 µs** | 54.0 µs ¹ | +| LMDB-wm | 1209k / 1190k | 1199k | 2.45 | 1.4 µs | 6.4 µs | +| ZeroDB-wm before | 819k / 810k | 814k | 1.66 | 4.6 µs | 11.6 µs | +| **ZeroDB-wm after** | 845k / 848k | 847k | 1.73 | **3.5 µs** | **10.3 µs** | + +after ÷ before: default 0.980 (0.922–1.040, noisy); wm **1.040** (1.033–1.047). + +¹ One rep only: rep 1 was 29.4 µs (better than before's 31.9 µs); rep 2 was +78.5 µs during a run that read 382 MiB from disk vs ~210 MiB for the others — a +page-cache eviction episode under the 2 GiB cap, not the annex. + +### YCSB B, durable (every commit fsynced) + +| config | ops/s (reps) | median | ÷ LMDB | write p50 | write p99 | +|---|---|---:|---:|---:|---:| +| LMDB | 238k / 239k | 238k | 1.00 | 167.8 µs | 178.1 µs | +| ZeroDB before | 238k / 241k | 239k | 1.00 | 161.2 µs | 178.1 µs | +| **ZeroDB after** | 258k / 258k | **258k** | **1.08** | **129.4 µs** | **148.8 µs** | + +after ÷ before: **1.078** (rep range 1.070–1.085). One page fewer written and +flushed per commit shows up directly when the flush is cheap (NVMe). + +## YCSB — x86-64 bench server (confirmation) + +Xeon E3-1230 v2 (4C/8T, turbo off, performance governor), 2× SATA SSD in +mdraid RAID 1, Debian 12. Same method, 45 s per run, 2 reps (durable: 1 rep). + +| workload | before → after ops/s | after ÷ before | write p50 before → after | +|---|---|---:|---| +| A no-sync | 254k → 261k | **1.027** (1.019–1.034) | 17.2 → 13.8 µs | +| A no-sync, wm | 336k → 357k | **1.065** (1.022–1.110) | 8.7 → 8.0 µs | +| B no-sync | 421k → 448k | **1.064** (1.046–1.083) | 18.2 → 13.9 µs | +| B no-sync, wm | 661k → 696k | **1.053** (1.052–1.054) | 8.9 → 6.9 µs | +| B durable | 188k → 189k | 1.006 | 1398 → 1316 µs | + +The durable row is flat here: this box's fsync is a ~1.4 ms SATA RAID flush, +which drowns per-commit CPU (the same reason the ladder's `commit/sync/*` rungs +measured flat on it). + +## Meilisearch — x86-64 bench server + +Meilisearch v1.53.1 built three times (stock LMDB, ZeroDB before, ZeroDB after), +`cargo xtask bench` on `movies.json`, `hackernews-add-new-documents.json` and +`search/movies.json`, 3 rounds with the starting engine rotated each round. +Median server-side time (`::meta::total` span) over all runs: + +| workload | LMDB | ZeroDB before | ZeroDB after | after ÷ before | after ÷ LMDB | +|---|---:|---:|---:|---:|---:| +| movies indexing (30 runs) | 4.850 s | 4.865 s | 4.879 s | 1.00 | 1.01 | +| hackernews incremental additions (9 runs) | 40.25 s | 40.74 s | 40.01 s | 0.98 | 0.99 | +| — of which `indexing::scheduler::commit` | 11.03 s | 10.86 s | 10.21 s | **0.94** | 0.93 | +| movies search (30 runs) | 15.4 ms | 14.8 ms | 14.4 ms | 0.97 | 0.94 | + +No regression. Bulk indexing is flat as expected (few, large commits); the +incremental-additions workload, which commits more often, gains 6% on its commit +span. + +## Verdict + +ADR-0022 is a real-case win where commits are frequent or durable on fast +storage, and neutral for Meilisearch bulk indexing: YCSB +2–8% throughput, write +p50 −14% to −25% in every configuration on both machines, no regression in any +Meilisearch workload. It is not a Meilisearch indexing speed-up and should not be +described as one. diff --git a/crates/zerodb-core/src/builder.rs b/crates/zerodb-core/src/builder.rs index 94222d1..fb5d363 100644 --- a/crates/zerodb-core/src/builder.rs +++ b/crates/zerodb-core/src/builder.rs @@ -355,6 +355,7 @@ fn finalize_image( last_pg, free_db: DBRecord::empty(), main_db, + fl_count: 0, }; m0.encode(&mut buf[0..psize as usize])?; let mut m1 = m0; @@ -894,6 +895,7 @@ impl EnvStream { last_pg, free_db: DBRecord::empty(), main_db, + fl_count: 0, }; meta.encode(&mut frame)?; sink.emit(0, &frame)?; diff --git a/crates/zerodb-core/src/check.rs b/crates/zerodb-core/src/check.rs index f8f8393..25bd510 100644 --- a/crates/zerodb-core/src/check.rs +++ b/crates/zerodb-core/src/check.rs @@ -39,9 +39,13 @@ struct Checker<'a> { /// otherwise choose ids that all collide (HashDoS). The engine's /// dirty-store hasher is unaffected — its keys are engine-authored. visited: HashSet, - /// Free page ids collected from every GC PIL → occurrence count - /// (INV-22/INV-24; SPEC 05 §9). Randomly seeded, see `visited`. + /// Free page ids collected from every GC PIL **and the selected meta's + /// free-list annex** → occurrence count (INV-22/INV-24/INV-28; SPEC 05 + /// §9/§2a). Randomly seeded, see `visited`. free: HashMap, + /// The selected meta's annex id count (INV-28/GC-29: `> 0` forbids a GC + /// tree entry keyed `BE(meta_txnid)`). + annex_count: usize, violations: Vec, } @@ -353,6 +357,15 @@ impl<'a> Checker<'a> { if txnid > self.meta_txnid { self.fail("INV-23", format!("GC entry keyed by future txn {txnid}")); } + // INV-28/GC-29 (format v2): a meta with a non-empty annex holds the + // newest freeing-txn's PIL itself — a tree entry under the same + // txnid would be the forbidden split placement (GC-32). + if txnid == self.meta_txnid && self.annex_count > 0 { + self.fail( + "INV-28", + format!("GC entry {txnid} coexists with a non-empty meta annex (GC-29/GC-32)"), + ); + } let pil: Vec = match val { LeafValue::Inline(v) => v.to_vec(), LeafValue::Overflow { head_pgno, dsize } => { @@ -522,9 +535,9 @@ pub fn check_image(bytes: &[u8], psize: u32) -> Vec { (Ok(a), Ok(b)) => (a, b), _ => return vec!["INV-1: meta slots undecodable".into()], }; - let meta = match select_meta(&v0, &v1, false) { - crate::page::MetaChoice::Both { meta, .. } - | crate::page::MetaChoice::OnlyOne { meta, .. } => meta, + let (meta, chosen) = match select_meta(&v0, &v1, false) { + crate::page::MetaChoice::Both { meta, chosen } + | crate::page::MetaChoice::OnlyOne { meta, chosen } => (meta, chosen), crate::page::MetaChoice::None => { return vec!["INV-2: no valid meta slot (both torn/foreign)".into()] } @@ -557,13 +570,34 @@ pub fn check_image(bytes: &[u8], psize: u32) -> Vec { return violations; } } + // Format v2 (ADR-0022): the selected meta's free-list annex ids are free + // pages (SPEC 05 §2a). Validate their GC-29 shape (the INV-25/26 + // analogues, labelled INV-28) and count them into the free set for the + // INV-22/24 partition. + let annex = MetaPage::read_annex(&bytes[chosen * ps..(chosen + 1) * ps], psize) + .expect("annex count was bounds-checked by MetaPage::validate"); + let mut free: HashMap = HashMap::new(); + let mut prev: Option = None; + for &id in &annex { + if id < FIRST_DATA_PGNO || id > meta.last_pg { + violations.push(format!("INV-28: annex free id {id} out of range")); + } + if let Some(p) = prev { + if p >= id { + violations.push("INV-28: annex ids not strictly ascending".into()); + } + } + prev = Some(id); + *free.entry(id).or_insert(0) += 1; + } let mut checker = Checker { bytes, psize, meta_txnid: meta.txnid, last_pg: meta.last_pg, visited: HashSet::new(), - free: HashMap::new(), + free, + annex_count: annex.len(), violations, }; checker.check_catalog_record("main_db", &meta.main_db); @@ -607,8 +641,10 @@ mod tests { let base = slot * PS as usize; let e_off = base + 120 + 32; // main_db record + entries offset img[e_off] = 99; - let crc = crate::page::crc32c(&img[base..base + 168]); - img[base + 168..base + 172].copy_from_slice(&crc.to_le_bytes()); + // Format v2 CRC: fl_count (offset 168) is 0 in builder images, + // so the coverage is exactly [0, 172); the CRC field sits at 172. + let crc = crate::page::crc32c(&img[base..base + 172]); + img[base + 172..base + 176].copy_from_slice(&crc.to_le_bytes()); } let v = check_image(&img, PS); assert!( diff --git a/crates/zerodb-core/src/env.rs b/crates/zerodb-core/src/env.rs index 82afe32..d4e4216 100644 --- a/crates/zerodb-core/src/env.rs +++ b/crates/zerodb-core/src/env.rs @@ -189,7 +189,7 @@ pub trait Backing: Send + Sync { /// state (SPEC 04 TXN-18). Readers `Arc`-clone the env's published snapshot at /// begin and never re-read a durable meta page (the slot a pinned txnid lived /// in is overwritten two commits later, TXN-63). -#[derive(Debug, Clone, Copy, PartialEq, Eq)] +#[derive(Debug, Clone, PartialEq, Eq)] pub struct Snapshot { /// The commit point of this snapshot. pub txnid: u64, @@ -199,17 +199,28 @@ pub struct Snapshot { pub main_db: DBRecord, /// Root/stats of the free (GC) DB. pub free_db: DBRecord, + /// This snapshot's meta free-list annex (SPEC 02 §3 format v2, ADR-0022; + /// SPEC 05 §2a): the pages freed by txn `txnid`, exactly the `fl_ids` of + /// its meta. Pinned **in the snapshot** — the meta slot `txnid & 1` is + /// overwritten by txn `txnid + 2` even while this snapshot stays pinned + /// (TXN-63), so holders (`copy`, the next writer) must never re-read the + /// slot. `Arc<[u64]>` keeps `Snapshot` cheap to clone. + pub free_annex: std::sync::Arc<[u64]>, } impl Snapshot { - /// The snapshot a validated meta page describes. + /// The snapshot a validated meta page describes, with `annex` as the + /// meta's free-list annex ids (read via [`MetaPage::read_annex`] from the + /// same validated slot buffer; `annex.len()` must equal `meta.fl_count`). #[must_use] - pub fn from_meta(meta: &MetaPage) -> Snapshot { + pub fn from_meta(meta: &MetaPage, annex: Vec) -> Snapshot { + debug_assert_eq!(annex.len(), meta.fl_count as usize); Snapshot { txnid: meta.txnid, last_pg: meta.last_pg, main_db: meta.main_db, free_db: meta.free_db, + free_annex: annex.into(), } } } @@ -1635,20 +1646,27 @@ pub fn open_with_backing_policy( let slot1 = read_slot(bytes, META_B_PGNO, ps, page_size)?; // Select the live snapshot (SPEC 02 §3.2 / SPEC 06 REC-2..5). - let meta = match select_meta(&slot0, &slot1, prev_snapshot) { - MetaChoice::Both { meta, .. } => meta, - MetaChoice::OnlyOne { meta, .. } => { + let (meta, chosen_slot) = match select_meta(&slot0, &slot1, prev_snapshot) { + MetaChoice::Both { meta, chosen } => (meta, chosen), + MetaChoice::OnlyOne { meta, chosen } => { if prev_snapshot { // REC-2† (ratified 2026-07-16): one valid slot + PREV_SNAPSHOT is // a hard error — there are not two committed snapshots to pick an // older from. return Err(Error::Mdb(MdbError::Invalid)); } - meta + (meta, chosen) } // REC-3: both invalid → unrecoverable. MetaChoice::None => return Err(Error::Mdb(MdbError::Invalid)), }; + // Format v2 (ADR-0022): read the selected slot's free-list annex ids — + // the one and only meta read, alongside the roots (TXN-18); the slot is + // overwritten two commits later, so the ids are pinned in the Snapshot. + // `fl_count` passed the §3.2 rule-5 bound + CRC; the ids' GC-29 shape is + // validated by the consumer before any id is handed out (GC-33). + let annex = MetaPage::read_annex(&bytes[chosen_slot * ps..(chosen_slot + 1) * ps], page_size) + .ok_or(Error::Mdb(MdbError::Invalid))?; // SPEC 06 REC-1a / SPEC 02 §3.2 step 6 (geometry validation): a slot can // carry a valid CRC and still name geometry the real file cannot back — a @@ -1699,7 +1717,7 @@ pub fn open_with_backing_policy( map_size, // Seed the published-snapshot cell from the durable meta — the one // and only time a meta *page* is read for roots (SPEC 04 TXN-18). - snap_cell: SnapshotCell::new(Arc::new(Snapshot::from_meta(&meta))), + snap_cell: SnapshotCell::new(Arc::new(Snapshot::from_meta(&meta, annex))), write_mutex: WriterLock::new(), commit_hook: Mutex::new(None), poisoned: AtomicBool::new(false), diff --git a/crates/zerodb-core/src/nested.rs b/crates/zerodb-core/src/nested.rs index 35eda73..5c1be14 100644 --- a/crates/zerodb-core/src/nested.rs +++ b/crates/zerodb-core/src/nested.rs @@ -171,6 +171,9 @@ impl TxnRead for NestedRoTxn<'_> { fn free_record(&self) -> &DBRecord { self.parent.free_record() } + fn free_annex_count(&self) -> u64 { + self.parent.free_annex_count() + } fn page_size(&self) -> u32 { self.parent.page_size() } diff --git a/crates/zerodb-core/src/page/crc32c.rs b/crates/zerodb-core/src/page/crc32c.rs index ba9f9b4..ab8471f 100644 --- a/crates/zerodb-core/src/page/crc32c.rs +++ b/crates/zerodb-core/src/page/crc32c.rs @@ -41,16 +41,36 @@ const fn build_table() -> [u32; 256] { table } +/// Fold `bytes` into a running (non-finalized) CRC32C state. The single source +/// of the byte-wise table step, shared by [`crc32c`] and [`crc32c_concat`] so +/// the two cannot drift if the table or algorithm ever changes. +#[inline] +fn crc_update(mut crc: u32, bytes: &[u8]) -> u32 { + for &byte in bytes { + let idx = ((crc ^ byte as u32) & 0xFF) as usize; + crc = (crc >> 8) ^ TABLE[idx]; + } + crc +} + /// Compute the CRC32C (Castagnoli) checksum of `data`. /// /// Returns the finalized checksum (post final-XOR), i.e. the value written to a /// meta page's `meta_crc` field. #[must_use] pub fn crc32c(data: &[u8]) -> u32 { + crc_update(0xFFFF_FFFF, data) ^ 0xFFFF_FFFF +} + +/// CRC32C over the concatenation of `parts`, as one stream — equal to +/// `crc32c` of the parts joined into a single buffer, without the join. +/// Used for the meta CRC's split coverage (SPEC 02 §3.3 as amended by +/// ADR-0022: `[0, 172)` then the annex ids, skipping the CRC field itself). +#[must_use] +pub fn crc32c_concat(parts: &[&[u8]]) -> u32 { let mut crc = 0xFFFF_FFFFu32; - for &byte in data { - let idx = ((crc ^ byte as u32) & 0xFF) as usize; - crc = (crc >> 8) ^ TABLE[idx]; + for part in parts { + crc = crc_update(crc, part); } crc ^ 0xFFFF_FFFF } diff --git a/crates/zerodb-core/src/page/meta.rs b/crates/zerodb-core/src/page/meta.rs index c4719a9..36dad22 100644 --- a/crates/zerodb-core/src/page/meta.rs +++ b/crates/zerodb-core/src/page/meta.rs @@ -8,11 +8,13 @@ //! §3.2's numbered list), and the double-buffer selection formula //! ([`select`]). -use super::crc32c::crc32c; +use super::crc32c::crc32c_concat; use super::geometry::validate_page_size; use super::header::CommonHeader; use super::raw::{read_u16, read_u32, read_u64, write_u16, write_u32, write_u64}; -use super::{PageError, FORMAT_VERSION, MAGIC, META_CONTENT_LEN, PGNO_INVALID, P_META}; +use super::{ + PageError, FORMAT_VERSION, MAGIC, META_ANNEX_OFF, META_CONTENT_LEN, PGNO_INVALID, P_META, +}; // Meta body field offsets (absolute within the page). const OFF_MAGIC: usize = 32; @@ -24,7 +26,27 @@ const OFF_LAST_PG: usize = 56; const OFF_BODY_TXNID: usize = 64; const OFF_FREE_DB: usize = 72; const OFF_MAIN_DB: usize = 120; -const OFF_META_CRC: usize = 168; +const OFF_FL_COUNT: usize = 168; +const OFF_META_CRC: usize = META_CONTENT_LEN; // 172 (SPEC 02 §3, format v2) + +/// Maximum number of free-list annex ids a meta page of `psize` bytes can +/// carry (SPEC 02 §3, ADR-0022): the ids start at [`META_ANNEX_OFF`] and run +/// to the end of the page. +#[must_use] +pub fn meta_annex_cap(psize: u32) -> usize { + (psize as usize).saturating_sub(META_ANNEX_OFF) / 8 +} + +/// The meta CRC32C with the ADR-0022 split coverage: `[0, 172)` followed by +/// the `8·fl_count` annex-id bytes at [`META_ANNEX_OFF`], skipping the CRC +/// field itself. `buf` must hold the whole page; `fl_count` must already be +/// bounds-checked against [`meta_annex_cap`]. +fn meta_crc_of(buf: &[u8], fl_count: usize) -> u32 { + crc32c_concat(&[ + &buf[..META_CONTENT_LEN], + &buf[META_ANNEX_OFF..META_ANNEX_OFF + 8 * fl_count], + ]) +} /// Size of a [`DBRecord`], in bytes. pub const DBRECORD_LEN: usize = 48; @@ -147,6 +169,10 @@ pub struct MetaPage { pub free_db: DBRecord, /// Root/stats of the main/catalog DB. pub main_db: DBRecord, + /// Free-list annex id count (SPEC 02 §3 format v2, ADR-0022). The ids + /// themselves stay in the page buffer (offset [`META_ANNEX_OFF`]) and are + /// read with [`MetaPage::read_annex`]; this decoded struct stays `Copy`. + pub fl_count: u32, } impl MetaPage { @@ -165,16 +191,42 @@ impl MetaPage { last_pg: 1, free_db: DBRecord::empty(), main_db: DBRecord::empty(), + fl_count: 0, } } /// Encode this meta into `buf` (which must be at least `page_size` bytes), - /// computing and writing the CRC and zeroing the reserved tail. + /// with an **empty** free-list annex, computing and writing the CRC and + /// zeroing the reserved tail. See [`MetaPage::encode_with_annex`]. /// /// # Errors /// /// [`PageError::InvalidPageSize`] or [`PageError::BufferTooSmall`]. pub fn encode(&self, buf: &mut [u8]) -> Result<(), PageError> { + self.encode_with_annex(buf, &[]) + } + + /// Encode this meta into `buf` with `annex` as the free-list annex ids + /// (SPEC 02 §3 format v2, ADR-0022; SPEC 05 §2a). This encoder enforces only + /// the capacity bound; it does **not** check the GC-29 shape (strictly + /// ascending, unique, in range), and it is sound for it not to, because the + /// shape is never trusted on the way *out* of a meta — it is re-checked by + /// `validate_pil_ids` before any id is handed out (SPEC 05 GC-33). The two + /// callers and their inputs: + /// - `RwTxn::freelist_save` → the commit path: produces a GC-29-shaped list + /// by construction, so what it encodes is well-formed. + /// - `copy::write_raw` → raw env copy: forwards the pinned snapshot's annex + /// verbatim. That annex passed only the open-time count+CRC bound, **not** + /// `validate_pil_ids`, so a hostile-but-CRC-valid source file is copied + /// faithfully (ids unchanged) — the copy is as trustworthy as its source + /// and no more, and the copy's own reader re-validates before any draw. + /// + /// # Errors + /// + /// [`PageError::InvalidPageSize`], [`PageError::BufferTooSmall`], or + /// [`PageError::BadValueSize`] if `annex` exceeds [`meta_annex_cap`] + /// (engine bug: the save's fit check owns that bound). + pub fn encode_with_annex(&self, buf: &mut [u8], annex: &[u64]) -> Result<(), PageError> { validate_page_size(self.page_size)?; let psize = self.page_size as usize; if buf.len() < psize { @@ -183,6 +235,14 @@ impl MetaPage { psize, }); } + if annex.len() > meta_annex_cap(self.page_size) { + return Err(PageError::BadValueSize(annex.len() as u64)); + } + debug_assert_eq!( + self.fl_count as usize, + annex.len(), + "MetaPage.fl_count must match the annex slice (single source: the ids)" + ); // Zero the whole page first so every reserved byte is 0. buf[..psize].fill(0); // Common header (zeros reserved0 + checksum; variant tail already 0). @@ -202,12 +262,40 @@ impl MetaPage { write_u64(buf, OFF_BODY_TXNID, self.txnid); self.free_db.write(buf, OFF_FREE_DB); self.main_db.write(buf, OFF_MAIN_DB); - // CRC over [0, 168). - let crc = crc32c(&buf[..META_CONTENT_LEN]); + // Free-list annex (format v2): count at 168, ids from 176. + write_u32(buf, OFF_FL_COUNT, annex.len() as u32); + for (i, &id) in annex.iter().enumerate() { + write_u64(buf, META_ANNEX_OFF + 8 * i, id); + } + // CRC over [0, 172) ∪ the annex ids (SPEC 02 §3.3). + let crc = meta_crc_of(buf, annex.len()); write_u32(buf, OFF_META_CRC, crc); Ok(()) } + /// Read the free-list annex ids out of a meta page buffer (SPEC 02 §3, + /// format v2). Returns `None` if `fl_count` exceeds the page's capacity — + /// callers treat that as a corrupt freelist (`MdbError::Invalid`), though + /// for a slot that passed [`MetaPage::validate`] the bound already held. + /// The ids' GC-29 shape (ascending, in range) is **not** checked here: + /// the consumer runs `validate_pil_ids` before any id is handed out + /// (SPEC 05 GC-33), exactly as for a tree PIL. + #[must_use] + pub fn read_annex(buf: &[u8], psize: u32) -> Option> { + if buf.len() < psize as usize { + return None; + } + let count = read_u32(buf, OFF_FL_COUNT) as usize; + if count > meta_annex_cap(psize) { + return None; + } + let mut ids = Vec::with_capacity(count); + for i in 0..count { + ids.push(read_u64(buf, META_ANNEX_OFF + 8 * i)); + } + Some(ids) + } + /// Validate `buf` as a meta slot, following SPEC 02 §3.2's numbered list /// (the single owner of the meta-slot validation predicate). Returns a rich /// [`MetaValidity`] verdict; validity failures are *data*, not errors. @@ -251,9 +339,15 @@ impl MetaPage { body: body_txnid, }); } - // Rule 5: CRC over [0, 168). + // Rule 5 (format v2, ADR-0022): the annex count is bounds-checked + // BEFORE the CRC — a hostile count must not drive the CRC read out + // of the page — then the CRC covers [0, 172) ∪ the annex ids. + let fl_count = read_u32(buf, OFF_FL_COUNT) as usize; + if fl_count > meta_annex_cap(psize) { + return Ok(MetaValidity::BadAnnexCount(fl_count as u32)); + } let stored = read_u32(buf, OFF_META_CRC); - let computed = crc32c(&buf[..META_CONTENT_LEN]); + let computed = meta_crc_of(buf, fl_count); if stored != computed { return Ok(MetaValidity::BadCrc { stored, computed }); } @@ -269,6 +363,7 @@ impl MetaPage { last_pg: read_u64(buf, OFF_LAST_PG), free_db: DBRecord::read(buf, OFF_FREE_DB), main_db: DBRecord::read(buf, OFF_MAIN_DB), + fl_count: fl_count as u32, })) } } @@ -285,6 +380,11 @@ pub enum MetaValidity { BadVersion(u32), /// `page_size` is not a power of two in range (the observed value). BadPageSize(u32), + /// `fl_count` exceeds the page's annex capacity (SPEC 02 §3.2 rule 5, + /// format v2) — the slot is **invalid and discarded** (as §3.3 / REC-8 + /// word it). The cause is either a torn write or a hostile count; both are + /// rejected identically, before the CRC read, so the name covers both. + BadAnnexCount(u32), /// Header txnid and body txnid disagree — a torn write (INV-2). TxnidMismatch { /// Header stamp (offset 8). diff --git a/crates/zerodb-core/src/page/mod.rs b/crates/zerodb-core/src/page/mod.rs index 2bbb2fb..bc56741 100644 --- a/crates/zerodb-core/src/page/mod.rs +++ b/crates/zerodb-core/src/page/mod.rs @@ -38,11 +38,14 @@ mod tree; // validation; it contains no unsafe operation. mod trust; -pub use crc32c::crc32c; +pub use crc32c::{crc32c, crc32c_concat}; pub(crate) use header::read_flags; pub(crate) use header::read_page_txnid; pub use header::{CommonHeader, PageRef}; -pub use meta::{select as select_meta, DBRecord, MetaChoice, MetaPage, MetaValidity, DBRECORD_LEN}; +pub use meta::{ + meta_annex_cap, select as select_meta, DBRecord, MetaChoice, MetaPage, MetaValidity, + DBRECORD_LEN, +}; pub use overflow::{write_overflow_head, OverflowRef}; pub use tree::{BranchMut, BranchRef, LeafMut, LeafRef, LeafValue}; pub use trust::FileTrust; @@ -54,8 +57,10 @@ pub use trust::FileTrust; /// File identifier, ASCII `"ZDB1"`. Stored as a byte array (endianness-free). pub const MAGIC: [u8; 4] = *b"ZDB1"; -/// On-disk format version. Bumped only on an incompatible change (ADR-0002 §D8). -pub const FORMAT_VERSION: u32 = 1; +/// On-disk format version. Bumped only on an incompatible change (ADR-0002 +/// §D8). Version 2 (ADR-0022): the meta free-list annex — `fl_count` at +/// offset 168, `meta_crc` moved to 172 with split coverage, annex ids at 176. +pub const FORMAT_VERSION: u32 = 2; /// Size of the common page header, in bytes. pub const HEADER_SIZE: usize = 32; @@ -92,8 +97,14 @@ pub const MIN_KEYS_LEAF: usize = 1; /// Minimum children a non-root branch may hold. pub const MIN_KEYS_BRANCH: usize = 2; -/// Number of leading bytes of a meta page covered by its CRC (SPEC 02 §3.3). -pub const META_CONTENT_LEN: usize = 168; +/// Number of leading bytes of a meta page covered by its CRC, before the +/// free-list annex ids (SPEC 02 §3.3 as amended by ADR-0022: the full +/// coverage is `[0, META_CONTENT_LEN) ∪ [META_ANNEX_OFF, … + 8·fl_count)`). +pub const META_CONTENT_LEN: usize = 172; + +/// Absolute byte offset of the meta free-list annex ids (SPEC 02 §3, +/// ADR-0022). The annex capacity is `(psize - META_ANNEX_OFF) / 8` ids. +pub const META_ANNEX_OFF: usize = 176; /// Smallest permitted page size, in bytes. pub const MIN_PAGE_SIZE: u32 = 4096; diff --git a/crates/zerodb-core/src/readers.rs b/crates/zerodb-core/src/readers.rs index c3f6ddb..069735d 100644 --- a/crates/zerodb-core/src/readers.rs +++ b/crates/zerodb-core/src/readers.rs @@ -454,6 +454,7 @@ mod tests { last_pg: 1, main_db: DBRecord::empty(), free_db: DBRecord::empty(), + free_annex: Arc::from(&[][..]), }) } @@ -586,6 +587,7 @@ mod loom_tests { last_pg: 1, main_db: DBRecord::empty(), free_db: DBRecord::empty(), + free_annex: Arc::from(&[][..]), }) } diff --git a/crates/zerodb-core/src/rotxn.rs b/crates/zerodb-core/src/rotxn.rs index c638ea4..374966a 100644 --- a/crates/zerodb-core/src/rotxn.rs +++ b/crates/zerodb-core/src/rotxn.rs @@ -77,6 +77,13 @@ pub trait TxnRead { fn validated_pages(&self) -> Option<&ValidatedPages<'_>> { None } + /// The meta free-list annex id count this txn observes (SPEC 05 §2a, + /// ADR-0022; feeds [`free_page_count`]). A reader reports its snapshot's + /// count; the writer reports its pool's live remainder (a working-state + /// view, like [`TxnRead::free_record`]). + fn free_annex_count(&self) -> u64 { + 0 + } } /// Which database a [`Database`] handle addresses (SPEC 02 §6). `Copy` so the @@ -265,6 +272,9 @@ impl TxnRead for RoTxn<'_> { fn free_record(&self) -> &DBRecord { &self.snap.free_db } + fn free_annex_count(&self) -> u64 { + self.snap.free_annex.len() as u64 + } fn page_size(&self) -> u32 { self.psize } @@ -516,7 +526,9 @@ pub fn free_page_count(txn: &T) -> Result { let source = txn.source(); let tree = Tree::new(source, txn.page_size(), rec.root, rec.depth); let mut cursor = tree.cursor(); - let mut total = 0u64; + // Format v2 (SPEC 05 §2a/GC-23): the snapshot's meta annex ids are free + // pages too; the snapshot carries their count, so no meta re-read. + let mut total = txn.free_annex_count(); let mut entry = cursor.first().map_err(map_page_err)?; while let Some((_key, val)) = entry { // GC-23: sum only each PIL's `count` prefix — no `Vec` decode and diff --git a/crates/zerodb-core/src/rwtxn.rs b/crates/zerodb-core/src/rwtxn.rs index 5ccda95..8d307d2 100644 --- a/crates/zerodb-core/src/rwtxn.rs +++ b/crates/zerodb-core/src/rwtxn.rs @@ -196,6 +196,30 @@ impl Drain { } } +/// Merge two strictly-ascending, mutually-disjoint id slices into one +/// strictly-ascending vec (GC-3/GC-4; the `freelist_save` step-(c) merge of +/// this txn's `freed` with the carried annex remainder, SPEC 05 §2a GC-30). +/// Disjointness holds by provenance — an annex id was free in the base +/// snapshot, a `freed` id was live in it — and is asserted in debug builds. +fn merge_sorted_unique(a: &[u64], b: &[u64]) -> Vec { + let mut out = Vec::with_capacity(a.len() + b.len()); + let (mut i, mut j) = (0, 0); + while i < a.len() && j < b.len() { + if a[i] < b[j] { + out.push(a[i]); + i += 1; + } else { + debug_assert_ne!(a[i], b[j], "annex/freed overlap (INV-24)"); + out.push(b[j]); + j += 1; + } + } + out.extend_from_slice(&a[i..]); + out.extend_from_slice(&b[j..]); + debug_assert!(out.windows(2).all(|w| w[0] < w[1]), "merge not ascending"); + out +} + /// A root-to-leaf descent path: `(pgno, ki)` per level (SPEC 03 §1 cursor /// shape). `ki` is the chosen child index on branches and the entry/insertion /// slot on the leaf. @@ -606,6 +630,18 @@ pub struct RwTxn<'env> { /// [`RwTxn::free_page`]. Pgno-hashed (issue #9): engine-authored keys, /// no HashDoS surface. reclaimed: HashSet, + /// The base meta's free-list annex pool (SPEC 05 §2a, ADR-0022): the + /// pages freed by txn `base.txnid`, parsed (and `validate_pil_ids`- + /// checked, GC-33) at begin when `base.free_annex > 0`. Drawn by + /// [`RwTxn::annex_draw`] (allocate step 2b, gated per GC-31); whatever + /// remains is folded into `freed` at the top of `freelist_save` (the + /// GC-30 carry) — by then the pool is empty, so GcSave never sees it. + annex: Drain, + /// `freelist_save`'s GC-32 placement decision: `true` = this txn's own + /// entry rides in the meta annex (C4 encodes `freed` as `fl_ids`; no + /// GC-tree entry `BE(txnid)` exists), `false` = the v1 tree arm ran (or + /// `freed` is empty) and the annex is empty. + annex_out: bool, /// GC-12 allocation restriction (ADR-0005 D2). alloc_mode: AllocMode, /// Entries of `drains` whose `remaining` shrank via an **in-save pool @@ -739,6 +775,19 @@ impl Env { if base.txnid >= crate::readers::MAX_COMMITTED_TXNID { return Err(Error::Mdb(MdbError::Invalid)); } + // The base meta's free-list annex pool (SPEC 05 §2a, ADR-0022). The + // ids are pinned in the snapshot (the slot is overwritten two + // commits later, TXN-63) and are never trusted (GC-33): they must + // pass the same ascending/range validation as a tree PIL before any + // can be handed out — open-time slots only had the count bound and + // the CRC checked. + let annex = if base.free_annex.is_empty() { + Drain::new(Vec::new()) + } else { + let ids = base.free_annex.to_vec(); + validate_pil_ids(&ids, base.last_pg)?; + Drain::new(ids) + }; Ok(RwTxn { txnid: base.txnid + 1, // TXN-2 (+ the reuse note in the module docs) next_pgno: base.last_pg + 1, @@ -764,6 +813,8 @@ impl Env { loose: Vec::new(), drains: BTreeMap::new(), reclaimed: HashSet::default(), + annex, + annex_out: false, alloc_mode: AllocMode::Normal, save_touched: std::collections::BTreeSet::new(), errored: false, @@ -905,6 +956,11 @@ impl TxnRead for RwTxn<'_> { fn free_record(&self) -> &DBRecord { &self.free_db } + fn free_annex_count(&self) -> u64 { + // Working-state view (like `free_record`): the base annex pool's + // live remainder — drawn ids are no longer free. + Drain::len(&self.annex) as u64 + } fn page_size(&self) -> u32 { self.psize } @@ -1196,6 +1252,13 @@ impl<'env> RwTxn<'env> { if let Some(start) = self.gc_reclaim(n)? { return Ok(start); } + // GC-16 step 2b (SPEC 05 §2a, ADR-0022): the base meta's + // free-list annex pool, after the tree (the annex's + // freeing-txn is the newest possible, so tree-first keeps + // the GC-18/19 oldest-first order). + if let Some(start) = self.annex_draw(n) { + return Ok(start); + } } AllocMode::GcSave => { // GC-12 (as amended, ADR-0005): never *read* the GC tree @@ -1210,6 +1273,14 @@ impl<'env> RwTxn<'env> { if let Some(start) = self.save_pool_draw(n) { return Ok(start); } + // The carried annex pool (SPEC 05 §2a GC-30): gate-checked + // like every annex draw; the step-(c) placement re-reads + // `annex.live()` after every put, so a draw here can never + // leave a handed-out id in the persisted set — the exact + // GC-12-amended pool argument. + if let Some(start) = self.annex_draw(n) { + return Ok(start); + } } } let mp = map_pages(self.env.inner().map_size(), self.psize); @@ -1299,6 +1370,41 @@ impl<'env> RwTxn<'env> { Some(start) } + /// Allocate step 2b (SPEC 05 §2a GC-31, ADR-0022): draw from the base + /// meta's free-list annex pool. The pool's freeing-txnid is `base.txnid` + /// (= `writer_txnid − 1`), so the gate `F ≤ oldest_reader()` admits the + /// draw exactly when no live reader is pinned **below** the base — a + /// reader pinned *at* the base is safe (pages freed by `base` left its + /// trees), and the crash-fallback meta is `base` itself, whose annex + /// still lists the drawn ids as free (the TXN-62 / GC-12 pool argument + /// verbatim). Deterministic like GC-19: smallest id / first contiguous + /// run of the (ascending) pool. Serves both allocation modes: during + /// ops as step 2b, and inside `freelist_save` as the carried pool + /// (GC-30 — the step-(c) placement re-reads `annex.live()` after every + /// tree write, so an in-save draw is always re-accounted). + // Kept out of line: `allocate`'s hot shape is the loose pop; the draw + // arms stay separate so their size never moves it (PERF-GAP B13). + #[inline(never)] + fn annex_draw(&mut self, n: u64) -> Option { + if self.annex.len() == 0 { + return None; + } + let f = self.base.txnid; + if f > self.oldest_reader() { + return None; // GC-31 gate: a reader is pinned below the base. + } + let start = find_run(self.annex.live(), n)?; + // M1.8 shadow check, as in every other draw arm. + #[cfg(debug_assertions)] + self.debug_assert_gate(f); + self.annex.take_run(start, n); + for p in start..start + n { + let first_time = self.reclaimed.insert(p); + debug_assert!(first_time, "page {p} reclaimed twice (INV-24)"); + } + Some(start) + } + /// The GC reuse gate (SPEC 05 GC-18, SPEC 04 TXN-20/21): `min(smallest /// live reader snapshot txnid, writer_txnid − 1)`. The reader term is the /// M1.8 lock-free reader-table SeqCst scan @@ -4008,6 +4114,26 @@ impl<'env> RwTxn<'env> { debug_assert!(self.alloc_mode == AllocMode::Normal); self.alloc_mode = AllocMode::GcSave; debug_assert!(self.save_touched.is_empty()); + // GC-30 carry (SPEC 05 §2a, ADR-0022): the base annex's unconsumed + // remainder must land in this txn's own entry (annex or tree) — meta + // `base` is overwritten by txn `base + 2`, so dropping it leaks. + // The remainder stays a live **in-save pool** (the GC-12-amended + // rule, exactly like the drain pool): its ids passed the gate with + // `F = base.txnid`, so the save's own allocations may consume them — + // without this, every save's GC-tree ops extend the file while the + // carried surplus sits unreachable in `freed`, an unbounded ratchet + // (caught by churn_parity_general during the ADR-0022 spike). The + // step-(c) placement merges `freed ∪ annex.live()` and re-loops if + // the put itself drew from the pool, so no persisted set ever lists + // a handed-out page. The pool is folded into `freed` only after the + // fixed point (annex arm) or spent into the tree entry (tree arm). + // + // GC-32 placement: this txn's own entry goes to the meta annex when + // the final merged set fits; once the tree arm runs it is sticky for + // this save (no flip-flop if a later drain rewrite grows `freed`). + let annex_cap = crate::page::meta_annex_cap(self.psize); + let mut tree_arm = false; + let mut annex_mode = false; // Every ops-drained entry needs its GC-20 rewrite at least once. let mut pending: std::collections::BTreeSet = self.drains.keys().copied().collect(); // Release-active bound on the *inner* rewrite loop (GC-13 guard, ADR @@ -4065,17 +4191,34 @@ impl<'env> RwTxn<'env> { self.put_pil(f, &remaining)?; } } - // (c) this txn's own entry under BE(writer_txnid). + // (c) this txn's own entry: the meta annex when it fits (GC-32), + // else the GC tree under BE(writer_txnid) as in format v1. self.freed.sort_unstable(); let pre_dedup = self.freed.len(); self.freed.dedup(); debug_assert_eq!(self.freed.len(), pre_dedup, "page double-freed (GC-4)"); let before = self.freed.len(); - if before > 0 { - let ids = self.freed.clone(); - self.put_pil(self.txnid, &ids)?; - if self.freed.len() != before { - continue; // the put freed committed GC pages; the PIL must grow + let annex_before = Drain::len(&self.annex); + let total = before + annex_before; + if total > 0 { + if !tree_arm && total <= annex_cap { + // GC-32 annex arm: `freed ∪ annex.live()` IS the annex + // (merged after the loop; C4 encodes it as `fl_ids`). + // The stash writes nothing into the tree, allocates + // nothing and frees nothing, so the fixed point below + // simply loses the step-(c) feedback edge. + annex_mode = true; + } else { + tree_arm = true; // sticky (GC-32) + annex_mode = false; + let ids = merge_sorted_unique(&self.freed, self.annex.live()); + self.put_pil(self.txnid, &ids)?; + if self.freed.len() != before || Drain::len(&self.annex) != annex_before { + // The put COW-freed committed GC pages (PIL must + // grow) or drew from the annex pool (the written + // PIL lists a handed-out page): rewrite next round. + continue; + } } } // Fixed-point checks: no entry awaits a re-rewrite (anti-leak — a @@ -4092,6 +4235,15 @@ impl<'env> RwTxn<'env> { } self.freed.extend(std::mem::take(&mut self.loose)); } + // GC-30/GC-32 settle: in the annex arm the persisted annex is the + // merged set, materialized into `freed` for C4/C6; in the tree arm + // the pool's ids were written into the tree entry `BE(txnid)`. The + // pool is spent either way (no allocation runs after C1). + if annex_mode { + self.freed = merge_sorted_unique(&self.freed, self.annex.live()); + } + self.annex = Drain::new(Vec::new()); + self.annex_out = annex_mode; Ok(()) } @@ -4194,6 +4346,10 @@ impl<'env> RwTxn<'env> { // slot; the intact `N-1` slot is the crash fallback). ----- let last_pg = self.next_pgno - 1; let slot = self.txnid & 1; + // GC-32 (ADR-0022): in annex mode this txn's freed PIL rides in the + // meta itself — `freed` is final after C1 and already sorted/deduped + // (GC-3/4); `freelist_save` guaranteed it fits the cap. + let annex: &[u64] = if self.annex_out { &self.freed } else { &[] }; let meta = MetaPage { pgno: slot, txnid: self.txnid, @@ -4205,9 +4361,10 @@ impl<'env> RwTxn<'env> { last_pg, free_db: self.free_db, main_db: self.main_db, + fl_count: annex.len() as u32, }; let mut buf = vec![0u8; psize as usize]; - meta.encode(&mut buf).map_err(corrupt)?; + meta.encode_with_annex(&mut buf, annex).map_err(corrupt)?; inner.backing_ref().write_at_page(slot, psize, &buf)?; inner.run_hook(HookPoint::H3); // Crash here: recovered snapshot is `N-1` (meta missing/torn → CRC @@ -4270,6 +4427,11 @@ impl<'env> RwTxn<'env> { last_pg, main_db: self.main_db, free_db: self.free_db, + free_annex: if self.annex_out { + Arc::from(&self.freed[..]) + } else { + Arc::from(&[][..]) + }, })); Ok(()) // Drop of `self` releases the write mutex — the tail of C6. diff --git a/crates/zerodb-core/tests/page_edges.rs b/crates/zerodb-core/tests/page_edges.rs index abc2c76..2ec62ca 100644 --- a/crates/zerodb-core/tests/page_edges.rs +++ b/crates/zerodb-core/tests/page_edges.rs @@ -520,16 +520,18 @@ fn built_valid_meta(psize: u32) -> Vec { buf } -/// Every single-byte flip in the CRC-covered region [0, META_CONTENT_LEN) -/// must invalidate the slot; every single-byte flip in the excluded reserved -/// tail [172, psize) must NOT (SPEC 02 §3.3 — the sector-tear reality of -/// REC-22). The `meta_crc` field itself [168, 172) is not part of the hashed -/// range but flipping it desyncs stored-vs-recomputed CRC, so it also -/// invalidates. +/// Every single-byte flip in the CRC-covered region [0, META_CONTENT_LEN = +/// 172) — which since format v2 includes `fl_count` at [168, 172) (ADR-0022; +/// a silently-zeroed annex count must tear the slot) — must invalidate the +/// slot; every single-byte flip in the excluded reserved tail +/// [176, psize) must NOT, for an empty-annex meta (SPEC 02 §3.3 — the +/// sector-tear reality of REC-22). The `meta_crc` field itself [172, 176) is +/// not part of the hashed range but flipping it desyncs stored-vs-recomputed +/// CRC, so it also invalidates. fn meta_validation_rejection_matrix(psize: u32) { let buf = built_valid_meta(psize); - for i in 0..168usize { + for i in 0..172usize { let mut b = buf.clone(); b[i] ^= 0xFF; let v = MetaPage::validate(&b, psize).unwrap(); @@ -538,7 +540,7 @@ fn meta_validation_rejection_matrix(psize: u32) { "byte {i} (CRC-covered) flip must invalidate meta" ); } - for i in 168..172usize { + for i in 172..176usize { let mut b = buf.clone(); b[i] ^= 0xFF; let v = MetaPage::validate(&b, psize).unwrap(); @@ -549,9 +551,11 @@ fn meta_validation_rejection_matrix(psize: u32) { } // Sample the reserved tail rather than iterating every byte (up to ~65KB) // to stay within the ~30s test-runtime budget; every sampled byte must - // NOT invalidate, documenting that this range is unprotected by the CRC. - let stride = ((psize as usize - 172) / 40).max(1); - for i in (172..psize as usize).step_by(stride) { + // NOT invalidate, documenting that this range is unprotected by the CRC + // when the annex is empty (with a non-empty annex the ids' bytes ARE + // covered — the spec02_format annex tests lock that). + let stride = ((psize as usize - 176) / 40).max(1); + for i in (176..psize as usize).step_by(stride) { let mut b = buf.clone(); b[i] ^= 0xFF; let v = MetaPage::validate(&b, psize).unwrap(); diff --git a/crates/zerodb-core/tests/spec02_format.rs b/crates/zerodb-core/tests/spec02_format.rs index b0d26a5..aaaf91e 100644 --- a/crates/zerodb-core/tests/spec02_format.rs +++ b/crates/zerodb-core/tests/spec02_format.rs @@ -16,15 +16,16 @@ const PSIZE: u32 = 4096; // §3.5 — creation meta, slot 0 (page size 4096, txnid 0, map_size 1 MiB) // --------------------------------------------------------------------------- -/// Build the expected first-168 bytes of the §3.5 creation meta. -fn expected_meta_content() -> [u8; 168] { - let mut e = [0u8; 168]; +/// Build the expected first-172 bytes of the §3.5 creation meta (format v2, +/// ADR-0022: `fl_count` at 168 — zero for a creation meta — then the CRC). +fn expected_meta_content() -> [u8; 172] { + let mut e = [0u8; 172]; // off 0: pgno = 0 (all zero) // off 8: txnid = 0 (all zero) e[16] = 0x08; // flags = P_META (0x0008) // off 18..32 reserved 0 e[32..36].copy_from_slice(b"ZDB1"); // magic 5A 44 42 31 - e[36] = 0x01; // format_version = 1 + e[36] = 0x02; // format_version = 2 (ADR-0022) e[40] = 0x00; e[41] = 0x10; // page_size = 4096 (0x0000_1000) // off 44 env_flags = 0 @@ -39,6 +40,7 @@ fn expected_meta_content() -> [u8; 168] { *b = 0xFF; // main_db.root = PGNO_INVALID } // off 128..168 main_db stats 0 + // off 168..172 fl_count = 0 (empty free-list annex) e } @@ -48,21 +50,22 @@ fn meta_creation_slot0_bytes() { let mut buf = vec![0u8; PSIZE as usize]; meta.encode(&mut buf).unwrap(); - // Lock [0, 168) byte-for-byte. + // Lock [0, 172) byte-for-byte. assert_eq!( - &buf[..168], + &buf[..172], &expected_meta_content()[..], - "meta content [0,168)" + "meta content [0,172)" ); - // The CRC field must be exactly CRC32C over [0, 168). - let crc = page::crc32c(&buf[..168]); - let stored = u32::from_le_bytes([buf[168], buf[169], buf[170], buf[171]]); - assert_eq!(stored, crc, "meta_crc must cover [0,168)"); + // The CRC field must be exactly CRC32C over [0, 172) — fl_count is 0, so + // no annex ids extend the coverage (SPEC 02 §3.3, format v2). + let crc = page::crc32c(&buf[..172]); + let stored = u32::from_le_bytes([buf[172], buf[173], buf[174], buf[175]]); + assert_eq!(stored, crc, "meta_crc must cover [0,172) when fl_count = 0"); - // The reserved tail [172, psize) must be all zero (excluded from CRC). + // The reserved tail [176, psize) must be all zero (excluded from CRC). assert!( - buf[172..].iter().all(|&b| b == 0), + buf[176..].iter().all(|&b| b == 0), "reserved tail must be zero" ); @@ -94,7 +97,122 @@ fn meta_slots_identical_at_creation() { // Only the pgno field (offset 0) differs between the slots. assert_eq!(b0[0], 0); assert_eq!(b1[0], 1); - assert_eq!(&b0[8..168], &b1[8..168], "bodies identical apart from pgno"); + assert_eq!(&b0[8..172], &b1[8..172], "bodies identical apart from pgno"); +} + +// --------------------------------------------------------------------------- +// §3 format v2 — the meta free-list annex (ADR-0022; SPEC 05 §2a GC-29/33) +// --------------------------------------------------------------------------- + +#[test] +fn meta_annex_encode_validate_read_roundtrip() { + let ids: Vec = vec![2, 5, 6, 7, 40]; + let mut meta = MetaPage::create(0, PSIZE, 1024 * 1024); + meta.last_pg = 64; + meta.fl_count = ids.len() as u32; + let mut buf = vec![0u8; PSIZE as usize]; + meta.encode_with_annex(&mut buf, &ids).unwrap(); + + // Field placement: fl_count at 168 (LE), ids from 176 (LE u64 each). + assert_eq!(&buf[168..172], &(ids.len() as u32).to_le_bytes()); + for (i, id) in ids.iter().enumerate() { + assert_eq!(&buf[176 + 8 * i..184 + 8 * i], &id.to_le_bytes()); + } + // CRC coverage = [0, 172) ∪ the annex ids (SPEC 02 §3.3), one stream. + let crc = page::crc32c_concat(&[&buf[..172], &buf[176..176 + 8 * ids.len()]]); + assert_eq!(&buf[172..176], &crc.to_le_bytes()); + // Equivalence with the single-shot CRC over the joined bytes. + let mut joined = buf[..172].to_vec(); + joined.extend_from_slice(&buf[176..176 + 8 * ids.len()]); + assert_eq!(crc, page::crc32c(&joined)); + + // Validates and decodes; the ids read back exactly. + match MetaPage::validate(&buf, PSIZE).unwrap() { + MetaValidity::Valid(decoded) => { + assert_eq!(decoded.fl_count, ids.len() as u32); + assert_eq!(decoded, meta); + } + other => panic!("expected Valid, got {other:?}"), + } + assert_eq!(MetaPage::read_annex(&buf, PSIZE).unwrap(), ids); +} + +#[test] +fn meta_annex_torn_id_rejected_by_crc() { + let ids: Vec = (2..60).collect(); + let mut meta = MetaPage::create(0, PSIZE, 1024 * 1024); + meta.last_pg = 64; + meta.fl_count = ids.len() as u32; + let mut buf = vec![0u8; PSIZE as usize]; + meta.encode_with_annex(&mut buf, &ids).unwrap(); + // Flip one byte inside the LAST annex id — far past META_CONTENT_LEN. + buf[176 + 8 * (ids.len() - 1)] ^= 0xFF; + assert!( + matches!( + MetaPage::validate(&buf, PSIZE).unwrap(), + MetaValidity::BadCrc { .. } + ), + "a torn annex id must tear the slot (REC-8 annex note)" + ); +} + +#[test] +fn meta_annex_zeroed_count_rejected_by_crc() { + // The leak guard: a torn write that silently zeroes `fl_count` (losing + // the freed-page record) must invalidate the slot, because fl_count is + // inside the main CRC coverage (SPEC 02 §3.3). + let ids: Vec = vec![2, 3, 4]; + let mut meta = MetaPage::create(0, PSIZE, 1024 * 1024); + meta.last_pg = 8; + meta.fl_count = ids.len() as u32; + let mut buf = vec![0u8; PSIZE as usize]; + meta.encode_with_annex(&mut buf, &ids).unwrap(); + buf[168..172].fill(0); // fl_count := 0, stale CRC left in place + buf[176..176 + 24].fill(0); // ids zeroed too (torn-to-zeros tail) + assert!( + matches!( + MetaPage::validate(&buf, PSIZE).unwrap(), + MetaValidity::BadCrc { .. } + ), + "zeroing the annex must tear the slot, never read as 'no annex'" + ); +} + +#[test] +fn meta_annex_count_over_cap_rejected_before_crc() { + let meta = MetaPage::create(0, PSIZE, 1024 * 1024); + let mut buf = vec![0u8; PSIZE as usize]; + meta.encode(&mut buf).unwrap(); + // A hostile count far past the page's capacity must be refused by the + // bound check (rule 5, before the CRC read), not read out of bounds. + buf[168..172].copy_from_slice(&u32::MAX.to_le_bytes()); + assert!(matches!( + MetaPage::validate(&buf, PSIZE).unwrap(), + MetaValidity::BadAnnexCount(_) + )); + assert_eq!(MetaPage::read_annex(&buf, PSIZE), None); + // Encoding more ids than fit is a typed error, never a panic. + let too_many: Vec = (2..2 + page::meta_annex_cap(PSIZE) as u64 + 1).collect(); + let mut m2 = MetaPage::create(0, PSIZE, 1024 * 1024); + m2.fl_count = too_many.len() as u32; + assert!(m2.encode_with_annex(&mut buf, &too_many).is_err()); +} + +#[test] +fn meta_annex_cap_arithmetic() { + // (psize − 176) / 8, per SPEC 02 §3. + assert_eq!(page::meta_annex_cap(4096), (4096 - 176) / 8); + assert_eq!(page::meta_annex_cap(65536), (65536 - 176) / 8); + // A full-cap annex encodes and validates. + let cap = page::meta_annex_cap(PSIZE); + let ids: Vec = (2..2 + cap as u64).collect(); + let mut meta = MetaPage::create(0, PSIZE, 1024 * 1024); + meta.last_pg = ids.last().copied().unwrap() + 1; + meta.fl_count = cap as u32; + let mut buf = vec![0u8; PSIZE as usize]; + meta.encode_with_annex(&mut buf, &ids).unwrap(); + assert!(MetaPage::validate(&buf, PSIZE).unwrap().is_valid()); + assert_eq!(MetaPage::read_annex(&buf, PSIZE).unwrap(), ids); } // --------------------------------------------------------------------------- diff --git a/crates/zerodb-oracle/tests/crash_harness_smoke.rs b/crates/zerodb-oracle/tests/crash_harness_smoke.rs index 2b29a75..96d776d 100644 --- a/crates/zerodb-oracle/tests/crash_harness_smoke.rs +++ b/crates/zerodb-oracle/tests/crash_harness_smoke.rs @@ -152,7 +152,18 @@ fn regression_nometasync_reclaim_clobber_window_is_characterized() { // it never touches `Op` — so the mode guard below could not catch it). // The assertions are unchanged; only the vehicle was replaced, by searching // for a seed that still satisfies all four of them. - const SEED: u64 = 11_834_834_180_059_103_290; + // + // RE-PINNED 2026-10-03 (was 11_834_834_180_059_103_290): ADR-0022 (the + // meta free-list annex, format v2) changes which pages a txn reclaims + // when — small freed sets ride the meta and are reused one hop earlier — + // so the old seed's cut stopped producing a stale-fallback image (0 + // stale; nothing was violated, the vehicle was lost again). Same re-pin + // protocol: seed 198 satisfies all four assertions (12 verified, 4 stale + // fallbacks, NoMetaSync, no violation). The clobber window itself is + // unchanged by the annex — the reclaimed-page TXN-62/GC-18 reasoning is + // identical whether the freed list lived in the tree or the meta. + // Re-pin accepted by the maintainer (Quentin, 2026-10-05). + const SEED: u64 = 198; assert_eq!( gen_spec(SEED).mode, Mode::NoMetaSync, diff --git a/crates/zerodb-tools/tests/data_file_probe.rs b/crates/zerodb-tools/tests/data_file_probe.rs index b5bb348..2f2ce50 100644 --- a/crates/zerodb-tools/tests/data_file_probe.rs +++ b/crates/zerodb-tools/tests/data_file_probe.rs @@ -92,7 +92,7 @@ fn stat_prints_the_engine_identification_line() { .expect("stat must print an engine line"); assert!(line.contains("zerodb"), "engine line: {line}"); assert!(line.contains("ZDB1"), "engine line: {line}"); - assert!(line.contains("format_version 1"), "engine line: {line}"); + assert!(line.contains("format_version 2"), "engine line: {line}"); assert!(line.contains("data.mdb"), "engine line: {line}"); } diff --git a/crates/zerodb/src/copy.rs b/crates/zerodb/src/copy.rs index 3cf7668..e8a78cb 100644 --- a/crates/zerodb/src/copy.rs +++ b/crates/zerodb/src/copy.rs @@ -388,6 +388,10 @@ fn write_raw( // Synthesize both meta slots from the pinned snapshot — the copy is a // self-contained env at snapshot T, freelist preserved (SPEC 02 §3). + // Format v2 (ADR-0022): the snapshot's free-list annex rides along — + // the ids are pinned in `Snapshot` (the live env's slot may already + // hold a newer meta, TXN-63), and dropping them would leak those pages + // in the copy (INV-22/INV-28). let mut metas = vec![0u8; 2 * ps]; let mut meta = MetaPage { pgno: 0, @@ -400,10 +404,13 @@ fn write_raw( last_pg: snap.last_pg, free_db: snap.free_db, main_db: snap.main_db, + fl_count: snap.free_annex.len() as u32, }; - meta.encode(&mut metas[0..ps]).map_err(corrupt)?; + meta.encode_with_annex(&mut metas[0..ps], &snap.free_annex) + .map_err(corrupt)?; meta.pgno = 1; - meta.encode(&mut metas[ps..2 * ps]).map_err(corrupt)?; + meta.encode_with_annex(&mut metas[ps..2 * ps], &snap.free_annex) + .map_err(corrupt)?; file.write_all_at(&metas, base)?; // Data pages in chunks, so progress is observable. The chunk is sized diff --git a/crates/zerodb/tests/loose_page_and_trailing_shrink.rs b/crates/zerodb/tests/loose_page_and_trailing_shrink.rs index f709135..1702716 100644 --- a/crates/zerodb/tests/loose_page_and_trailing_shrink.rs +++ b/crates/zerodb/tests/loose_page_and_trailing_shrink.rs @@ -98,11 +98,12 @@ fn same_txn_alloc_then_free_overflow_reused_leaves_only_the_leaf_cow() { .len(); assert_eq!( size_after, - size_before + 2 * PS as u64, + size_before + PS as u64, "the whole 5-page overflow run must be reclaimed (GC-7/8/10); the only \ - durable growth allowed is the one-page leaf COW plus the one GC leaf \ - page recording the freed old leaf (M1.5 freelist_save, SPEC 05 GC-11 — \ - pre-M1.5 this txn leaked the old leaf instead of listing it)" + durable growth allowed is the one-page leaf COW — the freed old leaf \ + is recorded in the meta free-list annex (SPEC 05 §2a GC-29, ADR-0022, \ + format v2), no longer in a GC tree leaf (format v1 grew one extra GC \ + leaf page here; pre-M1.5 the old leaf leaked instead of being listed)" ); let rtxn = env.read_txn().unwrap(); @@ -155,11 +156,11 @@ fn trailing_loose_pages_shrink_next_pgno() { .len(); assert_eq!( size_after, - size_seed + 2 * PS as u64, + size_seed + PS as u64, "a 10-page trailing alloc-then-free must shrink back to just the one \ - unavoidable leaf-COW page (GC-10) plus the one GC leaf page recording \ - the freed old leaf (M1.5 freelist_save), not leak any of the 10 \ - overflow pages" + unavoidable leaf-COW page (GC-10) — the freed old leaf is recorded in \ + the meta free-list annex (SPEC 05 §2a GC-29, ADR-0022, format v2), not \ + in a GC tree leaf — and must not leak any of the 10 overflow pages" ); // The env is still fully usable afterward (the rolled-back pgnos are diff --git a/crates/zerodb/tests/meta_annex_gc.rs b/crates/zerodb/tests/meta_annex_gc.rs new file mode 100644 index 0000000..940f25b --- /dev/null +++ b/crates/zerodb/tests/meta_annex_gc.rs @@ -0,0 +1,429 @@ +//! Meta free-list annex behavior (SPEC 05 §2a GC-29..33, SPEC 02 §3 format +//! v2, ADR-0022): steady-state placement, the GC-31 reader gate, the GC-30 +//! carry, the GC-32 spill-to-tree, copy, and reopen — each asserted through +//! observable state (file size, `check_image`, `free_page_count`, the meta +//! bytes) rather than internals. + +use std::path::{Path, PathBuf}; +use std::sync::atomic::{AtomicU64, Ordering}; + +use zerodb::{check, free_page_count, CompactionOption, CopyToFile, Env, EnvFlags, EnvOpenOptions}; +use zerodb_core::page::{MetaPage, MetaValidity, PGNO_INVALID}; + +static COUNTER: AtomicU64 = AtomicU64::new(0); + +struct TempDir { + path: PathBuf, +} + +impl TempDir { + fn new() -> TempDir { + let pid = std::process::id(); + let seq = COUNTER.fetch_add(1, Ordering::Relaxed); + let path = std::env::temp_dir().join(format!("zerodb-annex-{pid}-{seq}")); + std::fs::create_dir_all(&path).expect("create temp dir"); + TempDir { path } + } + fn path(&self) -> &Path { + &self.path + } +} + +impl Drop for TempDir { + fn drop(&mut self) { + let _ = std::fs::remove_dir_all(&self.path); + } +} + +const PS: u32 = 4096; +const MAP: usize = 32 << 20; + +fn open(dir: &Path, flags: EnvFlags) -> Env { + let mut opts = EnvOpenOptions::new(); + opts.map_size(MAP); + opts.page_size(PS); + opts.flags(flags); + opts.open(dir).expect("open env") +} + +/// Non-vacuousness tripwire for the WRITE_MAP twins (ADR-0021/ADR-0022 +/// convention, see `nested_fanout.rs::assert_mode`): the txn's dirty-page +/// realization must match the env flags it was opened with, so a twin can +/// never silently run the default heap-staged path instead of in-place +/// WRITE_MAP. +fn assert_mode(wtxn: &zerodb::RwTxn<'_>, flags: EnvFlags) { + assert_eq!( + wtxn.dirty_in_map_mode(), + flags.contains(EnvFlags::WRITE_MAP), + "dirty-page realization does not match the env flags" + ); +} + +fn data_file(dir: &Path) -> PathBuf { + dir.join(zerodb::DATA_FILE_NAME) +} + +fn assert_clean(dir: &Path) { + let bytes = std::fs::read(data_file(dir)).expect("read data file"); + let v = check::check_image(&bytes, PS); + assert!(v.is_empty(), "invariant violations: {v:#?}"); +} + +fn file_size(dir: &Path) -> u64 { + std::fs::metadata(data_file(dir)).unwrap().len() +} + +/// Growth metric that stays meaningful under both dirty-page realizations. +/// The default backing grows the file by exactly what each commit needs, so +/// `file_size` tracks page reuse vs. extension precisely. Under `WRITE_MAP` +/// the file is `set_len`d to the full `map_size` at map time (see +/// `zerodb-core::env`'s read-annex comment on the geometry-validation length +/// check), so `file_size` is pinned from the first commit on and can never +/// show growth — the meta's logical page high-water mark (`last_pg`) is the +/// mode-independent stand-in: it only advances when a commit extends past +/// the previous high water, and stays put when a commit's allocations are all +/// satisfied by annex/pool reuse. +fn high_water(dir: &Path, flags: EnvFlags) -> u64 { + if !flags.contains(EnvFlags::WRITE_MAP) { + return file_size(dir); + } + let bytes = std::fs::read(data_file(dir)).unwrap(); + let ps = PS as usize; + let pick = |b: &[u8]| match MetaPage::validate(b, PS).unwrap() { + MetaValidity::Valid(m) => Some(m), + _ => None, + }; + let m0 = pick(&bytes[..ps]); + let m1 = pick(&bytes[ps..2 * ps]); + let m = match (m0, m1) { + (Some(a), Some(b)) => { + if a.txnid >= b.txnid { + a + } else { + b + } + } + (Some(a), None) => a, + (None, Some(b)) => b, + (None, None) => panic!("no valid meta slot"), + }; + m.last_pg +} + +/// Read the LIVE meta's `(txnid, fl_count, free_db_root)` straight from the +/// data file (both slots validated, higher txnid wins — SPEC 02 §3.2). +fn live_meta(dir: &Path) -> (u64, u32, u64) { + let bytes = std::fs::read(data_file(dir)).unwrap(); + let ps = PS as usize; + let pick = |b: &[u8]| match MetaPage::validate(b, PS).unwrap() { + MetaValidity::Valid(m) => Some(m), + _ => None, + }; + let m0 = pick(&bytes[..ps]); + let m1 = pick(&bytes[ps..2 * ps]); + let m = match (m0, m1) { + (Some(a), Some(b)) => { + if a.txnid >= b.txnid { + a + } else { + b + } + } + (Some(a), None) => a, + (None, Some(b)) => b, + (None, None) => panic!("no valid meta slot"), + }; + (m.txnid, m.fl_count, m.free_db.root) +} + +/// GC-29/GC-32 steady state: small-churn commits keep the whole free list in +/// the meta annex — the GC tree never materializes — and page reuse keeps the +/// file size flat. +#[test] +fn steady_state_annex_only_gc_tree_stays_empty() { + steady_state_annex_only_gc_tree_stays_empty_inner(EnvFlags::EMPTY); +} + +/// WRITE_MAP in-place twin (ADR-0021 M3 convention): the same steady-state +/// churn, but with dirty pages realized in the writable map. The annex draw +/// (ids reused from the meta's loose pool, not fresh-extended pages — see the +/// flat `file_size` assertions below) is orthogonal to WRITE_MAP's dirty-page +/// realization, so it must behave identically; `assert_mode` is the tripwire +/// that the twin actually engaged in-place writes instead of silently running +/// the default heap-staged path. +#[test] +fn steady_state_annex_only_gc_tree_stays_empty_writemap_in_place() { + steady_state_annex_only_gc_tree_stays_empty_inner(EnvFlags::WRITE_MAP); +} + +fn steady_state_annex_only_gc_tree_stays_empty_inner(flags: EnvFlags) { + let dir = TempDir::new(); + let env = open(dir.path(), flags); + let db = env.main_database(); + + let val = vec![0xABu8; 256]; + for i in 0u64..60 { + let mut w = env.write_txn().unwrap(); + assert_mode(&w, flags); + db.put(&mut w, &i.to_be_bytes(), &val).unwrap(); + w.commit().unwrap(); + } + let (_, fl_mid, root_mid) = live_meta(dir.path()); + assert!(fl_mid > 0, "steady-state commits must carry a meta annex"); + assert_eq!( + root_mid, PGNO_INVALID, + "small churn must never materialize the GC tree (GC-29/GC-32)" + ); + + // Overwrite churn: every commit frees the COW'd path and reuses the + // previous commit's annex pages. After a few settling commits the file + // must not grow at all. + for i in 0u64..10 { + let mut w = env.write_txn().unwrap(); + db.put(&mut w, &i.to_be_bytes(), &val).unwrap(); + w.commit().unwrap(); + } + let size_settled = high_water(dir.path(), flags); + for i in 0u64..50 { + let mut w = env.write_txn().unwrap(); + db.put(&mut w, &i.to_be_bytes(), &val).unwrap(); + w.commit().unwrap(); + } + assert_eq!( + high_water(dir.path(), flags), + size_settled, + "annex reuse must keep overwrite churn at zero growth (file size \ + under the default backing, last_pg under WRITE_MAP)" + ); + assert_clean(dir.path()); + + // free_page_count sees the annex ids (GC-23 as amended). + let rtxn = env.read_txn().unwrap(); + let (_, fl_now, root_now) = live_meta(dir.path()); + assert_eq!(root_now, PGNO_INVALID); + assert_eq!(free_page_count(&rtxn).unwrap(), u64::from(fl_now)); +} + +/// GC-31 gate + GC-30 carry: a reader pinned below the base blocks annex +/// draws (the file grows while it lives, nothing leaks), and after it +/// releases, the carried ids are reclaimed and growth stops. +#[test] +fn reader_gate_blocks_annex_draw_then_carry_reclaims() { + reader_gate_blocks_annex_draw_then_carry_reclaims_inner(EnvFlags::EMPTY); +} + +/// WRITE_MAP in-place twin (ADR-0021 M3 convention) of the parked-reader +/// case: the GC-31 gate and GC-30 carry are meta free-list bookkeeping, +/// independent of whether dirty pages are realized in the writable map, so +/// the same growth-then-settle shape must hold. `assert_mode` tripwires that +/// the twin really ran in-place. +#[test] +fn reader_gate_blocks_annex_draw_then_carry_reclaims_writemap_in_place() { + reader_gate_blocks_annex_draw_then_carry_reclaims_inner(EnvFlags::WRITE_MAP); +} + +fn reader_gate_blocks_annex_draw_then_carry_reclaims_inner(flags: EnvFlags) { + let dir = TempDir::new(); + let env = open(dir.path(), flags); + let db = env.main_database(); + let val = vec![0x44u8; 256]; + + for i in 0u64..8 { + let mut w = env.write_txn().unwrap(); + db.put(&mut w, &i.to_be_bytes(), &val).unwrap(); + w.commit().unwrap(); + } + // Pin NOW; the next commit's base will be newer than this pin, so every + // later annex (and tree entry) is gated off while the reader lives. + let pin = env.read_txn().unwrap(); + { + let mut w = env.write_txn().unwrap(); + assert_mode(&w, flags); + db.put(&mut w, &0u64.to_be_bytes(), &val).unwrap(); + w.commit().unwrap(); + } + let size_pinned_base = high_water(dir.path(), flags); + for i in 0u64..6 { + let mut w = env.write_txn().unwrap(); + db.put(&mut w, &i.to_be_bytes(), &val).unwrap(); + w.commit().unwrap(); + } + let size_during = high_water(dir.path(), flags); + assert!( + size_during > size_pinned_base, + "with a reader pinned below every base, commits must extend (file \ + size under the default backing, last_pg under WRITE_MAP) — the \ + GC-31 gate refuses the annex — got no growth, so the gate leaked" + ); + assert_clean(dir.path()); + drop(pin); + // Carried ids (GC-30: each commit re-listed the still-free remainder) + // become reclaimable; churn must stop growing the file. + for _ in 0..3 { + let mut w = env.write_txn().unwrap(); + db.put(&mut w, &1u64.to_be_bytes(), &val).unwrap(); + w.commit().unwrap(); + } + let size_settled = high_water(dir.path(), flags); + for i in 0u64..8 { + let mut w = env.write_txn().unwrap(); + db.put(&mut w, &i.to_be_bytes(), &val).unwrap(); + w.commit().unwrap(); + } + assert_eq!( + high_water(dir.path(), flags), + size_settled, + "after the reader releases, carried annex ids must satisfy churn \ + with no further growth" + ); + assert_clean(dir.path()); +} + +/// GC-32 all-or-nothing spill: a commit freeing more ids than the annex cap +/// writes the whole PIL to the GC tree (annex empty), and those pages are +/// reclaimable afterwards. +#[test] +fn over_cap_freed_set_spills_whole_to_gc_tree() { + let dir = TempDir::new(); + let env = open(dir.path(), EnvFlags::EMPTY); + let db = env.main_database(); + let cap = zerodb_core::page::meta_annex_cap(PS) as u64; // 490 at 4 KiB + + // One huge overflow value, committed... + { + let mut w = env.write_txn().unwrap(); + db.put( + &mut w, + b"huge", + &vec![0x7Au8; (cap as usize + 30) * PS as usize], + ) + .unwrap(); + w.commit().unwrap(); + } + // ...then freed in the next txn: freed > cap ⇒ the tree arm. + { + let mut w = env.write_txn().unwrap(); + assert!(db.delete(&mut w, b"huge").unwrap()); + w.commit().unwrap(); + } + let (_, fl, root) = live_meta(dir.path()); + assert_eq!( + fl, 0, + "an over-cap freed set must leave the annex empty (GC-32)" + ); + assert_ne!(root, PGNO_INVALID, "the spill must land in the GC tree"); + assert_clean(dir.path()); + { + let rtxn = env.read_txn().unwrap(); + assert!(free_page_count(&rtxn).unwrap() > cap); + } + + // The spilled pages are drawn back out (tree first, GC-18): re-inserting + // a similar value must not grow the file beyond its high-water. + let size_spilled = file_size(dir.path()); + { + let mut w = env.write_txn().unwrap(); + db.put(&mut w, b"huge2", &vec![0x7Bu8; cap as usize * PS as usize]) + .unwrap(); + w.commit().unwrap(); + } + assert!( + file_size(dir.path()) <= size_spilled, + "re-insert must reuse the spilled run, not extend" + ); + assert_clean(dir.path()); +} + +/// Non-compact copy carries the pinned snapshot's annex (the live slot may +/// already hold a newer meta — the ids are pinned in the Snapshot), so the +/// copy is leak-free and reports the same free count. The compact copy drops +/// free pages by construction and must also come out clean. +#[test] +fn copy_preserves_annex_compact_drops_it() { + let dir = TempDir::new(); + let env = open(dir.path(), EnvFlags::EMPTY); + let db = env.main_database(); + let val = vec![0x55u8; 256]; + for i in 0u64..20 { + let mut w = env.write_txn().unwrap(); + db.put(&mut w, &i.to_be_bytes(), &val).unwrap(); + w.commit().unwrap(); + } + let (_, fl, _) = live_meta(dir.path()); + assert!(fl > 0, "fixture must have a non-empty annex"); + let src_free = { + let rtxn = env.read_txn().unwrap(); + free_page_count(&rtxn).unwrap() + }; + + let raw = TempDir::new(); + let raw_path = raw.path().join("copy-raw.zdb"); + env.copy_to_file(&raw_path, CompactionOption::Disabled) + .unwrap(); + let bytes = std::fs::read(&raw_path).unwrap(); + let v = check::check_image(&bytes, PS); + assert!(v.is_empty(), "raw copy violations: {v:#?}"); + let copy_env = { + let dirp = raw.path().join("raw-env"); + std::fs::create_dir_all(&dirp).unwrap(); + std::fs::copy(&raw_path, dirp.join(zerodb::DATA_FILE_NAME)).unwrap(); + open(&dirp, EnvFlags::EMPTY) + }; + { + let rtxn = copy_env.read_txn().unwrap(); + assert_eq!( + free_page_count(&rtxn).unwrap(), + src_free, + "raw copy must preserve the freelist, annex included" + ); + assert_eq!( + copy_env + .main_database() + .get(&rtxn, &3u64.to_be_bytes()) + .unwrap(), + Some(val.as_slice()) + ); + } + + let compact_path = raw.path().join("copy-compact.zdb"); + env.copy_to_file(&compact_path, CompactionOption::Enabled) + .unwrap(); + let bytes = std::fs::read(&compact_path).unwrap(); + let v = check::check_image(&bytes, PS); + assert!(v.is_empty(), "compact copy violations: {v:#?}"); +} + +/// Reopen re-parses the annex from the selected slot (GC-33 validation at the +/// writer's begin): reuse keeps working across a close/open cycle. +#[test] +fn reopen_reparses_annex_and_reuses() { + let dir = TempDir::new(); + { + let env = open(dir.path(), EnvFlags::EMPTY); + let db = env.main_database(); + let val = vec![0x66u8; 256]; + for i in 0u64..20 { + let mut w = env.write_txn().unwrap(); + db.put(&mut w, &i.to_be_bytes(), &val).unwrap(); + w.commit().unwrap(); + } + } + let (_, fl, _) = live_meta(dir.path()); + assert!(fl > 0); + let size_before = file_size(dir.path()); + + let env = open(dir.path(), EnvFlags::EMPTY); + let db = env.main_database(); + let val = vec![0x66u8; 256]; + for i in 0u64..10 { + let mut w = env.write_txn().unwrap(); + db.put(&mut w, &i.to_be_bytes(), &val).unwrap(); + w.commit().unwrap(); + } + assert_eq!( + file_size(dir.path()), + size_before, + "the reopened env must draw from the persisted annex, not extend" + ); + assert_clean(dir.path()); +} diff --git a/docs/DECISIONS.md b/docs/DECISIONS.md index 44731d1..8526ae5 100644 --- a/docs/DECISIONS.md +++ b/docs/DECISIONS.md @@ -24,3 +24,4 @@ | 0019 | [Durable meta write through an O_DSYNC descriptor — one barrier per durable commit, as LMDB's `me_mfd`](adr/0019-meta-write-dsync.md) | Accepted (all platforms; LMDB's failed-write scrub adopted) | Phase 3 (durable-commit cost) | | 0020 | [Shrink the 32-byte common page header (24 bytes, keeping pgno and the writer stamp)](adr/0020-compact-page-header.md) | Spike approved; format change pending its numbers | Phase 3 (density) | | 0021 | [True in-place WRITE_MAP — dirty pages in the map, no heap staging, no commit write-back](adr/0021-writemap-in-place.md) | Accepted (2026-10-02, Quentin); spike first — fair WRITE_MAP comparison shows ZeroDB-writemap 0.42–0.53× of LMDB-writemap | Phase 3 (PERF-GAP B5, issue #13) | +| 0022 | [Meta free-list annex — the per-commit freed PIL rides inside the (CRC-covered) meta page; GC tree becomes the spill/cold path; FORMAT_VERSION 2](adr/0022-meta-freelist-annex.md) | Accepted (2026-10-05, Quentin; format v2 ratified) | perf lever B12 (free-list save); feeds PLAN 3.1 | diff --git a/docs/DIVERGENCES.md b/docs/DIVERGENCES.md index 2914a10..9081cd3 100644 --- a/docs/DIVERGENCES.md +++ b/docs/DIVERGENCES.md @@ -33,6 +33,7 @@ Rules: | D-019 | Page-validation policy (ADR-0014) | LMDB never validates page contents: cell pointers, key sizes, value spans and overflow references are used as read (`mdb_page_get` checks only `pgno < mt_next_pgno`; the descent checks only that the bottom page is a leaf). A corrupt page is undefined behaviour. | ZeroDB validates by default (a corrupt page is a typed error) and adds an **opt-in** trusting policy, `FileTrust::trust_contents()` (an `unsafe` constructor passed to the safe `EnvOpenOptions::file_trust` setter), under which map pages skip the per-cell walk and behave like LMDB's. Page-number bounds, page type, meta validation and free-list id checks stay on in both policies. Overflow runs differ: the validating default checks each run's header; the trusting policy reads an overflow value without its run header (as `mdb_node_read` does), keeping only the snapshot high-water bound, so a corrupt run can return wrong bytes but never read outside the committed map (ADR-0014 amendment, 2026-09-29). A ZeroDB extension, not a behaviour change: the default is unchanged and no consumer sees the option unless it calls it. | APPROVED — opt-in, default validating | Quentin de Quelen (chat, 2026-09-26: "for the 'trusting the file' it would be great if it could be an option"; 2026-09-28: "do the read path now") | | D-020 | Sequential-writes fast path (ADR-0015) | `mdb_put` sets up a fresh cursor and descends from the root on every call; `MDB_APPEND` re-descends through `mdb_cursor_last`. | An **opt-in** env option (`EnvOpenOptions::sequential_writes`, default off) with a per-database runtime override (`Env::set_sequential_writes`) lets a write txn remember each tree's rightmost-leaf path and append there without descending when a key sorts after every key in that leaf (SPEC 03 §6.6). Results and committed files are identical with it on or off; only speed changes (ascending/APPEND loads faster, random-key writes 3–5 % slower). A ZeroDB extension, off unless called. | APPROVED — opt-in, default off | Quentin de Quelen (chat, 2026-09-28: "can the sequential writes be an option of the env"; env option plus per-database override: "perfect") | | D-021 | Configurable dirty limit (ADR-0017) | The spill threshold is a compile-time constant (`MDB_IDL_UM_MAX`, 131,072 pages); `mdb_page_spill` writes 1/8 of the dirty list once the txn's dirty room runs low. | ZeroDB spills the same way with the same default (SPEC 04 §6.3a) and adds an **opt-in** env option, `EnvOpenOptions::max_dirty_bytes`, to set the limit in bytes. A ZeroDB extension, not a behaviour change: unset, the limit is LMDB's, and results and committed files never depend on it. | APPROVED — opt-in, default LMDB's | Quentin de Quelen (chat, 2026-09-30: "go continue, I should be closer to LMDB in memory usage") | +| D-022 | Free-list placement (ADR-0022, meta annex) | The freed-page list of every commit is written into FREE_DBI (the GC B-tree) under `BE(txnid)` unconditionally; the GC tree is never empty under churn and each commit COWs a GC leaf. | Each commit's freed list rides inside its meta page's reserved bytes (the "annex"); the GC tree stays empty in steady state and is used only when a commit's freed set exceeds the annex cap or a reader parks a carried list. **On-disk format only** (`FORMAT_VERSION` 1→2); heed-level behaviour, allocation-order determinism and all query results are identical. Tools-observable (`zerodb stat` shows an empty FREE_DBI in steady state; churn files are ~1 page/commit smaller). Not a heed/LMDB *behaviour* divergence — filed here only for the format paper trail. | APPROVED | Quentin, 2026-10-05 (ADR-0022 ratification) | ### M1.13 adapter-boundary re-impositions (2026-07-17) diff --git a/docs/SPEC/02-pages.md b/docs/SPEC/02-pages.md index 9bb5402..968e9d2 100644 --- a/docs/SPEC/02-pages.md +++ b/docs/SPEC/02-pages.md @@ -73,7 +73,7 @@ LMDB struct is transliterated. | `FILL_THRESHOLD_PERMILLE` | `250` | 25.0 % — below this a page is a merge/borrow candidate (SPEC 03). | | `MIN_KEYS_LEAF` | `1` | Min entries a non-root leaf may hold after delete. | | `MIN_KEYS_BRANCH` | `2` | Min children a non-root branch may hold. | -| `META_CONTENT_LEN` | `168` | Number of leading bytes of a meta page covered by its CRC (see §3). | +| `META_CONTENT_LEN` | `172` | Number of leading bytes of a meta page covered by its CRC, before the annex ids (see §3/§3.3; format v2, ADR-0022). | **Page-type flags** (`u16`, in the common header `flags` field; a page has exactly one of the first four structural bits set): @@ -177,8 +177,10 @@ body** begins at offset 32: | 64 | 8 | `txnid` | u64 | Commit txnid (== header `txnid`). | | 72 | 48 | `free_db` | DBRecord | Root/stats of the free (GC) DB — `FREE_DBI`. §3.1. | | 120 | 48 | `main_db` | DBRecord | Root/stats of the main/catalog DB — `MAIN_DBI`. §3.1. | -| 168 | 4 | `meta_crc` | u32 | **Mandatory CRC32C** over bytes `[0, 168)` of this page (see §3.3). | -| 172 | psize−172 | reserved | — | MUST be 0; **not** covered by the CRC. | +| 168 | 4 | `fl_count` | u32 | **Free-list annex** id count (ADR-0022, format v2): `0 ≤ fl_count ≤ (psize − 176) / 8`. See the annex row below and SPEC 05 §2a. | +| 172 | 4 | `meta_crc` | u32 | **Mandatory CRC32C** over bytes `[0, 172) ∪ [176, 176 + 8·fl_count)` of this page (see §3.3). | +| 176 | 8·fl_count | `fl_ids` | u64[] | The annex: page ids freed by **this meta's txnid** (the PIL format v1 stored in the GC tree under `BE(txnid)`), little-endian, strictly ascending, unique, each in `[FIRST_DATA_PGNO, last_pg]`. Semantics are owned by SPEC 05 §2a (GC-29..33). | +| 176+8·fl_count | to psize | reserved | — | MUST be 0; **not** covered by the CRC. | > **SPEC 04 interface note.** This layout fixes the fields the commit pipeline > reads/writes: `txnid`, `last_pg`, `map_size`, `free_db.root`, `main_db.root`. @@ -230,8 +232,18 @@ At env open the engine reads both slots and validates each independently: 4. Header `txnid` (offset 8) equals body `txnid` (offset 64), else the slot is inconsistent and is **discarded** (INV-2). This guards a torn write that updated one copy but not the other. -5. `meta_crc` matches the recomputed CRC32C over `[0, 168)` (§3.3). A slot that - fails the CRC is **torn** and is discarded. +5. `fl_count ≤ (psize − 176) / 8` (the annex fits the page; checked **before** + the CRC so a hostile count cannot drive an out-of-bounds CRC read). Here + `psize` is the **env's expected page size** (the one validation is called + with), not the slot's own `page_size` field — using the expected size is + conservative even when a hostile slot claims a larger `page_size`, since the + CRC read stays within the buffer the env actually mapped. Else + the slot is discarded; then `meta_crc` matches the recomputed CRC32C over + `[0, 172) ∪ [176, 176 + 8·fl_count)` (§3.3). A slot that fails either is + **torn** and is discarded. (The annex *ids*' ordering/range are validated + where they are consumed — the writer's parse and the checker, SPEC 05 + GC-33 — not here; a CRC-valid slot with hostile ids must fail typed at + draw, exactly like a hostile tree PIL.) This numbered list is the **single owner** of the meta-slot validation predicate; SPEC 06 REC-1 references it rather than restating it. @@ -273,14 +285,21 @@ Selection among the *CRC-valid* slots: ### §3.3 — CRC32C coverage (exact byte range) `meta_crc` = CRC32C (Castagnoli, polynomial `0x1EDC6F41`, reflected input/ -output, init `0xFFFFFFFF`, final XOR `0xFFFFFFFF`) computed over **exactly the -first `META_CONTENT_LEN = 168` bytes of the meta page** — absolute offsets -`[0, 168)`. That range covers the common header (including the reserved bytes -18–31, which MUST be 0) and the meta body through the end of `main_db`. The -`meta_crc` field itself (offset 168) and the reserved tail (`[172, psize)`) are -**excluded**. Reserved bytes inside the covered range MUST be zero so the CRC is -deterministic. Implementation of CRC32C is ADR-0002 §D1 (software table in -Phase 1; ARMv8 `crc32c` instructions in Phase 3.9). +output, init `0xFFFFFFFF`, final XOR `0xFFFFFFFF`) computed over **the first +`META_CONTENT_LEN = 172` bytes of the meta page followed by the annex ids** — +absolute offsets `[0, 172) ∪ [176, 176 + 8·fl_count)`, as one CRC stream in +that byte order (ADR-0022, format v2; v1 covered `[0, 168)`). The covered +prefix spans the common header (including the reserved bytes 18–31, which +MUST be 0), the meta body through `main_db`, and `fl_count`; the `meta_crc` +field itself (offset 172) and the reserved tail past the annex are +**excluded**. Covering `fl_count` inside the main CRC is load-bearing: a torn +meta write cannot silently zero the annex (which would leak its free pages) +without tearing the slot as a whole — the torn-meta guarantee (REC-8) extends +to the annex. Validation reads `fl_count` *before* the CRC is checked, so it +is bounds-checked first (`fl_count ≤ (psize − 176) / 8`, else the slot is +invalid) to bound the CRC read. Reserved bytes inside the covered range MUST +be zero so the CRC is deterministic. Implementation of CRC32C is ADR-0002 §D1 +(software table in Phase 1; ARMv8 `crc32c` instructions in Phase 3.9). ### §3.4 — Env-creation protocol (both slots initialised, empty DB) @@ -291,7 +310,8 @@ Creating a new env writes **both** meta slots (pages 0 and 1) as valid, identica 2. Write slot 0 and slot 1 identically: `txnid = 0`, `magic`, `format_version`, `page_size`, `map_size`, `env_flags = 0`, `last_pg = 1` (pages 0 and 1 exist; no data page yet), `free_db` and `main_db` both **empty** - (`root = PGNO_INVALID`, all stats 0, `depth = 0`), fresh `meta_crc` on each. + (`root = PGNO_INVALID`, all stats 0, `depth = 0`), `fl_count = 0` (empty + free-list annex), fresh `meta_crc` on each. 3. `fsync(data)` so both slots are durable. 4. `fsync(parent dir)` so the data file's **directory entry** is durable (added 2026-07-22, issue #46): without it, a crash shortly after creation @@ -320,7 +340,7 @@ off 18 : 00 00 reserved0 off 20 : 00 00 00 00 checksum (data-page CRC; unused, 0) off 24 : 00 00 00 00 00 00 00 00 reserved (meta variant tail) off 32 : 5A 44 42 31 magic "ZDB1" -off 36 : 01 00 00 00 format_version = 1 +off 36 : 02 00 00 00 format_version = 2 off 40 : 00 10 00 00 page_size = 4096 (0x1000) off 44 : 00 00 00 00 env_flags = 0 off 48 : 00 00 10 00 00 00 00 00 map_size = 1 MiB (0x100000) @@ -336,8 +356,10 @@ off 152 : 00 00 00 00 00 00 00 00 main_db.entries = 0 off 160 : 00 00 main_db.depth = 0 (empty) off 162 : 00 00 main_db.flags = 0 off 164 : 00 00 00 00 main_db.leaf2_ksize = 0 -off 168 : meta_crc = CRC32C over bytes [0,168) -off 172 .. 4096 : 00 reserved tail (excluded from CRC) +off 168 : 00 00 00 00 fl_count = 0 (empty free-list annex) +off 172 : meta_crc = CRC32C over bytes [0,172) + (fl_count = 0: no annex ids follow) +off 176 .. 4096 : 00 reserved tail (excluded from CRC) ``` If the first commit (txn 1) then inserts one key, it allocates the root leaf at @@ -625,8 +647,11 @@ owns; the *page* format is exactly §2/§4): overflow run exactly like any large value (§5); there is no bespoke spill format. The in-tree ordering of the ids within the list is SPEC 05's choice. -The check tool treats GC entries as the authority for "free" in the -reachability-xor-freeness invariant (INV-10, INV-14). +Since format v2 (ADR-0022) the **newest** freeing-txn's PIL normally rides in +the live meta's free-list annex (§3, SPEC 05 §2a) instead of this tree; the +tree holds the spill/cold entries. The check tool treats GC entries **plus the +selected meta's annex ids** as the authority for "free" in the +reachability-xor-freeness invariant (INV-10, INV-14, INV-28). --- diff --git a/docs/SPEC/04-txn-mvcc.md b/docs/SPEC/04-txn-mvcc.md index 241bad1..f4035df 100644 --- a/docs/SPEC/04-txn-mvcc.md +++ b/docs/SPEC/04-txn-mvcc.md @@ -1004,10 +1004,10 @@ every hook; this section defines the steps and their ordering. Both write modes |---|------|-----------|----------------------------------| | C0 | Assert `child_count == 0`; compute freed-page list. | — | last committed meta `N−1` (unchanged). | | C1a | **catalog write-back** (M1.6, SPEC 02 §6.1): write each dirty named-DB working record back into the main tree as its `F_SUBDATA` catalog entry. Before `freelist_save` (matches LMDB's sub-DB flush order) so the pages this COWs/frees are captured by C1. | — | `N−1` (all changes still in the dirty set). | - | C1 | **freelist_save** (SPEC 05 §4): write this txn's freed pages into the GC DB, dirtying GC pages into the dirty set (loop-until-stable, SPEC 05 GC-11). | H0 | `N−1` (all changes still in dirty set, nothing written). | + | C1 | **freelist_save** (SPEC 05 §4): apply drain rewrites to the GC DB (dirtying GC pages into the dirty set, loop-until-stable, SPEC 05 GC-11) and place this txn's freed set — into the meta free-list annex when it fits (format v2, SPEC 05 §2a GC-32: stashed for C4, no tree write), else into the GC DB under `BE(N)`. | H0 | `N−1` (all changes still in dirty set, nothing written). | | C2 | **write dirty pages** to their pgnos (`pwrite` each dirty frame; memcpy-into-map under heap-staged WRITE_MAP, TXN-45a; a no-op per in-map frame under in-place WRITE_MAP, TXN-45b — those bytes were stored at allocation/edit time, under the same TXN-62 invariant). Not yet durable. | **H1** | `N−1` live; new pages sit in free/beyond-HWM slots the `N−1` tree does not reference (TXN-62). Partial/torn data pages are unreferenced garbage. | | C3 | **fsync(data)** — flush all data pages (skipped under `NO_SYNC`/`MAP_ASYNC`, SPEC 06 REC-6). | **H2** | `N−1` live; txn `N`'s data fully durable but unreferenced (no meta points at it). | - | C4 | **write meta** to slot `N&1` (the *older* slot), with `txnid = writer_txnid` and a fresh CRC (SPEC 02 §3.3). Not yet durable. | **H3** | Either `N−1` (meta `N` not yet reached disk, or reached but torn → CRC rejects it → older slot wins) or `N` (meta reached disk intact). Never a torn meta accepted. | + | C4 | **write meta** to slot `N&1` (the *older* slot), with `txnid = writer_txnid`, the C1-stashed free-list annex (format v2, SPEC 02 §3) and a fresh CRC covering both (SPEC 02 §3.3). Not yet durable. | **H3** | Either `N−1` (meta `N` not yet reached disk, or reached but torn → CRC rejects it → older slot wins) or `N` (meta reached disk intact). Never a torn meta accepted. | | C5 | **fsync(meta)** — flush the meta page (skipped under `NO_META_SYNC`/`NO_SYNC`: the meta is written but not fsynced this commit, SPEC 06 REC-6/REC-7/REC-9). | **H4** | txn `N` durable and live. | | C6 | Publish the new snapshot (TXN-18/TXN-19): swap the `Arc`, then `commit_point.store(writer_txnid, SeqCst)`; update `last_committed_txnid`; release the write mutex. | — | commit complete; new readers see `N`. | diff --git a/docs/SPEC/05-gc.md b/docs/SPEC/05-gc.md index 4f01bfb..072a855 100644 --- a/docs/SPEC/05-gc.md +++ b/docs/SPEC/05-gc.md @@ -18,9 +18,10 @@ documented as-is (§8) — the redesign is Phase 3.1 (ADR-first), **not** here. Clean-room note: the fork's `mdb_page_alloc`, `mdb_find_oldest`, `mdb_freelist_save`, and the loose-page list were read to understand the *algorithm* (CLAUDE.md rule 4); the design below is ZeroDB's own. Normative rules -are a single clean **GC-1 … GC-28** namespace (contiguous ids, no split-id -suffixes; GC-28 is the extended-page durability rule, placed with the allocation -topic in §5 though it is the highest id — see §10); new check-tool invariants +are a single clean **GC-1 … GC-33** namespace (contiguous ids, no split-id +suffixes; GC-28 is the extended-page durability rule, placed with the +allocation topic in §5 out of id order, and GC-29..33 are the format-v2 meta +free-list annex rules of §2a, ADR-0022 — see §10); new check-tool invariants continue SPEC 03's numbering from **INV-22**. This document **answers ADR-0002 open questions OQ1 and OQ2** (GC-2, GC-15, and @@ -96,6 +97,97 @@ writer_txnid − 1)`. --- +## §2a — The meta free-list annex (format v2, ADR-0022; GC-29..GC-33) + +**[Added 2026-10-03, ADR-0022 — implemented as a spike on relayed maintainer +approval; awaiting direct ratification.]** Since FORMAT_VERSION 2 the +**newest** freeing-txn's PIL normally lives in the committing meta page +itself (`fl_count`/`fl_ids`, SPEC 02 §3), not in the GC tree. The annex is a +pure *placement* change: its ids are exactly the PIL that v1 stored under +`BE(meta.txnid)`, with identical ordering, validation and gate rules. The GC +tree remains and holds the spill/cold entries; everything in §§4–6 continues +to govern it unchanged. + +- **GC-29 — annex = hoisted own-entry.** The annex of meta `N` is the PIL of + pages freed by txn `N` (freeing-txnid = the meta's own txnid; no separate + field): strictly ascending, unique, each in `[FIRST_DATA_PGNO, last_pg]` + (GC-3/GC-4 verbatim). `fl_count = 0` ⇔ txn `N` contributed no entry. A meta + with `fl_count > 0` has **no** GC-tree entry keyed `BE(N)` (see GC-32). +- **GC-30 — the carry (anti-leak across slots) + the in-save pool.** Meta + `N`'s slot is overwritten by txn `N + 2`, so the annex must never be the + only record of a free page once `N` stops being the base. Every + **committing** txn `W` (base `B = W − 1`) accounts for every live id of + `B`'s annex: ids drawn during the txn are in `reclaimed` (reused, or loose + if re-freed); the unconsumed remainder stays a live, gated **in-save + allocation pool** through `freelist_save` (a fourth GC-12 source, exactly + the drain-pool rule: its ids passed the gate with `F = B`), and the + step-(c) placement persists the **merge** `freed ∪ annex.live()` into + `W`'s annex or `W`'s tree entry — re-looping whenever the step-(c) put + itself drew from the pool, so no persisted set ever lists a handed-out + page. *Why a pool and not an upfront fold:* ids folded into `freed` are + un-allocatable in-save (they are pending free-list content), so an + upfront fold starves the save's own GC-tree allocations into `extend` on + every commit while the carried surplus grows in lockstep — an unbounded + file ratchet, observed as a `churn_parity_general` boundedness failure + during the ADR-0022 spike. A txn that commits nothing (the + unchanged-commit short-circuit) writes no meta, so `B` stays the base and + its annex stays live. Consequence: the INV-22 partition — with the annex + counted as free (INV-28) — holds for **every** committed meta. The carry + re-tags the persisted remainder under `W` (freeing-txnid moves forward); + that only ever *delays* reclaim eligibility, so the gate proofs are + unaffected. GC-13 gains a fourth bounded quantity: the annex pool is + loaded once at begin and never refilled, and every in-save draw strictly + shrinks it. +- **GC-31 — annex draw gate.** The annex pool is drawable iff + `B ≤ oldest_reader()` (TXN-20/21 verbatim with `F = B`): a reader pinned at + `B` cannot reach pages freed *by* `B` (they left `B`'s trees), and the + crash-fallback meta for writer `W` is `B` itself, whose annex still lists + the drawn ids as free and whose trees do not reference them — the TXN-62 / + in-save pool-draw argument (GC-12) verbatim. That "fallback is `B` itself" + premise holds as stated only in the durable modes (SPEC 06 REC-6/REC-12). + Under the relaxed-durability modes (`NO_META_SYNC` / `NO_SYNC` / + `MAP_ASYNC`) the crash-fallback can be older than `B` — the pre-existing, + characterized REC-10/REC-11 window — and under in-place `WRITE_MAP` an + uncommitted or aborted txn's in-map writes can reach disk with no commit + step at all, which falls in that same characterized window; see SPEC 06 + for the full argument and the window's taxonomy. +- **GC-32 — all-or-nothing placement.** At `freelist_save`, if the final + deduped `freed` set fits the annex cap (`(psize − 176) / 8` ids) it is + **stashed for the meta encode** and GC-11 step (c) performs **no** tree + put; otherwise the whole set goes to the tree under `BE(W)` as in v1 and + the annex is empty. One freeing-txn's pages are never split between the + two. Once a save has taken the tree arm it keeps it for that save (a later + drain-rewrite shrinking `freed` must not flip the placement mid-loop). + The stash encodes no tree write, allocates nothing and frees nothing, so + the GC-11 fixed point loses its step-(c) feedback edge in the annex arm + and the GC-13 termination argument is unchanged (it only removes an edge). +- **GC-33 — never trusted.** Annex ids are on-disk data: `fl_count` is + bounds-checked and the ids are covered by the meta CRC at slot validation + (SPEC 02 §3.2 rule 5 / §3.3), and `validate_pil_ids` (ascending + range, + the GC-18 validation rule) runs when the writer parses the base annex — + before any id can be handed out. A violating annex fails the parse with + `MdbError::Invalid` (corrupt freelist), exactly like a hostile tree PIL. + +**Allocation integration** (amends GC-16): the annex pool is step **2b**, +tried after the GC-tree draw and before extend. `B` is the newest possible +freeing-txnid (every tree key is `≤ B`, and `= B` only when `B` spilled, in +which case its annex is empty — GC-32), so tree-first keeps the GC-18/19 +oldest-first scan order and single-page determinism. Within the pool, draws +are GC-19-deterministic: smallest id for `n = 1`, first contiguous run for +`n > 1`. Drawn ids enter `reclaimed`. Inside `freelist_save` (GcSave mode) +the pool remains drawable as GC-12 source (4) — after the loose list and +the drain pool, before extend — under the same gate; GC-30 owns the +re-accounting that keeps the persisted set draw-free. + +**Steady-state effect** (the ADR-0022 motivation): with no parked readers and +small per-commit frees, the GC tree is empty, `freelist_save` performs zero +tree operations (no GC-leaf COW, no descent, no extra C2 page), and the next +txn's reclaim is an O(1) pool draw — the B12 free-list-save and GC-scan costs +collapse to the annex encode (~8 bytes/id inside the meta CRC/write that +every commit already pays). + +--- + ## §3 — Loose pages (allocated and freed within one txn) - **GC-7** — A page this txn both **dirtied/allocated and then freed** (e.g. a leaf @@ -140,9 +232,16 @@ in-loop allocator so the loop cannot leak GC entries. condition.]** This skeleton is normative — an implementation must be reconstructible from it alone: + **[Amended 2026-10-03, ADR-0022/GC-30/GC-32: the base annex remainder + stays a live in-save pool (GC-12 source 4); step (c) persists the merge + `freed ∪ annex.live()` — to the meta annex (no tree put) whenever the + merged set fits the cap (all-or-nothing, sticky per save), else to the + tree as in v1 — and re-loops if the put drew from the pool.]** + ``` freelist_save(txn): alloc_mode = GC_SAVE # GC-12 scope: whole procedure + # base_annex.remainder stays drawable (GC-30 pool); merged in step (c) pending = keys(txn.drains) # every ops-drained entry needs # its GC-20 rewrite at least once # (b) trailing-loose shrink FIRST (GC-10): release never-written @@ -167,14 +266,28 @@ in-loop allocator so the loop cannot leak GC entries. if remaining.is_empty(): drains.remove(F); delete(FREE_DBI, BE(F)) else: put(FREE_DBI, BE(F), encode(remaining)) - # (c) this txn's own entry under key = BE(writer_txnid). + # (c) this txn's own entry: the merge freed ∪ annex.live() goes to + # the meta annex when it fits (GC-32), else to the tree under + # key = BE(writer_txnid) as in v1. sort_dedup(freed_pgs) # GC-3/GC-4 before = freed_pgs.len() - if before > 0: - put(FREE_DBI, BE(writer_txnid), encode(freed_pgs), RESERVE) # GC-2 - if freed_pgs.len() != before: - continue # the put COW-freed committed - # GC pages: PIL must grow + annex_before = annex.live().len() + if before + annex_before > 0: + if not tree_arm and before + annex_before <= annex_cap: + annex_mode = true # merged after the loop; no + # tree write, cannot free or + # allocate anything + else: + tree_arm = true # sticky for this save (GC-32) + annex_mode = false + ids = merge(freed_pgs, annex.live()) + put(FREE_DBI, BE(writer_txnid), encode(ids), RESERVE) # GC-2 + if freed_pgs.len() != before or annex.live().len() != annex_before: + continue # the put COW-freed committed + # GC pages (PIL must grow) or + # drew from the annex pool + # (written PIL lists a + # handed-out page) # FIXED POINT = all three, together: # 1. no pending rewrite (no drain-pool draw escaped a re-rewrite) # 2. freed_pgs stable across the put (the write freed nothing new) @@ -293,6 +406,15 @@ in-loop allocator so the loop cannot leak GC entries. if let Some(run) = gc_reclaim(n, oldest_reader()): return run + # 2b. The base meta's free-list annex pool (format v2, §2a GC-31): + # gated by base_txnid <= oldest_reader(); GC-19-deterministic + # within the pool. Tried AFTER the tree: the annex's freeing-txn + # is the newest possible, so tree-first preserves oldest-first. + # (Inside freelist_save the pool stays drawable as GC-12 source + # (4), after loose and the drain pool — GC-30.) + if let Some(run) = annex_pool_draw(n, oldest_reader()): + return run + # 3. Extend the file: hand out [next_pgno .. next_pgno+n), then # next_pgno += n. Fails MapFull per GC-17. return extend(n) @@ -420,6 +542,21 @@ in-loop allocator so the loop cannot leak GC entries. total += pil.count # the u64 count prefix (GC-3) return total ``` + Since format v2 the snapshot's meta annex ids are free pages too: + `free_page_count` adds the snapshot's `fl_count` (carried on the published + snapshot, so no meta re-read) to the tree walk's total. + + *Writer working-state note (format v2).* Called on the live write txn before + its `freelist_save`, the two free-page sources report at **different + granularities**: the base annex contributes its *live remainder* (ids already + drawn this txn are excluded immediately, since a draw removes them from the + in-save pool), while the GC tree still contributes each drained entry's full + `count` prefix until the save rewrites that entry. The value is therefore a + conservative working-state estimate, not the post-commit free count; the exact + per-snapshot value is the read-txn definition above, evaluated against a + committed meta. No consumer reads `free_page_count` mid-write-txn (Meilisearch + uses `non_free_pages_size` under a read snapshot, GC-24); the drift is + documented, not relied upon. It runs under an ordinary read snapshot (SPEC 04 §3), counting the **sum of PIL counts**, not the number of GC entries. Each entry contributes only its 8-byte `count` prefix (`pil_count`): the walk reads the prefix and checks it against the @@ -492,6 +629,12 @@ The check tool (`zerodb-tools check`, M1.12) treats the GC DB as the authority f exceeds the high-water. - **INV-26** — **PIL sorted & counted.** Each PIL's `count` prefix equals the number of ids that follow, and the ids are strictly ascending (GC-3/GC-4). +- **INV-28** — **Annex accounting (format v2, ADR-0022).** The selected + meta's annex ids join the free set for INV-22/24/25 (reachable XOR free, + listed once across annex ∪ all tree PILs, in range), the annex is strictly + ascending and `fl_count`-consistent (the INV-26 analogue; the length half + is the SPEC 02 §3.2 rule-5 bound), and `fl_count > 0` implies no GC-tree + entry keyed `BE(meta.txnid)` (GC-29/GC-32). - **INV-27** — **`non_free_pages_size` consistency** (amended 2026-09-29). Without `WRITE_MAP`, the pages of a committed image partition exactly into the two meta slots, user-tree pages, GC-tree pages and free pages, so @@ -517,9 +660,11 @@ The check tool (`zerodb-tools check`, M1.12) treats the GC DB as the authority f | coalescing / fragmentation | GC-21 | PLAN 1.5 tolerance | | disk-usage APIs | GC-22..24, INV-27 | SPEC 00 rows 18/19 | | huge-txn weakness | GC-25..27 | PLAN 3.1 (ADR-first) | -| check invariants | INV-22..27 | SPEC 03 §11 | +| meta free-list annex (format v2) | GC-29..33, INV-28 | SPEC 02 §3/§3.3, ADR-0022 | +| check invariants | INV-22..28 | SPEC 03 §11 | -**Rule count: GC-1 … GC-28 (28 normative rules, contiguous ids, no split-id -suffixes) + INV-22 … INV-27 (6 new check invariants).** OQ1 answered by **GC-2** +**Rule count: GC-1 … GC-33 (33 normative rules, contiguous ids, no split-id +suffixes; GC-29..33 are the format-v2 annex rules of §2a, ADR-0022) + +INV-22 … INV-28 (7 check invariants).** OQ1 answered by **GC-2** (big-endian GC keys); OQ2 answered by **GC-15** (track `next_pgno`, persist `last_pg`; no meta-field change). Both recorded in the ADR-0002 amendment. diff --git a/docs/SPEC/06-recovery.md b/docs/SPEC/06-recovery.md index bda138e..7b69039 100644 --- a/docs/SPEC/06-recovery.md +++ b/docs/SPEC/06-recovery.md @@ -29,9 +29,11 @@ this section defines the **recovery decision** and its error taxonomy. - **REC-1** — **Validate each slot independently.** The validation predicate is owned by **SPEC 02 §3.2** (the numbered list: `magic`, `format_version`, - `page_size` ∈ {power of two, 4096–65536}, header `txnid` == body `txnid`, and — - mandatory — `meta_crc` over `[0,168)`, SPEC 02 §3.3). REC-1 does **not** restate - the list; it references SPEC 02 §3.2 as the single owner. A slot failing any + `page_size` ∈ {power of two, 4096–65536}, header `txnid` == body `txnid`, the + annex-count bound, and — mandatory — `meta_crc` over + `[0,172) ∪ [176, 176 + 8·fl_count)`, SPEC 02 §3.3 as amended by ADR-0022). + REC-1 does **not** restate the list; it references SPEC 02 §3.2 as the + single owner. A slot failing any check is **invalid** (torn or foreign) and is discarded from selection. - **REC-1a** — **Geometry validation of the selected slot** (added 2026-09-09, security review H1; predicate owned by SPEC 02 §3.2 step 6). A CRC-valid slot @@ -182,6 +184,19 @@ open, REC-1/REC-2 select snapshot `X` and the check tool (SPEC 03 §11 + SPEC 05 Phase 3.9); it is **not** claimed to detect every torn meta — sector-aligned tears are caught by the double-buffer/txnid design, not the CRC. + **Format-v2 annex note (ADR-0022).** The CRC-covered region is + `[0,172) ∪ [176, 176 + 8·fl_count)` and may extend past the first 512-byte + sector when the free-list annex is large (`fl_count > 42` at 512-byte + sectors). The sub-sector argument is unchanged; additionally, a + **sector-aligned** tear that lands *inside* the covered annex now leaves the + CRC inconsistent and the slot is rejected — which is required, not + incidental: a half-written (or silently zeroed) annex would leak that + commit's freed pages while the rest of the slot looked complete. The + "complete old or complete new, both CRC-valid" outcome of a sector-aligned + tear therefore holds only when the covered region fits sector 0; otherwise + the tear degrades to a rejected slot and the intact other slot wins — + strictly more conservative, never less. + --- ## §3 — Durability flags: crash windows (SPEC 01 §S6 lattice) diff --git a/docs/adr/0022-meta-freelist-annex.md b/docs/adr/0022-meta-freelist-annex.md new file mode 100644 index 0000000..8d3e03f --- /dev/null +++ b/docs/adr/0022-meta-freelist-annex.md @@ -0,0 +1,215 @@ +# ADR-0022: Meta free-list annex — the per-commit freed PIL rides in the meta page + +- Status: **Accepted (2026-10-05, Quentin)** — format change (`FORMAT_VERSION` + 1→2) directly ratified per CLAUDE.md rule 6, together with the crash-harness + seed re-pin (198). Implemented first as a spike on approval relayed + 2026-10-02/03; re-gated on main after ADR-0021 (in-place WRITE_MAP) merged, + with WRITE_MAP in-place twins of the annex battery. +- Result (2026-10-03): **measured win, gate green, spec-review clean.** Bench + server 3-column A/B (x86-64, turbo off, 5 interleaved rounds, CODEGEN_UNITS=1, + BASE = main 84582e8): `commit/batch/n1` 2.04× → **1.60×** LMDB (−22%, + after÷before 0.780), `commit/batch/n100` 1.09× → 1.02×; `n10k` and both sync + rungs flat (overhead amortized / fsync-bound — as predicted). Full gate: test + 630/0, miri 0-fail, crash-test-quick all durability modes, loom 9/0, stress + 180s 2/0, fuzz-quick 2.1M clean. Adversarial spec-review found no technical + blockers; its should-fix/nit items are addressed in the branch. +- Real-case check (2026-10-05, three columns LMDB / before `8b066a6` / after, + `benches/results/2026-10-05-meta-annex-real-case.md`): YCSB on Graviton4 NVMe + durable B **1.078×** before (258k vs 239k ops/s, now 1.08× LMDB), no-sync + 0.98–1.04×; on the x86 bench server no-sync **1.03–1.07×**, durable flat + (SATA-flush-bound). Write p50 −14% to −25% in every configuration on both + machines. Meilisearch: no regression — movies indexing 1.00×, hackernews + incremental additions 0.98× (commit span 0.94×), search 0.97×. Scope of the + claim: frequent or durable commits; not a Meilisearch bulk-indexing speed-up. +- Milestone: perf track "non-copy per-commit CPU" lever #1 (cheaper free-list + save), PERF-GAP-VS-LMDB §B12; forward-looking toward PLAN 3.1. +- Date: 2026-10-03 +- Numbering note: ADR-0021 (WRITE_MAP in-place) lives on its own open branch; + this ADR takes 0022 to avoid colliding with it. + +## Context + +B12 (bench server, x86-64, 4 KiB, 20k single-put NO_SYNC commits) attributes +the largest single slice of ZeroDB's +3.1 µs commit-CPU gap vs the LMDB fork +to the free-list save: ~1.6 µs/commit. Local census re-measurement in this +worktree (macOS aarch64, 4 KiB DB pages, temporary sub-phase instrumentation) +confirms both the magnitude and the shape: + +- `freelist_save` total ≈ 1.2–1.5 µs/commit; +- per commit it executes ~2 `put_pil` + ~1 `delete_tree` on the GC tree + (steady state: this txn's own entry, one partial-drain GC-20 rewrite, one + fully-drained entry delete); +- of that, ~0.3 µs is one real COW touch of the GC leaf (frame + full-page + copy), ~0.8–0.9 µs is the generic tree-op machinery of the three ops, + ~0.1–0.2 µs is encode/sort/descent (descent is *not* the problem: the GC + tree is one leaf; `search_path` is ~90 ns total), and the GC-11 loop runs + exactly 1.00 iterations/commit; +- the save also dirties the GC leaf, so commit C2 writes one extra page per + commit (ZeroDB 5.93 pages vs LMDB 4.96 in B12). + +The format-stable ceiling is low. LMDB's own save does the same three +B-tree ops on the same flat txnid→PIL format and costs ~0.5–0.7 µs; matching +it means shaving generic per-op overhead (separate levers: touch cost B12-1, +lazy validation A2(b)) and buys at most ~0.5 µs here. The three tree ops per +commit are *inherent to the format*: the freed list of txn N must be durable +and crash-consistent with meta N, and in the flat-freelist format the only +place for it is the GC tree, whose mutation is itself a COW tree write. + +With the format constraint lifted, the structural observation is: **the meta +page already is the one page every commit must write and CRC**. At 4 KiB the +meta's fixed content ends at byte 172; the remaining ~3.9 KiB are reserved +zeros. A typical commit frees a handful of pages (~5 in the census; a PIL of +c ids is 8c bytes). The freed list of txn N can therefore ride *inside* meta +N — written by the same `write_at_page`, covered by the same CRC, atomic with +the same meta — and the GC tree drops out of the per-commit path entirely. + +Prior art: LMDB has no equivalent (its meta is ~112 bytes of a page and it +keeps the freelist in FREE_DBI unconditionally — and pays the same tree-op +cost we do). libmdbx likewise keeps a GC table but batches differently. This +is a deliberate, documented **non-LMDB lever** (observable heed-level +semantics unchanged; allocation-order determinism preserved), in the same +spirit as the ratified rightmost-leaf finger (ADR-0015 note in rwtxn.rs): +LMDB's own technique for this path cannot go below LMDB's own cost, and the +B12 target is to *beat* that bucket, with PLAN 3.1 already pointing at a +freelist redesign. + +## Options + +### Option A — format-stable micro-optimisation (rejected) + +One search+touch for the whole save (all steady-state keys land in the same +GC leaf), direct `LeafMut` edits, kill the two per-save `Vec` clones +(`freed.clone()`, `remaining.to_vec()`). + +- Pros: no format change, no recovery surface. +- Cons: measured ceiling ~0.3–0.5 µs of the 1.6; duplicates leaf-edit logic in + the riskiest file; still COWs + rewrites a GC leaf and still writes the + extra page at C2; becomes dead code when 3.1 lands. + +### Option B — freed-PIL annex in the meta page (chosen) + +Meta format (FORMAT_VERSION 1 → 2; offsets `[0, 168)` unchanged): + +| Off | Size | Name | Meaning | +|----:|-----:|------|---------| +| 168 | 4 | `fl_count` (u32) | Number of annex ids. `0 ≤ fl_count ≤ (psize − 176) / 8`. | +| 172 | 4 | `meta_crc` (u32) | CRC32C over `[0, 172) ∪ [176, 176 + 8·fl_count)`. | +| 176 | 8·c | `fl_ids` | LE u64 page ids, strictly ascending, each in `[FIRST_DATA_PGNO, last_pg]`. | +| rest | — | reserved | MUST be 0; not CRC-covered. | + +Semantics: the annex is **exactly the PIL that version 1 stored in the GC +tree under key `BE(meta.txnid)`** — the pages freed by that commit — hoisted +into the meta. Freeing-txnid = the meta's own txnid; no separate field. + +- **Write side (commit C1).** `freelist_save` keeps the *unconsumed + remainder* of the base meta's annex as a live, gated **in-save pool** (a + fourth GC-12 source, exactly the drain-pool rule — an upfront fold into + `freed` was the spike's first cut and ratcheted the file unboundedly, + because folded ids are un-allocatable in-save and the save's own GC-tree + ops then extend on every commit; caught by `churn_parity_general`). The + GC-11 loop for drain rewrites (GC-20) runs unchanged. Step (c) persists + the **merge** `freed ∪ annex.live()`: if it fits the annex cap it is + **stashed for the meta encode** and no tree put happens (the stash cannot + allocate or free pages, so the GC-11/13 fixed point is reached with no + step-(c) feedback); otherwise the whole merged list goes to the tree under + `BE(writer_txnid)` exactly as today and the annex is empty — re-looping if + the put itself drew from the pool, so the persisted set never lists a + handed-out page. All-or-nothing — one freeing-txn's pages are never split + between annex and tree. Once the tree arm has run in a save, it stays the + arm for that save (no flip-flop if a drain rewrite later shrinks `freed`). +- **Read side (allocation).** The writer parses the base meta's annex at txn + begin (one u32 read when empty; `validate_pil_ids` on the ids otherwise — + the freelist is still never trusted, GC-18). `allocate()` draw order + becomes: loose → GC tree (`gc_reclaim`, unchanged) → **annex pool** → + extend. The annex's freeing-txn (`base.txnid`) is the *newest* possible F, + so trying the tree first preserves the GC-18/19 oldest-first determinism. + The annex draw is gated exactly like any GC draw: `F ≤ oldest_reader()` + (here: no live reader pinned below base). Drawn ids enter `reclaimed`. + Inside `freelist_save` (GcSave mode) the annex pool **is** consulted, as + the carried pool (GC-30): after the drain pool (`save_pool_draw`) is + exhausted, `allocate()`'s GcSave arm falls through to the same + `annex_draw`, gated the same way. This is safe only because the Write-side + step (c) placement re-reads `annex.live()` after every tree put before + deciding what to persist, so a draw here can never leave a handed-out id in + the persisted set — the GC-12-amended anti-leak argument applies to this + draw exactly as it does to `save_pool_draw` (`crates/zerodb-core/src/ + rwtxn.rs`, the `AllocMode::GcSave` arm of `allocate`, ~line 1277). +- **Crash safety.** The annex is CRC-covered by the meta it belongs to: a + torn meta write that corrupts the annex tears the whole slot, which the + §3.2 selection discards (fallback to N−1, whose own annex+tree state is + per-snapshot consistent). The TXN-62 argument for reusing annex ids is the + in-save pool-draw argument verbatim: an id in meta B's annex was freed *by* + txn B, is absent from B's trees, and B is the crash-fallback meta for the + writer W = B+1 — so overwriting it before W's meta lands is safe, and any + reader that pins after the draw pins `≥ B`. + +- Pros: steady-state save does **zero tree ops** (no GC-leaf COW, no descent, + no node edits, no extra C2 page); next txn's reclaim is an O(1) pool draw + instead of a cursor scan + PIL decode; CRC/encode cost grows by 8 bytes/id + (~40 B/commit typical). +- Cons: format break (version bump; pre-release, no consumers — approved); + recovery/check/tools must learn the annex; a new carry invariant to hold; + under a parked reader the carried list is re-encoded into every meta until + it exceeds the cap and spills (bounded by the cap, ≤ ~3.9 KiB at 4 KiB psize). + +### Option C — O_DSYNC-style / batching tricks + +Not applicable: the cost is CPU in tree ops, not flush count (B12 is NO_SYNC). + +## Decision + +Option B, as a **spike** (CLAUDE.md rule 7: format change ⇒ smallest change +that tests the riskiest assumption — here, that the annex removes the +measured CPU while recovery, INV-22 and bounded growth stay provable). +Always-on with the FORMAT_VERSION bump (an opt-in annex would double the +format test matrix for no consumer benefit; "parity by default" governs +API-visible behavior, which is unchanged). + +## Invariants (normative; SPEC 05 §2a / SPEC 02 §3 own the final wording) + +- **GC-29 (annex = hoisted own-entry).** The annex of meta N holds exactly + the ids version 1 would have written under `BE(N)`, strictly ascending, + unique, in `[FIRST_DATA_PGNO, last_pg]`; `fl_count = 0` means no entry. + A meta with `fl_count > 0` has **no** GC-tree entry keyed `BE(N)`. +- **GC-30 (carry + in-save pool).** Every committing txn W accounts for + every live id of its base's annex: drawn ids are in `reclaimed` (reused or + loose), and the remainder — drawable in-save as a gated pool — is merged + into the step-(c) persisted set (W's annex or W's tree entry). + Consequence: the per-image partition (INV-22, with the annex counted as + free — INV-28) holds for every committed meta. +- **GC-31 (gate).** Annex draws require `base.txnid ≤ oldest_reader()`; the + TXN-20/TXN-62 proofs apply unchanged with F = base.txnid. +- **GC-32 (all-or-nothing).** One freeing-txn's pages are never split between + its annex and a tree entry. +- **GC-33 (never trusted).** Annex ids pass `validate_pil_ids` before any id + is handed out, exactly like a tree PIL (GC-18 validation rule); `fl_count` + is bounds-checked and the ids CRC-checked at slot validation. +- **INV-28 (check).** The checker counts annex ids of the selected meta into + the free set: reachable XOR (tree-free ∪ annex-free), no double listing + (extends INV-22/24/25/26 to the annex). + +## Consequences + +- `FORMAT_VERSION` 2; version-1 files are rejected at open (`BadVersion`) — + pre-release, sanctioned. `migrate-from-lmdb` output moves to v2 implicitly + (it writes through the engine). +- SPEC 02 §3 (meta table, §3.3 CRC coverage, §3.5 example note), SPEC 05 + (new §2a + GC-16/18/23 amendments, GC-29..33, INV-28), SPEC 06 (REC-8 note: + the annex is inside the torn-meta guarantee) updated in this change. +- `Snapshot` carries the annex count so `free_page_count()` stays exact per + snapshot (GC-23 note). +- New tests: meta annex codec roundtrip + torn-annex rejection + cap bounds; + annex reclaim/carry/spill behavior (incl. reader-gate block); checker + annex accounting; crash harness runs unchanged on the new format. +- The census' `allocate` cost also drops in steady state (the GC cursor scan + disappears when the tree is empty). That is a *consequence* of this format + change, not the separate allocate lever being folded in. + +## Open questions for human review + +1. ~~Ratify the format change itself (CLAUDE.md rule 6).~~ **Resolved + 2026-10-05:** ratified directly by Quentin. +2. ~~Cap policy: full `(psize − 176)/8` (chosen) vs a smaller policy cap.~~ + **Resolved 2026-10-05:** the full cap is accepted with the ADR. +3. Whether PLAN 3.1 should absorb this as its first stage (the annex is the + hot tier of any future freelist redesign) or whether 3.1 supersedes it.