From 5860788a45545e71a6d64cce9d1ec10c66776fcd Mon Sep 17 00:00:00 2001 From: Quentin de Quelen Date: Sat, 3 Oct 2026 15:07:22 +0200 Subject: [PATCH 1/5] =?UTF-8?q?ADR-0022=20spike:=20meta=20free-list=20anne?= =?UTF-8?q?x=20=E2=80=94=20hoist=20the=20per-commit=20freed=20PIL=20into?= =?UTF-8?q?=20the=20meta=20page?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The B12 census attributes the largest single slice of ZeroDB's per-commit CPU gap vs the LMDB fork (~1.6 µs) to the free-list save: `freelist_save` executes ~2 `put_pil` + ~1 `delete_tree` on the GC tree every commit, each a COW tree write, and dirties one extra page written at C2. With the on-disk format constraint lifted (pre-release, no consumers), the freed list of txn N now rides inside meta N's reserved bytes — written by the same `write_at_page`, covered by the same CRC, atomic with the same meta — so the steady-state save does zero GC-tree ops and the next txn's reclaim is an O(1) pool draw instead of a cursor scan + PIL decode. FORMAT_VERSION 1 -> 2 (v1 files rejected at open, sanctioned pre-release). Invariants GC-29..33 + INV-28; crash safety via the meta CRC + torn-meta fallback; an in-save pool bounds file growth under a parked reader (churn_parity_general guards it). SPEC 02/05/06 updated; ADR-0022 added. Spike per CLAUDE.md rule 7 (format change => smallest change that tests the riskiest assumption). Awaiting direct human ratification of the format change before merge (rule 6). Gate green: test 630/0, miri 0-fail, crash-test-quick all modes, loom 9/0, stress 180s 2/0, fuzz-quick 2.1M runs clean. --- PROGRESS.md | 2 + crates/zerodb-core/src/builder.rs | 2 + crates/zerodb-core/src/check.rs | 52 ++- crates/zerodb-core/src/env.rs | 34 +- crates/zerodb-core/src/nested.rs | 3 + crates/zerodb-core/src/page/crc32c.rs | 16 + crates/zerodb-core/src/page/meta.rs | 104 +++++- crates/zerodb-core/src/page/mod.rs | 23 +- crates/zerodb-core/src/readers.rs | 2 + crates/zerodb-core/src/rotxn.rs | 14 +- crates/zerodb-core/src/rwtxn.rs | 176 ++++++++- crates/zerodb-core/tests/page_edges.rs | 26 +- crates/zerodb-core/tests/spec02_format.rs | 146 +++++++- .../tests/crash_harness_smoke.rs | 12 +- crates/zerodb-tools/tests/data_file_probe.rs | 2 +- crates/zerodb/src/copy.rs | 11 +- .../tests/loose_page_and_trailing_shrink.rs | 17 +- crates/zerodb/tests/meta_annex_gc.rs | 343 ++++++++++++++++++ docs/DECISIONS.md | 1 + docs/SPEC/02-pages.md | 59 ++- docs/SPEC/04-txn-mvcc.md | 4 +- docs/SPEC/05-gc.md | 150 +++++++- docs/SPEC/06-recovery.md | 21 +- docs/adr/0022-meta-freelist-annex.md | 194 ++++++++++ 24 files changed, 1303 insertions(+), 111 deletions(-) create mode 100644 crates/zerodb/tests/meta_annex_gc.rs create mode 100644 docs/adr/0022-meta-freelist-annex.md diff --git a/PROGRESS.md b/PROGRESS.md index d52ca59..da85bf8 100644 --- a/PROGRESS.md +++ b/PROGRESS.md @@ -153,3 +153,5 @@ ADR-0019 PARKED, 2026-10-01 — durable meta write via an O_DSYNC fd (2 → 1 fd ADR-0020 SPIKE, 2026-10-01 — 24-byte page-header spike approved for measurement only; the format change is not approved and awaits the spike's numbers. Branch `qdequele/zerodb-header24-spike`. ADR-0021 PRODUCTION PASS, 2026-10-02 — true in-place WRITE_MAP hardened on the spike branch (not merged): B1 unsafe-fn broker with the sole sanctioned call in zerodb-core::dirty; B2 superseded by a stronger M1 miri finding (under Stacked Borrows a cached whole-map reference is UB after any in-place write — the in-map writer now borrows its map view lazily per access, like readers; spill-time re-derivation kept for the heap-staged paths); B3 map-aware fault backend (brokered regions journaled, sealed per sync, image cuts open the real writable map; vacuousness tripwire + smoke pin); B4 loom justified as no-new-model (ADR-0021 §hardening), 180 s stress writemap-in-place variant, M1.13 erased-cursor audit + miri battery in heed-zerodb; M1 zerodb_io::testmap::TestWriteMap (test-backing feature) + writemap_in_place_miri battery; M2 release-mode TXN-62 typed guard in allocate; M3 abort-after-spill / nested-fanout / put_reserved batteries parameterized over WRITE_MAP. SPEC 04 TXN-71/C5a/TXN-45b and SPEC 06 REC-20 amended in the same change. + +ADR-0022 SPIKE, 2026-10-03 — meta free-list annex (format v2): the per-commit freed PIL rides inside the CRC-covered meta page; the GC tree is the spill/cold path (over-cap sets, reader-gated carries). Steady-state freelist_save drops from ~2 GC-tree puts + 1 delete + 1 leaf COW per commit to zero tree ops (~1.2–1.5 µs → ~0.1 µs local 4 KiB census; one fewer dirty page per commit), and churn steady-state files shrink ~12% vs v1. First-cut lesson: folding the carried annex into `freed` up front starved the save's own allocations into per-commit extends (unbounded ratchet, caught by churn_parity_general); the carried remainder must stay a gated in-save pool (GC-12 source 4) merged at placement. Pinned crash seed re-pinned (198) per its documented protocol. Implemented on coordinator-relayed approval; ADR-0022 awaits direct ratification before merge. Worktree branch, not merged. diff --git a/crates/zerodb-core/src/builder.rs b/crates/zerodb-core/src/builder.rs index 94222d1..fb5d363 100644 --- a/crates/zerodb-core/src/builder.rs +++ b/crates/zerodb-core/src/builder.rs @@ -355,6 +355,7 @@ fn finalize_image( last_pg, free_db: DBRecord::empty(), main_db, + fl_count: 0, }; m0.encode(&mut buf[0..psize as usize])?; let mut m1 = m0; @@ -894,6 +895,7 @@ impl EnvStream { last_pg, free_db: DBRecord::empty(), main_db, + fl_count: 0, }; meta.encode(&mut frame)?; sink.emit(0, &frame)?; diff --git a/crates/zerodb-core/src/check.rs b/crates/zerodb-core/src/check.rs index f8f8393..25bd510 100644 --- a/crates/zerodb-core/src/check.rs +++ b/crates/zerodb-core/src/check.rs @@ -39,9 +39,13 @@ struct Checker<'a> { /// otherwise choose ids that all collide (HashDoS). The engine's /// dirty-store hasher is unaffected — its keys are engine-authored. visited: HashSet, - /// Free page ids collected from every GC PIL → occurrence count - /// (INV-22/INV-24; SPEC 05 §9). Randomly seeded, see `visited`. + /// Free page ids collected from every GC PIL **and the selected meta's + /// free-list annex** → occurrence count (INV-22/INV-24/INV-28; SPEC 05 + /// §9/§2a). Randomly seeded, see `visited`. free: HashMap, + /// The selected meta's annex id count (INV-28/GC-29: `> 0` forbids a GC + /// tree entry keyed `BE(meta_txnid)`). + annex_count: usize, violations: Vec, } @@ -353,6 +357,15 @@ impl<'a> Checker<'a> { if txnid > self.meta_txnid { self.fail("INV-23", format!("GC entry keyed by future txn {txnid}")); } + // INV-28/GC-29 (format v2): a meta with a non-empty annex holds the + // newest freeing-txn's PIL itself — a tree entry under the same + // txnid would be the forbidden split placement (GC-32). + if txnid == self.meta_txnid && self.annex_count > 0 { + self.fail( + "INV-28", + format!("GC entry {txnid} coexists with a non-empty meta annex (GC-29/GC-32)"), + ); + } let pil: Vec = match val { LeafValue::Inline(v) => v.to_vec(), LeafValue::Overflow { head_pgno, dsize } => { @@ -522,9 +535,9 @@ pub fn check_image(bytes: &[u8], psize: u32) -> Vec { (Ok(a), Ok(b)) => (a, b), _ => return vec!["INV-1: meta slots undecodable".into()], }; - let meta = match select_meta(&v0, &v1, false) { - crate::page::MetaChoice::Both { meta, .. } - | crate::page::MetaChoice::OnlyOne { meta, .. } => meta, + let (meta, chosen) = match select_meta(&v0, &v1, false) { + crate::page::MetaChoice::Both { meta, chosen } + | crate::page::MetaChoice::OnlyOne { meta, chosen } => (meta, chosen), crate::page::MetaChoice::None => { return vec!["INV-2: no valid meta slot (both torn/foreign)".into()] } @@ -557,13 +570,34 @@ pub fn check_image(bytes: &[u8], psize: u32) -> Vec { return violations; } } + // Format v2 (ADR-0022): the selected meta's free-list annex ids are free + // pages (SPEC 05 §2a). Validate their GC-29 shape (the INV-25/26 + // analogues, labelled INV-28) and count them into the free set for the + // INV-22/24 partition. + let annex = MetaPage::read_annex(&bytes[chosen * ps..(chosen + 1) * ps], psize) + .expect("annex count was bounds-checked by MetaPage::validate"); + let mut free: HashMap = HashMap::new(); + let mut prev: Option = None; + for &id in &annex { + if id < FIRST_DATA_PGNO || id > meta.last_pg { + violations.push(format!("INV-28: annex free id {id} out of range")); + } + if let Some(p) = prev { + if p >= id { + violations.push("INV-28: annex ids not strictly ascending".into()); + } + } + prev = Some(id); + *free.entry(id).or_insert(0) += 1; + } let mut checker = Checker { bytes, psize, meta_txnid: meta.txnid, last_pg: meta.last_pg, visited: HashSet::new(), - free: HashMap::new(), + free, + annex_count: annex.len(), violations, }; checker.check_catalog_record("main_db", &meta.main_db); @@ -607,8 +641,10 @@ mod tests { let base = slot * PS as usize; let e_off = base + 120 + 32; // main_db record + entries offset img[e_off] = 99; - let crc = crate::page::crc32c(&img[base..base + 168]); - img[base + 168..base + 172].copy_from_slice(&crc.to_le_bytes()); + // Format v2 CRC: fl_count (offset 168) is 0 in builder images, + // so the coverage is exactly [0, 172); the CRC field sits at 172. + let crc = crate::page::crc32c(&img[base..base + 172]); + img[base + 172..base + 176].copy_from_slice(&crc.to_le_bytes()); } let v = check_image(&img, PS); assert!( diff --git a/crates/zerodb-core/src/env.rs b/crates/zerodb-core/src/env.rs index 82afe32..d4e4216 100644 --- a/crates/zerodb-core/src/env.rs +++ b/crates/zerodb-core/src/env.rs @@ -189,7 +189,7 @@ pub trait Backing: Send + Sync { /// state (SPEC 04 TXN-18). Readers `Arc`-clone the env's published snapshot at /// begin and never re-read a durable meta page (the slot a pinned txnid lived /// in is overwritten two commits later, TXN-63). -#[derive(Debug, Clone, Copy, PartialEq, Eq)] +#[derive(Debug, Clone, PartialEq, Eq)] pub struct Snapshot { /// The commit point of this snapshot. pub txnid: u64, @@ -199,17 +199,28 @@ pub struct Snapshot { pub main_db: DBRecord, /// Root/stats of the free (GC) DB. pub free_db: DBRecord, + /// This snapshot's meta free-list annex (SPEC 02 §3 format v2, ADR-0022; + /// SPEC 05 §2a): the pages freed by txn `txnid`, exactly the `fl_ids` of + /// its meta. Pinned **in the snapshot** — the meta slot `txnid & 1` is + /// overwritten by txn `txnid + 2` even while this snapshot stays pinned + /// (TXN-63), so holders (`copy`, the next writer) must never re-read the + /// slot. `Arc<[u64]>` keeps `Snapshot` cheap to clone. + pub free_annex: std::sync::Arc<[u64]>, } impl Snapshot { - /// The snapshot a validated meta page describes. + /// The snapshot a validated meta page describes, with `annex` as the + /// meta's free-list annex ids (read via [`MetaPage::read_annex`] from the + /// same validated slot buffer; `annex.len()` must equal `meta.fl_count`). #[must_use] - pub fn from_meta(meta: &MetaPage) -> Snapshot { + pub fn from_meta(meta: &MetaPage, annex: Vec) -> Snapshot { + debug_assert_eq!(annex.len(), meta.fl_count as usize); Snapshot { txnid: meta.txnid, last_pg: meta.last_pg, main_db: meta.main_db, free_db: meta.free_db, + free_annex: annex.into(), } } } @@ -1635,20 +1646,27 @@ pub fn open_with_backing_policy( let slot1 = read_slot(bytes, META_B_PGNO, ps, page_size)?; // Select the live snapshot (SPEC 02 §3.2 / SPEC 06 REC-2..5). - let meta = match select_meta(&slot0, &slot1, prev_snapshot) { - MetaChoice::Both { meta, .. } => meta, - MetaChoice::OnlyOne { meta, .. } => { + let (meta, chosen_slot) = match select_meta(&slot0, &slot1, prev_snapshot) { + MetaChoice::Both { meta, chosen } => (meta, chosen), + MetaChoice::OnlyOne { meta, chosen } => { if prev_snapshot { // REC-2† (ratified 2026-07-16): one valid slot + PREV_SNAPSHOT is // a hard error — there are not two committed snapshots to pick an // older from. return Err(Error::Mdb(MdbError::Invalid)); } - meta + (meta, chosen) } // REC-3: both invalid → unrecoverable. MetaChoice::None => return Err(Error::Mdb(MdbError::Invalid)), }; + // Format v2 (ADR-0022): read the selected slot's free-list annex ids — + // the one and only meta read, alongside the roots (TXN-18); the slot is + // overwritten two commits later, so the ids are pinned in the Snapshot. + // `fl_count` passed the §3.2 rule-5 bound + CRC; the ids' GC-29 shape is + // validated by the consumer before any id is handed out (GC-33). + let annex = MetaPage::read_annex(&bytes[chosen_slot * ps..(chosen_slot + 1) * ps], page_size) + .ok_or(Error::Mdb(MdbError::Invalid))?; // SPEC 06 REC-1a / SPEC 02 §3.2 step 6 (geometry validation): a slot can // carry a valid CRC and still name geometry the real file cannot back — a @@ -1699,7 +1717,7 @@ pub fn open_with_backing_policy( map_size, // Seed the published-snapshot cell from the durable meta — the one // and only time a meta *page* is read for roots (SPEC 04 TXN-18). - snap_cell: SnapshotCell::new(Arc::new(Snapshot::from_meta(&meta))), + snap_cell: SnapshotCell::new(Arc::new(Snapshot::from_meta(&meta, annex))), write_mutex: WriterLock::new(), commit_hook: Mutex::new(None), poisoned: AtomicBool::new(false), diff --git a/crates/zerodb-core/src/nested.rs b/crates/zerodb-core/src/nested.rs index 35eda73..5c1be14 100644 --- a/crates/zerodb-core/src/nested.rs +++ b/crates/zerodb-core/src/nested.rs @@ -171,6 +171,9 @@ impl TxnRead for NestedRoTxn<'_> { fn free_record(&self) -> &DBRecord { self.parent.free_record() } + fn free_annex_count(&self) -> u64 { + self.parent.free_annex_count() + } fn page_size(&self) -> u32 { self.parent.page_size() } diff --git a/crates/zerodb-core/src/page/crc32c.rs b/crates/zerodb-core/src/page/crc32c.rs index ba9f9b4..4cc3362 100644 --- a/crates/zerodb-core/src/page/crc32c.rs +++ b/crates/zerodb-core/src/page/crc32c.rs @@ -55,6 +55,22 @@ pub fn crc32c(data: &[u8]) -> u32 { crc ^ 0xFFFF_FFFF } +/// CRC32C over the concatenation of `parts`, as one stream — equal to +/// `crc32c` of the parts joined into a single buffer, without the join. +/// Used for the meta CRC's split coverage (SPEC 02 §3.3 as amended by +/// ADR-0022: `[0, 172)` then the annex ids, skipping the CRC field itself). +#[must_use] +pub fn crc32c_concat(parts: &[&[u8]]) -> u32 { + let mut crc = 0xFFFF_FFFFu32; + for part in parts { + for &byte in *part { + let idx = ((crc ^ byte as u32) & 0xFF) as usize; + crc = (crc >> 8) ^ TABLE[idx]; + } + } + crc ^ 0xFFFF_FFFF +} + #[cfg(test)] mod tests { use super::crc32c; diff --git a/crates/zerodb-core/src/page/meta.rs b/crates/zerodb-core/src/page/meta.rs index c4719a9..41b4efa 100644 --- a/crates/zerodb-core/src/page/meta.rs +++ b/crates/zerodb-core/src/page/meta.rs @@ -8,11 +8,13 @@ //! §3.2's numbered list), and the double-buffer selection formula //! ([`select`]). -use super::crc32c::crc32c; +use super::crc32c::crc32c_concat; use super::geometry::validate_page_size; use super::header::CommonHeader; use super::raw::{read_u16, read_u32, read_u64, write_u16, write_u32, write_u64}; -use super::{PageError, FORMAT_VERSION, MAGIC, META_CONTENT_LEN, PGNO_INVALID, P_META}; +use super::{ + PageError, FORMAT_VERSION, MAGIC, META_ANNEX_OFF, META_CONTENT_LEN, PGNO_INVALID, P_META, +}; // Meta body field offsets (absolute within the page). const OFF_MAGIC: usize = 32; @@ -24,7 +26,27 @@ const OFF_LAST_PG: usize = 56; const OFF_BODY_TXNID: usize = 64; const OFF_FREE_DB: usize = 72; const OFF_MAIN_DB: usize = 120; -const OFF_META_CRC: usize = 168; +const OFF_FL_COUNT: usize = 168; +const OFF_META_CRC: usize = META_CONTENT_LEN; // 172 (SPEC 02 §3, format v2) + +/// Maximum number of free-list annex ids a meta page of `psize` bytes can +/// carry (SPEC 02 §3, ADR-0022): the ids start at [`META_ANNEX_OFF`] and run +/// to the end of the page. +#[must_use] +pub fn meta_annex_cap(psize: u32) -> usize { + (psize as usize).saturating_sub(META_ANNEX_OFF) / 8 +} + +/// The meta CRC32C with the ADR-0022 split coverage: `[0, 172)` followed by +/// the `8·fl_count` annex-id bytes at [`META_ANNEX_OFF`], skipping the CRC +/// field itself. `buf` must hold the whole page; `fl_count` must already be +/// bounds-checked against [`meta_annex_cap`]. +fn meta_crc_of(buf: &[u8], fl_count: usize) -> u32 { + crc32c_concat(&[ + &buf[..META_CONTENT_LEN], + &buf[META_ANNEX_OFF..META_ANNEX_OFF + 8 * fl_count], + ]) +} /// Size of a [`DBRecord`], in bytes. pub const DBRECORD_LEN: usize = 48; @@ -147,6 +169,10 @@ pub struct MetaPage { pub free_db: DBRecord, /// Root/stats of the main/catalog DB. pub main_db: DBRecord, + /// Free-list annex id count (SPEC 02 §3 format v2, ADR-0022). The ids + /// themselves stay in the page buffer (offset [`META_ANNEX_OFF`]) and are + /// read with [`MetaPage::read_annex`]; this decoded struct stays `Copy`. + pub fl_count: u32, } impl MetaPage { @@ -165,16 +191,32 @@ impl MetaPage { last_pg: 1, free_db: DBRecord::empty(), main_db: DBRecord::empty(), + fl_count: 0, } } /// Encode this meta into `buf` (which must be at least `page_size` bytes), - /// computing and writing the CRC and zeroing the reserved tail. + /// with an **empty** free-list annex, computing and writing the CRC and + /// zeroing the reserved tail. See [`MetaPage::encode_with_annex`]. /// /// # Errors /// /// [`PageError::InvalidPageSize`] or [`PageError::BufferTooSmall`]. pub fn encode(&self, buf: &mut [u8]) -> Result<(), PageError> { + self.encode_with_annex(buf, &[]) + } + + /// Encode this meta into `buf` with `annex` as the free-list annex ids + /// (SPEC 02 §3 format v2, ADR-0022; SPEC 05 §2a). The caller guarantees + /// the GC-29 shape (strictly ascending, unique, in range) — `freelist_save` + /// produces exactly that; this encoder only enforces the capacity bound. + /// + /// # Errors + /// + /// [`PageError::InvalidPageSize`], [`PageError::BufferTooSmall`], or + /// [`PageError::BadValueSize`] if `annex` exceeds [`meta_annex_cap`] + /// (engine bug: the save's fit check owns that bound). + pub fn encode_with_annex(&self, buf: &mut [u8], annex: &[u64]) -> Result<(), PageError> { validate_page_size(self.page_size)?; let psize = self.page_size as usize; if buf.len() < psize { @@ -183,6 +225,14 @@ impl MetaPage { psize, }); } + if annex.len() > meta_annex_cap(self.page_size) { + return Err(PageError::BadValueSize(annex.len() as u64)); + } + debug_assert_eq!( + self.fl_count as usize, + annex.len(), + "MetaPage.fl_count must match the annex slice (single source: the ids)" + ); // Zero the whole page first so every reserved byte is 0. buf[..psize].fill(0); // Common header (zeros reserved0 + checksum; variant tail already 0). @@ -202,12 +252,40 @@ impl MetaPage { write_u64(buf, OFF_BODY_TXNID, self.txnid); self.free_db.write(buf, OFF_FREE_DB); self.main_db.write(buf, OFF_MAIN_DB); - // CRC over [0, 168). - let crc = crc32c(&buf[..META_CONTENT_LEN]); + // Free-list annex (format v2): count at 168, ids from 176. + write_u32(buf, OFF_FL_COUNT, annex.len() as u32); + for (i, &id) in annex.iter().enumerate() { + write_u64(buf, META_ANNEX_OFF + 8 * i, id); + } + // CRC over [0, 172) ∪ the annex ids (SPEC 02 §3.3). + let crc = meta_crc_of(buf, annex.len()); write_u32(buf, OFF_META_CRC, crc); Ok(()) } + /// Read the free-list annex ids out of a meta page buffer (SPEC 02 §3, + /// format v2). Returns `None` if `fl_count` exceeds the page's capacity — + /// callers treat that as a corrupt freelist (`MdbError::Invalid`), though + /// for a slot that passed [`MetaPage::validate`] the bound already held. + /// The ids' GC-29 shape (ascending, in range) is **not** checked here: + /// the consumer runs `validate_pil_ids` before any id is handed out + /// (SPEC 05 GC-33), exactly as for a tree PIL. + #[must_use] + pub fn read_annex(buf: &[u8], psize: u32) -> Option> { + if buf.len() < psize as usize { + return None; + } + let count = read_u32(buf, OFF_FL_COUNT) as usize; + if count > meta_annex_cap(psize) { + return None; + } + let mut ids = Vec::with_capacity(count); + for i in 0..count { + ids.push(read_u64(buf, META_ANNEX_OFF + 8 * i)); + } + Some(ids) + } + /// Validate `buf` as a meta slot, following SPEC 02 §3.2's numbered list /// (the single owner of the meta-slot validation predicate). Returns a rich /// [`MetaValidity`] verdict; validity failures are *data*, not errors. @@ -251,9 +329,15 @@ impl MetaPage { body: body_txnid, }); } - // Rule 5: CRC over [0, 168). + // Rule 5 (format v2, ADR-0022): the annex count is bounds-checked + // BEFORE the CRC — a hostile count must not drive the CRC read out + // of the page — then the CRC covers [0, 172) ∪ the annex ids. + let fl_count = read_u32(buf, OFF_FL_COUNT) as usize; + if fl_count > meta_annex_cap(psize) { + return Ok(MetaValidity::BadAnnexCount(fl_count as u32)); + } let stored = read_u32(buf, OFF_META_CRC); - let computed = crc32c(&buf[..META_CONTENT_LEN]); + let computed = meta_crc_of(buf, fl_count); if stored != computed { return Ok(MetaValidity::BadCrc { stored, computed }); } @@ -269,6 +353,7 @@ impl MetaPage { last_pg: read_u64(buf, OFF_LAST_PG), free_db: DBRecord::read(buf, OFF_FREE_DB), main_db: DBRecord::read(buf, OFF_MAIN_DB), + fl_count: fl_count as u32, })) } } @@ -285,6 +370,9 @@ pub enum MetaValidity { BadVersion(u32), /// `page_size` is not a power of two in range (the observed value). BadPageSize(u32), + /// `fl_count` exceeds the page's annex capacity (SPEC 02 §3.2 rule 5, + /// format v2) — the slot is invalid (torn or hostile). + BadAnnexCount(u32), /// Header txnid and body txnid disagree — a torn write (INV-2). TxnidMismatch { /// Header stamp (offset 8). diff --git a/crates/zerodb-core/src/page/mod.rs b/crates/zerodb-core/src/page/mod.rs index 2bbb2fb..bc56741 100644 --- a/crates/zerodb-core/src/page/mod.rs +++ b/crates/zerodb-core/src/page/mod.rs @@ -38,11 +38,14 @@ mod tree; // validation; it contains no unsafe operation. mod trust; -pub use crc32c::crc32c; +pub use crc32c::{crc32c, crc32c_concat}; pub(crate) use header::read_flags; pub(crate) use header::read_page_txnid; pub use header::{CommonHeader, PageRef}; -pub use meta::{select as select_meta, DBRecord, MetaChoice, MetaPage, MetaValidity, DBRECORD_LEN}; +pub use meta::{ + meta_annex_cap, select as select_meta, DBRecord, MetaChoice, MetaPage, MetaValidity, + DBRECORD_LEN, +}; pub use overflow::{write_overflow_head, OverflowRef}; pub use tree::{BranchMut, BranchRef, LeafMut, LeafRef, LeafValue}; pub use trust::FileTrust; @@ -54,8 +57,10 @@ pub use trust::FileTrust; /// File identifier, ASCII `"ZDB1"`. Stored as a byte array (endianness-free). pub const MAGIC: [u8; 4] = *b"ZDB1"; -/// On-disk format version. Bumped only on an incompatible change (ADR-0002 §D8). -pub const FORMAT_VERSION: u32 = 1; +/// On-disk format version. Bumped only on an incompatible change (ADR-0002 +/// §D8). Version 2 (ADR-0022): the meta free-list annex — `fl_count` at +/// offset 168, `meta_crc` moved to 172 with split coverage, annex ids at 176. +pub const FORMAT_VERSION: u32 = 2; /// Size of the common page header, in bytes. pub const HEADER_SIZE: usize = 32; @@ -92,8 +97,14 @@ pub const MIN_KEYS_LEAF: usize = 1; /// Minimum children a non-root branch may hold. pub const MIN_KEYS_BRANCH: usize = 2; -/// Number of leading bytes of a meta page covered by its CRC (SPEC 02 §3.3). -pub const META_CONTENT_LEN: usize = 168; +/// Number of leading bytes of a meta page covered by its CRC, before the +/// free-list annex ids (SPEC 02 §3.3 as amended by ADR-0022: the full +/// coverage is `[0, META_CONTENT_LEN) ∪ [META_ANNEX_OFF, … + 8·fl_count)`). +pub const META_CONTENT_LEN: usize = 172; + +/// Absolute byte offset of the meta free-list annex ids (SPEC 02 §3, +/// ADR-0022). The annex capacity is `(psize - META_ANNEX_OFF) / 8` ids. +pub const META_ANNEX_OFF: usize = 176; /// Smallest permitted page size, in bytes. pub const MIN_PAGE_SIZE: u32 = 4096; diff --git a/crates/zerodb-core/src/readers.rs b/crates/zerodb-core/src/readers.rs index c3f6ddb..069735d 100644 --- a/crates/zerodb-core/src/readers.rs +++ b/crates/zerodb-core/src/readers.rs @@ -454,6 +454,7 @@ mod tests { last_pg: 1, main_db: DBRecord::empty(), free_db: DBRecord::empty(), + free_annex: Arc::from(&[][..]), }) } @@ -586,6 +587,7 @@ mod loom_tests { last_pg: 1, main_db: DBRecord::empty(), free_db: DBRecord::empty(), + free_annex: Arc::from(&[][..]), }) } diff --git a/crates/zerodb-core/src/rotxn.rs b/crates/zerodb-core/src/rotxn.rs index c638ea4..374966a 100644 --- a/crates/zerodb-core/src/rotxn.rs +++ b/crates/zerodb-core/src/rotxn.rs @@ -77,6 +77,13 @@ pub trait TxnRead { fn validated_pages(&self) -> Option<&ValidatedPages<'_>> { None } + /// The meta free-list annex id count this txn observes (SPEC 05 §2a, + /// ADR-0022; feeds [`free_page_count`]). A reader reports its snapshot's + /// count; the writer reports its pool's live remainder (a working-state + /// view, like [`TxnRead::free_record`]). + fn free_annex_count(&self) -> u64 { + 0 + } } /// Which database a [`Database`] handle addresses (SPEC 02 §6). `Copy` so the @@ -265,6 +272,9 @@ impl TxnRead for RoTxn<'_> { fn free_record(&self) -> &DBRecord { &self.snap.free_db } + fn free_annex_count(&self) -> u64 { + self.snap.free_annex.len() as u64 + } fn page_size(&self) -> u32 { self.psize } @@ -516,7 +526,9 @@ pub fn free_page_count(txn: &T) -> Result { let source = txn.source(); let tree = Tree::new(source, txn.page_size(), rec.root, rec.depth); let mut cursor = tree.cursor(); - let mut total = 0u64; + // Format v2 (SPEC 05 §2a/GC-23): the snapshot's meta annex ids are free + // pages too; the snapshot carries their count, so no meta re-read. + let mut total = txn.free_annex_count(); let mut entry = cursor.first().map_err(map_page_err)?; while let Some((_key, val)) = entry { // GC-23: sum only each PIL's `count` prefix — no `Vec` decode and diff --git a/crates/zerodb-core/src/rwtxn.rs b/crates/zerodb-core/src/rwtxn.rs index 5ccda95..8d307d2 100644 --- a/crates/zerodb-core/src/rwtxn.rs +++ b/crates/zerodb-core/src/rwtxn.rs @@ -196,6 +196,30 @@ impl Drain { } } +/// Merge two strictly-ascending, mutually-disjoint id slices into one +/// strictly-ascending vec (GC-3/GC-4; the `freelist_save` step-(c) merge of +/// this txn's `freed` with the carried annex remainder, SPEC 05 §2a GC-30). +/// Disjointness holds by provenance — an annex id was free in the base +/// snapshot, a `freed` id was live in it — and is asserted in debug builds. +fn merge_sorted_unique(a: &[u64], b: &[u64]) -> Vec { + let mut out = Vec::with_capacity(a.len() + b.len()); + let (mut i, mut j) = (0, 0); + while i < a.len() && j < b.len() { + if a[i] < b[j] { + out.push(a[i]); + i += 1; + } else { + debug_assert_ne!(a[i], b[j], "annex/freed overlap (INV-24)"); + out.push(b[j]); + j += 1; + } + } + out.extend_from_slice(&a[i..]); + out.extend_from_slice(&b[j..]); + debug_assert!(out.windows(2).all(|w| w[0] < w[1]), "merge not ascending"); + out +} + /// A root-to-leaf descent path: `(pgno, ki)` per level (SPEC 03 §1 cursor /// shape). `ki` is the chosen child index on branches and the entry/insertion /// slot on the leaf. @@ -606,6 +630,18 @@ pub struct RwTxn<'env> { /// [`RwTxn::free_page`]. Pgno-hashed (issue #9): engine-authored keys, /// no HashDoS surface. reclaimed: HashSet, + /// The base meta's free-list annex pool (SPEC 05 §2a, ADR-0022): the + /// pages freed by txn `base.txnid`, parsed (and `validate_pil_ids`- + /// checked, GC-33) at begin when `base.free_annex > 0`. Drawn by + /// [`RwTxn::annex_draw`] (allocate step 2b, gated per GC-31); whatever + /// remains is folded into `freed` at the top of `freelist_save` (the + /// GC-30 carry) — by then the pool is empty, so GcSave never sees it. + annex: Drain, + /// `freelist_save`'s GC-32 placement decision: `true` = this txn's own + /// entry rides in the meta annex (C4 encodes `freed` as `fl_ids`; no + /// GC-tree entry `BE(txnid)` exists), `false` = the v1 tree arm ran (or + /// `freed` is empty) and the annex is empty. + annex_out: bool, /// GC-12 allocation restriction (ADR-0005 D2). alloc_mode: AllocMode, /// Entries of `drains` whose `remaining` shrank via an **in-save pool @@ -739,6 +775,19 @@ impl Env { if base.txnid >= crate::readers::MAX_COMMITTED_TXNID { return Err(Error::Mdb(MdbError::Invalid)); } + // The base meta's free-list annex pool (SPEC 05 §2a, ADR-0022). The + // ids are pinned in the snapshot (the slot is overwritten two + // commits later, TXN-63) and are never trusted (GC-33): they must + // pass the same ascending/range validation as a tree PIL before any + // can be handed out — open-time slots only had the count bound and + // the CRC checked. + let annex = if base.free_annex.is_empty() { + Drain::new(Vec::new()) + } else { + let ids = base.free_annex.to_vec(); + validate_pil_ids(&ids, base.last_pg)?; + Drain::new(ids) + }; Ok(RwTxn { txnid: base.txnid + 1, // TXN-2 (+ the reuse note in the module docs) next_pgno: base.last_pg + 1, @@ -764,6 +813,8 @@ impl Env { loose: Vec::new(), drains: BTreeMap::new(), reclaimed: HashSet::default(), + annex, + annex_out: false, alloc_mode: AllocMode::Normal, save_touched: std::collections::BTreeSet::new(), errored: false, @@ -905,6 +956,11 @@ impl TxnRead for RwTxn<'_> { fn free_record(&self) -> &DBRecord { &self.free_db } + fn free_annex_count(&self) -> u64 { + // Working-state view (like `free_record`): the base annex pool's + // live remainder — drawn ids are no longer free. + Drain::len(&self.annex) as u64 + } fn page_size(&self) -> u32 { self.psize } @@ -1196,6 +1252,13 @@ impl<'env> RwTxn<'env> { if let Some(start) = self.gc_reclaim(n)? { return Ok(start); } + // GC-16 step 2b (SPEC 05 §2a, ADR-0022): the base meta's + // free-list annex pool, after the tree (the annex's + // freeing-txn is the newest possible, so tree-first keeps + // the GC-18/19 oldest-first order). + if let Some(start) = self.annex_draw(n) { + return Ok(start); + } } AllocMode::GcSave => { // GC-12 (as amended, ADR-0005): never *read* the GC tree @@ -1210,6 +1273,14 @@ impl<'env> RwTxn<'env> { if let Some(start) = self.save_pool_draw(n) { return Ok(start); } + // The carried annex pool (SPEC 05 §2a GC-30): gate-checked + // like every annex draw; the step-(c) placement re-reads + // `annex.live()` after every put, so a draw here can never + // leave a handed-out id in the persisted set — the exact + // GC-12-amended pool argument. + if let Some(start) = self.annex_draw(n) { + return Ok(start); + } } } let mp = map_pages(self.env.inner().map_size(), self.psize); @@ -1299,6 +1370,41 @@ impl<'env> RwTxn<'env> { Some(start) } + /// Allocate step 2b (SPEC 05 §2a GC-31, ADR-0022): draw from the base + /// meta's free-list annex pool. The pool's freeing-txnid is `base.txnid` + /// (= `writer_txnid − 1`), so the gate `F ≤ oldest_reader()` admits the + /// draw exactly when no live reader is pinned **below** the base — a + /// reader pinned *at* the base is safe (pages freed by `base` left its + /// trees), and the crash-fallback meta is `base` itself, whose annex + /// still lists the drawn ids as free (the TXN-62 / GC-12 pool argument + /// verbatim). Deterministic like GC-19: smallest id / first contiguous + /// run of the (ascending) pool. Serves both allocation modes: during + /// ops as step 2b, and inside `freelist_save` as the carried pool + /// (GC-30 — the step-(c) placement re-reads `annex.live()` after every + /// tree write, so an in-save draw is always re-accounted). + // Kept out of line: `allocate`'s hot shape is the loose pop; the draw + // arms stay separate so their size never moves it (PERF-GAP B13). + #[inline(never)] + fn annex_draw(&mut self, n: u64) -> Option { + if self.annex.len() == 0 { + return None; + } + let f = self.base.txnid; + if f > self.oldest_reader() { + return None; // GC-31 gate: a reader is pinned below the base. + } + let start = find_run(self.annex.live(), n)?; + // M1.8 shadow check, as in every other draw arm. + #[cfg(debug_assertions)] + self.debug_assert_gate(f); + self.annex.take_run(start, n); + for p in start..start + n { + let first_time = self.reclaimed.insert(p); + debug_assert!(first_time, "page {p} reclaimed twice (INV-24)"); + } + Some(start) + } + /// The GC reuse gate (SPEC 05 GC-18, SPEC 04 TXN-20/21): `min(smallest /// live reader snapshot txnid, writer_txnid − 1)`. The reader term is the /// M1.8 lock-free reader-table SeqCst scan @@ -4008,6 +4114,26 @@ impl<'env> RwTxn<'env> { debug_assert!(self.alloc_mode == AllocMode::Normal); self.alloc_mode = AllocMode::GcSave; debug_assert!(self.save_touched.is_empty()); + // GC-30 carry (SPEC 05 §2a, ADR-0022): the base annex's unconsumed + // remainder must land in this txn's own entry (annex or tree) — meta + // `base` is overwritten by txn `base + 2`, so dropping it leaks. + // The remainder stays a live **in-save pool** (the GC-12-amended + // rule, exactly like the drain pool): its ids passed the gate with + // `F = base.txnid`, so the save's own allocations may consume them — + // without this, every save's GC-tree ops extend the file while the + // carried surplus sits unreachable in `freed`, an unbounded ratchet + // (caught by churn_parity_general during the ADR-0022 spike). The + // step-(c) placement merges `freed ∪ annex.live()` and re-loops if + // the put itself drew from the pool, so no persisted set ever lists + // a handed-out page. The pool is folded into `freed` only after the + // fixed point (annex arm) or spent into the tree entry (tree arm). + // + // GC-32 placement: this txn's own entry goes to the meta annex when + // the final merged set fits; once the tree arm runs it is sticky for + // this save (no flip-flop if a later drain rewrite grows `freed`). + let annex_cap = crate::page::meta_annex_cap(self.psize); + let mut tree_arm = false; + let mut annex_mode = false; // Every ops-drained entry needs its GC-20 rewrite at least once. let mut pending: std::collections::BTreeSet = self.drains.keys().copied().collect(); // Release-active bound on the *inner* rewrite loop (GC-13 guard, ADR @@ -4065,17 +4191,34 @@ impl<'env> RwTxn<'env> { self.put_pil(f, &remaining)?; } } - // (c) this txn's own entry under BE(writer_txnid). + // (c) this txn's own entry: the meta annex when it fits (GC-32), + // else the GC tree under BE(writer_txnid) as in format v1. self.freed.sort_unstable(); let pre_dedup = self.freed.len(); self.freed.dedup(); debug_assert_eq!(self.freed.len(), pre_dedup, "page double-freed (GC-4)"); let before = self.freed.len(); - if before > 0 { - let ids = self.freed.clone(); - self.put_pil(self.txnid, &ids)?; - if self.freed.len() != before { - continue; // the put freed committed GC pages; the PIL must grow + let annex_before = Drain::len(&self.annex); + let total = before + annex_before; + if total > 0 { + if !tree_arm && total <= annex_cap { + // GC-32 annex arm: `freed ∪ annex.live()` IS the annex + // (merged after the loop; C4 encodes it as `fl_ids`). + // The stash writes nothing into the tree, allocates + // nothing and frees nothing, so the fixed point below + // simply loses the step-(c) feedback edge. + annex_mode = true; + } else { + tree_arm = true; // sticky (GC-32) + annex_mode = false; + let ids = merge_sorted_unique(&self.freed, self.annex.live()); + self.put_pil(self.txnid, &ids)?; + if self.freed.len() != before || Drain::len(&self.annex) != annex_before { + // The put COW-freed committed GC pages (PIL must + // grow) or drew from the annex pool (the written + // PIL lists a handed-out page): rewrite next round. + continue; + } } } // Fixed-point checks: no entry awaits a re-rewrite (anti-leak — a @@ -4092,6 +4235,15 @@ impl<'env> RwTxn<'env> { } self.freed.extend(std::mem::take(&mut self.loose)); } + // GC-30/GC-32 settle: in the annex arm the persisted annex is the + // merged set, materialized into `freed` for C4/C6; in the tree arm + // the pool's ids were written into the tree entry `BE(txnid)`. The + // pool is spent either way (no allocation runs after C1). + if annex_mode { + self.freed = merge_sorted_unique(&self.freed, self.annex.live()); + } + self.annex = Drain::new(Vec::new()); + self.annex_out = annex_mode; Ok(()) } @@ -4194,6 +4346,10 @@ impl<'env> RwTxn<'env> { // slot; the intact `N-1` slot is the crash fallback). ----- let last_pg = self.next_pgno - 1; let slot = self.txnid & 1; + // GC-32 (ADR-0022): in annex mode this txn's freed PIL rides in the + // meta itself — `freed` is final after C1 and already sorted/deduped + // (GC-3/4); `freelist_save` guaranteed it fits the cap. + let annex: &[u64] = if self.annex_out { &self.freed } else { &[] }; let meta = MetaPage { pgno: slot, txnid: self.txnid, @@ -4205,9 +4361,10 @@ impl<'env> RwTxn<'env> { last_pg, free_db: self.free_db, main_db: self.main_db, + fl_count: annex.len() as u32, }; let mut buf = vec![0u8; psize as usize]; - meta.encode(&mut buf).map_err(corrupt)?; + meta.encode_with_annex(&mut buf, annex).map_err(corrupt)?; inner.backing_ref().write_at_page(slot, psize, &buf)?; inner.run_hook(HookPoint::H3); // Crash here: recovered snapshot is `N-1` (meta missing/torn → CRC @@ -4270,6 +4427,11 @@ impl<'env> RwTxn<'env> { last_pg, main_db: self.main_db, free_db: self.free_db, + free_annex: if self.annex_out { + Arc::from(&self.freed[..]) + } else { + Arc::from(&[][..]) + }, })); Ok(()) // Drop of `self` releases the write mutex — the tail of C6. diff --git a/crates/zerodb-core/tests/page_edges.rs b/crates/zerodb-core/tests/page_edges.rs index abc2c76..2ec62ca 100644 --- a/crates/zerodb-core/tests/page_edges.rs +++ b/crates/zerodb-core/tests/page_edges.rs @@ -520,16 +520,18 @@ fn built_valid_meta(psize: u32) -> Vec { buf } -/// Every single-byte flip in the CRC-covered region [0, META_CONTENT_LEN) -/// must invalidate the slot; every single-byte flip in the excluded reserved -/// tail [172, psize) must NOT (SPEC 02 §3.3 — the sector-tear reality of -/// REC-22). The `meta_crc` field itself [168, 172) is not part of the hashed -/// range but flipping it desyncs stored-vs-recomputed CRC, so it also -/// invalidates. +/// Every single-byte flip in the CRC-covered region [0, META_CONTENT_LEN = +/// 172) — which since format v2 includes `fl_count` at [168, 172) (ADR-0022; +/// a silently-zeroed annex count must tear the slot) — must invalidate the +/// slot; every single-byte flip in the excluded reserved tail +/// [176, psize) must NOT, for an empty-annex meta (SPEC 02 §3.3 — the +/// sector-tear reality of REC-22). The `meta_crc` field itself [172, 176) is +/// not part of the hashed range but flipping it desyncs stored-vs-recomputed +/// CRC, so it also invalidates. fn meta_validation_rejection_matrix(psize: u32) { let buf = built_valid_meta(psize); - for i in 0..168usize { + for i in 0..172usize { let mut b = buf.clone(); b[i] ^= 0xFF; let v = MetaPage::validate(&b, psize).unwrap(); @@ -538,7 +540,7 @@ fn meta_validation_rejection_matrix(psize: u32) { "byte {i} (CRC-covered) flip must invalidate meta" ); } - for i in 168..172usize { + for i in 172..176usize { let mut b = buf.clone(); b[i] ^= 0xFF; let v = MetaPage::validate(&b, psize).unwrap(); @@ -549,9 +551,11 @@ fn meta_validation_rejection_matrix(psize: u32) { } // Sample the reserved tail rather than iterating every byte (up to ~65KB) // to stay within the ~30s test-runtime budget; every sampled byte must - // NOT invalidate, documenting that this range is unprotected by the CRC. - let stride = ((psize as usize - 172) / 40).max(1); - for i in (172..psize as usize).step_by(stride) { + // NOT invalidate, documenting that this range is unprotected by the CRC + // when the annex is empty (with a non-empty annex the ids' bytes ARE + // covered — the spec02_format annex tests lock that). + let stride = ((psize as usize - 176) / 40).max(1); + for i in (176..psize as usize).step_by(stride) { let mut b = buf.clone(); b[i] ^= 0xFF; let v = MetaPage::validate(&b, psize).unwrap(); diff --git a/crates/zerodb-core/tests/spec02_format.rs b/crates/zerodb-core/tests/spec02_format.rs index b0d26a5..aaaf91e 100644 --- a/crates/zerodb-core/tests/spec02_format.rs +++ b/crates/zerodb-core/tests/spec02_format.rs @@ -16,15 +16,16 @@ const PSIZE: u32 = 4096; // §3.5 — creation meta, slot 0 (page size 4096, txnid 0, map_size 1 MiB) // --------------------------------------------------------------------------- -/// Build the expected first-168 bytes of the §3.5 creation meta. -fn expected_meta_content() -> [u8; 168] { - let mut e = [0u8; 168]; +/// Build the expected first-172 bytes of the §3.5 creation meta (format v2, +/// ADR-0022: `fl_count` at 168 — zero for a creation meta — then the CRC). +fn expected_meta_content() -> [u8; 172] { + let mut e = [0u8; 172]; // off 0: pgno = 0 (all zero) // off 8: txnid = 0 (all zero) e[16] = 0x08; // flags = P_META (0x0008) // off 18..32 reserved 0 e[32..36].copy_from_slice(b"ZDB1"); // magic 5A 44 42 31 - e[36] = 0x01; // format_version = 1 + e[36] = 0x02; // format_version = 2 (ADR-0022) e[40] = 0x00; e[41] = 0x10; // page_size = 4096 (0x0000_1000) // off 44 env_flags = 0 @@ -39,6 +40,7 @@ fn expected_meta_content() -> [u8; 168] { *b = 0xFF; // main_db.root = PGNO_INVALID } // off 128..168 main_db stats 0 + // off 168..172 fl_count = 0 (empty free-list annex) e } @@ -48,21 +50,22 @@ fn meta_creation_slot0_bytes() { let mut buf = vec![0u8; PSIZE as usize]; meta.encode(&mut buf).unwrap(); - // Lock [0, 168) byte-for-byte. + // Lock [0, 172) byte-for-byte. assert_eq!( - &buf[..168], + &buf[..172], &expected_meta_content()[..], - "meta content [0,168)" + "meta content [0,172)" ); - // The CRC field must be exactly CRC32C over [0, 168). - let crc = page::crc32c(&buf[..168]); - let stored = u32::from_le_bytes([buf[168], buf[169], buf[170], buf[171]]); - assert_eq!(stored, crc, "meta_crc must cover [0,168)"); + // The CRC field must be exactly CRC32C over [0, 172) — fl_count is 0, so + // no annex ids extend the coverage (SPEC 02 §3.3, format v2). + let crc = page::crc32c(&buf[..172]); + let stored = u32::from_le_bytes([buf[172], buf[173], buf[174], buf[175]]); + assert_eq!(stored, crc, "meta_crc must cover [0,172) when fl_count = 0"); - // The reserved tail [172, psize) must be all zero (excluded from CRC). + // The reserved tail [176, psize) must be all zero (excluded from CRC). assert!( - buf[172..].iter().all(|&b| b == 0), + buf[176..].iter().all(|&b| b == 0), "reserved tail must be zero" ); @@ -94,7 +97,122 @@ fn meta_slots_identical_at_creation() { // Only the pgno field (offset 0) differs between the slots. assert_eq!(b0[0], 0); assert_eq!(b1[0], 1); - assert_eq!(&b0[8..168], &b1[8..168], "bodies identical apart from pgno"); + assert_eq!(&b0[8..172], &b1[8..172], "bodies identical apart from pgno"); +} + +// --------------------------------------------------------------------------- +// §3 format v2 — the meta free-list annex (ADR-0022; SPEC 05 §2a GC-29/33) +// --------------------------------------------------------------------------- + +#[test] +fn meta_annex_encode_validate_read_roundtrip() { + let ids: Vec = vec![2, 5, 6, 7, 40]; + let mut meta = MetaPage::create(0, PSIZE, 1024 * 1024); + meta.last_pg = 64; + meta.fl_count = ids.len() as u32; + let mut buf = vec![0u8; PSIZE as usize]; + meta.encode_with_annex(&mut buf, &ids).unwrap(); + + // Field placement: fl_count at 168 (LE), ids from 176 (LE u64 each). + assert_eq!(&buf[168..172], &(ids.len() as u32).to_le_bytes()); + for (i, id) in ids.iter().enumerate() { + assert_eq!(&buf[176 + 8 * i..184 + 8 * i], &id.to_le_bytes()); + } + // CRC coverage = [0, 172) ∪ the annex ids (SPEC 02 §3.3), one stream. + let crc = page::crc32c_concat(&[&buf[..172], &buf[176..176 + 8 * ids.len()]]); + assert_eq!(&buf[172..176], &crc.to_le_bytes()); + // Equivalence with the single-shot CRC over the joined bytes. + let mut joined = buf[..172].to_vec(); + joined.extend_from_slice(&buf[176..176 + 8 * ids.len()]); + assert_eq!(crc, page::crc32c(&joined)); + + // Validates and decodes; the ids read back exactly. + match MetaPage::validate(&buf, PSIZE).unwrap() { + MetaValidity::Valid(decoded) => { + assert_eq!(decoded.fl_count, ids.len() as u32); + assert_eq!(decoded, meta); + } + other => panic!("expected Valid, got {other:?}"), + } + assert_eq!(MetaPage::read_annex(&buf, PSIZE).unwrap(), ids); +} + +#[test] +fn meta_annex_torn_id_rejected_by_crc() { + let ids: Vec = (2..60).collect(); + let mut meta = MetaPage::create(0, PSIZE, 1024 * 1024); + meta.last_pg = 64; + meta.fl_count = ids.len() as u32; + let mut buf = vec![0u8; PSIZE as usize]; + meta.encode_with_annex(&mut buf, &ids).unwrap(); + // Flip one byte inside the LAST annex id — far past META_CONTENT_LEN. + buf[176 + 8 * (ids.len() - 1)] ^= 0xFF; + assert!( + matches!( + MetaPage::validate(&buf, PSIZE).unwrap(), + MetaValidity::BadCrc { .. } + ), + "a torn annex id must tear the slot (REC-8 annex note)" + ); +} + +#[test] +fn meta_annex_zeroed_count_rejected_by_crc() { + // The leak guard: a torn write that silently zeroes `fl_count` (losing + // the freed-page record) must invalidate the slot, because fl_count is + // inside the main CRC coverage (SPEC 02 §3.3). + let ids: Vec = vec![2, 3, 4]; + let mut meta = MetaPage::create(0, PSIZE, 1024 * 1024); + meta.last_pg = 8; + meta.fl_count = ids.len() as u32; + let mut buf = vec![0u8; PSIZE as usize]; + meta.encode_with_annex(&mut buf, &ids).unwrap(); + buf[168..172].fill(0); // fl_count := 0, stale CRC left in place + buf[176..176 + 24].fill(0); // ids zeroed too (torn-to-zeros tail) + assert!( + matches!( + MetaPage::validate(&buf, PSIZE).unwrap(), + MetaValidity::BadCrc { .. } + ), + "zeroing the annex must tear the slot, never read as 'no annex'" + ); +} + +#[test] +fn meta_annex_count_over_cap_rejected_before_crc() { + let meta = MetaPage::create(0, PSIZE, 1024 * 1024); + let mut buf = vec![0u8; PSIZE as usize]; + meta.encode(&mut buf).unwrap(); + // A hostile count far past the page's capacity must be refused by the + // bound check (rule 5, before the CRC read), not read out of bounds. + buf[168..172].copy_from_slice(&u32::MAX.to_le_bytes()); + assert!(matches!( + MetaPage::validate(&buf, PSIZE).unwrap(), + MetaValidity::BadAnnexCount(_) + )); + assert_eq!(MetaPage::read_annex(&buf, PSIZE), None); + // Encoding more ids than fit is a typed error, never a panic. + let too_many: Vec = (2..2 + page::meta_annex_cap(PSIZE) as u64 + 1).collect(); + let mut m2 = MetaPage::create(0, PSIZE, 1024 * 1024); + m2.fl_count = too_many.len() as u32; + assert!(m2.encode_with_annex(&mut buf, &too_many).is_err()); +} + +#[test] +fn meta_annex_cap_arithmetic() { + // (psize − 176) / 8, per SPEC 02 §3. + assert_eq!(page::meta_annex_cap(4096), (4096 - 176) / 8); + assert_eq!(page::meta_annex_cap(65536), (65536 - 176) / 8); + // A full-cap annex encodes and validates. + let cap = page::meta_annex_cap(PSIZE); + let ids: Vec = (2..2 + cap as u64).collect(); + let mut meta = MetaPage::create(0, PSIZE, 1024 * 1024); + meta.last_pg = ids.last().copied().unwrap() + 1; + meta.fl_count = cap as u32; + let mut buf = vec![0u8; PSIZE as usize]; + meta.encode_with_annex(&mut buf, &ids).unwrap(); + assert!(MetaPage::validate(&buf, PSIZE).unwrap().is_valid()); + assert_eq!(MetaPage::read_annex(&buf, PSIZE).unwrap(), ids); } // --------------------------------------------------------------------------- diff --git a/crates/zerodb-oracle/tests/crash_harness_smoke.rs b/crates/zerodb-oracle/tests/crash_harness_smoke.rs index 2b29a75..0b22552 100644 --- a/crates/zerodb-oracle/tests/crash_harness_smoke.rs +++ b/crates/zerodb-oracle/tests/crash_harness_smoke.rs @@ -152,7 +152,17 @@ fn regression_nometasync_reclaim_clobber_window_is_characterized() { // it never touches `Op` — so the mode guard below could not catch it). // The assertions are unchanged; only the vehicle was replaced, by searching // for a seed that still satisfies all four of them. - const SEED: u64 = 11_834_834_180_059_103_290; + // + // RE-PINNED 2026-10-03 (was 11_834_834_180_059_103_290): ADR-0022 (the + // meta free-list annex, format v2) changes which pages a txn reclaims + // when — small freed sets ride the meta and are reused one hop earlier — + // so the old seed's cut stopped producing a stale-fallback image (0 + // stale; nothing was violated, the vehicle was lost again). Same re-pin + // protocol: seed 198 satisfies all four assertions (12 verified, 4 stale + // fallbacks, NoMetaSync, no violation). The clobber window itself is + // unchanged by the annex — the reclaimed-page TXN-62/GC-18 reasoning is + // identical whether the freed list lived in the tree or the meta. + const SEED: u64 = 198; assert_eq!( gen_spec(SEED).mode, Mode::NoMetaSync, diff --git a/crates/zerodb-tools/tests/data_file_probe.rs b/crates/zerodb-tools/tests/data_file_probe.rs index b5bb348..2f2ce50 100644 --- a/crates/zerodb-tools/tests/data_file_probe.rs +++ b/crates/zerodb-tools/tests/data_file_probe.rs @@ -92,7 +92,7 @@ fn stat_prints_the_engine_identification_line() { .expect("stat must print an engine line"); assert!(line.contains("zerodb"), "engine line: {line}"); assert!(line.contains("ZDB1"), "engine line: {line}"); - assert!(line.contains("format_version 1"), "engine line: {line}"); + assert!(line.contains("format_version 2"), "engine line: {line}"); assert!(line.contains("data.mdb"), "engine line: {line}"); } diff --git a/crates/zerodb/src/copy.rs b/crates/zerodb/src/copy.rs index 3cf7668..e8a78cb 100644 --- a/crates/zerodb/src/copy.rs +++ b/crates/zerodb/src/copy.rs @@ -388,6 +388,10 @@ fn write_raw( // Synthesize both meta slots from the pinned snapshot — the copy is a // self-contained env at snapshot T, freelist preserved (SPEC 02 §3). + // Format v2 (ADR-0022): the snapshot's free-list annex rides along — + // the ids are pinned in `Snapshot` (the live env's slot may already + // hold a newer meta, TXN-63), and dropping them would leak those pages + // in the copy (INV-22/INV-28). let mut metas = vec![0u8; 2 * ps]; let mut meta = MetaPage { pgno: 0, @@ -400,10 +404,13 @@ fn write_raw( last_pg: snap.last_pg, free_db: snap.free_db, main_db: snap.main_db, + fl_count: snap.free_annex.len() as u32, }; - meta.encode(&mut metas[0..ps]).map_err(corrupt)?; + meta.encode_with_annex(&mut metas[0..ps], &snap.free_annex) + .map_err(corrupt)?; meta.pgno = 1; - meta.encode(&mut metas[ps..2 * ps]).map_err(corrupt)?; + meta.encode_with_annex(&mut metas[ps..2 * ps], &snap.free_annex) + .map_err(corrupt)?; file.write_all_at(&metas, base)?; // Data pages in chunks, so progress is observable. The chunk is sized diff --git a/crates/zerodb/tests/loose_page_and_trailing_shrink.rs b/crates/zerodb/tests/loose_page_and_trailing_shrink.rs index f709135..1702716 100644 --- a/crates/zerodb/tests/loose_page_and_trailing_shrink.rs +++ b/crates/zerodb/tests/loose_page_and_trailing_shrink.rs @@ -98,11 +98,12 @@ fn same_txn_alloc_then_free_overflow_reused_leaves_only_the_leaf_cow() { .len(); assert_eq!( size_after, - size_before + 2 * PS as u64, + size_before + PS as u64, "the whole 5-page overflow run must be reclaimed (GC-7/8/10); the only \ - durable growth allowed is the one-page leaf COW plus the one GC leaf \ - page recording the freed old leaf (M1.5 freelist_save, SPEC 05 GC-11 — \ - pre-M1.5 this txn leaked the old leaf instead of listing it)" + durable growth allowed is the one-page leaf COW — the freed old leaf \ + is recorded in the meta free-list annex (SPEC 05 §2a GC-29, ADR-0022, \ + format v2), no longer in a GC tree leaf (format v1 grew one extra GC \ + leaf page here; pre-M1.5 the old leaf leaked instead of being listed)" ); let rtxn = env.read_txn().unwrap(); @@ -155,11 +156,11 @@ fn trailing_loose_pages_shrink_next_pgno() { .len(); assert_eq!( size_after, - size_seed + 2 * PS as u64, + size_seed + PS as u64, "a 10-page trailing alloc-then-free must shrink back to just the one \ - unavoidable leaf-COW page (GC-10) plus the one GC leaf page recording \ - the freed old leaf (M1.5 freelist_save), not leak any of the 10 \ - overflow pages" + unavoidable leaf-COW page (GC-10) — the freed old leaf is recorded in \ + the meta free-list annex (SPEC 05 §2a GC-29, ADR-0022, format v2), not \ + in a GC tree leaf — and must not leak any of the 10 overflow pages" ); // The env is still fully usable afterward (the rolled-back pgnos are diff --git a/crates/zerodb/tests/meta_annex_gc.rs b/crates/zerodb/tests/meta_annex_gc.rs new file mode 100644 index 0000000..70c1be4 --- /dev/null +++ b/crates/zerodb/tests/meta_annex_gc.rs @@ -0,0 +1,343 @@ +//! Meta free-list annex behavior (SPEC 05 §2a GC-29..33, SPEC 02 §3 format +//! v2, ADR-0022): steady-state placement, the GC-31 reader gate, the GC-30 +//! carry, the GC-32 spill-to-tree, copy, and reopen — each asserted through +//! observable state (file size, `check_image`, `free_page_count`, the meta +//! bytes) rather than internals. + +use std::path::{Path, PathBuf}; +use std::sync::atomic::{AtomicU64, Ordering}; + +use zerodb::{check, free_page_count, CompactionOption, CopyToFile, Env, EnvOpenOptions}; +use zerodb_core::page::{MetaPage, MetaValidity, PGNO_INVALID}; + +static COUNTER: AtomicU64 = AtomicU64::new(0); + +struct TempDir { + path: PathBuf, +} + +impl TempDir { + fn new() -> TempDir { + let pid = std::process::id(); + let seq = COUNTER.fetch_add(1, Ordering::Relaxed); + let path = std::env::temp_dir().join(format!("zerodb-annex-{pid}-{seq}")); + std::fs::create_dir_all(&path).expect("create temp dir"); + TempDir { path } + } + fn path(&self) -> &Path { + &self.path + } +} + +impl Drop for TempDir { + fn drop(&mut self) { + let _ = std::fs::remove_dir_all(&self.path); + } +} + +const PS: u32 = 4096; +const MAP: usize = 32 << 20; + +fn open(dir: &Path) -> Env { + let mut opts = EnvOpenOptions::new(); + opts.map_size(MAP); + opts.page_size(PS); + opts.open(dir).expect("open env") +} + +fn data_file(dir: &Path) -> PathBuf { + dir.join(zerodb::DATA_FILE_NAME) +} + +fn assert_clean(dir: &Path) { + let bytes = std::fs::read(data_file(dir)).expect("read data file"); + let v = check::check_image(&bytes, PS); + assert!(v.is_empty(), "invariant violations: {v:#?}"); +} + +fn file_size(dir: &Path) -> u64 { + std::fs::metadata(data_file(dir)).unwrap().len() +} + +/// Read the LIVE meta's `(txnid, fl_count, free_db_root)` straight from the +/// data file (both slots validated, higher txnid wins — SPEC 02 §3.2). +fn live_meta(dir: &Path) -> (u64, u32, u64) { + let bytes = std::fs::read(data_file(dir)).unwrap(); + let ps = PS as usize; + let pick = |b: &[u8]| match MetaPage::validate(b, PS).unwrap() { + MetaValidity::Valid(m) => Some(m), + _ => None, + }; + let m0 = pick(&bytes[..ps]); + let m1 = pick(&bytes[ps..2 * ps]); + let m = match (m0, m1) { + (Some(a), Some(b)) => { + if a.txnid >= b.txnid { + a + } else { + b + } + } + (Some(a), None) => a, + (None, Some(b)) => b, + (None, None) => panic!("no valid meta slot"), + }; + (m.txnid, m.fl_count, m.free_db.root) +} + +/// GC-29/GC-32 steady state: small-churn commits keep the whole free list in +/// the meta annex — the GC tree never materializes — and page reuse keeps the +/// file size flat. +#[test] +fn steady_state_annex_only_gc_tree_stays_empty() { + let dir = TempDir::new(); + let env = open(dir.path()); + let db = env.main_database(); + + let val = vec![0xABu8; 256]; + for i in 0u64..60 { + let mut w = env.write_txn().unwrap(); + db.put(&mut w, &i.to_be_bytes(), &val).unwrap(); + w.commit().unwrap(); + } + let (_, fl_mid, root_mid) = live_meta(dir.path()); + assert!(fl_mid > 0, "steady-state commits must carry a meta annex"); + assert_eq!( + root_mid, PGNO_INVALID, + "small churn must never materialize the GC tree (GC-29/GC-32)" + ); + + // Overwrite churn: every commit frees the COW'd path and reuses the + // previous commit's annex pages. After a few settling commits the file + // must not grow at all. + for i in 0u64..10 { + let mut w = env.write_txn().unwrap(); + db.put(&mut w, &i.to_be_bytes(), &val).unwrap(); + w.commit().unwrap(); + } + let size_settled = file_size(dir.path()); + for i in 0u64..50 { + let mut w = env.write_txn().unwrap(); + db.put(&mut w, &i.to_be_bytes(), &val).unwrap(); + w.commit().unwrap(); + } + assert_eq!( + file_size(dir.path()), + size_settled, + "annex reuse must keep overwrite churn at zero file growth" + ); + assert_clean(dir.path()); + + // free_page_count sees the annex ids (GC-23 as amended). + let rtxn = env.read_txn().unwrap(); + let (_, fl_now, root_now) = live_meta(dir.path()); + assert_eq!(root_now, PGNO_INVALID); + assert_eq!(free_page_count(&rtxn).unwrap(), u64::from(fl_now)); +} + +/// GC-31 gate + GC-30 carry: a reader pinned below the base blocks annex +/// draws (the file grows while it lives, nothing leaks), and after it +/// releases, the carried ids are reclaimed and growth stops. +#[test] +fn reader_gate_blocks_annex_draw_then_carry_reclaims() { + let dir = TempDir::new(); + let env = open(dir.path()); + let db = env.main_database(); + let val = vec![0x44u8; 256]; + + for i in 0u64..8 { + let mut w = env.write_txn().unwrap(); + db.put(&mut w, &i.to_be_bytes(), &val).unwrap(); + w.commit().unwrap(); + } + // Pin NOW; the next commit's base will be newer than this pin, so every + // later annex (and tree entry) is gated off while the reader lives. + let pin = env.read_txn().unwrap(); + { + let mut w = env.write_txn().unwrap(); + db.put(&mut w, &0u64.to_be_bytes(), &val).unwrap(); + w.commit().unwrap(); + } + let size_pinned_base = file_size(dir.path()); + for i in 0u64..6 { + let mut w = env.write_txn().unwrap(); + db.put(&mut w, &i.to_be_bytes(), &val).unwrap(); + w.commit().unwrap(); + } + let size_during = file_size(dir.path()); + assert!( + size_during > size_pinned_base, + "with a reader pinned below every base, commits must extend the file \ + (the GC-31 gate refuses the annex) — got no growth, so the gate leaked" + ); + assert_clean(dir.path()); + drop(pin); + // Carried ids (GC-30: each commit re-listed the still-free remainder) + // become reclaimable; churn must stop growing the file. + for _ in 0..3 { + let mut w = env.write_txn().unwrap(); + db.put(&mut w, &1u64.to_be_bytes(), &val).unwrap(); + w.commit().unwrap(); + } + let size_settled = file_size(dir.path()); + for i in 0u64..8 { + let mut w = env.write_txn().unwrap(); + db.put(&mut w, &i.to_be_bytes(), &val).unwrap(); + w.commit().unwrap(); + } + assert_eq!( + file_size(dir.path()), + size_settled, + "after the reader releases, carried annex ids must satisfy churn" + ); + assert_clean(dir.path()); +} + +/// GC-32 all-or-nothing spill: a commit freeing more ids than the annex cap +/// writes the whole PIL to the GC tree (annex empty), and those pages are +/// reclaimable afterwards. +#[test] +fn over_cap_freed_set_spills_whole_to_gc_tree() { + let dir = TempDir::new(); + let env = open(dir.path()); + let db = env.main_database(); + let cap = zerodb_core::page::meta_annex_cap(PS) as u64; // 490 at 4 KiB + + // One huge overflow value, committed... + { + let mut w = env.write_txn().unwrap(); + db.put( + &mut w, + b"huge", + &vec![0x7Au8; (cap as usize + 30) * PS as usize], + ) + .unwrap(); + w.commit().unwrap(); + } + // ...then freed in the next txn: freed > cap ⇒ the tree arm. + { + let mut w = env.write_txn().unwrap(); + assert!(db.delete(&mut w, b"huge").unwrap()); + w.commit().unwrap(); + } + let (_, fl, root) = live_meta(dir.path()); + assert_eq!( + fl, 0, + "an over-cap freed set must leave the annex empty (GC-32)" + ); + assert_ne!(root, PGNO_INVALID, "the spill must land in the GC tree"); + assert_clean(dir.path()); + { + let rtxn = env.read_txn().unwrap(); + assert!(free_page_count(&rtxn).unwrap() > cap); + } + + // The spilled pages are drawn back out (tree first, GC-18): re-inserting + // a similar value must not grow the file beyond its high-water. + let size_spilled = file_size(dir.path()); + { + let mut w = env.write_txn().unwrap(); + db.put(&mut w, b"huge2", &vec![0x7Bu8; cap as usize * PS as usize]) + .unwrap(); + w.commit().unwrap(); + } + assert!( + file_size(dir.path()) <= size_spilled, + "re-insert must reuse the spilled run, not extend" + ); + assert_clean(dir.path()); +} + +/// Non-compact copy carries the pinned snapshot's annex (the live slot may +/// already hold a newer meta — the ids are pinned in the Snapshot), so the +/// copy is leak-free and reports the same free count. The compact copy drops +/// free pages by construction and must also come out clean. +#[test] +fn copy_preserves_annex_compact_drops_it() { + let dir = TempDir::new(); + let env = open(dir.path()); + let db = env.main_database(); + let val = vec![0x55u8; 256]; + for i in 0u64..20 { + let mut w = env.write_txn().unwrap(); + db.put(&mut w, &i.to_be_bytes(), &val).unwrap(); + w.commit().unwrap(); + } + let (_, fl, _) = live_meta(dir.path()); + assert!(fl > 0, "fixture must have a non-empty annex"); + let src_free = { + let rtxn = env.read_txn().unwrap(); + free_page_count(&rtxn).unwrap() + }; + + let raw = TempDir::new(); + let raw_path = raw.path().join("copy-raw.zdb"); + env.copy_to_file(&raw_path, CompactionOption::Disabled) + .unwrap(); + let bytes = std::fs::read(&raw_path).unwrap(); + let v = check::check_image(&bytes, PS); + assert!(v.is_empty(), "raw copy violations: {v:#?}"); + let copy_env = { + let dirp = raw.path().join("raw-env"); + std::fs::create_dir_all(&dirp).unwrap(); + std::fs::copy(&raw_path, dirp.join(zerodb::DATA_FILE_NAME)).unwrap(); + open(&dirp) + }; + { + let rtxn = copy_env.read_txn().unwrap(); + assert_eq!( + free_page_count(&rtxn).unwrap(), + src_free, + "raw copy must preserve the freelist, annex included" + ); + assert_eq!( + copy_env + .main_database() + .get(&rtxn, &3u64.to_be_bytes()) + .unwrap(), + Some(val.as_slice()) + ); + } + + let compact_path = raw.path().join("copy-compact.zdb"); + env.copy_to_file(&compact_path, CompactionOption::Enabled) + .unwrap(); + let bytes = std::fs::read(&compact_path).unwrap(); + let v = check::check_image(&bytes, PS); + assert!(v.is_empty(), "compact copy violations: {v:#?}"); +} + +/// Reopen re-parses the annex from the selected slot (GC-33 validation at the +/// writer's begin): reuse keeps working across a close/open cycle. +#[test] +fn reopen_reparses_annex_and_reuses() { + let dir = TempDir::new(); + { + let env = open(dir.path()); + let db = env.main_database(); + let val = vec![0x66u8; 256]; + for i in 0u64..20 { + let mut w = env.write_txn().unwrap(); + db.put(&mut w, &i.to_be_bytes(), &val).unwrap(); + w.commit().unwrap(); + } + } + let (_, fl, _) = live_meta(dir.path()); + assert!(fl > 0); + let size_before = file_size(dir.path()); + + let env = open(dir.path()); + let db = env.main_database(); + let val = vec![0x66u8; 256]; + for i in 0u64..10 { + let mut w = env.write_txn().unwrap(); + db.put(&mut w, &i.to_be_bytes(), &val).unwrap(); + w.commit().unwrap(); + } + assert_eq!( + file_size(dir.path()), + size_before, + "the reopened env must draw from the persisted annex, not extend" + ); + assert_clean(dir.path()); +} diff --git a/docs/DECISIONS.md b/docs/DECISIONS.md index 44731d1..978daf6 100644 --- a/docs/DECISIONS.md +++ b/docs/DECISIONS.md @@ -24,3 +24,4 @@ | 0019 | [Durable meta write through an O_DSYNC descriptor — one barrier per durable commit, as LMDB's `me_mfd`](adr/0019-meta-write-dsync.md) | Accepted (all platforms; LMDB's failed-write scrub adopted) | Phase 3 (durable-commit cost) | | 0020 | [Shrink the 32-byte common page header (24 bytes, keeping pgno and the writer stamp)](adr/0020-compact-page-header.md) | Spike approved; format change pending its numbers | Phase 3 (density) | | 0021 | [True in-place WRITE_MAP — dirty pages in the map, no heap staging, no commit write-back](adr/0021-writemap-in-place.md) | Accepted (2026-10-02, Quentin); spike first — fair WRITE_MAP comparison shows ZeroDB-writemap 0.42–0.53× of LMDB-writemap | Phase 3 (PERF-GAP B5, issue #13) | +| 0022 | [Meta free-list annex — the per-commit freed PIL rides inside the (CRC-covered) meta page; GC tree becomes the spill/cold path; FORMAT_VERSION 2](adr/0022-meta-freelist-annex.md) | Spike implemented on relayed approval; awaits direct ratification | perf lever B12 (free-list save); feeds PLAN 3.1 | diff --git a/docs/SPEC/02-pages.md b/docs/SPEC/02-pages.md index 9bb5402..19baffd 100644 --- a/docs/SPEC/02-pages.md +++ b/docs/SPEC/02-pages.md @@ -73,7 +73,7 @@ LMDB struct is transliterated. | `FILL_THRESHOLD_PERMILLE` | `250` | 25.0 % — below this a page is a merge/borrow candidate (SPEC 03). | | `MIN_KEYS_LEAF` | `1` | Min entries a non-root leaf may hold after delete. | | `MIN_KEYS_BRANCH` | `2` | Min children a non-root branch may hold. | -| `META_CONTENT_LEN` | `168` | Number of leading bytes of a meta page covered by its CRC (see §3). | +| `META_CONTENT_LEN` | `172` | Number of leading bytes of a meta page covered by its CRC, before the annex ids (see §3/§3.3; format v2, ADR-0022). | **Page-type flags** (`u16`, in the common header `flags` field; a page has exactly one of the first four structural bits set): @@ -177,8 +177,10 @@ body** begins at offset 32: | 64 | 8 | `txnid` | u64 | Commit txnid (== header `txnid`). | | 72 | 48 | `free_db` | DBRecord | Root/stats of the free (GC) DB — `FREE_DBI`. §3.1. | | 120 | 48 | `main_db` | DBRecord | Root/stats of the main/catalog DB — `MAIN_DBI`. §3.1. | -| 168 | 4 | `meta_crc` | u32 | **Mandatory CRC32C** over bytes `[0, 168)` of this page (see §3.3). | -| 172 | psize−172 | reserved | — | MUST be 0; **not** covered by the CRC. | +| 168 | 4 | `fl_count` | u32 | **Free-list annex** id count (ADR-0022, format v2): `0 ≤ fl_count ≤ (psize − 176) / 8`. See the annex row below and SPEC 05 §2a. | +| 172 | 4 | `meta_crc` | u32 | **Mandatory CRC32C** over bytes `[0, 172) ∪ [176, 176 + 8·fl_count)` of this page (see §3.3). | +| 176 | 8·fl_count | `fl_ids` | u64[] | The annex: page ids freed by **this meta's txnid** (the PIL format v1 stored in the GC tree under `BE(txnid)`), little-endian, strictly ascending, unique, each in `[FIRST_DATA_PGNO, last_pg]`. Semantics are owned by SPEC 05 §2a (GC-29..33). | +| 176+8·fl_count | to psize | reserved | — | MUST be 0; **not** covered by the CRC. | > **SPEC 04 interface note.** This layout fixes the fields the commit pipeline > reads/writes: `txnid`, `last_pg`, `map_size`, `free_db.root`, `main_db.root`. @@ -230,8 +232,14 @@ At env open the engine reads both slots and validates each independently: 4. Header `txnid` (offset 8) equals body `txnid` (offset 64), else the slot is inconsistent and is **discarded** (INV-2). This guards a torn write that updated one copy but not the other. -5. `meta_crc` matches the recomputed CRC32C over `[0, 168)` (§3.3). A slot that - fails the CRC is **torn** and is discarded. +5. `fl_count ≤ (psize − 176) / 8` (the annex fits the page; checked **before** + the CRC so a hostile count cannot drive an out-of-bounds CRC read), else + the slot is discarded; then `meta_crc` matches the recomputed CRC32C over + `[0, 172) ∪ [176, 176 + 8·fl_count)` (§3.3). A slot that fails either is + **torn** and is discarded. (The annex *ids*' ordering/range are validated + where they are consumed — the writer's parse and the checker, SPEC 05 + GC-33 — not here; a CRC-valid slot with hostile ids must fail typed at + draw, exactly like a hostile tree PIL.) This numbered list is the **single owner** of the meta-slot validation predicate; SPEC 06 REC-1 references it rather than restating it. @@ -273,14 +281,21 @@ Selection among the *CRC-valid* slots: ### §3.3 — CRC32C coverage (exact byte range) `meta_crc` = CRC32C (Castagnoli, polynomial `0x1EDC6F41`, reflected input/ -output, init `0xFFFFFFFF`, final XOR `0xFFFFFFFF`) computed over **exactly the -first `META_CONTENT_LEN = 168` bytes of the meta page** — absolute offsets -`[0, 168)`. That range covers the common header (including the reserved bytes -18–31, which MUST be 0) and the meta body through the end of `main_db`. The -`meta_crc` field itself (offset 168) and the reserved tail (`[172, psize)`) are -**excluded**. Reserved bytes inside the covered range MUST be zero so the CRC is -deterministic. Implementation of CRC32C is ADR-0002 §D1 (software table in -Phase 1; ARMv8 `crc32c` instructions in Phase 3.9). +output, init `0xFFFFFFFF`, final XOR `0xFFFFFFFF`) computed over **the first +`META_CONTENT_LEN = 172` bytes of the meta page followed by the annex ids** — +absolute offsets `[0, 172) ∪ [176, 176 + 8·fl_count)`, as one CRC stream in +that byte order (ADR-0022, format v2; v1 covered `[0, 168)`). The covered +prefix spans the common header (including the reserved bytes 18–31, which +MUST be 0), the meta body through `main_db`, and `fl_count`; the `meta_crc` +field itself (offset 172) and the reserved tail past the annex are +**excluded**. Covering `fl_count` inside the main CRC is load-bearing: a torn +meta write cannot silently zero the annex (which would leak its free pages) +without tearing the slot as a whole — the torn-meta guarantee (REC-8) extends +to the annex. Validation reads `fl_count` *before* the CRC is checked, so it +is bounds-checked first (`fl_count ≤ (psize − 176) / 8`, else the slot is +invalid) to bound the CRC read. Reserved bytes inside the covered range MUST +be zero so the CRC is deterministic. Implementation of CRC32C is ADR-0002 §D1 +(software table in Phase 1; ARMv8 `crc32c` instructions in Phase 3.9). ### §3.4 — Env-creation protocol (both slots initialised, empty DB) @@ -291,7 +306,8 @@ Creating a new env writes **both** meta slots (pages 0 and 1) as valid, identica 2. Write slot 0 and slot 1 identically: `txnid = 0`, `magic`, `format_version`, `page_size`, `map_size`, `env_flags = 0`, `last_pg = 1` (pages 0 and 1 exist; no data page yet), `free_db` and `main_db` both **empty** - (`root = PGNO_INVALID`, all stats 0, `depth = 0`), fresh `meta_crc` on each. + (`root = PGNO_INVALID`, all stats 0, `depth = 0`), `fl_count = 0` (empty + free-list annex), fresh `meta_crc` on each. 3. `fsync(data)` so both slots are durable. 4. `fsync(parent dir)` so the data file's **directory entry** is durable (added 2026-07-22, issue #46): without it, a crash shortly after creation @@ -320,7 +336,7 @@ off 18 : 00 00 reserved0 off 20 : 00 00 00 00 checksum (data-page CRC; unused, 0) off 24 : 00 00 00 00 00 00 00 00 reserved (meta variant tail) off 32 : 5A 44 42 31 magic "ZDB1" -off 36 : 01 00 00 00 format_version = 1 +off 36 : 02 00 00 00 format_version = 2 off 40 : 00 10 00 00 page_size = 4096 (0x1000) off 44 : 00 00 00 00 env_flags = 0 off 48 : 00 00 10 00 00 00 00 00 map_size = 1 MiB (0x100000) @@ -336,8 +352,10 @@ off 152 : 00 00 00 00 00 00 00 00 main_db.entries = 0 off 160 : 00 00 main_db.depth = 0 (empty) off 162 : 00 00 main_db.flags = 0 off 164 : 00 00 00 00 main_db.leaf2_ksize = 0 -off 168 : meta_crc = CRC32C over bytes [0,168) -off 172 .. 4096 : 00 reserved tail (excluded from CRC) +off 168 : 00 00 00 00 fl_count = 0 (empty free-list annex) +off 172 : meta_crc = CRC32C over bytes [0,172) + (fl_count = 0: no annex ids follow) +off 176 .. 4096 : 00 reserved tail (excluded from CRC) ``` If the first commit (txn 1) then inserts one key, it allocates the root leaf at @@ -625,8 +643,11 @@ owns; the *page* format is exactly §2/§4): overflow run exactly like any large value (§5); there is no bespoke spill format. The in-tree ordering of the ids within the list is SPEC 05's choice. -The check tool treats GC entries as the authority for "free" in the -reachability-xor-freeness invariant (INV-10, INV-14). +Since format v2 (ADR-0022) the **newest** freeing-txn's PIL normally rides in +the live meta's free-list annex (§3, SPEC 05 §2a) instead of this tree; the +tree holds the spill/cold entries. The check tool treats GC entries **plus the +selected meta's annex ids** as the authority for "free" in the +reachability-xor-freeness invariant (INV-10, INV-14, INV-28). --- diff --git a/docs/SPEC/04-txn-mvcc.md b/docs/SPEC/04-txn-mvcc.md index 241bad1..f4035df 100644 --- a/docs/SPEC/04-txn-mvcc.md +++ b/docs/SPEC/04-txn-mvcc.md @@ -1004,10 +1004,10 @@ every hook; this section defines the steps and their ordering. Both write modes |---|------|-----------|----------------------------------| | C0 | Assert `child_count == 0`; compute freed-page list. | — | last committed meta `N−1` (unchanged). | | C1a | **catalog write-back** (M1.6, SPEC 02 §6.1): write each dirty named-DB working record back into the main tree as its `F_SUBDATA` catalog entry. Before `freelist_save` (matches LMDB's sub-DB flush order) so the pages this COWs/frees are captured by C1. | — | `N−1` (all changes still in the dirty set). | - | C1 | **freelist_save** (SPEC 05 §4): write this txn's freed pages into the GC DB, dirtying GC pages into the dirty set (loop-until-stable, SPEC 05 GC-11). | H0 | `N−1` (all changes still in dirty set, nothing written). | + | C1 | **freelist_save** (SPEC 05 §4): apply drain rewrites to the GC DB (dirtying GC pages into the dirty set, loop-until-stable, SPEC 05 GC-11) and place this txn's freed set — into the meta free-list annex when it fits (format v2, SPEC 05 §2a GC-32: stashed for C4, no tree write), else into the GC DB under `BE(N)`. | H0 | `N−1` (all changes still in dirty set, nothing written). | | C2 | **write dirty pages** to their pgnos (`pwrite` each dirty frame; memcpy-into-map under heap-staged WRITE_MAP, TXN-45a; a no-op per in-map frame under in-place WRITE_MAP, TXN-45b — those bytes were stored at allocation/edit time, under the same TXN-62 invariant). Not yet durable. | **H1** | `N−1` live; new pages sit in free/beyond-HWM slots the `N−1` tree does not reference (TXN-62). Partial/torn data pages are unreferenced garbage. | | C3 | **fsync(data)** — flush all data pages (skipped under `NO_SYNC`/`MAP_ASYNC`, SPEC 06 REC-6). | **H2** | `N−1` live; txn `N`'s data fully durable but unreferenced (no meta points at it). | - | C4 | **write meta** to slot `N&1` (the *older* slot), with `txnid = writer_txnid` and a fresh CRC (SPEC 02 §3.3). Not yet durable. | **H3** | Either `N−1` (meta `N` not yet reached disk, or reached but torn → CRC rejects it → older slot wins) or `N` (meta reached disk intact). Never a torn meta accepted. | + | C4 | **write meta** to slot `N&1` (the *older* slot), with `txnid = writer_txnid`, the C1-stashed free-list annex (format v2, SPEC 02 §3) and a fresh CRC covering both (SPEC 02 §3.3). Not yet durable. | **H3** | Either `N−1` (meta `N` not yet reached disk, or reached but torn → CRC rejects it → older slot wins) or `N` (meta reached disk intact). Never a torn meta accepted. | | C5 | **fsync(meta)** — flush the meta page (skipped under `NO_META_SYNC`/`NO_SYNC`: the meta is written but not fsynced this commit, SPEC 06 REC-6/REC-7/REC-9). | **H4** | txn `N` durable and live. | | C6 | Publish the new snapshot (TXN-18/TXN-19): swap the `Arc`, then `commit_point.store(writer_txnid, SeqCst)`; update `last_committed_txnid`; release the write mutex. | — | commit complete; new readers see `N`. | diff --git a/docs/SPEC/05-gc.md b/docs/SPEC/05-gc.md index 4f01bfb..4b52d4b 100644 --- a/docs/SPEC/05-gc.md +++ b/docs/SPEC/05-gc.md @@ -18,9 +18,10 @@ documented as-is (§8) — the redesign is Phase 3.1 (ADR-first), **not** here. Clean-room note: the fork's `mdb_page_alloc`, `mdb_find_oldest`, `mdb_freelist_save`, and the loose-page list were read to understand the *algorithm* (CLAUDE.md rule 4); the design below is ZeroDB's own. Normative rules -are a single clean **GC-1 … GC-28** namespace (contiguous ids, no split-id -suffixes; GC-28 is the extended-page durability rule, placed with the allocation -topic in §5 though it is the highest id — see §10); new check-tool invariants +are a single clean **GC-1 … GC-33** namespace (contiguous ids, no split-id +suffixes; GC-28 is the extended-page durability rule, placed with the +allocation topic in §5 out of id order, and GC-29..33 are the format-v2 meta +free-list annex rules of §2a, ADR-0022 — see §10); new check-tool invariants continue SPEC 03's numbering from **INV-22**. This document **answers ADR-0002 open questions OQ1 and OQ2** (GC-2, GC-15, and @@ -96,6 +97,90 @@ writer_txnid − 1)`. --- +## §2a — The meta free-list annex (format v2, ADR-0022; GC-29..GC-33) + +**[Added 2026-10-03, ADR-0022 — implemented as a spike on relayed maintainer +approval; awaiting direct ratification.]** Since FORMAT_VERSION 2 the +**newest** freeing-txn's PIL normally lives in the committing meta page +itself (`fl_count`/`fl_ids`, SPEC 02 §3), not in the GC tree. The annex is a +pure *placement* change: its ids are exactly the PIL that v1 stored under +`BE(meta.txnid)`, with identical ordering, validation and gate rules. The GC +tree remains and holds the spill/cold entries; everything in §§4–6 continues +to govern it unchanged. + +- **GC-29 — annex = hoisted own-entry.** The annex of meta `N` is the PIL of + pages freed by txn `N` (freeing-txnid = the meta's own txnid; no separate + field): strictly ascending, unique, each in `[FIRST_DATA_PGNO, last_pg]` + (GC-3/GC-4 verbatim). `fl_count = 0` ⇔ txn `N` contributed no entry. A meta + with `fl_count > 0` has **no** GC-tree entry keyed `BE(N)` (see GC-32). +- **GC-30 — the carry (anti-leak across slots) + the in-save pool.** Meta + `N`'s slot is overwritten by txn `N + 2`, so the annex must never be the + only record of a free page once `N` stops being the base. Every + **committing** txn `W` (base `B = W − 1`) accounts for every live id of + `B`'s annex: ids drawn during the txn are in `reclaimed` (reused, or loose + if re-freed); the unconsumed remainder stays a live, gated **in-save + allocation pool** through `freelist_save` (a fourth GC-12 source, exactly + the drain-pool rule: its ids passed the gate with `F = B`), and the + step-(c) placement persists the **merge** `freed ∪ annex.live()` into + `W`'s annex or `W`'s tree entry — re-looping whenever the step-(c) put + itself drew from the pool, so no persisted set ever lists a handed-out + page. *Why a pool and not an upfront fold:* ids folded into `freed` are + un-allocatable in-save (they are pending free-list content), so an + upfront fold starves the save's own GC-tree allocations into `extend` on + every commit while the carried surplus grows in lockstep — an unbounded + file ratchet, observed as a `churn_parity_general` boundedness failure + during the ADR-0022 spike. A txn that commits nothing (the + unchanged-commit short-circuit) writes no meta, so `B` stays the base and + its annex stays live. Consequence: the INV-22 partition — with the annex + counted as free (INV-28) — holds for **every** committed meta. The carry + re-tags the persisted remainder under `W` (freeing-txnid moves forward); + that only ever *delays* reclaim eligibility, so the gate proofs are + unaffected. GC-13 gains a fourth bounded quantity: the annex pool is + loaded once at begin and never refilled, and every in-save draw strictly + shrinks it. +- **GC-31 — annex draw gate.** The annex pool is drawable iff + `B ≤ oldest_reader()` (TXN-20/21 verbatim with `F = B`): a reader pinned at + `B` cannot reach pages freed *by* `B` (they left `B`'s trees), and the + crash-fallback meta for writer `W` is `B` itself, whose annex still lists + the drawn ids as free and whose trees do not reference them — the TXN-62 / + in-save pool-draw argument (GC-12) verbatim. +- **GC-32 — all-or-nothing placement.** At `freelist_save`, if the final + deduped `freed` set fits the annex cap (`(psize − 176) / 8` ids) it is + **stashed for the meta encode** and GC-11 step (c) performs **no** tree + put; otherwise the whole set goes to the tree under `BE(W)` as in v1 and + the annex is empty. One freeing-txn's pages are never split between the + two. Once a save has taken the tree arm it keeps it for that save (a later + drain-rewrite shrinking `freed` must not flip the placement mid-loop). + The stash encodes no tree write, allocates nothing and frees nothing, so + the GC-11 fixed point loses its step-(c) feedback edge in the annex arm + and the GC-13 termination argument is unchanged (it only removes an edge). +- **GC-33 — never trusted.** Annex ids are on-disk data: `fl_count` is + bounds-checked and the ids are covered by the meta CRC at slot validation + (SPEC 02 §3.2 rule 5 / §3.3), and `validate_pil_ids` (ascending + range, + the GC-18 validation rule) runs when the writer parses the base annex — + before any id can be handed out. A violating annex fails the parse with + `MdbError::Invalid` (corrupt freelist), exactly like a hostile tree PIL. + +**Allocation integration** (amends GC-16): the annex pool is step **2b**, +tried after the GC-tree draw and before extend. `B` is the newest possible +freeing-txnid (every tree key is `≤ B`, and `= B` only when `B` spilled, in +which case its annex is empty — GC-32), so tree-first keeps the GC-18/19 +oldest-first scan order and single-page determinism. Within the pool, draws +are GC-19-deterministic: smallest id for `n = 1`, first contiguous run for +`n > 1`. Drawn ids enter `reclaimed`. Inside `freelist_save` (GcSave mode) +the pool remains drawable as GC-12 source (4) — after the loose list and +the drain pool, before extend — under the same gate; GC-30 owns the +re-accounting that keeps the persisted set draw-free. + +**Steady-state effect** (the ADR-0022 motivation): with no parked readers and +small per-commit frees, the GC tree is empty, `freelist_save` performs zero +tree operations (no GC-leaf COW, no descent, no extra C2 page), and the next +txn's reclaim is an O(1) pool draw — the B12 free-list-save and GC-scan costs +collapse to the annex encode (~8 bytes/id inside the meta CRC/write that +every commit already pays). + +--- + ## §3 — Loose pages (allocated and freed within one txn) - **GC-7** — A page this txn both **dirtied/allocated and then freed** (e.g. a leaf @@ -140,9 +225,16 @@ in-loop allocator so the loop cannot leak GC entries. condition.]** This skeleton is normative — an implementation must be reconstructible from it alone: + **[Amended 2026-10-03, ADR-0022/GC-30/GC-32: the base annex remainder + stays a live in-save pool (GC-12 source 4); step (c) persists the merge + `freed ∪ annex.live()` — to the meta annex (no tree put) whenever the + merged set fits the cap (all-or-nothing, sticky per save), else to the + tree as in v1 — and re-loops if the put drew from the pool.]** + ``` freelist_save(txn): alloc_mode = GC_SAVE # GC-12 scope: whole procedure + # base_annex.remainder stays drawable (GC-30 pool); merged in step (c) pending = keys(txn.drains) # every ops-drained entry needs # its GC-20 rewrite at least once # (b) trailing-loose shrink FIRST (GC-10): release never-written @@ -167,14 +259,28 @@ in-loop allocator so the loop cannot leak GC entries. if remaining.is_empty(): drains.remove(F); delete(FREE_DBI, BE(F)) else: put(FREE_DBI, BE(F), encode(remaining)) - # (c) this txn's own entry under key = BE(writer_txnid). + # (c) this txn's own entry: the merge freed ∪ annex.live() goes to + # the meta annex when it fits (GC-32), else to the tree under + # key = BE(writer_txnid) as in v1. sort_dedup(freed_pgs) # GC-3/GC-4 before = freed_pgs.len() - if before > 0: - put(FREE_DBI, BE(writer_txnid), encode(freed_pgs), RESERVE) # GC-2 - if freed_pgs.len() != before: - continue # the put COW-freed committed - # GC pages: PIL must grow + annex_before = annex.live().len() + if before + annex_before > 0: + if not tree_arm and before + annex_before <= annex_cap: + annex_mode = true # merged after the loop; no + # tree write, cannot free or + # allocate anything + else: + tree_arm = true # sticky for this save (GC-32) + annex_mode = false + ids = merge(freed_pgs, annex.live()) + put(FREE_DBI, BE(writer_txnid), encode(ids), RESERVE) # GC-2 + if freed_pgs.len() != before or annex.live().len() != annex_before: + continue # the put COW-freed committed + # GC pages (PIL must grow) or + # drew from the annex pool + # (written PIL lists a + # handed-out page) # FIXED POINT = all three, together: # 1. no pending rewrite (no drain-pool draw escaped a re-rewrite) # 2. freed_pgs stable across the put (the write freed nothing new) @@ -293,6 +399,15 @@ in-loop allocator so the loop cannot leak GC entries. if let Some(run) = gc_reclaim(n, oldest_reader()): return run + # 2b. The base meta's free-list annex pool (format v2, §2a GC-31): + # gated by base_txnid <= oldest_reader(); GC-19-deterministic + # within the pool. Tried AFTER the tree: the annex's freeing-txn + # is the newest possible, so tree-first preserves oldest-first. + # (Inside freelist_save the pool stays drawable as GC-12 source + # (4), after loose and the drain pool — GC-30.) + if let Some(run) = annex_pool_draw(n, oldest_reader()): + return run + # 3. Extend the file: hand out [next_pgno .. next_pgno+n), then # next_pgno += n. Fails MapFull per GC-17. return extend(n) @@ -420,6 +535,9 @@ in-loop allocator so the loop cannot leak GC entries. total += pil.count # the u64 count prefix (GC-3) return total ``` + Since format v2 the snapshot's meta annex ids are free pages too: + `free_page_count` adds the snapshot's `fl_count` (carried on the published + snapshot, so no meta re-read) to the tree walk's total. It runs under an ordinary read snapshot (SPEC 04 §3), counting the **sum of PIL counts**, not the number of GC entries. Each entry contributes only its 8-byte `count` prefix (`pil_count`): the walk reads the prefix and checks it against the @@ -492,6 +610,12 @@ The check tool (`zerodb-tools check`, M1.12) treats the GC DB as the authority f exceeds the high-water. - **INV-26** — **PIL sorted & counted.** Each PIL's `count` prefix equals the number of ids that follow, and the ids are strictly ascending (GC-3/GC-4). +- **INV-28** — **Annex accounting (format v2, ADR-0022).** The selected + meta's annex ids join the free set for INV-22/24/25 (reachable XOR free, + listed once across annex ∪ all tree PILs, in range), the annex is strictly + ascending and `fl_count`-consistent (the INV-26 analogue; the length half + is the SPEC 02 §3.2 rule-5 bound), and `fl_count > 0` implies no GC-tree + entry keyed `BE(meta.txnid)` (GC-29/GC-32). - **INV-27** — **`non_free_pages_size` consistency** (amended 2026-09-29). Without `WRITE_MAP`, the pages of a committed image partition exactly into the two meta slots, user-tree pages, GC-tree pages and free pages, so @@ -517,9 +641,11 @@ The check tool (`zerodb-tools check`, M1.12) treats the GC DB as the authority f | coalescing / fragmentation | GC-21 | PLAN 1.5 tolerance | | disk-usage APIs | GC-22..24, INV-27 | SPEC 00 rows 18/19 | | huge-txn weakness | GC-25..27 | PLAN 3.1 (ADR-first) | -| check invariants | INV-22..27 | SPEC 03 §11 | +| meta free-list annex (format v2) | GC-29..33, INV-28 | SPEC 02 §3/§3.3, ADR-0022 | +| check invariants | INV-22..28 | SPEC 03 §11 | -**Rule count: GC-1 … GC-28 (28 normative rules, contiguous ids, no split-id -suffixes) + INV-22 … INV-27 (6 new check invariants).** OQ1 answered by **GC-2** +**Rule count: GC-1 … GC-33 (33 normative rules, contiguous ids, no split-id +suffixes; GC-29..33 are the format-v2 annex rules of §2a, ADR-0022) + +INV-22 … INV-28 (7 check invariants).** OQ1 answered by **GC-2** (big-endian GC keys); OQ2 answered by **GC-15** (track `next_pgno`, persist `last_pg`; no meta-field change). Both recorded in the ADR-0002 amendment. diff --git a/docs/SPEC/06-recovery.md b/docs/SPEC/06-recovery.md index bda138e..7b69039 100644 --- a/docs/SPEC/06-recovery.md +++ b/docs/SPEC/06-recovery.md @@ -29,9 +29,11 @@ this section defines the **recovery decision** and its error taxonomy. - **REC-1** — **Validate each slot independently.** The validation predicate is owned by **SPEC 02 §3.2** (the numbered list: `magic`, `format_version`, - `page_size` ∈ {power of two, 4096–65536}, header `txnid` == body `txnid`, and — - mandatory — `meta_crc` over `[0,168)`, SPEC 02 §3.3). REC-1 does **not** restate - the list; it references SPEC 02 §3.2 as the single owner. A slot failing any + `page_size` ∈ {power of two, 4096–65536}, header `txnid` == body `txnid`, the + annex-count bound, and — mandatory — `meta_crc` over + `[0,172) ∪ [176, 176 + 8·fl_count)`, SPEC 02 §3.3 as amended by ADR-0022). + REC-1 does **not** restate the list; it references SPEC 02 §3.2 as the + single owner. A slot failing any check is **invalid** (torn or foreign) and is discarded from selection. - **REC-1a** — **Geometry validation of the selected slot** (added 2026-09-09, security review H1; predicate owned by SPEC 02 §3.2 step 6). A CRC-valid slot @@ -182,6 +184,19 @@ open, REC-1/REC-2 select snapshot `X` and the check tool (SPEC 03 §11 + SPEC 05 Phase 3.9); it is **not** claimed to detect every torn meta — sector-aligned tears are caught by the double-buffer/txnid design, not the CRC. + **Format-v2 annex note (ADR-0022).** The CRC-covered region is + `[0,172) ∪ [176, 176 + 8·fl_count)` and may extend past the first 512-byte + sector when the free-list annex is large (`fl_count > 42` at 512-byte + sectors). The sub-sector argument is unchanged; additionally, a + **sector-aligned** tear that lands *inside* the covered annex now leaves the + CRC inconsistent and the slot is rejected — which is required, not + incidental: a half-written (or silently zeroed) annex would leak that + commit's freed pages while the rest of the slot looked complete. The + "complete old or complete new, both CRC-valid" outcome of a sector-aligned + tear therefore holds only when the covered region fits sector 0; otherwise + the tear degrades to a rejected slot and the intact other slot wins — + strictly more conservative, never less. + --- ## §3 — Durability flags: crash windows (SPEC 01 §S6 lattice) diff --git a/docs/adr/0022-meta-freelist-annex.md b/docs/adr/0022-meta-freelist-annex.md new file mode 100644 index 0000000..6502732 --- /dev/null +++ b/docs/adr/0022-meta-freelist-annex.md @@ -0,0 +1,194 @@ +# ADR-0022: Meta free-list annex — the per-commit freed PIL rides in the meta page + +- Status: **Spike** — implemented for measurement on maintainer approval + relayed 2026-10-02/03 (format-change + session constraints lifted for the + commit-CPU lever track). The approval reached this change through the + coordinating session, not directly from the maintainer; per CLAUDE.md rule 6 + this ADR still **awaits direct human ratification before merge**. Nothing is + merged or pushed; the spike exists so the bench server can measure the win. +- Milestone: perf track "non-copy per-commit CPU" lever #1 (cheaper free-list + save), PERF-GAP-VS-LMDB §B12; forward-looking toward PLAN 3.1. +- Date: 2026-10-03 +- Numbering note: ADR-0021 (WRITE_MAP in-place) lives on its own open branch; + this ADR takes 0022 to avoid colliding with it. + +## Context + +B12 (bench server, x86-64, 4 KiB, 20k single-put NO_SYNC commits) attributes +the largest single slice of ZeroDB's +3.1 µs commit-CPU gap vs the LMDB fork +to the free-list save: ~1.6 µs/commit. Local census re-measurement in this +worktree (macOS aarch64, 4 KiB DB pages, temporary sub-phase instrumentation) +confirms both the magnitude and the shape: + +- `freelist_save` total ≈ 1.2–1.5 µs/commit; +- per commit it executes ~2 `put_pil` + ~1 `delete_tree` on the GC tree + (steady state: this txn's own entry, one partial-drain GC-20 rewrite, one + fully-drained entry delete); +- of that, ~0.3 µs is one real COW touch of the GC leaf (frame + full-page + copy), ~0.8–0.9 µs is the generic tree-op machinery of the three ops, + ~0.1–0.2 µs is encode/sort/descent (descent is *not* the problem: the GC + tree is one leaf; `search_path` is ~90 ns total), and the GC-11 loop runs + exactly 1.00 iterations/commit; +- the save also dirties the GC leaf, so commit C2 writes one extra page per + commit (ZeroDB 5.93 pages vs LMDB 4.96 in B12). + +The format-stable ceiling is low. LMDB's own save does the same three +B-tree ops on the same flat txnid→PIL format and costs ~0.5–0.7 µs; matching +it means shaving generic per-op overhead (separate levers: touch cost B12-1, +lazy validation A2(b)) and buys at most ~0.5 µs here. The three tree ops per +commit are *inherent to the format*: the freed list of txn N must be durable +and crash-consistent with meta N, and in the flat-freelist format the only +place for it is the GC tree, whose mutation is itself a COW tree write. + +With the format constraint lifted, the structural observation is: **the meta +page already is the one page every commit must write and CRC**. At 4 KiB the +meta's fixed content ends at byte 172; the remaining ~3.9 KiB are reserved +zeros. A typical commit frees a handful of pages (~5 in the census; a PIL of +c ids is 8c bytes). The freed list of txn N can therefore ride *inside* meta +N — written by the same `write_at_page`, covered by the same CRC, atomic with +the same meta — and the GC tree drops out of the per-commit path entirely. + +Prior art: LMDB has no equivalent (its meta is ~112 bytes of a page and it +keeps the freelist in FREE_DBI unconditionally — and pays the same tree-op +cost we do). libmdbx likewise keeps a GC table but batches differently. This +is a deliberate, documented **non-LMDB lever** (observable heed-level +semantics unchanged; allocation-order determinism preserved), in the same +spirit as the ratified rightmost-leaf finger (ADR-0015 note in rwtxn.rs): +LMDB's own technique for this path cannot go below LMDB's own cost, and the +B12 target is to *beat* that bucket, with PLAN 3.1 already pointing at a +freelist redesign. + +## Options + +### Option A — format-stable micro-optimisation (rejected) + +One search+touch for the whole save (all steady-state keys land in the same +GC leaf), direct `LeafMut` edits, kill the two per-save `Vec` clones +(`freed.clone()`, `remaining.to_vec()`). + +- Pros: no format change, no recovery surface. +- Cons: measured ceiling ~0.3–0.5 µs of the 1.6; duplicates leaf-edit logic in + the riskiest file; still COWs + rewrites a GC leaf and still writes the + extra page at C2; becomes dead code when 3.1 lands. + +### Option B — freed-PIL annex in the meta page (chosen) + +Meta format (FORMAT_VERSION 1 → 2; offsets `[0, 168)` unchanged): + +| Off | Size | Name | Meaning | +|----:|-----:|------|---------| +| 168 | 4 | `fl_count` (u32) | Number of annex ids. `0 ≤ fl_count ≤ (psize − 176) / 8`. | +| 172 | 4 | `meta_crc` (u32) | CRC32C over `[0, 172) ∪ [176, 176 + 8·fl_count)`. | +| 176 | 8·c | `fl_ids` | LE u64 page ids, strictly ascending, each in `[FIRST_DATA_PGNO, last_pg]`. | +| rest | — | reserved | MUST be 0; not CRC-covered. | + +Semantics: the annex is **exactly the PIL that version 1 stored in the GC +tree under key `BE(meta.txnid)`** — the pages freed by that commit — hoisted +into the meta. Freeing-txnid = the meta's own txnid; no separate field. + +- **Write side (commit C1).** `freelist_save` keeps the *unconsumed + remainder* of the base meta's annex as a live, gated **in-save pool** (a + fourth GC-12 source, exactly the drain-pool rule — an upfront fold into + `freed` was the spike's first cut and ratcheted the file unboundedly, + because folded ids are un-allocatable in-save and the save's own GC-tree + ops then extend on every commit; caught by `churn_parity_general`). The + GC-11 loop for drain rewrites (GC-20) runs unchanged. Step (c) persists + the **merge** `freed ∪ annex.live()`: if it fits the annex cap it is + **stashed for the meta encode** and no tree put happens (the stash cannot + allocate or free pages, so the GC-11/13 fixed point is reached with no + step-(c) feedback); otherwise the whole merged list goes to the tree under + `BE(writer_txnid)` exactly as today and the annex is empty — re-looping if + the put itself drew from the pool, so the persisted set never lists a + handed-out page. All-or-nothing — one freeing-txn's pages are never split + between annex and tree. Once the tree arm has run in a save, it stays the + arm for that save (no flip-flop if a drain rewrite later shrinks `freed`). +- **Read side (allocation).** The writer parses the base meta's annex at txn + begin (one u32 read when empty; `validate_pil_ids` on the ids otherwise — + the freelist is still never trusted, GC-18). `allocate()` draw order + becomes: loose → GC tree (`gc_reclaim`, unchanged) → **annex pool** → + extend. The annex's freeing-txn (`base.txnid`) is the *newest* possible F, + so trying the tree first preserves the GC-18/19 oldest-first determinism. + The annex draw is gated exactly like any GC draw: `F ≤ oldest_reader()` + (here: no live reader pinned below base). Drawn ids enter `reclaimed`. + Inside `freelist_save` (GcSave mode) the annex pool is never consulted — + it was folded into `freed` before the loop, so GC-12/13 are untouched. +- **Crash safety.** The annex is CRC-covered by the meta it belongs to: a + torn meta write that corrupts the annex tears the whole slot, which the + §3.2 selection discards (fallback to N−1, whose own annex+tree state is + per-snapshot consistent). The TXN-62 argument for reusing annex ids is the + in-save pool-draw argument verbatim: an id in meta B's annex was freed *by* + txn B, is absent from B's trees, and B is the crash-fallback meta for the + writer W = B+1 — so overwriting it before W's meta lands is safe, and any + reader that pins after the draw pins `≥ B`. + +- Pros: steady-state save does **zero tree ops** (no GC-leaf COW, no descent, + no node edits, no extra C2 page); next txn's reclaim is an O(1) pool draw + instead of a cursor scan + PIL decode; CRC/encode cost grows by 8 bytes/id + (~40 B/commit typical). +- Cons: format break (version bump; pre-release, no consumers — approved); + recovery/check/tools must learn the annex; a new carry invariant to hold; + under a parked reader the carried list is re-encoded into every meta until + it exceeds the cap and spills (bounded by the cap, ≤ ~3.9 KiB at 4 KiB psize). + +### Option C — O_DSYNC-style / batching tricks + +Not applicable: the cost is CPU in tree ops, not flush count (B12 is NO_SYNC). + +## Decision + +Option B, as a **spike** (CLAUDE.md rule 7: format change ⇒ smallest change +that tests the riskiest assumption — here, that the annex removes the +measured CPU while recovery, INV-22 and bounded growth stay provable). +Always-on with the FORMAT_VERSION bump (an opt-in annex would double the +format test matrix for no consumer benefit; "parity by default" governs +API-visible behavior, which is unchanged). + +## Invariants (normative; SPEC 05 §2a / SPEC 02 §3 own the final wording) + +- **GC-29 (annex = hoisted own-entry).** The annex of meta N holds exactly + the ids version 1 would have written under `BE(N)`, strictly ascending, + unique, in `[FIRST_DATA_PGNO, last_pg]`; `fl_count = 0` means no entry. + A meta with `fl_count > 0` has **no** GC-tree entry keyed `BE(N)`. +- **GC-30 (carry + in-save pool).** Every committing txn W accounts for + every live id of its base's annex: drawn ids are in `reclaimed` (reused or + loose), and the remainder — drawable in-save as a gated pool — is merged + into the step-(c) persisted set (W's annex or W's tree entry). + Consequence: the per-image partition (INV-22, with the annex counted as + free — INV-28) holds for every committed meta. +- **GC-31 (gate).** Annex draws require `base.txnid ≤ oldest_reader()`; the + TXN-20/TXN-62 proofs apply unchanged with F = base.txnid. +- **GC-32 (all-or-nothing).** One freeing-txn's pages are never split between + its annex and a tree entry. +- **GC-33 (never trusted).** Annex ids pass `validate_pil_ids` before any id + is handed out, exactly like a tree PIL (GC-18 validation rule); `fl_count` + is bounds-checked and the ids CRC-checked at slot validation. +- **INV-28 (check).** The checker counts annex ids of the selected meta into + the free set: reachable XOR (tree-free ∪ annex-free), no double listing + (extends INV-22/24/25/26 to the annex). + +## Consequences + +- `FORMAT_VERSION` 2; version-1 files are rejected at open (`BadVersion`) — + pre-release, sanctioned. `migrate-from-lmdb` output moves to v2 implicitly + (it writes through the engine). +- SPEC 02 §3 (meta table, §3.3 CRC coverage, §3.5 example note), SPEC 05 + (new §2a + GC-16/18/23 amendments, GC-29..33, INV-28), SPEC 06 (REC-8 note: + the annex is inside the torn-meta guarantee) updated in this change. +- `Snapshot` carries the annex count so `free_page_count()` stays exact per + snapshot (GC-23 note). +- New tests: meta annex codec roundtrip + torn-annex rejection + cap bounds; + annex reclaim/carry/spill behavior (incl. reader-gate block); checker + annex accounting; crash harness runs unchanged on the new format. +- The census' `allocate` cost also drops in steady state (the GC cursor scan + disappears when the tree is empty). That is a *consequence* of this format + change, not the separate allocate lever being folded in. + +## Open questions for human review + +1. Ratify the format change itself (CLAUDE.md rule 6 — this ADR was + implemented as a spike on relayed approval; it needs your direct sign-off + before merge). +2. Cap policy: full `(psize − 176)/8` (chosen) vs a smaller policy cap to + bound the parked-reader re-encode cost earlier. +3. Whether PLAN 3.1 should absorb this as its first stage (the annex is the + hot tier of any future freelist redesign) or whether 3.1 supersedes it. From 8d6a2495d0b87197744cb926e9e3c7fd66a5eb19 Mon Sep 17 00:00:00 2001 From: Quentin de Quelen Date: Sat, 3 Oct 2026 15:39:08 +0200 Subject: [PATCH 2/5] ADR-0022: address spec-review follow-ups + record measured result MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Non-blocking items from the adversarial spec-review (no technical blockers were found): - encode_with_annex doc: name both callers (freelist_save, copy::write_raw) and the trust model — the raw-copy path forwards a snapshot annex validated only by open-time count+CRC, re-validated by the copy's own reader before any draw (no soundness hole; GC-33). - SPEC 05 GC-23: document the writer's mid-txn free_page_count granularity (annex = live remainder vs tree = full count prefixes) as a conservative working-state estimate; no consumer reads it mid-write-txn. - DIVERGENCES D-022 (PENDING): free-list placement is tools-observable (empty steady-state GC tree, smaller churn files); heed behaviour unchanged. - crc32c / crc32c_concat share one crc_update inner loop (drift guard, N1). - SPEC 02 §3.2 rule 5: pin that fl_count is bounded against the env's expected psize, not the slot's own page_size (N4). BadAnnexCount doc wording (N2). - ADR status: record the measured win, green gate, and clean review. Gate re-confirmed: fmt/clippy clean, cargo test --workspace 630/0. --- crates/zerodb-core/src/page/crc32c.rs | 24 ++++++++++++++---------- crates/zerodb-core/src/page/meta.rs | 20 ++++++++++++++++---- docs/DIVERGENCES.md | 1 + docs/SPEC/02-pages.md | 6 +++++- docs/SPEC/05-gc.md | 12 ++++++++++++ docs/adr/0022-meta-freelist-annex.md | 9 +++++++++ 6 files changed, 57 insertions(+), 15 deletions(-) diff --git a/crates/zerodb-core/src/page/crc32c.rs b/crates/zerodb-core/src/page/crc32c.rs index 4cc3362..ab8471f 100644 --- a/crates/zerodb-core/src/page/crc32c.rs +++ b/crates/zerodb-core/src/page/crc32c.rs @@ -41,18 +41,25 @@ const fn build_table() -> [u32; 256] { table } +/// Fold `bytes` into a running (non-finalized) CRC32C state. The single source +/// of the byte-wise table step, shared by [`crc32c`] and [`crc32c_concat`] so +/// the two cannot drift if the table or algorithm ever changes. +#[inline] +fn crc_update(mut crc: u32, bytes: &[u8]) -> u32 { + for &byte in bytes { + let idx = ((crc ^ byte as u32) & 0xFF) as usize; + crc = (crc >> 8) ^ TABLE[idx]; + } + crc +} + /// Compute the CRC32C (Castagnoli) checksum of `data`. /// /// Returns the finalized checksum (post final-XOR), i.e. the value written to a /// meta page's `meta_crc` field. #[must_use] pub fn crc32c(data: &[u8]) -> u32 { - let mut crc = 0xFFFF_FFFFu32; - for &byte in data { - let idx = ((crc ^ byte as u32) & 0xFF) as usize; - crc = (crc >> 8) ^ TABLE[idx]; - } - crc ^ 0xFFFF_FFFF + crc_update(0xFFFF_FFFF, data) ^ 0xFFFF_FFFF } /// CRC32C over the concatenation of `parts`, as one stream — equal to @@ -63,10 +70,7 @@ pub fn crc32c(data: &[u8]) -> u32 { pub fn crc32c_concat(parts: &[&[u8]]) -> u32 { let mut crc = 0xFFFF_FFFFu32; for part in parts { - for &byte in *part { - let idx = ((crc ^ byte as u32) & 0xFF) as usize; - crc = (crc >> 8) ^ TABLE[idx]; - } + crc = crc_update(crc, part); } crc ^ 0xFFFF_FFFF } diff --git a/crates/zerodb-core/src/page/meta.rs b/crates/zerodb-core/src/page/meta.rs index 41b4efa..36dad22 100644 --- a/crates/zerodb-core/src/page/meta.rs +++ b/crates/zerodb-core/src/page/meta.rs @@ -207,9 +207,19 @@ impl MetaPage { } /// Encode this meta into `buf` with `annex` as the free-list annex ids - /// (SPEC 02 §3 format v2, ADR-0022; SPEC 05 §2a). The caller guarantees - /// the GC-29 shape (strictly ascending, unique, in range) — `freelist_save` - /// produces exactly that; this encoder only enforces the capacity bound. + /// (SPEC 02 §3 format v2, ADR-0022; SPEC 05 §2a). This encoder enforces only + /// the capacity bound; it does **not** check the GC-29 shape (strictly + /// ascending, unique, in range), and it is sound for it not to, because the + /// shape is never trusted on the way *out* of a meta — it is re-checked by + /// `validate_pil_ids` before any id is handed out (SPEC 05 GC-33). The two + /// callers and their inputs: + /// - `RwTxn::freelist_save` → the commit path: produces a GC-29-shaped list + /// by construction, so what it encodes is well-formed. + /// - `copy::write_raw` → raw env copy: forwards the pinned snapshot's annex + /// verbatim. That annex passed only the open-time count+CRC bound, **not** + /// `validate_pil_ids`, so a hostile-but-CRC-valid source file is copied + /// faithfully (ids unchanged) — the copy is as trustworthy as its source + /// and no more, and the copy's own reader re-validates before any draw. /// /// # Errors /// @@ -371,7 +381,9 @@ pub enum MetaValidity { /// `page_size` is not a power of two in range (the observed value). BadPageSize(u32), /// `fl_count` exceeds the page's annex capacity (SPEC 02 §3.2 rule 5, - /// format v2) — the slot is invalid (torn or hostile). + /// format v2) — the slot is **invalid and discarded** (as §3.3 / REC-8 + /// word it). The cause is either a torn write or a hostile count; both are + /// rejected identically, before the CRC read, so the name covers both. BadAnnexCount(u32), /// Header txnid and body txnid disagree — a torn write (INV-2). TxnidMismatch { diff --git a/docs/DIVERGENCES.md b/docs/DIVERGENCES.md index 2914a10..5550085 100644 --- a/docs/DIVERGENCES.md +++ b/docs/DIVERGENCES.md @@ -33,6 +33,7 @@ Rules: | D-019 | Page-validation policy (ADR-0014) | LMDB never validates page contents: cell pointers, key sizes, value spans and overflow references are used as read (`mdb_page_get` checks only `pgno < mt_next_pgno`; the descent checks only that the bottom page is a leaf). A corrupt page is undefined behaviour. | ZeroDB validates by default (a corrupt page is a typed error) and adds an **opt-in** trusting policy, `FileTrust::trust_contents()` (an `unsafe` constructor passed to the safe `EnvOpenOptions::file_trust` setter), under which map pages skip the per-cell walk and behave like LMDB's. Page-number bounds, page type, meta validation and free-list id checks stay on in both policies. Overflow runs differ: the validating default checks each run's header; the trusting policy reads an overflow value without its run header (as `mdb_node_read` does), keeping only the snapshot high-water bound, so a corrupt run can return wrong bytes but never read outside the committed map (ADR-0014 amendment, 2026-09-29). A ZeroDB extension, not a behaviour change: the default is unchanged and no consumer sees the option unless it calls it. | APPROVED — opt-in, default validating | Quentin de Quelen (chat, 2026-09-26: "for the 'trusting the file' it would be great if it could be an option"; 2026-09-28: "do the read path now") | | D-020 | Sequential-writes fast path (ADR-0015) | `mdb_put` sets up a fresh cursor and descends from the root on every call; `MDB_APPEND` re-descends through `mdb_cursor_last`. | An **opt-in** env option (`EnvOpenOptions::sequential_writes`, default off) with a per-database runtime override (`Env::set_sequential_writes`) lets a write txn remember each tree's rightmost-leaf path and append there without descending when a key sorts after every key in that leaf (SPEC 03 §6.6). Results and committed files are identical with it on or off; only speed changes (ascending/APPEND loads faster, random-key writes 3–5 % slower). A ZeroDB extension, off unless called. | APPROVED — opt-in, default off | Quentin de Quelen (chat, 2026-09-28: "can the sequential writes be an option of the env"; env option plus per-database override: "perfect") | | D-021 | Configurable dirty limit (ADR-0017) | The spill threshold is a compile-time constant (`MDB_IDL_UM_MAX`, 131,072 pages); `mdb_page_spill` writes 1/8 of the dirty list once the txn's dirty room runs low. | ZeroDB spills the same way with the same default (SPEC 04 §6.3a) and adds an **opt-in** env option, `EnvOpenOptions::max_dirty_bytes`, to set the limit in bytes. A ZeroDB extension, not a behaviour change: unset, the limit is LMDB's, and results and committed files never depend on it. | APPROVED — opt-in, default LMDB's | Quentin de Quelen (chat, 2026-09-30: "go continue, I should be closer to LMDB in memory usage") | +| D-022 | Free-list placement (ADR-0022, meta annex) | The freed-page list of every commit is written into FREE_DBI (the GC B-tree) under `BE(txnid)` unconditionally; the GC tree is never empty under churn and each commit COWs a GC leaf. | Each commit's freed list rides inside its meta page's reserved bytes (the "annex"); the GC tree stays empty in steady state and is used only when a commit's freed set exceeds the annex cap or a reader parks a carried list. **On-disk format only** (`FORMAT_VERSION` 1→2); heed-level behaviour, allocation-order determinism and all query results are identical. Tools-observable (`zerodb stat` shows an empty FREE_DBI in steady state; churn files are ~1 page/commit smaller). Not a heed/LMDB *behaviour* divergence — filed here only for the format paper trail. | **PENDING — ADR-0022 is a spike awaiting direct ratification of the format change (CLAUDE.md rule 6); not merged** | (unratified as of 2026-10-03) | ### M1.13 adapter-boundary re-impositions (2026-07-17) diff --git a/docs/SPEC/02-pages.md b/docs/SPEC/02-pages.md index 19baffd..968e9d2 100644 --- a/docs/SPEC/02-pages.md +++ b/docs/SPEC/02-pages.md @@ -233,7 +233,11 @@ At env open the engine reads both slots and validates each independently: inconsistent and is **discarded** (INV-2). This guards a torn write that updated one copy but not the other. 5. `fl_count ≤ (psize − 176) / 8` (the annex fits the page; checked **before** - the CRC so a hostile count cannot drive an out-of-bounds CRC read), else + the CRC so a hostile count cannot drive an out-of-bounds CRC read). Here + `psize` is the **env's expected page size** (the one validation is called + with), not the slot's own `page_size` field — using the expected size is + conservative even when a hostile slot claims a larger `page_size`, since the + CRC read stays within the buffer the env actually mapped. Else the slot is discarded; then `meta_crc` matches the recomputed CRC32C over `[0, 172) ∪ [176, 176 + 8·fl_count)` (§3.3). A slot that fails either is **torn** and is discarded. (The annex *ids*' ordering/range are validated diff --git a/docs/SPEC/05-gc.md b/docs/SPEC/05-gc.md index 4b52d4b..10812f6 100644 --- a/docs/SPEC/05-gc.md +++ b/docs/SPEC/05-gc.md @@ -538,6 +538,18 @@ in-loop allocator so the loop cannot leak GC entries. Since format v2 the snapshot's meta annex ids are free pages too: `free_page_count` adds the snapshot's `fl_count` (carried on the published snapshot, so no meta re-read) to the tree walk's total. + + *Writer working-state note (format v2).* Called on the live write txn before + its `freelist_save`, the two free-page sources report at **different + granularities**: the base annex contributes its *live remainder* (ids already + drawn this txn are excluded immediately, since a draw removes them from the + in-save pool), while the GC tree still contributes each drained entry's full + `count` prefix until the save rewrites that entry. The value is therefore a + conservative working-state estimate, not the post-commit free count; the exact + per-snapshot value is the read-txn definition above, evaluated against a + committed meta. No consumer reads `free_page_count` mid-write-txn (Meilisearch + uses `non_free_pages_size` under a read snapshot, GC-24); the drift is + documented, not relied upon. It runs under an ordinary read snapshot (SPEC 04 §3), counting the **sum of PIL counts**, not the number of GC entries. Each entry contributes only its 8-byte `count` prefix (`pil_count`): the walk reads the prefix and checks it against the diff --git a/docs/adr/0022-meta-freelist-annex.md b/docs/adr/0022-meta-freelist-annex.md index 6502732..efe60c8 100644 --- a/docs/adr/0022-meta-freelist-annex.md +++ b/docs/adr/0022-meta-freelist-annex.md @@ -6,6 +6,15 @@ coordinating session, not directly from the maintainer; per CLAUDE.md rule 6 this ADR still **awaits direct human ratification before merge**. Nothing is merged or pushed; the spike exists so the bench server can measure the win. +- Result (2026-10-03): **measured win, gate green, spec-review clean.** Bench + server 3-column A/B (x86-64, turbo off, 5 interleaved rounds, CODEGEN_UNITS=1, + BASE = main 84582e8): `commit/batch/n1` 2.04× → **1.60×** LMDB (−22%, + after÷before 0.780), `commit/batch/n100` 1.09× → 1.02×; `n10k` and both sync + rungs flat (overhead amortized / fsync-bound — as predicted). Full gate: test + 630/0, miri 0-fail, crash-test-quick all durability modes, loom 9/0, stress + 180s 2/0, fuzz-quick 2.1M clean. Adversarial spec-review found no technical + blockers; its should-fix/nit items are addressed in the branch. Open for human + ratification (questions 1–3 below). - Milestone: perf track "non-copy per-commit CPU" lever #1 (cheaper free-list save), PERF-GAP-VS-LMDB §B12; forward-looking toward PLAN 3.1. - Date: 2026-10-03 From 0dadf38c5091e2b084b96b1621ab20d01e17428b Mon Sep 17 00:00:00 2001 From: Quentin de Quelen Date: Mon, 5 Oct 2026 11:23:28 +0200 Subject: [PATCH 3/5] ADR-0022: WRITE_MAP in-place twins for the annex battery + doc fixes MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Add WRITE_MAP in-place twins of steady_state_annex_only_gc_tree_stays_empty and reader_gate_blocks_annex_draw_then_carry_reclaims (the parked-reader case), following the ADR-0021 M3 parameterized-body convention from nested_fanout.rs/put_reserved_adversarial.rs: an `assert_mode` tripwire on `dirty_in_map_mode()` plus a growth-metric tripwire. The growth metric needed a WRITE_MAP-aware rework: file_size is pinned to map_size from the first commit under WRITE_MAP (the file is set_len'd at map time), so it can never show the extend-vs-reuse distinction the reader-gate test depends on. Added `high_water()`, which falls back to the meta's logical `last_pg` high-water mark under WRITE_MAP — this surfaced and fixed a false "gate leaked" failure in the twin that was purely a test-metric artifact, not an engine bug (confirmed against zerodb-core::env's own WRITE_MAP set_len(map_size) comment). Skipped the optional annex-draw counter in the crash harness (item 2): no counter for annex draws during freelist_save exists today, unlike dirty_in_map_mode (already a test hook) or the fault journal's map_regions (already tracked by the broker) — adding one would mean new public API/ atomics on RwTxn, which is out of scope here. Doc fixes: docs/adr/0022-meta-freelist-annex.md's Read-side paragraph wrongly claimed the annex pool is never consulted inside freelist_save; corrected to match the Write-side paragraph, SPEC 05 GC-30, and rwtxn.rs's AllocMode::GcSave arm (annex_draw is the carried-pool source after save_pool_draw). SPEC 05 GC-31 now scopes the "crash-fallback meta for writer W is B itself" claim to the durable modes, cross-referencing SPEC 06 for the REC-10/REC-11 window and the in-place WRITE_MAP case where an uncommitted/aborted txn's in-map writes can reach disk without a commit step. Co-Authored-By: Claude Opus 5.5 --- crates/zerodb/tests/meta_annex_gc.rs | 124 +++++++++++++++++++++++---- docs/SPEC/05-gc.md | 9 +- docs/adr/0022-meta-freelist-annex.md | 11 ++- 3 files changed, 122 insertions(+), 22 deletions(-) diff --git a/crates/zerodb/tests/meta_annex_gc.rs b/crates/zerodb/tests/meta_annex_gc.rs index 70c1be4..940f25b 100644 --- a/crates/zerodb/tests/meta_annex_gc.rs +++ b/crates/zerodb/tests/meta_annex_gc.rs @@ -7,7 +7,7 @@ use std::path::{Path, PathBuf}; use std::sync::atomic::{AtomicU64, Ordering}; -use zerodb::{check, free_page_count, CompactionOption, CopyToFile, Env, EnvOpenOptions}; +use zerodb::{check, free_page_count, CompactionOption, CopyToFile, Env, EnvFlags, EnvOpenOptions}; use zerodb_core::page::{MetaPage, MetaValidity, PGNO_INVALID}; static COUNTER: AtomicU64 = AtomicU64::new(0); @@ -38,13 +38,27 @@ impl Drop for TempDir { const PS: u32 = 4096; const MAP: usize = 32 << 20; -fn open(dir: &Path) -> Env { +fn open(dir: &Path, flags: EnvFlags) -> Env { let mut opts = EnvOpenOptions::new(); opts.map_size(MAP); opts.page_size(PS); + opts.flags(flags); opts.open(dir).expect("open env") } +/// Non-vacuousness tripwire for the WRITE_MAP twins (ADR-0021/ADR-0022 +/// convention, see `nested_fanout.rs::assert_mode`): the txn's dirty-page +/// realization must match the env flags it was opened with, so a twin can +/// never silently run the default heap-staged path instead of in-place +/// WRITE_MAP. +fn assert_mode(wtxn: &zerodb::RwTxn<'_>, flags: EnvFlags) { + assert_eq!( + wtxn.dirty_in_map_mode(), + flags.contains(EnvFlags::WRITE_MAP), + "dirty-page realization does not match the env flags" + ); +} + fn data_file(dir: &Path) -> PathBuf { dir.join(zerodb::DATA_FILE_NAME) } @@ -59,6 +73,43 @@ fn file_size(dir: &Path) -> u64 { std::fs::metadata(data_file(dir)).unwrap().len() } +/// Growth metric that stays meaningful under both dirty-page realizations. +/// The default backing grows the file by exactly what each commit needs, so +/// `file_size` tracks page reuse vs. extension precisely. Under `WRITE_MAP` +/// the file is `set_len`d to the full `map_size` at map time (see +/// `zerodb-core::env`'s read-annex comment on the geometry-validation length +/// check), so `file_size` is pinned from the first commit on and can never +/// show growth — the meta's logical page high-water mark (`last_pg`) is the +/// mode-independent stand-in: it only advances when a commit extends past +/// the previous high water, and stays put when a commit's allocations are all +/// satisfied by annex/pool reuse. +fn high_water(dir: &Path, flags: EnvFlags) -> u64 { + if !flags.contains(EnvFlags::WRITE_MAP) { + return file_size(dir); + } + let bytes = std::fs::read(data_file(dir)).unwrap(); + let ps = PS as usize; + let pick = |b: &[u8]| match MetaPage::validate(b, PS).unwrap() { + MetaValidity::Valid(m) => Some(m), + _ => None, + }; + let m0 = pick(&bytes[..ps]); + let m1 = pick(&bytes[ps..2 * ps]); + let m = match (m0, m1) { + (Some(a), Some(b)) => { + if a.txnid >= b.txnid { + a + } else { + b + } + } + (Some(a), None) => a, + (None, Some(b)) => b, + (None, None) => panic!("no valid meta slot"), + }; + m.last_pg +} + /// Read the LIVE meta's `(txnid, fl_count, free_db_root)` straight from the /// data file (both slots validated, higher txnid wins — SPEC 02 §3.2). fn live_meta(dir: &Path) -> (u64, u32, u64) { @@ -90,13 +141,30 @@ fn live_meta(dir: &Path) -> (u64, u32, u64) { /// file size flat. #[test] fn steady_state_annex_only_gc_tree_stays_empty() { + steady_state_annex_only_gc_tree_stays_empty_inner(EnvFlags::EMPTY); +} + +/// WRITE_MAP in-place twin (ADR-0021 M3 convention): the same steady-state +/// churn, but with dirty pages realized in the writable map. The annex draw +/// (ids reused from the meta's loose pool, not fresh-extended pages — see the +/// flat `file_size` assertions below) is orthogonal to WRITE_MAP's dirty-page +/// realization, so it must behave identically; `assert_mode` is the tripwire +/// that the twin actually engaged in-place writes instead of silently running +/// the default heap-staged path. +#[test] +fn steady_state_annex_only_gc_tree_stays_empty_writemap_in_place() { + steady_state_annex_only_gc_tree_stays_empty_inner(EnvFlags::WRITE_MAP); +} + +fn steady_state_annex_only_gc_tree_stays_empty_inner(flags: EnvFlags) { let dir = TempDir::new(); - let env = open(dir.path()); + let env = open(dir.path(), flags); let db = env.main_database(); let val = vec![0xABu8; 256]; for i in 0u64..60 { let mut w = env.write_txn().unwrap(); + assert_mode(&w, flags); db.put(&mut w, &i.to_be_bytes(), &val).unwrap(); w.commit().unwrap(); } @@ -115,16 +183,17 @@ fn steady_state_annex_only_gc_tree_stays_empty() { db.put(&mut w, &i.to_be_bytes(), &val).unwrap(); w.commit().unwrap(); } - let size_settled = file_size(dir.path()); + let size_settled = high_water(dir.path(), flags); for i in 0u64..50 { let mut w = env.write_txn().unwrap(); db.put(&mut w, &i.to_be_bytes(), &val).unwrap(); w.commit().unwrap(); } assert_eq!( - file_size(dir.path()), + high_water(dir.path(), flags), size_settled, - "annex reuse must keep overwrite churn at zero file growth" + "annex reuse must keep overwrite churn at zero growth (file size \ + under the default backing, last_pg under WRITE_MAP)" ); assert_clean(dir.path()); @@ -140,8 +209,22 @@ fn steady_state_annex_only_gc_tree_stays_empty() { /// releases, the carried ids are reclaimed and growth stops. #[test] fn reader_gate_blocks_annex_draw_then_carry_reclaims() { + reader_gate_blocks_annex_draw_then_carry_reclaims_inner(EnvFlags::EMPTY); +} + +/// WRITE_MAP in-place twin (ADR-0021 M3 convention) of the parked-reader +/// case: the GC-31 gate and GC-30 carry are meta free-list bookkeeping, +/// independent of whether dirty pages are realized in the writable map, so +/// the same growth-then-settle shape must hold. `assert_mode` tripwires that +/// the twin really ran in-place. +#[test] +fn reader_gate_blocks_annex_draw_then_carry_reclaims_writemap_in_place() { + reader_gate_blocks_annex_draw_then_carry_reclaims_inner(EnvFlags::WRITE_MAP); +} + +fn reader_gate_blocks_annex_draw_then_carry_reclaims_inner(flags: EnvFlags) { let dir = TempDir::new(); - let env = open(dir.path()); + let env = open(dir.path(), flags); let db = env.main_database(); let val = vec![0x44u8; 256]; @@ -155,20 +238,22 @@ fn reader_gate_blocks_annex_draw_then_carry_reclaims() { let pin = env.read_txn().unwrap(); { let mut w = env.write_txn().unwrap(); + assert_mode(&w, flags); db.put(&mut w, &0u64.to_be_bytes(), &val).unwrap(); w.commit().unwrap(); } - let size_pinned_base = file_size(dir.path()); + let size_pinned_base = high_water(dir.path(), flags); for i in 0u64..6 { let mut w = env.write_txn().unwrap(); db.put(&mut w, &i.to_be_bytes(), &val).unwrap(); w.commit().unwrap(); } - let size_during = file_size(dir.path()); + let size_during = high_water(dir.path(), flags); assert!( size_during > size_pinned_base, - "with a reader pinned below every base, commits must extend the file \ - (the GC-31 gate refuses the annex) — got no growth, so the gate leaked" + "with a reader pinned below every base, commits must extend (file \ + size under the default backing, last_pg under WRITE_MAP) — the \ + GC-31 gate refuses the annex — got no growth, so the gate leaked" ); assert_clean(dir.path()); drop(pin); @@ -179,16 +264,17 @@ fn reader_gate_blocks_annex_draw_then_carry_reclaims() { db.put(&mut w, &1u64.to_be_bytes(), &val).unwrap(); w.commit().unwrap(); } - let size_settled = file_size(dir.path()); + let size_settled = high_water(dir.path(), flags); for i in 0u64..8 { let mut w = env.write_txn().unwrap(); db.put(&mut w, &i.to_be_bytes(), &val).unwrap(); w.commit().unwrap(); } assert_eq!( - file_size(dir.path()), + high_water(dir.path(), flags), size_settled, - "after the reader releases, carried annex ids must satisfy churn" + "after the reader releases, carried annex ids must satisfy churn \ + with no further growth" ); assert_clean(dir.path()); } @@ -199,7 +285,7 @@ fn reader_gate_blocks_annex_draw_then_carry_reclaims() { #[test] fn over_cap_freed_set_spills_whole_to_gc_tree() { let dir = TempDir::new(); - let env = open(dir.path()); + let env = open(dir.path(), EnvFlags::EMPTY); let db = env.main_database(); let cap = zerodb_core::page::meta_annex_cap(PS) as u64; // 490 at 4 KiB @@ -255,7 +341,7 @@ fn over_cap_freed_set_spills_whole_to_gc_tree() { #[test] fn copy_preserves_annex_compact_drops_it() { let dir = TempDir::new(); - let env = open(dir.path()); + let env = open(dir.path(), EnvFlags::EMPTY); let db = env.main_database(); let val = vec![0x55u8; 256]; for i in 0u64..20 { @@ -281,7 +367,7 @@ fn copy_preserves_annex_compact_drops_it() { let dirp = raw.path().join("raw-env"); std::fs::create_dir_all(&dirp).unwrap(); std::fs::copy(&raw_path, dirp.join(zerodb::DATA_FILE_NAME)).unwrap(); - open(&dirp) + open(&dirp, EnvFlags::EMPTY) }; { let rtxn = copy_env.read_txn().unwrap(); @@ -313,7 +399,7 @@ fn copy_preserves_annex_compact_drops_it() { fn reopen_reparses_annex_and_reuses() { let dir = TempDir::new(); { - let env = open(dir.path()); + let env = open(dir.path(), EnvFlags::EMPTY); let db = env.main_database(); let val = vec![0x66u8; 256]; for i in 0u64..20 { @@ -326,7 +412,7 @@ fn reopen_reparses_annex_and_reuses() { assert!(fl > 0); let size_before = file_size(dir.path()); - let env = open(dir.path()); + let env = open(dir.path(), EnvFlags::EMPTY); let db = env.main_database(); let val = vec![0x66u8; 256]; for i in 0u64..10 { diff --git a/docs/SPEC/05-gc.md b/docs/SPEC/05-gc.md index 10812f6..072a855 100644 --- a/docs/SPEC/05-gc.md +++ b/docs/SPEC/05-gc.md @@ -143,7 +143,14 @@ to govern it unchanged. `B` cannot reach pages freed *by* `B` (they left `B`'s trees), and the crash-fallback meta for writer `W` is `B` itself, whose annex still lists the drawn ids as free and whose trees do not reference them — the TXN-62 / - in-save pool-draw argument (GC-12) verbatim. + in-save pool-draw argument (GC-12) verbatim. That "fallback is `B` itself" + premise holds as stated only in the durable modes (SPEC 06 REC-6/REC-12). + Under the relaxed-durability modes (`NO_META_SYNC` / `NO_SYNC` / + `MAP_ASYNC`) the crash-fallback can be older than `B` — the pre-existing, + characterized REC-10/REC-11 window — and under in-place `WRITE_MAP` an + uncommitted or aborted txn's in-map writes can reach disk with no commit + step at all, which falls in that same characterized window; see SPEC 06 + for the full argument and the window's taxonomy. - **GC-32 — all-or-nothing placement.** At `freelist_save`, if the final deduped `freed` set fits the annex cap (`(psize − 176) / 8` ids) it is **stashed for the meta encode** and GC-11 step (c) performs **no** tree diff --git a/docs/adr/0022-meta-freelist-annex.md b/docs/adr/0022-meta-freelist-annex.md index efe60c8..e3f2881 100644 --- a/docs/adr/0022-meta-freelist-annex.md +++ b/docs/adr/0022-meta-freelist-annex.md @@ -119,8 +119,15 @@ into the meta. Freeing-txnid = the meta's own txnid; no separate field. so trying the tree first preserves the GC-18/19 oldest-first determinism. The annex draw is gated exactly like any GC draw: `F ≤ oldest_reader()` (here: no live reader pinned below base). Drawn ids enter `reclaimed`. - Inside `freelist_save` (GcSave mode) the annex pool is never consulted — - it was folded into `freed` before the loop, so GC-12/13 are untouched. + Inside `freelist_save` (GcSave mode) the annex pool **is** consulted, as + the carried pool (GC-30): after the drain pool (`save_pool_draw`) is + exhausted, `allocate()`'s GcSave arm falls through to the same + `annex_draw`, gated the same way. This is safe only because the Write-side + step (c) placement re-reads `annex.live()` after every tree put before + deciding what to persist, so a draw here can never leave a handed-out id in + the persisted set — the GC-12-amended anti-leak argument applies to this + draw exactly as it does to `save_pool_draw` (`crates/zerodb-core/src/ + rwtxn.rs`, the `AllocMode::GcSave` arm of `allocate`, ~line 1277). - **Crash safety.** The annex is CRC-covered by the meta it belongs to: a torn meta write that corrupts the annex tears the whole slot, which the §3.2 selection discards (fallback to N−1, whose own annex+tree state is From 1ef755b4b39b73e80a65012fb88839ad8ca8a80d Mon Sep 17 00:00:00 2001 From: Quentin de Quelen Date: Mon, 5 Oct 2026 15:07:16 +0200 Subject: [PATCH 4/5] =?UTF-8?q?ADR-0022:=20accepted=20=E2=80=94=20format?= =?UTF-8?q?=20v2=20and=20crash=20seed=20198=20ratified=20(Quentin,=202026-?= =?UTF-8?q?10-05)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Status Spike → Accepted; open questions 1–2 resolved (format change, full annex cap); D-022 APPROVED; DECISIONS/PROGRESS updated with the re-gate on main after ADR-0021. Co-Authored-By: Claude Opus 5.5 --- PROGRESS.md | 2 +- .../tests/crash_harness_smoke.rs | 1 + docs/DECISIONS.md | 2 +- docs/DIVERGENCES.md | 2 +- docs/adr/0022-meta-freelist-annex.md | 20 +++++++++---------- 5 files changed, 13 insertions(+), 14 deletions(-) diff --git a/PROGRESS.md b/PROGRESS.md index da85bf8..b24a168 100644 --- a/PROGRESS.md +++ b/PROGRESS.md @@ -154,4 +154,4 @@ ADR-0020 SPIKE, 2026-10-01 — 24-byte page-header spike approved for measuremen ADR-0021 PRODUCTION PASS, 2026-10-02 — true in-place WRITE_MAP hardened on the spike branch (not merged): B1 unsafe-fn broker with the sole sanctioned call in zerodb-core::dirty; B2 superseded by a stronger M1 miri finding (under Stacked Borrows a cached whole-map reference is UB after any in-place write — the in-map writer now borrows its map view lazily per access, like readers; spill-time re-derivation kept for the heap-staged paths); B3 map-aware fault backend (brokered regions journaled, sealed per sync, image cuts open the real writable map; vacuousness tripwire + smoke pin); B4 loom justified as no-new-model (ADR-0021 §hardening), 180 s stress writemap-in-place variant, M1.13 erased-cursor audit + miri battery in heed-zerodb; M1 zerodb_io::testmap::TestWriteMap (test-backing feature) + writemap_in_place_miri battery; M2 release-mode TXN-62 typed guard in allocate; M3 abort-after-spill / nested-fanout / put_reserved batteries parameterized over WRITE_MAP. SPEC 04 TXN-71/C5a/TXN-45b and SPEC 06 REC-20 amended in the same change. -ADR-0022 SPIKE, 2026-10-03 — meta free-list annex (format v2): the per-commit freed PIL rides inside the CRC-covered meta page; the GC tree is the spill/cold path (over-cap sets, reader-gated carries). Steady-state freelist_save drops from ~2 GC-tree puts + 1 delete + 1 leaf COW per commit to zero tree ops (~1.2–1.5 µs → ~0.1 µs local 4 KiB census; one fewer dirty page per commit), and churn steady-state files shrink ~12% vs v1. First-cut lesson: folding the carried annex into `freed` up front starved the save's own allocations into per-commit extends (unbounded ratchet, caught by churn_parity_general); the carried remainder must stay a gated in-save pool (GC-12 source 4) merged at placement. Pinned crash seed re-pinned (198) per its documented protocol. Implemented on coordinator-relayed approval; ADR-0022 awaits direct ratification before merge. Worktree branch, not merged. +ADR-0022 SPIKE, 2026-10-03 — meta free-list annex (format v2): the per-commit freed PIL rides inside the CRC-covered meta page; the GC tree is the spill/cold path (over-cap sets, reader-gated carries). Steady-state freelist_save drops from ~2 GC-tree puts + 1 delete + 1 leaf COW per commit to zero tree ops (~1.2–1.5 µs → ~0.1 µs local 4 KiB census; one fewer dirty page per commit), and churn steady-state files shrink ~12% vs v1. First-cut lesson: folding the carried annex into `freed` up front starved the save's own allocations into per-commit extends (unbounded ratchet, caught by churn_parity_general); the carried remainder must stay a gated in-save pool (GC-12 source 4) merged at placement. Pinned crash seed re-pinned (198) per its documented protocol. Implemented on coordinator-relayed approval; ratified directly 2026-10-05 (format v2 + seed 198), re-gated on main after ADR-0021 merged (test 653/0, miri 206/0, loom 9/0, stress 3/0, crash-test-quick clean, fuzz-quick clean). diff --git a/crates/zerodb-oracle/tests/crash_harness_smoke.rs b/crates/zerodb-oracle/tests/crash_harness_smoke.rs index 0b22552..96d776d 100644 --- a/crates/zerodb-oracle/tests/crash_harness_smoke.rs +++ b/crates/zerodb-oracle/tests/crash_harness_smoke.rs @@ -162,6 +162,7 @@ fn regression_nometasync_reclaim_clobber_window_is_characterized() { // fallbacks, NoMetaSync, no violation). The clobber window itself is // unchanged by the annex — the reclaimed-page TXN-62/GC-18 reasoning is // identical whether the freed list lived in the tree or the meta. + // Re-pin accepted by the maintainer (Quentin, 2026-10-05). const SEED: u64 = 198; assert_eq!( gen_spec(SEED).mode, diff --git a/docs/DECISIONS.md b/docs/DECISIONS.md index 978daf6..8526ae5 100644 --- a/docs/DECISIONS.md +++ b/docs/DECISIONS.md @@ -24,4 +24,4 @@ | 0019 | [Durable meta write through an O_DSYNC descriptor — one barrier per durable commit, as LMDB's `me_mfd`](adr/0019-meta-write-dsync.md) | Accepted (all platforms; LMDB's failed-write scrub adopted) | Phase 3 (durable-commit cost) | | 0020 | [Shrink the 32-byte common page header (24 bytes, keeping pgno and the writer stamp)](adr/0020-compact-page-header.md) | Spike approved; format change pending its numbers | Phase 3 (density) | | 0021 | [True in-place WRITE_MAP — dirty pages in the map, no heap staging, no commit write-back](adr/0021-writemap-in-place.md) | Accepted (2026-10-02, Quentin); spike first — fair WRITE_MAP comparison shows ZeroDB-writemap 0.42–0.53× of LMDB-writemap | Phase 3 (PERF-GAP B5, issue #13) | -| 0022 | [Meta free-list annex — the per-commit freed PIL rides inside the (CRC-covered) meta page; GC tree becomes the spill/cold path; FORMAT_VERSION 2](adr/0022-meta-freelist-annex.md) | Spike implemented on relayed approval; awaits direct ratification | perf lever B12 (free-list save); feeds PLAN 3.1 | +| 0022 | [Meta free-list annex — the per-commit freed PIL rides inside the (CRC-covered) meta page; GC tree becomes the spill/cold path; FORMAT_VERSION 2](adr/0022-meta-freelist-annex.md) | Accepted (2026-10-05, Quentin; format v2 ratified) | perf lever B12 (free-list save); feeds PLAN 3.1 | diff --git a/docs/DIVERGENCES.md b/docs/DIVERGENCES.md index 5550085..9081cd3 100644 --- a/docs/DIVERGENCES.md +++ b/docs/DIVERGENCES.md @@ -33,7 +33,7 @@ Rules: | D-019 | Page-validation policy (ADR-0014) | LMDB never validates page contents: cell pointers, key sizes, value spans and overflow references are used as read (`mdb_page_get` checks only `pgno < mt_next_pgno`; the descent checks only that the bottom page is a leaf). A corrupt page is undefined behaviour. | ZeroDB validates by default (a corrupt page is a typed error) and adds an **opt-in** trusting policy, `FileTrust::trust_contents()` (an `unsafe` constructor passed to the safe `EnvOpenOptions::file_trust` setter), under which map pages skip the per-cell walk and behave like LMDB's. Page-number bounds, page type, meta validation and free-list id checks stay on in both policies. Overflow runs differ: the validating default checks each run's header; the trusting policy reads an overflow value without its run header (as `mdb_node_read` does), keeping only the snapshot high-water bound, so a corrupt run can return wrong bytes but never read outside the committed map (ADR-0014 amendment, 2026-09-29). A ZeroDB extension, not a behaviour change: the default is unchanged and no consumer sees the option unless it calls it. | APPROVED — opt-in, default validating | Quentin de Quelen (chat, 2026-09-26: "for the 'trusting the file' it would be great if it could be an option"; 2026-09-28: "do the read path now") | | D-020 | Sequential-writes fast path (ADR-0015) | `mdb_put` sets up a fresh cursor and descends from the root on every call; `MDB_APPEND` re-descends through `mdb_cursor_last`. | An **opt-in** env option (`EnvOpenOptions::sequential_writes`, default off) with a per-database runtime override (`Env::set_sequential_writes`) lets a write txn remember each tree's rightmost-leaf path and append there without descending when a key sorts after every key in that leaf (SPEC 03 §6.6). Results and committed files are identical with it on or off; only speed changes (ascending/APPEND loads faster, random-key writes 3–5 % slower). A ZeroDB extension, off unless called. | APPROVED — opt-in, default off | Quentin de Quelen (chat, 2026-09-28: "can the sequential writes be an option of the env"; env option plus per-database override: "perfect") | | D-021 | Configurable dirty limit (ADR-0017) | The spill threshold is a compile-time constant (`MDB_IDL_UM_MAX`, 131,072 pages); `mdb_page_spill` writes 1/8 of the dirty list once the txn's dirty room runs low. | ZeroDB spills the same way with the same default (SPEC 04 §6.3a) and adds an **opt-in** env option, `EnvOpenOptions::max_dirty_bytes`, to set the limit in bytes. A ZeroDB extension, not a behaviour change: unset, the limit is LMDB's, and results and committed files never depend on it. | APPROVED — opt-in, default LMDB's | Quentin de Quelen (chat, 2026-09-30: "go continue, I should be closer to LMDB in memory usage") | -| D-022 | Free-list placement (ADR-0022, meta annex) | The freed-page list of every commit is written into FREE_DBI (the GC B-tree) under `BE(txnid)` unconditionally; the GC tree is never empty under churn and each commit COWs a GC leaf. | Each commit's freed list rides inside its meta page's reserved bytes (the "annex"); the GC tree stays empty in steady state and is used only when a commit's freed set exceeds the annex cap or a reader parks a carried list. **On-disk format only** (`FORMAT_VERSION` 1→2); heed-level behaviour, allocation-order determinism and all query results are identical. Tools-observable (`zerodb stat` shows an empty FREE_DBI in steady state; churn files are ~1 page/commit smaller). Not a heed/LMDB *behaviour* divergence — filed here only for the format paper trail. | **PENDING — ADR-0022 is a spike awaiting direct ratification of the format change (CLAUDE.md rule 6); not merged** | (unratified as of 2026-10-03) | +| D-022 | Free-list placement (ADR-0022, meta annex) | The freed-page list of every commit is written into FREE_DBI (the GC B-tree) under `BE(txnid)` unconditionally; the GC tree is never empty under churn and each commit COWs a GC leaf. | Each commit's freed list rides inside its meta page's reserved bytes (the "annex"); the GC tree stays empty in steady state and is used only when a commit's freed set exceeds the annex cap or a reader parks a carried list. **On-disk format only** (`FORMAT_VERSION` 1→2); heed-level behaviour, allocation-order determinism and all query results are identical. Tools-observable (`zerodb stat` shows an empty FREE_DBI in steady state; churn files are ~1 page/commit smaller). Not a heed/LMDB *behaviour* divergence — filed here only for the format paper trail. | APPROVED | Quentin, 2026-10-05 (ADR-0022 ratification) | ### M1.13 adapter-boundary re-impositions (2026-07-17) diff --git a/docs/adr/0022-meta-freelist-annex.md b/docs/adr/0022-meta-freelist-annex.md index e3f2881..7713a9c 100644 --- a/docs/adr/0022-meta-freelist-annex.md +++ b/docs/adr/0022-meta-freelist-annex.md @@ -1,11 +1,10 @@ # ADR-0022: Meta free-list annex — the per-commit freed PIL rides in the meta page -- Status: **Spike** — implemented for measurement on maintainer approval - relayed 2026-10-02/03 (format-change + session constraints lifted for the - commit-CPU lever track). The approval reached this change through the - coordinating session, not directly from the maintainer; per CLAUDE.md rule 6 - this ADR still **awaits direct human ratification before merge**. Nothing is - merged or pushed; the spike exists so the bench server can measure the win. +- Status: **Accepted (2026-10-05, Quentin)** — format change (`FORMAT_VERSION` + 1→2) directly ratified per CLAUDE.md rule 6, together with the crash-harness + seed re-pin (198). Implemented first as a spike on approval relayed + 2026-10-02/03; re-gated on main after ADR-0021 (in-place WRITE_MAP) merged, + with WRITE_MAP in-place twins of the annex battery. - Result (2026-10-03): **measured win, gate green, spec-review clean.** Bench server 3-column A/B (x86-64, turbo off, 5 interleaved rounds, CODEGEN_UNITS=1, BASE = main 84582e8): `commit/batch/n1` 2.04× → **1.60×** LMDB (−22%, @@ -201,10 +200,9 @@ API-visible behavior, which is unchanged). ## Open questions for human review -1. Ratify the format change itself (CLAUDE.md rule 6 — this ADR was - implemented as a spike on relayed approval; it needs your direct sign-off - before merge). -2. Cap policy: full `(psize − 176)/8` (chosen) vs a smaller policy cap to - bound the parked-reader re-encode cost earlier. +1. ~~Ratify the format change itself (CLAUDE.md rule 6).~~ **Resolved + 2026-10-05:** ratified directly by Quentin. +2. ~~Cap policy: full `(psize − 176)/8` (chosen) vs a smaller policy cap.~~ + **Resolved 2026-10-05:** the full cap is accepted with the ADR. 3. Whether PLAN 3.1 should absorb this as its first stage (the annex is the hot tier of any future freelist redesign) or whether 3.1 supersedes it. From ad47caeb2445f5fa96f3932087c671dc521bc9c7 Mon Sep 17 00:00:00 2001 From: Quentin de Quelen Date: Mon, 5 Oct 2026 17:28:09 +0200 Subject: [PATCH 5/5] ADR-0022: real-case three-column results; README benches refreshed MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit YCSB on Graviton4 NVMe and the x86 bench server, plus Meilisearch's xtask workloads, LMDB / main (8b066a6) / annex. Durable YCSB B +7.8% on NVMe, no-sync +3–7% on x86, write p50 −14% to −25% everywhere; Meilisearch flat (commit span −6% on incremental additions). README gains a current-tree ZeroDB-vs-LMDB table and labels the 2026-09-30 field table as such. Co-Authored-By: Claude Opus 5.5 --- README.md | 30 ++++- .../2026-10-05-meta-annex-real-case.md | 109 ++++++++++++++++++ docs/adr/0022-meta-freelist-annex.md | 11 +- 3 files changed, 144 insertions(+), 6 deletions(-) create mode 100644 benches/results/2026-10-05-meta-annex-real-case.md diff --git a/README.md b/README.md index 73cbdb7..18d92be 100644 --- a/README.md +++ b/README.md @@ -60,11 +60,33 @@ bytes. Full list: [`docs/COMPATIBILITY.md`](docs/COMPATIBILITY.md). ## Performance -On its real consumers, ZeroDB runs at **LMDB-level performance** (Meilisearch -indexing 1.00×, search 1.03×; hannoy search 0.95×). Against the field — the +On its real consumers, ZeroDB runs at **LMDB-level performance**: Meilisearch +indexing 1.01× (movies) and 0.99× (incremental hackernews additions), search +0.94× — time relative to LMDB, lower is better; hannoy search 0.95×. + +**ZeroDB vs LMDB, current tree** — YCSB on a Graviton4 with local NVMe, +10 M × 128 B, kops/s (higher is better). `WRITE_MAP` is opt-in on both engines +(LMDB's `MDB_WRITEMAP`, ZeroDB's in-place `WRITE_MAP`): + +| Workload | LMDB | zerodb | LMDB `WRITE_MAP` | zerodb `WRITE_MAP` | +|---|---:|---:|---:|---:| +| YCSB A (no-sync, 2 GiB cap) | 256 | 211 | 545 | 408 | +| YCSB B (no-sync, 2 GiB cap) | 490 | 355 | 1,199 | 847 | +| YCSB B (fsync) | 238 | **258** | — | — | + +Durable writes are ahead of LMDB (1.08×); no-sync writes trail it by 18–27%, and +`WRITE_MAP` makes ZeroDB 1.6–1.7× faster than default LMDB, still behind LMDB's +own `WRITE_MAP`. Details: +[`benches/results/2026-10-05-meta-annex-real-case.md`](benches/results/2026-10-05-meta-annex-real-case.md). +_(2026-10-05, 2 reps.)_ + +**Against the field** — the [rust-storage-bench](https://github.com/marvin-j97/rust-storage-bench) suite behind the *fjall 3* article, on a Graviton4 with local NVMe — throughput -in kops/s (higher is better; per-row winner in **bold**): +in kops/s (higher is better; per-row winner in **bold**). This run predates +in-place `WRITE_MAP` and the meta free-list annex, and uses a different setup +(8 GiB cap, 2 GiB cache, 100 B values), so its numbers don't compare with the +table above: | Workload | LMDB | zerodb | fjall 3 | rocksdb | redb | sqlite | |---|---:|---:|---:|---:|---:|---:| @@ -74,7 +96,7 @@ in kops/s (higher is better; per-row winner in **bold**): | feed | **55** | 42 | 37 | 31 | 16 | 38 | | 100 M keys | **93** | 79 | 76 | 65 | 15 | 36 | -ZeroDB tracks LMDB closely — at parity on durable writes (YCSB B), a little ahead +In that run ZeroDB tracked LMDB closely — at parity on durable writes (YCSB B), a little ahead on 4 KB values, behind on the write-heavy no-sync mix (its weakest path). The LSM engines (fjall, rocksdb) take the raw write-throughput rows; the B-trees (LMDB and ZeroDB) keep read p99 in microseconds where the LSMs run to hundreds. Per-engine diff --git a/benches/results/2026-10-05-meta-annex-real-case.md b/benches/results/2026-10-05-meta-annex-real-case.md new file mode 100644 index 0000000..a65fcf9 --- /dev/null +++ b/benches/results/2026-10-05-meta-annex-real-case.md @@ -0,0 +1,109 @@ +# Meta free-list annex (ADR-0022) — real-case three-column check — 2026-10-05 + +The commit-ladder A/B for ADR-0022 only showed a win on tiny no-sync commits +(`commit/batch/n1` −22%; `n10k` and the fsync rungs flat). This run checks it on +real workloads: YCSB through rust-storage-bench, and Meilisearch's own +`cargo xtask bench` workloads. + +Three columns throughout: + +| column | tree | +|---|---| +| **LMDB** | heed 0.22.1 (the Meilisearch LMDB fork) — drift control | +| **ZeroDB before** | `8b066a6` — main with ADR-0021 (in-place `WRITE_MAP`), no annex | +| **ZeroDB after** | `b01b75b` — `8b066a6` + ADR-0022 (the annex), nothing else | + +`*-wm` rows open the environment with `WRITE_MAP` (LMDB `MDB_WRITEMAP`, ZeroDB +in-place `WRITE_MAP`). + +## YCSB — Graviton4 NVMe (the target hardware) + +AWS `m8gd.xlarge` (Graviton4, 4 vCPU, 16 GiB), Ubuntu 24.04 aarch64, local NVMe +instance store, ext4 `noatime`. rust-storage-bench (with the `zerodb` backend), +10 M items × 128 B, json corpus, 8 MiB engine cache, 60 s per run; no-sync runs in +a `MemoryMax=2G` scope, durable runs uncapped. Fresh data dir and dropped page +cache before every run. 2 reps, the second in reverse order. + +### YCSB A (50/50 read/update), no-sync + +| config | ops/s (reps) | median | ÷ LMDB | write p50 | write p99 | +|---|---|---:|---:|---:|---:| +| LMDB | 255k / 257k | 256k | 1.00 | 6.4 µs | 13.9 µs | +| ZeroDB before | 207k / 207k | 207k | 0.81 | 14.0 µs | 27.7 µs | +| **ZeroDB after** | 213k / 209k | 211k | 0.82 | **12.0 µs** | **24.8 µs** | +| LMDB-wm | 545k / 545k | 545k | 2.13 | 1.3 µs | 4.1 µs | +| ZeroDB-wm before | 421k / 404k | 412k | 1.61 | 4.4 µs | 9.3 µs | +| **ZeroDB-wm after** | 412k / 405k | 408k | 1.60 | **3.3 µs** | **7.8 µs** | + +after ÷ before: default **1.018** (rep range 1.008–1.028); wm 0.991 (0.963–1.020). + +### YCSB B (95/5 read/update), no-sync + +| config | ops/s (reps) | median | ÷ LMDB | write p50 | write p99 | +|---|---|---:|---:|---:|---:| +| LMDB | 492k / 487k | 490k | 1.00 | 7.1 µs | 15.4 µs | +| ZeroDB before | 368k / 358k | 363k | 0.74 | 15.1 µs | 31.9 µs | +| **ZeroDB after** | 372k / 339k | 355k | 0.73 | **13.1 µs** | 54.0 µs ¹ | +| LMDB-wm | 1209k / 1190k | 1199k | 2.45 | 1.4 µs | 6.4 µs | +| ZeroDB-wm before | 819k / 810k | 814k | 1.66 | 4.6 µs | 11.6 µs | +| **ZeroDB-wm after** | 845k / 848k | 847k | 1.73 | **3.5 µs** | **10.3 µs** | + +after ÷ before: default 0.980 (0.922–1.040, noisy); wm **1.040** (1.033–1.047). + +¹ One rep only: rep 1 was 29.4 µs (better than before's 31.9 µs); rep 2 was +78.5 µs during a run that read 382 MiB from disk vs ~210 MiB for the others — a +page-cache eviction episode under the 2 GiB cap, not the annex. + +### YCSB B, durable (every commit fsynced) + +| config | ops/s (reps) | median | ÷ LMDB | write p50 | write p99 | +|---|---|---:|---:|---:|---:| +| LMDB | 238k / 239k | 238k | 1.00 | 167.8 µs | 178.1 µs | +| ZeroDB before | 238k / 241k | 239k | 1.00 | 161.2 µs | 178.1 µs | +| **ZeroDB after** | 258k / 258k | **258k** | **1.08** | **129.4 µs** | **148.8 µs** | + +after ÷ before: **1.078** (rep range 1.070–1.085). One page fewer written and +flushed per commit shows up directly when the flush is cheap (NVMe). + +## YCSB — x86-64 bench server (confirmation) + +Xeon E3-1230 v2 (4C/8T, turbo off, performance governor), 2× SATA SSD in +mdraid RAID 1, Debian 12. Same method, 45 s per run, 2 reps (durable: 1 rep). + +| workload | before → after ops/s | after ÷ before | write p50 before → after | +|---|---|---:|---| +| A no-sync | 254k → 261k | **1.027** (1.019–1.034) | 17.2 → 13.8 µs | +| A no-sync, wm | 336k → 357k | **1.065** (1.022–1.110) | 8.7 → 8.0 µs | +| B no-sync | 421k → 448k | **1.064** (1.046–1.083) | 18.2 → 13.9 µs | +| B no-sync, wm | 661k → 696k | **1.053** (1.052–1.054) | 8.9 → 6.9 µs | +| B durable | 188k → 189k | 1.006 | 1398 → 1316 µs | + +The durable row is flat here: this box's fsync is a ~1.4 ms SATA RAID flush, +which drowns per-commit CPU (the same reason the ladder's `commit/sync/*` rungs +measured flat on it). + +## Meilisearch — x86-64 bench server + +Meilisearch v1.53.1 built three times (stock LMDB, ZeroDB before, ZeroDB after), +`cargo xtask bench` on `movies.json`, `hackernews-add-new-documents.json` and +`search/movies.json`, 3 rounds with the starting engine rotated each round. +Median server-side time (`::meta::total` span) over all runs: + +| workload | LMDB | ZeroDB before | ZeroDB after | after ÷ before | after ÷ LMDB | +|---|---:|---:|---:|---:|---:| +| movies indexing (30 runs) | 4.850 s | 4.865 s | 4.879 s | 1.00 | 1.01 | +| hackernews incremental additions (9 runs) | 40.25 s | 40.74 s | 40.01 s | 0.98 | 0.99 | +| — of which `indexing::scheduler::commit` | 11.03 s | 10.86 s | 10.21 s | **0.94** | 0.93 | +| movies search (30 runs) | 15.4 ms | 14.8 ms | 14.4 ms | 0.97 | 0.94 | + +No regression. Bulk indexing is flat as expected (few, large commits); the +incremental-additions workload, which commits more often, gains 6% on its commit +span. + +## Verdict + +ADR-0022 is a real-case win where commits are frequent or durable on fast +storage, and neutral for Meilisearch bulk indexing: YCSB +2–8% throughput, write +p50 −14% to −25% in every configuration on both machines, no regression in any +Meilisearch workload. It is not a Meilisearch indexing speed-up and should not be +described as one. diff --git a/docs/adr/0022-meta-freelist-annex.md b/docs/adr/0022-meta-freelist-annex.md index 7713a9c..8d3e03f 100644 --- a/docs/adr/0022-meta-freelist-annex.md +++ b/docs/adr/0022-meta-freelist-annex.md @@ -12,8 +12,15 @@ rungs flat (overhead amortized / fsync-bound — as predicted). Full gate: test 630/0, miri 0-fail, crash-test-quick all durability modes, loom 9/0, stress 180s 2/0, fuzz-quick 2.1M clean. Adversarial spec-review found no technical - blockers; its should-fix/nit items are addressed in the branch. Open for human - ratification (questions 1–3 below). + blockers; its should-fix/nit items are addressed in the branch. +- Real-case check (2026-10-05, three columns LMDB / before `8b066a6` / after, + `benches/results/2026-10-05-meta-annex-real-case.md`): YCSB on Graviton4 NVMe + durable B **1.078×** before (258k vs 239k ops/s, now 1.08× LMDB), no-sync + 0.98–1.04×; on the x86 bench server no-sync **1.03–1.07×**, durable flat + (SATA-flush-bound). Write p50 −14% to −25% in every configuration on both + machines. Meilisearch: no regression — movies indexing 1.00×, hackernews + incremental additions 0.98× (commit span 0.94×), search 0.97×. Scope of the + claim: frequent or durable commits; not a Meilisearch bulk-indexing speed-up. - Milestone: perf track "non-copy per-commit CPU" lever #1 (cheaper free-list save), PERF-GAP-VS-LMDB §B12; forward-looking toward PLAN 3.1. - Date: 2026-10-03