Skip to content

fix: concurrent cache misses share one backing request - #1711

Draft
lucasmcdonald3 wants to merge 15 commits into
masterfrom
lucmcdon/coalesce-cache-misses
Draft

lucasmcdonald3 wants to merge 15 commits into
masterfrom
lucmcdon/coalesce-cache-misses

Conversation

@lucasmcdonald3

@lucasmcdonald3 lucasmcdonald3 commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Issue #, if available: #1663, #1665. Supersedes #1664.

Description of changes:

Concurrent cache misses share one backing request.

Hierarchical keyring

  • Before: N concurrent misses for the same branch key make N keystore calls.
  • After: Concurrent misses share one keystore call. One caller makes the call and the others wait for it to finish. Once the call returns, all callers continue in parallel with the single result. (This is the same behavior as in the MPL's StormTracking cache.)

Also more changes to align behavior with the MPL:

  • Entries can refresh before expiry (gracePeriod variable)
  • Waiting callers time out after 10 seconds

Caching CMM

  • Before: N concurrent misses for the same cache entry make N backing calls. (This could create N data keys on encrypt).
  • After: Concurrent misses share one backing call. One caller makes the call and the others wait for it to finish. Once the call returns, all callers continue in parallel with the single result.

Also add the "Waiting callers time out after 10 seconds" behavior.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

Check any applicable:

  • Were any files moved? Moving files changes their URL, which breaks all hyperlinks to the files.

Yosef Bensimchon and others added 7 commits October 6, 2026 19:53
When many encrypt/decrypt operations run concurrently against a cold
cache for the same branch key, each cache miss independently queried the
keystore, firing N DynamoDB GetItem and N KMS Decrypt calls instead of
one. getBranchKeyMaterials now shares a single in-flight request per
cache entry id, evicting it on settle so the cryptographic materials
cache keeps ownership of caching and TTL. A rejected request is evicted
too, so the next call retries rather than sharing the failure.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Track in-flight branch key requests per cache, so keyrings sharing a
cache share requests. Coalesce concurrent caching CMM misses the same
way; waiting encrypts read through the cache so each use counts
against the entry's limits.

Co-authored-by: Yosef Bensimchon <yosef.bensimchon@xero.com>
Copy branch key materials before sharing them with waiting callers, so
a concurrent eviction cannot zero them first. In the caching CMM, let
waiters request in parallel when the response cannot serve them.
Port the MPL's StormTracker to the hierarchical keyring: refresh entries
in their grace period, let another caller fetch after graceInterval,
fail waiters after inFlightTTL, and limit concurrent fetches to fanOut.

In the caching CMM, callers join an in-flight request only while it has
room in its limits, so a burst runs its requests in parallel. Waiters
time out after inFlightTTL and retry a failed request once.
@lucasmcdonald3 lucasmcdonald3 changed the title fix: coalesce concurrent cache misses fix: dedupe concurrent cache misses Oct 9, 2026
@lucasmcdonald3 lucasmcdonald3 changed the title fix: dedupe concurrent cache misses fix: concurrent cache misses share one backing request Oct 9, 2026
Lucas McDonald added 7 commits October 9, 2026 13:49
Early refresh of cached branch keys is now an opt-in gracePeriod
keyring option (seconds, default 0). With the default, a cached
branch key is used until it expires and the keystore is not called,
matching the behavior before storm tracking.

When a keystore fetch fails, callers waiting on it now fail right
away with the keystore's error instead of timing out after
inFlightTTL with a generic error. Failed fetches no longer count
toward fanOut, so failing branch keys cannot block fetches for
other branch keys.
Remove a cache key's in-flight list once it is empty, and drop stuck
requests nobody can join, so the map does not grow with every cache
key ever used.

Move SHARED_REQUESTS and the in-flight map to an internal module that
the package index does not export.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant