Repository navigation
Add Payjoin Receiver Support (BIP 77) - #746
Camillarhi wants to merge 1 commit into
Conversation
|
👋 I see @tnull was un-assigned. |
1db2b08 to
49017bd
Compare
|
We've merged the async persistence PR you mentioned. You might want to build your draft PR on the merged commit from there on until we cut you a release. |
Thanks for letting me know. I'll build on the merged commit. |
8f6ba65 to
6499918
Compare
|
Are you stuck? Did something in our library break CI @Camillarhi |
Not stuck at all. I was just closing out some other PRs. Still working on this one, I'll let you know when it's ready. |
12a41ad to
eb97832
Compare
120b089 to
8cc2a31
Compare
5fe1a5c to
0411ada
Compare
|
Marking this as ready for review. The core receiver flow is implemented, including session persistence, PSBT handling, input contribution, mempool monitoring for payjoin transactions, and node restart recovery. Two things still pending that I'll follow up with:
Happy to get early feedback on the overall approach in the meantime. |
|
As mentioned elsewhere, we'll defer this to the 0.9 milestone. For now removed the review requests to silence the 5-fold notifications every day. Still ofc. intend to get back to this soon though. |
231c5bb to
5a9bf22
Compare
4802d17 to
eecb3ad
Compare
f68d6f7 to
2d48b5e
Compare
|
@spacebear21 notified me out of band that version 1.0 has been released. This PR has been updated and now builds on the released 1.0.0. The 1.0 API moved a fair bit from the version the PR was originally based on, and I have handled all the changes it introduced. I also tested end-to-end on regtest against a Polar bitcoind, with payjoin-cli as the sender. @bc1cindy @spacebear21 @DanGould @zealsham, ready for another look since some changes have been made since the last review due to the version bump. |
|
I did a cursory review using Opus and Sol and found a couple larger design considerations that probably should be reworked before doing a low level review. There were also some straightforward implementation issues that I think are secondary to the design considerations. For now I think we should focus on resolving the design considerations before proceeding with cleaning up the implementation issues. With that being said, please let me know if there is anything here is inaccurate. Design considerationsPayjoin is not feature gatedI guess whether Payjoin should be feature gated is something that would require input from LDK team, however I would imagine that it probably should be. Even if the LDK team determines that Payjoin should not be feature gated, the current implementation breaks other currently supported features like cargo check --no-default-features --features uniffi-defaultit fails with errors. Bitcoin core is mandatory dependency (which is not explicitly enforced at build time)Currently Payjoin requires the Bitcoin core backend during the LDK node construction and throws an error if not found. This basically makes Bitcoin core a silent dependency of Payjoin that is only found at run time instead of build time when the Payjoin configuration exists. This feels like a bad anti-pattern. The main part of this concern is that nodes that rely on Esplora or Electrum backends cannot use Payjoin. Additionally, this also makes Payjoin unusable by the uniffi bindings (as mentioned in previous section). My suggestionThe best suggestion I can think of is to enable the
What do you think of this approach? I think the 2 main requirements that implementing Payjoin in LDK node should have (other than not breaking anything) are:
As for whether Payjoin should be feature gated I think I lean towards not feature gating it, or at least making it a default feature, however I would defer to LDK maintainers on that. Implementation detailsAnchor reserve not protected from being used in Payjoin txAnchor reserve UTXO can potentially be selected as an input to the Payjoin tx. Lack of UTXO lockingContributed UTXOs are never locked or reserved. A concurrent Mid-session errors strand the fallback for ~24hNon-transient Relay requests have no timeout and allow very large responsespost_request() uses a bare bitreq request: bitreq::post(req.url)
.with_header(...)
.with_body(...)
.send_async()See src/payment/payjoin/manager.rs:282-291. bitreq has no timeout by default and permits bodies up to 1 GiB by default. A stalled or hostile relay can hold a session task forever. Because resume_payjoin_sessions() waits for the whole JoinSet, one stuck session also prevents later 15-second resume rounds. These requests need a timeout, a small protocol-specific body limit, and explicit status validation. The scheduler does unbounded work every 15 secondsEvery cycle:
There is no session limit, concurrency limit, per-session backoff, or next-attempt timestamp. Applications can create unused URIs for 24 hours, making this an easy local resource exhaustion path. A per-session worker or persisted next-action schedule with bounded concurrency would fit better. The public API hides the session lifecyclereceive() returns only a URI: pub async fn receive(...) -> Result<String, Error>There is no public session ID, status query, cancellation method, or user-facing event. Background failures are only logged. The application cannot reliably correlate a URI with the resulting payment or distinguish an active session from a failed one. Returning a session handle containing the URI and ID would leave room for status() and cancel(). Disabling Payjoin strands existing sessionsThe manager is constructed only when config.payjoin_config is present. If a node issued Payjoin URIs and restarts without that config, its active persisted sessions are ignored. They will not be polled, expired, or closed with fallback. Existing active sessions should at least cause a warning or build error when Payjoin is disabled. |
|
Thanks for the review @xstoicunicornx. I've gone through each point below. I agree with most of it and will address those, but there are a few I'd push back on.
Agreed, I'll add the gate.
It isn't a compile-time constraint, but it is enforced at node construction. On requiring Bitcoin Core at all, that was settled earlier in this review. @tnull opened an issue on the
I don't think a best-effort check gets us what the real one does.
It's the same shape as
There's a PR in progress for this: #1037. If it lands first I'll rebase onto it.
The fallback isn't stranded. At expiry, the replay returns I also don't think closing the session and broadcasting on a non-transient error is the right fix. The errors that escape here mostly aren't permanent. For example, The real problem is that the retry is hot and unbounded, every 15 seconds for up to 24 hours with no backoff. I'll add per-session backoff instead.
I'll set a timeout, a protocol-sized body limit, and status validation.
Agreed. I'll add bounded concurrency alongside the per-session backoff above.
On cancellation, that's a fair point, and I'll work on adding it here. I will update the
I will look into a good approach for this. |
Implements the receiver side of the BIP 77 Payjoin v2 protocol, allowing LDK Node users to receive payjoin payments via a payjoin directory and OHTTP relay. - Adds a `PayjoinPayment` handler exposing a `receive()` method that returns a BIP 21 URI the sender can use to initiate the payjoin flow. The full receiver state machine is implemented covering all `ReceiveSession` states: polling the directory, validating the sender's proposal, contributing inputs, finalizing the PSBT, and monitoring the mempool. - Session state is persisted via `KVStorePayjoinReceiverPersister` and survives node restarts through event log replay. Sender inputs are tracked by `OutPoint` across polling attempts to prevent replay attacks. The sender's fallback transaction is broadcast on cancellation or failure to ensure the receiver still gets paid. - Adds `PaymentKind::Payjoin` to the payment store, `PayjoinConfig` for configuring the payjoin directory and OHTTP relay via `Builder::set_payjoin_config`, and background tasks for session resumption every 15 seconds and cleanup of terminal sessions after 24 hours.
Looked at this more closely. No funds are at risk: the sender still holds their original and can broadcast it, and ours are untouched. If Payjoin is re-enabled, the resume path replays the session, sees it expired, and closes it, broadcasting the fallback if its inputs are still unspent. The only residue is a stale session row, so I don't think it needs handling beyond that. |
So the more I think about it the more I actually think the spendable sats check is unnecessary when selecting inputs since the guard should really be placed if/when adding/modifying outputs, but otherwise the wallet balance should strictly be increasing. The only harm I can see is that the anchor reserve sats will be unavailable until the payjoin tx confirms, but as you said this is similar shape to More thorough review incoming! |
There was a problem hiding this comment.
Overall I think shape looks pretty good. I really mostly focused on the payjoin bits for now since I think there was already a good amount of feedback just from those parts. Let me know if anything I said is not quite right. I think might be about ready to add tests for next review.
There was recently an async compatible receiver v2 api added to PDK which might be more ergonomic to use with async function calls. Every place where we previously required callback for validation/signing now has a non-callback counterpart. The only exception to this is the Monitor typestate which is currently in progress. This is not yet in the latest release but we should be cutting the next minor release which includes it fairly soon. Might be worth just pinning to latest commit for now while we iterate on this PR?
Also claude found an issue where starting a fresh process and payjoin works, but calling stop() and then start() on the same node object causes payjoin receive to go dead and stay dead for the life of that object. Below is the verbose explanation (sorry I just don't know that I can explain better/in my own words). Let me know if the assessment is wrong.
Stale stop_receiver in PayjoinManager
Severity: blocker (liveness)
Affects: PR 746, commit b84175b, also present in 2d48b5e
Files: src/builder.rs:2516, src/payment/payjoin/manager.rs, src/lib.rs:859
Summary
PayjoinManager holds a watch::Receiver created at build time and clones it on every
background tick. Because Receiver::clone inherits the source receiver's last-seen version
rather than the current one, the clone is permanently stale once stop() has fired. After one
stop() / start() cycle, every resume tick aborts all payjoin session tasks before they make
progress.
The code
Build time, builder.rs:2516, inside PayjoinManager::new(...):
stop_sender.subscribe(), // stored as PayjoinManager.stop_receiver, lives foreverEvery tick, manager.rs, in resume_payjoin_sessions:
let mut interrupt = self.stop_receiver.clone();
tokio::select! {
_ = async { /* drain join_set */ } => { ... }
_ = interrupt.changed() => {
join_set.abort_all();
log_info!(self.logger, "Resumed payjoin sessions were interrupted.");
}
}Why the clone is the bug
A watch::Receiver carries a "last seen version". changed() resolves when the shared version
differs from it. The two ways of obtaining a receiver differ in which version they start from
(tokio 1.53.1):
// watch.rs:1387 Sender::subscribe -> CURRENT version
let version = shared.state.load().version();
// watch.rs:1004 Receiver::clone -> the SOURCE receiver's version
let version = self.version;PayjoinManager.stop_receiver is created once at build with version V0. Nothing ever calls
changed() or borrow_and_update() on it, so its version stays pinned at V0 for the life of
the Node.
Sequence:
- Build.
stop_receiverversion = V0. stop()callsstop_sender.send(()). Shared version becomes V1.stop_receiverstill says V0.start()again. The resume task is respawned with a correct fresh receiver atlib.rs:859
(self.stop_sender.subscribe()), so the outer loop is fine.- Each tick calls
self.stop_receiver.clone(), inheriting V0. V1 differs from V0, so
interrupt.changed()is already ready. select!picks it andabort_all()fires before any session task makes progress. Every tick,
forever.
Every other task in start() calls self.stop_sender.subscribe() directly inside start(), so
they all pick up the current version and behave correctly. The payjoin manager is the only
subscriber created at build time and kept across restarts.
Scope
Requires one completed stop() followed by start() on the same Node object. That is a
supported path, deliberately so: start_inner calls
runtime.allow_cancellable_background_task_spawns() at lib.rs:326, which exists specifically
to start a new task-tracker generation after stop() closed the previous one.
Fresh-process restart is unaffected, since a new Node means a new watch channel at V0. So this
hits long-lived embedders that cycle the node (mobile apps backgrounding, a service reconnecting)
rather than anything that reloads from disk.
Damage is liveness only, not corruption. Persist-before-advance means the event log is
authoritative, so an aborted task resumes from where it was persisted. Sessions simply never
advance: no polling, no proposal, no fallback broadcast, until the 24h directory expiry
eventually trips is_expired(). The payer waits out the full expiry.
Fix
Drop the field and subscribe inside the function:
let mut interrupt = self.stop_sender.subscribe();This requires the manager to hold the Sender rather than a Receiver. Alternatively, keep the
Receiver and make the clone current with borrow_and_update() before the select. Subscribing
from a sender is cleaner and matches what the rest of start() already does.
| bitcoin-payment-instructions = { git = "https://github.com/jkczyz/bitcoin-payment-instructions", rev = "a74cc28da239f29efce14e6aff5978717b739647", optional = true } | ||
| bitcoin-payment-instructions = { git = "https://github.com/benthecarman/bitcoin-payment-instructions", rev = "632f2ce8de7ea2035d5c83d7d745a52ea3e1fe70", optional = true } | ||
|
|
||
| payjoin = { version = "1.0.0", default-features = false, features = ["v2", "io"] } |
There was a problem hiding this comment.
Payjoin should be made an optional feature (but can still be the default). We shouldn't force builds that wish to not use payjoin to pull in all payjoin dependencies. Additionally not making it feature gated causes these to error:
cargo check --no-default-features --features uniffi-default
cargo check --no-default-features --features chain-bitcoind,storage-sqlite
All payjoin related fixtures should be gated by payjoin feature.
| pub fn payjoin_payment(&self) -> Result<PayjoinPayment, Error> { | ||
| let manager = self.payjoin_manager.as_ref().ok_or(Error::PayjoinNotConfigured)?; | ||
| Ok(PayjoinPayment::new(Arc::clone(manager), Arc::clone(&self.is_running))) | ||
| } |
There was a problem hiding this comment.
Probably better to stay consistent with the rest of the methods and return infallible handler:
| pub fn payjoin_payment(&self) -> Result<PayjoinPayment, Error> { | |
| let manager = self.payjoin_manager.as_ref().ok_or(Error::PayjoinNotConfigured)?; | |
| Ok(PayjoinPayment::new(Arc::clone(manager), Arc::clone(&self.is_running))) | |
| } | |
| pub fn payjoin_payment(&self) -> PayjoinPayment { | |
| PayjoinPayment::new(self.payjoin_manager.clone(), Arc::clone(&self.is_running)) | |
| } |
If no PayjoinConfig was set on the builder, its methods return Error::PayjoinNotConfigured.
See eebc347 for the tweeks needed to PayjoinPayment.
| }, | ||
| }; | ||
|
|
||
| pending.close().save_async(&persister).await?; |
There was a problem hiding this comment.
We should broadcasting the pending.fallback_tx() first before closing here, maybe just call self.close_session_with_fallback? Otherwise just don't close and let the normal processing loop broadcast the fallback tx and close the PendingFallback state?
| pub async fn get_session(&self) -> Result<Option<PayjoinSession>, Error> { | ||
| self.payjoin_session_store.get(&self.session_id).await | ||
| } |
There was a problem hiding this comment.
| pub async fn get_session(&self) -> Result<Option<PayjoinSession>, Error> { | |
| self.payjoin_session_store.get(&self.session_id).await | |
| } | |
| pub async fn get_session(&self) -> Result<PayjoinSession, Error> { | |
| self.payjoin_session_store.get(&self.session_id).await?.ok_or(Error::InvalidPaymentId) | |
| } |
This would make call sites simpler?
| let mut session = self | ||
| .payjoin_session_store | ||
| .get(&self.session_id) | ||
| .await? | ||
| .ok_or(Error::InvalidPaymentId)?; |
There was a problem hiding this comment.
| let mut session = self | |
| .payjoin_session_store | |
| .get(&self.session_id) | |
| .await? | |
| .ok_or(Error::InvalidPaymentId)?; | |
| let mut session = self.get_session().await?; |
I think you can just do this if you take the above changes to get_session. Also applies to load and close.
| Ok(PayjoinReceiveSession { uri: pj_uri.to_string(), session_id }) | ||
| } | ||
|
|
||
| fn relay_order<'a>(&self, payjoin_config: &'a PayjoinConfig) -> Result<Vec<&'a str>, Error> { |
There was a problem hiding this comment.
I don't think this needs to return a whole list of relays? It is effectively being used as "pick one relay randomly per request" anyway so might as well just return a single randomly selected relay.
I would also suggest using a separate mailroom manager for managing relay selection, as this also allows for marking relays that have failed. See the payjoin-cli implementation: https://github.com/payjoin/rust-payjoin/blob/master/payjoin-cli/src/app/v2/ohttp.rs.
| match self.handle_error(error, persister).await? { | ||
| // Retry on the next tick instead of hammering the relay | ||
| // when a transient failure occurs. | ||
| ReceiveSession::HasReplyableError(_) => return Ok(()), |
There was a problem hiding this comment.
Should there be a max number of attempts for handling the replyable error?
| let cur_anchor_reserve_sats = | ||
| total_anchor_channels_reserve_sats(&self.channel_manager, &self.config); | ||
| let spendable_amount_sats = | ||
| self.wallet.get_spendable_amount_sats(cur_anchor_reserve_sats).unwrap_or(0); |
There was a problem hiding this comment.
As mentioned in previous comment, I don't know that this is a necessary/appropriate check here. If anything we would want to check if/when contributing outputs because that is when we might potentially be sending funds out of the wallet.
| .check_broadcast_suitability(None, |tx| { | ||
| tokio::task::block_in_place(|| { | ||
| tokio::runtime::Handle::current() | ||
| .block_on(self.chain_source.can_broadcast_transaction(tx)) | ||
| }) | ||
| .map_err(|e| ImplementationError::from(e.to_string().as_str())) | ||
| }) |
There was a problem hiding this comment.
It should not be necessary to block on this. There is now an async compatible/non-blocking interface. See
Receiver<UncheckedOriginalPayload>::extract_tx_to_check_broadcast_suitability and Receiver<UncheckedOriginalPayload>::apply_broadcast_suitability.
|
|
||
| self.payjoin_session_store.insert_or_update(session).await?; | ||
| }, | ||
| ReceiverSessionOutcome::Aborted => { |
There was a problem hiding this comment.
Does this not need to update session.status?
This PR adds support for receiving payjoin payments in LDK Node. This is currently a work in progress and implements the receiver side of the payjoin protocol.
KVStorePayjoinReceiverPersisterto handle session persistencePayjoinas aPaymentKindto the payment storeNote on persistence: The payjoin library currently only supports synchronous persistence, but they're working on adding async support(payjoin/rust-payjoin#1235). This PR sets up the persistence structure (
KVStorePayjoinReceiverPersister), which will be updated to use async operations once the upstream PR is merged.This PR partially fixes #177 and fixes #1019