mirror of
https://github.com/netbirdio/netbird.git
synced 2026-10-05 21:19:08 +02:00
[management] Keep an established claim when its re-read is inconclusive
SaveProxy upserts on the proxy ID, so on a reconnect the row Connect just wrote is the claim the account has held since its first connect, and the session guard on DeleteProxy matches because the upsert wrote the new session. Withdrawing that row whenever the post-write re-read errored surrendered an established claim on a transient store error — a window in which any other account could take the address — where the pre-existing code left the row untouched. An inconclusive re-read still refuses the connect, but marks the session disconnected instead of deleting the row; only a conclusive answer that the address is claimed withdraws it. The write-then-re-read argument moves to the doc of the exported ErrClusterAddressUnavailable, where the API needs it, and both helpers point there instead of carrying it twice. Store-backed tests drive the re-read through the real queries — a reconnect keeps its row, the account's own pin is not a competing claim, another account's is — since the whole path now depends on the store excluding the account's own claims. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Sa3DsBDP3VciAi4PPG17L6
This commit is contained in:
co-authored by
Claude Fable 5.1
parent
e6c69f674d
commit
d80f0ff031
@@ -92,28 +92,23 @@ func (m *Manager) Connect(ctx context.Context, proxyID, sessionID, clusterAddres
|
||||
return p, nil
|
||||
}
|
||||
|
||||
// confirmClusterAddressClaim re-asks, once the proxy's row is committed,
|
||||
// whether the account may hold the address, and withdraws the row if not.
|
||||
//
|
||||
// The connect path checks IsClusterAddressAvailable before Connect, but that
|
||||
// read and the write here are separate statements: another claim — a foreign
|
||||
// proxy row, or another account's agent network gateway pin — can land in
|
||||
// between, and its own check would not have seen this row yet either.
|
||||
// Re-reading after the write closes that window from this side, and the
|
||||
// gateway bootstrap does the same from its side: both claimants write before
|
||||
// they re-read, so of two concurrent claims at least one re-reads after the
|
||||
// other has committed and backs off. Each statement runs autocommit, so that
|
||||
// re-read sees every commit before it on sqlite, postgres and mysql alike.
|
||||
// Both may back off, which costs a reconnect; neither keeps a claim the other
|
||||
// holds, which is the invariant. No lock spans the proxies and settings
|
||||
// tables portably, and a claims table would be more machinery than the
|
||||
// property needs.
|
||||
//
|
||||
// The row is withdrawn on an inconclusive re-read too: a claim that cannot be
|
||||
// confirmed must not stand, and the proxy reconnects on its own.
|
||||
// confirmClusterAddressClaim re-reads availability once the proxy's row is
|
||||
// committed and withdraws the row if the claim is lost; see
|
||||
// proxy.ErrClusterAddressUnavailable for why the re-read is what closes the
|
||||
// race with a concurrent claim. An inconclusive re-read refuses the connect
|
||||
// but only marks the row disconnected: SaveProxy upserts on the proxy ID, so
|
||||
// on a reconnect the row is a claim the account already held, and a transient
|
||||
// store error must not surrender it.
|
||||
func (m *Manager) confirmClusterAddressClaim(ctx context.Context, p *proxy.Proxy, accountID string) error {
|
||||
available, err := m.IsClusterAddressAvailable(ctx, p.ClusterAddress, accountID)
|
||||
if err == nil && available {
|
||||
if err != nil {
|
||||
if discErr := m.store.DisconnectProxy(ctx, p.ID, p.SessionID); discErr != nil {
|
||||
log.WithContext(ctx).Errorf("failed to mark proxy %s session %s disconnected after an inconclusive claim check on %s: %v",
|
||||
p.ID, p.SessionID, p.ClusterAddress, discErr)
|
||||
}
|
||||
return fmt.Errorf("confirm claim on cluster address %s: %w", p.ClusterAddress, err)
|
||||
}
|
||||
if available {
|
||||
return nil
|
||||
}
|
||||
|
||||
@@ -121,9 +116,6 @@ func (m *Manager) confirmClusterAddressClaim(ctx context.Context, p *proxy.Proxy
|
||||
log.WithContext(ctx).Errorf("failed to withdraw proxy %s session %s after losing the claim on %s: %v",
|
||||
p.ID, p.SessionID, p.ClusterAddress, delErr)
|
||||
}
|
||||
if err != nil {
|
||||
return fmt.Errorf("confirm claim on cluster address %s: %w", p.ClusterAddress, err)
|
||||
}
|
||||
log.WithContext(ctx).Warnf("cluster address %s was claimed while proxy %s registered for account %s, withdrawing its row",
|
||||
p.ClusterAddress, p.ID, accountID)
|
||||
return fmt.Errorf("cluster address %s: %w", p.ClusterAddress, proxy.ErrClusterAddressUnavailable)
|
||||
|
||||
Reference in New Issue
Block a user