* [client] Coalesce peer list change notifications to the mobile listener
Every peer state change spawned a goroutine to call the platform
listener. During a reconnect storm this pinned hundreds of OS threads
in JNI and let the UI call back into the engine from each of them.
Deliver peer list changes from a single goroutine per listener and
collapse pending changes into the latest count.
* [client] Cover a pending wake-up when the peer list deliverer is replaced
The replacement test waited for the old callback to finish before
swapping listeners, so it never exercised the stop check that runs
after a wake-up. Block the old callback, queue a peer list change and
swap while it is blocked, then assert the old listener never sees the
new count.
* [client] Signal peer list deliverer exit and wait for it in the test
The replacement test sampled the old listener after a fixed sleep, so a
late stale delivery could slip past it. Close a done channel when the
deliverer goroutine returns and let the test wait on it instead.
* [client] Drop the test-only peer list deliverer exit channel
The done channel was only read by the replacement test. Production code
cannot wait on it, since joining the deliverer would block on a mobile
callback. The tests now drive the deliverer loop directly and check that
setListener and removeListener close its stop channel.
On network changes the client restarted the whole engine. That is heavy-handed and slow: it tears down working state to recover from a transition the engine could handle itself. This replaces the restart with proper network event handling.
Suspend the retry loops while no network is available. Instead of burning through backoff intervals against an unreachable network, the reconnection loops park until the OS reports a usable network again.
Reconnect immediately on a network switch. When the OS hands us a new network, connections bound to the old one are swept and re-dialed right away, rather than waiting for a timeout to notice they are dead.
The Status recorder used to fire notifier callbacks while holding d.mux:
- notifyPeerListChanged / notifyPeerStateChangeListeners ran from inside
the locked section of every Update*/AddPeerStateRoute/etc.
- notifyAddressChanged ran from UpdateLocalPeerState and CleanLocalPeerState
while d.mux was held.
- onConnectionChanged was registered with a defer above defer d.mux.Unlock,
so it executed before the mutex was released in the Mark*Connected/
Disconnected helpers.
- notifyPeerStateChangeListeners did a blocking channel send under d.mux,
so a slow subscriber stalled every other d.mux holder.
A listener that re-enters the recorder (e.g. calls GetFullStatus from
within a callback) deadlocks against d.mux, and any callback that takes
longer than expected stalls every other state query for its duration.
Capture the values needed for notification under the lock, release d.mux,
then call the notifier. Build per-peer router-state snapshots inside the
lock and dispatch them via dispatchRouterPeers afterwards. The router-peer
channel send stays blocking, but now happens outside d.mux so a slow
consumer cannot stall any other d.mux holder, and no peer state
transitions are silently dropped.
The notifier itself is unchanged: its internal state was already protected
by its own locks, and the field d.notifier is set once in NewRecorder and
never reassigned, so reading it without d.mux is safe.
Also fix a pre-existing race in Test_notifier_RemoveListener /
Test_notifier_SetListener: setListener spawns a goroutine that writes
listener.peers, but the tests read listener.peers without waiting for it.
Refactored updateServerStates and calculateState
added some checks to ensure we are not sending connecting on context canceled
removed some state updates from the RunClient function