* [relay] Bind listeners before serving to fix the shutdown race
A shutdown signal that arrives while the relay is still starting races
the listener goroutines. Server.Shutdown read the listener's server
field while Listen was writing it, which the race detector reported,
and when Shutdown won it saw a nil server, returned as if nothing was
running, and the listener then bound and served with nothing left to
stop it. Server.Listen never returned and the process hung on exit.
The listener lifecycle is now split into Bind and Serve. Server.Listen
binds every listener under its mutex before spawning the accept loops,
so the fields Shutdown reads are written before the goroutines exist.
A closed flag on the server makes a Listen that runs after Shutdown
return without binding. A bind failure on one listener shuts down the
ones already bound and surfaces the error at once instead of holding
it in a channel until the surviving listener exits.
* [relay] Mark the server closed before shutting down the relay
Shutdown set the closed flag only after the relay had finished closing
peers, so a Listen that started during that window could bind sockets
and start serving on a server that was already going down. The flag is
now set under listenerMux before the relay shutdown, and Accept closes
connections it receives once the relay is closed instead of leaving
them to the client's timeout.
The tests now cover the public Listen bind failure and a QUIC listener
shut down while blocked in Accept, and they pick ports that are free
for both TCP and UDP, reporting a bind error instead of a timeout.
* Fix tests
* [relay] Test the Listen and Shutdown race and the QUIC bind rollback
The existing tests order Listen and Shutdown deterministically, so the
race the fix targets was never exercised. A new test fires both from a
shared start channel across repeated rounds so either side can take the
lock first, and fails on the hang the old code produced.
The rollback test now drives Server.Listen with the real ws and quic
listeners and a UDP blocker, so the ws socket is bound and released when
the quic bind fails. Both tests pass a real TLS config because a nil one
only yields a quic listener in the devcert build.
* [relay] Bound the Shutdown wait in the concurrent Listen and Shutdown test
* [relay] Keep the default WS port for an empty listen address
net.Listen picks a random port for an empty address, while ListenAndServe
used :http or :https. Apply the same default in Bind so an empty address
keeps listening where it did before.
* [relay] Release the WS socket when Serve fails
ServeTLS can return before it takes ownership of the listener, for example
when no certificate is configured, leaving the socket opened in Bind bound
until Shutdown. Close it in Serve on any error other than a server close.
* [relay] Drain the relay outside the listener lock
Shutdown held listenerMux while the relay closed its peers gracefully, so
ListenerProtocols, and with it the healthcheck, blocked for the whole drain.
Mark the server closed and take the listeners under the lock, then drain and
stop them after releasing it. The closed flag still keeps Listen from
registering new listeners.
* [relay] Replace net.Conn with context-aware Conn interface for relay transports
Introduce a listener.Conn interface with context-based Read/Write methods,
replacing net.Conn throughout the relay server. This enables proper timeout
propagation (e.g. handshake timeout) without goroutine-based workarounds
and removes unused LocalAddr/SetDeadline methods from WS and QUIC conns.
* [relay] Refactor Peer context management to ensure proper cleanup
Integrate context creation (`context.WithCancel`) directly in `NewPeer` and remove redundant initialization in `Work`. Add `ctxCancel` calls to ensure context is properly canceled during `Close` operations.
Health-check connections now send a properly formatted auth message
with a well-known peer ID instead of immediately closing. The server
recognizes this peer ID and handles the connection gracefully with a
debug log instead of error logs.
Replaces string-based exposed address handling with URL-based InstanceURL() (type url.URL) across relay/server and relay/healthcheck; adds SchemeREL/SchemeRELS constants; updates getInstanceURL to return *url.URL with scheme and TLS validation; adjusts WS dialing and health-check logic to use URL fields.
* fix(relay): use exposed address for healthcheck TLS validation
Healthcheck was using listen address (0.0.0.0) instead of exposed address
(domain name) for certificate validation, causing validation to always fail.
Now correctly uses the exposed address where the TLS certificate is valid,
matching real client connection behavior.
* - store exposedAddress directly in Relay struct instead of parsing on every call
- remove unused parseHostPort() function
- remove unused ListenAddress() method from ServiceChecker interface
- improve error logging with address context
* [relay/healthcheck] Remove QUIC health check logic, update WebSocket validation flow
Refactored health check logic by removing QUIC-specific connection validation and simplifying logic for WebSocket protocol. Adjusted certificate validation flow and improved handling of exposed addresses.
* [relay/healthcheck] Fix certificate validation status during health check
---------
Co-authored-by: Maycon Santos <mlsmaycon@gmail.com>
Avoid invalid disconnection notifications in case the closed race dials.
In this PR resolve multiple race condition questions. Easier to understand the fix based on commit by commit.
- Remove store dependency from notifier
- Enforce the notification orders
- Fix invalid disconnection notification
- Ensure the order of the events on the consumer side
- Clients now subscribe to peer status changes.
- The server manages and maintains these subscriptions.
- Replaced raw string peer IDs with a custom peer ID type for better type safety and clarity.
* Move the handshake logic to separated struct
- The server will response to the client after it ready to process the peer
- Preload the response messages
* Fix deprecated lint issue
* Fix error handling
* [relay-server] Relay measure auth time (#2675)
Measure the Relay client's authentication time
This update adds new relay integration for NetBird clients. The new relay is based on web sockets and listens on a single port.
- Adds new relay implementation with websocket with single port relaying mechanism
- refactor peer connection logic, allowing upgrade and downgrade from/to P2P connection
- peer connections are faster since it connects first to relay and then upgrades to P2P
- maintains compatibility with old clients by not using the new relay
- updates infrastructure scripts with new relay service