* [relay] Bind listeners before serving to fix the shutdown race
A shutdown signal that arrives while the relay is still starting races
the listener goroutines. Server.Shutdown read the listener's server
field while Listen was writing it, which the race detector reported,
and when Shutdown won it saw a nil server, returned as if nothing was
running, and the listener then bound and served with nothing left to
stop it. Server.Listen never returned and the process hung on exit.
The listener lifecycle is now split into Bind and Serve. Server.Listen
binds every listener under its mutex before spawning the accept loops,
so the fields Shutdown reads are written before the goroutines exist.
A closed flag on the server makes a Listen that runs after Shutdown
return without binding. A bind failure on one listener shuts down the
ones already bound and surfaces the error at once instead of holding
it in a channel until the surviving listener exits.
* [relay] Mark the server closed before shutting down the relay
Shutdown set the closed flag only after the relay had finished closing
peers, so a Listen that started during that window could bind sockets
and start serving on a server that was already going down. The flag is
now set under listenerMux before the relay shutdown, and Accept closes
connections it receives once the relay is closed instead of leaving
them to the client's timeout.
The tests now cover the public Listen bind failure and a QUIC listener
shut down while blocked in Accept, and they pick ports that are free
for both TCP and UDP, reporting a bind error instead of a timeout.
* Fix tests
* [relay] Test the Listen and Shutdown race and the QUIC bind rollback
The existing tests order Listen and Shutdown deterministically, so the
race the fix targets was never exercised. A new test fires both from a
shared start channel across repeated rounds so either side can take the
lock first, and fails on the hang the old code produced.
The rollback test now drives Server.Listen with the real ws and quic
listeners and a UDP blocker, so the ws socket is bound and released when
the quic bind fails. Both tests pass a real TLS config because a nil one
only yields a quic listener in the devcert build.
* [relay] Bound the Shutdown wait in the concurrent Listen and Shutdown test
* [relay] Keep the default WS port for an empty listen address
net.Listen picks a random port for an empty address, while ListenAndServe
used :http or :https. Apply the same default in Bind so an empty address
keeps listening where it did before.
* [relay] Release the WS socket when Serve fails
ServeTLS can return before it takes ownership of the listener, for example
when no certificate is configured, leaving the socket opened in Bind bound
until Shutdown. Close it in Serve on any error other than a server close.
* [relay] Drain the relay outside the listener lock
Shutdown held listenerMux while the relay closed its peers gracefully, so
ListenerProtocols, and with it the healthcheck, blocked for the whole drain.
Mark the server closed and take the listeners under the lock, then drain and
stop them after releasing it. The closed flag still keeps Listen from
registering new listeners.
- introduce variables to avoid publishing latest docker tags and installers
- Refactor .goreleaser.yaml to simplify docker configurations and add environment-driven flags
- removed management debug containers (it was doing only log var)
- Stopped building arm v6 32bits in favor of v7 32 bits for services (not client)
- Add target argument to docker files
* [relay] Replace net.Conn with context-aware Conn interface for relay transports
Introduce a listener.Conn interface with context-based Read/Write methods,
replacing net.Conn throughout the relay server. This enables proper timeout
propagation (e.g. handshake timeout) without goroutine-based workarounds
and removes unused LocalAddr/SetDeadline methods from WS and QUIC conns.
* [relay] Refactor Peer context management to ensure proper cleanup
Integrate context creation (`context.WithCancel`) directly in `NewPeer` and remove redundant initialization in `Work`. Add `ctxCancel` calls to ensure context is properly canceled during `Close` operations.
* Unified NetBird combined server (Management, Signal, Relay, STUN) as a single executable with richer YAML configuration, validation, and defaults.
* Official Dockerfile/image for single-container deployment.
* Optional in-process profiling endpoint for diagnostics.
* Multiplexing to route HTTP/gRPC/WebSocket traffic via one port; runtime hooks to inject custom handlers.
* **Chores**
* Updated deployment scripts, compose files, and reverse-proxy templates to target the combined server; added example configs and getting-started updates.
Health-check connections now send a properly formatted auth message
with a well-known peer ID instead of immediately closing. The server
recognizes this peer ID and handles the connection gracefully with a
debug log instead of error logs.
Replaces string-based exposed address handling with URL-based InstanceURL() (type url.URL) across relay/server and relay/healthcheck; adds SchemeREL/SchemeRELS constants; updates getInstanceURL to return *url.URL with scheme and TLS validation; adjusts WS dialing and health-check logic to use URL fields.
* fix(relay): use exposed address for healthcheck TLS validation
Healthcheck was using listen address (0.0.0.0) instead of exposed address
(domain name) for certificate validation, causing validation to always fail.
Now correctly uses the exposed address where the TLS certificate is valid,
matching real client connection behavior.
* - store exposedAddress directly in Relay struct instead of parsing on every call
- remove unused parseHostPort() function
- remove unused ListenAddress() method from ServiceChecker interface
- improve error logging with address context
* [relay/healthcheck] Remove QUIC health check logic, update WebSocket validation flow
Refactored health check logic by removing QUIC-specific connection validation and simplifying logic for WebSocket protocol. Adjusted certificate validation flow and improved handling of exposed addresses.
* [relay/healthcheck] Fix certificate validation status during health check
---------
Co-authored-by: Maycon Santos <mlsmaycon@gmail.com>
The health check endpoint listens on a dedicated HTTP server.
By default, it is available at 0.0.0.0:9000/health. This can be configured using the --health-listen-address flag.
The results are cached for 3 seconds to avoid excessive calls.
The health check performs the following:
Checks the number of active listeners.
Validates each listener via WebSocket and QUIC dials, including TLS certificate verification.
This will allow running netbird commands (including debugging) against the daemon and provide a flow similar to non-container usages.
It will by default both log to file and stderr so it can be handled more uniformly in container-native environments.
Avoid invalid disconnection notifications in case the closed race dials.
In this PR resolve multiple race condition questions. Easier to understand the fix based on commit by commit.
- Remove store dependency from notifier
- Enforce the notification orders
- Fix invalid disconnection notification
- Ensure the order of the events on the consumer side
Fix nil pointer in Relay conn address
Meanwhile, we create a relayed net.Conn struct instance, it is possible to set the relayedURL to nil.
panic: value method github.com/netbirdio/netbird/relay/client.RelayAddr.String called using nil *RelayAddr pointer
Fix relayed URL variable protection
Protect the channel closing
- Clients now subscribe to peer status changes.
- The server manages and maintains these subscriptions.
- Replaced raw string peer IDs with a custom peer ID type for better type safety and clarity.
* fix: set TLS ServerName for hostname-based QUIC connections
When connecting to a relay server by hostname, certificates are
validated against the IP address instead of the hostname.
This change sets ServerName in the TLS config when connecting
via hostname, ensuring proper certificate validation.
* use default port if port is missing in URL string
Fix WireGuard watcher related issues
- Fix race handling between TURN and Relayed reconnection
- Move the WgWatcher logic to separate struct
- Handle timeouts in a more defensive way
- Fix initial Relay client reconnection to the home server
The nhooyr.io/websocket package was renamed to github.com/coder/websocket when
the project was transferred to "coder" as the new maintainer.
Use the new import path and update go.mod and go.sum accordingly.
Signed-off-by: Christian Stewart <christian@aperture.us>
Fixes an issue on macOS where the server throws errors with default settings:
failed to write transport message to: DATAGRAM frame too large.
Further investigation is required to optimize MTU-related values.
When the remote peer switches the Relay instance then must to close the proxy connection to the old instance.
It can cause issues when the remote peer switch connects to the Relay instance multiple times and then reconnects to an instance it had previously connected to.
The cleanup loop did not manage those situations well when a connection failed or
the connection success but the code did not add a peer connection to it yet.
- in the cleanup loop check if a connection failed to a server
- after adding a foreign server connection force to keep it a minimum 5 sec