fix: correct ICE candidate field semantics in troubleshooting docs (#933)

* fix: correct ICE candidate field semantics in troubleshooting docs

The ICE candidate (Local/Remote) field shows the selected pair only, so
every relayed connection reads -/- on both peers, even when STUN worked
and srflx candidates were gathered (symmetric NAT case). Lab-verified on
client 0.76.3 against NetBird Cloud.

- relayed-connections: fix the relay/relay sample to -/-; reword the
  candidate table (the '-' row named one cause for a symptom with two);
  replace the 'weaker side' heuristic, which cannot work when both sides
  show '-', with a client-log method that splits the two causes; mark
  'relay' as the legacy TURN fallback path
- troubleshooting-client: replace the 0.27.4 status sample with 0.76.3
  output (Direct/Routes fields are gone; Relay server address and
  Networks exist now) and update the field explanations to match

* docs: show Networks line with routed subnet and exit node in status sample

Verified output: a client's status -d lists the routes a routing peer
serves it under that peer's block (Networks: 0.0.0.0/0, 10.0.0.0/24),
while the trailing summary Networks line stays '-' unless this device
routes something itself. Adds a Networks field explanation.

* docs: accuracy and readability pass on status field docs

- status sample: kernel WireGuard interface implies Linux, so the sample
  host is linux/amd64 now
- clearer table cells (candidate pair, not connection; per-row P2P
  implications) and untangled the log-reading sentence
- consistent phrasing in the Relay server address and Networks field
  explanations

* docs: expand srflx and prflx abbreviations in the candidate table

srflx (server-reflexive) and prflx (peer-reflexive) were used without
ever being expanded anywhere in the help pages.

* docs: mark turn.netbird.io as fallback-only on Ports & Firewalls

Since v0.29.0 (relay integration, #2244) clients relay through
*.relay.netbird.io and contact TURN only when that relay is unreachable
or the remote peer runs an older client. The rule stays recommended as
the last relay path.

* docs: include bare relay.netbird.io in the fallback-only note

Clients bootstrap against relay.netbird.io before being assigned a
regional *.relay.netbird.io host, so both belong in the reference.

* docs: trim the keep-this-rule advice from the fallback note

* docs: call the TURN relay legacy in the fallback note

Matches the 'Legacy fallback' wording in the relayed-connections
candidate table.

* docs: srflx plus failed checks does not skip the firewall steps

A gathered srflx candidate proves STUN discovery only; the failed
connectivity checks can still be a host firewall or a destination-scoped
egress policy, so clear Steps 4-5 on both peers before concluding
symmetric NAT.
This commit is contained in:
Jack Carter
2026-08-18 16:17:30 +02:00
committed by GitHub
parent 967e4956b3
commit d9c21cb130
3 changed files with 53 additions and 32 deletions

View File

@@ -52,6 +52,7 @@ NetBird usually won't need open ports, but sometimes you or your IT team needs t
* Note that `nftables` resolves hostnames only when the ruleset is loaded, pinning the rule to the IPs resolved at that moment. Since the pool is dynamic and geo-distributed, reload the ruleset periodically or keep the allowlist updated by other means.
* Relay service (UDP/TCP):
* **Endpoint**: turn.netbird.io
* **Legacy fallback only**: clients v0.29.0 and later relay through the NetBird relay service below (`relay.netbird.io` and `*.relay.netbird.io`, TCP/443) and use the legacy TURN relay only when that relay is unreachable or the other peer runs an older client.
* **Port range**: UDP/80,443 and TCP/443-65535
* **IPv4**: The list is dynamic and geo-distributed; we advise you to check the nearest cluster with the following command:
* `nslookup turn.netbird.io`

View File

@@ -91,52 +91,58 @@ This will output the following information:
```shell
Peers detail:
server-a.netbird.cloud:
NetBird IP: 100.75.232.118/32
NetBird IP: 100.75.232.118
Public key: kndklnsakldvnsld+XeRF4CLr/lcNF+DSdkd/t0nZHDqmE=
Status: Connected
-- detail --
Connection type: P2P
Direct: true
ICE candidate (Local/Remote): host/host
ICE candidate endpoints (Local/Remote): 10.128.0.35:51820/10.128.0.54:51820
Relay server address: rels://us-nyc-2.relay.netbird.io:443
Last connection update: 20 seconds ago
Last Wireguard handshake: 19 seconds ago
Last WireGuard handshake: 19 seconds ago
Transfer status (received/sent) 6.1 KiB/20.6 KiB
Quantum resistance: false
Routes: 10.0.0.0/24
Networks: 0.0.0.0/0, 10.0.0.0/24
Latency: 37.503682ms
server-b.netbird.cloud:
NetBird IP: 100.75.226.48/32
NetBird IP: 100.75.226.48
Public key: Mi6jtrK5Tokndklnsakldvnsld+XeRF4CLr/lcNF+DSdkd=
Status: Connected
-- detail --
Connection type: Relayed
Direct: false
ICE candidate (Local/Remote): relay/host
ICE candidate endpoints (Local/Remote): 108.54.10.33:60434/10.128.0.12:51820
ICE candidate (Local/Remote): -/-
ICE candidate endpoints (Local/Remote): -/-
Relay server address: rels://us-nyc-2.relay.netbird.io:443
Last connection update: 20 seconds ago
Last Wireguard handshake: 18 seconds ago
Last WireGuard handshake: 18 seconds ago
Transfer status (received/sent) 6.1 KiB/20.6 KiB
Quantum resistance: false
Routes: -
Latency: 37.503682ms
Networks: -
Latency: 89.503682ms
OS: darwin/amd64
Daemon version: 0.27.4
CLI version: 0.27.4
OS: linux/amd64
Daemon version: 0.76.3
CLI version: 0.76.3
Profile: default
Management: Connected to https://api.netbird.io:443
Signal: Connected to https://signal.netbird.io:443
Relays:
[stun:turn.netbird.io:5555] is Available
[stun:stun.netbird.io:443] is Available
[stun:stun.netbird.io:5555] is Available
[turns:turn.netbird.io:443?transport=tcp] is Available
[rels://us-nyc-2.relay.netbird.io:443] is Available
Nameservers:
[8.8.8.8:53, 8.8.4.4:53] for [.] is Available
FQDN: maycons-mbp-2.netbird.cloud
FQDN: my-workstation.netbird.cloud
NetBird IP: 100.75.143.239/16
Interface type: Kernel
Wireguard port: 51820
Quantum resistance: false
Routes: -
Lazy connection: false
SSH Server: Disabled
Networks: -
Peers count: 2/2 Connected
```
@@ -150,13 +156,17 @@ As for peers, the status reports the following fields:
`P2P` or `Relayed`. A relayed connection indicates a network limitation that prevents a direct connection between the peers. To diagnose and fix a relayed connection, see [Troubleshooting relayed connections](/help/troubleshooting-relayed-connections).
### Direct
`true` or `false`. `true` indicates a direct connection between the peers without a local proxy, which is common when the local peer is allocating the relay connection.
### ICE candidate (Local/Remote)
For example `relay/host`, where `relay` is the local ICE candidate type and `host` is the remote ICE candidate type. Use `Connection type` above to tell whether the selected path is direct (`P2P`) or `Relayed`.
The candidate pair ICE selected for the tunnel: the local peer's candidate type, then the remote peer's. `host/host` means both sides connect over local interface addresses; `srflx` on either side means that address was discovered via STUN. On a relayed connection the field reads `-/-` on both peers, because ICE never selected a pair; that is expected, not an extra fault. The full breakdown, including the causes behind `-/-`, is in [Troubleshooting relayed connections](/help/troubleshooting-relayed-connections#reading-the-ice-candidates).
### Relay server address
The NetBird relay (`rels://…`) available to this peer connection. It appears for P2P connections too; it carries traffic only when `Connection type` is `Relayed`.
### Networks
The ranges this peer routes for this device: network resources such as an office subnet (`10.0.0.0/24`) and, when the peer is the device's exit node, the default route (`0.0.0.0/0`). `-` means the peer routes nothing for this device.
### Last WireGuard handshake

View File

@@ -30,7 +30,7 @@ Find the peer in question and look at the **Connection type** field:
Status: Connected
-- detail --
Connection type: Relayed
ICE candidate (Local/Remote): relay/relay
ICE candidate (Local/Remote): -/-
ICE candidate endpoints (Local/Remote): -/-
Relay server address: rels://us-nyc-2.relay.netbird.io:443
Last WireGuard handshake: 25 seconds ago
@@ -41,7 +41,7 @@ Find the peer in question and look at the **Connection type** field:
| Field | What it tells you |
|---|---|
| `Connection type: Relayed` | Traffic flows through a relay server instead of directly between the peers |
| `ICE candidate (Local/Remote)` | How each side is connecting, the key diagnostic, explained below |
| `ICE candidate (Local/Remote)` | The candidate pair ICE selected; reads `-/-` on a relayed connection, explained below |
| `Relay server address` | Which relay server carries the connection |
| `Last WireGuard handshake` | A recent handshake means the tunnel itself is healthy, relayed or not |
@@ -80,17 +80,27 @@ The authoritative endpoint and port list is in [Ports & Firewalls](/about-netbir
## Reading the ICE candidates
The `ICE candidate (Local/Remote)` field shows how each side of the selected connection is reachable. It's the fastest way to tell which peer to investigate:
The `ICE candidate (Local/Remote)` field shows the candidate pair ICE *selected*: the addresses the tunnel actually uses. It reports successful outcomes only. If hole punching didn't complete, there is no selected pair, and the field reads `-/-`.
| Candidate | Meaning | Implication |
|---|---|---|
| `host` | A local interface address | Direct connectivity, P2P possible |
| `srflx` | Public address discovered via STUN | NAT traversal worked on this side |
| `prflx` | Address discovered during connectivity checks | P2P possible |
| `relay` | A relay allocation | Hole punching failed on this side |
| `-` | No candidate established | STUN unreachable or UDP blocked on this side |
| `host` | A local interface address | P2P over directly reachable addresses |
| `srflx` | Server-reflexive: this side's public address, discovered via STUN | P2P; hole punching worked through this side's NAT |
| `prflx` | Peer-reflexive: an address learned during the connectivity checks themselves | P2P; the working address surfaced mid-checks |
| `relay` | A TURN relay allocation | Legacy fallback: appears only when a peer's connection to the NetBird relay is down and TURN carries the traffic instead |
| `-` | No candidate selected, ICE did not complete | Expected on every relayed connection |
The most useful pattern: **when one side shows `srflx` or `host` and the other shows `relay` or `-`, focus your troubleshooting on the weaker side.** That peer's network is the one blocking the direct path.
The mistake this field invites: reading `-` as "STUN is unreachable or UDP is blocked". That is one way to get here (this side never gathered a public candidate). The other is a symmetric NAT: STUN works, both sides gather `srflx` candidates, and the connectivity checks still fail because each NAT hands out a different port per destination, so the advertised addresses are never the ones packets actually arrive from. Both causes end in the same `-/-`, and **both peers show it**: a relayed connection never displays a good side and a bad side, so this field alone cannot tell you which network to fix. The [decision flow](#the-decision-flow) below can.
The client log can also split the two causes directly. Raise the log level with `netbird debug log level debug` (or use the log inside a [debug bundle](/help/troubleshooting-client#debug-bundle)) and look for the candidate and state lines:
```
discovered local candidate udp4 srflx 203.0.113.10:51791 ...
ICE ConnectionState has changed to Checking
ICE ConnectionState has changed to Failed
```
If `srflx` lines appear and the state still cycles from `Checking` to `Failed`, STUN discovery works and something is blocking the direct path itself: a host firewall ([Step 4](#step-4-is-a-host-firewall-in-the-way)), an egress policy that allows UDP only to the NetBird endpoints, or a symmetric NAT. Clear the first two on both peers ([Step 5](#step-5-repeat-on-the-other-peer)) before concluding symmetric NAT ([Step 1](#step-1-environment-triage), [Step 6](#step-6-conclude-or-escalate)). If no `srflx` lines appear at all, this side never learned its public address, so STUN or UDP is blocked: go to [Step 3](#step-3-is-stun-reachable). A strong tell for symmetric NAT: the `srflx` port differs per STUN destination across those lines.
## The decision flow
@@ -182,7 +192,7 @@ If instead something looks wrong but you can't place it, collect evidence and es
A remote engineer's laptop reaches `build-server` in the office, but `netbird status -d` on the laptop shows the connection is relayed and latency is poor. Working the flow:
1. **Confirm.** The laptop shows `Connection type: Relayed` and `ICE candidate (Local/Remote): srflx/relay`. The laptop's own side reached STUN fine (`srflx`); the weak side is the server.
1. **Confirm.** The laptop shows `Connection type: Relayed` and `ICE candidate (Local/Remote): -/-`, like every relayed connection. The field can't say which side is the blocker, so work the flow on both peers.
2. **Triage.** The laptop is on home fiber, the server on the office LAN. Neither is mobile, CGNAT, or behind a cloud NAT gateway, so this should be fixable. Continue.
3. **Control plane, on the server.** `curl` to the Management API and `nc -zv signal.netbird.io 443` both succeed.
4. **STUN, on the server.** The `Relays:` section shows `[stun:stun.netbird.io:3478] is Unavailable, reason: stun request: context deadline exceeded`. The office egress firewall is dropping outbound UDP.