diff --git a/src/pages/manage/networks/how-routing-peers-work.mdx b/src/pages/manage/networks/how-routing-peers-work.mdx index 97bcf466..b6d98164 100644 --- a/src/pages/manage/networks/how-routing-peers-work.mdx +++ b/src/pages/manage/networks/how-routing-peers-work.mdx @@ -98,19 +98,45 @@ The routing peer must have direct network reachability to the resources it serve ## High availability -Multiple routing peers can serve the same network or route. Behavior depends on the **metric** you assign each routing peer in the dashboard. There are no tunable thresholds — the metric is the only high availability control. +Multiple routing peers can serve the same network or route. Each peer that uses the network (a user's laptop, for example) picks one of those routing peers and sends all of that network's traffic through it until something makes it switch. The **metric** you assign each routing peer in the dashboard is the main control. There are no tunable thresholds. + +### How a peer picks a routing peer + +A peer compares the routing peers it is connected to, in this order: + +1. **Metric.** The lowest metric wins, whatever the connection type or latency. +2. **Connection type.** Among equal metrics, a direct (peer-to-peer) connection beats a relayed one. +3. **Latency.** Among equal metrics and the same connection type, the lowest latency wins. To prevent flapping, a peer only moves to a routing peer whose latency is at least 10 ms lower than its current one. + +Latency is measured once, when the connection to a routing peer is set up, and is not re-measured while that connection stays up. + +### What makes a peer switch + +A peer repeats this comparison whenever a routing peer's connection state changes or the network's configuration is updated. In practice it switches when: + +- **Its routing peer goes offline.** The peer moves to the next routing peer in the order above, normally within seconds. +- **A routing peer with a lower metric comes back online.** The peer moves back as soon as that routing peer reconnects, even while its connection is still relayed. +- **Its routing peer's connection falls back from direct to relayed.** If another routing peer with the same metric still has a direct connection, the peer moves to it, although nothing went offline. With different metrics this does not happen, because metric outranks connection type. + +### What a switch does to open connections + +With masquerade on (the default), every switch resets established TCP connections through the network, whatever caused it: a failover, a failback, or a fallback to relayed. The new routing peer translates the traffic to a different source address, so applications such as SSH must reconnect. Expect the connection to stall for several seconds before the reset, while the peer notices the change. With masquerade off, connections are not translated and behave differently: see [High availability with masquerade off](/manage/networks/masquerade#high-availability-with-masquerade-off). + +If users report dropped connections while no routing peer went offline, check whether the routing peers share a metric and whether their connections fall back to relayed. [Troubleshooting relayed connections](/help/troubleshooting-relayed-connections) covers why a connection is relayed. ### Primary / failover (different metrics) -The lower-metric peer carries all traffic. The higher-metric peer is held in reserve and only takes over when the primary becomes unreachable. Failover is automatic: clients start sending traffic through the standby once the primary is seen as unreachable, normally within seconds. When the primary comes back online, clients switch back to it. With masquerade on, the default, established TCP connections through the previous peer reset and applications must reconnect, because the standby translates them to a different source address. With masquerade off they are not translated, and behave differently: see [High availability with masquerade off](/manage/networks/masquerade#high-availability-with-masquerade-off). +Use different metrics when you want every peer on one known routing peer, with the others in reserve. Peers stay on the primary even when its connection falls back to relayed, but they fail back to it as soon as it returns, which resets connections a second time with masquerade on. To give routing peers different metrics, add them to the Network individually; routing peers added as a group share one metric. -**Example.** Routing Peer A has a lower metric than Routing Peer B. When Peer A goes down, all traffic fails over to Peer B. When Peer A comes back online, all traffic switches back to Peer A. +**Example.** Routing Peer A has a lower metric than Routing Peer B. Traffic uses Peer A, including while its connection is relayed. When Peer A goes down, traffic fails over to Peer B, and moves back as soon as Peer A reconnects. -### Latency switching (equal metrics) +### Nearest routing peer (equal metrics) -When two routing peers share the same metric, each client picks the one with lower latency. A switch only happens when the latency difference exceeds **20 ms**, which prevents flapping between peers that are roughly equivalent. +Use equal metrics when routing peers are spread across regions and you want each peer to use the closest one: for example, an EU-based and a US-based routing peer serving the same network, with users on both continents. Each peer picks the lowest-latency routing peer when it connects, though a direct connection still beats a relayed one whatever the latency. It does not move to another routing peer later just because latency changes, and routing peers within 10 ms of each other are effectively tied. The trade-off: a fallback to relayed can move peers to another routing peer and reset their connections. -Useful when routing peers are geographically distributed and you want each client to land on the closest one automatically — for example, an EU-based peer and a US-based peer serving the same network, with users on both continents. +### Finding switches in the client log + +The client logs each switch at the default log level as `New chosen route is …`, followed by the chosen routing peer's WireGuard public key and its score. The comparison behind the decision, including the score of the routing peer it kept, is logged only at the debug level. ### Failure domains diff --git a/src/pages/manage/networks/index.mdx b/src/pages/manage/networks/index.mdx index eef84250..67bd6894 100644 --- a/src/pages/manage/networks/index.mdx +++ b/src/pages/manage/networks/index.mdx @@ -165,7 +165,7 @@ This single policy grants the `Development` group access to **every** resource i Before you depend on a Network in production, work through these: -- **High availability.** Add more than one routing peer to the same Network for redundancy. Added individually, each peer gets its own metric: a lower metric is the primary and a higher one the failover, while equal metrics balance traffic by latency. Added as a group, the peers share one metric, so they balance by latency only and can't act as primary and failover. Keep highly available peers in different failure domains. See [High availability](/manage/networks/how-routing-peers-work#high-availability). +- **High availability.** Add more than one routing peer to the same Network for redundancy. Added individually, each routing peer gets its own metric: the lowest is the primary and higher ones are failovers. With equal metrics, each peer picks one when it connects, by connection type and then latency; this is not load balancing. Added as a group, the routing peers share one metric, so they can't act as primary and failover. Keep highly available peers in different failure domains. See [High availability](/manage/networks/how-routing-peers-work#high-availability). - **Monitoring.** Enable the **Routing Peer Disconnected** event in [Notifications](/manage/settings/notifications) to get alerted by email, webhook, or Slack when a routing peer goes offline. - **Masquerade.** On by default and the simplest option. Turn it off only when you need source-IP visibility, and only on Linux routing peers, as that's the only platform where it can be disabled. Disabling it requires a return route in the destination network, and makes high availability something you have to arrange rather than something you get. See [Masquerade](/manage/networks/masquerade). - **Internal DNS.** Domain resources resolve on the routing peer, so it must be able to resolve the name. If it already can, nothing more is needed; if it can't, distribute a nameserver to the routing peer's group. See [Internal DNS Servers](/manage/dns/internal-dns-servers). diff --git a/src/pages/use-cases/remote-access/exit-nodes.mdx b/src/pages/use-cases/remote-access/exit-nodes.mdx index 4aba92ee..7e4033a1 100644 --- a/src/pages/use-cases/remote-access/exit-nodes.mdx +++ b/src/pages/use-cases/remote-access/exit-nodes.mdx @@ -148,7 +148,7 @@ On an east-coast device, run `netbird status -d`: the peer carrying the `0.0.0.0 ### What this does and does not do -This is nearest-exit selection plus automatic failover, not load balancing: each device follows its own lowest-latency exit node, and traffic is not spread evenly across the exit nodes. See [Latency switching](/manage/networks/how-routing-peers-work#latency-switching-equal-metrics) for the selection mechanism. +This is nearest-exit selection plus automatic failover, not load balancing: each device follows its own lowest-latency exit node, and traffic is not spread evenly across the exit nodes. See [Nearest routing peer](/manage/networks/how-routing-peers-work#nearest-routing-peer-equal-metrics) for the selection mechanism. Latency-based selection is also not a guarantee of *which* exit node a given user lands on. If a rule must always hold, such as certain users always exiting in a specific country, use separate exit nodes with per-group distribution, or posture checks as in [Geo-Based Exit Node Routing](#geo-based-exit-node-routing) below. Make sure each device receives at most one exit node with [Auto Apply](#exit-node-selection-and-auto-apply) enabled: enable it on one exit node per distribution group, and keep those groups from overlapping — more than one auto-applied exit node on a device is the failure mode described in step 2 above. Offer any additional exit nodes with Auto Apply disabled; users can switch between exit nodes manually in the client either way.