Riccardo Manfrin 164d92e78d [client] Stop dumping the whole device to clear one peer endpoint (#7632)
* [client] Stop dumping the whole device to clear one peer endpoint

Clearing a peer's endpoint has to remove and re-add the peer, because neither the
netlink API nor the wireguard-go UAPI can clear an endpoint in place. To keep the
peer's allowed IPs across that dance, RemoveEndpointAddress read them back from the
device: a full wgctrl.Device() dump on the kernel path, a full IpcGet plus text parse
on the userspace one. Both cost a round trip proportional to the entire network map,
both run under the interface lock, and both run on every relay and ICE transition.
On a routing peer with ~15700 peers that is megabytes of netlink traffic per
transition, at a measured 713 transitions per minute, with every other configuration
operation queued behind it. RemoveAllowedIP paid the same price for the same reason.

The allowed IPs cannot come from the caller: peer.Conn knows the peer's own overlay
addresses, while the routed prefixes are attached separately by the route manager's
refcounter, so a caller-supplied set would silently drop every route behind the peer.

The configurer is the only writer of its device's peer set, so it can keep an
authoritative mirror of what it configured and answer from memory instead. The mirror
is fed by every operation that changes a peer's allowed IPs and reset by a device
reconfiguration that replaces the peer set. A peer the mirror has not seen, which is
what an out-of-band reconfiguration leaves behind, still falls back to reading the
device and seeds the mirror from it.

Prefixes are unmapped on the way in, so a v4-mapped address compares equal to the
plain v4 prefix for the same network rather than registering as a second entry.

Measured on a userspace device, allocations to clear one endpoint:

  peers      64     256    1024    4096
  before   1452       -   21617       -
  after      91      91      91      91

* [client] Keep update-only allowed IP adds out of the peer mirror

AddAllowedIP configures the device with update_only, which is a silent no-op when
the peer does not exist, so its success says nothing about whether the device took
the prefix. Recording it unconditionally let the mirror hold a peer the device had
dropped, and RemoveEndpointAddress re-adds a peer without update_only: clearing the
endpoint of such a peer recreated it, carrying allowed IPs the device never held.
Allowed IPs are unique per device, so the recreated peer takes those prefixes away
from the peer that legitimately holds them.

This is not a theoretical window. Under lazy connections a routing peer's device
entry is torn down and re-created on the idle transition, and a routed prefix
re-added during that window is lost exactly because of update_only (#6863).

Allowed IP adds now merge only onto a peer the store already knows, which mirrors
the device: the operations that can create a peer record it, the update-only ones
do not. A peer missing from the store still falls back to reading the device.

* [client] Hand a prefix over to its new owner in the peer mirror

An allowed IP belongs to exactly one peer: configuring a prefix on a peer takes it
away from whichever peer held it before, and the configurer leaves that handover to
the device rather than removing the prefix from the previous holder itself, which is
what UpdatePeer's "wg will handle duplicated peer IP" refers to. The mirror recorded
the prefix on the new peer while leaving it listed under the old one, so clearing the
old peer's endpoint rewrote its allowed IPs from that stale list and took the prefix
back from the peer that now owns it. Traffic for the routed prefix then went to the
wrong peer. Reading the device before each write used to rule this out.

The store now tracks the owner of each prefix and performs the same handover, so
rewriting one peer's list cannot reclaim a prefix another peer holds.

Prefixes are also masked on the way in. A device stores them masked, so a caller
passing host bits would otherwise fail to match what a device fallback seeded and
could never remove that prefix by value. Conversion back from the device now keys
the v4-mapped decision on the mask width as well, so a genuine v6 prefix inside the
mapped range stays v6 instead of being dropped as an invalid v4 prefix.

* [client] Keep a mapped v6 prefix below /96 out of the v4 form

normalizePrefix unmapped any v4-mapped address before masking it, keeping the
original prefix length. For a genuine v6 prefix inside the mapped range, such as
::ffff:0:0/64, that pairs a v4 address with a v6 sized mask: netip.PrefixFrom
returns an invalid prefix and Masked turns it into the zero prefix. The store then
held a prefix whose Bits is -1, which cannot reproduce the allowed IP the device
was given, so re-adding the peer after an endpoint removal could fail once the
peer had already been removed.

Masking now comes first, and it also decides the address family: only a prefix at
least 96 bits long keeps the mapped marker through the mask, so anything shorter
inside that range is v6 and stays v6.

* [client] Record a peer created by a preshared key write

Setting a preshared key without updateOnly creates the peer when it is absent, and
Rosenpass applies a peer's first key exactly that way, since applyKeyLocked passes
the peer's initialized flag. The store ignored that operation, so the peer could
exist on the device while the store treated it as unknown.

An update-only allowed IP add on such a peer then succeeded on the device, which
moved the prefix away from its previous holder, while the store skipped the peer
and left the previous holder still claiming it. Clearing that holder's endpoint
rewrote it from the stale claim and took the prefix back, leaving the peer that
owns the route with nothing.

Every device operation that can create a peer now records it, which is the same
rule the update-only operations already follow from the other side.

* [client] Match a peer on the parsed key instead of its base64 form

getPeer scanned the device comparing Key.String to the caller's key. wgtypes.Key
is a 32 byte array, so it compares directly, while String base64 encodes it into a
fresh allocation on every iteration. The scan therefore allocated once per peer on
the device to find a single peer, and on a large network that is tens of thousands
of allocations per lookup.

The key is parsed once up front and the arrays are compared. Behaviour is
unchanged: the callers already parse the same key before reaching here, so the new
parse error is unreachable in practice and only guards the helper on its own.

* [client] Normalize prefixes on their way to the device

Prefixes were normalized when recorded but not when written, so a caller's raw prefix
reached the device while a different form was kept for it. The conversion is also where
a mapped prefix goes wrong: net.IPNet prints a v4-mapped address as v4 but takes the
length from its 16 byte mask, so ::ffff:10.1.2.3/64 is handed to a userspace device as
10.1.2.3/0 — an allowed IP matching every v4 address, on a peer that was meant to carry
one /64.

prefixesToIPNets now normalizes, and the two hand-built conversions in AddAllowedIP go
through it, so there is a single place where a prefix is turned into something a device
is given and it cannot disagree with what is recorded for it.

* [client] Parse the endpoint before configuring the peer

The userspace UpdatePeer parsed the endpoint address after the device had already been
configured, and returned on a parse failure. The device was then left holding a peer
that neither the activity recorder nor the allowed IP store had been told about, so the
peer was invisible to the wake path and the prefix handover for its allowed IPs never
happened, leaving the previous holder still claiming them.

The parse now happens before anything is written, so the only failure left after the
device is touched is one the caller cannot cause.

* [client] Keep the record when a peer removal fails

The two configurers disagreed: the kernel one dropped its record only once the device
had accepted the removal, the userspace one dropped it either way. Removing a peer is a
single device write, so a failure leaves the peer exactly as it was, with the allowed IPs
the record still describes. Dropping it there asserts nothing useful and only sends the
next caller to read the whole device back for an answer it already had.

The userspace one now follows the kernel and returns early on failure.

* [client] Write down what the allowed IP store does not guarantee

Two properties were relied on without being stated. The store's lock covers its map and
not the device write beside it, so consistency between the two rests on callers being
serialized, which WGIface does with its mutex; anyone removing that would have no way to
learn it mattered. And the fallback to the device only covers a peer the store has never
seen, so a peer first recorded from empty while the device already held prefixes keeps
only what was recorded, and the next endpoint removal drops the rest.

* [client] Key the allowed IP store on the parsed peer key

The store keyed on the textual key, so a lookup compared 44 byte strings while the
callers all held the parsed key already and the configurer had to carry both forms.
wgtypes.Key is a 32 byte array and compares directly, which is what getPeer was changed
to do for the same reason.

The store and its helpers now take wgtypes.Key, the callers pass the key they parsed on
entry, and the textual form survives only where something outside speaks it: parseStatus
reports peers that way, so the userspace fallback converts once for its scan.

* [client] Document the configurer methods the store changed

The exported configurer methods now carry what the allowed IP store made true of them:
when the mirror is reset, that a peer update merges its prefixes and takes them from
their previous owner, that an update-only add on an absent peer does nothing, and what
each side does with its record when a device write fails — where the two configurers
differ, since the userspace one reports a prefix it does not have and the kernel one
treats it as a no-op. mergeLocked states the lock its callers must already hold.

Docstrings that only restated the name of a test are left out; the tests explain the
scenario they set up in the body, where the explanation belongs.
2026-09-28 18:03:42 +02:00

Start using NetBird at netbird.io
See Documentation
Join our Slack channel or our Community forum


🚀 We are hiring! Join us at https://netbird.io/careers

🤖 NetBird Agent Network (Beta)

Identity-aware access control for AI agents — keyless access to LLM APIs and private resources over the encrypted NetBird tunnel. See agent-network/ or read the docs at netbird.ai.

NetBird combines a configuration-free peer-to-peer private network and a centralized access control system in a single platform, making it easy to create secure private networks for your organization or home.

Connect. NetBird creates a WireGuard-based overlay network that automatically connects your machines over an encrypted tunnel, leaving behind the hassle of opening ports, complex firewall rules, VPN gateways, and so forth.

Secure. NetBird enables secure remote access by applying granular access policies while allowing you to manage them intuitively from a single place. Works universally on any infrastructure.

https://github.com/user-attachments/assets/10cec749-bb56-4ab3-97af-4e38850108d2

Self-host NetBird (video)

Watch the video

Key features

Connectivity Management Security Automation Platforms
✓ Kernel WireGuard ✓ Admin Web UI ✓ SSO & MFA support ✓ Public API ✓ Linux
✓ Peer-to-peer connections ✓ Auto peer discovery and configuration ✓ Access control: groups & rules ✓ Setup keys for bulk provisioning ✓ macOS
✓ Connection relay fallback ✓ IdP integrations ✓ Activity logging ✓ Self-hosting quickstart script ✓ Windows
✓ Routes to external networks ✓ Private DNS ✓ Traffic events ✓ IdP groups sync with JWT ✓ Android
✓ Domain-based DNS routes ✓ Custom DNS zones ✓ Device posture checks ✓ Terraform provider ✓ Android TV
✓ Exit nodes ✓ Multiuser support ✓ Peer-to-peer encryption ✓ Ansible collection ✓ iOS
✓ IPv6 dual-stack overlay ✓ Multi-account profile switching ✓ SSH with central access policies ✓ Apple TV
✓ Browser SSH & RDP ✓ Quantum-resistance with Rosenpass ✓ FreeBSD
✓ Reverse proxy with auto-TLS ✓ Periodic re-authentication ✓ pfSense
✓ OPNsense
✓ MikroTik RouterOS
✓ OpenWRT
✓ Synology
✓ TrueNAS
✓ Proxmox
✓ Raspberry Pi
✓ Serverless
✓ Container

Quickstart with NetBird Cloud

Quickstart with self-hosted NetBird

This is the quickest way to try self-hosted NetBird. It should take around 5 minutes to get started if you already have a public domain and a VM. Follow the Advanced guide with a custom identity provider for installations with different IdPs.

Infrastructure requirements:

  • A Linux VM with at least 1 CPU and 2 GB of memory.
  • The VM should be publicly accessible on TCP ports 80 and 443 and UDP port 3478.
  • A public domain name pointing to the VM.

Software requirements:

Steps

  • Download and run the installation script:
export NETBIRD_DOMAIN=netbird.example.com; curl -fsSL https://github.com/netbirdio/netbird/releases/latest/download/getting-started.sh | bash

A bit on NetBird internals

  • Every machine in the network runs the NetBird agent, which manages WireGuard.
  • Every agent connects to the Management Service, which holds network state, manages peer IPs, and distributes updates to agents.
  • Agents use ICE (via pion/ice) to discover connection candidates for peer-to-peer connections.
  • Candidates are discovered with the help of STUN servers.
  • Agents negotiate a connection through the Signal Service, exchanging end-to-end encrypted messages with candidates.
  • When NAT traversal fails (e.g. mobile carrier-grade NAT) and a direct p2p connection isn't possible, the system falls back to a Relay Service and a secure WireGuard tunnel is established through it.

NetBird high-level architecture diagram

See a complete architecture overview for details.

Reporting bugs and requesting features

NetBird uses a discussion-first workflow. Bug reports and feature requests start in Discussions, not as issues.

What you want to do Where to go
Report a bug, regression, or unexpected behavior Issue Triage
Request a feature or share an idea Ideas & Feature Requests
Ask about setup, configuration, or self-hosting Q&A / Support
Report a security vulnerability Security policy, never a public thread

Our team and maintainers triage discussions, ask follow-up questions, check for duplicates, and reproduce bugs. Validated reports are promoted to issues. This keeps the issue tracker a clear answer to one question: what is the team working on.

Please search existing discussions and issues first, including closed ones. If something similar already exists, upvote it and add your details there instead of opening a duplicate.

For bug reports, include your NetBird version, operating system, deployment type (Cloud, self-hosted, Kubernetes, or Docker), reproduction steps, expected and actual behavior, and a debug bundle where relevant:

netbird version
netbird status -d -A
netbird debug for 1m -A -S -U

-U uploads the bundle and prints a file key you can paste instead of attaching the archive. -A anonymizes the output, which matters on a public thread. It masks most identifying details but is not full redaction, so read the bundle before posting it. Two levels are available:

Level How to select What it masks
default -A / --anonymize Public IP addresses, IPv6 ULA addresses, MAC addresses, and domains other than netbird.io, netbird.cloud, netbird.selfhosted, and netbird.stage. IPv4 private, CGNAT, and link-local ranges are kept
strict --anonymize-level strict (implies -A) The above, plus IPv4 private, CGNAT, and link-local ranges, peer names in front of netbird.cloud, netbird.selfhosted, and netbird.stage, and WireGuard public keys. Labels under netbird.io are kept, since it only hosts infrastructure

See collecting a debug bundle and the CLI reference for details.

See How to use Discussions, Issues, and Pull Requests for the full workflow, or SUPPORT.md for a shorter version.

Contributing

Contributions are welcome. Read CONTRIBUTING.md first. NetBird works ticket first, anything that changes behavior needs an issue the team has agreed on before you open a pull request.

Community projects

Note: The main branch may be in an unstable or even broken state during development. For stable versions, see releases.

Support acknowledgement

In November 2022, NetBird joined the StartUpSecure program sponsored by the Federal Ministry of Education and Research of the Federal Republic of Germany. Together with the CISPA Helmholtz Center for Information Security, NetBird brings security best practices and simplicity to private networking.

CISPA_Logo_BLACK_EN_RZ_RGB (1)

Acknowledgements

We build on open source technologies like WireGuard®, Pion ICE, and Rosenpass. We greatly appreciate the work these projects are doing, and we'd love it if you could support them too (e.g., by starring or contributing).

This repository is licensed under the BSD-3-Clause license, which applies to all parts of the repository except for the directories management/, signal/ and relay/. Those directories are licensed under the GNU Affero General Public License version 3.0 (AGPLv3). See the respective LICENSE files inside each directory.

WireGuard and the WireGuard logo are registered trademarks of Jason A. Donenfeld.

Languages
Go 94.5%
TypeScript 2.9%
Shell 1.5%
HTML 0.5%
Go Template 0.2%
Other 0.2%