mirror of
https://github.com/netbirdio/netbird.git
synced 2026-09-29 10:09:07 +02:00
* [client] Stop dumping the whole device to clear one peer endpoint Clearing a peer's endpoint has to remove and re-add the peer, because neither the netlink API nor the wireguard-go UAPI can clear an endpoint in place. To keep the peer's allowed IPs across that dance, RemoveEndpointAddress read them back from the device: a full wgctrl.Device() dump on the kernel path, a full IpcGet plus text parse on the userspace one. Both cost a round trip proportional to the entire network map, both run under the interface lock, and both run on every relay and ICE transition. On a routing peer with ~15700 peers that is megabytes of netlink traffic per transition, at a measured 713 transitions per minute, with every other configuration operation queued behind it. RemoveAllowedIP paid the same price for the same reason. The allowed IPs cannot come from the caller: peer.Conn knows the peer's own overlay addresses, while the routed prefixes are attached separately by the route manager's refcounter, so a caller-supplied set would silently drop every route behind the peer. The configurer is the only writer of its device's peer set, so it can keep an authoritative mirror of what it configured and answer from memory instead. The mirror is fed by every operation that changes a peer's allowed IPs and reset by a device reconfiguration that replaces the peer set. A peer the mirror has not seen, which is what an out-of-band reconfiguration leaves behind, still falls back to reading the device and seeds the mirror from it. Prefixes are unmapped on the way in, so a v4-mapped address compares equal to the plain v4 prefix for the same network rather than registering as a second entry. Measured on a userspace device, allocations to clear one endpoint: peers 64 256 1024 4096 before 1452 - 21617 - after 91 91 91 91 * [client] Keep update-only allowed IP adds out of the peer mirror AddAllowedIP configures the device with update_only, which is a silent no-op when the peer does not exist, so its success says nothing about whether the device took the prefix. Recording it unconditionally let the mirror hold a peer the device had dropped, and RemoveEndpointAddress re-adds a peer without update_only: clearing the endpoint of such a peer recreated it, carrying allowed IPs the device never held. Allowed IPs are unique per device, so the recreated peer takes those prefixes away from the peer that legitimately holds them. This is not a theoretical window. Under lazy connections a routing peer's device entry is torn down and re-created on the idle transition, and a routed prefix re-added during that window is lost exactly because of update_only (#6863). Allowed IP adds now merge only onto a peer the store already knows, which mirrors the device: the operations that can create a peer record it, the update-only ones do not. A peer missing from the store still falls back to reading the device. * [client] Hand a prefix over to its new owner in the peer mirror An allowed IP belongs to exactly one peer: configuring a prefix on a peer takes it away from whichever peer held it before, and the configurer leaves that handover to the device rather than removing the prefix from the previous holder itself, which is what UpdatePeer's "wg will handle duplicated peer IP" refers to. The mirror recorded the prefix on the new peer while leaving it listed under the old one, so clearing the old peer's endpoint rewrote its allowed IPs from that stale list and took the prefix back from the peer that now owns it. Traffic for the routed prefix then went to the wrong peer. Reading the device before each write used to rule this out. The store now tracks the owner of each prefix and performs the same handover, so rewriting one peer's list cannot reclaim a prefix another peer holds. Prefixes are also masked on the way in. A device stores them masked, so a caller passing host bits would otherwise fail to match what a device fallback seeded and could never remove that prefix by value. Conversion back from the device now keys the v4-mapped decision on the mask width as well, so a genuine v6 prefix inside the mapped range stays v6 instead of being dropped as an invalid v4 prefix. * [client] Keep a mapped v6 prefix below /96 out of the v4 form normalizePrefix unmapped any v4-mapped address before masking it, keeping the original prefix length. For a genuine v6 prefix inside the mapped range, such as ::ffff:0:0/64, that pairs a v4 address with a v6 sized mask: netip.PrefixFrom returns an invalid prefix and Masked turns it into the zero prefix. The store then held a prefix whose Bits is -1, which cannot reproduce the allowed IP the device was given, so re-adding the peer after an endpoint removal could fail once the peer had already been removed. Masking now comes first, and it also decides the address family: only a prefix at least 96 bits long keeps the mapped marker through the mask, so anything shorter inside that range is v6 and stays v6. * [client] Record a peer created by a preshared key write Setting a preshared key without updateOnly creates the peer when it is absent, and Rosenpass applies a peer's first key exactly that way, since applyKeyLocked passes the peer's initialized flag. The store ignored that operation, so the peer could exist on the device while the store treated it as unknown. An update-only allowed IP add on such a peer then succeeded on the device, which moved the prefix away from its previous holder, while the store skipped the peer and left the previous holder still claiming it. Clearing that holder's endpoint rewrote it from the stale claim and took the prefix back, leaving the peer that owns the route with nothing. Every device operation that can create a peer now records it, which is the same rule the update-only operations already follow from the other side. * [client] Match a peer on the parsed key instead of its base64 form getPeer scanned the device comparing Key.String to the caller's key. wgtypes.Key is a 32 byte array, so it compares directly, while String base64 encodes it into a fresh allocation on every iteration. The scan therefore allocated once per peer on the device to find a single peer, and on a large network that is tens of thousands of allocations per lookup. The key is parsed once up front and the arrays are compared. Behaviour is unchanged: the callers already parse the same key before reaching here, so the new parse error is unreachable in practice and only guards the helper on its own. * [client] Normalize prefixes on their way to the device Prefixes were normalized when recorded but not when written, so a caller's raw prefix reached the device while a different form was kept for it. The conversion is also where a mapped prefix goes wrong: net.IPNet prints a v4-mapped address as v4 but takes the length from its 16 byte mask, so ::ffff:10.1.2.3/64 is handed to a userspace device as 10.1.2.3/0 — an allowed IP matching every v4 address, on a peer that was meant to carry one /64. prefixesToIPNets now normalizes, and the two hand-built conversions in AddAllowedIP go through it, so there is a single place where a prefix is turned into something a device is given and it cannot disagree with what is recorded for it. * [client] Parse the endpoint before configuring the peer The userspace UpdatePeer parsed the endpoint address after the device had already been configured, and returned on a parse failure. The device was then left holding a peer that neither the activity recorder nor the allowed IP store had been told about, so the peer was invisible to the wake path and the prefix handover for its allowed IPs never happened, leaving the previous holder still claiming them. The parse now happens before anything is written, so the only failure left after the device is touched is one the caller cannot cause. * [client] Keep the record when a peer removal fails The two configurers disagreed: the kernel one dropped its record only once the device had accepted the removal, the userspace one dropped it either way. Removing a peer is a single device write, so a failure leaves the peer exactly as it was, with the allowed IPs the record still describes. Dropping it there asserts nothing useful and only sends the next caller to read the whole device back for an answer it already had. The userspace one now follows the kernel and returns early on failure. * [client] Write down what the allowed IP store does not guarantee Two properties were relied on without being stated. The store's lock covers its map and not the device write beside it, so consistency between the two rests on callers being serialized, which WGIface does with its mutex; anyone removing that would have no way to learn it mattered. And the fallback to the device only covers a peer the store has never seen, so a peer first recorded from empty while the device already held prefixes keeps only what was recorded, and the next endpoint removal drops the rest. * [client] Key the allowed IP store on the parsed peer key The store keyed on the textual key, so a lookup compared 44 byte strings while the callers all held the parsed key already and the configurer had to carry both forms. wgtypes.Key is a 32 byte array and compares directly, which is what getPeer was changed to do for the same reason. The store and its helpers now take wgtypes.Key, the callers pass the key they parsed on entry, and the textual form survives only where something outside speaks it: parseStatus reports peers that way, so the userspace fallback converts once for its scan. * [client] Document the configurer methods the store changed The exported configurer methods now carry what the allowed IP store made true of them: when the mirror is reset, that a peer update merges its prefixes and takes them from their previous owner, that an update-only add on an absent peer does nothing, and what each side does with its record when a device write fails — where the two configurers differ, since the userspace one reports a prefix it does not have and the kernel one treats it as a no-op. mergeLocked states the lock its callers must already hold. Docstrings that only restated the name of a test are left out; the tests explain the scenario they set up in the body, where the explanation belongs.
374 lines
10 KiB
Go
374 lines
10 KiB
Go
//go:build (linux && !android) || freebsd
|
|
|
|
package configurer
|
|
|
|
import (
|
|
"fmt"
|
|
"net"
|
|
"net/netip"
|
|
"slices"
|
|
"time"
|
|
|
|
log "github.com/sirupsen/logrus"
|
|
"golang.zx2c4.com/wireguard/wgctrl"
|
|
"golang.zx2c4.com/wireguard/wgctrl/wgtypes"
|
|
|
|
"github.com/netbirdio/netbird/monotime"
|
|
)
|
|
|
|
type KernelConfigurer struct {
|
|
deviceName string
|
|
statsCache *statsCache
|
|
allowedIPs *allowedIPStore
|
|
}
|
|
|
|
// NewKernelConfigurer creates a configurer with an empty allowed IP mirror
|
|
// and a statistics cache for the named kernel device.
|
|
func NewKernelConfigurer(deviceName string) *KernelConfigurer {
|
|
c := &KernelConfigurer{
|
|
deviceName: deviceName,
|
|
allowedIPs: newAllowedIPStore(),
|
|
}
|
|
c.statsCache = newStatsCache(statsCacheTTL, c.fetchStats)
|
|
return c
|
|
}
|
|
|
|
// ConfigureInterface sets the device key, port and firewall mark, replacing all peers.
|
|
// The allowed IP mirror is reset only after the device accepts the configuration.
|
|
func (c *KernelConfigurer) ConfigureInterface(privateKey string, port int) error {
|
|
log.Debugf("adding Wireguard private key")
|
|
key, err := wgtypes.ParseKey(privateKey)
|
|
if err != nil {
|
|
return err
|
|
}
|
|
fwmark := getFwmark()
|
|
config := wgtypes.Config{
|
|
PrivateKey: &key,
|
|
ReplacePeers: true,
|
|
FirewallMark: &fwmark,
|
|
ListenPort: &port,
|
|
}
|
|
|
|
err = c.configure(config)
|
|
if err != nil {
|
|
return fmt.Errorf(`received error "%w" while configuring interface %s with port %d`, err, c.deviceName, port)
|
|
}
|
|
|
|
c.allowedIPs.reset()
|
|
return nil
|
|
}
|
|
|
|
// SetPresharedKey sets the preshared key for a peer.
|
|
// If updateOnly is true, only updates the existing peer; if false, creates or updates.
|
|
func (c *KernelConfigurer) SetPresharedKey(peerKey string, psk wgtypes.Key, updateOnly bool) error {
|
|
parsedPeerKey, err := wgtypes.ParseKey(peerKey)
|
|
if err != nil {
|
|
return err
|
|
}
|
|
|
|
cfg := buildPresharedKeyConfig(parsedPeerKey, psk, updateOnly)
|
|
if err := c.configure(cfg); err != nil {
|
|
return err
|
|
}
|
|
|
|
// Without updateOnly this creates the peer when it is absent, so the store has to
|
|
// know about it even though no allowed IP was configured.
|
|
if !updateOnly {
|
|
c.allowedIPs.ensure(parsedPeerKey)
|
|
}
|
|
return nil
|
|
}
|
|
|
|
// UpdatePeer creates or updates a peer, merging allowed IPs with its existing set.
|
|
// Prefixes assigned to this peer are transferred from their previous owners.
|
|
func (c *KernelConfigurer) UpdatePeer(peerKey string, allowedIps []netip.Prefix, keepAlive time.Duration, endpoint *net.UDPAddr, preSharedKey *wgtypes.Key) error {
|
|
peerKeyParsed, err := wgtypes.ParseKey(peerKey)
|
|
if err != nil {
|
|
return err
|
|
}
|
|
peer := wgtypes.PeerConfig{
|
|
PublicKey: peerKeyParsed,
|
|
ReplaceAllowedIPs: false,
|
|
// don't replace allowed ips, wg will handle duplicated peer IP
|
|
AllowedIPs: prefixesToIPNets(allowedIps),
|
|
PersistentKeepaliveInterval: &keepAlive,
|
|
Endpoint: endpoint,
|
|
PresharedKey: preSharedKey,
|
|
}
|
|
|
|
config := wgtypes.Config{
|
|
Peers: []wgtypes.PeerConfig{peer},
|
|
}
|
|
err = c.configure(config)
|
|
if err != nil {
|
|
return fmt.Errorf(`received error "%w" while updating peer on interface %s with settings: allowed ips %s, endpoint %s`, err, c.deviceName, allowedIps, endpoint.String())
|
|
}
|
|
|
|
c.allowedIPs.add(peerKeyParsed, allowedIps)
|
|
return nil
|
|
}
|
|
|
|
// RemoveEndpointAddress clears the endpoint of a peer while keeping it configured.
|
|
// Neither the netlink API nor the userspace one can clear an endpoint in place, so the peer
|
|
// is removed and re-added with the allowed IPs it already had.
|
|
func (c *KernelConfigurer) RemoveEndpointAddress(peerKey string) error {
|
|
peerKeyParsed, err := wgtypes.ParseKey(peerKey)
|
|
if err != nil {
|
|
return err
|
|
}
|
|
|
|
allowedIPs, err := c.peerAllowedIPs(peerKeyParsed)
|
|
if err != nil {
|
|
return err
|
|
}
|
|
|
|
removePeerCfg := wgtypes.PeerConfig{
|
|
PublicKey: peerKeyParsed,
|
|
Remove: true,
|
|
}
|
|
|
|
if err := c.configure(wgtypes.Config{Peers: []wgtypes.PeerConfig{removePeerCfg}}); err != nil {
|
|
return fmt.Errorf("remove peer %s from interface %s: %w", peerKey, c.deviceName, err)
|
|
}
|
|
|
|
reAddPeerCfg := wgtypes.PeerConfig{
|
|
PublicKey: peerKeyParsed,
|
|
AllowedIPs: prefixesToIPNets(allowedIPs),
|
|
ReplaceAllowedIPs: true,
|
|
}
|
|
|
|
if err := c.configure(wgtypes.Config{Peers: []wgtypes.PeerConfig{reAddPeerCfg}}); err != nil {
|
|
c.allowedIPs.forget(peerKeyParsed)
|
|
return fmt.Errorf(
|
|
"re-add peer %s to interface %s with allowed IPs %v: %w",
|
|
peerKey, c.deviceName, allowedIPs, err,
|
|
)
|
|
}
|
|
|
|
return nil
|
|
}
|
|
|
|
// RemovePeer removes a peer and forgets its allowed IPs after a successful device write.
|
|
func (c *KernelConfigurer) RemovePeer(peerKey string) error {
|
|
peerKeyParsed, err := wgtypes.ParseKey(peerKey)
|
|
if err != nil {
|
|
return err
|
|
}
|
|
|
|
peer := wgtypes.PeerConfig{
|
|
PublicKey: peerKeyParsed,
|
|
Remove: true,
|
|
}
|
|
|
|
config := wgtypes.Config{
|
|
Peers: []wgtypes.PeerConfig{peer},
|
|
}
|
|
err = c.configure(config)
|
|
if err != nil {
|
|
return fmt.Errorf(`received error "%w" while removing peer %s from interface %s`, err, peerKey, c.deviceName)
|
|
}
|
|
|
|
c.allowedIPs.forget(peerKeyParsed)
|
|
return nil
|
|
}
|
|
|
|
// AddAllowedIP adds a prefix to an existing peer; an absent peer is a silent no-op.
|
|
func (c *KernelConfigurer) AddAllowedIP(peerKey string, allowedIP netip.Prefix) error {
|
|
peerKeyParsed, err := wgtypes.ParseKey(peerKey)
|
|
if err != nil {
|
|
return err
|
|
}
|
|
peer := wgtypes.PeerConfig{
|
|
PublicKey: peerKeyParsed,
|
|
UpdateOnly: true,
|
|
ReplaceAllowedIPs: false,
|
|
AllowedIPs: prefixesToIPNets([]netip.Prefix{allowedIP}),
|
|
}
|
|
|
|
config := wgtypes.Config{
|
|
Peers: []wgtypes.PeerConfig{peer},
|
|
}
|
|
err = c.configure(config)
|
|
if err != nil {
|
|
return fmt.Errorf(`received error "%w" while adding allowed Ip to peer on interface %s with settings: allowed ips %s`, err, c.deviceName, allowedIP)
|
|
}
|
|
|
|
c.allowedIPs.addExisting(peerKeyParsed, []netip.Prefix{allowedIP})
|
|
return nil
|
|
}
|
|
|
|
// RemoveAllowedIP removes a prefix while preserving the peer's other allowed IPs.
|
|
// A prefix not assigned to the peer is a no-op.
|
|
func (c *KernelConfigurer) RemoveAllowedIP(peerKey string, allowedIP netip.Prefix) error {
|
|
peerKeyParsed, err := wgtypes.ParseKey(peerKey)
|
|
if err != nil {
|
|
return fmt.Errorf("parse peer key: %w", err)
|
|
}
|
|
|
|
currentAllowedIPs, err := c.peerAllowedIPs(peerKeyParsed)
|
|
if err != nil {
|
|
return err
|
|
}
|
|
|
|
idx := slices.Index(currentAllowedIPs, normalizePrefix(allowedIP))
|
|
if idx < 0 {
|
|
return nil
|
|
}
|
|
newAllowedIPs := slices.Delete(currentAllowedIPs, idx, idx+1)
|
|
|
|
peer := wgtypes.PeerConfig{
|
|
PublicKey: peerKeyParsed,
|
|
UpdateOnly: true,
|
|
ReplaceAllowedIPs: true,
|
|
AllowedIPs: prefixesToIPNets(newAllowedIPs),
|
|
}
|
|
|
|
config := wgtypes.Config{
|
|
Peers: []wgtypes.PeerConfig{peer},
|
|
}
|
|
if err := c.configure(config); err != nil {
|
|
return fmt.Errorf("remove allowed IP %s on interface %s: %w", allowedIP, c.deviceName, err)
|
|
}
|
|
|
|
c.allowedIPs.set(peerKeyParsed, newAllowedIPs)
|
|
return nil
|
|
}
|
|
|
|
// peerAllowedIPs returns the allowed IPs configured for a peer, reading them from the device
|
|
// only for a peer the store has not seen. Dumping the device costs a netlink round trip
|
|
// proportional to the whole network map, and this runs on every relay and ICE transition.
|
|
func (c *KernelConfigurer) peerAllowedIPs(peerKey wgtypes.Key) ([]netip.Prefix, error) {
|
|
if prefixes, ok := c.allowedIPs.get(peerKey); ok {
|
|
return prefixes, nil
|
|
}
|
|
|
|
existingPeer, err := c.getPeer(c.deviceName, peerKey)
|
|
if err != nil {
|
|
return nil, fmt.Errorf("get peer: %w", err)
|
|
}
|
|
|
|
prefixes := ipNetsToPrefixes(existingPeer.AllowedIPs)
|
|
c.allowedIPs.set(peerKey, prefixes)
|
|
return prefixes, nil
|
|
}
|
|
|
|
// getPeer scans the device for one peer. wgtypes.Key is an array, so the comparison is a
|
|
// plain equality: Key.String would base64 encode into a fresh allocation for every peer.
|
|
func (c *KernelConfigurer) getPeer(ifaceName string, peerPubKey wgtypes.Key) (wgtypes.Peer, error) {
|
|
wg, err := wgctrl.New()
|
|
if err != nil {
|
|
return wgtypes.Peer{}, fmt.Errorf("wgctl: %w", err)
|
|
}
|
|
defer func() {
|
|
err = wg.Close()
|
|
if err != nil {
|
|
log.Errorf("Got error while closing wgctl: %v", err)
|
|
}
|
|
}()
|
|
|
|
wgDevice, err := wg.Device(ifaceName)
|
|
if err != nil {
|
|
return wgtypes.Peer{}, fmt.Errorf("get device %s: %w", ifaceName, err)
|
|
}
|
|
for _, peer := range wgDevice.Peers {
|
|
if peer.PublicKey == peerPubKey {
|
|
return peer, nil
|
|
}
|
|
}
|
|
return wgtypes.Peer{}, ErrPeerNotFound
|
|
}
|
|
|
|
func (c *KernelConfigurer) configure(config wgtypes.Config) error {
|
|
wg, err := wgctrl.New()
|
|
if err != nil {
|
|
return err
|
|
}
|
|
defer func() {
|
|
if err := wg.Close(); err != nil {
|
|
log.Errorf("Failed to close wgctrl client: %v", err)
|
|
}
|
|
}()
|
|
|
|
return wg.ConfigureDevice(c.deviceName, config)
|
|
}
|
|
|
|
func (c *KernelConfigurer) Close() {
|
|
}
|
|
|
|
func (c *KernelConfigurer) FullStats() (*Stats, error) {
|
|
wg, err := wgctrl.New()
|
|
if err != nil {
|
|
return nil, fmt.Errorf("wgctl: %w", err)
|
|
}
|
|
defer func() {
|
|
err = wg.Close()
|
|
if err != nil {
|
|
log.Errorf("Got error while closing wgctl: %v", err)
|
|
}
|
|
}()
|
|
|
|
wgDevice, err := wg.Device(c.deviceName)
|
|
if err != nil {
|
|
return nil, fmt.Errorf("get device %s: %w", c.deviceName, err)
|
|
}
|
|
fullStats := &Stats{
|
|
DeviceName: wgDevice.Name,
|
|
PublicKey: wgDevice.PublicKey.String(),
|
|
ListenPort: wgDevice.ListenPort,
|
|
FWMark: wgDevice.FirewallMark,
|
|
Peers: []Peer{},
|
|
}
|
|
|
|
for _, p := range wgDevice.Peers {
|
|
peer := Peer{
|
|
PublicKey: p.PublicKey.String(),
|
|
AllowedIPs: p.AllowedIPs,
|
|
TxBytes: p.TransmitBytes,
|
|
RxBytes: p.ReceiveBytes,
|
|
LastHandshake: p.LastHandshakeTime,
|
|
PresharedKey: [32]byte(p.PresharedKey),
|
|
}
|
|
if p.Endpoint != nil {
|
|
peer.Endpoint = *p.Endpoint
|
|
}
|
|
fullStats.Peers = append(fullStats.Peers, peer)
|
|
}
|
|
return fullStats, nil
|
|
}
|
|
|
|
func (c *KernelConfigurer) GetStats() (map[string]WGStats, error) {
|
|
return c.statsCache.get()
|
|
}
|
|
|
|
func (c *KernelConfigurer) LastActivities() map[string]monotime.Time {
|
|
return nil
|
|
}
|
|
|
|
func (c *KernelConfigurer) fetchStats() (map[string]WGStats, error) {
|
|
stats := make(map[string]WGStats)
|
|
wg, err := wgctrl.New()
|
|
if err != nil {
|
|
return nil, fmt.Errorf("wgctl: %w", err)
|
|
}
|
|
defer func() {
|
|
err = wg.Close()
|
|
if err != nil {
|
|
log.Errorf("Got error while closing wgctl: %v", err)
|
|
}
|
|
}()
|
|
|
|
wgDevice, err := wg.Device(c.deviceName)
|
|
if err != nil {
|
|
return nil, fmt.Errorf("get device %s: %w", c.deviceName, err)
|
|
}
|
|
|
|
for _, peer := range wgDevice.Peers {
|
|
stats[peer.PublicKey.String()] = WGStats{
|
|
LastHandshake: peer.LastHandshakeTime,
|
|
TxBytes: peer.TransmitBytes,
|
|
RxBytes: peer.ReceiveBytes,
|
|
}
|
|
}
|
|
return stats, nil
|
|
}
|