Commit Graph
143 Commits
Author SHA1 Message Date
f0a40e4395 [client, management] Harden the certificate posture client and keep challenge nonces fresh (#8052)
* implement certificate posture check

* log signal address

* add keychain and cert store support

* read the console user's keychain through a user session helper

A root daemon cannot reach a login keychain: securityd is per session and a
key ACL needs a session to prompt in, so dropping uid is not enough. The
daemon now answers certificate challenges from the System keychain itself,
where MDM installs device identities, and launches "netbird posture
cert-proof" into the console user's desktop session with launchctl asuser
for the login keychain. Only the signature and the chain cross back, never
the private key.

The console user comes from SCDynamicStoreCopyConsoleUser, bound with purego
like the keychain calls. The login window reports no user, root, or
"loginwindow", and all three are treated as no keychain to read, so a Mac at
the lock screen sends device proofs alone.

Adds info logging across the path: the keychain search list, per class query
status and item counts, the chain built per candidate, and the verification
error for every rejected candidate. A run that sends nothing now says why.

README.md documents the trust model, the console user limitation and how to
read the logs.

* read the signed-in user's certificate store on Windows

A service reads LocalMachine\MY, where AD and Intune enrol device
certificates. CurrentUser\MY lives in the signed-in user's registry hive
with keys protected against their profile, and a service that opens it does
not fail: "current user" resolves to HKU\S-1-5-18, so it silently reads the
service account's own empty store. The service therefore reads the machine
store itself and launches "netbird posture cert-proof" with the session
token for the rest, mirroring the macOS console user helper.

Windows lets a privileged service assume a user identity, so the token goes
straight into the child process and no external tooling is involved.
CREATE_NO_WINDOW keeps a console window from flashing on the desktop every
sync. In-process impersonation would also work but is per OS thread while
goroutines migrate, so the child process avoids that class of bug.

Session selection prefers the physical console and falls back to any active
session, so remote desktop and VDI hosts are covered. WTSQueryUserToken
needs SE_TCB_NAME, so a user-run client skips the helper and reads the
machine store alone.

SystemStore takes a store location, gaining NewUserStore alongside
NewSystemStore and the per candidate logging macOS already had. The request
building and proof merging move to helper_spawn.go, shared by both
platforms, and helperStore picks what the helper reads per platform.

* start TPM support

* split goreleaser to support pkcs11 and exclude on docker

* update goreleaser

* go mod tidy

* add tpm pin to netbird config

* split cert and key location and allow key lookup on tpm

* add unsupported flag for mobile devices

* Isolate the cert proof helper from the service environment and cap its output

* Read the PKCS#11 token PIN from NB_TPM_PIN instead of the profile config

* Bound certificate proof collection so a stuck token or keychain cannot hold the sync loop

* Stop retrying a PKCS#11 PIN the token rejected

* Log certificate posture details at debug level

* Sign only nonces and peer keys of the size management issues

* Skip certificate files whose key belongs to another certificate

* Bound PKCS#11 driver sizes, pin template values, and log out only a login the session owns

* Never pass NULL to CFRelease and skip unreadable keychain identities

* Keep the macOS keychain code out of iOS and the PKCS#11 driver out of Android

* Find a chain to each challenge's CAs through every intermediate the store holds

* Require a token label whenever a PKCS#11 PIN is set

* Read user certificates only from the session of the active profile's owner

* Collect certificate proofs again when the owner's session changes and report lost proofs

* Test the PKCS#11 build against SoftHSM in CI and warn once where the build has no driver

* Document where an inline PKCS#11 PIN is stored and how it is protected

* Refuse PKCS#11 URIs that this client cannot honour instead of widening the match

* Trust certificate and key files only when no other user can write or redirect them

* Explain a Windows certificate whose key only a legacy CryptoAPI provider holds

* Use platform absolute module paths in tests and add a real owner session test for Windows

* Match the Windows profile owner by name instead of resolving it through the domain controller

* Keep the certificate stores and TPM library out of the WebAssembly build

* [client] Read TSS2 key files on go-tpm, checked against the library it replaces

The TSS2 parser was the only reason this repository depended on a crypto suite
whose own build tooling it inherits. The replacement sits on go-tpm, which was
already a direct dependency and is in fact what that suite calls underneath, so
this removes a wrapper rather than porting onto a different library: the load,
the derived storage root key and the signing commands are the same calls.

Swapping a parser on the one path a customer actually runs is not something to
assert, so the two are held side by side for this commit. One test feeds the
replacement bytes the old library wrote and requires the same key type, empty
auth flag, parent handle, blobs and decoded public key; the other feeds both the
fixtures the tests are built on, so those are the shape the format calls for and
not merely the shape the new parser reads. The scaffolding goes away with the
dependency in the commit that follows.

The encoder behind the fixtures is written out separately from the parser under
test, so an encoder bug and a decoder bug cannot cancel each other out.

* [client] Drop go.step.sm/crypto and the repo-wide upgrades it imposed

The TSS2 parser was the only thing in the repository that used this module, and
it brought 302 modules into the graph to do it — 35 of them linters, along with
Google Cloud KMS and IAM, the AWS SDK and a terminal styling library. Those are
the module's own development dependencies, which minimal version selection turns
into floors in ours, and they are the whole reason gRPC, protobuf, the AWS SDK,
OpenTelemetry, logrus and five x/ packages had moved. Management, signal, relay
and proxy inherited every one of them for a feature none of them runs.

Removing the import is not enough, because tidy never downgrades: the raised
floors stay written in go.mod. Each one is pinned back to the version main had,
then tidy is left to raise again whatever something still genuinely needs. It
raised nothing: all 43 are back where they were, and go-tpm was already in the
graph at the same version, so the certificate feature now costs no new module at
all.

The differential tests go with it. They existed to check the swap against the
library while both were present, and there is nothing left to compare against.

* [client] Clear the lint findings only the macOS and Windows runners see

golangci-lint analyses one build at a time, so running it on Linux says nothing
about the two platforms CI also lints. Against those builds the feature's
packages reported eight findings, and the structural one is Config.dir: it is
dead on macOS and Windows because neither reads a directory at all, their
collectors take the configuration and discard it. Moving the method beside its
only callers makes that visible in the layout instead of in a linter, and leaves
the gap itself — no file or token store on those platforms — where it belongs,
as something to decide rather than something to silence.

An absent key file beside a certificate was reported as a nil signer with a nil
error, which the caller then had to recognise by its nilness. It is a sentinel
now, so the meaning is in the error rather than in the absence of one.

The rest follow the standard library: the elliptic coordinates and the private
scalar come from the encoding helpers rather than the deprecated big.Int fields,
and an error string loses its trailing colon.

Lint is clean on linux, darwin and windows; the hardware TPM path was exercised
separately against a real device and passes.

* Accept the TSS2 emptyAuth boolean OpenSSL writes and persistent parents on 32-bit builds

* Count the certificates field in the peer meta store test

* Check the store directory before listing it, refuse group-writable files, and reject a URI with two PIN sources

* Share a PKCS#11 login between sessions and send each PIN at most once at a time

* Collect certificate proofs again when the meta sync carrying them failed

* Use no Windows user store when a domainless owner matches accounts of several domains

* Use no user certificate store when the active profile's owner cannot be read

* Document the PIN sources on CertPKCS11URI and keep the README PIN example off the command line

* Test that the PKCS#11 URI stays out of the debug bundle and run the wrong-PIN test only on a disposable token

* Refuse a TPM PSS signature request for the maximum salt length

* Add the certificate fields to the network map golden data

* Retry posture checks whose meta sync timed out instead of dropping them

* Start no system info gathering while a timed-out one is still running

* Guard the applied posture checks across goroutines and keep refreshing proofs while a pending update times out

* Log what a successful certificate proof helper wrote to stderr

* Send recollected certificate proofs to management only when the proven chains changed

* Explain a macOS keychain key whose access list does not allow netbird

* Kill the whole macOS certificate helper process group when it times out

* End sudo option parsing before the macOS certificate helper binary

* Hold off system info gathering only while a timed-out one is still running

* Collect certificate proofs on the posture watcher instead of under the sync lock

* Read the certificate store directory and PKCS#11 URI from the daemon environment, not the profile config

* Install the RPM sysconfig file readable by root only and show the certificate posture variables

* Move the certificate posture README into the package doc and the docs site

* Name NB_CERT_PKCS11_URI in the PIN-without-token error

* Keep the file check results of the latest-started system info refresh

* Give the full import command for a keychain key netbird may not use, and correct the package doc

* Restrict the service environment file to root on every package install

* Search only the System keychain in the macOS daemon and only the login keychain in the user helper

* Let the certificate proof helper read the PKCS#11 token from the environment on Linux

* Ask a macOS user's keychain again only after an hour when it proved nothing

* Clear the lint findings in certificate posture

* Hold off the keychain helper only after a completed or timed-out run, independent of CA order

* Keep free functions out of the method lists of PKCS11Store, URI and Challenger

* Name the post-install permission helper in snake case and shorten the sysconfig certificate block

* Drop the certificate store directory from certproof.Config, which only NB_CERT_STORE_DIR sets

* [management] Renew certificate challenge nonces on quiet accounts

A certificate challenge nonce is accepted for its own window and the one before
it, and it only reaches a peer attached to a network map. An account where
nothing changes sends no map, so after a day the peer re-sends the nonce it
still holds, verification rejects its whole proof set, and the certificates
stored for it are dropped. It fails the certificate check and loses every policy
gated on it until some unrelated change happens to push a map. The outage
repairs itself in seconds, which is what makes it expensive: it is intermittent,
it only hits stable networks, and it is not reproducible on demand.

Push the account's peers an update often enough that the nonce they hold is
never close to expiring. Only accounts whose posture checks actually ask for a
certificate are tracked, so a deployment without the feature does no extra work.

The refresh runs from one goroutine over a map of accounts rather than a timer
per account: the period is hours, so one pass every few minutes costs nothing
next to it, and there is no timer to re-arm when an account that falls due
sooner appears. Each account's first run is offset by a hash of its ID, because
the challenge window is global and an instance restart would otherwise arm every
account in the same moment.

The push carries no administrative change, so it is counted as a refresh rather
than an update and stays out of the figures that track what was edited.

(cherry picked from commit 7ad4a0df37)

* [management] Make the certificate challenge window one knob to turn

Renewal was timed against the window in two different ways: the period derived
from it, the sweep interval did not. Shortening the window to watch a renewal in
an end-to-end run would have left the refresher still looking for due accounts
every quarter of an hour, so nothing would have been renewed in time and the
test would have reported the feature broken.

Derive the sweep from the period, within bounds that keep a very short window
from spinning and a normal one from checking less often than is useful, and
allow the window itself to be set through the environment so a run can take
seconds instead of half a day. A value that cannot be parsed or falls outside
the bounds keeps the default, because a window nobody intended is a security
property nobody chose, and an override is logged at warning level since it sets
how long a device keeps passing the check after its key is gone.

Every instance has to be given the same value: the window is part of the nonce,
so instances that disagree reject each other's.

(cherry picked from commit 0e38fcf409)

* [management] Pin the property that makes per-peer nonce state unnecessary

A nonce carries the window it was minted in, not the instant, and is accepted
for that window and the one before it. So a peer re-stamped at least once per
window can never be left holding one outside the accepted pair, whenever it was
last served and however much life its own nonce had left. That is the whole
reason management tracks nothing per peer, and it was resting on an argument
rather than a test.

The phases are part of the property, not decoration: accounts are deliberately
given a refresh phase of their own, so the guarantee has to hold off the window
boundary too. The negative case shows why that matters — a cadence of exactly
two windows lands inside the grace window when it is aligned to the boundary and
leaves a gap when it is not.

(cherry picked from commit dee68facfd)

* [management] Renew challenges only for the peers that answer one

The refresh pushed an update to every connected peer of the account, while only
the peers a certificate check applies to carry a nonce. On an account where a
handful of peers sit behind the check and the rest do not, everyone was woken
several times a day to be handed a map that changed nothing for them.

Push to the sources of the enabled policies whose posture checks include a
certificate check, which is exactly the set that is sent a challenge.

Resolving the set the other way round than the gRPC layer does is the risk here:
a peer the refresh forgets stops being renewed and falls out of its policies
silently, which is the failure this whole mechanism exists to prevent. So the
selection is held against processPeerPostureChecks, the per-peer rule that
decides who receives a challenge in the first place, by a test that asks both
the same question and requires the same answer.

(cherry picked from commit dc4d0e0274)

* [management] Derive certificate challenge nonces from the stored encryption key

The nonce secret came from the server's WireGuard key, which is generated afresh
in every process and never persisted. A nonce carries no state, so the only
thing that lets one instance verify what another issued is deriving the same
secret — and that premise, written in the comment above the challenger, was not
met: every instance had its own key.

A peer reconnecting after a restart therefore presented a nonce minted under the
previous secret, verification failed with a mismatch, its whole proof set was
rejected and the certificates stored for it were dropped until it signed again.
Reproduced three times on the lab, each one logging "nonce was not issued to
this peer", which only a changed secret produces. On a single instance it costs
seconds of lost policy access per restart; across instances it is not transient
at all, because every reconnect that lands elsewhere is rejected the same way.

Derive from the data store encryption key instead: it is generated once, written
back to the configuration and read by every instance, so it survives restarts
and is shared. Where none is configured the secret falls back to the WireGuard
key with a warning — degraded but still unpredictable, which is the property
that matters most: a peer able to guess it could mint the nonces of future
windows, sign them while its key is present and keep passing after it is gone.

The challenger is now built once and passed to the two places that need it,
rather than re-derived per message.

(cherry picked from commit 278f2f3807)

* [management] Register an account for renewal where its nonce is issued

Renewal was armed when a peer connected or when a posture check was saved, both
of which ask the store whether the account has a certificate check. That misses
the case it most needs to catch: the check is created through one instance while
the peers are connected to another, so the instance serving them never learns it
has anything to renew and their nonce expires. It also charged a query to every
peer connect in every account, including the ones that will never use the
feature, which a fleet reconnecting after a restart pays all at once.

Register where the nonce is actually stamped instead. A nonce is verified from a
shared secret and so travels between instances, but the renewal that keeps it
fresh cannot: only the instance holding a peer's stream can push to it. Issuing
and renewing now line up by construction — an instance renews exactly the
accounts it has issued nonces for — and an instance that never issues one has
nothing to renew, so there is no case left to miss.

The registration is a map insert with no store access, which is what lets it sit
on a path taken by every login and every initial sync.

Reported by Viktor Liu, who also proposed registering at the point of issue.

(cherry picked from commit 2d16dd7d7cf54762f2e64c5630ea092f32ef63ab)

* [management] Register for renewal on pushed updates, not only on connect

Registering where the nonce is stamped only covered the login and the initial
sync, which both happen when a peer opens a stream. That left out the path the
mechanism exists for.

On the cloud the network map controller is wrapped so that an update publishes
to an event bus instead of pushing locally: an instance handling a REST change
broadcasts, and every instance holding a peer of that account pushes to its own.
Those pushes stamp a nonce through the update handler, and nothing there
registered, so an instance learned about an account only when one of its peers
happened to reconnect. For a quiet fleet that is the original bug: the check is
created, the peers are told about it, and nobody renews what they were told.

Registering on the pushed update closes it, and is the difference between
stamping and marking a peer connected — one happens on every push, the other
only when a stream opens. Reported by Viktor Liu; the broadcast that makes it
work was pointed out by Pascal Fischer.

(cherry picked from commit 59efe8d93e53bacdf57cb546f4ab2c19dc4eddab)

* [management] Let the challenge refresh loop stop with the manager that owns it

The loop was started on a context explicitly detached from the caller's, so
nothing could ever stop it. Production is unaffected either way, since
BuildManager is called with context.Background(), but a test that builds a
manager leaked a sweeping goroutine for the rest of the run, and a shutdown
path added later would have had no way to reach it.

Take the manager's context as the request buffer built on the line above
already does. The test pins the contract the loop offers, so a detached
context cannot come back inside Start either.

* [management] Bound one account's challenge refresh so it cannot starve the rest

Resolving which peers answer a challenge reads the store three times, and the
refresher sweeps accounts one after another on a single goroutine. A read that
never returns held the sweep for the life of the process, so every other
account on the instance stopped being renewed and its peers fell out of the
policies gated on the check: one account's bad luck became an outage for all
of them.

Give each refresh the sweep interval it is allowed to occupy, capped at 30s so
a 12-hour window does not grant minutes to a query that should take
milliseconds. A refresh that runs out of time keeps its account tracked, since
a deadline says nothing about whether that account still has a certificate
check.

* [management] Send challenge refreshes down the path the rest of management uses

The refresh dispatched through UpdateAffectedPeers, the one variant that takes
no reason, so it was missing from the update counters and coalesced with
nothing. An administrator editing a policy while the sweep ran made the
account's network map twice over, and UpdateOperationRefresh, added for
exactly this caller, was never referenced.

Buffer it with a posture_check/refresh reason instead. The periodic push is
now visible in the metrics as what it is, distinct from an edit, and the send
detaches from the sweep deadline on its own, so that deadline bounds the store
reads it was meant for.

* [management] Keep the certificate challenge comments to what the history does not say

Four of these ran to three and four times the comment budget, the longest at
992 characters. Most of the excess argued against designs that were never
written or explained a bug that no longer exists in the code, which is what
the commit that fixed it is for.

What is left is the part a reader cannot recover from the code: that the
nonce secret has to be persisted and unpredictable, that stamping and
renewing are decided together because only the serving instance can push, and
that the target rule is the inverse of processPeerPostureChecks.

* Keep the newest posture checks pending whatever made their meta sync fail

* Report no lost certificate when the engine stops during a proof collection

* Share the proof collection single-flight across engine restarts

* Close a PKCS#11 module that loads but cannot be used

* Fix the pending checks comments

* Renew certificate challenges only for the peers streamed to this instance

* Ignore a challenge stamp from an older sync stream of the same peer

* Kill the Windows certificate proof helper with its whole process tree

* Expect the challenge untrack in the session ownership test

* Drop an invalid certificate proof without discarding the valid ones

* Start a system info gathering beside one that has been stuck for ten timeouts

---------

Co-authored-by: pascal <pascal@netbird.io>
Co-authored-by: mlsmaycon <mlsmaycon@gmail.com>
Co-authored-by: riccardom <riccardomanfrin@gmail.com>
2026-10-09 15:42:18 +02:00
Pascal Fischerandmlsmaycon 53a14551c8 [client, management] implement certificate posture check (#7535)
Co-authored-by: mlsmaycon <mlsmaycon@gmail.com>
2026-10-09 14:57:00 +02:00
Riccardo Manfrin 3e85e40be2 [client] Cache the WireGuard interface check shared by ICE agents (#8001)
* [client] Take a WireGuard detector through the interface filter

The interface filter answers whether an interface is a WireGuard device by opening
a wgctrl client and asking for it, and it does that for every interface it is given.
Nothing about that call is tied to the caller, so it can be answered by a shared
object instead of being repeated, but the filter has no way to receive one.

InterfaceFilter and the constructors that build one now take a detector, and the ICE
config carries it so that every agent can be handed the same one. Nobody supplies a
detector yet: a nil one probes on every call, which is what the filter did before, so
this changes no behaviour.

* [client] Share one WireGuard detector across every ICE agent

Creating an ICE agent builds two interface filters, one for the agent and one for
the transport net it sits on, and each is asked about every host interface. For an
interface the disallow list does not settle, answering means opening a wgctrl client,
which builds a kernel and a userspace client and resolves the netlink family, and
then a round trip that usually just reports the device does not exist. An agent is
created per peer connection attempt, so on a large network that runs constantly:
on a routing peer with ~16000 peers it measured 2.40s of a 66.59s CPU profile, 3.6%,
split evenly between opening the client and the round trip.

The engine now owns a detector and passes it to every agent through the ICE config,
so the answer for an interface is reused instead of being asked again for each agent.
It is kept for a second, short enough that a WireGuard interface appearing is picked
up before ICE settles on candidates over it.

The callers that build one filter and keep it, the relay and the UDP mux, keep
passing nil and so keep probing, which costs them nothing at their rate.

* [client] Recheck the WireGuard cache inside the singleflight group

A caller that saw an expired entry could enter the singleflight group
after another caller had already refreshed the entry and left it, and
probe the interface a second time. Read the cache again inside the group
before probing.

This also makes the concurrent probe test independent of scheduling: a
late caller finds the fresh entry instead of starting a new probe.

* [client] Drop expired WireGuard detector entries

The detector lives as long as the engine and kept an entry for every
interface name it was ever asked about. On hosts that churn interfaces,
such as container veths, the map only grew. Remove expired entries when
a new answer is stored; the map holds a few dozen names at most, so the
sweep is cheap and runs at most once per interface per TTL.

* [client] Skip the disallow-list filter test on iOS

InterfaceFilter does not apply the disallow list on iOS, so the subtest
reaches the probe there and its no-probe assertion cannot hold.
2026-10-09 11:03:34 +02:00
Zoltan Papp 1c7d87d5fc [client,android] Generate debug bundle to file (#7528)
* [client] Add a debug bundle file export to the Android bridge

The Android app can only upload a debug bundle and hand the user a key.
Users who want to inspect what leaves their device before sharing it
have no way to get the zip itself. Add DebugBundleFile, which generates
the bundle into the cache directory and returns its path instead of
uploading; the app copies it wherever the user chose and removes it.

DebugBundle keeps its behavior. Both entry points share the unexported
debugBundle with an upload switch, so the body stays where it was and
merges cleanly with the MDM overlay change on main.

Because the file variant leaves the zip to the caller and the upload
variant only removes it after the upload finishes, a process killed in
between leaves a zip behind in the cache. Remove stale bundles before
generating a new one: RemoveStaleBundles deletes zips matching the
generator's pattern that are older than an hour. Remote debug jobs write
to the same directory, so younger files are treated as still in use.

* Preserve network map for debug bundle on Android

* [client] Keep exported Android debug bundles out of the stale cleanup

DebugBundleFile hands the zip to the caller, but the file kept the
netbird.debug.*.zip name that RemoveStaleBundles matches, so a later
debug run could delete it once it was older than an hour. Rename the
exported bundle to netbird.debug-file.*.zip after generation so the
cleanup only ever touches bundles no caller owns.

* [client] Warn when a stale debug bundle cannot be removed

A failed removal means bundles pile up in the cache directory, so log it
at Warn instead of Debug. A file that is already gone was removed by a
concurrent cleanup and is skipped silently.

* [client] Drop the outdated debugBundle comment

The comment still said the file variant leaves the zip in place, but it
is renamed by debug.ExportBundle since the stale-cleanup change.

* [client] Test that the network map reaches the debug bundle

Cover both halves of the path Android now relies on: the engine keeps
the latest sync response once persistence is enabled, and the bundle
generator writes it to network_map.json (anonymized or not) and omits
the file when there is no sync response.

* [client] Remove abandoned exported debug bundles after a day

An exported bundle is owned by the caller, but if the app is killed
before it copies and deletes the file, nothing ever removes it from the
cache directory. Let RemoveStaleBundles also match exported bundles,
with a 24 hour max age instead of the caller-provided one, so a bundle
that is still being saved survives while an abandoned one goes.
2026-10-05 14:29:28 +02:00
Riccardo Manfrin 9f8ddc7131 [client] Discover interfaces lazily in stdnet instead of at construction (#7346)
* [client] Discover interfaces lazily in stdnet instead of at construction

stdnet.NewNet and NewNetWithDiscover ended with

    return n, n.UpdateInterfaces()

handing back a non-nil *Net together with the discovery error. Three of the
five call sites (Engine.newWgIface, ice.NewAgent, SingleSocketUDPMux) logged
the error and kept using the instance, which is only safe as long as the
instance still works after a failed discovery.

That stopped being true when Interfaces() gained a lazily refreshed cache:
updateInterfaces sets lastUpdate only on success, so after a failed
construction the 30s cache guard never holds and Interfaces() returns an
error rather than the empty list it used to return. Feeding such an instance
to pion is worse than passing nothing at all - ice.NewAgent falls back to its
own stdnet when Net is nil, and the interface blacklist is applied separately
through AgentConfig.InterfaceFilter, so the fallback loses nothing. Instead,
a transient discovery failure (the Android bridge at boot, or an interface
disappearing between net.Interfaces() and Interface.Addrs()) turned into a
hard "error getting local interfaces" from ice.NewAgent, and aborted the STUN
and TURN probes, which never even need the interface list.

Since the accessors already refresh a stale cache on demand, the eager
discovery in the constructors is redundant: drop it, make both constructors
infallible, and let the discovery error surface at the call that actually
needs the interfaces. UpdateInterfaces had no callers left and is not part of
transport.Net, so it is removed along with it.

InterfaceByIndex and InterfaceByName read the cached slice directly and never
refreshed it, so they would have kept reporting ErrInterfaceNotFound forever
on an instance whose first discovery failed. They now go through the same
refresh path as Interfaces().

* [client] Warm the stdnet interface cache at construction

Moving discovery to first use regressed the privileged suites on the three
platforms that always build an ICE bind: Darwin, FreeBSD and Windows time out
in TestWGIface_UpdateAddr, TestRecreation, TestEngine_SSH and
TestEngine_MultiplePeers, while Linux stays green because a host with the
WireGuard kernel module takes the kernel-device branch and never drives the
mux that asks for interfaces.

interfaceFilter probes with wgctrl every interface the disallow list does not
already exclude. Discovering at construction ran that probe before the caller
had an overlay interface of its own; discovering at first use runs it after,
so on a userspace WireGuard platform the probe reaches the UAPI socket of the
same process. The tests reach it because they construct with a nil disallow
list, where the client passes DefaultInterfaceBlacklist and its own interface
is excluded by prefix.

Restore the original timing with an explicit warm-up. The constructors stay
infallible and the error is still reported by the accessor that needs the
interfaces, so the contract this branch is about is unchanged.

* Revert "[client] Warm the stdnet interface cache at construction"

This reverts commit 947e25288f.

* [client] Give the privileged tests the interface blacklist the client uses

The suites that create a WireGuard interface construct stdnet with a nil
disallow list, which the client never does: Engine passes
profilemanager.DefaultInterfaceBlacklist, whose "wt" and "utun" prefixes
exclude the overlay interface before the filter reaches its wgctrl probe.

With an empty list every interface reaches that probe, the one the test has
just created included, and on a userspace WireGuard platform the probe talks
to the UAPI socket of the same process. That is why Darwin, FreeBSD and
Windows timed out here while Linux, which takes the kernel-device branch on a
host with the module loaded, stayed green.

Pass the blacklist in both suites so they exercise the configuration the
client ships. client/iface declares the prefixes locally because
profilemanager imports it.

Also cover the constructors directly: the existing tests build the struct
literal, so nothing asserted that NewNet and NewNetWithDiscover leave the
cache cold.

* [client] Pass the blacklist in the remaining tests that build an interface

Same reason as the previous commit, four call sites it missed: engine_test,
the route manager and systemops suites, and the privileged DNS server suite
all construct stdnet with a nil disallow list and then create a WireGuard
interface. TestAddVPNRoute surfaced it on FreeBSD once the earlier two files
stopped timing out first.

client/internal/dns declares the prefixes locally; profilemanager imports
that package, so it cannot import profilemanager back.
2026-10-02 15:24:39 +02:00
Viktor Liu 8edc120370 [client] Replace the eBPF WireGuard proxy with loopback endpoint addressing (#7316) 2026-09-30 10:41:49 +02:00
Zoltan Papp 825389818c [client] Gather fresh system info on every management sync stream connect (#7409)
* Gather fresh system info on every management sync stream connect

The engine collected the peer meta once at start and reused the same
Info for every Sync stream reconnect, so a mobile network switch that
redials management kept reporting the old local network addresses.
The peer network range posture check was then evaluated against stale
data until the client restarted.

Sync now takes a gatherer that runs at each stream connect. The
gatherer is cheap: GetInfo plus the cached posture check file results,
kept in the new system.InfoSource, which the engine refreshes whenever
the checks list changes. No process enumeration runs on the reconnect
path.

Also fix the management mock server calling itself instead of SyncFunc.

* Evaluate the login response posture checks before the first sync connect

The engine starts with the checks the login response carried, and the
first sync stream request used to send their evaluated file results.
After moving the gather into InfoSource, the stream opened with an empty
cache and the first sync response did not refill it, because its checks
equal the ones the engine already holds. Desktop peers therefore never
reported process or file posture results.

Seed the cache once before the first connect, where the old gather ran,
so a timed out evaluation still falls through to the address-only info.

* Harden the sync info source against nil callbacks and shared slices

A nil getInfo opens the stream without metadata, as a nil sysInfo did
before. The cached posture results are a copy, so the Info returned by
Refresh cannot alias the snapshot later Current calls report. The
exclusion test asserts the remaining address count so it cannot pass
vacuously on a single-address host.

* Retry a posture check refresh that timed out or failed to sync

The checks list was recorded before the gather ran, so once the gather
timed out or SyncMeta failed, the next sync response carrying the same
list matched the recorded one and nothing retried. The peer kept
reporting the previous posture results until the list changed again.

Record the checks only after the meta reached management, so a failed
cycle is repeated on the next sync response.

* Log the skipped posture refresh, let the mock Sync return errors and deflake the reconnect test

* Drop the nil guard around the sync info callback

* Send the refreshed info on the first sync connect instead of gathering it twice
2026-09-04 15:07:01 +02:00
Viktor Liu 51095cb986 [client, management] Support per-peer lazy connection state and default proxy peers to lazy (#6762)
* Support per-peer lazy connection state and default proxy peers to lazy

* Classify forward targets from incoming config in lazy exclusion

* Set IsUserspaceBind mock so lazy manager starts in engine test

* Skip lazy exclude reconciliation when the set is unchanged

* Keep cached lazy flag when a sync carries no peer config
2026-08-26 09:33:51 +02:00
Viktor Liu 4ef65294e9 [client] Reinject captured first packet on lazy connection activation (#6572) 2026-06-30 11:22:25 +02:00
Zoltan Pappandcoderabbitai[bot] 2d7b309004 [client] Categorize privileged tests behind a build tag and run them in Docker (#6425)
* [client] categorize root/system-mutating tests behind a privileged build tag

Tests that need root or mutate host state (nftables/iptables/DNS, TUN/WireGuard
interfaces, routes, eBPF, SSH/service install) are now gated behind a
//go:build privileged tag. The default `go test ./client/...` runs as a non-root
user with no sudo and leaves host networking untouched; mixed files were split so
pure-logic tests stay in the default suite.

A self-hosting ory/dockertest/v4 harness (client/testutil/privileged) runs the
privileged suite inside a --privileged --cap-add=NET_ADMIN container via
`make test-privileged`; a DOCKER_CI=true guard skips the spawn when already inside
the container. Added `make test-unit` for the host-safe run.

* [client] add PRIV_RUN/PRIV_PKGS filters to the privileged test harness

The dockertest harness now reads two optional env vars when building the
in-container `go test` command: PRIV_RUN adds a -run test-name filter and
PRIV_PKGS overrides the package list. Both empty reproduce the full privileged
suite, so CI and `make test-privileged` behave as before. Lets a developer run a
single privileged test in the container, e.g.:

  PRIV_RUN=TestNftablesManager PRIV_PKGS=./client/firewall/nftables/... make test-privileged

* [client] fix unused-helper lint after the privileged test split

Splitting privileged tests into *_privileged_test.go left their shared helpers in
the untagged files, so in the default (no-tag) build they had no callers and
golangci-lint flagged them as unused.

Moved the privileged-only helpers into the privileged files next to their callers
(generateDummyHandler; createEngine/startSignal/startManagement/getConnectedPeers/
getPeers + kaep/kasp; (*mockDaemon).setJWTToken). Annotated the shared routing-test
fixtures that must stay untagged for cross-platform compilation with //nolint:unused
(systemops_bsd expected* vars, ensureIPv6DefaultRoute on bsd/windows,
loopbackIfaceWindows), matching the existing linux variant.

* [client] fix privileged test CI failures and run the harness on macOS

The host-safe unit run dropped sudo but two privileged test groups were
never tagged, and the Docker privileged job silently never ran the suite:

- Gate the ssh/server PrivilegeDropper command-construction tests behind
  the privileged tag (they require root to target a different UID); split
  them into executor_unix_privileged_test.go.
- Tag sharedsock raw-socket tests privileged (need CAP_NET_RAW).
- Fix the Docker job command: nested single quotes around the build tags
  closed the sh -c wrapper early, dropping the go list package set and the
  privileged tag, so go test ran on the empty repo root. Use double quotes.

Make the self-hosting harness usable from a dev Mac:

- Build it on darwin as well as linux; it only drives Docker.
- Resolve the active docker context endpoint into DOCKER_HOST when the
  default /var/run/docker.sock is absent (Docker Desktop, Colima, OrbStack).
- Rename the misspelled containerGoModache constant to containerGoModCache.

* Update client/internal/engine_privileged_test.go

Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>

* Update client/internal/routemanager/systemops/systemops_linux_test.go

Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>

* Update client/internal/routemanager/systemops/systemops_windows_test.go

Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>

* Update client/server/server_privileged_test.go

Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>

* [ci] Run privileged-tagged tests on darwin, windows and freebsd

The privileged build tag split moved root/system-mutating tests behind
//go:build privileged, but only the linux docker job was given the tag.
The native darwin (sudo), windows (PsExec64 -s) and freebsd VM runners
already have the required privileges, so add the privileged tag there too
to keep CI running the same set of tests as before the split.

* [ci] Exclude dockertest harness from the darwin privileged run

The privileged tag now compiles client/testutil/privileged on darwin, whose
TestRunPrivilegedSuiteInDocker spawns a container the macOS runner has no
Docker for. Exclude the harness package from the darwin list, matching the
linux job, so the privileged tests run in place without a container spawn.

---------

Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2026-06-28 16:15:54 +02:00
Zoltan Papp ac9529ea8c [client] Fix engine lifecyrcle race (#6443)
* [client] always clean up on Engine.Start failure via defer

The rosenpass init paths (NewManager/Run) returned without calling
e.close(), leaking the WireGuard interface and other partially
initialized state on failure. Per-branch cleanup was easy to miss when
adding new early returns.

Convert Start to a named error return and tear down via a single defer
that calls e.close() whenever err != nil, removing the scattered
per-branch close() calls (including the redundant one in initFirewall).

* [client] make Engine single-use and guard against double Start

Create the run context once in NewEngine instead of in Start. This
keeps e.cancel valid for the engine's whole lifetime, so Stop can
cancel a Start that is blocked waiting on the network while holding
syncMsgMux: Stop now cancels before taking the lock, unblocking that
Start so it can release the mutex.

Reject re-entry into Start: a non-nil wgInterface means a prior Start
already ran (ErrEngineAlreadyStarted), and a cancelled run context
means the engine was stopped (ErrEngineAlreadyStopped). Both checks run
before the cleanup defer so a duplicate call cannot tear down the
running engine's state.

* [client] let engine context unblock WaitStreamConnected

WaitStreamConnected only watched the signal client's own context, which
derives from the parent engineCtx rather than the engine's run context.
A Start blocked here (signal stream not yet up) could therefore not be
released by Engine.Stop, since Stop only cancels the engine's run
context.

Pass a context into WaitStreamConnected and select on it too, and have
the engine pass e.ctx, so Stop cancelling e.ctx unblocks a parked Start.
Update the Client interface, the mock, and callers accordingly.

* [client] fix Start/Stop race by making the run loop own engine shutdown

ConnectClient.Stop stopped the engine directly while the run loop's
backoff cycle could still be starting an engine, so Engine.close raced
Engine.Start (e.g. firewall setup reading wgInterface while close nils
it). embed.Client.Start's rollback only avoided a deadlock by cancelling
before Stop; the race itself remained and was caught by -race.

Make the run loop the sole owner of engine shutdown: derive the run
context in NewConnectClient, and have Stop cancel it and wait for the
loop to exit (skipping the wait when the loop never ran) instead of
calling engine.Stop. The loop now always stops the engine on its way
out, dropping the unsynchronised wgInterface check it used to guard that
call. Self-calls from within the loop use runCancel to avoid waiting on
themselves.

embed keeps a defensive pre-Stop cancel(); the daemon's cleanupConnection
gets a TODO to adopt Stop() rather than stopping the engine in parallel.

* [client] init context state in engine tests

Engine tests built the engine context with context.WithCancel(
context.Background()), omitting CtxInitState. Now that the run context
is created in the constructor, the wgIfaceMonitor goroutine can reach
triggerClientRestart during teardown, which calls CtxGetState and
panics on the missing state. Real entry points (up, embed, service)
always CtxInitState; only the tests skipped it.

* [client] interrupt connect backoff on context cancel

The run loop retried with a raw ExponentialBackOff, so a backoff sleep
ignored context cancellation. Now that ConnectClient.Stop waits for the
run loop to exit, a cancel landing during a sleep would block Stop for
the full interval (up to MaxInterval). Wrap the backoff with the run
context so Retry returns promptly on cancel; the retry budget itself
(MaxElapsedTime) is unchanged.

* [client] bound WaitStreamConnected in signal client tests

The tests waited on WaitStreamConnected with context.Background() and the
client's own context was also Background, so a stream that never connects
would hang until the suite timeout. Pass a 5s timeout context and assert
StreamConnected afterwards so the tests fail fast with a clear reason.

* [client] fix WaitStreamConnected stale-channel race

The StreamConnected check and the wait-channel creation took the mutex
separately, so notifyStreamConnected could set the status and close/clear
connectedCh in between: the waiter then created a fresh channel nobody
would ever close and blocked forever. Also, the status read was unlocked
while notify wrote it under the mutex (a data race). Do the check and the
channel fetch in one locked section; drop the now-unused
getStreamStatusChan helper. Pre-existing bug, not introduced by this branch.

* [client] abort Start if context cancelled while waiting for signal stream

receiveSignalEvents blocks in WaitStreamConnected until the signal stream
connects or the context is cancelled. If Stop cancelled e.ctx while Start
was parked there, Start kept going: it started the remaining subsystems on
a cancelled context and marked a shutting-down engine as started. Return
the context error from receiveSignalEvents and propagate it from Start, so
the deferred cleanup runs and the cancellation reaches the caller.

* [client] clean up all started components on Start failure

Start's failure defer only called close(), which covers the wg interface,
firewall, rosenpass and port forwarding but leaves connMgr, srWatcher,
route/DNS/flow/state managers and the monitor goroutines running. A late
failure (e.g. the context-cancelled check after the signal stream) thus
leaked them.

Extract Stop's locked teardown into stopLocked (caller holds syncMsgMux,
does not wait on shutdownWg) and call it from both Stop and Start's defer.
The defer also cancels the run context first so goroutines started before
the failure unwind. Teardown order is unchanged.
2026-06-22 13:52:57 +02:00
Bethuel Mmbaga 14af179556 [management] Refactor management server bootstrap (#6256) 2026-05-26 17:44:28 +03:00
Viktor Liu 205ebcfda2 [management, client] Add IPv6 overlay support (#5631) 2026-05-07 11:33:37 +02:00
Bethuel Mmbaga df197d5001 [management] Prevent JWT reuse during peer login (#6002) 2026-04-29 15:04:27 +03:00
Maycon Santos 53b04e512a [management] Reuse a single cache store across all management server consumers (#5889)
* Add support for legacy IDP cache environment variable

* Centralize cache store creation to reuse a single Redis connection pool

Each cache consumer (IDP cache, token store, PKCE store, secrets manager,
EDR validator) was independently calling NewStore, creating separate Redis
clients with their own connection pools — up to 1400 potential connections
from a single management server process.

Introduce a shared CacheStore() singleton on BaseServer that creates one
store at boot and injects it into all consumers. Consumer constructors now
receive a store.StoreInterface instead of creating their own.

For Redis mode, all consumers share one connection pool (1000 max conns).
For in-memory mode, all consumers share one GoCache instance.

* Update management-integrations module to latest version

* sync go.sum

* Export `GetAddrFromEnv` to allow reuse across packages

* Update management-integrations module version in go.mod and go.sum

* Update management-integrations module version in go.mod and go.sum
2026-04-16 16:04:53 +02:00
Zoltan Papp 0efef671d7 [client] Unexport GetServerPublicKey, add HealthCheck method (#5735)
* Unexport GetServerPublicKey, add HealthCheck method

Internalize server key fetching into Login, Register,
GetDeviceAuthorizationFlow, and GetPKCEAuthorizationFlow methods,
removing the need for callers to fetch and pass the key separately.

Replace the exported GetServerPublicKey with a HealthCheck() error
method for connection validation, keeping IsHealthy() bool for
non-blocking background monitoring.

Fix test encryption to use correct key pairs (client public key as
remotePubKey instead of server private key).

* Refactor `doMgmLogin` to return only error, removing unused response
2026-04-07 12:18:21 +02:00
Zoltan PappandViktor Liu 91f0d5cefd [client] Feature/client metrics (#5512)
* Add client metrics

* Add client metrics system with OpenTelemetry and VictoriaMetrics support

Implements a comprehensive client metrics system to track peer connection
stages and performance. The system supports multiple backend implementations
(OpenTelemetry, VictoriaMetrics, and no-op) and tracks detailed connection
stage durations from creation through WireGuard handshake.

Key changes:
- Add metrics package with pluggable backend implementations
- Implement OpenTelemetry metrics backend
- Implement VictoriaMetrics metrics backend
- Add no-op metrics implementation for disabled state
- Track connection stages: creation, semaphore, signaling, connection ready, and WireGuard handshake
- Move WireGuard watcher functionality to conn.go
- Refactor engine to integrate metrics tracking
- Add metrics export endpoint in debug server

* Add signaling metrics tracking for initial and reconnection attempts

* Reset connection stage timestamps during reconnections to exclude unnecessary metrics tracking

* Delete otel lib from client

* Update unit tests

* Invoke callback on handshake success in WireGuard watcher

* Add Netbird version tracking to client metrics

Integrate Netbird version into VictoriaMetrics backend and metrics labels. Update `ClientMetrics` constructor and metric name formatting to include version information.

* Add sync duration tracking to client metrics

Introduce `RecordSyncDuration` for measuring sync message processing time. Update all metrics implementations (VictoriaMetrics, no-op) to support the new method. Refactor `ClientMetrics` to use `AgentInfo` for static agent data.

* Remove no-op metrics implementation and simplify ClientMetrics constructor

Eliminate unused `noopMetrics` and refactor `ClientMetrics` to always use the VictoriaMetrics implementation. Update associated logic to reflect these changes.

* Add total duration tracking for connection attempts

Calculate total duration for both initial connections and reconnections, accounting for different timestamp scenarios. Update `Export` method to include Prometheus HELP comments.

* Add metrics push support to VictoriaMetrics integration

* [client] anchor connection metrics to first signal received

* Remove creation_to_semaphore connection stage metric

The semaphore queuing stage (Created → SemaphoreAcquired) is no longer
tracked. Connection metrics now start from SignalingReceived. Updated
docs and Grafana dashboard accordingly.

* [client] Add remote push config for metrics with version-based eligibility

Introduce remoteconfig.Manager that fetches a remote JSON config to control
metrics push interval and restrict pushing to a specific agent version
range. When NB_METRICS_INTERVAL is set, remote config is bypassed
entirely for local override.

* [client] Add WASM-compatible NewClientMetrics implementation

Replace NewClientMetrics in metrics.go with a WASM-specific stub in metrics_js.go, returning nil for compatibility with JS builds. Simplify method usage for WASM targets.

* Add missing file

* Update default case in DeploymentType.String to return "unknown" instead of "selfhosted"

* [client] Rework metrics to use timestamped samples instead of histograms

Replace cumulative Prometheus histograms with timestamped point-in-time
samples that are pushed once and cleared. This fixes metrics for sparse
events (connections/syncs that happen once at startup) where rate() and
increase() produced incorrect or empty results.

Changes:
- Switch from VictoriaMetrics histogram library to raw Prometheus text
  format with explicit millisecond timestamps
- Reset samples after successful push (no resending stale data)
- Rename connection_to_handshake → connection_to_wg_handshake
- Add netbird_peer_connection_count metric for ICE vs Relay tracking
- Simplify dashboard: point-based scatter plots, donut pie chart
- Add maxStalenessInterval=1m to VictoriaMetrics to prevent forward-fill
- Fix deployment_type Unknown returning "selfhosted" instead of "unknown"
- Fix inverted shouldPush condition in push.go

* [client] Add InfluxDB metrics backend alongside VictoriaMetrics

Add influxdb.go with timestamped line protocol export for sparse
one-shot events. Restore victoria.go to use proper Prometheus
histograms. Update Grafana dashboards, add InfluxDB datasource,
and update docs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* [client] Fix metrics issues and update dev docker setup

- Fix StopPush not clearing push state, preventing restart
- Fix race condition reading currentConnPriority without lock in recordConnectionMetrics
- Fix stale comment referencing old metrics server URL
- Update docker-compose for InfluxDB: add scoped tokens, .env config, init scripts
- Rename docker-compose.victoria.yml to docker-compose.yml

* [client] Add anonymised peer tracking to pushed metrics

Introduce peer_id and connection_pair_id tags to InfluxDB metrics.
Public keys are hashed (truncated SHA-256) for anonymisation. The
connection pair ID is deterministic regardless of which side computes
it, enabling deduplication of reconnections in the ICE vs Relay
dashboard. Also pin Grafana to v11.6.0 for file-based provisioning
and fix datasource UID references.

* Remove unused dependencies from go.mod and go.sum

* Refactor InfluxDB ingest pipeline: extract validation logic

- Move line validation logic to `validateLine` and `validateField` helper functions.
- Improve error handling with structured validation and clearer separation of concerns.
- Add stderr redirection for error messages in `create-tokens.sh`.

* Set non-root user in Dockerfile for Ingest service

* Fix Windows CI: command line too long

* Remove Victoria metrics

* Add hashed peer ID as Authorization header in metrics push

* Revert influxdb in docker compose

* Enable gzip compression and authorization validation for metrics push and ingest

* Reducate code of complexity

* Update debug documentation to include metrics.txt description

* Increase `maxBodySize` limit to 50 MB and update gzip reader wrapping logic

* Refactor deployment type detection to use URL parsing for improved accuracy

* Update readme

* Throttle remote config retries on fetch failure

* Preserve first WG handshake timestamp, ignore rekeys

* Skip adding empty metrics.txt to debug bundle in debug mode

* Update default metrics server URL to https://ingest.netbird.io

* Atomic metrics export-and-reset to prevent sample loss between Export and Reset calls

* Fix doc

* Refactor Push configuration to improve clarity and enforce minimum push interval

* Remove `minPushInterval` and update push interval validation logic

* Revert ExportAndReset, it is acceptable data loss

* Fix metrics review issues: rename env var, remove stale infra, add tests

- Rename NB_METRICS_ENABLED to NB_METRICS_PUSH_ENABLED to clarify that
  collection is always active (for debug bundles) and only push is opt-in
- Change default config URL from staging to production (ingest.netbird.io)
- Delete broken Prometheus dashboard (used non-existent metric names)
- Delete unused VictoriaMetrics datasource config
- Replace committed .env with .env.example containing placeholder values
- Wire Grafana admin credentials through env vars in docker-compose
- Make metricsStages a pointer to prevent reset-vs-write race on reconnect
- Fix typed-nil interface in debug bundle path (GetClientMetrics)
- Use deterministic field order in InfluxDB Export (sorted keys)
- Replace Authorization header with X-Peer-ID for metrics push
- Fix ingest server timeout to use time.Second instead of float
- Fix gzip double-close, stale comments, trim log levels
- Add tests for influxdb.go and MetricsStages

* Add login duration metric, ingest tag validation, and duration bounds

- Add netbird_login measurement recording login/auth duration to management
  server, with success/failure result tag
- Validate InfluxDB tags against per-measurement allowlists in ingest server
  to prevent arbitrary tag injection
- Cap all duration fields (*_seconds) at 300s instead of only total_seconds
- Add ingest server tests for tag/field validation, bounds, and auth

* Add arch tag to all metrics

* Fix Grafana dashboard: add arch to drop columns, add login panels

* Validate NB_METRICS_SERVER_URL is an absolute HTTP(S) URL

* Address review comments: fix README wording, update stale comments

* Clarify env var precedence does not bypass remote config eligibility

* Remove accidentally committed pprof files

---------

Co-authored-by: Viktor Liu <viktor@netbird.io>
2026-03-22 12:45:41 +01:00
Zoltan Papp fe9b844511 [client] refactor auto update workflow (#5448)
Auto-update logic moved out of the UI into a dedicated updatemanager.Manager service that runs in the connection layer. The
UI no longer polls or checks for updates independently.
The update manager supports three modes driven by the management server's auto-update policy:
No policy set by mgm: checks GitHub for the latest version and notifies the user (previous behavior, now centralized)
mgm enforces update: the "About" menu triggers installation directly instead of just downloading the file — user still initiates the action
mgm forces update: installation proceeds automatically without user interaction
updateManager lifecycle is now owned by daemon, giving the daemon server direct control via a new TriggerUpdate RPC
Introduces EngineServices struct to group external service dependencies passed to NewEngine, reducing its argument count from 11 to 4
2026-03-13 17:01:28 +01:00
Viktor Liu d4f7df271a [cllient] Don't track ebpf traffic in conntrack (#5166) 2026-01-27 11:04:23 +01:00
Diego Romar b3a2992a10 [client/android] - Fix Rosenpass connectivity for Android peers (#5044)
* [client] Add WGConfigurer interface

To allow Rosenpass to work both with kernel
WireGuard via wgctrl (default behavior) and
userspace WireGuard via IPC on Android/iOS
using WGUSPConfigurer

* [client] Remove Rosenpass debug logs

* [client] Return simpler peer configuration in outputKey method

ConfigureDevice, the method previously used in
outputKey via wgClient to update the device's
properties, is now defined in the WGConfigurer
interface and implemented both in kernel_unix and
usp configurers.

PresharedKey datatype was also changed from
boolean to [32]byte to compare it
to the original NetBird PSK, so that Rosenpass
may replace it with its own when necessary.

* [client] Remove unused field

* [client] Replace usage of WGConfigurer

Replaced with preshared key setter interface,
which only defines a method to set / update the preshared key.

Logic has been migrated from rosenpass/netbird_handler to client/iface.

* [client] Use same default peer keepalive value when setting preshared keys

* [client] Store PresharedKeySetter iface in rosenpass manager

To avoid no-op if SetInterface is called before generateConfig

* [client] Add mutex usage in rosenpass netbird handler

* [client] change implementation setting Rosenpass preshared key

Instead of providing a method to configure a device (device/interface.go),
it forwards the new parameters to the configurer (either
kernel_unix.go / usp.go).

This removes dependency on reading FullStats, and makes use of a common
method (buildPresharedKeyConfig in configurer/common.go) to build a
minimal WG config that only sets/updates the PSK.

netbird_handler.go now keeps s list of initializedPeers to choose whether
to set the value of "UpdateOnly" when calling iface.SetPresharedKey.

* [client] Address possible race condition

Between outputKey calls and peer removal; it
checks again if the peer still exists in the
peers map before inserting it in the
initializedPeers map.

* [client] Add psk Rosenpass-initialized check

On client/internal/peer/conn.go, the presharedKey
function would always return the current key
set in wgConfig.presharedKey.

This would eventually overwrite a key set
by Rosenpass if the feature is active.

The purpose here is to set a handler that will
check if a given peer has its psk initialized
by Rosenpass to skip updating the psk
via updatePeer (since it calls presharedKey
method in conn.go).

* Add missing updateOnly flag setup for usp peers

* Change common.go buildPresharedKeyConfig signature

PeerKey datatype changed from string to
wgTypes.Key. Callers are responsible for parsing
a peer key with string datatype.
2026-01-20 13:26:51 -03:00
Zoltan Papp 58daa674ef [Management/Client] Trigger debug bundle runs from API/Dashboard (#4592) (#4832)
This PR adds the ability to trigger debug bundle generation remotely from the Management API/Dashboard.
2026-01-19 11:22:16 +01:00
Misha Bragin e586c20e36 [management, infrastructure, idp] Simplified IdP Management - Embedded IdP (#5008)
Embed Dex as a built-in IdP to simplify self-hosting setup.
Adds an embedded OIDC Identity Provider (Dex) with local user management and optional external IdP connectors (Google/GitHub/OIDC/SAML), plus device-auth flow for CLI login. Introduces instance onboarding/setup endpoints (including owner creation), field-level encryption for sensitive user data, a streamlined self-hosting provisioning script, and expanded APIs + test coverage for IdP management.

more at https://github.com/netbirdio/netbird/pull/5008#issuecomment-3718987393
2026-01-07 14:52:32 +01:00
Zoltan Papp 011cc81678 [client, management] auto-update (#4732) 2025-12-19 19:57:39 +01:00
Pascal Fischer 7193bd2da7 [management] Refactor network map controller (#4789) 2025-12-02 12:34:28 +01:00
Diego Romar 32146e576d [android] allow selection/deselection of network resources on android peers (#4607) 2025-11-21 13:36:33 +01:00
Pascal Fischer 3351b38434 [management] pass config to controller (#4807) 2025-11-19 11:52:18 +01:00
Viktor Liu d71a82769c [client,management] Rewrite the SSH feature (#4015) 2025-11-17 17:10:41 +01:00
Viktor Liu 9cc9462cd5 [client] Use stdnet with a context to avoid DNS deadlocks (#4781) 2025-11-13 20:16:45 +01:00
Pascal Fischer cc97cffff1 [management] move network map logic into new design (#4774) 2025-11-13 12:09:46 +01:00
Zoltan Papp 4d33567888 [client] Remove endpoint address on peer disconnect, retain status for activity recording (#4228)
* When a peer disconnects, remove the endpoint address to avoid sending traffic to a non-existent address, but retain the status for the activity recorder.
2025-10-08 03:12:16 +02:00
Viktor Liu b5daec3b51 [client,signal,management] Add browser client support (#4415) 2025-10-01 20:10:11 +02:00
Zoltan Papp 9e81e782e5 [client] Fix/v4 stun routing (#4430)
Deduplicate STUN package sending.
Originally, because every peer shared the same UDP address, the library could not distinguish which STUN message was associated with which candidate. As a result, the Pion library responded from all candidates for every STUN message.
2025-09-11 10:08:54 +02:00
Bethuel Mmbaga 5113c70943 [management] Extends integration and peers manager (#4450) 2025-09-06 13:13:49 +03:00
Bethuel Mmbaga a8dcff69c2 [management] Add peers manager to integrations (#4405) 2025-09-04 23:07:03 +03:00
Viktor Liu d4c067f0af [client] Don't deactivate upstream resolvers on failure (#4128) 2025-08-29 17:40:05 +02:00
Viktor Liu f063866ce8 [client] Add flag to configure MTU (#4213) 2025-08-26 16:00:14 +02:00
Pascal Fischer b3056d0937 [management] Use DI containers for server bootstrapping (#4343) 2025-08-15 17:14:48 +02:00
Bethuel Mmbaga a4e8647aef [management] Enable flow groups (#4230)
Adds the ability to limit traffic events logging to specific peer groups
2025-08-13 00:00:40 +03:00
Viktor Liu 1d5e871bdf [misc] Move shared components to shared directory (#4286)
Moved the following directories:

```
  - management/client → shared/management/client
  - management/domain → shared/management/domain
  - management/proto → shared/management/proto
  - signal/client → shared/signal/client
  - signal/proto → shared/signal/proto
  - relay/client → shared/relay/client
  - relay/auth → shared/relay/auth
```

and adjusted import paths
2025-08-05 15:22:58 +02:00
hakansa cb8b6ca59b [client] Feat: Support Multiple Profiles (#3980)
[client] Feat: Support Multiple Profiles (#3980)
2025-07-25 16:54:46 +03:00
Krzysztof Nazarewski (kdn) af8687579b client: container: support CLI with entrypoint addition (#4126)
This will allow running netbird commands (including debugging) against the daemon and provide a flow similar to non-container usages.

It will by default both log to file and stderr so it can be handled more uniformly in container-native environments.
2025-07-25 11:44:30 +02:00
Pascal Fischer cb1e437785 [client] handle order of check when checking order of files in isChecksEqual (#4219) 2025-07-24 21:00:51 +02:00
Viktor Liu d6ed9c037e [client] Fix bind exclusion routes (#4154) 2025-07-21 12:13:21 +02:00
Maycon Santos 08fd460867 [management] Add validate flow response (#4172)
This PR adds a validate flow response feature to the management server by integrating an IntegratedValidator component. The main purpose is to enable validation of PKCE authorization flows through an integrated validator interface.

- Adds a new ValidateFlowResponse method to the IntegratedValidator interface
- Integrates the validator into the management server to validate PKCE authorization flows
- Updates dependency version for management-integrations
2025-07-18 12:18:52 +02:00
Pedro Maia Costa e67f44f47c [client] fix test (#4156) 2025-07-16 12:09:38 +02:00
Zoltan Papp 3e6eede152 [client] Fix elapsed time calculation when machine is in sleep mode (#4140) 2025-07-12 11:10:45 +02:00
Zoltan Papp fbb1b55beb [client] refactor lazy detection (#4050)
This PR introduces a new inactivity package responsible for monitoring peer activity and notifying when peers become inactive.
Introduces a new Signal message type to close the peer connection after the idle timeout is reached.
Periodically checks the last activity of registered peers via a Bind interface.
Notifies via a channel when peers exceed a configurable inactivity threshold.
Default settings
DefaultInactivityThreshold is set to 15 minutes, with a minimum allowed threshold of 1 minute.

Limitations
This inactivity check does not support kernel WireGuard integration. In kernel–user space communication, the user space side will always be responsible for closing the connection.
2025-07-04 19:52:27 +02:00
Ali Amer d9402168ad [management] Add option to disable default all-to-all policy (#3970)
This PR introduces a new configuration option `DisableDefaultPolicy` that prevents the creation of the default all-to-all policy when new accounts are created. This is useful for automation scenarios where explicit policies are preferred.
### Key Changes:
- Added DisableDefaultPolicy flag to the management server config
- Modified account creation logic to respect this flag
- Updated all test cases to explicitly pass the flag (defaulting to false to maintain backward compatibility)
- Propagated the flag through the account manager initialization chain

### Testing:

- Verified default behavior remains unchanged when flag is false
- Confirmed no default policy is created when flag is true
- All existing tests pass with the new parameter
2025-07-02 02:41:59 +02:00
Viktor Liu 6127a01196 [client] Remove strings from allowed IPs (#3920) 2025-06-10 14:26:28 +02:00
Viktor Liu 3c535cdd2b [client] Add lazy connections to routed networks (#3908) 2025-06-08 14:10:34 +02:00