* implement certificate posture check
* log signal address
* add keychain and cert store support
* read the console user's keychain through a user session helper
A root daemon cannot reach a login keychain: securityd is per session and a
key ACL needs a session to prompt in, so dropping uid is not enough. The
daemon now answers certificate challenges from the System keychain itself,
where MDM installs device identities, and launches "netbird posture
cert-proof" into the console user's desktop session with launchctl asuser
for the login keychain. Only the signature and the chain cross back, never
the private key.
The console user comes from SCDynamicStoreCopyConsoleUser, bound with purego
like the keychain calls. The login window reports no user, root, or
"loginwindow", and all three are treated as no keychain to read, so a Mac at
the lock screen sends device proofs alone.
Adds info logging across the path: the keychain search list, per class query
status and item counts, the chain built per candidate, and the verification
error for every rejected candidate. A run that sends nothing now says why.
README.md documents the trust model, the console user limitation and how to
read the logs.
* read the signed-in user's certificate store on Windows
A service reads LocalMachine\MY, where AD and Intune enrol device
certificates. CurrentUser\MY lives in the signed-in user's registry hive
with keys protected against their profile, and a service that opens it does
not fail: "current user" resolves to HKU\S-1-5-18, so it silently reads the
service account's own empty store. The service therefore reads the machine
store itself and launches "netbird posture cert-proof" with the session
token for the rest, mirroring the macOS console user helper.
Windows lets a privileged service assume a user identity, so the token goes
straight into the child process and no external tooling is involved.
CREATE_NO_WINDOW keeps a console window from flashing on the desktop every
sync. In-process impersonation would also work but is per OS thread while
goroutines migrate, so the child process avoids that class of bug.
Session selection prefers the physical console and falls back to any active
session, so remote desktop and VDI hosts are covered. WTSQueryUserToken
needs SE_TCB_NAME, so a user-run client skips the helper and reads the
machine store alone.
SystemStore takes a store location, gaining NewUserStore alongside
NewSystemStore and the per candidate logging macOS already had. The request
building and proof merging move to helper_spawn.go, shared by both
platforms, and helperStore picks what the helper reads per platform.
* start TPM support
* split goreleaser to support pkcs11 and exclude on docker
* update goreleaser
* go mod tidy
* add tpm pin to netbird config
* split cert and key location and allow key lookup on tpm
* add unsupported flag for mobile devices
* Isolate the cert proof helper from the service environment and cap its output
* Read the PKCS#11 token PIN from NB_TPM_PIN instead of the profile config
* Bound certificate proof collection so a stuck token or keychain cannot hold the sync loop
* Stop retrying a PKCS#11 PIN the token rejected
* Log certificate posture details at debug level
* Sign only nonces and peer keys of the size management issues
* Skip certificate files whose key belongs to another certificate
* Bound PKCS#11 driver sizes, pin template values, and log out only a login the session owns
* Never pass NULL to CFRelease and skip unreadable keychain identities
* Keep the macOS keychain code out of iOS and the PKCS#11 driver out of Android
* Find a chain to each challenge's CAs through every intermediate the store holds
* Require a token label whenever a PKCS#11 PIN is set
* Read user certificates only from the session of the active profile's owner
* Collect certificate proofs again when the owner's session changes and report lost proofs
* Test the PKCS#11 build against SoftHSM in CI and warn once where the build has no driver
* Document where an inline PKCS#11 PIN is stored and how it is protected
* Refuse PKCS#11 URIs that this client cannot honour instead of widening the match
* Trust certificate and key files only when no other user can write or redirect them
* Explain a Windows certificate whose key only a legacy CryptoAPI provider holds
* Use platform absolute module paths in tests and add a real owner session test for Windows
* Match the Windows profile owner by name instead of resolving it through the domain controller
* Keep the certificate stores and TPM library out of the WebAssembly build
* [client] Read TSS2 key files on go-tpm, checked against the library it replaces
The TSS2 parser was the only reason this repository depended on a crypto suite
whose own build tooling it inherits. The replacement sits on go-tpm, which was
already a direct dependency and is in fact what that suite calls underneath, so
this removes a wrapper rather than porting onto a different library: the load,
the derived storage root key and the signing commands are the same calls.
Swapping a parser on the one path a customer actually runs is not something to
assert, so the two are held side by side for this commit. One test feeds the
replacement bytes the old library wrote and requires the same key type, empty
auth flag, parent handle, blobs and decoded public key; the other feeds both the
fixtures the tests are built on, so those are the shape the format calls for and
not merely the shape the new parser reads. The scaffolding goes away with the
dependency in the commit that follows.
The encoder behind the fixtures is written out separately from the parser under
test, so an encoder bug and a decoder bug cannot cancel each other out.
* [client] Drop go.step.sm/crypto and the repo-wide upgrades it imposed
The TSS2 parser was the only thing in the repository that used this module, and
it brought 302 modules into the graph to do it — 35 of them linters, along with
Google Cloud KMS and IAM, the AWS SDK and a terminal styling library. Those are
the module's own development dependencies, which minimal version selection turns
into floors in ours, and they are the whole reason gRPC, protobuf, the AWS SDK,
OpenTelemetry, logrus and five x/ packages had moved. Management, signal, relay
and proxy inherited every one of them for a feature none of them runs.
Removing the import is not enough, because tidy never downgrades: the raised
floors stay written in go.mod. Each one is pinned back to the version main had,
then tidy is left to raise again whatever something still genuinely needs. It
raised nothing: all 43 are back where they were, and go-tpm was already in the
graph at the same version, so the certificate feature now costs no new module at
all.
The differential tests go with it. They existed to check the swap against the
library while both were present, and there is nothing left to compare against.
* [client] Clear the lint findings only the macOS and Windows runners see
golangci-lint analyses one build at a time, so running it on Linux says nothing
about the two platforms CI also lints. Against those builds the feature's
packages reported eight findings, and the structural one is Config.dir: it is
dead on macOS and Windows because neither reads a directory at all, their
collectors take the configuration and discard it. Moving the method beside its
only callers makes that visible in the layout instead of in a linter, and leaves
the gap itself — no file or token store on those platforms — where it belongs,
as something to decide rather than something to silence.
An absent key file beside a certificate was reported as a nil signer with a nil
error, which the caller then had to recognise by its nilness. It is a sentinel
now, so the meaning is in the error rather than in the absence of one.
The rest follow the standard library: the elliptic coordinates and the private
scalar come from the encoding helpers rather than the deprecated big.Int fields,
and an error string loses its trailing colon.
Lint is clean on linux, darwin and windows; the hardware TPM path was exercised
separately against a real device and passes.
* Accept the TSS2 emptyAuth boolean OpenSSL writes and persistent parents on 32-bit builds
* Count the certificates field in the peer meta store test
* Check the store directory before listing it, refuse group-writable files, and reject a URI with two PIN sources
* Share a PKCS#11 login between sessions and send each PIN at most once at a time
* Collect certificate proofs again when the meta sync carrying them failed
* Use no Windows user store when a domainless owner matches accounts of several domains
* Use no user certificate store when the active profile's owner cannot be read
* Document the PIN sources on CertPKCS11URI and keep the README PIN example off the command line
* Test that the PKCS#11 URI stays out of the debug bundle and run the wrong-PIN test only on a disposable token
* Refuse a TPM PSS signature request for the maximum salt length
* Add the certificate fields to the network map golden data
* Retry posture checks whose meta sync timed out instead of dropping them
* Start no system info gathering while a timed-out one is still running
* Guard the applied posture checks across goroutines and keep refreshing proofs while a pending update times out
* Log what a successful certificate proof helper wrote to stderr
* Send recollected certificate proofs to management only when the proven chains changed
* Explain a macOS keychain key whose access list does not allow netbird
* Kill the whole macOS certificate helper process group when it times out
* End sudo option parsing before the macOS certificate helper binary
* Hold off system info gathering only while a timed-out one is still running
* Collect certificate proofs on the posture watcher instead of under the sync lock
* Read the certificate store directory and PKCS#11 URI from the daemon environment, not the profile config
* Install the RPM sysconfig file readable by root only and show the certificate posture variables
* Move the certificate posture README into the package doc and the docs site
* Name NB_CERT_PKCS11_URI in the PIN-without-token error
* Keep the file check results of the latest-started system info refresh
* Give the full import command for a keychain key netbird may not use, and correct the package doc
* Restrict the service environment file to root on every package install
* Search only the System keychain in the macOS daemon and only the login keychain in the user helper
* Let the certificate proof helper read the PKCS#11 token from the environment on Linux
* Ask a macOS user's keychain again only after an hour when it proved nothing
* Clear the lint findings in certificate posture
* Hold off the keychain helper only after a completed or timed-out run, independent of CA order
* Keep free functions out of the method lists of PKCS11Store, URI and Challenger
* Name the post-install permission helper in snake case and shorten the sysconfig certificate block
* Drop the certificate store directory from certproof.Config, which only NB_CERT_STORE_DIR sets
* [management] Renew certificate challenge nonces on quiet accounts
A certificate challenge nonce is accepted for its own window and the one before
it, and it only reaches a peer attached to a network map. An account where
nothing changes sends no map, so after a day the peer re-sends the nonce it
still holds, verification rejects its whole proof set, and the certificates
stored for it are dropped. It fails the certificate check and loses every policy
gated on it until some unrelated change happens to push a map. The outage
repairs itself in seconds, which is what makes it expensive: it is intermittent,
it only hits stable networks, and it is not reproducible on demand.
Push the account's peers an update often enough that the nonce they hold is
never close to expiring. Only accounts whose posture checks actually ask for a
certificate are tracked, so a deployment without the feature does no extra work.
The refresh runs from one goroutine over a map of accounts rather than a timer
per account: the period is hours, so one pass every few minutes costs nothing
next to it, and there is no timer to re-arm when an account that falls due
sooner appears. Each account's first run is offset by a hash of its ID, because
the challenge window is global and an instance restart would otherwise arm every
account in the same moment.
The push carries no administrative change, so it is counted as a refresh rather
than an update and stays out of the figures that track what was edited.
(cherry picked from commit 7ad4a0df37)
* [management] Make the certificate challenge window one knob to turn
Renewal was timed against the window in two different ways: the period derived
from it, the sweep interval did not. Shortening the window to watch a renewal in
an end-to-end run would have left the refresher still looking for due accounts
every quarter of an hour, so nothing would have been renewed in time and the
test would have reported the feature broken.
Derive the sweep from the period, within bounds that keep a very short window
from spinning and a normal one from checking less often than is useful, and
allow the window itself to be set through the environment so a run can take
seconds instead of half a day. A value that cannot be parsed or falls outside
the bounds keeps the default, because a window nobody intended is a security
property nobody chose, and an override is logged at warning level since it sets
how long a device keeps passing the check after its key is gone.
Every instance has to be given the same value: the window is part of the nonce,
so instances that disagree reject each other's.
(cherry picked from commit 0e38fcf409)
* [management] Pin the property that makes per-peer nonce state unnecessary
A nonce carries the window it was minted in, not the instant, and is accepted
for that window and the one before it. So a peer re-stamped at least once per
window can never be left holding one outside the accepted pair, whenever it was
last served and however much life its own nonce had left. That is the whole
reason management tracks nothing per peer, and it was resting on an argument
rather than a test.
The phases are part of the property, not decoration: accounts are deliberately
given a refresh phase of their own, so the guarantee has to hold off the window
boundary too. The negative case shows why that matters — a cadence of exactly
two windows lands inside the grace window when it is aligned to the boundary and
leaves a gap when it is not.
(cherry picked from commit dee68facfd)
* [management] Renew challenges only for the peers that answer one
The refresh pushed an update to every connected peer of the account, while only
the peers a certificate check applies to carry a nonce. On an account where a
handful of peers sit behind the check and the rest do not, everyone was woken
several times a day to be handed a map that changed nothing for them.
Push to the sources of the enabled policies whose posture checks include a
certificate check, which is exactly the set that is sent a challenge.
Resolving the set the other way round than the gRPC layer does is the risk here:
a peer the refresh forgets stops being renewed and falls out of its policies
silently, which is the failure this whole mechanism exists to prevent. So the
selection is held against processPeerPostureChecks, the per-peer rule that
decides who receives a challenge in the first place, by a test that asks both
the same question and requires the same answer.
(cherry picked from commit dc4d0e0274)
* [management] Derive certificate challenge nonces from the stored encryption key
The nonce secret came from the server's WireGuard key, which is generated afresh
in every process and never persisted. A nonce carries no state, so the only
thing that lets one instance verify what another issued is deriving the same
secret — and that premise, written in the comment above the challenger, was not
met: every instance had its own key.
A peer reconnecting after a restart therefore presented a nonce minted under the
previous secret, verification failed with a mismatch, its whole proof set was
rejected and the certificates stored for it were dropped until it signed again.
Reproduced three times on the lab, each one logging "nonce was not issued to
this peer", which only a changed secret produces. On a single instance it costs
seconds of lost policy access per restart; across instances it is not transient
at all, because every reconnect that lands elsewhere is rejected the same way.
Derive from the data store encryption key instead: it is generated once, written
back to the configuration and read by every instance, so it survives restarts
and is shared. Where none is configured the secret falls back to the WireGuard
key with a warning — degraded but still unpredictable, which is the property
that matters most: a peer able to guess it could mint the nonces of future
windows, sign them while its key is present and keep passing after it is gone.
The challenger is now built once and passed to the two places that need it,
rather than re-derived per message.
(cherry picked from commit 278f2f3807)
* [management] Register an account for renewal where its nonce is issued
Renewal was armed when a peer connected or when a posture check was saved, both
of which ask the store whether the account has a certificate check. That misses
the case it most needs to catch: the check is created through one instance while
the peers are connected to another, so the instance serving them never learns it
has anything to renew and their nonce expires. It also charged a query to every
peer connect in every account, including the ones that will never use the
feature, which a fleet reconnecting after a restart pays all at once.
Register where the nonce is actually stamped instead. A nonce is verified from a
shared secret and so travels between instances, but the renewal that keeps it
fresh cannot: only the instance holding a peer's stream can push to it. Issuing
and renewing now line up by construction — an instance renews exactly the
accounts it has issued nonces for — and an instance that never issues one has
nothing to renew, so there is no case left to miss.
The registration is a map insert with no store access, which is what lets it sit
on a path taken by every login and every initial sync.
Reported by Viktor Liu, who also proposed registering at the point of issue.
(cherry picked from commit 2d16dd7d7cf54762f2e64c5630ea092f32ef63ab)
* [management] Register for renewal on pushed updates, not only on connect
Registering where the nonce is stamped only covered the login and the initial
sync, which both happen when a peer opens a stream. That left out the path the
mechanism exists for.
On the cloud the network map controller is wrapped so that an update publishes
to an event bus instead of pushing locally: an instance handling a REST change
broadcasts, and every instance holding a peer of that account pushes to its own.
Those pushes stamp a nonce through the update handler, and nothing there
registered, so an instance learned about an account only when one of its peers
happened to reconnect. For a quiet fleet that is the original bug: the check is
created, the peers are told about it, and nobody renews what they were told.
Registering on the pushed update closes it, and is the difference between
stamping and marking a peer connected — one happens on every push, the other
only when a stream opens. Reported by Viktor Liu; the broadcast that makes it
work was pointed out by Pascal Fischer.
(cherry picked from commit 59efe8d93e53bacdf57cb546f4ab2c19dc4eddab)
* [management] Let the challenge refresh loop stop with the manager that owns it
The loop was started on a context explicitly detached from the caller's, so
nothing could ever stop it. Production is unaffected either way, since
BuildManager is called with context.Background(), but a test that builds a
manager leaked a sweeping goroutine for the rest of the run, and a shutdown
path added later would have had no way to reach it.
Take the manager's context as the request buffer built on the line above
already does. The test pins the contract the loop offers, so a detached
context cannot come back inside Start either.
* [management] Bound one account's challenge refresh so it cannot starve the rest
Resolving which peers answer a challenge reads the store three times, and the
refresher sweeps accounts one after another on a single goroutine. A read that
never returns held the sweep for the life of the process, so every other
account on the instance stopped being renewed and its peers fell out of the
policies gated on the check: one account's bad luck became an outage for all
of them.
Give each refresh the sweep interval it is allowed to occupy, capped at 30s so
a 12-hour window does not grant minutes to a query that should take
milliseconds. A refresh that runs out of time keeps its account tracked, since
a deadline says nothing about whether that account still has a certificate
check.
* [management] Send challenge refreshes down the path the rest of management uses
The refresh dispatched through UpdateAffectedPeers, the one variant that takes
no reason, so it was missing from the update counters and coalesced with
nothing. An administrator editing a policy while the sweep ran made the
account's network map twice over, and UpdateOperationRefresh, added for
exactly this caller, was never referenced.
Buffer it with a posture_check/refresh reason instead. The periodic push is
now visible in the metrics as what it is, distinct from an edit, and the send
detaches from the sweep deadline on its own, so that deadline bounds the store
reads it was meant for.
* [management] Keep the certificate challenge comments to what the history does not say
Four of these ran to three and four times the comment budget, the longest at
992 characters. Most of the excess argued against designs that were never
written or explained a bug that no longer exists in the code, which is what
the commit that fixed it is for.
What is left is the part a reader cannot recover from the code: that the
nonce secret has to be persisted and unpredictable, that stamping and
renewing are decided together because only the serving instance can push, and
that the target rule is the inverse of processPeerPostureChecks.
* Keep the newest posture checks pending whatever made their meta sync fail
* Report no lost certificate when the engine stops during a proof collection
* Share the proof collection single-flight across engine restarts
* Close a PKCS#11 module that loads but cannot be used
* Fix the pending checks comments
* Renew certificate challenges only for the peers streamed to this instance
* Ignore a challenge stamp from an older sync stream of the same peer
* Kill the Windows certificate proof helper with its whole process tree
* Expect the challenge untrack in the session ownership test
* Drop an invalid certificate proof without discarding the valid ones
* Start a system info gathering beside one that has been stuck for ten timeouts
---------
Co-authored-by: pascal <pascal@netbird.io>
Co-authored-by: mlsmaycon <mlsmaycon@gmail.com>
Co-authored-by: riccardom <riccardomanfrin@gmail.com>
* [client] Add a debug bundle file export to the Android bridge
The Android app can only upload a debug bundle and hand the user a key.
Users who want to inspect what leaves their device before sharing it
have no way to get the zip itself. Add DebugBundleFile, which generates
the bundle into the cache directory and returns its path instead of
uploading; the app copies it wherever the user chose and removes it.
DebugBundle keeps its behavior. Both entry points share the unexported
debugBundle with an upload switch, so the body stays where it was and
merges cleanly with the MDM overlay change on main.
Because the file variant leaves the zip to the caller and the upload
variant only removes it after the upload finishes, a process killed in
between leaves a zip behind in the cache. Remove stale bundles before
generating a new one: RemoveStaleBundles deletes zips matching the
generator's pattern that are older than an hour. Remote debug jobs write
to the same directory, so younger files are treated as still in use.
* Preserve network map for debug bundle on Android
* [client] Keep exported Android debug bundles out of the stale cleanup
DebugBundleFile hands the zip to the caller, but the file kept the
netbird.debug.*.zip name that RemoveStaleBundles matches, so a later
debug run could delete it once it was older than an hour. Rename the
exported bundle to netbird.debug-file.*.zip after generation so the
cleanup only ever touches bundles no caller owns.
* [client] Warn when a stale debug bundle cannot be removed
A failed removal means bundles pile up in the cache directory, so log it
at Warn instead of Debug. A file that is already gone was removed by a
concurrent cleanup and is skipped silently.
* [client] Drop the outdated debugBundle comment
The comment still said the file variant leaves the zip in place, but it
is renamed by debug.ExportBundle since the stale-cleanup change.
* [client] Test that the network map reaches the debug bundle
Cover both halves of the path Android now relies on: the engine keeps
the latest sync response once persistence is enabled, and the bundle
generator writes it to network_map.json (anonymized or not) and omits
the file when there is no sync response.
* [client] Remove abandoned exported debug bundles after a day
An exported bundle is owned by the caller, but if the app is killed
before it copies and deletes the file, nothing ever removes it from the
cache directory. Let RemoveStaleBundles also match exported bundles,
with a 24 hour max age instead of the caller-provided one, so a bundle
that is still being saved survives while an abandoned one goes.
This introduces a disabled-by-default allow-remote-jobs setting that
controls whether the management server may run jobs (such as debug
bundles) on a peer. The flag propagates end to end: through client
configuration, the daemon SetConfig and Login requests, authentication,
and system info, up to management, where it is stored on the peer and
exposed on the peers API as remote_jobs_allowed. The client refuses any
management-requested job unless the peer has opted in. Because enabling
remote jobs crosses the user-to-root boundary, turning it on requires
privilege, mirroring the SSH-server gate. Administrators can enforce the
setting through MDM policy on both macOS and Windows, and MDM can also
override the debug-bundle upload URL. The change ships policy
documentation and generated profile templates, and adds configuration,
conflict, and enforcement tests covering the opt-in, privilege, and MDM
paths.
- **Wails v3 application** (`client/ui`) with a React + TypeScript + Tailwind frontend replacing the Fyne UI: main connection view, exit-node switcher, networks/peers browser with detail panels, profile management, settings (general, network, SSH, security, troubleshooting, appearance), debug-bundle creation, and a first-run welcome flow.
- **Internationalization**: go-i18n bundle with 9 locales (en, de, es, fr, hu, it, pt, ru, zh-CN) shared between the tray and the frontend.
- **New system tray** implementation with per-platform theme-aware icons, including a native XEmbed host for Linux (`xembed_tray_linux.c`) and a Linux theme watcher.
- **Session handling**: auth session watcher (`client/internal/auth/sessionwatch`), pending login flow, session-expiration dialog and tray notifications, and `netbird login` improvements.
- **Daemon API extensions** (`daemon.proto`): status stream subscription, event stream, networks/exit-node selection endpoints, and richer full status — with probe throttling on the daemon side to protect against UI-driven request storms.
- **UI preferences store** persisted per profile, autostart management via the daemon (single source of truth in HKCU on Windows).
- **Build system**: Taskfile-based builds per platform (macOS, Linux, Windows), Docker cross-compilation images, MSIX/NSIS/nfpm/AppImage packaging, and a new `frontend-ui` CI workflow.
Co-authored-by: Zoltan Papp <zoltan.pmail@gmail.com>
Co-authored-by: Eduard Gert <kontakt@eduardgert.de>
Co-authored-by: braginini <bangvalo@gmail.com>
Co-authored-by: Pascal Fischer <32096965+pascal-fischer@users.noreply.github.com>
Co-authored-by: riccardom <riccardomanfrin@gmail.com>
* Add iOS debug bundle support in Go
Thread cacheDir through NewClient -> RunOniOS -> MobileDependency.TempDir
so the iOS client can pass its sandbox-writable cache directory for
debug bundle zip file creation instead of os.TempDir().
Move log collection into platform-dispatched addPlatformLog():
- iOS: adds the file-based Go client log (with rotation, stderr/stdout
companions and anonymization handled by addLogfile) plus the Swift app
log (swift-log.log) written by the iOS app into the same log directory
- Other non-Android platforms: existing file-based log + systemd fallback
Narrow the debug_nonandroid.go build tag to !android && !ios so iOS no
longer attempts the systemd journal fallback.
Add a DebugBundle() entry point to the iOS Go client that generates a
bundle, uploads it and returns the upload key. It works with or without
a running engine: when the engine is up it reuses the live config, sync
response and client metrics; otherwise it loads the config from disk (or
the preloaded tvOS config). Guard the live config/ConnectClient behind a
state mutex since DebugBundle may run on a different thread.
* Include the iOS state file in the debug bundle
addStateFile() resolved the state path via ServiceManager.GetStatePath(),
which on iOS points at a hard-coded default that does not exist in the app
sandbox, so the state file was silently skipped.
Add an optional StatePath to GeneratorDependencies and use it when set,
falling back to the ServiceManager default otherwise. The iOS DebugBundle
passes the client's actual state file path (the App Group profile state),
matching the Android bundle which includes the state file.
* ios: enable sync response persistence for debug bundle
Turn on sync response persistence before starting the engine so
DebugBundle can include the network map. On iOS the store is disk-backed
(see syncstore) to keep the map out of the constrained process memory.
* ios: pass log file path through NewClient constructor (#6393)
Add logFilePath field to Client struct and expose it as a parameter
in NewClient so callers provide the Go log path at construction time.
Wire it into DebugBundle via GeneratorDependencies.LogPath so the
debug bundle includes client.log and swift-log.log regardless of
whether the bundle is triggered by the app or the management server.
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* ios: pass log file path to engine for remote debug bundles
RunOniOS started the engine with an empty LogPath, so EngineConfig.LogPath
was never set. Management-triggered (jobs) debug bundles read the log path
from the engine config, so they collected no client logs (client.log,
rotated logs, swift-log.log). The GUI path was unaffected because it passes
c.logFilePath directly to the bundle generator.
Thread c.logFilePath through RunOniOS into the engine config so remote
bundles include the client logs too.
---------
Co-authored-by: evgeniyChepelev <68751844+evgeniyChepelev@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* Initial scaffolding
* Applies MDM override
* Unit tests
* Helpers business logic
* Return error if trying to modify any config that is gated by MDM
* Add ManagedFields to returned config over GetConfig
* Adds initial 101 MDM policy business logic testing
* gRPC MDM changes
* MDM Name scoping for clarity
* Implements windows loading of MDM policy
* Adds missing WGPort config
* Cleanup setupKey to align to linear
* Align split tunnel code
* Adds some log
* Prefix every log with MDM
* Adds debug config cobra command
This can be useful for troubleshooting and checking config
now that its resolution is not trivial
defaults > config > env cars > CLI/UI > MDM
* Adds MDM 1m diff checker & reloader
* Adds also up/start after cancel
* Publishes event for UI to sync upon MDM changes
* Add events to resync UI to actual config
This also provide fixup for UI no aligning to changed config when coming from cli up with config flags.
* UI behavior conflicts relaxation
UI sends full config snapshot with all values. It doesn't
make sense to block it if the values are aligned with the
values constrained by the MDM policy. It's just simplier
to allow values that are compliant. (this goes for the CLI
as well at this point)
* Lock toggle Settngs
* Advanced Settings locking
* Fixup presharedkey
* Apply MDM locks
* Toggle gray in/out for Advanced Settings
* Adds support for disabling of Profiles and UpdateSettings feature flags
* Adds Gate Login as well when --disable-update-settings=true is given to service
This commit tries to settle things with an old PR-4237 which had relaxed
the case where the SetConfig returned an `Unavailable` code error.
Under this circumnstance the PR allowed the upFunc to just emit a warning and
progress further with the login gRPC. Since the login call is consuming
the --management-url coming from the `up` command, it might be possible
to abuse the "Unavailable" code to inject a management URL that is different
from the configured one even though the --disable-update-settings is set
to true (?)
* Evaluate disable-update-settings errors only when there's an actual override
* [UI] Fixup advanced Settings
* [UI] Fixup for preshared key
* [UI] Fixup for profile enable/disable toggle
We need to align the initial state to evaluate the delta in case.
The initial state has to be "true" since the profile starts visible.
Then we receive MDM and transition the cache bool value to the actual
MDM imposed state
* Enforces disable networks
* [UI] Aligns to "enable/disable once on change only"
* Fixup: MDM wins. always
* Removes --disable-advanced-settings
It was a typo in our meetings. the actual thing is --disable-update-settings
* [PROTO] Removes --disable-advanced-settings
* [UI] Removes --disable-advanced-settings
* Pins feat profile retrieval to notif event
* [UI] Fix for "hide" not working when propagating to parent with children
* Adds dep for reading plist files
* Introduces support for darwing plist loading
* Tests MDM config reload via ticker
* [PROVISIONING] ADMX/ADML/PS/bash scripts/templates
* CI fixes
- Add docstrings to `mdm_integration`
- refactor for cognitive complexity
- mod tidy
* Linting
* Add docstrings to `mdm_integration`
* nil,nil is no policy and no error. Allow it
* nil,nil is no policy and no error. Allow it
* exclude MDM profile adminstrated keys data from debug bundle
* Fixes Rosenpass left disable after MDM unlock
* Partial revert coderabbit added docstrings
* Renaming fix
* Avoid locking on clientRunning bool when the connection is aborted for whatever reason
We want to just signal this through the giveUpChan, we will manage the signal from
the waiter side and in case set it to false there. THis way we avoid locking,
which should allow the MDM down+wait_for_term_chan_signal_+up procedure
clientRunning is used to signal two different conditions here:
1. the initialization procedure is over (we have an engine)
2. the connection being up (or being attempted)
Probably these two functionalities should not alias, and the failure of the second condition
(because of any error) should just drive a reconnection (currently it's not happening,
and we silently go idle).
OR, mor probably, the two things are the SAME and there should not exist a case where
we did the "Up" initialization and connection attempt but we are not still attempting it.
* Moves test helper at te very bottom
* Addresses github comments
* No lock no copy
* Prevents engine not stopping within 10 secs from being paired by another instance
We instead juts SKIP updating the policy, so
1. the MDM ticker will kick in 1 minute time,
2. find the policy misaligned,
3. enter the onMDMPolicyChange,
4. find the s.clientRunning == true
(because it is set to false only in server cleanupConnection,
and not by s.actCancel())
5. call s.actCancel() again if not nil
6. immediately return from <-s.clientGiveUpChan
7. finally call s.restartEngineForMDMLocked()
* Since we ARE running there should be a config
If the config was cancelled midflight, connect will abort later on
* DisableAutoConnect should not stop a running connection.
DisableAutoConnect should just avoid the connection attempts *when the service starts*.
If we are started and we are up and running, DisableAutoConnect should not kick in.
Another PR will follow about this topic
* Removes unused vars
* Moves callback into Run method arg
* align comment to removal of DisableAutoConnect
DisableAutoConnect should just avoid the connection attempts *when the service starts*.
If we are started and we are up and running, DisableAutoConnect should not kick in
* Removes unused managed_fields data.
This was initially used to drive the UI but approach changed
to reload config/features upon notifications which makes this data redundant.
* Reorder stuff
* Unexport unrequired vars/functions
PoliciesEqual → policiesEqual
AllKeys → allKeys
* Adds list of MDM managed fields in the debug bundle
* Adds heuristic to detect an edge case on Linux where a system has configured logrotate as a separate service to rotate log files which would mangle our client log files. If we detect logrotate being configured for netbird, we disable our rotation.
* Adds new env var to disable log rotation: NB_LOG_DISABLE_ROTATION
* Adds compressed and plain logrotate files to debug bundle.
* Replaces lumberjack with timberjack (maintained fork with bug fixes and extra features).
* Clarifies which daemon version is running in the bundle stats.
* Change logging for client service status to console
* Add client metrics
* Add client metrics system with OpenTelemetry and VictoriaMetrics support
Implements a comprehensive client metrics system to track peer connection
stages and performance. The system supports multiple backend implementations
(OpenTelemetry, VictoriaMetrics, and no-op) and tracks detailed connection
stage durations from creation through WireGuard handshake.
Key changes:
- Add metrics package with pluggable backend implementations
- Implement OpenTelemetry metrics backend
- Implement VictoriaMetrics metrics backend
- Add no-op metrics implementation for disabled state
- Track connection stages: creation, semaphore, signaling, connection ready, and WireGuard handshake
- Move WireGuard watcher functionality to conn.go
- Refactor engine to integrate metrics tracking
- Add metrics export endpoint in debug server
* Add signaling metrics tracking for initial and reconnection attempts
* Reset connection stage timestamps during reconnections to exclude unnecessary metrics tracking
* Delete otel lib from client
* Update unit tests
* Invoke callback on handshake success in WireGuard watcher
* Add Netbird version tracking to client metrics
Integrate Netbird version into VictoriaMetrics backend and metrics labels. Update `ClientMetrics` constructor and metric name formatting to include version information.
* Add sync duration tracking to client metrics
Introduce `RecordSyncDuration` for measuring sync message processing time. Update all metrics implementations (VictoriaMetrics, no-op) to support the new method. Refactor `ClientMetrics` to use `AgentInfo` for static agent data.
* Remove no-op metrics implementation and simplify ClientMetrics constructor
Eliminate unused `noopMetrics` and refactor `ClientMetrics` to always use the VictoriaMetrics implementation. Update associated logic to reflect these changes.
* Add total duration tracking for connection attempts
Calculate total duration for both initial connections and reconnections, accounting for different timestamp scenarios. Update `Export` method to include Prometheus HELP comments.
* Add metrics push support to VictoriaMetrics integration
* [client] anchor connection metrics to first signal received
* Remove creation_to_semaphore connection stage metric
The semaphore queuing stage (Created → SemaphoreAcquired) is no longer
tracked. Connection metrics now start from SignalingReceived. Updated
docs and Grafana dashboard accordingly.
* [client] Add remote push config for metrics with version-based eligibility
Introduce remoteconfig.Manager that fetches a remote JSON config to control
metrics push interval and restrict pushing to a specific agent version
range. When NB_METRICS_INTERVAL is set, remote config is bypassed
entirely for local override.
* [client] Add WASM-compatible NewClientMetrics implementation
Replace NewClientMetrics in metrics.go with a WASM-specific stub in metrics_js.go, returning nil for compatibility with JS builds. Simplify method usage for WASM targets.
* Add missing file
* Update default case in DeploymentType.String to return "unknown" instead of "selfhosted"
* [client] Rework metrics to use timestamped samples instead of histograms
Replace cumulative Prometheus histograms with timestamped point-in-time
samples that are pushed once and cleared. This fixes metrics for sparse
events (connections/syncs that happen once at startup) where rate() and
increase() produced incorrect or empty results.
Changes:
- Switch from VictoriaMetrics histogram library to raw Prometheus text
format with explicit millisecond timestamps
- Reset samples after successful push (no resending stale data)
- Rename connection_to_handshake → connection_to_wg_handshake
- Add netbird_peer_connection_count metric for ICE vs Relay tracking
- Simplify dashboard: point-based scatter plots, donut pie chart
- Add maxStalenessInterval=1m to VictoriaMetrics to prevent forward-fill
- Fix deployment_type Unknown returning "selfhosted" instead of "unknown"
- Fix inverted shouldPush condition in push.go
* [client] Add InfluxDB metrics backend alongside VictoriaMetrics
Add influxdb.go with timestamped line protocol export for sparse
one-shot events. Restore victoria.go to use proper Prometheus
histograms. Update Grafana dashboards, add InfluxDB datasource,
and update docs.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* [client] Fix metrics issues and update dev docker setup
- Fix StopPush not clearing push state, preventing restart
- Fix race condition reading currentConnPriority without lock in recordConnectionMetrics
- Fix stale comment referencing old metrics server URL
- Update docker-compose for InfluxDB: add scoped tokens, .env config, init scripts
- Rename docker-compose.victoria.yml to docker-compose.yml
* [client] Add anonymised peer tracking to pushed metrics
Introduce peer_id and connection_pair_id tags to InfluxDB metrics.
Public keys are hashed (truncated SHA-256) for anonymisation. The
connection pair ID is deterministic regardless of which side computes
it, enabling deduplication of reconnections in the ICE vs Relay
dashboard. Also pin Grafana to v11.6.0 for file-based provisioning
and fix datasource UID references.
* Remove unused dependencies from go.mod and go.sum
* Refactor InfluxDB ingest pipeline: extract validation logic
- Move line validation logic to `validateLine` and `validateField` helper functions.
- Improve error handling with structured validation and clearer separation of concerns.
- Add stderr redirection for error messages in `create-tokens.sh`.
* Set non-root user in Dockerfile for Ingest service
* Fix Windows CI: command line too long
* Remove Victoria metrics
* Add hashed peer ID as Authorization header in metrics push
* Revert influxdb in docker compose
* Enable gzip compression and authorization validation for metrics push and ingest
* Reducate code of complexity
* Update debug documentation to include metrics.txt description
* Increase `maxBodySize` limit to 50 MB and update gzip reader wrapping logic
* Refactor deployment type detection to use URL parsing for improved accuracy
* Update readme
* Throttle remote config retries on fetch failure
* Preserve first WG handshake timestamp, ignore rekeys
* Skip adding empty metrics.txt to debug bundle in debug mode
* Update default metrics server URL to https://ingest.netbird.io
* Atomic metrics export-and-reset to prevent sample loss between Export and Reset calls
* Fix doc
* Refactor Push configuration to improve clarity and enforce minimum push interval
* Remove `minPushInterval` and update push interval validation logic
* Revert ExportAndReset, it is acceptable data loss
* Fix metrics review issues: rename env var, remove stale infra, add tests
- Rename NB_METRICS_ENABLED to NB_METRICS_PUSH_ENABLED to clarify that
collection is always active (for debug bundles) and only push is opt-in
- Change default config URL from staging to production (ingest.netbird.io)
- Delete broken Prometheus dashboard (used non-existent metric names)
- Delete unused VictoriaMetrics datasource config
- Replace committed .env with .env.example containing placeholder values
- Wire Grafana admin credentials through env vars in docker-compose
- Make metricsStages a pointer to prevent reset-vs-write race on reconnect
- Fix typed-nil interface in debug bundle path (GetClientMetrics)
- Use deterministic field order in InfluxDB Export (sorted keys)
- Replace Authorization header with X-Peer-ID for metrics push
- Fix ingest server timeout to use time.Second instead of float
- Fix gzip double-close, stale comments, trim log levels
- Add tests for influxdb.go and MetricsStages
* Add login duration metric, ingest tag validation, and duration bounds
- Add netbird_login measurement recording login/auth duration to management
server, with success/failure result tag
- Validate InfluxDB tags against per-measurement allowlists in ingest server
to prevent arbitrary tag injection
- Cap all duration fields (*_seconds) at 300s instead of only total_seconds
- Add ingest server tests for tag/field validation, bounds, and auth
* Add arch tag to all metrics
* Fix Grafana dashboard: add arch to drop columns, add login panels
* Validate NB_METRICS_SERVER_URL is an absolute HTTP(S) URL
* Address review comments: fix README wording, update stale comments
* Clarify env var precedence does not bypass remote config eligibility
* Remove accidentally committed pprof files
---------
Co-authored-by: Viktor Liu <viktor@netbird.io>
Auto-update logic moved out of the UI into a dedicated updatemanager.Manager service that runs in the connection layer. The
UI no longer polls or checks for updates independently.
The update manager supports three modes driven by the management server's auto-update policy:
No policy set by mgm: checks GitHub for the latest version and notifies the user (previous behavior, now centralized)
mgm enforces update: the "About" menu triggers installation directly instead of just downloading the file — user still initiates the action
mgm forces update: installation proceeds automatically without user interaction
updateManager lifecycle is now owned by daemon, giving the daemon server direct control via a new TriggerUpdate RPC
Introduces EngineServices struct to group external service dependencies passed to NewEngine, reducing its argument count from 11 to 4
This will allow running netbird commands (including debugging) against the daemon and provide a flow similar to non-container usages.
It will by default both log to file and stderr so it can be handled more uniformly in container-native environments.
With the lazy connection feature, the peer will connect to target peers on-demand. The trigger can be any IP traffic.
This feature can be enabled with the NB_ENABLE_EXPERIMENTAL_LAZY_CONN environment variable.
When the engine receives a network map, it binds a free UDP port for every remote peer, and the system configures WireGuard endpoints for these ports. When traffic appears on a UDP socket, the system removes this listener and starts the peer connection procedure immediately.
Key changes
Fix slow netbird status -d command
Move from engine.go file to conn_mgr.go the peer connection related code
Refactor the iface interface usage and moved interface file next to the engine code
Add new command line flag and UI option to enable feature
The peer.Conn struct is reusable after it has been closed.
Change connection states
Connection states
Idle: The peer is not attempting to establish a connection. This typically means it's in a lazy state or the remote peer is expired.
Connecting: The peer is actively trying to establish a connection. This occurs when the peer has entered an active state and is continuously attempting to reach the remote peer.
Connected: A successful peer-to-peer connection has been established and communication is active.