Commit Graph
1469 Commits
Author SHA1 Message Date
Edward 51a9d32cfc [client] Warn on missing translations and check placeholders (#8090) 2026-10-06 16:59:26 +02:00
Pascal Fischer 2623feeb5b [management] remove ingress ports (#8062) 2026-10-06 16:35:31 +02:00
Zoltan Papp d56e6fc5f4 [client, android, ios] Coalesce peer list change notifications to the mobile listener (#7546)
* [client] Coalesce peer list change notifications to the mobile listener

Every peer state change spawned a goroutine to call the platform
listener. During a reconnect storm this pinned hundreds of OS threads
in JNI and let the UI call back into the engine from each of them.
Deliver peer list changes from a single goroutine per listener and
collapse pending changes into the latest count.

* [client] Cover a pending wake-up when the peer list deliverer is replaced

The replacement test waited for the old callback to finish before
swapping listeners, so it never exercised the stop check that runs
after a wake-up. Block the old callback, queue a peer list change and
swap while it is blocked, then assert the old listener never sees the
new count.

* [client] Signal peer list deliverer exit and wait for it in the test

The replacement test sampled the old listener after a fixed sleep, so a
late stale delivery could slip past it. Close a done channel when the
deliverer goroutine returns and let the test wait on it instead.

* [client] Drop the test-only peer list deliverer exit channel

The done channel was only read by the replacement test. Production code
cannot wait on it, since joining the deliverer would block on a mobile
callback. The tests now drive the deliverer loop directly and check that
setListener and removeListener close its stop channel.
2026-10-06 15:41:01 +02:00
Alberto Xosé Méndez Taboada c5aa55d2b9 [client] feat(i18n): Add Galician (gl) localization for desktop client (#7041)
* Add Galician (gl) localization for desktop client

* [client] Address review feedback and sync latest i18n keys for Galician

Fix terms per review feedback (usarase, creen, auditar, JWT cache, poscuántico, etc.) and translate newly added upstream i18n keys.

* [client] Fix translations, polish terminology, and sync latest keys for Galician

* [client] Fix i18n parity against branch source of truth and drop undeclared uk

* [ci] Skip Galician locales in codespell check

* [client] Add latest error keys for Galician and revert _index.json formatting

Add error.settings_locked and error.settings_managed_by_mdm to Galician locale for 100% key parity with en/common.json. Revert multi-line formatting in _index.json to keep a single-line entry for Galician.
2026-10-06 15:05:01 +02:00
Eduard GertandEdward f5707c3485 [client] Add RTL layout support to the desktop UI (#8076)
* [client] Add RTL layout support to the desktop UI

<html dir> follows the active language, Radix primitives get the
matching dir, physical spacing/positioning uses logical utilities, and
directional icons, animations, arrow-key navigation and tooltip sides
flip in RTL.

* [client] Address RTL review feedback

Force LTR with isolation for monospace values, let truncated names
take their direction from their content, and keep the profile name
field in the UI direction.

* [client] Keep translated monospace labels in their own direction

The forced-LTR rule for monospace values now skips elements with an
explicit dir, translated development labels use dir=auto, and
monospace values rendered through TruncatedText opt into LTR.

---------

Co-authored-by: Edward <43848523+thomashacker@users.noreply.github.com>
2026-10-06 12:08:11 +02:00
Riccardo Manfrin a8dff998ef [client] Gate settings updates on value, not on field presence (#7398)
* [client] Gate settings updates on value, not on field presence

The update-settings kill switch (--disable-update-settings /
NB_DISABLE_UPDATE_SETTINGS / the MDM DisableUpdateSettings key) forbids
changing settings, but it decided what a "change" was by looking at
whether a field was present in the request. The CLI fills the whole
config surface of SetConfigRequest and LoginRequest from its flags and
environment on every `netbird up` (setupSetConfigReq in cmd/up.go), so a
client configured by environment restates its own configuration on every
start and tripped the gate every time.

SetConfig only warned about that, but Login carries the same fields and
was gated the same way, and Login runs inside the CLI's backoff loop: the
daemon answered every attempt with codes.Unavailable, `netbird up` never
completed, and a container with NB_DISABLE_UPDATE_SETTINGS plus any
config env var (NB_MANAGEMENT_URL, for one) could not come up at all.

Both gates now compare values. Config.WouldChange is the dry-run half of
UpdateConfig: it runs the very same diff logic (Config.apply) against a
copy of the stored config, so the gate cannot drift from what an actual
update would do, nor go stale when a field is added. A request that
restates what the profile already holds changes nothing and is allowed; a
request that diverges is refused exactly as before, and a dry run that
cannot be evaluated fails closed. A profile with no config on disk yet is
judged against the config the daemon would create for it.

For Login, the compared input comes from loginOverridesInput, which
persistLoginOverrides also uses to perform the write, so the gate judges
precisely the two fields a login can persist (management URL, pre-shared
key) and no field it ignores.

Two adjacent defects surfaced while making the comparison exact:

- Config.apply compared URLs as raw strings, so the same endpoint spelled
  without its default port ("https://api.netbird.io" vs
  "https://api.netbird.io:443") counted as a new value and rewrote the
  config. It now compares the parsed forms.
- UpdateConfig did not collapse the redacted pre-shared key, unlike
  UpdateOrCreateConfig and DirectUpdateConfig, so a UI round-trip of the
  mask replaced the stored key with asterisks.

The CLI warning for a refused SetConfig said the method was not available
in the daemon, which sent people looking for a version mismatch that was
not there; it now reports the refusal.

* [client] Do not write the profile config while only reading it to decide

The update-settings gate needs the stored config to decide whether a request
changes anything, so the previous commit moved that read ahead of the refusal.
The read is not side-effect free: profilemanager.GetConfig writes the config
back whenever apply() has to fill in a default the file was missing. A request
that the gate then refuses had therefore already rewritten the profile file.

PeekConfig is GetConfig without that write-back. The returned config is still
normalized in memory, which is what the decision needs; the file is left
exactly as it was found. Every caller of storedConfigAtPath feeds a gate that
can refuse, so they all peek.

Note for reviewers: the daemon still normalizes the file on startup and on
every real update, so nothing depends on a read performing that write.

* [client] Compare service URLs as endpoints, not as strings

Three places in one request path each had their own notion of "same
management URL": the config layer compared the parsed URLs as strings, the
privileged-change gate compared scheme + host + effective port, and the MDM
conflict check compared strings after filling in the default port. Only the
middle one was right.

A string comparison answers the wrong question. "https://api.netbird.io",
"https://api.netbird.io/" and "https://API.netbird.io:443" are one endpoint
written three ways, so a client restating its own management URL with a
trailing slash — a normal way to write it — was still read as a client asking
to be repointed, and the update-settings gate refused it. The MDM check had
the same flaw against the enforced value.

profilemanager.SameServiceURL is now the single comparison: same scheme, same
host case-insensitively as DNS names are, same effective port. The config
layer, the privileged-change gate and the MDM conflict check all defer to it,
so there is one answer to "did this URL change?" instead of three.

* [client] Stop the config dry run from generating throwaway keys

The dry run's baseline for a profile with no config file yet went through
createNewConfig, and apply() generates a WireGuard and an SSH key whenever it
finds those fields empty. The baseline is compared against and discarded, so
every evaluation minted a keypair it threw away — and logged "generated new
Wireguard key". The CLI retries Login in a backoff loop, so a first `netbird up`
on a fresh profile filled the daemon log with what reads like peer-key rotation.

The baseline now starts from the shared skeleton with placeholder keys, so
apply() has nothing to generate. No ConfigInput field maps to either key, so
the comparison is unaffected.

* [client] Cover the login the update-settings gate used to refuse

The gate's decision procedure was tested directly, but no test drove the Login
RPC that the refusal actually broke: the CLI retries Login in a backoff loop,
so a refused no-op login is what kept a client configured by environment from
ever coming up. The handler-level coverage stopped at the refusal case, which
passes on the pre-fix code too.

This test fails on the pre-fix daemon with "update settings are disabled" and
passes now. Past the gate the handler does real work the test does not stand
up, so it asserts only that the refusal did not happen.

* [client] Re-take the update-settings decision under the config lock

Login checks twice on purpose: the first check refuses the ordinary case
early, and authorizeAndPrepareLogin re-takes the authoritative one under
guardedConfigMu because the first is unsynchronized against a concurrent
privileged request. The update-settings decision is now equally
value-dependent — it compares the request against the stored config — but it
was taken only in the first, unlocked check.

So a login that was a no-op when it was checked could be written after a
concurrent writer had repointed the profile, which is exactly the window the
lock exists to close. The decision is now re-taken alongside the privilege one,
which also makes it the last read before persistLoginOverrides writes.

The test drives that interleaving through the existing afterLoginPreCheck seam
and fails without the re-check.

* [client] Drop an unreachable guard and fix two stale comments

- loginOverridesInput's nil-message guard cannot be reached: Login
  dereferences the message well before it, in storedLoginConfig.
- The docstring above afterLoginPreCheck described persistLoginOverrides,
  which lives further down the file and now carries its own.
- UpdateConfig's comment named DirectUpdateConfig; the function is
  DirectUpdateOrCreateConfig.

* [client] Make config reads pure and provision the identity explicitly

Reading a config wrote it back. profilemanager.readConfig persisted whatever
apply() had filled in, and ReadConfig created and wrote the file outright when
it was absent, so every reader was quietly a writer: a gate deciding whether to
refuse a request, a UI listing profiles, a mobile getter reading one preference.
The previous commit worked around that with a PeekConfig variant, which left
two read functions with opposite side effects and the antipattern still there
for everyone else.

Only one thing in a read genuinely had to be persisted: apply() generated the
WireGuard and SSH keys when it found them empty, and a generated key cannot be
recomputed — losing it means the peer comes back with a different identity and
registers again. Everything else apply() fills in is a deterministic default
that the next read recomputes anyway.

So identity provisioning is now its own step, Config.EnsureIdentity, and the
callers that provision write the result out themselves, in the open:

- Server.getConfig, the daemon's provisioning point;
- the CLI's foreground login, which is about to dial management;
- update() / directUpdate(), the config write paths — a stored profile can
  legitimately carry no identity, since a mobile logout clears the keys in
  place, and the next write is what has to mint a new one.

ReadConfig and GetConfig no longer write anything, PeekConfig is gone, and the
dry-run baseline no longer needs placeholder keys to keep apply() from minting
real ones.

One deliberate leftover: readConfig still calls util.EnforcePermission, which
chmods a config file whose permissions are too broad. It changes no content and
is idempotent, and dropping it would leave a legacy file world-readable until
its first write.

* [client] Name the two config readers for what they do

ReadConfig and GetConfig differed in one thing — what happens when the file is
absent — and neither name said which was which:

- ReadConfig      -> ReadOrGenerateConfig  (reads it, or generates one in memory)
- GetConfig       -> GetExistingConfig     (reads it, or fails)

Three comments went with them:

- GetConfig's said "return with Config and if it was created. Errors out if it
  does not exist", which described a bool it does not return and a creation it
  never performs.
- ReadConfig's explained that it does not write, which is what a reader is
  supposed to do anyway.
- Server.getConfig's said it "errors out if it does not exist", which it does
  not — it resolves a default config, and now provisions the identity too.

* [client] Do not panic on a config with no sync message version

apply() wrote the incoming sync message version through the stored pointer,
without checking it was there: a config that carries no version yet made it
dereference nil. Reachable from the update-settings dry run, which runs inside
a request handler — where failing closed is the worst acceptable outcome, and a
panic is not one.

The field is now reassigned like every other optional one, which also means
apply() no longer mutates anything the caller still holds through a pointer, so
the dry run's copy has one less field to detach.

Reported by cubic-dev-ai on PR #7398.

* [client] Compare the client certificate paths before reporting a change

apply() assigned the incoming mTLS certificate and key paths and set updated
unconditionally, without comparing them to what the config already held. It is
the same presence-instead-of-value mistake this branch set out to fix, one
layer down: a caller restating its own certificate paths was reported as
changing them, which trips the value-aware update-settings gate.

Reported by cubic-dev-ai on PR #7398.

* [client] Address the remaining bot findings on PR #7398

- Login logged the active-profile-state error and returned the same cause; the
  repo's guidelines call for one or the other, and the wrapped error is the one
  that carries context. (CodeRabbit)
- `netbird up` reported a codes.Unavailable SetConfig failure as "the daemon
  refused the settings update", but that code also covers a daemon that became
  unreachable. It now reports what the daemon said without asserting why.
  (cubic-dev-ai)
- TestLogin_ChangingTheManagementURLIsRefused asserted the error and nothing
  else, while "refused before it can touch daemon state" is the contract. It now
  checks the stored management URL, the in-progress login and the active profile,
  matching its SetConfig counterpart. (cubic-dev-ai)

* [client] Keep the peer identity out of a read that finds no file

ReadOrGenerateConfig resolves a default config when the profile has no file
yet, and createNewConfig was minting the WireGuard and SSH keys while doing so.
That defeated the provisioning pair it was meant to serve: the CLI's foreground
login calls EnsureIdentity to find out whether it has to persist the keys, got
generated == false because the read had already generated them, and so never
wrote them out. The login then dialed management with an identity that only
existed in memory, and the next login registered a second peer.

createNewConfig no longer provisions. createProvisionedConfig is the variant
that does, and the callers whose contract is "usable as it comes back" use it:
CreateInMemoryConfig, whose callers connect with the result, and the two
create-and-write branches. A read gets a config with no identity, so the
caller's own EnsureIdentity reports the work and triggers the write.

Reported by CodeRabbit and cubic-dev-ai on PR #7398, both on the same defect.

* [client] Stop the gate test from dialing the real management server

TestLogin_RestatingTheStoredConfigPassesTheGate asserts that the gate lets a
no-op login through, and the handler then went on to do the login for real:
isLoginRequired builds an auth client when isLoginRequiredFn is unset, so the
test dialed the profile's management URL — api.netbird.io:443. It took 1.05s
locally and would hang on a runner with no egress, for a fact about the gate
that needs no network at all.

Stubbed like the login_outcome tests do. The test now runs in 0.00s.

Reported by cubic-dev-ai on PR #7398.

* [client] Keep the admin panel path part of its identity

The endpoint comparison introduced for the management URL was applied to the
admin URL too, and that one is opened in a browser rather than dialed over
gRPC: a panel served under /netbird is not the panel served at the root. So a
config whose admin URL differed only by path reported no change, and the new
path was never persisted — a custom panel URL could not be updated at all.

SameServiceURLIncludingPath adds what a URL carries past its endpoint (path,
query, fragment, userinfo) while still treating equivalent spellings as equal:
a missing path and "/" are the same root, and so is a trailing slash. The
management URL keeps the endpoint-only comparison, since only the endpoint is
ever dialed.

Ports are also normalized numerically now, so ":0443" and ":443" are one port.

Reported by cubic-dev-ai on PR #7398 (two findings).

* [client] Treat a profile with no identity as already deregistered

Two findings on the same consequence of pure reads: a profile can legitimately
carry no keys, because logging out clears them in place.

- sendLogoutRequestWithConfig went straight to wgtypes.ParseKey and failed with
  "incorrect key size: 0" on the second logout of the same profile. There is
  nothing to deregister for a peer that was never registered, so it returns
  cleanly. Before pure reads this case was hidden: the read minted a key and
  the daemon dialed management with one it had never seen.
- The mobile logout read the config with the generating reader right after
  checking the file exists. The two are not atomic, so a profile removed in
  between was resolved from the defaults and recreated by the write that
  follows. It uses the existing-file reader now.

Reported by cubic-dev-ai and CodeRabbit on PR #7398.

* [client] Fail `netbird up` when the daemon refuses the settings update

With the update-settings kill switch on, `netbird up --enable-rosenpass`
connected and said almost nothing: SetConfig refused the change, the CLI
downgraded that to a warning, and Login carries no rosenpass field to apply, so
the flag was silently dropped. The setting stayed disabled, which is the point
of the switch, but the caller was never told their request had been ignored.

The refusal now travels as codes.FailedPrecondition instead of
codes.Unavailable, and the CLI fails on it. Unavailable means "the daemon
cannot serve this call", which is why the CLI downgraded it and why
client/ui/services reads it as an unreachable daemon — both wrong for a daemon
that answered and refused. FailedPrecondition also matches what the MDM gate
already returns for a managed field, so both refusals are now one class of
error, and it is added to the login backoff's early-exit codes so a refused
login stops instead of retrying for 30s.

This does not put the container back in the deadlock: with the value-aware
gate, a client restating its own configuration is not refused at all, so
nothing reaches this path unless a real change was asked for.

* [client] Name the reader storedConfigAtPath actually calls

The purity note still said profilemanager.GetConfig, which the rename two
commits later turned into GetExistingConfig.

Reported by cubic-dev-ai on PR #7398.

* [client] Restore the gofmt alignment of the error constants

The comment added above errUpdateSettingsDisabled in the previous commit split
the const block's alignment group, so gofmt wants the two constants above it
re-aligned. CI runs gofmt, so this would have failed the lint job.

* [client] Let an unprivileged caller log out a profile with no identity

The empty-key check sat behind requirePrivilegeForDeregistration, so an
unprivileged logout of an identity-less profile was refused with
PermissionDenied instead of completing as the no-op it is. And it was refused
for most profiles, not a corner case: the gate arms whenever the SSH server is
enabled, and sshServerEnabled reads an absent ServerSSHAllowed as enabled, so
every legacy profile qualifies.

The check now runs first. What the gate protects against is handing this
machine's registered key to another management server; with no key there is
nothing to hand over and nothing to protect.

Reported by CodeRabbit and cubic-dev-ai on PR #7398, both on the same defect.

* [client] Stop `netbird login` from retrying a refusal for 30 seconds

`netbird up` and `netbird login` both run Login through the backoff cycle, and
each carried its own copy of the list of codes that end it. Only up.go learned
about codes.FailedPrecondition, so a refused `netbird login` kept retrying and
then reported "login backoff cycle failed" instead of what the daemon said.

terminalLoginError is now that list, once, next to WithBackOff — the duplicated
copies are what let the two commands disagree in the first place.

Reported by cubic-dev-ai on PR #7398.

* [client] Answer terminalLoginError's nil case on its own terms

A successful Login reaches terminalLoginError with a nil error, and nothing
covered that. It happens to work on grpc v1.80.0 — gstatus.FromError(nil)
answers (nil, true), and Status.Code tolerates a nil receiver by returning
codes.OK, which is not in the terminal set — but that is a chain of internal
details to be relying on for the common path, and none of it was asserted.

Now the nil error is handled where it is obvious, and the table covers it.

Reported by CodeRabbit on PR #7398, which called it a panic; measured on
v1.80.0 it is not one. The gap was the untested reliance, not a crash.

* [client] Treat an unset optional field as its default when diffing a config

Seven Config fields mean "the effective default" when they hold no value:
the five SSH toggles, the SSH JWT cache TTL, and the network monitor. Every
consumer already reads a nil as that default, but apply() diffed them by
presence — `config.X == nil || *input.X != *config.X` — so an input restating
the default counted as a change.

That made the update-settings gate refuse `netbird up` outright. The CLI
sends every flag whose value came from an environment variable
(SetFlagsFromEnvVars goes through pflag's FlagSet.Set, which marks the flag
Changed), and the config a plain login writes leaves all seven unset, so a
container configured with, say, NB_ENABLE_SSH_ROOT=false restated a default
the file held as null on every start and was answered with
FailedPrecondition.

apply() now resolves the seven up front, the way it already did for
ServerSSHAllowed and RemoteJobsAllowed, which also repairs such a profile on
its next write. With the values named, the comparisons below diff values
instead of presence, so their nil branches are gone.

The network monitor keeps its platform default — on for windows and darwin —
and naming it as false elsewhere is what createEngineConfig already read a
nil to be. getJWTCacheTTL reaches the same 0 through its own default, and
Android's GetEnableSSH* getters already answered nil with false.

* [client] Normalize the config before diffing it in WouldChange

apply() reports two different things through one bool: an input that changed
a value, and a field it had to fill in because the config carried none. The
update-settings gate reads that bool as "the caller asked for a change", so
any config still missing a default answered a request that asks for nothing
with a refusal.

Readers already hand out normalized configs — readConfig applies an empty
input for exactly this reason — which is why the gate got away with it. But a
handler that refuses a request must not depend on where its caller obtained
the config, and it must not start reading "this profile predates a field" as
"the caller asked for a change" the day someone adds one with a default.

WouldChange now runs the filling-in as a pass of its own and discards its
verdict, so the pass that answers the caller measures only what the input
did.

* [client] Stop the last config write that skipped normalization

Every path that creates or updates a profile config goes through apply(),
which resolves an optional field to its default — except RenameProfile,
which read the file with a bare json.Unmarshal, set the name, and wrote it
straight back. That copied whatever the file held, so a config written by a
client that stored these fields as null kept them null. It could not
introduce a null, only carry one forward, but renaming a profile is a poor
place to leave a half-resolved config behind. It now reads through
GetExistingConfig, which normalizes what it hands out.

The tests state the invariant the fix completes, over the *bool fields of
Config listed by reflection so a field added later is covered without
touching them: none may come out of apply() unset, and no write may store
one as null. An optional bool that can be nil, true or false forces every
reader to invent the meaning of nil, and makes a diff of the config compare
presence rather than value — which is exactly what refused `netbird up` for
a client restating its own defaults.

SyncMessageVersion stays a genuine three-state field and is not covered: it
is an *int whose absence means the client pins no version, and it travels to
management that way.

* [client] Refuse a serialized config that carries no peer identity

ConfigFromJSON still promised a "fully initialized" config after this PR
moved key generation out of apply() into EnsureIdentity, but identity stopped
being one of the defaults it applies. Its two callers both connect with what
they get back: the iOS SDK's Client.SetConfigFromJSON keeps it as the
preloaded config Run() uses on tvOS, and Auth.SetConfigFromJSON as the config
it authenticates with.

No caller feeds it a document without keys today — every stored document
comes from Auth.GetConfigJSON, whose config is provisioned by
DirectUpdateOrCreateConfig or CreateInMemoryConfig, and the tvOS app only
ever edits fields of a document it already has. This is a safety net for the
next caller, not a live bug.

Provisioning the identity here would be the wrong net. Neither caller can
hand a generated key back to the store the document came from — Client
exports no config at all — so the peer would connect under an identity
nothing persists and register anew on every launch, which is the failure the
EnsureIdentity split exists to prevent. A document with no identity means
nobody has logged in yet, and saying so is the only useful answer.

Both keys are required because both are dead ends when missing: an empty
WireGuard key fails the management login on its size, and an empty SSH key
fails ssh.GeneratePublicKey in ConnectClient before the engine starts.

* [client] Say that the null-on-disk fixture is synthesized, not written

The test comment described the null state in the present tense — "the config
a plain login writes leaves every one of them unset" — which was true before
this branch and is not any more: apply() now resolves those fields, so a
login writes them set. unsetOnDisk puts the null state back deliberately, to
stand in for a profile an older client wrote. Comments only.

* [client] Gather the optional-field defaults into one function

Resolving an unset optional field was spread over five places: the two
values newConfigSkeleton pre-sets, the block this branch added for the SSH
toggles, the network monitor's own if, the `else if` tails of
ServerSSHAllowed and RemoteJobsAllowed, and a trailing if for
DisableNotifications several hundred lines further down. Reading apply() left
no single answer to "what does this field default to, and who decides".

They now live in Config.resolveUnsetDefaults, which apply() calls before it
compares anything — the ordering being the point, since it is what lets
every comparison below diff values instead of presence. The comparisons for
ServerSSHAllowed, RemoteJobsAllowed and DisableNotifications lose their
`config.X == nil ||` clauses accordingly, as the other six already had.

newConfigSkeleton keeps its two, and that is the one asymmetry worth naming:
ServerSSHAllowed defaults to false for a new profile and to true for a
legacy one, and it only works because the skeleton runs first. The doc
comment says so, where before it was implied by the order of two distant
blocks.

Pure refactor. Verified as one: for the four fields whose branches moved,
plus two that did not and the JWT TTL, all 63 combinations of stored value
(nil/false/true) against input value (absent/false/true) produce byte-
identical resolved values and `updated` verdicts before and after.

* [client] Resolve the merge conflicts left in the tree

262ce8c3b landed with the conflict markers still in it, so client/server and
the iOS SDK did not compile. Four regions, resolved as follows.

client/server/mdm.go — main moved the MDM conflict-check machinery into the
mdm package (mdm.ResolveConflicts, mdm.ConflictBool, mdm.ConflictURL, ...).
This branch had edited the local copies, which are now dead: dropped, along
with the profilemanager import that only the local conflictURL needed.

client/server/server.go, Login gate — this branch's value-aware gate stays
(the point of the PR: refuse a real divergence, let a restatement through),
so main's presence-based `loginRequestHasConfigOverrides` block goes; that
helper no longer exists here anyway. Main's other change in the same lines
is real and kept: the MDM policy now comes from the daemon-owned
s.mdmLoader.Load() instead of the package-level loadMDMPolicy, which main
removed. The stale call right below the conflict was the reason the file
would not have compiled even with the markers gone.

client/server/server.go, getConfig — both sides add something and both are
needed. The identity is provisioned and persisted first, then the MDM
overlay is applied, so what reaches disk stays the profile's own config: the
overlay is runtime-only and re-derived on every load.

client/ios/NetBirdSDK/client.go — main reworked SetConfigFromJSON to store
the JSON and re-parse it on each load, which is the shape kept; the parse is
now only a validity check, and this branch's reason for it (a document with
no peer identity is refused, not just an unparseable one) moves into that
comment.

client/server/update_settings_gate_test.go — follows the sentinel constant
to its new home, mdm.PreSharedKeyRedactedSentinel.

* [client] Reuse util's service-URL comparison instead of a second copy

The endpoint-comparison rules this branch introduced now live in util (PR
#7472 moved them there so the MDM conflict check could stop comparing URLs
as strings). Keeping a copy here is what produced that bug in the first
place: two implementations of "is this the same endpoint?" drift, and the
one that drifts starts refusing a URL that addresses the very server it
already points at.

So SameServiceURL delegates the port normalization to util.ServiceURLPort
and drops the local one, and SameServiceURLIncludingPath — endpoint plus
path, for the admin panel URL, which is opened rather than dialed — is
util.SameServiceURL plus the query, fragment and userinfo it adds on top,
so the local path normalization goes too.

What stays here is the distinction util does not make: SameServiceURL is
endpoint-only, because a management URL is dialed and only its host and port
are, while util.SameServiceURL includes the path.

Pure refactor. Verified as one: all 198 pairs of a 14-spelling matrix
(default and zero-padded ports, host case, trailing slash, path, query,
fragment, userinfo, both schemes, nil operands) answer identically for both
functions before and after.

* [client] Give a newly added profile its identity (review item 1)

AddProfile writes the config it builds straight to disk, but built it with
createNewConfig, which stopped generating the peer's keys when identity
generation moved out of apply() into EnsureIdentity. The profile file landed
with an empty PrivateKey and SSHKey.

Nothing lost the keys permanently — the daemon's own getConfig provisions and
persists them on first use — but every reader that does not write got a
config that cannot connect in the meantime, which is exactly the set this
branch grew: the update-settings gate deciding whether to refuse a request,
and the mobile SDKs loading a stored profile.

createProvisionedConfig exists for callers that persist or connect, and this
is one; before the split, createNewConfig produced the keys here too.

* [client] Let a logged-out profile deserialize again (review item 2)

ConfigFromJSON refused a document with no WireGuard or SSH key. A config
legitimately has none between a logout and the next login: mobile
LogoutProfile clears both in place and writes the profile back, so the peer
re-registers on the next login instead of returning as itself.

So the refusal broke the mobile flows it was meant to protect. On iOS and
tvOS the stored JSON of a logged-out profile stopped loading through
Client.SetConfigFromJSON and Auth.SetConfigFromJSON, and copyConfig — which
round-trips a Config through JSON to take an in-memory copy before applying
the MDM overlay — failed on the same document. Where the old code silently
minted a key, this returned an error, which is worse for logout and profile
switching alike: neither is asking to connect.

The deserializer now stays out of the identity question in both directions:
it does not generate one (a read cannot hand back keys nothing will write
down) and does not refuse one that is absent. Whoever goes on to connect is
where an absent identity has to be answered — and it already is, by the
login path that provisions and persists.

ErrConfigWithoutIdentity goes with it; nothing else used it.

* [client] Fold the scheme case here too, like util does (review item 6)

profilemanager.SameServiceURL compared the scheme with ==, util.SameServiceURL
with EqualFold. No observable difference — net/url lowercases the scheme when
it parses, and both functions take parsed URLs — but two functions of the same
name with two different rules is a trap for whoever reads one and assumes the
other.

* [client] Classify the daemon's refusals in the GUI (review item 3)

FailedPrecondition reached the classifier unmatched, so a refusal showed as
"Operation failed". It is the code both of the daemon's deliberate refusals
carry: the update-settings kill switch, and a field an MDM policy manages.

Both are now named — settings_locked and settings_managed_by_mdm, matched on
the message the daemon composes — and FailedPrecondition itself falls back to
change_refused, so a refusal the daemon grows later still reads as a refusal
rather than a failure.

Only the English strings are added. Bundle.Translate falls back to the
default language for a missing key, so other locales show English until the
usual translation pass, rather than the bare "error.<code>" the classifier
would otherwise surface.

Note: the package needs GTK4/WebKit to build, which this machine has not, so
the test is type-checked (go vet, GOOS=windows) but was not executed locally;
CI's Linux job runs it.

* [client] Cover the mobile profile round trip: create, logout, reload

Both mobile regressions this branch's review turned up lived on the same
path, and neither was visible from the desktop client: a profile created
without an identity, and a logged-out profile that would no longer
deserialize. The desktop never meets the second one — it is mobile logout
that clears the peer's keys in place, so the next login registers a new peer
instead of bringing the old one back.

The test walks a profile through the round its user puts it through —
created, logged out, loaded again, switched away from and back — and loads it
at each step the way the SDKs do: read the stored config, serialize it, load
it back. That is Client.SetConfigFromJSON storing the document for tvOS,
Auth.SetConfigFromJSON authenticating with it, and copyConfig taking an
in-memory copy before the MDM overlay.

Verified to fail on each regression separately: restoring the bare
constructor in AddProfile fails it with "a new profile was written with no
identity", and restoring the identity check in ConfigFromJSON fails it at
"load the profile back".

client/mobile already had the coverage for the first one in
TestLogoutProfile_DisableProfiles — which arrived from main with the MDM
work, and which I had not been running.

* [client] Name only the refusals, not every FailedPrecondition

The classifier gained a blanket FailedPrecondition -> change_refused fallback
so a refusal would stop reading as "Operation failed". It reaches too far:
the daemon returns that code for two dozen states that are not settings
refusals — "not logged in", "client is not running", "another capture is
already running", "session can no longer be extended, log in again to
reconnect" — and errorClassifier is shared with the session and connection
services, not just the settings save.

So the user was told the service had refused their change while what they
actually had to do was log in again. The two refusals the daemon composes
stay named by their message; everything else goes back to the generic
message, which says nothing rather than something wrong.

Reported by cubic on the PR.

* [client] Say what each assertion was checking in the mobile test

AGENTS.md asks for a context message on comparison and boolean assertions,
and four of the ones added with this test had none, so a failure would have
read as a bare Empty/Equal with no hint of which step of the round trip broke.

Reported by cubic on the PR.

* [client] Translate the two new error strings into every locale

The GUI classifier gained error.settings_locked and
error.settings_managed_by_mdm, and only the English strings were added: the
bundle falls back to the default language for a missing key, so nothing would
have shown a bare "error.<code>" to a user.

CI disagrees, and it is right to: check-translations.mjs requires every
locale to carry the full English key set, so English-only fails the gate
rather than degrading quietly.

The ten locales now carry both strings. These are my translations, not a
localization pass — worth a second pass by whoever owns the language, in
particular for the phrasing of "an administrator has locked them".

The uk file also loses two lines of stray 8-space indentation, normalized by
rewriting the file; no key or value changed with it.

* [client] Persist the profile before overlaying MDM on it (review item)

`netbird login` read the config, applied the MDM policy on top, and only then
provisioned the identity and wrote the result out. On a profile with no
identity yet — a first login — that write persisted the enforced values into
the user's own config file: an MDM-managed management URL or pre-shared key
became indistinguishable from one the user set, and stayed behind once the
policy was withdrawn.

Provisioning and its write now come first, and the overlay is applied to the
in-memory config afterwards, where it belongs: it is re-derived on every load
and never meant to reach disk from here. Server.getConfig already orders the
two this way; the two paths now agree.

Reported by cubic on the PR.

* [client] Assert against the stored config, not a resolved default (review item)

The login-gate test read the profile back with ReadOrGenerateConfig, which
resolves a default config in memory when the file is missing — and that
default's management URL is the very value the assertion checks. An erased or
mislocated profile would have passed the test instead of failing it.

The file is written by the test itself, so GetExistingConfig is the right
reader: it errors when the file is gone.

Reported by cubic on the PR.

* [client] Keep the mTLS pair off the gate's dry run (review item)

WouldChange runs the real apply() against a throwaway copy, and apply() loads
the client mTLS certificate and key from disk whenever the config names them.
So every gated SetConfig and Login read the pair — twice per request, once for
the normalization pass and once for the verdict — including requests that were
about to be refused or that changed nothing, and logged an error per request
when the files were missing. The gate used to be presence-based and never
called apply(), so this was new work on a request path.

The loaded pair feeds the connection and never the comparison: nothing in
apply() reads it back, and it does not move the `updated` verdict. A config
built only to be compared against now says so, and apply() skips the load for
it.

Reported by cubic on the PR.

* Makes it explicit that RenameProfile does write on disk

* [client] Provision the peer identity under the config lock (review item)

Login took the authoritative update-settings and privilege decisions under
guardedConfigMu, then released it and called getConfig, which mints the peer's
identity and writes the config out. Between that read and that write, a
SetConfig holding the same lock could land a change and answer its caller —
and then be overwritten by the config the login had already read.

The window is narrow: getConfig only writes when the profile has no identity
or no file, so in practice a first login racing a settings change on the same
profile. It is also narrower than before this branch, where the write happened
inside the reader on every read that filled in a default.

Provisioning now runs where the decision it belongs to runs: at the end of
authorizeAndPrepareLogin, with the lock already held, next to
persistLoginOverrides, which writes there too. No lock is taken that was not
held before, so the documented guardedConfigMu-then-mutex order is untouched.

getConfig keeps its behaviour by calling the same extracted helper; on the
login path it now finds the identity already there and writes nothing. The
other callers are unchanged, and still provision outside any lock — a
concurrent SetConfig is not part of their flow.

Reported by cubic on the PR.

* [client] Declare the probe marker to the debug-bundle field check

TestAddConfig_AllFieldsCovered walks Config by reflection and fails until every
field is either rendered in the debug bundle or listed as excluded with a
reason. The probe marker added for the gate's dry run was neither, so the
client unit suite went red on every platform.

It is excluded: it marks a throwaway copy built to be compared against and
discarded, so it is never set on a config anyone runs with, and rendering it
would only ever print false.

* [client] Provision the peer identity on the iOS login path

Key generation used to happen inside apply(), so a config loaded from JSON with
no keys got them in memory on the way in, the login worked, and the app stored
the result. This branch moved generation into EnsureIdentity, and nothing in
the iOS SDK called it.

The consequence lands on the flow the mobile logout sets up: logout clears both
keys in place so the next login registers a new peer. The app then hands that
keyless JSON to Auth.SetConfigFromJSON, and the login calls auth.NewAuth with
an empty WireGuard key, which fails on key size before the SSO flow starts —
the user cannot sign back in.

Auth.setBaseConfig now provisions, which covers both entry points (NewAuth and
SetConfigFromJSON). It mints on the base config, the one GetConfigJSON returns
for the caller to persist, and writes it to disk itself when the profile has a
file — non-atomically, like NewAuth's own write, since the tvOS App Group
sandbox blocks temp-file-and-rename.

Not covered by a test: the package builds only under GOOS=ios, which the test
jobs do not run. Verified by building and vetting for GOOS=ios/arm64.

Reported by pappz in review.

* [client] Name the resolving reader for what it does, not what it makes

ReadOrGenerateConfig reads the profile config and falls back to the defaults in
memory when there is no file. "Generate" reads as "produces and stores", which
is the opposite of the property the rename it came from was meant to advertise:
the read is pure, writes nothing and mints no identity.

ReadConfigOrDefault says the same without the side effect, and pairs with
GetExistingConfig, which fails where this one falls back. Its doc comment now
states the absence of a write rather than only the fallback.

Pure rename; the two remaining mentions of the pre-branch name ReadConfig in
the tests go with it.

Reported by pappz in review.

* [client] Read an emptied NAT list as the absent one it matches

apply() compared NATExternalIPs with reflect.DeepEqual, which calls a nil
slice and an empty slice different. Both mean the same thing — no NAT
mappings — and the two meet on a perfectly ordinary start: a profile stores
the absent list as JSON null and reads it back nil, while `netbird up` sends
CleanNATExternalIPs, an empty list, whenever NB_EXTERNAL_IP_MAP is set to
nothing, which a deployment template does by default.

So the gate saw a change where nothing changed and refused the request with
FailedPrecondition. That is the same deadlock this branch exists to remove,
reached through another field: a container with the kill switch on could not
come up, and `netbird up` reported "the daemon refused the settings update".

The DNS label list next to it already used slices.Equal, which treats nil and
empty as the same list. The NAT list now does too, and the last use of
reflect in the package goes with it.

Reported by pappz in review.
2026-10-06 11:39:22 +02:00
Nicolas Frati ab79aebd88 [misc] Share one license collection script across the UBI images (#8055)
* [self-hosted] Add a UBI image variant for the combined server

OpenShift and other Red Hat environments expect UBI-based images that run
as an arbitrary non-root UID. The proxy and rootless client already ship
-ubi variants; this adds the same for netbird-server, published as
<version>-ubi and ubi-latest for amd64 and arm64.

* [self-hosted] Check the license output path before creating temp files

The existing-output exit ran before the cleanup trap was registered, so it
left the two mktemp files behind.

* [self-hosted] Certify the netbird-server UBI image

Adds netbird-server to the Red Hat certification components. Its Partner
Connect component ID goes in the REDHAT_CERT_ID_NETBIRD_SERVER repository
variable.

* [misc] Share one license collection script across the UBI images

The client, proxy and combined images each carried a near-identical copy of
collect-licenses.sh, and the signal and relay variants would add two more.
The copies differed only in the Go package, build tags, component license
and the proxy's web licenses, which are now options of one script in
release_files/. The client gains the staged write the others already had.
2026-10-06 10:25:57 +02:00
Zoltan Papp 2b5293687f [client] Skip late session warnings on desktop and schedule them in the app on Android (#7548)
* Skip session warnings that fire after their window

The warning timers run on the monotonic clock, which does not advance
while an Android device is suspended. A timer armed for T-10 or T-2 can
therefore fire long after the window it was armed for, delivering a
"session expires soon" notification once that window is already gone.

Gate both callbacks on the wall clock at fire time: the T-10 warning is
skipped once the final-warning window has been reached, and the final
warning is skipped once the deadline itself has passed. Both set their
edge guard before returning so a skipped warning cannot fire again for
the same deadline.

* Harden the late-warning guards

Clamp a non-positive final lead to zero in the T-10 guard so a disabled
final warning cannot move the cutoff past the deadline, matching how
armTimerLocked already treats it.

Strip the monotonic reading from both sides of the comparison so the
guard measures wall-clock time regardless of how the caller built the
deadline. The production deadline comes from a protobuf timestamp and
has no monotonic reading; this keeps the guard correct for callers that
derive one from time.Now.

* Log the deadline and lateness on skipped warnings

Include the deadline and how far past the cutoff the timer fired, so a
debug bundle shows how long the device was suspended.

* Inject the clock into the late-warning guard and cover it with tests

The guard read time.Now internally, so the skip paths were reachable
only through a deadline already in the past and the boundary depended
on real time. Extract the comparison into isLate and read the time
through a nowFn field, so tests can place a resume anywhere around the
deadline without sleeping.

* Send the final warning when the T-10 timer fires inside its window

A suspend between roughly eight and ten minutes long made the T-10
timer fire inside the final-warning window and the final timer fire
after the deadline, so both were skipped and a user who resumed with
time left got no warning at all. When the T-10 timer fires late but
before the deadline, send the final warning in its place and mark it
fired so the delayed final timer does not repeat it.

* Respect dismissal when promoting a late warning to the final one

fireFinal skips the final warning once the user dismissed the deadline,
but the promoted path did not, so a dismissed deadline could still get
a final warning. Check the dismissal first, and give each skip reason
its own log line so an already-fired final warning no longer logs a
negative lateness.

* Add a deadline-only mode to the session watcher

Android will schedule its own expiry warnings from the deadline, so
the engine must not arm the T-10 and T-2 timers there. NewDeadlineOnly
keeps the deadline validation, the recorder propagation and the
logging, and skips only the timers, so the status snapshot the app
reads stays correct and an out-of-range deadline is still rejected.

* Use the deadline-only watcher on Android and drop the warning callbacks

The warning timers run on the monotonic clock, which does not advance
while the device sleeps, so a warning armed for T-10 could fire long
after its window. The app now schedules the warnings itself with
WorkManager, anchored to the wall clock, from the deadline it reads
through SessionExpiresAtUnix on every OnStateChanged.

Wire the deadline-only watcher into the android build and remove the
event-driven path from the gomobile surface: OnSessionExpiring, the
event subscription behind it and DismissSessionWarning, which the app
never called.

* Describe the late-warning guard without naming Android

The guard stays for the desktop builds, where a timer can also stall
across a sleep. Android no longer arms the timers at all.
2026-10-05 16:31:43 +02:00
Riccardo Manfrin f175e402c7 [client] stop offering to every peer when the relay transport drops (#7092)
* [client] stop offering to every peer when the relay transport drops

The relay transport is shared: one connection per relay server carries the
streams of every peer using it. When it drops, each of those peers gets a
Disconnected verdict from evalConnStatus even when ICE is still carrying its
traffic, because peerUsesRelay comes from HasRelayAddress(), which only reports
that management offered relay servers, not that we are connected to one. The
guard answers Disconnected with the aggressive retry, so every peer starts
sending offers over signal for a transport that no offer can restore: the relay
client's own guard is what reconnects it.

Feed relayManager.Ready() into the status inputs and return PartiallyConnected
when ICE is up and the missing side is the shared transport. That is the
existing "one path works, the other does not" branch, which retries three times
and then hourly instead of walking the exponential ladder forever.

Peers are not left waiting for the hourly tick: when the transport comes back,
Manager.onServerConnected notifies srWatcher, the guard resets the ticker to
800ms and iceState.reset() clears the hourly mode.

The verdict is unchanged when the transport is up but this peer is unreachable
over relay - it may have moved to another server, and only an offer carries its
new relay address - and in force-relay mode, where relay is the only transport.

* Renaming according to actual meanings

* Don't consider an in progress ICE as "partially connected"

when the relay is not..

* Aligns tests

* Address wrong comments
2026-10-05 15:59:31 +02:00
Zoltan PappandClaude Opus 5 bc44cdc37a [client] Fix browser login popup show from go (#7408)
* [client] Show the SSO login popup and open the browser from Go

The browser-login popup was created hidden and relied on its own webview
to size and show itself and to launch the external browser. On macOS a
hidden WKWebView gets throttled or suspended (App Nap / hidden-window
throttling), so on the first-use path nothing appeared and the browser
never opened, leaving the session-expiration dialog disabled until the
PKCE flow timed out. Reproduced by freezing the popup's WebContent
process: the old code showed nothing, the new code shows the popup and
opens the browser within 30 ms regardless of the webview state.

Show and focus the popup from Go right after creation and launch the
browser from Go on both the create and reuse paths. The popup's frontend
no longer shows or focuses itself, so the browser keeps the foreground
once it activates. This also fixes the reuse path, where a fragment-only
SetURL kept the mounted React tree and the once-only guard skipped
opening the browser for the new URI. Browser launch failures surface in
the error dialog instead of being swallowed.

* [client] Show every dialog window from Go once its frontend has painted

Dialog windows (browser-login, session-expiration, install-progress,
welcome, error) were created hidden and made visible only by their own
webview's Show call after sizing. A hidden WKWebView on macOS can be
throttled or suspended before that code runs, which left the window
hidden forever. The main and settings windows already avoided this with
the painted event plus a fallback timer, but that timer was armed on
WindowRuntimeReady, which a frozen webview never reaches either.

Route all dialogs through the same mechanism: the auto-size hook emits
the painted event instead of showing the window, Go shows and focuses
it on that event, and a fallback timer armed at creation shows it after
3 s regardless. The browser-login popup opens the browser in an
after-show callback so the browser still lands in front of the popup,
also on the fallback path.

* [client] Tie install-progress hidden-window restore to the current popup

CloseInstallProgress nils s.installProgress before calling w.Close(), so a
replacement popup can open before the old window's WindowClosing event runs.
The old callback then restored the windows the replacement had just hidden,
because the restore sat outside the identity check.

Guard the restore with the same check the state reset uses, and restore from
CloseInstallProgress itself so the programmatic close path still re-shows the
hidden windows — mirroring how CloseBrowserLogin already handles it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [client] Correlate painted reports with the window generation that sent them

A painted report carried only the window name, so a late report from a popup
that was already closed and replaced marked its replacement ready. The
replacement was then shown before its own frontend had rendered, which is the
blank-dialog case this flow exists to prevent.

Each dialog start URL now carries a monotonic generation token, echoed back by
ReadySignal, and a report whose token no longer matches the live window is
dropped.

* [client] Separate a window being painted from its frontend being mounted

One flag gated both showing a window and emitting to it, so the fallback timer
set it for a frontend that had not subscribed yet: the queued events were
flushed into a window that could not hear them, losing the login trigger and
the settings tab selection.

Showing is now gated on painted and emitting on mounted, and only a real
frontend report sets mounted. The fallback timer also moved to its own helper
so the runtime-ready hook can rearm it, giving the frontend a full budget to
mount rather than sharing one with webview boot.

* [client] Tag hidden windows with the popup that hid them

Windows hidden while a popup owned the screen went into one untagged list, so
whichever popup closed first restored all of them and emptied the list. An
install started during SSO login re-showed the main window the login popup had
deliberately hidden, and left the login popup with nothing to restore.

Each entry now records the popup that hid it, and a restore releases only that
popup's own entries. This also subsumes the manual filtering CloseRenewFlow did
to keep its own session-expiration window from being re-shown.

* [client] Cover the hidden-window bookkeeping with tests

application.Window carries unexported methods, so the hide/restore paths could
not be faked and the earlier tests could only assert which entries survived a
restore, never which windows were actually shown.

The bookkeeping now goes through hideableWindow, the four methods it needs,
with the window enumeration and the main-window raise behind seams that are nil
in production. That makes the case the owner tag exists for testable end to
end: an install started during SSO login restores only the login popup it hid,
and leaves the main window hidden until the login popup itself closes.

* [client] Report the first paint from unstamped windows too

The main and settings windows carry no generation token, so ReadySignal
saw an empty generation that already matched the ref's initial value and
never emitted the painted event. Those windows only became visible through
the fallback timer, and their frontend was never marked mounted, so the
login trigger and the requested settings tab stayed queued.

Start the ref from null so the first report goes out regardless of the
generation value.

* [client] Hand covered windows over when a popup closes under another

Closing the browser-login popup while the install-progress popup was
still up restored the main window the login had hidden, even though the
install popup was meant to own the screen until it finished. The owner
tag on each hidden entry only stops a popup from restoring another's
windows; it says nothing about what to do with its own when a second
popup still covers them.

Track which popups currently own the screen and, on restore, re-tag the
entries another live popup covers to that popup instead of showing them.
A popup is never handed its own window, so closing the popup on top still
brings the one below back.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-10-05 15:32:48 +02:00
Zoltan Papp 6b3cfbabd2 [client] Resolve the Android network route peer by HA unique ID instead of scanning the full status (#7705)
* [client] Build the Android network list from a single peer snapshot

Networks() called GetFullStatus() once per network to find the peer
that serves it, copying every peer state and taking eight recorder
locks each time. With 100+ peers and the UI calling Networks() from
every peer list change, this queued hundreds of callers on the status
recorder lock. Take one snapshot per call and index it by route.

* [client] Track the active route peer by HA unique ID in the status recorder

The Android network list resolved the route owner by scanning peer state
route keys. Those keys are the handler string: a prefix for static routes
and the domain pattern for dynamic ones, so the prefix-based lookup never
matched dynamic routes, and two networks sharing a prefix resolved to the
same owner.

The route watcher now records the chosen route peer under the route's HA
unique ID in the status recorder, and the Android binding looks the owner
up by that ID. The key is unique per network and independent of the
handler string format, so both anomalies are gone. The prefix-based
routeOwners helper is removed.

* [client] Record the active route peer before notifying listeners

AddPeerStateRoute and RemovePeerStateRoute fire the peer list change
callback and wake the status subscribers. The active route peer mapping
was written after those calls, so a Networks() call landing in between
found no mapping for the network and fell back to the first connected
peer, or kept showing the previous peer on removal. Nothing re-notified
after the mapping write, so the wrong peer stayed until the next peer
list change.

Write and delete the mapping before the notifying calls so a listener
reacting to the notification always reads the current owner.
2026-10-05 15:11:08 +02:00
Zoltan Papp 1c7d87d5fc [client,android] Generate debug bundle to file (#7528)
* [client] Add a debug bundle file export to the Android bridge

The Android app can only upload a debug bundle and hand the user a key.
Users who want to inspect what leaves their device before sharing it
have no way to get the zip itself. Add DebugBundleFile, which generates
the bundle into the cache directory and returns its path instead of
uploading; the app copies it wherever the user chose and removes it.

DebugBundle keeps its behavior. Both entry points share the unexported
debugBundle with an upload switch, so the body stays where it was and
merges cleanly with the MDM overlay change on main.

Because the file variant leaves the zip to the caller and the upload
variant only removes it after the upload finishes, a process killed in
between leaves a zip behind in the cache. Remove stale bundles before
generating a new one: RemoveStaleBundles deletes zips matching the
generator's pattern that are older than an hour. Remote debug jobs write
to the same directory, so younger files are treated as still in use.

* Preserve network map for debug bundle on Android

* [client] Keep exported Android debug bundles out of the stale cleanup

DebugBundleFile hands the zip to the caller, but the file kept the
netbird.debug.*.zip name that RemoveStaleBundles matches, so a later
debug run could delete it once it was older than an hour. Rename the
exported bundle to netbird.debug-file.*.zip after generation so the
cleanup only ever touches bundles no caller owns.

* [client] Warn when a stale debug bundle cannot be removed

A failed removal means bundles pile up in the cache directory, so log it
at Warn instead of Debug. A file that is already gone was removed by a
concurrent cleanup and is skipped silently.

* [client] Drop the outdated debugBundle comment

The comment still said the file variant leaves the zip in place, but it
is renamed by debug.ExportBundle since the stale-cleanup change.

* [client] Test that the network map reaches the debug bundle

Cover both halves of the path Android now relies on: the engine keeps
the latest sync response once persistence is enabled, and the bundle
generator writes it to network_map.json (anonymized or not) and omits
the file when there is no sync response.

* [client] Remove abandoned exported debug bundles after a day

An exported bundle is owned by the caller, but if the app is killed
before it copies and deletes the file, nothing ever removes it from the
cache directory. Let RemoveStaleBundles also match exported bundles,
with a 24 hour max age instead of the caller-provided one, so a bundle
that is still being saved survives while an abandoned one goes.
2026-10-05 14:29:28 +02:00
Riccardo Manfrin 9f8ddc7131 [client] Discover interfaces lazily in stdnet instead of at construction (#7346)
* [client] Discover interfaces lazily in stdnet instead of at construction

stdnet.NewNet and NewNetWithDiscover ended with

    return n, n.UpdateInterfaces()

handing back a non-nil *Net together with the discovery error. Three of the
five call sites (Engine.newWgIface, ice.NewAgent, SingleSocketUDPMux) logged
the error and kept using the instance, which is only safe as long as the
instance still works after a failed discovery.

That stopped being true when Interfaces() gained a lazily refreshed cache:
updateInterfaces sets lastUpdate only on success, so after a failed
construction the 30s cache guard never holds and Interfaces() returns an
error rather than the empty list it used to return. Feeding such an instance
to pion is worse than passing nothing at all - ice.NewAgent falls back to its
own stdnet when Net is nil, and the interface blacklist is applied separately
through AgentConfig.InterfaceFilter, so the fallback loses nothing. Instead,
a transient discovery failure (the Android bridge at boot, or an interface
disappearing between net.Interfaces() and Interface.Addrs()) turned into a
hard "error getting local interfaces" from ice.NewAgent, and aborted the STUN
and TURN probes, which never even need the interface list.

Since the accessors already refresh a stale cache on demand, the eager
discovery in the constructors is redundant: drop it, make both constructors
infallible, and let the discovery error surface at the call that actually
needs the interfaces. UpdateInterfaces had no callers left and is not part of
transport.Net, so it is removed along with it.

InterfaceByIndex and InterfaceByName read the cached slice directly and never
refreshed it, so they would have kept reporting ErrInterfaceNotFound forever
on an instance whose first discovery failed. They now go through the same
refresh path as Interfaces().

* [client] Warm the stdnet interface cache at construction

Moving discovery to first use regressed the privileged suites on the three
platforms that always build an ICE bind: Darwin, FreeBSD and Windows time out
in TestWGIface_UpdateAddr, TestRecreation, TestEngine_SSH and
TestEngine_MultiplePeers, while Linux stays green because a host with the
WireGuard kernel module takes the kernel-device branch and never drives the
mux that asks for interfaces.

interfaceFilter probes with wgctrl every interface the disallow list does not
already exclude. Discovering at construction ran that probe before the caller
had an overlay interface of its own; discovering at first use runs it after,
so on a userspace WireGuard platform the probe reaches the UAPI socket of the
same process. The tests reach it because they construct with a nil disallow
list, where the client passes DefaultInterfaceBlacklist and its own interface
is excluded by prefix.

Restore the original timing with an explicit warm-up. The constructors stay
infallible and the error is still reported by the accessor that needs the
interfaces, so the contract this branch is about is unchanged.

* Revert "[client] Warm the stdnet interface cache at construction"

This reverts commit 947e25288f.

* [client] Give the privileged tests the interface blacklist the client uses

The suites that create a WireGuard interface construct stdnet with a nil
disallow list, which the client never does: Engine passes
profilemanager.DefaultInterfaceBlacklist, whose "wt" and "utun" prefixes
exclude the overlay interface before the filter reaches its wgctrl probe.

With an empty list every interface reaches that probe, the one the test has
just created included, and on a userspace WireGuard platform the probe talks
to the UAPI socket of the same process. That is why Darwin, FreeBSD and
Windows timed out here while Linux, which takes the kernel-device branch on a
host with the module loaded, stayed green.

Pass the blacklist in both suites so they exercise the configuration the
client ships. client/iface declares the prefixes locally because
profilemanager imports it.

Also cover the constructors directly: the existing tests build the struct
literal, so nothing asserted that NewNet and NewNetWithDiscover leave the
cache cold.

* [client] Pass the blacklist in the remaining tests that build an interface

Same reason as the previous commit, four call sites it missed: engine_test,
the route manager and systemops suites, and the privileged DNS server suite
all construct stdnet with a nil disallow list and then create a WireGuard
interface. TestAddVPNRoute surfaced it on FreeBSD once the earlier two files
stopped timing out first.

client/internal/dns declares the prefixes locally; profilemanager imports
that package, so it cannot import profilemanager back.
2026-10-02 15:24:39 +02:00
Zoltan Papp 5ceca6e500 [client] Report both peers' state when the connect test times out (#7944)
Test_ConnectPeers fails every few weeks on the Linux runner with a bare
"waiting for peer handshake timeout after 30s". The failing logs show
both kernel devices up and both peers configured within a second, then
nothing for 30 s, which is six retries of the 5 s handshake retransmit
and so a condition that lasted the whole window rather than a race.
The failure cannot be reproduced locally and the log cannot tell
whether initiations were sent, whether they arrived, or whether only
one direction worked.

On timeout the test now prints each device's view of its peer, the
endpoint, the byte counters and the last handshake, so the next
failure says which of those it is. The comment also states that the
peers are kernel devices on the runner and that the first initiation
of each side is always lost to the other side not knowing the peer
yet.
2026-10-01 21:01:56 +02:00
Zoltan PappandDaniele Casciani 3906295446 [client] Add Homebrew cask e2e test (#7618)
* [client] Migrate macOS cask template to Homebrew install steps

Homebrew deprecated the postflight and uninstall_preflight cask stanzas
in favour of the declarative *_steps DSL, so every brew command that
evaluates netbirdio/tap now prints deprecation warnings. Once the
deprecation becomes a disable the generated cask stops loading and
netbird-ui can no longer be installed or upgraded through Homebrew.

The *_steps blocks take JSON-serialisable steps run in a sandbox rather
than arbitrary Ruby, so system_command is re-expressed as run/remove.
The two postflight blocks merge into one because a cask carries only a
single instance, preserving the original order. set_permissions moves
from a hardcoded /Applications to base: :appdir, matching what the
installer invocation already did. The launchctl fallbacks keep their
tolerant semantics through must_succeed: false, and remove is a no-op
when the plist is absent.

(cherry picked from commit df3756151f)

* [client] Test the macOS Homebrew cask on a disposable runner

The cask template only runs on real macOS with Homebrew, sudo and
launchd, so changes to it have never been exercised before merge. This
job installs the rendered cask on a GitHub macOS runner, walks the
uninstall through a running, stopped and missing daemon, and reinstalls
over the tap's published legacy cask, which is the path every existing
user takes on their next upgrade.

The fixture is the published cask itself rather than a pinned version
and checksums, so the test follows each release instead of breaking at
the next one. The installer scripts inside the signed archives are not
under test, which is why their paths are left out of the trigger.

* [client] Address SonarCloud findings in the Homebrew cask test

Positional parameters move into local variables and the scenario switch
gains an explicit default, so an unknown scenario fails instead of
silently running the plain install and uninstall path.

* [client] Make the Homebrew cask test deterministic with a stub bundle

The released installer script opens the UI as root, which never returns
on a headless runner, so a test that installs the published archive
hangs until the job timeout. The cask itself never looks past two script
paths and a version argument, so the test now builds a stub bundle on
the runner, serves it from a local HTTP server and renders the template
against it. The scripts ship without the executable bit, which turns the
0755 check into proof that set_permissions ran, and the stub records the
version and uid it received. The published archives are still downloaded
to assert the two script paths exist, and the published cask still
supplies the legacy stanzas for the reinstall scenario.

* [client] Drop the launchctl stderr check from the Homebrew cask test

The test asserts what the cask template promises: install, uninstall and
no deprecation warnings. Whether the uninstall steps print launchctl
errors is a review remark on the template, not part of that contract.

* [client] Retry the daemon start in the Homebrew cask test stub

A reinstall runs the previous cask's bootout and the new postflight
within a second of each other. launchd is still tearing the old daemon
down at that point, so loading the same label again fails with EIO. The
stub now retries the start for up to fifteen seconds, and the test still
verifies afterwards that the daemon reached the running state.

---------

Co-authored-by: Daniele Casciani <d.casciani@genogra.com>
2026-10-01 21:00:08 +02:00
Viktor Liu 6425b04200 [client] Only treat LocalSystem as a privileged identity by SID on Windows (#7889) 2026-10-01 22:01:10 +09:00
Viktor Liu 6c453a0f97 [client] Add debug cpu start and stop commands (#7749)
* Add debug cpu start and stop commands to profile the daemon without a restart

* Restore test globals on every exit and stop the daemon in the cpu profile test

* Add a no-updown flag to debug for

* Enable sync response persistence with --no-updown and reset flags between debug test runs

* Reset flags of every command between debug test runs

* Reset slice flags with Replace in the debug test helper

* Explain a running CPU profile in debug for and document cpu start and no-updown limits
2026-09-30 17:37:49 +02:00
Maycon Santos e72be6698f [client] Keep the advertised ICE session ID when following a remote restart (#7814)
A worker that saw a new remote session ID rebuilt its agent and also
picked a new local ID. On the answer path nothing carries that ID back,
so the next offer made the remote see a changed session, rebuild, and
answer with yet another ID. Two peers kept tearing down working ICE
connections on every offer and answer; nearly every answer in the
affected logs carried a new remote session ID.

Only a local restart changes the local ID now: a failed negotiation, as
before, and an explicit Close, which previously kept the old ID and left
the remote answering from a negotiation this side had abandoned.
Following a remote restart keeps the ID the remote already knows, so the
pair settles after one rebuild, also against peers that still pick a
new ID when following a restart.
2026-09-30 16:00:14 +02:00
Edward 96bfcc3600 [client] Classify Windows local accounts by NetBIOS name (#7628) 2026-09-30 11:10:14 +02:00
Viktor Liu 8edc120370 [client] Replace the eBPF WireGuard proxy with loopback endpoint addressing (#7316) 2026-09-30 10:41:49 +02:00
Eduard Gert 7ff709f565 [client] Keep the delete-profile dialog open until the delete finishes (#7752)
* [client] Keep the delete-profile dialog open until the delete finishes

Confirming a profile deletion closed the dialog straight away and left the
daemon call running in the background. A slow delete then looked like nothing
had happened: the dialog was gone, the profile was still listed, and the row
only disappeared whenever the refresh landed.

The confirm dialog now owns the action. It stays open with the confirm button
spinning, closes once the call resolves, and surfaces a failure after it is
gone rather than behind it. Nothing on the daemon path carries a deadline, so
the wait is bounded in the dialog instead: Cancel comes back after five seconds
and the wait is abandoned at thirty, which keeps a hung daemon from trapping
the user in a modal that cannot be dismissed.

A loading button keeps its own variant colours rather than the disabled skin,
which dimmed the spinner to grey on the danger variant, and blocks input
through aria-disabled and a click guard instead.

* [client] Match the theme and anonymize pickers to the language switcher

* [client] Settle a confirm dialog only from the run that opened it

Cancelling at the stall point leaves the action running, and the provider is
mounted once: take() read whichever settler the ref held when the stale run
finally finished. Open another prompt in the meantime and that run answered it
— a hung delete that later resolved confirmed a profile switch nobody accepted,
and its timeout closed the new dialog with an unrelated error.

Each run now remembers the settler it was dispatched for and settles only while
the ref still points at it. A late arrival finds a stranger there and answers
nothing.

* [client] Add the shared Select the theme and anonymize pickers use

The picker rework landed without the component both pickers import, so the
branch did not compile. Add it, and name it for what it is: a select of a few
options, with nothing settings-specific about it, so it sits with the other
input controls rather than under a name that discourages reuse.

* [client] Drop the Select header comment

* [client] Fix cubic comments

* [client] Name the Select trigger with the option it is showing

* [client] Name the language trigger with the language it is showing

* [client] Hold the confirm dialog for 15s before offering cancel
2026-09-29 17:08:33 +02:00
Riccardo Manfrin 164d92e78d [client] Stop dumping the whole device to clear one peer endpoint (#7632)
* [client] Stop dumping the whole device to clear one peer endpoint

Clearing a peer's endpoint has to remove and re-add the peer, because neither the
netlink API nor the wireguard-go UAPI can clear an endpoint in place. To keep the
peer's allowed IPs across that dance, RemoveEndpointAddress read them back from the
device: a full wgctrl.Device() dump on the kernel path, a full IpcGet plus text parse
on the userspace one. Both cost a round trip proportional to the entire network map,
both run under the interface lock, and both run on every relay and ICE transition.
On a routing peer with ~15700 peers that is megabytes of netlink traffic per
transition, at a measured 713 transitions per minute, with every other configuration
operation queued behind it. RemoveAllowedIP paid the same price for the same reason.

The allowed IPs cannot come from the caller: peer.Conn knows the peer's own overlay
addresses, while the routed prefixes are attached separately by the route manager's
refcounter, so a caller-supplied set would silently drop every route behind the peer.

The configurer is the only writer of its device's peer set, so it can keep an
authoritative mirror of what it configured and answer from memory instead. The mirror
is fed by every operation that changes a peer's allowed IPs and reset by a device
reconfiguration that replaces the peer set. A peer the mirror has not seen, which is
what an out-of-band reconfiguration leaves behind, still falls back to reading the
device and seeds the mirror from it.

Prefixes are unmapped on the way in, so a v4-mapped address compares equal to the
plain v4 prefix for the same network rather than registering as a second entry.

Measured on a userspace device, allocations to clear one endpoint:

  peers      64     256    1024    4096
  before   1452       -   21617       -
  after      91      91      91      91

* [client] Keep update-only allowed IP adds out of the peer mirror

AddAllowedIP configures the device with update_only, which is a silent no-op when
the peer does not exist, so its success says nothing about whether the device took
the prefix. Recording it unconditionally let the mirror hold a peer the device had
dropped, and RemoveEndpointAddress re-adds a peer without update_only: clearing the
endpoint of such a peer recreated it, carrying allowed IPs the device never held.
Allowed IPs are unique per device, so the recreated peer takes those prefixes away
from the peer that legitimately holds them.

This is not a theoretical window. Under lazy connections a routing peer's device
entry is torn down and re-created on the idle transition, and a routed prefix
re-added during that window is lost exactly because of update_only (#6863).

Allowed IP adds now merge only onto a peer the store already knows, which mirrors
the device: the operations that can create a peer record it, the update-only ones
do not. A peer missing from the store still falls back to reading the device.

* [client] Hand a prefix over to its new owner in the peer mirror

An allowed IP belongs to exactly one peer: configuring a prefix on a peer takes it
away from whichever peer held it before, and the configurer leaves that handover to
the device rather than removing the prefix from the previous holder itself, which is
what UpdatePeer's "wg will handle duplicated peer IP" refers to. The mirror recorded
the prefix on the new peer while leaving it listed under the old one, so clearing the
old peer's endpoint rewrote its allowed IPs from that stale list and took the prefix
back from the peer that now owns it. Traffic for the routed prefix then went to the
wrong peer. Reading the device before each write used to rule this out.

The store now tracks the owner of each prefix and performs the same handover, so
rewriting one peer's list cannot reclaim a prefix another peer holds.

Prefixes are also masked on the way in. A device stores them masked, so a caller
passing host bits would otherwise fail to match what a device fallback seeded and
could never remove that prefix by value. Conversion back from the device now keys
the v4-mapped decision on the mask width as well, so a genuine v6 prefix inside the
mapped range stays v6 instead of being dropped as an invalid v4 prefix.

* [client] Keep a mapped v6 prefix below /96 out of the v4 form

normalizePrefix unmapped any v4-mapped address before masking it, keeping the
original prefix length. For a genuine v6 prefix inside the mapped range, such as
::ffff:0:0/64, that pairs a v4 address with a v6 sized mask: netip.PrefixFrom
returns an invalid prefix and Masked turns it into the zero prefix. The store then
held a prefix whose Bits is -1, which cannot reproduce the allowed IP the device
was given, so re-adding the peer after an endpoint removal could fail once the
peer had already been removed.

Masking now comes first, and it also decides the address family: only a prefix at
least 96 bits long keeps the mapped marker through the mask, so anything shorter
inside that range is v6 and stays v6.

* [client] Record a peer created by a preshared key write

Setting a preshared key without updateOnly creates the peer when it is absent, and
Rosenpass applies a peer's first key exactly that way, since applyKeyLocked passes
the peer's initialized flag. The store ignored that operation, so the peer could
exist on the device while the store treated it as unknown.

An update-only allowed IP add on such a peer then succeeded on the device, which
moved the prefix away from its previous holder, while the store skipped the peer
and left the previous holder still claiming it. Clearing that holder's endpoint
rewrote it from the stale claim and took the prefix back, leaving the peer that
owns the route with nothing.

Every device operation that can create a peer now records it, which is the same
rule the update-only operations already follow from the other side.

* [client] Match a peer on the parsed key instead of its base64 form

getPeer scanned the device comparing Key.String to the caller's key. wgtypes.Key
is a 32 byte array, so it compares directly, while String base64 encodes it into a
fresh allocation on every iteration. The scan therefore allocated once per peer on
the device to find a single peer, and on a large network that is tens of thousands
of allocations per lookup.

The key is parsed once up front and the arrays are compared. Behaviour is
unchanged: the callers already parse the same key before reaching here, so the new
parse error is unreachable in practice and only guards the helper on its own.

* [client] Normalize prefixes on their way to the device

Prefixes were normalized when recorded but not when written, so a caller's raw prefix
reached the device while a different form was kept for it. The conversion is also where
a mapped prefix goes wrong: net.IPNet prints a v4-mapped address as v4 but takes the
length from its 16 byte mask, so ::ffff:10.1.2.3/64 is handed to a userspace device as
10.1.2.3/0 — an allowed IP matching every v4 address, on a peer that was meant to carry
one /64.

prefixesToIPNets now normalizes, and the two hand-built conversions in AddAllowedIP go
through it, so there is a single place where a prefix is turned into something a device
is given and it cannot disagree with what is recorded for it.

* [client] Parse the endpoint before configuring the peer

The userspace UpdatePeer parsed the endpoint address after the device had already been
configured, and returned on a parse failure. The device was then left holding a peer
that neither the activity recorder nor the allowed IP store had been told about, so the
peer was invisible to the wake path and the prefix handover for its allowed IPs never
happened, leaving the previous holder still claiming them.

The parse now happens before anything is written, so the only failure left after the
device is touched is one the caller cannot cause.

* [client] Keep the record when a peer removal fails

The two configurers disagreed: the kernel one dropped its record only once the device
had accepted the removal, the userspace one dropped it either way. Removing a peer is a
single device write, so a failure leaves the peer exactly as it was, with the allowed IPs
the record still describes. Dropping it there asserts nothing useful and only sends the
next caller to read the whole device back for an answer it already had.

The userspace one now follows the kernel and returns early on failure.

* [client] Write down what the allowed IP store does not guarantee

Two properties were relied on without being stated. The store's lock covers its map and
not the device write beside it, so consistency between the two rests on callers being
serialized, which WGIface does with its mutex; anyone removing that would have no way to
learn it mattered. And the fallback to the device only covers a peer the store has never
seen, so a peer first recorded from empty while the device already held prefixes keeps
only what was recorded, and the next endpoint removal drops the rest.

* [client] Key the allowed IP store on the parsed peer key

The store keyed on the textual key, so a lookup compared 44 byte strings while the
callers all held the parsed key already and the configurer had to carry both forms.
wgtypes.Key is a 32 byte array and compares directly, which is what getPeer was changed
to do for the same reason.

The store and its helpers now take wgtypes.Key, the callers pass the key they parsed on
entry, and the textual form survives only where something outside speaks it: parseStatus
reports peers that way, so the userspace fallback converts once for its scan.

* [client] Document the configurer methods the store changed

The exported configurer methods now carry what the allowed IP store made true of them:
when the mirror is reset, that a peer update merges its prefixes and takes them from
their previous owner, that an update-only add on an absent peer does nothing, and what
each side does with its record when a device write fails — where the two configurers
differ, since the userspace one reports a prefix it does not have and the kernel one
treats it as a no-op. mergeLocked states the lock its callers must already hold.

Docstrings that only restated the name of a test are left out; the tests explain the
scenario they set up in the body, where the explanation belongs.
2026-09-28 18:03:42 +02:00
Riccardo Manfrin aa1e66cc88 [client] Raise the daemon IPC receive limit above gRPC's 4 MB default (#7676)
Connections to the daemon were left on gRPC's own defaults, which cap a received
message at 4 MB. A detailed status carries an entry per peer, so on a large
deployment the response outgrows that cap and the command fails outright:

  netbird status -d
  Error: status failed: grpc: received message larger than max (4287609 vs. 4194304)

The limit is raised where the daemon dial options are built, so every caller
inherits it: the CLI, the desktop UI, the JSON gateway, and the SSH client and
proxy. It is overridable through NB_DAEMON_GRPC_MAX_MSG_SIZE for a deployment
that outgrows the new default too, mirroring what the management client already
does with NB_MANAGEMENT_GRPC_MAX_MSG_SIZE, and reusing its 16 MB default.

Only the receive direction needs raising. Requests to the daemon are small, and
gRPC does not cap the send side by default, so the daemon could already send a
response the caller then refused to read.
2026-09-25 11:17:34 +02:00
Riccardo Manfrin c7f610e6cd [misc, android] Build and lint the mobile Go code in CI (#7641)
Nothing in CI compiles the files behind //go:build android or //go:build ios.
The android bridge builds 7 of its 23 files on linux and skips client.go; the
iOS SDK is not built at all. The linter matrix picks a GOOS by picking a runner
OS, so it loads the same file set as the host build and never sees them either.
A type error in client/android/client.go therefore passes every check on its
PR, merges, and is discovered by netbirdio/android-client after sync-tag.yml
fires trigger_android_bump on the release tag.

The new Mobile workflow cross-compiles ./client/android/... for the GOARCH
values gomobile ships and ./client/ios/..., and vets the android bridge. The
new Android and iOS lint jobs run golangci-lint with GOOS/GOARCH in the job
env. No NDK, Xcode or gomobile is needed: these are library packages, so the
compiler type-checks them without a link step, and the dependency graph drags
in the android/ios-tagged files across client/iface, client/internal/dns and
client/internal/routemanager with them.

Linting those files for the first time surfaces one gosec G101 on the SSH
password-required marker. It is a sentinel string the Java side matches on,
not a credential, so it is suppressed at the declaration.
2026-09-25 10:21:20 +02:00
Theodor Midtlien 4c19226342 [client] Use POSIX style file read/write of json for windows (#7631)
* Use POSIX-like file read/write of json for windows + tests: allow renaming an open file.
2026-09-24 10:23:50 +02:00
Riccardo Manfrin 6e17f50040 [client] Validate the saved service parameters and pin the netsh lookup (#7584)
* [client] Export the only-owner-writable path check from elevate

Pure refactor, no behavior change: the existing checkOnlyOwnerWritable gets a
thin exported wrapper so callers outside the elevation path can reuse it. No
call site changes here.

* [client] Validate the saved service parameters before applying them

The install reads <stateDir>/service.json and applies it to the service it then
registers: its arguments, its config path and its environment. The restricted
ACL that saveServiceParams puts on the state directory is applied when the file
is written, which is not necessarily before the file is first read, so the
install now checks the file rather than assuming it.

A file whose ownership or permissions are not the ones saveServiceParams
produces is treated as absent, and the install proceeds with its defaults. The
check covers the directories above the file as well, so what is checked is what
is read.

* [client] Restrict which environment variables the service is registered with

--service-env, and the service.json it persists to, accepted any name. A small
set of them decides how a process resolves the executables and libraries it
loads, and the daemon needs none of those: it now refuses them when they are
passed explicitly, and drops them with a warning when they come back from a
service.json written by an older version, so an upgrade does not fail over a
variable nobody needs.

* [client] Resolve netsh by absolute path

The lookup consulted PATH first and fell back to System32, in both the copy the
userspace firewall uses and the one that tears the interface down. It now asks
Windows for the system directory, so the resolution no longer depends on the
environment the service happens to be started with.

* [client] Move the System32 lookup into a package both callers share

Pure refactor, no behavior change: client/iface and client/firewall/uspfilter
carried a copy each of the same function, and neither imports the other, so the
body moves to client/internal/wincmd — alongside winregistry, which is where
the client's other Windows-only helper already lives. Both call sites now read
wincmd.System32("netsh").

* [client] Cover the System32 lookup with a test

Asserts what the previous commits changed: the lookup is absolute, and neither
PATH nor %SystemRoot% moves it.

* [client] Refuse the loader environment families by prefix

Review follow-up on the previous commit:

- LD_* and DYLD_* are now refused whole rather than name by name. Their members
  differ per platform and libc and grow with new OS releases, so a list of them
  is out of date as soon as it is written — DYLD_FALLBACK_LIBRARY_PATH and
  DYLD_FALLBACK_FRAMEWORK_PATH were already missing from it.
- The names are folded to upper case only on Windows, where a variable is the
  same one however it is spelled. Elsewhere the environment is case-sensitive,
  so Path and PATH are two variables and only the exact spelling is the one that
  is read; the fold refused the wrong one.
- TEMP and TMP stay in the denylist, but the rationale and the message now say
  what they actually decide: where the service writes, not what it loads.
2026-09-23 21:47:48 +02:00
Riccardo Manfrin cb7ca8ef3f [client,management] Skip route firewall rule computation when no firewall (#7624)
* [client,management] Skip route firewall rule computation when no firewall

A peer that runs with the firewall disabled has no ACL manager and no
firewall to program, so nothing ever reads RoutesFirewallRules: the only
consumers are acl.Manager, which is reached solely when e.acl is set, and
the legacy-management probe in updateNetworkMap, which is guarded by a
non-nil firewall.

Building those rules is the most expensive part of a sync on a peer that
routes many network resources. On a 15k-peer deployment a debug bundle
showed getPeerNetworkResourceFirewallRules accounting for 62% of the
allocations of Calculate, and Calculate for effectively all of the
allocations of handleSync, which was taking 3.2s on average and holding
the engine lock for the duration.

Let the caller ask Calculate to leave the rules out. The client passes
its existing DisableFirewall setting; the management server keeps the
default and still produces them.

RoutesFirewallRulesIsEmpty is set from the resulting empty list, so a
receiver that would otherwise infer legacy management from an empty rule
set does not misread the skip.

* [client,management] Cover the skip flag through the envelope

Review feedback on #7624.

The components test compared only the length of the peer firewall rules, so
a change to their content would have passed while the message claimed they
came out unchanged. Compare the slices.

The skip path was also only exercised by setting the field directly on the
components, which bypasses the envelope conversion where
RoutesFirewallRulesIsEmpty is derived. That bit is what keeps the client from
reading skipped rules as a legacy management server, so it gets a test that
goes through EnvelopeToNetworkMap with the flag set.

* [management] Give the router a peer ACL so the rule comparison bites

Review feedback on #7624.

peer-router-1 appears in no peer ACL in the shared fixture, so its
FirewallRules came out empty and the equality assertion compared two empty
slices — it would have passed even if the peer rules were dropped entirely.

Add a policy covering the router and require the baseline to be non-empty
before comparing.
2026-09-23 11:50:44 +02:00
Viktor Liu f6109a3395 [client] Remove the empty GPO DNS policy store on Windows teardown (#7563) 2026-09-22 12:44:10 +02:00
Zoltan Papp bc0671fd21 [client] Fix peers not being notified when the relay connection drops (#7490)
* [relay] Signal relay disconnects through the conn context

AddCloseListener deduplicated listeners by comparing
reflect.ValueOf(callback).Pointer(). For a method value that pointer is
the address of the compiler-generated wrapper, not an identity bound to
the receiver, so every peer's w.onRelayClientDisconnected compared equal.

All peers on the home relay register under the same connectionURL key, so
only the first registration survived and the rest were silently dropped.
On a relay disconnect those peers were never notified: statusRelay stayed
connected and the reconnect guard never fired. The relayed net.Conn itself
was closed by closeAllConns, so nothing leaked, but the peer state machine
did not learn about it. Foreign relays had the same defect scoped to the
peers sharing that server.

Rather than fixing the deduplication, drop the peer-level listener registry
entirely. A relayed Conn now exposes Context(), cancelled when the
connection is torn down, with a cancellation cause naming the reason. This
is the same shape quic-go uses for its Conn and Stream types, and it
removes the whole class of problems around listener identity, lifetime and
deregistration: the signal belongs to the resource instead of a side table.

WorkerRelay watches that context in a goroutine whose lifetime matches the
connection. A watcher that wakes up for a superseded connection compares
the conn pointer against the current one and returns without touching the
state machine, so a fast relay reconnect cannot have a stale watcher tear
down the connection that replaced it.

Client.SetOnDisconnectListener stays: it is server-level and drives the
reconnect guard and foreign relay eviction, unrelated to peers.

handleRelayReady also checks the conn context, closing the race where the
relay dies between OpenConn and the readiness handoff and the peer would
otherwise build a WireGuard endpoint over a dead connection.

TestNotifierDoubleAdd covered the removed mechanism and is gone.
TestForeignAutoClose asserted nothing (both branches logged); it now waits
for the relay to leave the client map and fails if it does not.

* [relay] Fix build: return the concrete conn from Client.OpenConn

OpenConn now returns *Conn, but it still went through connContainer.netConn(),
which widens to net.Conn. The helper had one caller and only existed to produce
the interface value the signature no longer wants, so return container.conn
directly and drop it.

* [relay] Assert the local-close cancellation cause explicitly

The local-close test only rejected ErrServerDisconnected, so it would also
have passed for ErrPeerDisconnected or a bare context.Canceled. closeConn
cancels with net.ErrClosed, so assert that.

* [client] Ignore relay disconnects from superseded connections

The relayed conn watcher compared the conn pointer under relayLock, released
it, and only then tore the connection down. A new offer could install its
replacement in that window, so a watcher that validated the old pointer went
on to close the proxy of the connection that had already replaced it and
report the peer as disconnected while it was up.

Move the decision to where the teardown happens. Conn records which relayed
connection the current proxy was built from, and onRelayDisconnected takes the
connection the signal belongs to and drops it under conn.mu when it is no
longer the current one. Check and effect are now in the same critical section,
so the verdict cannot go stale before it is acted on.

This also covers the proxy read loops, whose disconnect listener took no
argument and had the same defect: it now names the connection it belongs to.
The WG timeout path keeps passing nil, since it deliberately tears down
whatever is current.

* [client] Bind the relayed conn reference to the proxy swap

relayedConnRef was set at the top of the readiness path, but wgProxyRelay only
changes at the end, in setRelayedProxy. The two failure returns in between —
newProxy and ConfigureWGEndpoint — left the reference pointing at a connection
that never became active while the old proxy was still installed. A disconnect
of that old, live relay would then be dismissed as belonging to a superseded
connection and never cleaned up.

Set the reference in setRelayedProxy, next to the proxy it belongs to. Both
success paths go through it and neither failure path does, so no failure branch
has to remember to roll anything back.
2026-09-21 17:00:37 +02:00
Zoltan Papp e70ec07320 Read the session deadline under the status read lock (#7550)
GetSessionExpiresAt took the exclusive lock for a plain field read, so
every caller queued behind writers and behind each other. The Android
SessionMonitor polls it from the main thread, and in the captured ANR
that is exactly where the main thread was blocked while hundreds of
peer-list callbacks held or waited on the same mutex.

d.mux is already an RWMutex and the other getters use RLock; this brings
the deadline read in line with them.
2026-09-15 12:05:57 +02:00
Zoltan Papp 2d28f9002a [client] Fix the Windows tray deadlock on re-entrant window creation (#7449)
* [client] Fix the Windows tray deadlock on re-entrant window creation

The Wails systray runs the left-click handler synchronously inside the
tray window procedure, and creating a window on a running app pumps a
nested Win32 message loop while WebView2 initialises. ensureWindow held
the non-reentrant createMu across that creation, so the second button-up
of a double click re-entered ShowWindow from the pump and blocked the
main thread on its own lock. A goroutine holding createMu while the main
thread pumped, and the Open* dialogs holding mu across NewWithOptions,
Show, Hide and InvokeSync, exposed the same inversion.

WindowManager now serialises creation with a per-slot creating flag and
queues the callers' operations until the window exists, and no Wails call
runs while mu is held. The tray click and second-instance handlers call
ShowWindow off the message loop.

* [client] Serialize window operations while a slot is being created

Callers arriving after the window is published but before the creator
has drained the queue took the existing-window fast path and could run
ahead of older queued operations, so a newer SetURL could be overwritten
by an older one. withWindow now queues every caller while the creating
flag is set and clears the flag only once the queue is seen empty under
the lock.

A factory panic or a nil window left the creating flag set and the slot
dead; creation and drain now reset that state on early exit.

hideOtherWindows records the windows it hid only when no restore ran
in between, tracked by a generation counter, and re-shows them otherwise,
so a restore racing the hide cannot strand hidden windows.

* [misc] Run the client/ui subpackage tests in CI

The three test workflows filtered the package list with a `/client/ui`
prefix match, which dropped the subpackages along with the package that
cannot compile without a frontend build. `services`, `preferences`,
`i18n` and `authsession` all carry Go-side unit tests that never ran,
including the window manager re-entrancy regression test.

Anchor the pattern so only `client/ui` itself is excluded. The linux leg
keeps the prefix match on 386, where only the 64-bit gtk4/webkitgtk dev
packages are installed and the Wails application package would fail to
link, and the alpine container job keeps it for the same reason.

* [misc] Run the client/ui subpackage tests on a gtk4 4.10 runner

The previous commit let the subpackages into the linux client job, where
client/ui/services failed to build: the wails runtime's linux cgo layer
uses GtkFileDialog, which arrived in gtk4 4.10, and the job's ubuntu-22.04
runner ships 4.6.

Move them to their own job pinned to ubuntu-24.04 and restore the linux
client job's original exclusion, leaving the 386 and privileged legs on
the runner they have used since 2024. The new job needs no build cache,
sudo or privileged tag, so it stays a few seconds long.

Darwin and Windows keep the anchored pattern from the previous commit and
already run these tests green, including the window manager re-entrancy
regression test on the platform the deadlock was reported on.

* [client] Defer a window close that lands while the window is still being created

WindowManager publishes a dialog's slot only after the factory returns,
and on Windows the factory blocks in the WebView2 embed pump. A Close*
arriving in that gap found a nil slot and returned without doing
anything, so the dialog appeared afterwards for a flow that had already
been cancelled. The pre-fix Open* dialog functions held mu across the
whole creation, which blocked a concurrent Close* until the slot was
set; removing that lock hold reopened this gap.

Close* now goes through closeWindow: while the slot is being created it
records a closer in pendingClose, and finishCreation runs that closer
before any queued operation, so a window that is going away is never
shown and Wails never sees a Show on a destroyed window, which would
recreate it. Ops queued behind a close are dropped; windowOp carries no
factory, so they cannot be replayed into a new creation, and the
frontend callers reissue on the next state change.

The browser-login slot uses the same restoring closer from both
CloseBrowserLogin and CloseRenewFlow, since the popup's WindowClosing
hook only restores on a user close. Where two closers race one
creation the first registered wins, so a later caller cannot replace a
restoring closer with one that does not restore.
2026-09-14 15:37:46 +02:00
Brandon Hopkins ec0c36b0e7 [client] Add light mode with system, light, and dark theme options (#7344)
* desktop UI light mode

* Theme review fixes plus macOS window outline fix

* Windows runtime chrome re-theming plus apply serialization

* Windows chrome threading and theme event ordering fixes

* Darken toggle and setting sidebar text

* resolve theme appearance, apply on UI thread

* read theme once per window

* Re-assert Windows dark opt-in after SetTheme

* split app-wide GTK theming from per-window chrome

* Update Wails dependency and checksums

* KDE tray icon panel fix

* Five review fixes: theme ordering, cgo dedup, KDE panel resolution

* Path guard hardening, toggle contrast, windows comment

* non-vacuous escape tests

* Default view edits

* Polish settings nav, controls, borders, and disc

* Profiles settings boarder, modals, and buttons

* Additional edits based on feedback

* Switch colors away from slight blue hue

* Update missing lang

* Fix vertical tab active view
2026-09-11 08:25:10 -07:00
Zoltan Papp 794956a7a3 [client] Fix relay instance address race (#7498)
Read the relay instance URL and IP atomically to prevent reconnects from mixing values from different connections. Extend existing connection and offer/answer logs with relay URLs and IPs to help trace mismatched advertisements.
2026-09-11 16:21:10 +02:00
Riccardo Manfrin a419e770d9 [client, proxy] Make the buffer-pool retune reachable while a device is stalled (#7452)
* [client] Track the WireGuard device on the engine as a lock-free handle

Add an atomic handle on the wg device next to wgInterface, stored once the
interface is up and cleared when it is closed. Nothing reads it yet, so this
is a pure addition with no behavior change; it exists so the next commit can
reach the device without taking syncMsgMux.

* [client] Retune the WireGuard buffer pool without the engine lock

SetPerformance took syncMsgMux before reaching the device. That lock is held
by handleSync while it adds and removes peers, and peer removal is exactly
what blocks when a device's buffer pool is exhausted: Peer.Stop waits on a
keepalive timer callback that is itself parked in WaitPool.Get. Raising the
cap is the way out of that state, so the call must not queue behind the lock
the stall is holding.

Read the device through the atomic handle instead. Device.SetPreallocatedBuffersPerPool
takes the pool's own lock and broadcasts, so the waiters wake up.

* [proxy] Extract the buffer-cap apply loop out of the perf handler

Pure move: the loop over the registered clients becomes applyBufferCap, with
the same sequential behavior and the same return values. Split out so the next
commit can change how it iterates without the diff also carrying the move.

* [proxy] Bound the perf endpoint so one wedged client cannot hold it

The apply loop was sequential and unbounded. embed.Client.SetPerformance goes
through the client lock, which Start holds for the whole of a startup, so a
single account that is busy or wedged delayed the new buffer cap for every
other account on the node -- on the endpoint whose whole purpose is to
un-wedge a node.

Apply to all clients concurrently and give the whole call a 5s budget.
Accounts that do not answer in time are reported in "failed" instead of
blocking the response.

* [client] Drop the device handle before closing the interface

close() cleared the atomic handle only after wgInterface.Close() returned, so a
concurrent SetPerformance could still load it, retune a device that is being
torn down, and report the change as applied for an engine that has stopped.
Clear it first, so the window closes before the teardown begins.

Reported by cubic on PR #7452.

* [proxy] Put the per-client retune behind a field

Pure refactor: applyBufferCap calls h.setPerformance instead of the client
method directly, and NewHandler wires it to setClientPerformance. Same call,
same behavior; the seam is what lets the next two commits be tested without a
live embedded client.

* [proxy] Do not report a finished retune as timed out

When the deadline fires, select chooses at random among the ready cases, so a
result already sitting in the buffered channel could be skipped and its account
reported as timed out even though the cap had been applied. Drain what is
buffered before declaring the rest pending.

Reported by cubic on PR #7452.

* [proxy] Keep one retune per account in flight

The 5s budget bounds how long the endpoint waits, not the work: SetPerformance
goes through the embedded client's lock, and on a wedged account Stop holds that
lock forever, so every retry left one more goroutine parked there.

Route each account through a single worker. A request that finds one already
running takes its result if it has landed, and otherwise reports the account
under "in_flight" instead of starting a second attempt. One stuck account now
costs one goroutine, no matter how often the endpoint is called.

Reported by CodeRabbit and cubic on PR #7452.

* [proxy] Make the retune budget a var

Pure refactor: perfApplyTimeout becomes a var so a test can shorten it instead
of waiting five seconds. Same value, same behavior in production.

* [proxy] Extract the buffered-result drain

Pure refactor: the loop that empties the results channel when the deadline
fires becomes collectBuffered. Same behavior; split out so it can be tested
on its own, which the inline version could not be without racing the deadline.

* [proxy] Cover the retune single-flight and the deadline drain

TestApplyBufferCapSingleFlightPerAccount fails without the worker registry:
five calls against a client stuck in its own lock start five blocked workers
instead of one.

TestCollectBufferedCountsResultsReadyAtTheDeadline pins the drain helper's
contract - buffered results counted, errors recorded, only unanswered accounts
left pending. It drives collectBuffered directly: through applyBufferCap the
two select cases race by construction, so an end-to-end version of it would
pass on the unfixed code about half the time.

* [proxy] Keep the worker alongside each pending account

Pure refactor: the pending set becomes a map to the account's worker instead of
an empty struct. Same membership and same behavior; the next commit needs the
worker to resolve an account whose result has not reached the channel yet.

* [proxy] Publish a retune result before releasing its slot

The worker sent its result last, after taking perfMu to remove itself from the
registry. That lock is taken once per account by every caller walking the fleet,
so a worker that finished on time could queue behind an apply over thousands of
accounts and land after the deadline. Send first, deregister after.

Reported by cubic on PR #7452.

* [proxy] Read the worker, not the clock, for a finished retune

Publishing earlier only narrows the window: a client that answers just before
the deadline can still be reported as timed out. At the deadline the workers
themselves are authoritative - a closed done channel means the retune finished
and w.err carries its outcome, ordered by the close. Consult them instead of
declaring every pending account timed out, and keep the timeout label for the
ones actually still running.

Reported by cubic on PR #7452.

* [proxy] Cover the finished-worker resolution at the deadline

Fails on the previous behavior with "applied = 0, want 1": every pending
account was labelled a timeout, including the one whose retune had already
completed.
2026-09-11 09:38:22 +02:00
Nicolas Frati 2f48dbea6a [client] Add a release-wired rootless UBI image variant (#7469)
* [client] Add a release-wired rootless UBI image variant

* [client] Add ARM64 to the rootless UBI image

* [client] Express license output validation as a guard
2026-09-10 21:41:59 +02:00
Zoltan PappandClaude Opus 5 9615d2ab16 [client] Report the remote jobs key in the MDM UI snapshot (#7485)
* [client] Report the remote jobs key in the MDM UI snapshot

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [client] Align the remote jobs snapshot key with the policy key

The snapshot field carried the JSON tag remoteJobsAllowed while the policy
key is allowRemoteJobs. GetConfigResponse.mDMManagedFields reports the raw
policy keys, and applyMDMRestrictions matches them against the struct's JSON
tags, so the field never turned true for a policy that set the key.

Every other field in Fields already uses its policy key as the JSON tag; this
was the only divergence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-10 12:06:30 +02:00
Nicolas Frati 15a684248c [client] Support arbitrary UIDs in rootless image (#7440)
* [client] Support arbitrary UIDs in rootless image

* [client] Keep rootless executables root-owned

* [client] Harden arbitrary UID image validation

* [client] Preserve executable access in rootless image

Keep the binary and entrypoint executable when deployments override the runtime group. Retain root ownership so non-root users cannot modify either file.

* [client] Verify rootless state reuse with a stable UID

Persisted profiles remain scoped to the creating UID. Verify same-UID container recreation without broadening application permissions, and document the Kubernetes volume permission behavior observed on OpenShift. Remove unused synthetic-user home metadata.

* [client] Separate image changes from invoking user fix

Keep this PR limited to resolving unmapped non-root invoking users. Move container permissions and their smoke test to a dependent image branch so they can be reviewed separately.

* [client] Restore invoking process user test

Retain coverage for successful current-user lookup without sudo. Numeric-identity fallback tests do not cover this existing behavior.
2026-09-10 12:02:19 +02:00
Viktor Liu d101f6cc46 [client] Redirect DNS port 53 with UDP and TCP DNAT instead of the eBPF forwarder (#7439) 2026-09-09 11:32:11 +02:00
Riccardo Manfrin d2e62e358a [client] Compare MDM-managed URLs as endpoints, not as strings (#7472)
A policy that enforces a management URL refuses any SetConfig or Login whose
URL differs from it. The comparison normalized only the default port, so
three ways of writing the very endpoint the policy names were reported as
conflicts:

  policy https://mgmt.example.com  vs  https://mgmt.example.com/     refused
                                       https://MGMT.example.com      refused
                                       https://mgmt.example.com:0443 refused

For an MDM-managed deployment whose stored or command-line URL is spelled
differently from the policy's value, that means every settings update is
refused with an MDMManagedFieldsViolation naming a field the caller did not
change. `netbird up --management-url https://MGMT.example.com` reproduces it.

The rules now live in util.SameServiceURL, and ConflictURL delegates: scheme
and host compared case-insensitively, the effective port normalized
numerically, a trailing slash ignored, and a path otherwise still part of the
identity so /other remains a divergence. Unparseable input falls back to
string equality.

util rather than either caller, because comparing two service URLs is
neither device management nor profile storage, and more than one place does
it: an MDM-enforced management URL against a requested one here, a stored
profile URL against a command-line one in profilemanager and the SSH gate.
Every copy of these rules that drifts turns an equivalent URL into a refused
request, which is how this one arose.

CanonicalURL is left alone: besides comparison it is the canonical value
handed to mdm.Restrictions and to the Android and iOS Preferences getters,
and normalizing what those return is a separate decision.
2026-09-08 16:32:16 +02:00
Riccardo Manfrin bb4de1d008 [client] Read MDM boolean keys delivered as JSON numbers (#7471)
encoding/json decodes every JSON number into float64, so the policy
values the mobile loaders produce never contain int or int64. GetBool
accepted both of those but not float64, so a managed boolean pushed as
1 or 0 — how some MDM consoles normalise flags — was reported as
unreadable while the key still counted as managed: the policy was not
applied, and the conflict gate rejected both values the user could pick
for that field.

The rejected-float assertion predates the JSON channel. It came with the
registry and plist loaders, where a real number for a flag is a
configuration mistake; on the JSON channel an integer is the only shape
a number can take. GetInt already accepts float64.
2026-09-08 15:06:46 +02:00
Riccardo ManfrinandZoltan Papp e14006ddc1 [client] mobile MDM bridge — iOS + Android setMDMPolicyFetcher entrypoint (#6435)
* MDM Android mobile wiring

* Removes dead code

* Removes static vars

* Now we need to apply MDM in the GetConfig

* You now need to explicitly call these around

* Adds iOS wiring

* Resolve merge conflicts from main

- login.go: keep both new imports (mdm + nbnet + server)
- ios/NetBirdSDK/client.go: additive struct-field merge (mdmLoader + stateMu/connectClient/config)
- setconfig_mdm_test.go: adopt new withMDMPolicy(t, s, policy) signature; fix stray old-signature call in TestSetConfig_MDMAllow_ManagementURLPortNormalized

* Convey MDM overlay config to Debug Bundle output

Aligns to other clients OSes behavior

* Solved conflict in client.go

* Fixup helper withMDMPolicy -> configWithMDM

* Fixup after merge

* Resolve merge conflicts

* [client] Move MDM enforcement logic into a shared Go layer (#7319)

The mobile bridges only carried the policy fetcher, leaving every
enforcement decision to the native apps: the desktop derived its UI
restrictions in the Wails service layer, the daemon kept the conflict
machinery in the server package, and both mobile bridges duplicated the
JSON fetch adapter. Anything the native side had to reimplement was a
place for iOS and Android to drift apart.

Enforcement now lives in client/mdm and is consumed identically by all
three platforms:

- conflicts.go holds the value-aware conflict checks lifted out of the
  daemon, so the same normalization (canonical URLs, PSK sentinel echo)
  applies wherever a config change is validated.
- restrictions.go derives the UI enforcement snapshot from a policy and
  renders it in the JSON shape the desktop frontend already consumes.
  The service-layer types become aliases, keeping one source of truth.
- jsonloader.go replaces the adapter that was copy-pasted into both
  bridges.
- changedetector.go moves change detection off the native side: the
  caller forwards the OS notification and asks whether the managed
  configuration actually changed, instead of diffing dictionaries
  itself.

The mobile bridges gain the enforcement the daemon already had. The
Preferences getters resolve managed keys from the policy, so a naive UI
shows the enforced value; Commit rejects a staged change that diverges
from a managed key; NewAuth resolves the managed management URL before
persisting the config and overlays the policy on it, so a login can no
longer run against a URL the policy forbids. Android's profile
mutations fail closed when disableProfiles is set.

NewAuth takes the fetcher as a required argument rather than keeping a
policy-blind overload: the apps consume this code as a submodule, so a
compile error at the bump is the point. The mobile PSK getter is
replaced by a presence check — the key has no reason to cross the
bridge, and not returning it means the native side needs no redaction
sentinel of its own.

* [client] Resolve the main merge conflicts in the MDM integration

The merge commit was recorded with the conflict markers still in the
tree. Resolve them so the branch builds again:

- client/ios/NetBirdSDK: keep both the mdm and mobile imports, and keep
  the mdmLoader/mdmDetector fields next to main's stateMu documentation.
- client/server/mdm.go: drop the conflict helpers main added locally,
  they already live in the client/mdm package on this branch, and keep
  the new checks main introduced (allowRemoteJobs, enableLocalMetrics,
  localMetricsAddress) as calls into the package-level helpers.
- client/mdm/conflicts.go: add ConflictStringPtr, the presence-aware
  string check main needs for the optional localMetricsAddress field.
- Port the two tests main added over the per-Server loader helper and the
  configWithMDM helper, both of which replaced the package-level policy
  injection this branch removed.

* [client] Reject explicit empty PSK when MDM enforces a pre-shared key

The SetConfig, Login and mobile Commit conflict checks collapsed the PSK
to a plain string, so an explicit empty value was indistinguishable from
an unset field and slipped past the MDM gate, clearing the persisted key.
Carry the optional field as a pointer through ConflictStringPtr, treating
only the redaction sentinel as a no-op echo. ConflictString had no other
callers and is removed.

* [client] Apply MDM overlay on the preloaded iOS config in Run

Run only overlaid the MDM policy when the config was loaded from file,
so the tvOS path fed by SetConfigFromJSON started with unmanaged
settings. Apply the overlay after the config source is selected, as the
other resolution sites already do.

* [client] Gate non-active profile logout behind the MDM profiles switch

The mobile ProfileManager let LogoutProfile clear credentials of any
profile even when disableProfiles was enforced. Follow the daemon's
validateProfileLogout semantics: logging out of the active profile is a
plain logout and stays allowed, logging out of any other profile is
profile management and is rejected under the policy.

* [client] Resolve the managed management URL through the MDM overlay on mobile

NewAuth on Android and iOS replaced the caller URL with the raw policy
value before persisting, so a malformed managed URL failed config
validation and blocked the login instead of being skipped with a warning
like the overlay does. Preferences.GetManagementURL likewise echoed the
raw policy string to the native UI even when the overlay had rejected it.

Follow the daemon: persist the caller URL, overlay the policy on the
resolved config, and report the overlaid ManagementURL as the effective
value.

* [client] Clean up MDM review leftovers

Drop the unused ChangeDetector.Current, point the stale LoadPolicy
comment references at Loader.Load, and move the profileEmail godoc back
above its function.

* [client] Check remote jobs and local metrics keys in the mobile MDM conflict gate

MDMConflicts skipped allowRemoteJobs, enableLocalMetrics and
localMetricsAddress even though the overlay applies all three and the
daemon gate already checks them, so a mobile Commit could persist values
diverging from the enforced policy. Align the list with the daemon.

* [client] Silence the deprecated PreSharedKey lint in the login conflict test

The legacy LoginRequest.PreSharedKey field is deliberately exercised by
the test, matching the nolint already carried by the production path.

* [client] Publish the mobile MDM loader and detector atomically

SetMDMPolicyFetcher wrote the loader and change detector as two plain
fields that Run, the OS-change callback and the restrictions getter read
from other threads without synchronization. Hold both behind a single
atomic pointer so a registration is published as one unit and readers
always observe a matching loader and detector pair; Preferences gets the
same treatment for its loader. Exported signatures are unchanged.

* [client] Report the MDM-overlaid remote jobs value from mobile Preferences

GetRemoteJobsAllowed returned the staged or persisted value even when
the policy manages allowRemoteJobs, so the native settings UI could show
a value the Commit gate would reject. Resolve it through the overlay like
GetManagementURL does.

* [client] Stop persisting the MDM-overlaid config after mobile logins

NewAuth already writes the config through UpdateOrCreateConfig before
the MDM policy is overlaid, and the login itself never mutates the
Config. The post-login WriteOutConfig calls therefore only rewrote the
same file with the enforced ManagementURL and PreSharedKey in it, so a
removed or changed policy kept acting through the persisted values.

* [client] Document that the MDM overlay on Config is not reversible

ApplyMDMPolicy promised that an empty Policy clears a prior overlay, but
applyMDMPolicy only resets the enforcement metadata and the runtime-only
upload URL; the enforced ManagementURL, PreSharedKey and flags stay. Every
lifecycle owner resolves the base Config again before applying, so state
that contract instead of the reversibility that was never implemented.

* [client] Re-resolve the tvOS preloaded config before every MDM overlay

The iOS Client kept the config parsed from SetConfigFromJSON and applied
the MDM overlay onto that same instance on every Run, IsLoginRequired
and DebugBundle, so a key removed from the policy stayed enforced. Store
the JSON instead and parse it per load through one loadConfig path.

Auth serialized the overlaid config from GetConfigJSON, which tvOS then
persisted to UserDefaults and fed back as the preload. Keep the resolved
config as the base, run the login on a JSON round-trip copy with the
overlay, and return the base from GetConfigJSON.

* [client] Serve the MDM-managed management URL without touching the config file on mobile

Preferences.GetManagementURL resolved a managed URL by reading and
overlaying the persisted config, so a corrupt file or the tvOS sandbox
turned an enforced URL into a read error. Return the canonical managed
value directly, the same string BuildRestrictions already hands to the
UI, and only fall back to the staged or persisted value when MDM does
not manage the key.

NewAuth validated the caller-supplied management URL before the overlay
ran, so a malformed or echoed value blocked or persisted under an MDM
policy that already dictates the URL. Ignore the caller value while the
key is managed; the login runs against the overlay either way.

* [client] Align the MDM loader docs with the fetcher precedence and make disableAdvancedView a tristate

NewLoader, PolicyFetcher and the darwin/windows loadPlatform docs claimed
the fetcher is unused on desktop, while every loader returns its values
when one is injected. That precedence is the seam the server tests rely
on across platforms, so the docs now describe it; production desktop
callers still pass nil and keep the registry / plist authoritative.

Fields.DisableAdvancedView collapsed "managed and false" into the same
JSON as "not managed", unlike AllowServerSSH and the daemon's optional
proto field. Carry it as a *bool so the UIs can tell the two apart; the
desktop reflect loop skips pointer fields already, and the mobile
decoders treat null as not managed.

* [client] Clean up MDM review nits

- ResolveConflicts treats a managed key whose ConflictCheck has no Check
  as a conflict instead of dereferencing nil.
- Ticker.Run and ChangeDetector.Changed share policyChanged so the diff
  semantics and the log line cannot drift apart.
- TestLoader_NilFetcherReturnsEmpty skips on windows/darwin, where a nil
  fetcher reads the real registry / plist.
- The profilemanager test loader checks GetInt before GetBool so integer
  keys survive the round trip, and the PSK tests use the exported
  redaction sentinel.

* [client] Fix int policy values coercing to bool in the MDM test helper

withMDMPolicy rebuilt the policy map by trying GetString, then GetBool,
then GetInt. Policy.GetBool accepts native ints (non-zero means true), so
an int-valued key such as wireguardPort round-tripped through the helper as
the bool true and GetInt was never reached. Try GetInt before GetBool, as
the profilemanager helper already does; GetInt does not coerce bools, so
booleans still fall through to GetBool.

No test sets an int key today, so this was latent: the first test to
exercise the wireguardPort conflict gate would have seen ConflictInt64
report a conflict for every value, including a matching one.

---------

Co-authored-by: Zoltan Papp <zoltan.pmail@gmail.com>
2026-09-08 11:53:07 +02:00
Zoltan Papp 15c0a2903d [client] Return the context error when the SSH handshake fails with it (#7426)
* [client] Return the context error when the SSH handshake fails on a context deadline

The handshake mapped the context deadline onto the socket but returned the
raw socket error. Which error surfaces depends on a race between the x/crypto
ssh readLoop and kexLoop goroutines: the kexLoop write fails with i/o timeout
and closes the conn, and the readLoop then reports use of closed network
connection. Callers checking errors.Is(err, context.DeadlineExceeded) never
matched, and TestSSHClient_ContextCancellation flaked on the FreeBSD job.

Handshake now wraps the context error when the context is done or its
deadline has passed. The deadline comparison is needed because the socket
deadline and the context timer fire independently, so ctx.Err() can still be
nil when the deadline-triggered socket error arrives.

* [client] Close the silent test server conn without racing t.Cleanup

The accept goroutine registered the conn close via t.Cleanup, which can run
after the test's cleanup list has already been drained, leaving the accepted
connection open. The goroutine now holds the conn until a cleanup-closed
channel signals the end of the test and closes it on the way out.

* [client] Bind the SSH handshake to the context instead of a socket deadline

Mapping only the context deadline onto the socket left context cancellation
unobserved: an in-flight handshake kept running until the deadline, and the
error classification had to guess whether a raw socket error was caused by
the deadline. Closing the conn from context.AfterFunc covers both deadline
and cancellation, and ctx.Err() is already set by the time the close-induced
error surfaces, so the time-based DeadlineExceeded attribution is no longer
needed. The stop() result guards the window between a successful handshake
and the AfterFunc firing so a closed conn is never handed back as a client.
2026-09-07 15:50:28 +02:00
Brad Ison 76ea72237f [management] Add Agent Network managed proxy to the API spec (#7433)
Defines the cloud-side managed gateway provisioning surface
(POST/GET /api/integrations/agent-network/managed-proxy) and its
response objects so clients consume generated types instead of
hand-written ones. POST is idempotent: 202 when the call starts (or
restarts) provisioning, 200 when a deployment already exists; 409
names an already-assigned endpoint the managed flow does not own and
503 signals temporarily exhausted endpoint allocation.
2026-09-04 17:11:27 +02:00
Zoltan Papp 5cb6b0d33b [client] Assign the Android TUN address as a host prefix (#7414)
Android 16+ local network protection derives the blocked prefixes from the interface address prefix. A /16 address turns the whole overlay into a local network, so apps without ACCESS_LOCAL_NETWORK cannot reach any peer. Pass the address as /32 and /128 and add the overlay networks to the route list that the Android side turns into VPN routes, on both the initial create and the renew path.
2026-09-04 15:10:36 +02:00
Zoltan Papp 825389818c [client] Gather fresh system info on every management sync stream connect (#7409)
* Gather fresh system info on every management sync stream connect

The engine collected the peer meta once at start and reused the same
Info for every Sync stream reconnect, so a mobile network switch that
redials management kept reporting the old local network addresses.
The peer network range posture check was then evaluated against stale
data until the client restarted.

Sync now takes a gatherer that runs at each stream connect. The
gatherer is cheap: GetInfo plus the cached posture check file results,
kept in the new system.InfoSource, which the engine refreshes whenever
the checks list changes. No process enumeration runs on the reconnect
path.

Also fix the management mock server calling itself instead of SyncFunc.

* Evaluate the login response posture checks before the first sync connect

The engine starts with the checks the login response carried, and the
first sync stream request used to send their evaluated file results.
After moving the gather into InfoSource, the stream opened with an empty
cache and the first sync response did not refill it, because its checks
equal the ones the engine already holds. Desktop peers therefore never
reported process or file posture results.

Seed the cache once before the first connect, where the old gather ran,
so a timed out evaluation still falls through to the address-only info.

* Harden the sync info source against nil callbacks and shared slices

A nil getInfo opens the stream without metadata, as a nil sysInfo did
before. The cached posture results are a copy, so the Info returned by
Refresh cannot alias the snapshot later Current calls report. The
exclusion test asserts the remaining address count so it cannot pass
vacuously on a single-address host.

* Retry a posture check refresh that timed out or failed to sync

The checks list was recorded before the gather ran, so once the gather
timed out or SyncMeta failed, the next sync response carrying the same
list matched the recorded one and nothing retried. The peer kept
reporting the previous posture results until the list changed again.

Record the checks only after the meta reached management, so a failed
cycle is repeated on the next sync response.

* Log the skipped posture refresh, let the mock Sync return errors and deflake the reconnect test

* Drop the nil guard around the sync info callback

* Send the refreshed info on the first sync connect instead of gathering it twice
2026-09-04 15:07:01 +02:00
Viktor Liu 7c1253004b [client] Renew the Android TUN only when the routes it carries change (#7396) 2026-09-03 12:22:19 +02:00
Viktor Liu bb233c72b6 [client] Rebuild the overlay listeners when the TUN is renewed (#7397) 2026-09-03 12:22:00 +02:00
Zoltan Papp fbd4730f0b [client] Expose the remote jobs opt-in in the Android and iOS SDK preferences (#7406)
Remote jobs (debug bundle requests from management) are gated behind
Config.RemoteJobsAllowed, which defaults to false and could only be
enabled through the CLI flag or an MDM policy. The mobile SDKs had no
way to set it, so the mobile clients always refused the job.

Add GetRemoteJobsAllowed/SetRemoteJobsAllowed to both mobile
Preferences types, following the existing ServerSSHAllowed accessors,
so the apps can offer a settings toggle for it.
2026-09-03 11:47:32 +02:00
Maxim Egorov b0e03038ed [client] Pick the probe port from the system in Test_freePort (#7404) 2026-09-03 10:03:24 +02:00
Viktor Liu 778b3b3264 [client] Unify peer and route ACL filtering with multi-source rules (#6322)
* Unify peer and route ACL filtering with multi-source peer rules

* Remove partial userspace firewall mode and open foreign chains via a table-less allower

* Snapshot iptables rule maps before persisting state

* Scope userspace firewall wildcard source rules per address family

* Install nftables peer filter and mangle rules in a single transaction

* Share the iptables jump rule spec between install and cleanup

* Fix legacy ACL source wildcard and keep rollback tracking on delete failure

* Fix CI: recognize multi-value port set lookups in tests and correct PeerIP lint suppression

* Fall back to per-prefix filter rules when ipset is unavailable

* Annotate legacy PeerIP usages in ACL tests and fix import formatting

* Keep firewall rule bookkeeping in step with the kernel on replace and teardown

* Release the routing reference when the route manager shuts down

* Keep set references and rule tracking consistent when a routing rule fails
2026-09-03 10:02:50 +02:00