Commit Graph

3457 Commits

Author SHA1 Message Date
Viktor Liu
be33e9eaf4 Merge branch 'main' into embedded-vnc 2026-09-02 19:10:03 +02:00
Viktor Liu
336fc9aaaa Drop the VNC port helpers and approval kind nothing calls 2026-09-02 18:11:36 +02:00
Pascal Fischer
8a5e940c84 [management] remove old math rand lib (#6836) 2026-09-02 17:52:51 +02:00
Riccardo Manfrin
7b22d55bf6 [client] Bind the cached SSH JWT to the local caller that obtained it (#7378)
* [client] Bind the cached SSH JWT to the local caller that obtained it

Record the identity that obtained the token and return it only to that
same identity, comparing the account alone: the group set and the
elevation flag describe what a token may do rather than who it belongs
to, and the same user may call once elevated and once not.

A control channel that carries no caller identity gets a miss on read
and stores nothing on write, matching how the other ipcauth consumers
fail closed.

Clear the entry when the session it speaks for ends: logout, down and
profile switch.

* [client] Cover the profile-switch path of the SSH JWT cache

The cache being correct buys nothing if a handler around it forgets to
clear it, and SwitchProfile had no test at all.

Point the profile globals at a temp dir holding a single default profile,
which is the one ActiveProfileState.FilePath resolves without consulting
the current OS user, and call SwitchProfile with no request so neither
the switch itself nor the profile-list event is involved.

* [client] Report the SSH JWT cache in the no-identity startup warning

daemonServerOptions already warns once, at startup, about what a control
channel with no caller identity gives up. Name the SSH JWT cache there
too, on both the TCP and the no-peer-identity-primitive paths.

The per-request logs in cachedJWT and WaitJWTToken drop to Debug: the
condition is expected and handled on such a channel, the caller simply
re-authenticates, and repeating it on every SSH authentication buried the
one message that is actionable.

* [client] Stop the local-metrics manager leaking out of the profile test

localmetrics.NewManager runs a goroutine until its context is done, and
the test handed it context.Background(), so the manager outlived the test
and stayed in the test binary for every case that followed.

* [client] Keep the cached SSH JWT across a down/up cycle

Clearing the cache in cleanupConnection also caught Down, which ends the
connection and not the session: the peer stays enrolled, `up` reconnects
without going back to the IdP, and the token still belongs to the same
NetBird identity. With a long cache TTL that cost the owner a fresh
device-code flow for nothing, since the owner binding is what keeps the
token away from other local accounts.

Clear it on the two paths where the session really ends and the next one
may belong to a different NetBird user: profile logout when the profile
is the active one, and active-profile logout. SwitchProfile already
cleared it on its own.

* [client] Resolve the merge conflict in the profile-logout cleanup

main extracted the inline profile-logout cleanup into
cleanupAfterProfileLogout, which this branch had edited in place to clear
the SSH JWT cache. Take main's helper and move the clear inside it.

The helper returns early when the profile that was deregistered is not
the active one, so the cache is still only cleared when the session that
owns the token actually ends.

* [client] Do not cache an SSH JWT obtained under a session that ended

WaitJWTToken polls the IdP with s.mutex released, and that wait can run
for as long as the user takes in the browser. A logout or a profile
switch in the meantime clears the cache, but the poll then completed and
stored its token anyway, so the entry the next session read belonged to
the previous one.

Give the cache a generation that clear advances. WaitJWTToken takes the generation
before the wait and hands it back to store, which keeps the token only
while the generation still matches.

The two mutexes are distinct, so this was never a data race and the race
detector could not have found it: the window is between two separately
locked sections.

* [client] Make the profile-switch test switch a profile

SwitchProfile with a nil request skips switchProfileIfNeeded, so the test
only covered the no-op path and would have passed with profile-transition
invalidation broken. Create a second profile and name it in the request,
then assert the active profile actually moved before checking the cache.

Also correct the comment on the Down test: the logout handlers do call
cleanupConnection. What changed is that clearing the cache is no longer
one of the things cleanupConnection does.

* [client] Take the SSH JWT cache generation when the flow is created

WaitJWTToken read the generation after validating the device code, but
the flow it belongs to is created earlier, in RequestJWTAuth, and
SwitchProfile does not reset s.oauthAuthFlow. A profile switch between
the two therefore advanced the generation before it was ever read: the
guard compared the new session against itself and let the token through,
which is the case it exists to stop.

Record the generation on the flow when RequestJWTAuth creates it, and
read it from there. The whole span from the request to the IdP answering
now counts as one session for the cache.

* [client] Correct two test comments the clear-on-Down change invalidated

Moving the clear out of cleanupConnection left two comments describing
the old behaviour: newTestServer said cleanupConnection clears the cache,
and the comment above TestJWTCache_ClearDropsTheEntry listed Down among
the callers of clear. Neither is true any more.

* [client] Read the SSH JWT cache generation before the IdP round trip

RequestJWTAuth read the generation where it stored the flow, which is
after RequestAuthInfo has talked to the IdP. A logout or a profile switch
during that call advanced the generation first, so the flow recorded the
new session's value and the later store was accepted: the window moved
rather than closed.

Read it with the config, under the same s.mutex section. SwitchProfile
holds that mutex across its own clear(), so the config and the generation
cannot be torn apart by a switch.
2026-09-02 14:04:09 +02:00
Zoltan Papp
c3cf7c0c37 [client] Clarify that metrics ingest X-Peer-ID is not a credential (#7363)
* [client] Clarify that metrics ingest X-Peer-ID is not a credential

The ingest endpoint is intentionally unauthenticated: it accepts telemetry
from peers of both cloud and self-hosted deployments, and for a self-hosted
peer there is no shared trust anchor to authenticate against. The X-Peer-ID
header is a correlation tag whose format check exists to bound InfluxDB tag
cardinality.

Both the function name (validateAuth) and the 401 response implied an
authentication control that was never there, which invites the reading that
the check can be bypassed. Rename it to validatePeerIDFormat and return 400,
matching the other input validation failures in the same handler. Document
the intent in the godoc and the infra README.

No behavioural change for clients: push.go classifies responses by 2xx range
rather than by status code, so 400 and 401 are handled identically.

* [client] Reject metrics ingest bodies whose peer_id tag disagrees with the header

validateTag checked tag names against the per-measurement allowlist but only
bounded the value length, so the peer_id tag was free-form text up to 64 bytes
and could differ from the X-Peer-ID header the request was accepted with.

Tie the two together: the tag value must equal the header value. Since the
header is already checked to be 16 hex characters, this transitively constrains
the tag to the same shape. Every client sends the same value in both places
(metrics.go feeds agentInfo.peerID to both push.SetPeerID and the body tags),
so well-behaved clients are unaffected.

The mismatch is rejected rather than silently overwritten: rewriting the value
would re-serialize caller-controlled text back into line protocol and would
hide misbehaving senders instead of surfacing them. Rejection also matches the
other input validation failures in the same handler, which all return 400.

This narrows the value space of the peer_id tag but does not by itself bound
InfluxDB series cardinality: a sender that puts the same arbitrary 16 hex
characters in both the header and the body still passes. Limiting that needs a
per-source rate limit in front of the service.

* [client] Document what the metrics ingest peer_id check does and does not bound

The README described X-Peer-ID as the correlation tag, but grouping is done by
the peer_id tag in the submitted line protocol: that is what is forwarded to
InfluxDB, while the header only serves as the value each tag is checked against.

It also claimed the format check bounds tag cardinality. It bounds the value
space of the tag, not the number of distinct series, so state that explicitly
and point out that series cardinality has to be limited outside this service.

* [client] Set timeouts on the metrics ingest HTTP server

The server ran on http.ListenAndServe with no timeouts, silenced with a
nolint:gosec for G114. Without ReadHeaderTimeout a client can hold a connection
open by sending headers slowly, and without ReadTimeout or IdleTimeout
connections accumulate on an endpoint that takes unauthenticated requests.

Construct an http.Server with explicit limits instead, which also drops the
nolint. Handler stays nil so the existing DefaultServeMux registrations are
unaffected.

WriteTimeout is deliberately larger than the 10s upstream client timeout: the
response is only written after the forward to InfluxDB completes, so a tighter
value would cut off the server's own valid response.

* Revert "[client] Document what the metrics ingest peer_id check does and does not bound"

This reverts commit 837a5d8dda.

* Revert "[client] Reject metrics ingest bodies whose peer_id tag disagrees with the header"

This reverts commit 91d4f6128.

Tying the body peer_id tag to the X-Peer-ID header assumed the two always agree,
but a profile switch breaks that. UpdateAgentInfo swaps agentInfo.peerID and
calls push.SetPeerID with the new value while leaving the sample buffer alone,
and the peer_id is baked into each buffered line at record time, so samples from
the previous profile ship under the new header.

The consequences compound: validateLineProtocol rejects the whole batch on the
first bad line, so fresh samples are dropped along with the stale ones, and
push.go only resets the buffer after a successful push, so the batch is retried
and fails again. Metrics from that client stay stuck until the old samples age
out of the buffer, up to maxSampleAge (5 days).

Deciding whether the previous profile's unsent samples may be discarded, or
whether the push has to be partitioned per identity, is a product call, so
restore the previous behaviour for now. The header keeps its format check;
the body peer_id tag goes back to being bounded only by maxTagValueLength.

* [client] Stop claiming the metrics ingest peer ID format check bounds tag cardinality

The X-Peer-ID header is never forwarded to InfluxDB; the stored peer_id
tag comes from the request body and is constrained only by the tag
allowlist and the maximum tag value length. Align the README and the
validatePeerIDFormat godoc with the actual behavior after the
header/body match check was reverted.
2026-09-02 12:36:47 +02:00
Zoltan Papp
ecbeba8e67 [management] Fix geolocation panics (#7382)
* [management] Return errors instead of panicking on malformed geolocation inputs

* [management] Reject empty date suffix in geolocation database filename
2026-09-02 12:36:03 +02:00
eYey
2f55965031 [client, android] Type the split tunnelling mode instead of storing a string (#7387)
gomobile carries only basic types, so the typed constants stay unexported and
the exported SplitTunnelMode* ints are what the Java side gets.
2026-09-02 10:00:00 +02:00
Maycon Santos
a1415dbc05 [client] Fix the ICEBind races that wedge interface creation (#7377)
* [client] Add tests for the ICEBind open and close races

Running many embedded clients in one process intermittently wedges interface
creation. A goroutine dump taken from 50 clients shows ten of them parked for
seven minutes in Device.IpcSet, in closeBindLocked waiting on
device.net.stopping.Wait, holding device.net while every other device
goroutine queues behind it on Device.Up.

Open writes s.closed and Close reads it with no synchronisation, and Close
also closes s.closedChan without the mutex that Open swaps it under. Two
Closes can both pass the check and close the same channel, and a Close racing
an Open can mark the bind closed while a live channel and live receive
functions remain, after which every later Close takes its early return and
runs neither close(closedChan) nor StdNetBind.Close. The receive functions
never stop, so stopping.Wait never returns.

These tests do not fix that. The first pins the contract closeBindLocked
depends on and passes today. The other two fail under -race, reporting the
races at the three sites above, and pass again once closed and closedChan are
guarded consistently.

* [client] Release parked receivers so reopening a bind cannot stall

receiveRelayed held closedChanMu for the whole of its blocking select, so a
parked receiver kept the read lock indefinitely and Open could never take the
write lock it needs to install a fresh closedChan. wireguard-go reaches Open
from Device.IpcSet and Device.Up with device.net held, so the stall took the
device lock with it: interface creation never finished, every other device
goroutine queued behind Device.Up, and Engine.Start never returned.

Callers now copy the channel under a short read lock and select on the copy.
Copying alone would stand a new trap in the same place, because an Open that
follows an Open leaves the previous generation parked on a channel no later
Close can reach, so Open now closes the outgoing channel before swapping it.

closed and closedChan are also updated together under that mutex. Read and
written apart, Close could see a stale closed and skip both close(closedChan)
and StdNetBind.Close, leaving every receive function running and wedging
closeBindLocked on device.net.stopping.Wait, or two Close calls could pass the
check together and close the same channel twice.

TestICEBindOpenDoesNotBlockOnParkedReceiver fails without this change, without
needing the race detector. The other three cover the surrounding contract and
report the state races under -race.

* [client] Make the bind lifecycle transition atomic and tighten its tests

Review caught that the previous commit moved the torn transition rather than
removing it. Open published the new generation before calling StdNetBind.Open,
so an Open rejected because the bind was already open had already signalled the
outgoing generation, and a Close arriving in that window could mark the bind
closed while the same call went on to install live sockets. Every later Close
then returned early and never shut them down.

Open now calls StdNetBind.Open first, so a failure leaves the current
generation untouched, and both Open and Close hold the lock across the whole
transition. Ordering is safe: StdNetBind.Open reaches muUDPMux through
createReceiverFn, and no path takes muUDPMux before closedChanMu.

The tests were also weaker than they read. The stress test claimed to cover a
stale channel but only ever raced two Closes, and the concurrency test left
overlap to goroutine start order. Both now gate their goroutines on a common
start, the stress test races an Open against the Closes, and both assert the
surviving generation channel is actually closed. Waiting on receive functions
to be entered replaces part of the sleep in the reopen probe, and teardown
bounds its Close so a regression fails the assertion instead of hanging.

Two of the four now fail without the fix and no race detector, the stress test
by reproducing close of a closed channel at the Close early return.

* [client] Fail the reopen probe when its teardown does not complete

closeBounded swallowed its timeout and the cleanup discarded what
receiversStopped returned, so the bounds added in the previous commit only
stopped teardown hanging. A wedged Close or a parked receiver would have left
the test green with a leaked goroutine, which is the failure this test exists
to catch.

closeBounded now reports whether Close returned, and cleanup fails the test on
either bound.
2026-09-02 00:27:05 +02:00
Pascal Fischer
e3d6c3d0eb [management] fix private services calc on new db path (#7383) 2026-09-01 20:20:30 +02:00
Theodor Midtlien
e5c0cdf958 [client] Stay connected with login command (#7384) 2026-09-01 18:04:11 +02:00
Maycon Santos
ebc259e30b [management,client] Gate remote jobs behind an admin opt-in with MDM support (#7153)
This introduces a disabled-by-default allow-remote-jobs setting that
controls whether the management server may run jobs (such as debug
bundles) on a peer. The flag propagates end to end: through client
configuration, the daemon SetConfig and Login requests, authentication,
and system info, up to management, where it is stored on the peer and
exposed on the peers API as remote_jobs_allowed. The client refuses any
management-requested job unless the peer has opted in. Because enabling
remote jobs crosses the user-to-root boundary, turning it on requires
privilege, mirroring the SSH-server gate. Administrators can enforce the
setting through MDM policy on both macOS and Windows, and MDM can also
override the debug-bundle upload URL. The change ships policy
documentation and generated profile templates, and adds configuration,
conflict, and enforcement tests covering the opt-in, privilege, and MDM
paths.
2026-09-01 17:53:41 +02:00
Maycon Santos
3027130f0f [management] Add Agent Network access roles and self-service endpoints (#7221)
Delegating Agent Network today means handing out full account admin, and
regular users cannot see their own usage or how to connect a local tool.

Add two roles on top of the existing agent_network permission
submodules. agent_network_admin owns the whole area (providers,
policies, guardrails, budgets, usage, logs, settings) with read-only
users, groups, peers, and account info needed to build policies, and
nothing else in the account. usage_viewer is the regular User baseline
plus read on the aggregated usage and cost overview: no provider
configuration, no policies, no request-level logs, which can contain
captured prompts. billing_admin gets a proper permission-map entry with
the User baseline so role resolution stops failing with role-not-found;
its plan and invoice permissions stay enforced cloud-side.

Add the self-service endpoints behind the "My Agent Network" view,
available to every authenticated user because both answers are scoped
strictly to the caller. GET /api/agent-network/me/setup returns the
account endpoint plus the providers and models the caller's own groups
authorize, computed with the same rules the proxy enforces: policy
filtering as in policy selection, model allowlist union intersected
with declared models, orphan and disabled providers omitted. Not set up
and no access are deliberately indistinguishable, and the response
carries display metadata only. GET /api/agent-network/me/consumption
returns the caller's own user-dimension counters.
2026-09-01 14:35:38 +02:00
evgeniyChepelev
652d5f3c15 [client] Reuse the profile's account for iOS SSO logins (#7193)
* [client] Reuse the profile's account for iOS SSO logins

Android reads the profile's stored account and passes it as the OIDC
login_hint, and records it again after a successful login. iOS did neither: it
called GetOAuthFlow with an empty hint, so a re-login was resolved by whatever
session the browser's cookie jar held rather than by the account the profile
belongs to. With a non-ephemeral browser session that is the wrong account as
soon as more than one is signed in.

Mirror client/android/login.go: hint from mobile.ReadProfileEmail before the
flow, mobile.WriteProfileEmail after Login succeeds. Storing after Login and
not before keeps a rejected token from leaving a hint that points at an
account which cannot be used.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* [client] Persist the account email on tvOS and on the device flow

Two paths left a profile with no account bound, so every later login went out
without a login_hint — the case this change exists to remove.

WriteProfileEmail went through util.WriteJsonWithRestrictedPermission, which
writes a temp file and renames it over the target. The tvOS App Group sandbox
blocks exactly that, which is why the config sitting next to this file is
written with DirectWriteOutConfig. On tvOS the email write therefore failed and
was dropped with a warning. Use DirectWriteJson: the file is rewritten whole
from a single key, so the only thing atomicity buys here is surviving a crash
mid-write, and a torn file reads back as "no email" and is replaced by the next
login.

The device authorization flow never populated TokenInfo.Email, unlike the PKCE
flow, so a client driven through it — Android TV and tvOS — bound no account at
all. Parse the ID token there too.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* [client] Report a failed close from DirectWriteJson

The deferred close assigned its error to err, but the return value was not
named, so the assignment went nowhere: a close that failed was logged and the
function still returned nil. The write is only durable once the file closes
cleanly, so every caller — the management config, the profile configs and the
profile account email — could be told the data landed when it had not.

Name the return so the assignment does what its shape always intended, and
report the failure once. When the body succeeded the close error is returned and
the caller logs it. When the body already failed, that error is the one that
explains the failure and is what the caller gets, which leaves the deferred log
as the only place the close failure can surface — at debug, per the logging
rules for close errors on writes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-01 13:29:26 +02:00
Daneyon Hansen
7a9582db16 [management,proxy] Add agentgateway integration (#7274)
* [management] Add agentgateway provider catalog entry

Allow Agent Network providers to target an operator-supplied agentgateway proxy while stamping trusted NetBird identity headers.

Signed-off-by: Daneyon Hansen <daneyon.hansen@solo.io>

* [proxy] Allow trusted Agent Network identity headers

Permit only the built-in identity injector to replace the two reserved agentgateway attribution headers while keeping them blocked for every other middleware.

Signed-off-by: Daneyon Hansen <daneyon.hansen@solo.io>

* [management,proxy] Add multi-vendor gateway routing

Let one Agent Network route declare multiple parser surfaces while preserving the existing singular vendor wire field.

Signed-off-by: Daneyon Hansen <daneyon.hansen@solo.io>

* [management] Update router test for model policies

Signed-off-by: Daneyon Hansen <daneyon.hansen@solo.io>

* [proxy] Cover reserved header policy

Signed-off-by: Daneyon Hansen <daneyon.hansen@solo.io>

* [management] Add agentgateway model discovery

Use agentgateway's OpenAI-compatible models endpoint and omit wildcard patterns until NetBird can authorize and price them consistently.

Signed-off-by: Daneyon Hansen <daneyon.hansen@solo.io>

---------

Signed-off-by: Daneyon Hansen <daneyon.hansen@solo.io>
2026-09-01 13:03:16 +02:00
Zoltan Papp
4749005a50 [client] Resolve profiles for the sudo invoking user instead of root (#7238)
* [client] Resolve profiles for the sudo invoking user instead of root

The SSH server flags force `netbird up` through sudo, but the CLI resolved
every per-user path with the process user. As root that reads root's own
(empty) local state, so a `sudo netbird up` silently switched the daemon from
the user's profile to the default one — cancelling any login already waiting
in the browser — and then ran an SSO login for the default profile's config.
Whichever account that login returned, the default profile's peer belongs to
someone else, so every attempt ended in "peer is already registered by a
different User or a Setup Key", with nothing telling the user why.

Resolve the acting user through SUDO_USER when running as root: the active
profile, the profile config paths and the stored account email now come from
the invoking user's directories. Privilege decisions are untouched — they stay
on the kernel credentials of the daemon connection, which an environment
variable can never influence; a forged SUDO_USER only selects a profile root
could select anyway.

The invoking user's directories are strictly read-only under sudo. Anything
root wrote there would be root-owned and break the user's own runs, so instead
of chowning files back, the local writes are skipped: the active-profile
bookkeeping and the account-email state simply do not update from a sudo run
(the daemon records the switch on its side; a skipped email write costs at
most one extra account prompt later).

Plain root — no sudo context — has no user to act for, so the ambiguity is
refused instead of guessed at: when the daemon's active profile differs from
what root resolves and no --profile was given, up fails with a message naming
both profiles, instead of silently switching the daemon and failing later with
the ownership error.

* [client] Act on the daemon-resolved profile and fail closed in the root guard

Under sudo the local active-profile mirror is not updated, so up/login
re-reading it after a profile switch acted on the previous profile; use
the daemon-resolved ID directly instead. The plain-root guard now runs
after the readiness wait, denies on lookup errors and empty responses,
and matches the owning username as well; an unowned profile (fresh
install) and a daemon predating the RPC stay allowed. Write-skip
decisions key off the sudo environment alone so a transient user lookup
failure cannot turn a run into writing root-owned files into the user's
directory, and RemoveProfileState honors the read-only rule too.

* [client] Return a wrapped error instead of double-reporting the dial failure

* [client] Read the profile from the daemon when the local mirror is not authoritative

Under sudo without --profile, `up` took the active profile from the invoking
user's local active_profile.txt mirror and drove the daemon to it. But that
mirror is never written under sudo (the SwitchProfile write is a no-op), so it
goes stale after any --profile run and silently switches the daemon back to the
mirror's default. The plain-root guard was meant to refuse exactly this
ambiguity but only ran for plain root, never for the sudo case the fix targets.

When there is no --profile and the mirror is not authoritative (sudo or plain
root), take the profile the daemon already holds for the invoking user instead
of the stale mirror: stay on the user's current profile when the daemon owns it
(or it is unowned, as on a fresh install), and refuse with a --profile hint when
the daemon is on another user's profile. A daemon predating the RPC keeps the
mirror-derived profile.

Reproduce (before this change):
1. As a non-root user misha, with the daemon installed and running:
     sudo netbird up --profile work
   misha connects on the `work` profile.
2. Because the local mirror write is skipped under sudo,
   ~misha/.config/netbird/active_profile.txt still says `default` (or is still
   absent, which also resolves to `default`).
3. Run a bare:
     sudo netbird up
   The CLI reads `default` from the frozen mirror and sends ProfileName=default;
   the daemon silently switches away from `work` and brings the tunnel up on
   `default` — a different account/peer than the one last chosen, with no
   warning. After this change step 3 stays on `work`.

* [client] Return a sentinel error instead of nil-nil for the missing daemon RPC

* [client] Load the extend-session hint from the resolved profile

* [client] Fail closed instead of reading root's config when the sudo user lookup fails

* [client] Fail closed in InvokingUser when the sudo user lookup fails

A previous change made baseConfigDir fail closed when SUDO_USER cannot be
resolved, but InvokingUser still fell through to user.Current(). Those two
guards disagreed: the active-profile mirror and the email state refused to
read root's directory, while every profile-path caller happily resolved as
root.

The consequence of a transient NSS failure under sudo was that
Profile.FilePath resolved through getConfigDirForUser("root"), creating
/var/lib/netbird/root and reading the profile JSON from there, and the CLI
sent Username "root" to the daemon in SetConfig and ListProfiles, so the
daemon resolved the same phantom namespace. The invoking user was silently
moved onto a root-owned profile instead of being told the lookup failed.

Fail closed at the single source of the fallback. getConfigDirForUser is
left alone on purpose: it is a pure path helper that also serves
daemon-supplied usernames, and under sudo with a successful lookup it must
still create the invoking user's own profile directory.
2026-09-01 12:50:03 +02:00
Riccardo Manfrin
352a1d348a [client] Do not log the WireGuard key on a parse failure (#7379)
The error already says what went wrong: an invalid base64 payload
reports the offending byte offset, and a wrong key size reports the
length. Passing the key itself adds nothing an operator can act on, and
the line is emitted at Error level, so it reaches every log sink and
every debug bundle.
2026-09-01 12:41:12 +02:00
Anton Groshev
922be0b8c2 Fix docs link (#7352) 2026-09-01 12:38:06 +02:00
Riccardo Manfrin
c170905bc9 [client] Allow logging out of the active profile when profiles are disabled (#7360)
* [client] Allow logging out of the active profile when profiles are disabled

A profile-addressed logout was refused outright when the profiles feature is
disabled: handleProfileLogout ran validateProfileOperation, which returned
Unavailable ("profiles are disabled, you cannot use this feature without
profiles enabled") before looking at which profile was targeted.

The desktop UI always addresses logout by profile — both the profile menu and
the session-expiration dialog send the active profile's ID — so a client with
profiles disabled could not log out at all; only a plain `netbird logout`,
which takes the profile-less path, still worked. Logging out of the profile the
daemon is already running is a deregistration, not profile management, and with
profiles disabled there is a single profile anyway, so every profile-addressed
logout is by definition an active-profile logout.

Replace validateProfileOperation with validateProfileLogout, which skips the
profiles-disabled check when the target is the active profile and keeps gating
logout of any other profile. This mirrors switchProfileIfNeeded, which already
gates only the branch that actually manages profiles. The dropped
allowActiveProfile parameter was always true, leaving canRemoveProfile
unreachable, so both are removed.

* [client] Compare the username and propagate state errors on profile logout

Review follow-ups on the logout gate:

Propagate the GetActiveProfileState failure instead of discarding it. A failed
lookup made the target look non-active, so a caller with profiles disabled got
"profiles are disabled" in place of the real error.

Compare the username along with the ID when deciding whether the target is the
active profile, matching switchProfileIfNeeded. Legacy profile IDs are display
names, so two users can hold the same ID in their own profile directories, and
an ID-only match let one user's logout pass the gate against the other user's
active profile. The default profile is shared and carries no username, so it
keeps matching on the ID alone.

Re-read the active profile before the connection teardown rather than reusing
the pre-flight snapshot. Login switches profiles under guardedConfigMu, which
the logout path does not hold, so a login that landed while the deregistration
was in flight would otherwise lose its fresh connection to a stale flag.

* [client] Address review on the profile logout gate

Pass the username down to logoutFromProfile and reuse the running config only
when the target is the active profile for that username. On an ID-only match a
legacy profile ID shared between two users made the connected-client path
deregister the active peer while its connection stayed up, which the gate fix
alone did not cover.

Split the setup-key-less branch of Login into beginSSOLogin, with the
reuse-the-pending-flow decision in pendingOAuthFlowResponse. Login's cognitive
complexity drops from 37 to 21 (gocognit), clearing the SonarQube report on
this file with no behaviour change.

Point the test fixture at an https URL, since the profiles a gated logout must
not touch only need to be unreachable, not plaintext.
2026-09-01 12:19:16 +02:00
Maycon Santos
1081ca006d [management,client] Add anonymize level and upload URL to remote debug bundle jobs (#7147)
This extends the management-requested remote debug-bundle job with two
new, optional parameters. anonymize_level selects how aggressively the
bundle is scrubbed: "default" keeps internal (private) IP ranges
readable, while "strict" also anonymizes private, CGNAT and link-local
addresses; the value is trimmed and lowercased, and an unknown level is
rejected at creation. upload_url lets an operator point the peer at a
specific upload service instead of the default one; it must be a
well-formed https URL with a host, and an empty value falls back to the
default upload server. Both fields flow through the job workload API and
are surfaced in the create-debug-job modal on the dashboard. Validation
is shared so the client executor and the management boundary agree on
what a valid upload URL is, preventing drift between the two checks.
2026-09-01 11:45:20 +02:00
Maxim Egorov
930a25319d [client] Keep the route selection on an invalid request and apply it on a partial one (#7292)
* [client] Keep the route selection when every requested ID is unavailable

A non-append SelectRoutes() wipes the current selection before applying
the requested one, but it validated the requested IDs only afterwards,
while already mutating. A request naming no available route at all left
every route deselected and returned an error - so a typo in a route ID
silently dropped the user's exit node, and the routes stayed applied
while the selector claimed nothing was selected.

Validate first and bail out before touching any state when nothing in
the request is available. A request with at least one available route
keeps applying the valid part and reporting the rest, and an empty
request still deselects everything, since that is the caller asking for
exactly that rather than a failed lookup.

* [client] Trim the new comments to the contributing guide's length budget

CONTRIBUTING.md caps comments at 90 characters per line and roughly 250
per comment. The three comments added by this PR were over both limits.
The test comments also restated their own test names, so they lose that
half and keep only the why.

* [client] Apply the route selection even when some IDs are unknown

SelectRoutes and DeselectRoutes returned the error before TriggerSelection,
so a request mixing valid and unknown network IDs changed the selector but
never reached the routing table. The valid routes read as selected while
`ip route` showed nothing.

Trigger the selection first and return the error afterwards. The inner
selectRoutes already applied the valid part of a partial request, only the
outer layer dropped it.

* [client] Publish the network selection event on a partial failure

Returning early on error was correct while an error meant nothing had
happened. A partial failure now changes the selection and the routing
table, so returning first left the change with no trace in the event log
or the UI, even though the new state had already been broadcast.

* [client] Cover the append and deselect-all paths of the selection guard

The append path was never destructive and behaves the same with or without
the early return, so that case is characterization rather than a regression
test. The deselect-all case is a real guard: the early return also skips
resetting deselectAll, so a typo no longer drops the "nothing selected,
including future networks" policy.

* [client] Pin that a fully invalid selection disturbs nothing

The selection is now applied on every request, including one where no ID is
known and the selector is left untouched. Nothing may be torn down or
reinstalled on that path.

* Revert "[client] Publish the network selection event on a partial failure"

This reverts commit 26219592.

The event would lie on the opposite path: when no requested ID is available
the selector is left untouched, so an unconditional publish reports a change
that never happened. Telling that case from a partial failure needs the
manager to report whether anything was applied, which is a new signal in its
API and does not belong in a PR about the selector guard. Follow-up instead.

---------

Co-authored-by: Riccardo Manfrin <3090891+riccardomanfrin@users.noreply.github.com>
2026-08-31 18:47:37 +02:00
Theodor Midtlien
12e8874517 [client, relay, management] Bump go version to 1.26 and go-quic to v0.62.0 (#7359)
* Bump go version to 1.26 and go-quic to v0.62.0
* Replace deprecated ecdsa public key assembly and add tests for jwt
* Update goversioninfo
* Pin go toolchain to 1.26.7
2026-08-31 18:01:14 +02:00
Laotree
24959e1ed9 [client] Drop agentConnecting whenever ICE session state clears (#7327)
* [client] Drop agentConnecting whenever ICE session state clears

Closing a WorkerICE raced a blocked dial goroutine: Close released the
agent while connect() was still inside Dial, and the goroutine's own
cleanup skipped its flag reset because w.agent no longer matched. With
agentConnecting stuck on true, evalConnStatus read the peer as
connected, the reconnection guard stopped sending offers and
same-session offers were dropped, so the peer could not recover without
a restart. An aborted recreate in OnNewOffer reaches the same wedged
state without any race.

Route every teardown path through one abandonNegotiation helper so the
agent and flag fields always clear together; Close now also cleans up
residual state left by an aborted recreate.

* [client] Drive the ICE teardown race test through the real dial goroutine

The regression test simulated the stale goroutine by calling closeAgent
directly, so it pinned the symptom rather than the mechanism. Rework it
to start a real negotiation, tear it down mid-flight and let the actual
goroutine run its own cleanup: with no remote responder the dial can
only fail once Close cancels it, so the interleaving stays deterministic
without sleeps or injection points.

Assert the full idle state that abandonNegotiation owns (agent nil,
connecting false, remote session ID empty) instead of only InProgress,
and make the stale-cleanup ownership test verify that the newer session
survives field by field.

* [client] Assert live remote session ID after stale ICE cleanup

The stale-cleanup test compared a snapshot captured before closeAgent
ran, so clearing the field during cleanup would have gone unnoticed.
Read the field under the mutex after the cleanup instead.

* [client] Give the ICE race tests a no-op signal client

The candidate callback fires from a real gather and dereferences the
signaler, so a nil one crashes the test package intermittently when
gather wins the race against Close. Build the worker with a stub
signal.Client instead.

* [client] Read the ICE dial cancel func from an argument in connect

The error paths read w.agentDialerCancel without holding muxAgent while
OnNewOffer rewrites the field for a newer negotiation, a data race the
new teardown test trips under -race. Reading a stale value also let an
old goroutine cancel another session's dial. Capture the cancel func at
goroutine spawn, like the dial context already is.

* [client] Guard the ICE dial success path against stale negotiations

The stale-cleanup guard in closeAgent only protected teardown. Its
success-path counterpart was missing: an older negotiation could complete
agentDial after a newer one replaced w.agent, then clear the newer
session's agentConnecting, record lastSuccess and publish its dead
connection via onICEConnectionIsReady.

Verify ownership under muxAgent twice: right after the dial returns, so a
stale goroutine drops its connection before touching a closed agent, and
again at the state-commit point, atomic with the agentConnecting and
lastSuccess writes, so a replacement arriving in the meantime cannot get
its state clobbered. Both paths close the stale connection and return
without modifying worker state. A regression test holds session A's dial
open until session B is installed, then releases it; the stale connection
must be discarded and B's agent, connecting flag and remote session ID
must survive.

* [client] Fix ICE teardown test leak and document the stale delivery window

A code review of the stale-negotiation guard found a leftover resource
leak in TestWorkerICE_StaleCloseAgentKeepsCurrentSession: session B is
never closed, so its ICE sockets and blocked dial goroutine live as long
as the test process. Register t.Cleanup(w.Close).

The delivery race flagged after the success-path guard is pre-existing
and self-correcting - the newer negotiation overwrites the transient
endpoint - so document it in the existing todo instead of locking the
callback, which would invert lock order against Conn.Close. Adjust the
teardown test comment to match the now-synchronous Close flag clearing.
2026-08-31 17:49:59 +02:00
Max
7ffbcb0016 [client] Add Ukrainian localization for desktop client (#7035) 2026-08-31 16:28:30 +02:00
Viktor Liu
bad63c14b8 Say whether an approval was refused, unanswered, or never shown 2026-08-31 15:15:37 +02:00
Viktor Liu
3af7764b09 Give the user a minute to answer an approval prompt 2026-08-31 15:10:13 +02:00
Viktor Liu
932c87ea41 Stop retrying an Accept that will not recover 2026-08-31 15:10:13 +02:00
Zoltan Papp
086d8ba507 [client] Close the session-expiration dialog only on renewal (#7337)
* [client] Close the session-expiration dialog only on an actual session renewal

The dialog auto-closed on any Connected status snapshot, but the daemon
emits Connected periodically regardless of session state, so the warning
popup disappeared on the next snapshot (~30s) with no chance to
re-authenticate. Close only when the snapshot's session deadline jumps
past the one the dialog was opened for, meaning the session was renewed
from another surface (tray action, CLI, main window).

* [client] Compare session renewals against the exact deadline in the expiration dialog

The dialog reconstructed its reference deadline from the relative seconds
URL parameter, which carries up to a second of truncation and mount
latency, forcing a renewal-detection margin wide enough to miss a renewal
made shortly after the previous login. Pass the absolute deadline (unix
ms) from both tray call sites - the extend flow's cached deadline and the
final warning's event metadata - so any forward jump in the snapshot
deadline closes the dialog; the seconds-derived fallback with a small
tolerance remains for an unknown deadline.

* [client] Derive the expiration dialog countdown from the deadline

The per-second decrement assumed the interval fires once a second, but
the webview's timers get suspended for tens of seconds under App Nap /
hidden-window throttling, leaving the displayed countdown behind the
wall clock by the suspended time. Recompute the remaining time from the
absolute deadline on every tick so the first tick after a suspension
shows the correct value.

* [client] Tolerate the warning deadline's second precision in the renewal check

The final-warning metadata formats the deadline as RFC3339 truncated to
whole seconds while the status snapshot keeps millisecond precision, so
an unchanged deadline could appear up to 999 ms newer than the exact URL
value and close the dialog on the first snapshot. Allow a sub-second
tolerance on the exact path; any real renewal jumps by at least seconds.
2026-08-31 10:32:36 +02:00
eYey
945b0b6be2 Store the Android split tunnelling settings per profile (#7349)
Which applications the tunnel carries belongs with the rest of a
profile's preferences rather than on the Android side, so the choice
follows the profile the user is on.

Adds a split tunnel store beside the SSH session store, over a
"split-tunnel" namespace holding the mode and both selections. The two
selections are kept apart because the platform applies an allow list or
a deny list and never both, and so that switching mode does not throw
away the picks made in the other one.

Only the store itself needs the android build tag; the rest stays
untagged so it is covered by the package's host tests.
2026-08-31 10:29:05 +02:00
Viktor Liu
3ce095719a Describe the approval no-match result and the cursor-skip key as they behave
Claude-Session: https://claude.ai/code/session_01QKDYfH4WKLbpNQHccpVo3P
2026-08-30 05:56:59 +02:00
Viktor Liu
12040b1be0 Refuse an ambiguous X display and settle the approval timeout race under one claim
Claude-Session: https://claude.ai/code/session_01QKDYfH4WKLbpNQHccpVo3P
2026-08-29 17:25:02 +02:00
Viktor Liu
fecd7cfff5 Attach to the X server on the active VT and keep a retryable DXGI frame from tearing down the capturer
Claude-Session: https://claude.ai/code/session_01QKDYfH4WKLbpNQHccpVo3P
2026-08-29 16:59:05 +02:00
Viktor Liu
15d6e3bea9 Keep DXGI on an idle desktop, honour the FreeBSD pitch when swizzling, close the handler-drain race 2026-08-29 16:25:07 +02:00
Viktor Liu
5f739ef0ac Reject an unsupported VNC session id and say why a cursor rect was skipped 2026-08-29 16:09:47 +02:00
Viktor Liu
d0d8813dcb Fail an approval request fast when no subscriber received the prompt 2026-08-29 16:00:36 +02:00
Viktor Liu
e104cef490 Require a real first DXGI frame and read the FreeBSD framebuffer at its true pitch 2026-08-29 15:51:19 +02:00
Viktor Liu
9ec27f9d7f Authenticate the daemon and its VNC agent to each other without sending the token 2026-08-29 15:45:09 +02:00
Viktor Liu
ba104a11ad Encode one framebuffer update at a single negotiated pixel format 2026-08-29 13:56:54 +02:00
Viktor Liu
8d6b7b6175 Leave the VNC approver nil when there is no broker to ask 2026-08-29 13:49:47 +02:00
Viktor Liu
9672e04ce6 Scope crash recovery to Linux, drain connection handlers before teardown 2026-08-29 13:43:12 +02:00
Viktor Liu
d8236002c7 Honour the negotiated pixel format for the cursor, enqueue key edges reliably, drop racy test writes 2026-08-29 12:40:14 +02:00
Viktor Liu
d196b23de6 Stop narrowing the shared runtime dir, close the injector on stop, allowlist VNC metrics 2026-08-29 10:29:17 +02:00
Viktor Liu
827098c3d2 Compare cursor serials by identity so returning to an earlier cursor updates 2026-08-29 10:25:09 +02:00
Viktor Liu
1ecd1dab8b Collect auth requirements for bidirectional source peers and fix follow-up review findings 2026-08-29 09:57:07 +02:00
Viktor Liu
312b73f771 Persist virtual session processes for crash recovery and identify them by start time 2026-08-29 09:48:45 +02:00
Viktor Liu
8b2db16c84 Route resource endpoints through the shared policy peer filter 2026-08-29 09:37:17 +02:00
Viktor Liu
fb9c0ef602 Make session key authorization atomic, unblock the encoder on teardown, drop agent privileges unconditionally 2026-08-29 09:12:33 +02:00
Viktor Liu
95e86deeb8 Resolve VNC authorized users on the components path and fix uinput, X11 and macOS input gaps 2026-08-29 09:04:02 +02:00
Viktor Liu
4a0fe09ced Merge branch 'main' into embedded-vnc 2026-08-29 08:42:02 +02:00
Viktor Liu
432d249cde Accept bare marker protocols, reject msb_right framebuffers, split the policy row conversion 2026-08-29 08:24:36 +02:00
Bethuel Mmbaga
11733fd718 [infrastructure] Improve domain, Docker Compose, and license validation in self-hosted scripts (#7339) 2026-08-28 18:11:57 +03:00