Commit Graph

3429 Commits

Author SHA1 Message Date
Viktor Liu
3ce095719a Describe the approval no-match result and the cursor-skip key as they behave
Claude-Session: https://claude.ai/code/session_01QKDYfH4WKLbpNQHccpVo3P
2026-08-30 05:56:59 +02:00
Viktor Liu
12040b1be0 Refuse an ambiguous X display and settle the approval timeout race under one claim
Claude-Session: https://claude.ai/code/session_01QKDYfH4WKLbpNQHccpVo3P
2026-08-29 17:25:02 +02:00
Viktor Liu
fecd7cfff5 Attach to the X server on the active VT and keep a retryable DXGI frame from tearing down the capturer
Claude-Session: https://claude.ai/code/session_01QKDYfH4WKLbpNQHccpVo3P
2026-08-29 16:59:05 +02:00
Viktor Liu
15d6e3bea9 Keep DXGI on an idle desktop, honour the FreeBSD pitch when swizzling, close the handler-drain race 2026-08-29 16:25:07 +02:00
Viktor Liu
5f739ef0ac Reject an unsupported VNC session id and say why a cursor rect was skipped 2026-08-29 16:09:47 +02:00
Viktor Liu
d0d8813dcb Fail an approval request fast when no subscriber received the prompt 2026-08-29 16:00:36 +02:00
Viktor Liu
e104cef490 Require a real first DXGI frame and read the FreeBSD framebuffer at its true pitch 2026-08-29 15:51:19 +02:00
Viktor Liu
9ec27f9d7f Authenticate the daemon and its VNC agent to each other without sending the token 2026-08-29 15:45:09 +02:00
Viktor Liu
ba104a11ad Encode one framebuffer update at a single negotiated pixel format 2026-08-29 13:56:54 +02:00
Viktor Liu
8d6b7b6175 Leave the VNC approver nil when there is no broker to ask 2026-08-29 13:49:47 +02:00
Viktor Liu
9672e04ce6 Scope crash recovery to Linux, drain connection handlers before teardown 2026-08-29 13:43:12 +02:00
Viktor Liu
d8236002c7 Honour the negotiated pixel format for the cursor, enqueue key edges reliably, drop racy test writes 2026-08-29 12:40:14 +02:00
Viktor Liu
d196b23de6 Stop narrowing the shared runtime dir, close the injector on stop, allowlist VNC metrics 2026-08-29 10:29:17 +02:00
Viktor Liu
827098c3d2 Compare cursor serials by identity so returning to an earlier cursor updates 2026-08-29 10:25:09 +02:00
Viktor Liu
1ecd1dab8b Collect auth requirements for bidirectional source peers and fix follow-up review findings 2026-08-29 09:57:07 +02:00
Viktor Liu
312b73f771 Persist virtual session processes for crash recovery and identify them by start time 2026-08-29 09:48:45 +02:00
Viktor Liu
8b2db16c84 Route resource endpoints through the shared policy peer filter 2026-08-29 09:37:17 +02:00
Viktor Liu
fb9c0ef602 Make session key authorization atomic, unblock the encoder on teardown, drop agent privileges unconditionally 2026-08-29 09:12:33 +02:00
Viktor Liu
95e86deeb8 Resolve VNC authorized users on the components path and fix uinput, X11 and macOS input gaps 2026-08-29 09:04:02 +02:00
Viktor Liu
4a0fe09ced Merge branch 'main' into embedded-vnc 2026-08-29 08:42:02 +02:00
Viktor Liu
432d249cde Accept bare marker protocols, reject msb_right framebuffers, split the policy row conversion 2026-08-29 08:24:36 +02:00
Bethuel Mmbaga
11733fd718 [infrastructure] Improve domain, Docker Compose, and license validation in self-hosted scripts (#7339) 2026-08-28 18:11:57 +03:00
Pascal Fischer
353251d886 [management] fix posture check evaluation for direct peers in policy definition (#7348) 2026-08-28 16:46:42 +02:00
Pascal Fischer
611a9291cd [management] fix posture check flip evaluation for affected peers calc (#7347) 2026-08-28 15:39:48 +02:00
Zoltan Papp
89c6e84a41 [client, ios] Fix context cancellation during restart (#7329)
* fix(mobile): stop the client synchronously so a restart cannot inherit a cancelled context

Original finding
----------------
A user reported that leaving home and switching from wifi to cellular killed
all Internet traffic until NetBird was turned off. A debug bundle captured the
failure (iOS, CLI 0.75.0, self-hosted management, generated 2026-08-18 01:17;
the incident is at 2026-08-17 22:37:38-51 UTC).

The bundle shows the whole sequence:

  22:37:38.255  management sync stream drops (keepalive ACK timeout)
  22:37:43.670  Swift: "Network type changed: wifi -> cellular" -> schedules a
                restart with a 1s debounce
  22:37:44.737  Go: "ensuring wg interface is removed, Netbird engine context
                cancelled" - engineCtx dies, every peer gets context canceled
  22:37:49.910  iface.go:238 "failed to remove WireGuard interface utun6:
                timeout when waiting for interface utun6 to be removed"
                -> the teardown stretches out for ~5s
  22:37:50.710  Swift: "restartClient: starting client", needsLogin=false
                (so this is NOT a login expiry)
  22:37:51.013  Go: connect.go:476 "exiting client retry loop due to
                unrecoverable error: context canceled" - the OLD run dies here
  22:37:51.333  Go: grpc.go:135 "failed creating connection to Management
                Service: context canceled" - the NEW start, 2ms after the old
                run finally exited
  22:37:51.334  Swift: "restartClient: start failed" -> widget disconnected
  then nothing for 15 minutes

The tunnel stayed installed with no engine behind it, so every packet was
black-holed. status.txt, generated ~14 hours later, still reads Management:
Disconnected / Signal: Disconnected / Peers count: 0/0 - the client never
recovered on its own.

Root cause
----------
Client.Stop() cancelled a shared ctxCancel field and returned immediately,
without waiting for the run loop to exit. The Swift stop{} completion handler
therefore fired while the Go teardown was still running (stretched out by the
utun6 removal timeout), and the start that followed landed on a context that
the outgoing run was about to cancel.

Two further paths wrote the same shared field. IsLoginRequired() and
LoginForMobile() each overwrote c.ctxCancel, so any call to them during a live
session discarded the running engine's cancel function. restartClient() calls
needsLoginCached() on exactly this path.

Changes
-------
- Stop() now drives the stored ConnectClient: ConnectClient.Stop() cancels the
  run context and blocks on runExited, so the caller's completion handler only
  fires once the run loop has really finished. The ctxCancel path stays as a
  fallback for when no ConnectClient exists yet (e.g. during LoginForMobile).
- Run() owns its cancel in a local variable, so a concurrent call that
  overwrites the shared field can no longer cancel this run's context through
  the deferred cleanup.
- IsLoginRequired() and LoginForMobile() use local cancels and leave the shared
  field alone. LoginForMobile's cancel moves into the deferred cleanup of the
  goroutine that outlives the call, so the OAuth token wait is not cut short.
- The Android SDK gets the same treatment. The structural defect is identical
  there, but the trigger is absent: Android has no automatic engine restart on
  a network type change, and no interface-removal timeout to stretch the
  teardown. This part is preventive, not a fix for an observed failure.

* fix(mobile): do not let a superseded startup publish its client

Review found a window the previous commit left open. Run stored its cancel
function and only published the ConnectClient later, after loading config and
constructing the client. A Stop landing inside that window found no
ConnectClient, cancelled the run and returned immediately. A new Run could then
publish its own client, and the cancelled older run — still executing — would
overwrite it with a client that was already being torn down. The next Stop
stopped that stale client and left the live one running with nothing tracking
it.

Runs now carry a generation. Run claims one before doing any work and publishes
its client only while the generation is still current; a superseded run returns
without touching the shared state. Stop bumps the generation, so any startup
still in flight is invalidated, then cancels it and waits for the run to exit
before returning (20s cap so a wedged teardown cannot block the caller
forever).

setState is gone: publishState replaces it at both call sites on each platform.

* fix(ios): add a non-waiting Stop for callers on a deadline

Stop now waits for the run loop to exit, which is what a restart needs but
wrong for stopTunnel: iOS gives NEPacketTunnelProvider only a few seconds
there before it kills the extension, and the wait can run to its 20s cap.
Waiting past the deadline earns a SIGKILL, so the next start inherits a dirty
state instead of the orderly shutdown the wait was meant to buy.

StopWithoutWait tears the client down and returns. ConnectClient.Stop blocks on
runExited with no cap of its own, so the non-waiting path runs it detached
rather than only skipping the runDone wait.

Android keeps a single blocking Stop: it has no equivalent deadline.

* fix(mobile): guard the run lifecycle with a single lock

Stop and beginRun each touched the same lifecycle state across two locks in
sequence: take stateMu, release it, then take ctxCancelLock. A run starting in
that gap installed its own cancel before Stop reached it, so Stop cancelled the
fresh run and left its own target running — the same class of defect this branch
exists to fix, this time in the locking rather than the state.

ctxCancel moves into the stateMu group, and both sides take their snapshot in
one critical section. ctxCancelLock then guarded nothing and is gone.

* fix(mobile): drop the run-generation machinery for a serialized lifecycle

The platform callers (Swift/Kotlin) always stop before starting and coalesce
restarts, so the generation counter guarded against call patterns that cannot
occur. Replace it with a single-run contract:

- startRun refuses a second Run while the previous one has not exited
- finishRun clears the published state on every exit path, including errors
- Stop cancels and waits for the run loop with a bounded timeout; it no
  longer calls ConnectClient.Stop, whose wait is unbounded
- concurrent Stops wait on the same exit channel instead of returning early
- a superseded startup no longer reports a clean nil exit

* revert(android): drop the run lifecycle changes

Android does not have the defect this PR fixes. On ux/ios-style-redesign the
EngineRestarter is gone: network changes are handled as events instead of an
engine restart, so nothing stops the client and starts it again.

The remaining stop() callers are all final teardowns on the main thread with a
framework deadline - the stop-engine broadcast receiver, onDestroy, onRevoke and
the binder's stopEngine. A Stop that waits for the run loop would risk an ANR
there for a race that cannot occur, so the fix stays iOS-only.

* fix(ios): make loginComplete race-free

The OAuth goroutine spawned by LoginForMobile sets loginComplete after the
call has returned to Swift, while the Swift side polls IsLoginComplete and
later calls ClearLoginComplete from its own thread. The plain bool made all
three unsynchronized: the store may never become visible to the poller, and
a Clear racing the store can be lost, leaving a stale true that makes the
next login look already complete.

Switch the field to atomic.Bool. It is a standalone flag rather than part of
the run lifecycle that stateMu guards, and it has to stay readable while the
login goroutine is still in flight.
2026-08-28 09:33:08 +02:00
Viktor Liu
f5d821ccb2 Describe session auth as SSH and VNC, cover the conflicting key path, fix de and ru wording 2026-08-27 21:08:20 +02:00
Viktor Liu
913a1c6033 Fix framebuffer layout handling, capturer lifecycle and macOS pointer state races 2026-08-27 21:08:20 +02:00
Viktor Liu
5986929dc3 Carry the VNC session key through the network map DB path 2026-08-27 21:08:20 +02:00
Viktor Liu
071bd94d14 Close view-only input, approval-responder and inbound-block gaps in the VNC path 2026-08-27 21:07:03 +02:00
Viktor Liu
4d7ca23184 Split the status summary into per-section helpers and scope retrackConn to service mode 2026-08-27 19:49:39 +02:00
Viktor Liu
8997670c15 Regenerate the daemon proto with the protoc version the tree was generated with 2026-08-27 19:40:35 +02:00
Viktor Liu
c804bee088 Put the approval prompt's Deny and Allow buttons on one row 2026-08-27 18:36:12 +02:00
Viktor Liu
64d808ee63 Split the VNC, routing and metrics flag copying out of the config request builders 2026-08-27 18:36:12 +02:00
Viktor Liu
90483da26b Address review findings on the VNC server, session auth and capture decoder 2026-08-27 18:33:52 +02:00
Riccardo Manfrin
6620219939 [client] Add catch-all NRPT rule when NetBird is the primary DNS resolver (#7071)
* [client] Add catch-all NRPT rule when NetBird is the primary DNS resolver

* Remove obvious comments

* Install the catch-all rule where the adapter's DNS is set

addDNSSetupForAll makes us the peer's main DNS forwarder, and the catch-all NRPT
rule is the other half of that same job: without it the adapter's NameServer only
adds one more resolver to the set Windows queries in parallel. Having the two in
one place says that, where a separate block at the end of applyDNSConfig read as
an afterthought.

The block could not simply move up: removeDNSMatchPolicies deletes the catch-all
key too, so installing the rule before it ran would have had the rule deleted
moments later. The cleanup now runs first, which is what it was always for - it
clears what the previous apply installed before this one installs anything - and
keeps being unconditional, so a leftover rule from an earlier run cannot survive
into a config that no longer wants it.

* Name the escape hatch after the behaviour it restores

NB_DISABLE_DNS_CATCHALL_NRPT described the mechanism it switches off. What an
operator reaching for it wants is the behaviour they had before, so name it that:
NB_USE_LEGACY_DNS_RESOLUTION, matching NB_USE_LEGACY_ROUTING, the only other
legacy switch in the client.

Not NB_WIN_LEGACY_FULL_TUNNEL_DNS_RESOLVE, as first suggested: the rule follows a
primary nameserver group, not a full tunnel, and putting FULL_TUNNEL in a public
variable name would carry that confusion for as long as the variable lives. No
OS prefix either, since nothing else in the client has one and this switch is
inert anywhere but Windows by construction.

Behaviour and default are unchanged: the catch-all rule is on unless the variable
says otherwise.

* Exempt .local from the catch-all rule

RFC 6762 reserves .local for multicast DNS and says unicast resolvers must not
answer for it. The catch-all rule hands it to us anyway, we forward it to
whatever upstream the primary nameserver group points at, and the answer comes
back NXDOMAIN for hosts that do exist - printers, NAS boxes, anything
announcing itself on the link. Confirmed on a Win11Pro VM: laptop.local resolves
with the client down and returns "Nome DNS inesistente" with it up, and the
client log shows the query arriving on the catch-all handler and being forwarded
to 1.1.1.1.

An NRPT rule that names a namespace and lists no servers is an exemption: the
DNS client resolves those names as it would with no rule at all. What that looks
like in the registry is not what it sounds like. Writing no server value and
clearing ConfigOptions produces a rule Windows treats as a no-op - it never
appears in Get-DnsClientNrptPolicy -Effective and the catch-all keeps the query.
The value has to be present and empty, with ConfigOptions still 0x8: the flag
says the server list is the meaningful part of the rule, and an empty list then
means "no server, resolve normally". Verified both encodings on the VM.

Installed together with the catch-all, since without one nothing captures .local
in the first place, and removed with it.

Exclusivity is unaffected elsewhere, and a more specific rule still wins - a
match domain under .local keeps resolving through NetBird, which is what a
legacy Active Directory domain named corp.local needs. Verified separately that
a match domain does take precedence over the catch-all: declaring fritz.box
against the local router restored laptop.fritz.box while the catch-all was in
force.

* Treat the root namespace as a match domain, not a special case

The catch-all had a function, a registry key and a call site of its own, which
made it look like a different mechanism. It is not: "." is an NRPT namespace like
any other, it just happens to match every name. So it goes into the match domain
list, and addDNSMatchPolicy writes it along with the rest — batching, GPO
variant, volatile keys and cleanup all come for free.

The .local exemption stays a rule of its own, and not for symmetry: it is the one
rule with a different server list, an empty one. Putting it in the same Name
value would give it our resolver and exempt nothing.

Windows expands a rule's Name value into one effective namespace each, so a rule
carrying {.example.com, .} still shows both as separate rows in
Get-DnsClientNrptPolicy -Effective. Nothing is lost for diagnosis by dropping the
dedicated key.

Suggested by Vik in review.

* Do not report a failed NRPT cleanup as success

removeRegistryKeyFromDNSPolicyConfig returned nil for every OpenKey error, so a
permission or registry failure was indistinguishable from a key that was never
there. Cleanup then reported success while the rule stayed in force — which is
how a rule outlives the interface it points at and keeps sending every query to
an address that no longer answers.

Distinguish the two, the way listNRPTRuleKeys already does for the policy store
root: a missing key is nothing to do, anything else reaches the caller.

restoreHostDNS now propagates that error instead of logging it. applyDNSConfig
keeps logging on purpose: there we are about to write fresh rules over whatever
survived, while restore is the path where a rule left behind is the whole
problem.

Also addresses review nits on the tests: doc comments on the two added cases,
reported Close and DeleteKey errors so a failed cleanup cannot contaminate the
next registry test, and a context message on the exemption's namespace assertion.
2026-08-27 16:18:49 +02:00
Viktor Liu
f7e186fbd0 Merge branch 'main' into embedded-vnc 2026-08-27 16:01:14 +02:00
Viktor Liu
63c26be72f [client] Add local Prometheus metrics endpoint (#6689) 2026-08-27 13:53:06 +02:00
Pascal Fischer
e06c17cf59 [management] network map from nmap data type (#6919)
Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>
Co-authored-by: Dmitri Dolguikh <dmitri.external@netbird.io>
2026-08-27 11:28:05 +02:00
Viktor Liu
473392a935 [client] Tolerate a still-locked updater binary when cleaning up after an update (#7286) 2026-08-26 20:03:56 +02:00
Viktor Liu
0bd1147ff0 [client] Keep NetBird traffic out of third-party fwmark rules (#7314) 2026-08-26 20:03:40 +02:00
Viktor Liu
f221347c7a [infrastructure] Trigger the dashboard wasm client bump on release tags (#7277) 2026-08-26 15:34:10 +02:00
dmitri-netbird
0a9ce7f797 [client] fix a flake in TestResolver_ConcurrentStaleHitsCollapseRefresh test (#7326)
* fix a flake in TestResolver_ConcurrentStaleHitsCollapseRefresh test

Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>

* use testify's eventually asserts

Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>

---------

Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>
2026-08-26 12:51:19 +02:00
Viktor Liu
7e8b4e1417 [client, proxy] Remove lazy connection exclusions and run Rosenpass on the embedded proxy (#6763)
* Run lazy connection manager for rosenpass peers

* Treat forward-target peers as normal lazy connections

* Run Rosenpass in permissive mode on the embedded proxy
2026-08-26 12:43:48 +02:00
dmitri-netbird
2621aaa619 [management, client] add protobuf breaking changes check (#7305)
* add protobuf breaking changes check

Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>

* disable path check for now

Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>

* enable breaking checks

Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>

* testing breaking change

Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>

* Revert "testing breaking change"

This reverts commit 05e6ef9b78.

Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>

* remove commented out proto paths

Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>

* disable pushes

Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>

* responded to feedback

Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>

* trigger workflow on changes to buf config or the workflow itself

Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>

* fix the workflow file name

Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>

* explicit config for actions

Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>

---------

Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>
2026-08-26 11:48:05 +02:00
Zoltan Papp
ed7d4de999 [client, ios] Migrate switft profile manager to go (#6528)
* [client] Add iOS NetBirdSDK profile manager binding

Mirror the Android profile manager in the iOS gomobile binding so the
core's ID-based profilemanager.ServiceManager owns profile state on iOS
too, instead of a parallel Swift reimplementation.

Adds client/ios/NetBirdSDK/profile_manager.go (//go:build ios): an
ID-based ProfileManager wrapping ServiceManager with iOS-specific path
handling (default profile at the container-root netbird.cfg, others as
profiles/<id>.json) and a gomobile-friendly API: List/Add/Switch/Rename/
Logout/Remove plus active config/state path accessors. The default
profile keeps the reserved "default" id and is never assigned a hex id.

* fix(ios): preserve profile name when saving config during auth

NewAuth built a fresh in-memory config from only the management URL, so
the SSO/setup-key save (DirectWriteOutConfig) overwrote the profile config
file the profile manager had just written, wiping the display name to ""
and forcing the UI to fall back to the profile ID. Load the existing config
when present and override only the management URL, keeping the name and keys.

* [client] Extract the mobile profile manager into client/mobile

The Android and iOS gomobile bindings carried two near-identical copies of
the profile manager. Move the shared implementation into a new client/mobile
package and reduce both bindings to thin adapters that only translate to
gomobile-friendly types (gomobile binds per package, so the Profile /
ProfileArray wrappers have to stay platform-side).

Also bring the account-email layer over to the shared package: an SSO login
records the account under <stem>.account.json so the next login can pass it
as an OIDC login_hint. Logout keeps it, profile removal drops it. The suffix
deliberately differs from .state.json, which the engine's state manager owns
in the same directory on mobile.

Adds profilemanager.Prefs (namespaced per-profile preference store) and its
cleanup in ServiceManager.RemoveProfile, exposed through the shared manager
as ProfilePrefs.
2026-08-26 09:42:50 +02:00
Viktor Liu
51095cb986 [client, management] Support per-peer lazy connection state and default proxy peers to lazy (#6762)
* Support per-peer lazy connection state and default proxy peers to lazy

* Classify forward targets from incoming config in lazy exclusion

* Set IsUserspaceBind mock so lazy manager starts in engine test

* Skip lazy exclude reconciliation when the set is unchanged

* Keep cached lazy flag when a sync carries no peer config
2026-08-26 09:33:51 +02:00
Viktor Liu
ccf8f43cb1 [client] Ask the OS for privileges when a guarded SSH setting is changed (#7066) 2026-08-25 20:15:16 +02:00
Zoltan Papp
15fff4c164 [client] Sweep connections on network loss via a shared netevents manager (#7254)
Losing the last network only flipped the availability state: the dead management, signal and relay sockets stayed silently connected until their own timeouts, so the client kept reporting Connected with no network at all.

Introduce client/netevents with a Manager that ties the availability state, the connection sweeper and the status recorder together, and move the netstate and netsweep packages under it (netsweep renamed to sweep). SetNetworkAvailable(false) now also sweeps the registered connections so their owners redial and the listener reaches the NoNetwork state.

The Android and iOS bindings own a Manager instance and inject it through the constructors; consumers hold the concrete *Manager whose nil zero value reports always-online and never sweeps, with interfaces kept only as parameter contracts. The relay guard settle wait moved into the Manager as WaitSettled, removing the netevents import from the relay package.
2026-08-25 18:43:19 +02:00
Viktor Liu
4b72019809 Merge branch 'main' into embedded-vnc 2026-08-25 18:26:29 +02:00
dmitri-netbird
c512bf25aa [management] handle nil ptr in sendInitialSync() when the peer is deleted (#7315)
* fix a nil-ptr error occuring in sendInitialSync when the peer being synced is deleted

Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>

* handle a nil ptr in GetPeerNetworkMapComponents

Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>

---------

Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>
2026-08-25 16:14:17 +02:00