Commit Graph
3514 Commits
Author SHA1 Message Date
Zoltan Papp 2d28f9002a [client] Fix the Windows tray deadlock on re-entrant window creation (#7449)
* [client] Fix the Windows tray deadlock on re-entrant window creation

The Wails systray runs the left-click handler synchronously inside the
tray window procedure, and creating a window on a running app pumps a
nested Win32 message loop while WebView2 initialises. ensureWindow held
the non-reentrant createMu across that creation, so the second button-up
of a double click re-entered ShowWindow from the pump and blocked the
main thread on its own lock. A goroutine holding createMu while the main
thread pumped, and the Open* dialogs holding mu across NewWithOptions,
Show, Hide and InvokeSync, exposed the same inversion.

WindowManager now serialises creation with a per-slot creating flag and
queues the callers' operations until the window exists, and no Wails call
runs while mu is held. The tray click and second-instance handlers call
ShowWindow off the message loop.

* [client] Serialize window operations while a slot is being created

Callers arriving after the window is published but before the creator
has drained the queue took the existing-window fast path and could run
ahead of older queued operations, so a newer SetURL could be overwritten
by an older one. withWindow now queues every caller while the creating
flag is set and clears the flag only once the queue is seen empty under
the lock.

A factory panic or a nil window left the creating flag set and the slot
dead; creation and drain now reset that state on early exit.

hideOtherWindows records the windows it hid only when no restore ran
in between, tracked by a generation counter, and re-shows them otherwise,
so a restore racing the hide cannot strand hidden windows.

* [misc] Run the client/ui subpackage tests in CI

The three test workflows filtered the package list with a `/client/ui`
prefix match, which dropped the subpackages along with the package that
cannot compile without a frontend build. `services`, `preferences`,
`i18n` and `authsession` all carry Go-side unit tests that never ran,
including the window manager re-entrancy regression test.

Anchor the pattern so only `client/ui` itself is excluded. The linux leg
keeps the prefix match on 386, where only the 64-bit gtk4/webkitgtk dev
packages are installed and the Wails application package would fail to
link, and the alpine container job keeps it for the same reason.

* [misc] Run the client/ui subpackage tests on a gtk4 4.10 runner

The previous commit let the subpackages into the linux client job, where
client/ui/services failed to build: the wails runtime's linux cgo layer
uses GtkFileDialog, which arrived in gtk4 4.10, and the job's ubuntu-22.04
runner ships 4.6.

Move them to their own job pinned to ubuntu-24.04 and restore the linux
client job's original exclusion, leaving the 386 and privileged legs on
the runner they have used since 2024. The new job needs no build cache,
sudo or privileged tag, so it stays a few seconds long.

Darwin and Windows keep the anchored pattern from the previous commit and
already run these tests green, including the window manager re-entrancy
regression test on the platform the deadlock was reported on.

* [client] Defer a window close that lands while the window is still being created

WindowManager publishes a dialog's slot only after the factory returns,
and on Windows the factory blocks in the WebView2 embed pump. A Close*
arriving in that gap found a nil slot and returned without doing
anything, so the dialog appeared afterwards for a flow that had already
been cancelled. The pre-fix Open* dialog functions held mu across the
whole creation, which blocked a concurrent Close* until the slot was
set; removing that lock hold reopened this gap.

Close* now goes through closeWindow: while the slot is being created it
records a closer in pendingClose, and finishCreation runs that closer
before any queued operation, so a window that is going away is never
shown and Wails never sees a Show on a destroyed window, which would
recreate it. Ops queued behind a close are dropped; windowOp carries no
factory, so they cannot be replayed into a new creation, and the
frontend callers reissue on the next state change.

The browser-login slot uses the same restoring closer from both
CloseBrowserLogin and CloseRenewFlow, since the popup's WindowClosing
hook only restores on a user close. Where two closers race one
creation the first registered wins, so a later caller cannot replace a
restoring closer with one that does not restore.
2026-09-14 15:37:46 +02:00
Zoltan Papp 0f797f89c1 [client] Bump wireguard-go to 8bf8fa968f1a (#7532)
- Make netTun.Close idempotent (#19)
- Fix keepalive pool block (#20)
- Keep timer paths non-blocking and bound staged packets per peer, kernel style (#21)

https://github.com/netbirdio/wireguard-go/pull/19
https://github.com/netbirdio/wireguard-go/pull/20
https://github.com/netbirdio/wireguard-go/pull/21
2026-09-14 15:19:09 +02:00
Mohd Quamar Tyagi 791401060d [management] Prevent deleting groups referenced by agent network budget rules (#7450)
`validateDeleteGroup` already refuses to delete a group that is still used by routes, policies, nameservers, setup keys, users, network routers, reverse proxy services, and agent network policies. Account-level agent network budget rules also store group IDs in `TargetGroups`, but that check was missing.

Deleting such a group left a dangling ID on the budget rule. `budgetRuleApplies` then never matched callers by group, so the spend cap silently stopped applying.

This adds `isGroupLinkedToAgentNetworkBudgetRule` and uses it in `validateDeleteGroup`, matching the existing helpers.
2026-09-12 12:37:23 +02:00
Philippe Vaucher 973173be9c [misc] Pull the MinIO test image from quay.io (#7516)
The minio/minio repository is no longer on Docker Hub: the repo and the
pinned tag both return 404 and anonymous pulls are refused, so
Test_S3HandlerGetUploadURL fails on every Linux CI run with "pull access
denied for minio/minio".

The same release is published at quay.io/minio/minio, digest
sha256:a1ea29fa28355559ef137d71fc570e508a214ec84ff8083e39bc5428980b015e,
so the testcontainers request points there and keeps the pinned tag.
2026-09-12 10:39:57 +02:00
Bethuel Mmbaga 82b1c7da22 [management] Harden OIDC issuer validation and discovery (#7435) 2026-09-11 19:09:13 +03:00
Maycon Santos b789ffbb9f [management] Expire unvalidated custom domain registrations (#7497)
Prevent unvalidated registrations from reserving domain names indefinitely.

Give pending registrations a 48-hour validation window and clean up expired entries at startup and every 60 minutes. Emit CustomDomainValidationExpired for each deletion and preserve registrations referenced by services.

Reject validation after expiry and prevent concurrent validation from recreating deleted registrations. Normalize domain names with the shared parser before registration.

Migrate existing pending registrations to receive a fresh 48-hour validation window.
2026-09-11 17:57:58 +02:00
Brandon Hopkins ec0c36b0e7 [client] Add light mode with system, light, and dark theme options (#7344)
* desktop UI light mode

* Theme review fixes plus macOS window outline fix

* Windows runtime chrome re-theming plus apply serialization

* Windows chrome threading and theme event ordering fixes

* Darken toggle and setting sidebar text

* resolve theme appearance, apply on UI thread

* read theme once per window

* Re-assert Windows dark opt-in after SetTheme

* split app-wide GTK theming from per-window chrome

* Update Wails dependency and checksums

* KDE tray icon panel fix

* Five review fixes: theme ordering, cgo dedup, KDE panel resolution

* Path guard hardening, toggle contrast, windows comment

* non-vacuous escape tests

* Default view edits

* Polish settings nav, controls, borders, and disc

* Profiles settings boarder, modals, and buttons

* Additional edits based on feedback

* Switch colors away from slight blue hue

* Update missing lang

* Fix vertical tab active view
2026-09-11 08:25:10 -07:00
dmitri-netbird f422c41654 [management] extract peer update logic and wrap it in tests (#7338)
* extract peer update loop into a dedicated struct and wrap it in tests

Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>

* make linter happy

Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>

---------

Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>
2026-09-11 17:07:26 +02:00
Zoltan Papp 794956a7a3 [client] Fix relay instance address race (#7498)
Read the relay instance URL and IP atomically to prevent reconnects from mixing values from different connections. Extend existing connection and offer/answer logs with relay URLs and IPs to help trace mismatched advertisements.
2026-09-11 16:21:10 +02:00
riccardom e0523efe5a [client] pqkem: make exchange logs discriminate offer/answer, role, transport, kind
The per-exchange lifecycle logs were all at LevelTrace (below the default field
level) and did not distinguish a signalling re-bootstrap from a data-path rekey,
which made diagnosing exchange activity in the field guesswork (e.g. telling an
ICE-retry-driven re-bootstrap storm apart from normal KEM rotation).

Promote the two pivotal events — an offer going out and a PSK converging — to
LevelDebug and tag every lifecycle log with four dimensions:
  - msg:  offer vs answer (implicit in the message)
  - role: initiator vs responder for this peer
  - via:  signal vs data-path (threaded into processOffer/processAnswer)
  - kind: bootstrap (a new connection / reconnect) vs rotation (a rekey)

kind is derived from the exchange's AckID (zero = bootstrap) on the responder
side and from viaSignal on the initiator side. No behavioural change.
2026-09-11 15:03:47 +02:00
riccardom 9ea3ddcc52 [client] pqkem: don't drop a responder's offer when the controller has no re-offer
pqControllerReoffer returned true whenever we are the controller running the
KEM, even when ShouldSendBootstrapOffer was false — which is the steady state
for every PQ peer past its bootstrap (exchange in awaitingRekey) and for
non-capable peers. handleRemoteOffer then returned early, dropping the ICE
credentials and relay info carried in the peer's offer, so a responder-initiated
reconnect (the normal path under lazy connections) stalled: the controller only
recovered on its own guard timing.

Return true only when a re-offer was actually sent; otherwise fall through to the
normal offer handling (sendAnswer + notifyListeners) so the connection can come
up on the retained PSK. A stale PSK still self-heals via the WG watcher, which
triggers a fresh signal offer that re-bootstraps.
2026-09-11 15:03:47 +02:00
riccardom aca68853d8 [client] pqkem: enforce exchange roles and re-bootstrap stale signal offers
Two convergence bugs surfaced by the security review:

- Role guard (finding B): processOffer accepted an offer even when we are the
  KEM initiator for the peer, and processAnswer accepted an answer when we are
  the responder. The KEM is unidirectional (initiator offers, responder
  answers), so a role-violating message is anomalous — a desync, a duplicate,
  or an injected/spoofed data-path packet. Processing it derived and committed
  a fresh PSK, overwriting a live one and silently dropping any in-flight
  exchange (whose retry loop then exited without raising a failure or
  re-bootstrapping). Reject offers when we are the initiator and answers when
  we are not; this drops only anomalous traffic and leaves the normal flow
  untouched.

- Re-bootstrap on signal re-negotiation (finding A): SignalOffer was idempotent
  in stateAwaitingRekey too, replaying the frozen bootstrap offer. After the
  responder restarted and lost its state it derived a different PSK from fresh
  material, which our awaitingRekey side then rejected — a permanent desync with
  no recovery (in strict mode the peer stays blocked). Make the idempotency
  apply only while a bootstrap is still in flight (awaitingAnswer); once a PSK
  is derived, a fresh signal offer starts a new exchange so both sides converge.
  The controller-double-offer case the idempotency guarded is already covered by
  ShouldSendBootstrapOffer. Reusing the cached offer also reused the same
  ephemeral keys across exchanges, reducing forward secrecy.

Both paths have a failing-without-the-fix regression test.
2026-09-11 15:03:47 +02:00
riccardom 91f1a724e2 Adds observability and fixes cross cases
- strict-kem vs strict-rp said "Connected + Quantum resistance: true" but it is actually blocked
- perm-kem vs perm-rp "Connected + Quantum resistance: true" but it's a classic WG link, without PQ safety
2026-09-11 15:03:47 +02:00
riccardom 1ee012160b Add comment to explain working boundaries 2026-09-11 14:48:54 +02:00
riccardom 3efd4c94a0 UDPv4 is sufficient! we always have 100.x overlay. IPv6 is additive
ipAllowed[0] is the v4
2026-09-11 14:48:54 +02:00
riccardom ecddbf2e55 Addresses CI fixes 2026-09-11 14:48:54 +02:00
riccardom 8a04fa4d3f Addresses CI fixes 2026-09-11 14:48:54 +02:00
riccardom 11555c3139 Addresses CI fixes 2026-09-11 14:48:54 +02:00
riccardom 1dc48bcc06 Addresses CI fixes 2026-09-11 14:48:54 +02:00
riccardom 82699afb0f [client] pqkem: split handshaker Listen into per-message handlers
Reduce Listen's cognitive complexity (SonarCloud S3776, 30 -> under 25) by
extracting the offer and answer cases into handleRemoteOffer/handleRemoteAnswer,
with shared onSignalReceived/notifyListeners helpers and a pqControllerReoffer
helper for the controller re-offer branch. No functional change.
2026-09-11 14:48:54 +02:00
riccardom 4c6735e8ae [client] pqkem: don't let the responder's delayed update revert the PSK
The responder configures WireGuard with endpoint=nil first, then a delayed
update (scheduleDelayedUpdate) applies the real endpoint after fallbackDelay.
It captured the preshared key at schedule time and re-applied it. With the
post-quantum exchange the PSK can change within that window (a fresher PSK
derived and applied via SetPresharedKey), so re-applying the captured one
reverted WireGuard to a key the remote peer no longer used — a mismatch that
stalled the handshake until the WGWatcher timeout forced a retry (~30s).

Pass a nil PSK in the delayed update so it only sets the endpoint and leaves
the current PSK in place; the latest SetPresharedKey wins.
2026-09-11 14:48:54 +02:00
riccardom 3b61222b07 [client] pqkem: gate controller re-offer to kick the KEM exactly once
When the controller receives the responder's (KEM-less) offer it replies with
its own KEM offer instead of answering, so the only transaction that brings the
tunnel up is the one that also carries the PSK. Guard that reply with
ShouldSendBootstrapOffer so it fires only when no exchange is in flight: without
it, every responder offer triggered another offer (an offer-per-offer runaway).
The whole behaviour is isolated to the KEM path (config.PQ != nil); non-PQ
connections answer as before.
2026-09-11 14:48:54 +02:00
riccardom 1870b104ba [client] pqkem: make signalling bootstrap idempotent through awaitingRekey
The controller sends its KEM offer both on its own guard event and in reply
to the responder's offer. SignalOffer was idempotent only while awaiting the
answer; once the answer arrived (awaitingRekey) a repeat call started a fresh
exchange with a different PSK, desyncing the two peers (one on the old PSK,
one on the new) so WireGuard derived misaligned transport keys and dropped all
data. Treat awaitingRekey as in-flight too and return the same offer.
2026-09-11 14:48:54 +02:00
riccardom 4ffa2c724f Introduces a forced imparity on MLKEM bootstrap.
To ensure two peers agree on a key, we need asymmetry. one peer is
the controller ("initiator") the other is the "responder".

Otherwise imagine two offers in parallel driving two answers at the same time

   A                   B
   | <----B-OFFER----- |
   | -----A-OFFER----> |
   |                   |
   |                   |
   ---------------------------------
  |****** ICE + WG Handshake ****** |
   ---------------------------------
   |                   |
   | <----B-ANSWER---- |
   | -----A-ANSWER---> |

PSK is derived on receive of offer, so A and B derive different PSKs.
When WG handshake takes place it picks misaligned PSKs.

So we impair the two nodes and only the offer of one of the two (the controller/initiator)
is allowed to progress and drive the answer (and carry the KEM material).

If a responder initiates an offer, we redo the offer towards it. This is oK
since the ICEworker don't treat offer/answers differently.
2026-09-11 14:48:54 +02:00
riccardom 3ed4b4fb43 Revert "Introduces a forced WG handshake on initial MLKEM bootstrap."
This reverts commit b5a72eca65.
2026-09-11 14:48:54 +02:00
riccardom a78d587e7a Introduces a forced WG handshake on initial MLKEM bootstrap.
To ensure two peers agree on a key, we need asymmetry. one peer is
the controller ("initiator") the other is the "responder".

Otherwise imagine two offers in parallel driving two answers at the same time

   A                   B
   | <----B-OFFER----- |
   | -----A-OFFER----> |
   |                   |
   |                   |
   ---------------------------------
  |****** ICE + WG Handshake ****** |
   ---------------------------------
   |                   |
   | <----B-ANSWER---- |
   | -----A-ANSWER---> |

PSK is derived on receive of offer, so A and B derive different PSKs.
When WG handshake takes place it picks misaligned PSKs.

So we impair the two nodes and only the offer of one of the two (the controller/initiator)
carries the KEM material.

This means that if the responder OFFER/ANSWER comes first, when the controller/initiator's one
completes (and the genuine PSK is shared between A and B, we need to force a new WG handshake with
the proper keys.
2026-09-11 14:48:54 +02:00
riccardom 2b27040fb0 Anticipates PSK before WG does handshake so it finds it to set it 2026-09-11 14:48:54 +02:00
riccardom b93e8d6676 Skip default port send in signal proto 2026-09-11 14:48:54 +02:00
riccardom cfaf4f1bab Allow non strict mode 2026-09-11 14:48:54 +02:00
riccardom 5a6de00f7a Adds PQ connection tests 2026-09-11 14:48:54 +02:00
riccardom cb5ee8efa8 Remove obvious comments; leave only the why of things 2026-09-11 14:48:54 +02:00
riccardom 202176ec23 Be more explicit on names that is a fake key to ensure we don't communicate with others in strict mode 2026-09-11 14:48:54 +02:00
riccardom bfa36521da Adds test to validate compromised keys are not accepted 2026-09-11 14:48:54 +02:00
riccardom d988525850 Prioritize Kem over RP 2026-09-11 14:48:54 +02:00
riccardom 63d29966fd pqkem: concurrency tests (recovery + race)
- RecoversViaResignalAfterDataPathBreak: a data-path rotation that can no longer
  converge raises OnRekeyFailed, and re-bootstrapping over signalling resyncs both
  peers on a fresh PSK even while the data path stays broken.
- ConcurrentRekeysNoRace: hammers the single-lock state machine with concurrent
  rotation clocks from many goroutines (run with -race) and asserts no split-brain
  via a final deterministic bootstrap.
2026-09-11 14:48:54 +02:00
riccardom d485567978 pqkem: recover from persistent rekey failure by re-bootstrapping over signal
OnRekeyFailed now re-runs the KEM bootstrap over Signal (conn.RequestReoffer ->
handshaker.SendOffer) instead of only logging: a fresh signalling offer starts a new
exchange that overwrites the stalled PSK on both sides, resyncing after a persistent
data-path desync. Chosen over a responder-side awaitingAck revert (which fights the
confirm-less ack timing) and a full tunnel teardown (heavier). The tunnel stays up on
the previous PSK meanwhile since Signal is independent of the broken data path.
2026-09-11 14:48:54 +02:00
riccardom e1d5551ba5 Discriminate initial from rekey failure 2026-09-11 14:48:54 +02:00
riccardom f322be1944 pqkem: strict (fail-closed) mode + wire status Quantum resistance
Strict mode (NB_PQ_MLKEM_STRICT, default off) closes the initial PQ-vulnerable
window (NET-1408): when enabled, conn.presharedKey programs a per-conn random
sentinel PSK until the ML-KEM exchange derives the real one, so no session can form
on a non-PQ key (the real PSK is pushed via SetPresharedKey once it converges).
Default stays opportunistic.

Also surface PQ status: the peer 'Quantum resistance' flag (RosenpassEnabled) is now
true when an ML-KEM PSK has been derived for the peer, not only for Rosenpass.
2026-09-11 14:48:54 +02:00
riccardom b1b2cc14c1 pqkem: rotate PSK in kernel mode instead of skipping
The idle-gate reads LastActivities, which only tracks per-peer data in userspace;
in kernel mode it is empty, so the gate treated every kernel peer as idle and
disabled data-path rotation entirely. Detect the bind via IsUserspaceBind and, in
kernel mode, report zero activity age (always 'active') so rotation runs on every
rekey. Lazy back-to-idle is already limited in kernel; the eBPF WG-activity
detection will later supply a real signal that excludes handshake/pqkem traffic.
2026-09-11 14:48:54 +02:00
riccardom 8c8abfb2b1 pqkem: derive PSK with HKDF-SHA256
Replace the raw SHA-256 concat combiner with HKDF-SHA256 (crypto/hkdf, Go 1.24):
IKM = ML-KEM_ss || X25519_ss (draft-ietf-tls-ecdhe-mlkem order), salt = the
domain-separation label, info = full transcript (offer || answer) || canonicalised
peer identities. Keeps the transcript + identity binding while using a proper KDF.
2026-09-11 14:48:54 +02:00
riccardom b1c0a2bffe Don't rotate PQ keys if data path is idle for ~90s (less than a WG handhshake time 2026-09-11 14:48:54 +02:00
riccardom 562c02fa42 Adds log tracepoints
- Add a trace slog level (NB_PQ_MLKEM_LOG_LEVEL=trace) and move the verbose
  per-exchange lifecycle logs (offer/answer/PSK/ack/rotation) to it, so debug
  stays quiet and troubleshooting is opt-in.
- Stop logging the raw preshared key; drop the temporary pqkem-dbg OnRemoteOffer/
  OnRemoteAnswer probes.
- Demote the per-handshake conn log to trace.
2026-09-11 14:48:54 +02:00
riccardom b92b14543d Renames SetRemotePort to SetRemoteAddr 2026-09-11 14:48:54 +02:00
riccardom 1c46b4e9df pqkem: clock data-path PSK rotation from WireGuard handshakes
Source OnDataPathRekeyed from the WGWatcher's per-handshake callback
(onWGCheckSuccess), which fires only on a fresh handshake, and OnDataPathDown
from the handshake-timeout path. A fresh handshake clocks the next chained
KEM exchange pushed over the data-path UDP transport.
2026-09-11 14:48:54 +02:00
riccardom e3ab585c56 pqkem: register data-path endpoint from signalling
Learn the peer's data-path endpoint from the signalling offer/answer: its WG
overlay IP combined with the advertised pq UDP port (SetRemotePort -> AddPeer).
Registering here is safe before the tunnel is up because sends only ever fire
once it is (clocked by OnDataPathRekeyed). RemovePeer is wired at peer teardown
(engine.removePeer), not on transient disconnect.
2026-09-11 14:48:54 +02:00
riccardom 7c276ce727 pqkem: apply derived PSK at WG peer-config time (pull) + keep push for rekey 2026-09-11 14:48:54 +02:00
riccardom 555310bd56 pqkem: carry KEM offer/answer over the signalling exchange 2026-09-11 14:48:54 +02:00
riccardom 56b315d954 pqkem: dedicated slog logger via NB_PQ_MLKEM_LOG_LEVEL 2026-09-11 14:48:54 +02:00
riccardom bac6af349b Homogeneous logs prefix 2026-09-11 14:48:54 +02:00
riccardom 3270fc5e18 Bit of renaming
peer -> peerAddrs
have types for remoteID and localID
t.Close log error
Manager SetTransport -> Start
2026-09-11 14:48:54 +02:00