mirror of
https://github.com/netbirdio/netbird.git
synced 2026-10-06 21:49:08 +02:00
f5707c348532a4b15ea39fb53264339f6045932c
3431
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f5707c3485 |
[client] Add RTL layout support to the desktop UI (#8076)
* [client] Add RTL layout support to the desktop UI <html dir> follows the active language, Radix primitives get the matching dir, physical spacing/positioning uses logical utilities, and directional icons, animations, arrow-key navigation and tooltip sides flip in RTL. * [client] Address RTL review feedback Force LTR with isolation for monospace values, let truncated names take their direction from their content, and keep the profile name field in the UI direction. * [client] Keep translated monospace labels in their own direction The forced-LTR rule for monospace values now skips elements with an explicit dir, translated development labels use dir=auto, and monospace values rendered through TruncatedText opt into LTR. --------- Co-authored-by: Edward <43848523+thomashacker@users.noreply.github.com> |
||
|
|
a8dff998ef |
[client] Gate settings updates on value, not on field presence (#7398)
* [client] Gate settings updates on value, not on field presence
The update-settings kill switch (--disable-update-settings /
NB_DISABLE_UPDATE_SETTINGS / the MDM DisableUpdateSettings key) forbids
changing settings, but it decided what a "change" was by looking at
whether a field was present in the request. The CLI fills the whole
config surface of SetConfigRequest and LoginRequest from its flags and
environment on every `netbird up` (setupSetConfigReq in cmd/up.go), so a
client configured by environment restates its own configuration on every
start and tripped the gate every time.
SetConfig only warned about that, but Login carries the same fields and
was gated the same way, and Login runs inside the CLI's backoff loop: the
daemon answered every attempt with codes.Unavailable, `netbird up` never
completed, and a container with NB_DISABLE_UPDATE_SETTINGS plus any
config env var (NB_MANAGEMENT_URL, for one) could not come up at all.
Both gates now compare values. Config.WouldChange is the dry-run half of
UpdateConfig: it runs the very same diff logic (Config.apply) against a
copy of the stored config, so the gate cannot drift from what an actual
update would do, nor go stale when a field is added. A request that
restates what the profile already holds changes nothing and is allowed; a
request that diverges is refused exactly as before, and a dry run that
cannot be evaluated fails closed. A profile with no config on disk yet is
judged against the config the daemon would create for it.
For Login, the compared input comes from loginOverridesInput, which
persistLoginOverrides also uses to perform the write, so the gate judges
precisely the two fields a login can persist (management URL, pre-shared
key) and no field it ignores.
Two adjacent defects surfaced while making the comparison exact:
- Config.apply compared URLs as raw strings, so the same endpoint spelled
without its default port ("https://api.netbird.io" vs
"https://api.netbird.io:443") counted as a new value and rewrote the
config. It now compares the parsed forms.
- UpdateConfig did not collapse the redacted pre-shared key, unlike
UpdateOrCreateConfig and DirectUpdateConfig, so a UI round-trip of the
mask replaced the stored key with asterisks.
The CLI warning for a refused SetConfig said the method was not available
in the daemon, which sent people looking for a version mismatch that was
not there; it now reports the refusal.
* [client] Do not write the profile config while only reading it to decide
The update-settings gate needs the stored config to decide whether a request
changes anything, so the previous commit moved that read ahead of the refusal.
The read is not side-effect free: profilemanager.GetConfig writes the config
back whenever apply() has to fill in a default the file was missing. A request
that the gate then refuses had therefore already rewritten the profile file.
PeekConfig is GetConfig without that write-back. The returned config is still
normalized in memory, which is what the decision needs; the file is left
exactly as it was found. Every caller of storedConfigAtPath feeds a gate that
can refuse, so they all peek.
Note for reviewers: the daemon still normalizes the file on startup and on
every real update, so nothing depends on a read performing that write.
* [client] Compare service URLs as endpoints, not as strings
Three places in one request path each had their own notion of "same
management URL": the config layer compared the parsed URLs as strings, the
privileged-change gate compared scheme + host + effective port, and the MDM
conflict check compared strings after filling in the default port. Only the
middle one was right.
A string comparison answers the wrong question. "https://api.netbird.io",
"https://api.netbird.io/" and "https://API.netbird.io:443" are one endpoint
written three ways, so a client restating its own management URL with a
trailing slash — a normal way to write it — was still read as a client asking
to be repointed, and the update-settings gate refused it. The MDM check had
the same flaw against the enforced value.
profilemanager.SameServiceURL is now the single comparison: same scheme, same
host case-insensitively as DNS names are, same effective port. The config
layer, the privileged-change gate and the MDM conflict check all defer to it,
so there is one answer to "did this URL change?" instead of three.
* [client] Stop the config dry run from generating throwaway keys
The dry run's baseline for a profile with no config file yet went through
createNewConfig, and apply() generates a WireGuard and an SSH key whenever it
finds those fields empty. The baseline is compared against and discarded, so
every evaluation minted a keypair it threw away — and logged "generated new
Wireguard key". The CLI retries Login in a backoff loop, so a first `netbird up`
on a fresh profile filled the daemon log with what reads like peer-key rotation.
The baseline now starts from the shared skeleton with placeholder keys, so
apply() has nothing to generate. No ConfigInput field maps to either key, so
the comparison is unaffected.
* [client] Cover the login the update-settings gate used to refuse
The gate's decision procedure was tested directly, but no test drove the Login
RPC that the refusal actually broke: the CLI retries Login in a backoff loop,
so a refused no-op login is what kept a client configured by environment from
ever coming up. The handler-level coverage stopped at the refusal case, which
passes on the pre-fix code too.
This test fails on the pre-fix daemon with "update settings are disabled" and
passes now. Past the gate the handler does real work the test does not stand
up, so it asserts only that the refusal did not happen.
* [client] Re-take the update-settings decision under the config lock
Login checks twice on purpose: the first check refuses the ordinary case
early, and authorizeAndPrepareLogin re-takes the authoritative one under
guardedConfigMu because the first is unsynchronized against a concurrent
privileged request. The update-settings decision is now equally
value-dependent — it compares the request against the stored config — but it
was taken only in the first, unlocked check.
So a login that was a no-op when it was checked could be written after a
concurrent writer had repointed the profile, which is exactly the window the
lock exists to close. The decision is now re-taken alongside the privilege one,
which also makes it the last read before persistLoginOverrides writes.
The test drives that interleaving through the existing afterLoginPreCheck seam
and fails without the re-check.
* [client] Drop an unreachable guard and fix two stale comments
- loginOverridesInput's nil-message guard cannot be reached: Login
dereferences the message well before it, in storedLoginConfig.
- The docstring above afterLoginPreCheck described persistLoginOverrides,
which lives further down the file and now carries its own.
- UpdateConfig's comment named DirectUpdateConfig; the function is
DirectUpdateOrCreateConfig.
* [client] Make config reads pure and provision the identity explicitly
Reading a config wrote it back. profilemanager.readConfig persisted whatever
apply() had filled in, and ReadConfig created and wrote the file outright when
it was absent, so every reader was quietly a writer: a gate deciding whether to
refuse a request, a UI listing profiles, a mobile getter reading one preference.
The previous commit worked around that with a PeekConfig variant, which left
two read functions with opposite side effects and the antipattern still there
for everyone else.
Only one thing in a read genuinely had to be persisted: apply() generated the
WireGuard and SSH keys when it found them empty, and a generated key cannot be
recomputed — losing it means the peer comes back with a different identity and
registers again. Everything else apply() fills in is a deterministic default
that the next read recomputes anyway.
So identity provisioning is now its own step, Config.EnsureIdentity, and the
callers that provision write the result out themselves, in the open:
- Server.getConfig, the daemon's provisioning point;
- the CLI's foreground login, which is about to dial management;
- update() / directUpdate(), the config write paths — a stored profile can
legitimately carry no identity, since a mobile logout clears the keys in
place, and the next write is what has to mint a new one.
ReadConfig and GetConfig no longer write anything, PeekConfig is gone, and the
dry-run baseline no longer needs placeholder keys to keep apply() from minting
real ones.
One deliberate leftover: readConfig still calls util.EnforcePermission, which
chmods a config file whose permissions are too broad. It changes no content and
is idempotent, and dropping it would leave a legacy file world-readable until
its first write.
* [client] Name the two config readers for what they do
ReadConfig and GetConfig differed in one thing — what happens when the file is
absent — and neither name said which was which:
- ReadConfig -> ReadOrGenerateConfig (reads it, or generates one in memory)
- GetConfig -> GetExistingConfig (reads it, or fails)
Three comments went with them:
- GetConfig's said "return with Config and if it was created. Errors out if it
does not exist", which described a bool it does not return and a creation it
never performs.
- ReadConfig's explained that it does not write, which is what a reader is
supposed to do anyway.
- Server.getConfig's said it "errors out if it does not exist", which it does
not — it resolves a default config, and now provisions the identity too.
* [client] Do not panic on a config with no sync message version
apply() wrote the incoming sync message version through the stored pointer,
without checking it was there: a config that carries no version yet made it
dereference nil. Reachable from the update-settings dry run, which runs inside
a request handler — where failing closed is the worst acceptable outcome, and a
panic is not one.
The field is now reassigned like every other optional one, which also means
apply() no longer mutates anything the caller still holds through a pointer, so
the dry run's copy has one less field to detach.
Reported by cubic-dev-ai on PR #7398.
* [client] Compare the client certificate paths before reporting a change
apply() assigned the incoming mTLS certificate and key paths and set updated
unconditionally, without comparing them to what the config already held. It is
the same presence-instead-of-value mistake this branch set out to fix, one
layer down: a caller restating its own certificate paths was reported as
changing them, which trips the value-aware update-settings gate.
Reported by cubic-dev-ai on PR #7398.
* [client] Address the remaining bot findings on PR #7398
- Login logged the active-profile-state error and returned the same cause; the
repo's guidelines call for one or the other, and the wrapped error is the one
that carries context. (CodeRabbit)
- `netbird up` reported a codes.Unavailable SetConfig failure as "the daemon
refused the settings update", but that code also covers a daemon that became
unreachable. It now reports what the daemon said without asserting why.
(cubic-dev-ai)
- TestLogin_ChangingTheManagementURLIsRefused asserted the error and nothing
else, while "refused before it can touch daemon state" is the contract. It now
checks the stored management URL, the in-progress login and the active profile,
matching its SetConfig counterpart. (cubic-dev-ai)
* [client] Keep the peer identity out of a read that finds no file
ReadOrGenerateConfig resolves a default config when the profile has no file
yet, and createNewConfig was minting the WireGuard and SSH keys while doing so.
That defeated the provisioning pair it was meant to serve: the CLI's foreground
login calls EnsureIdentity to find out whether it has to persist the keys, got
generated == false because the read had already generated them, and so never
wrote them out. The login then dialed management with an identity that only
existed in memory, and the next login registered a second peer.
createNewConfig no longer provisions. createProvisionedConfig is the variant
that does, and the callers whose contract is "usable as it comes back" use it:
CreateInMemoryConfig, whose callers connect with the result, and the two
create-and-write branches. A read gets a config with no identity, so the
caller's own EnsureIdentity reports the work and triggers the write.
Reported by CodeRabbit and cubic-dev-ai on PR #7398, both on the same defect.
* [client] Stop the gate test from dialing the real management server
TestLogin_RestatingTheStoredConfigPassesTheGate asserts that the gate lets a
no-op login through, and the handler then went on to do the login for real:
isLoginRequired builds an auth client when isLoginRequiredFn is unset, so the
test dialed the profile's management URL — api.netbird.io:443. It took 1.05s
locally and would hang on a runner with no egress, for a fact about the gate
that needs no network at all.
Stubbed like the login_outcome tests do. The test now runs in 0.00s.
Reported by cubic-dev-ai on PR #7398.
* [client] Keep the admin panel path part of its identity
The endpoint comparison introduced for the management URL was applied to the
admin URL too, and that one is opened in a browser rather than dialed over
gRPC: a panel served under /netbird is not the panel served at the root. So a
config whose admin URL differed only by path reported no change, and the new
path was never persisted — a custom panel URL could not be updated at all.
SameServiceURLIncludingPath adds what a URL carries past its endpoint (path,
query, fragment, userinfo) while still treating equivalent spellings as equal:
a missing path and "/" are the same root, and so is a trailing slash. The
management URL keeps the endpoint-only comparison, since only the endpoint is
ever dialed.
Ports are also normalized numerically now, so ":0443" and ":443" are one port.
Reported by cubic-dev-ai on PR #7398 (two findings).
* [client] Treat a profile with no identity as already deregistered
Two findings on the same consequence of pure reads: a profile can legitimately
carry no keys, because logging out clears them in place.
- sendLogoutRequestWithConfig went straight to wgtypes.ParseKey and failed with
"incorrect key size: 0" on the second logout of the same profile. There is
nothing to deregister for a peer that was never registered, so it returns
cleanly. Before pure reads this case was hidden: the read minted a key and
the daemon dialed management with one it had never seen.
- The mobile logout read the config with the generating reader right after
checking the file exists. The two are not atomic, so a profile removed in
between was resolved from the defaults and recreated by the write that
follows. It uses the existing-file reader now.
Reported by cubic-dev-ai and CodeRabbit on PR #7398.
* [client] Fail `netbird up` when the daemon refuses the settings update
With the update-settings kill switch on, `netbird up --enable-rosenpass`
connected and said almost nothing: SetConfig refused the change, the CLI
downgraded that to a warning, and Login carries no rosenpass field to apply, so
the flag was silently dropped. The setting stayed disabled, which is the point
of the switch, but the caller was never told their request had been ignored.
The refusal now travels as codes.FailedPrecondition instead of
codes.Unavailable, and the CLI fails on it. Unavailable means "the daemon
cannot serve this call", which is why the CLI downgraded it and why
client/ui/services reads it as an unreachable daemon — both wrong for a daemon
that answered and refused. FailedPrecondition also matches what the MDM gate
already returns for a managed field, so both refusals are now one class of
error, and it is added to the login backoff's early-exit codes so a refused
login stops instead of retrying for 30s.
This does not put the container back in the deadlock: with the value-aware
gate, a client restating its own configuration is not refused at all, so
nothing reaches this path unless a real change was asked for.
* [client] Name the reader storedConfigAtPath actually calls
The purity note still said profilemanager.GetConfig, which the rename two
commits later turned into GetExistingConfig.
Reported by cubic-dev-ai on PR #7398.
* [client] Restore the gofmt alignment of the error constants
The comment added above errUpdateSettingsDisabled in the previous commit split
the const block's alignment group, so gofmt wants the two constants above it
re-aligned. CI runs gofmt, so this would have failed the lint job.
* [client] Let an unprivileged caller log out a profile with no identity
The empty-key check sat behind requirePrivilegeForDeregistration, so an
unprivileged logout of an identity-less profile was refused with
PermissionDenied instead of completing as the no-op it is. And it was refused
for most profiles, not a corner case: the gate arms whenever the SSH server is
enabled, and sshServerEnabled reads an absent ServerSSHAllowed as enabled, so
every legacy profile qualifies.
The check now runs first. What the gate protects against is handing this
machine's registered key to another management server; with no key there is
nothing to hand over and nothing to protect.
Reported by CodeRabbit and cubic-dev-ai on PR #7398, both on the same defect.
* [client] Stop `netbird login` from retrying a refusal for 30 seconds
`netbird up` and `netbird login` both run Login through the backoff cycle, and
each carried its own copy of the list of codes that end it. Only up.go learned
about codes.FailedPrecondition, so a refused `netbird login` kept retrying and
then reported "login backoff cycle failed" instead of what the daemon said.
terminalLoginError is now that list, once, next to WithBackOff — the duplicated
copies are what let the two commands disagree in the first place.
Reported by cubic-dev-ai on PR #7398.
* [client] Answer terminalLoginError's nil case on its own terms
A successful Login reaches terminalLoginError with a nil error, and nothing
covered that. It happens to work on grpc v1.80.0 — gstatus.FromError(nil)
answers (nil, true), and Status.Code tolerates a nil receiver by returning
codes.OK, which is not in the terminal set — but that is a chain of internal
details to be relying on for the common path, and none of it was asserted.
Now the nil error is handled where it is obvious, and the table covers it.
Reported by CodeRabbit on PR #7398, which called it a panic; measured on
v1.80.0 it is not one. The gap was the untested reliance, not a crash.
* [client] Treat an unset optional field as its default when diffing a config
Seven Config fields mean "the effective default" when they hold no value:
the five SSH toggles, the SSH JWT cache TTL, and the network monitor. Every
consumer already reads a nil as that default, but apply() diffed them by
presence — `config.X == nil || *input.X != *config.X` — so an input restating
the default counted as a change.
That made the update-settings gate refuse `netbird up` outright. The CLI
sends every flag whose value came from an environment variable
(SetFlagsFromEnvVars goes through pflag's FlagSet.Set, which marks the flag
Changed), and the config a plain login writes leaves all seven unset, so a
container configured with, say, NB_ENABLE_SSH_ROOT=false restated a default
the file held as null on every start and was answered with
FailedPrecondition.
apply() now resolves the seven up front, the way it already did for
ServerSSHAllowed and RemoteJobsAllowed, which also repairs such a profile on
its next write. With the values named, the comparisons below diff values
instead of presence, so their nil branches are gone.
The network monitor keeps its platform default — on for windows and darwin —
and naming it as false elsewhere is what createEngineConfig already read a
nil to be. getJWTCacheTTL reaches the same 0 through its own default, and
Android's GetEnableSSH* getters already answered nil with false.
* [client] Normalize the config before diffing it in WouldChange
apply() reports two different things through one bool: an input that changed
a value, and a field it had to fill in because the config carried none. The
update-settings gate reads that bool as "the caller asked for a change", so
any config still missing a default answered a request that asks for nothing
with a refusal.
Readers already hand out normalized configs — readConfig applies an empty
input for exactly this reason — which is why the gate got away with it. But a
handler that refuses a request must not depend on where its caller obtained
the config, and it must not start reading "this profile predates a field" as
"the caller asked for a change" the day someone adds one with a default.
WouldChange now runs the filling-in as a pass of its own and discards its
verdict, so the pass that answers the caller measures only what the input
did.
* [client] Stop the last config write that skipped normalization
Every path that creates or updates a profile config goes through apply(),
which resolves an optional field to its default — except RenameProfile,
which read the file with a bare json.Unmarshal, set the name, and wrote it
straight back. That copied whatever the file held, so a config written by a
client that stored these fields as null kept them null. It could not
introduce a null, only carry one forward, but renaming a profile is a poor
place to leave a half-resolved config behind. It now reads through
GetExistingConfig, which normalizes what it hands out.
The tests state the invariant the fix completes, over the *bool fields of
Config listed by reflection so a field added later is covered without
touching them: none may come out of apply() unset, and no write may store
one as null. An optional bool that can be nil, true or false forces every
reader to invent the meaning of nil, and makes a diff of the config compare
presence rather than value — which is exactly what refused `netbird up` for
a client restating its own defaults.
SyncMessageVersion stays a genuine three-state field and is not covered: it
is an *int whose absence means the client pins no version, and it travels to
management that way.
* [client] Refuse a serialized config that carries no peer identity
ConfigFromJSON still promised a "fully initialized" config after this PR
moved key generation out of apply() into EnsureIdentity, but identity stopped
being one of the defaults it applies. Its two callers both connect with what
they get back: the iOS SDK's Client.SetConfigFromJSON keeps it as the
preloaded config Run() uses on tvOS, and Auth.SetConfigFromJSON as the config
it authenticates with.
No caller feeds it a document without keys today — every stored document
comes from Auth.GetConfigJSON, whose config is provisioned by
DirectUpdateOrCreateConfig or CreateInMemoryConfig, and the tvOS app only
ever edits fields of a document it already has. This is a safety net for the
next caller, not a live bug.
Provisioning the identity here would be the wrong net. Neither caller can
hand a generated key back to the store the document came from — Client
exports no config at all — so the peer would connect under an identity
nothing persists and register anew on every launch, which is the failure the
EnsureIdentity split exists to prevent. A document with no identity means
nobody has logged in yet, and saying so is the only useful answer.
Both keys are required because both are dead ends when missing: an empty
WireGuard key fails the management login on its size, and an empty SSH key
fails ssh.GeneratePublicKey in ConnectClient before the engine starts.
* [client] Say that the null-on-disk fixture is synthesized, not written
The test comment described the null state in the present tense — "the config
a plain login writes leaves every one of them unset" — which was true before
this branch and is not any more: apply() now resolves those fields, so a
login writes them set. unsetOnDisk puts the null state back deliberately, to
stand in for a profile an older client wrote. Comments only.
* [client] Gather the optional-field defaults into one function
Resolving an unset optional field was spread over five places: the two
values newConfigSkeleton pre-sets, the block this branch added for the SSH
toggles, the network monitor's own if, the `else if` tails of
ServerSSHAllowed and RemoteJobsAllowed, and a trailing if for
DisableNotifications several hundred lines further down. Reading apply() left
no single answer to "what does this field default to, and who decides".
They now live in Config.resolveUnsetDefaults, which apply() calls before it
compares anything — the ordering being the point, since it is what lets
every comparison below diff values instead of presence. The comparisons for
ServerSSHAllowed, RemoteJobsAllowed and DisableNotifications lose their
`config.X == nil ||` clauses accordingly, as the other six already had.
newConfigSkeleton keeps its two, and that is the one asymmetry worth naming:
ServerSSHAllowed defaults to false for a new profile and to true for a
legacy one, and it only works because the skeleton runs first. The doc
comment says so, where before it was implied by the order of two distant
blocks.
Pure refactor. Verified as one: for the four fields whose branches moved,
plus two that did not and the JWT TTL, all 63 combinations of stored value
(nil/false/true) against input value (absent/false/true) produce byte-
identical resolved values and `updated` verdicts before and after.
* [client] Resolve the merge conflicts left in the tree
|
||
|
|
d622e03d40 |
[misc] Use ASCII hyphens in the LICENSE header (#8065)
The first line of LICENSE spelled BSD-3-Clause with non-breaking hyphens (U+2011). The Windows NSIS installer shows LICENSE on its license page and reads it in the system ANSI code page, so the UTF-8 bytes rendered as "BSD‑3‑Clause". Plain hyphens also match the SPDX identifier and keep the file ASCII-only. |
||
|
|
ab79aebd88 |
[misc] Share one license collection script across the UBI images (#8055)
* [self-hosted] Add a UBI image variant for the combined server OpenShift and other Red Hat environments expect UBI-based images that run as an arbitrary non-root UID. The proxy and rootless client already ship -ubi variants; this adds the same for netbird-server, published as <version>-ubi and ubi-latest for amd64 and arm64. * [self-hosted] Check the license output path before creating temp files The existing-output exit ran before the cleanup trap was registered, so it left the two mktemp files behind. * [self-hosted] Certify the netbird-server UBI image Adds netbird-server to the Red Hat certification components. Its Partner Connect component ID goes in the REDHAT_CERT_ID_NETBIRD_SERVER repository variable. * [misc] Share one license collection script across the UBI images The client, proxy and combined images each carried a near-identical copy of collect-licenses.sh, and the signal and relay variants would add two more. The copies differed only in the Go package, build tags, component license and the proxy's web licenses, which are now options of one script in release_files/. The client gains the staged write the others already had. |
||
|
|
d6340ba0de |
[self-hosted] Add a UBI image variant for the combined server (#7953)
* [self-hosted] Add a UBI image variant for the combined server OpenShift and other Red Hat environments expect UBI-based images that run as an arbitrary non-root UID. The proxy and rootless client already ship -ubi variants; this adds the same for netbird-server, published as <version>-ubi and ubi-latest for amd64 and arm64. * [self-hosted] Check the license output path before creating temp files The existing-output exit ran before the cleanup trap was registered, so it left the two mktemp files behind. * [self-hosted] Certify the netbird-server UBI image Adds netbird-server to the Red Hat certification components. Its Partner Connect component ID goes in the REDHAT_CERT_ID_NETBIRD_SERVER repository variable. |
||
|
|
eb5a98c059 | [management] add tenant delete endpoint to openapi (#8054) | ||
|
|
2b5293687f |
[client] Skip late session warnings on desktop and schedule them in the app on Android (#7548)
* Skip session warnings that fire after their window The warning timers run on the monotonic clock, which does not advance while an Android device is suspended. A timer armed for T-10 or T-2 can therefore fire long after the window it was armed for, delivering a "session expires soon" notification once that window is already gone. Gate both callbacks on the wall clock at fire time: the T-10 warning is skipped once the final-warning window has been reached, and the final warning is skipped once the deadline itself has passed. Both set their edge guard before returning so a skipped warning cannot fire again for the same deadline. * Harden the late-warning guards Clamp a non-positive final lead to zero in the T-10 guard so a disabled final warning cannot move the cutoff past the deadline, matching how armTimerLocked already treats it. Strip the monotonic reading from both sides of the comparison so the guard measures wall-clock time regardless of how the caller built the deadline. The production deadline comes from a protobuf timestamp and has no monotonic reading; this keeps the guard correct for callers that derive one from time.Now. * Log the deadline and lateness on skipped warnings Include the deadline and how far past the cutoff the timer fired, so a debug bundle shows how long the device was suspended. * Inject the clock into the late-warning guard and cover it with tests The guard read time.Now internally, so the skip paths were reachable only through a deadline already in the past and the boundary depended on real time. Extract the comparison into isLate and read the time through a nowFn field, so tests can place a resume anywhere around the deadline without sleeping. * Send the final warning when the T-10 timer fires inside its window A suspend between roughly eight and ten minutes long made the T-10 timer fire inside the final-warning window and the final timer fire after the deadline, so both were skipped and a user who resumed with time left got no warning at all. When the T-10 timer fires late but before the deadline, send the final warning in its place and mark it fired so the delayed final timer does not repeat it. * Respect dismissal when promoting a late warning to the final one fireFinal skips the final warning once the user dismissed the deadline, but the promoted path did not, so a dismissed deadline could still get a final warning. Check the dismissal first, and give each skip reason its own log line so an already-fired final warning no longer logs a negative lateness. * Add a deadline-only mode to the session watcher Android will schedule its own expiry warnings from the deadline, so the engine must not arm the T-10 and T-2 timers there. NewDeadlineOnly keeps the deadline validation, the recorder propagation and the logging, and skips only the timers, so the status snapshot the app reads stays correct and an out-of-range deadline is still rejected. * Use the deadline-only watcher on Android and drop the warning callbacks The warning timers run on the monotonic clock, which does not advance while the device sleeps, so a warning armed for T-10 could fire long after its window. The app now schedules the warnings itself with WorkManager, anchored to the wall clock, from the deadline it reads through SessionExpiresAtUnix on every OnStateChanged. Wire the deadline-only watcher into the android build and remove the event-driven path from the gomobile surface: OnSessionExpiring, the event subscription behind it and DismissSessionWarning, which the app never called. * Describe the late-warning guard without naming Android The guard stays for the desktop builds, where a timer can also stall across a sleep. Android no longer arms the timers at all. |
||
|
|
f175e402c7 |
[client] stop offering to every peer when the relay transport drops (#7092)
* [client] stop offering to every peer when the relay transport drops The relay transport is shared: one connection per relay server carries the streams of every peer using it. When it drops, each of those peers gets a Disconnected verdict from evalConnStatus even when ICE is still carrying its traffic, because peerUsesRelay comes from HasRelayAddress(), which only reports that management offered relay servers, not that we are connected to one. The guard answers Disconnected with the aggressive retry, so every peer starts sending offers over signal for a transport that no offer can restore: the relay client's own guard is what reconnects it. Feed relayManager.Ready() into the status inputs and return PartiallyConnected when ICE is up and the missing side is the shared transport. That is the existing "one path works, the other does not" branch, which retries three times and then hourly instead of walking the exponential ladder forever. Peers are not left waiting for the hourly tick: when the transport comes back, Manager.onServerConnected notifies srWatcher, the guard resets the ticker to 800ms and iceState.reset() clears the hourly mode. The verdict is unchanged when the transport is up but this peer is unreachable over relay - it may have moved to another server, and only an offer carries its new relay address - and in force-relay mode, where relay is the only transport. * Renaming according to actual meanings * Don't consider an in progress ICE as "partially connected" when the relay is not.. * Aligns tests * Address wrong comments |
||
|
|
ad03081e1f |
[management] Refresh only affected peers on IPv6 settings changes (#8051)
Account settings changes refreshed every peer, even an IPv6 group toggle that re-addresses a few, and the refresh goroutine got the request context, so it could be cancelled when the handler returned. Group paths that reconcile IPv6 addresses only walked the changed group, so peers reaching a re-addressed peer through its other groups missed the new address. The IPv6 reconcile now returns the peers whose address changed and callers pass them as changed peers. An IPv6-only settings change dispatches affected peers, adding every IPv6 holder on a range change since the interface prefix comes from the range. IPv4 range and account-wide changes keep the full refresh with a detached context. |
||
|
|
bc44cdc37a |
[client] Fix browser login popup show from go (#7408)
* [client] Show the SSO login popup and open the browser from Go The browser-login popup was created hidden and relied on its own webview to size and show itself and to launch the external browser. On macOS a hidden WKWebView gets throttled or suspended (App Nap / hidden-window throttling), so on the first-use path nothing appeared and the browser never opened, leaving the session-expiration dialog disabled until the PKCE flow timed out. Reproduced by freezing the popup's WebContent process: the old code showed nothing, the new code shows the popup and opens the browser within 30 ms regardless of the webview state. Show and focus the popup from Go right after creation and launch the browser from Go on both the create and reuse paths. The popup's frontend no longer shows or focuses itself, so the browser keeps the foreground once it activates. This also fixes the reuse path, where a fragment-only SetURL kept the mounted React tree and the once-only guard skipped opening the browser for the new URI. Browser launch failures surface in the error dialog instead of being swallowed. * [client] Show every dialog window from Go once its frontend has painted Dialog windows (browser-login, session-expiration, install-progress, welcome, error) were created hidden and made visible only by their own webview's Show call after sizing. A hidden WKWebView on macOS can be throttled or suspended before that code runs, which left the window hidden forever. The main and settings windows already avoided this with the painted event plus a fallback timer, but that timer was armed on WindowRuntimeReady, which a frozen webview never reaches either. Route all dialogs through the same mechanism: the auto-size hook emits the painted event instead of showing the window, Go shows and focuses it on that event, and a fallback timer armed at creation shows it after 3 s regardless. The browser-login popup opens the browser in an after-show callback so the browser still lands in front of the popup, also on the fallback path. * [client] Tie install-progress hidden-window restore to the current popup CloseInstallProgress nils s.installProgress before calling w.Close(), so a replacement popup can open before the old window's WindowClosing event runs. The old callback then restored the windows the replacement had just hidden, because the restore sat outside the identity check. Guard the restore with the same check the state reset uses, and restore from CloseInstallProgress itself so the programmatic close path still re-shows the hidden windows — mirroring how CloseBrowserLogin already handles it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * [client] Correlate painted reports with the window generation that sent them A painted report carried only the window name, so a late report from a popup that was already closed and replaced marked its replacement ready. The replacement was then shown before its own frontend had rendered, which is the blank-dialog case this flow exists to prevent. Each dialog start URL now carries a monotonic generation token, echoed back by ReadySignal, and a report whose token no longer matches the live window is dropped. * [client] Separate a window being painted from its frontend being mounted One flag gated both showing a window and emitting to it, so the fallback timer set it for a frontend that had not subscribed yet: the queued events were flushed into a window that could not hear them, losing the login trigger and the settings tab selection. Showing is now gated on painted and emitting on mounted, and only a real frontend report sets mounted. The fallback timer also moved to its own helper so the runtime-ready hook can rearm it, giving the frontend a full budget to mount rather than sharing one with webview boot. * [client] Tag hidden windows with the popup that hid them Windows hidden while a popup owned the screen went into one untagged list, so whichever popup closed first restored all of them and emptied the list. An install started during SSO login re-showed the main window the login popup had deliberately hidden, and left the login popup with nothing to restore. Each entry now records the popup that hid it, and a restore releases only that popup's own entries. This also subsumes the manual filtering CloseRenewFlow did to keep its own session-expiration window from being re-shown. * [client] Cover the hidden-window bookkeeping with tests application.Window carries unexported methods, so the hide/restore paths could not be faked and the earlier tests could only assert which entries survived a restore, never which windows were actually shown. The bookkeeping now goes through hideableWindow, the four methods it needs, with the window enumeration and the main-window raise behind seams that are nil in production. That makes the case the owner tag exists for testable end to end: an install started during SSO login restores only the login popup it hid, and leaves the main window hidden until the login popup itself closes. * [client] Report the first paint from unstamped windows too The main and settings windows carry no generation token, so ReadySignal saw an empty generation that already matched the ref's initial value and never emitted the painted event. Those windows only became visible through the fallback timer, and their frontend was never marked mounted, so the login trigger and the requested settings tab stayed queued. Start the ref from null so the first report goes out regardless of the generation value. * [client] Hand covered windows over when a popup closes under another Closing the browser-login popup while the install-progress popup was still up restored the main window the login had hidden, even though the install popup was meant to own the screen until it finished. The owner tag on each hidden entry only stops a popup from restoring another's windows; it says nothing about what to do with its own when a second popup still covers them. Track which popups currently own the screen and, on restore, re-tag the entries another live popup covers to that popup instead of showing them. A popup is never handed its own window, so closing the popup on top still brings the one below back. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6b3cfbabd2 |
[client] Resolve the Android network route peer by HA unique ID instead of scanning the full status (#7705)
* [client] Build the Android network list from a single peer snapshot Networks() called GetFullStatus() once per network to find the peer that serves it, copying every peer state and taking eight recorder locks each time. With 100+ peers and the UI calling Networks() from every peer list change, this queued hundreds of callers on the status recorder lock. Take one snapshot per call and index it by route. * [client] Track the active route peer by HA unique ID in the status recorder The Android network list resolved the route owner by scanning peer state route keys. Those keys are the handler string: a prefix for static routes and the domain pattern for dynamic ones, so the prefix-based lookup never matched dynamic routes, and two networks sharing a prefix resolved to the same owner. The route watcher now records the chosen route peer under the route's HA unique ID in the status recorder, and the Android binding looks the owner up by that ID. The key is unique per network and independent of the handler string format, so both anomalies are gone. The prefix-based routeOwners helper is removed. * [client] Record the active route peer before notifying listeners AddPeerStateRoute and RemovePeerStateRoute fire the peer list change callback and wake the status subscribers. The active route peer mapping was written after those calls, so a Networks() call landing in between found no mapping for the network and fell back to the first connected peer, or kept showing the previous peer on removal. Nothing re-notified after the mapping write, so the wrong peer stayed until the next peer list change. Write and delete the mapping before the notifying calls so a listener reacting to the notification always reads the current owner. |
||
|
|
1c7d87d5fc |
[client,android] Generate debug bundle to file (#7528)
* [client] Add a debug bundle file export to the Android bridge The Android app can only upload a debug bundle and hand the user a key. Users who want to inspect what leaves their device before sharing it have no way to get the zip itself. Add DebugBundleFile, which generates the bundle into the cache directory and returns its path instead of uploading; the app copies it wherever the user chose and removes it. DebugBundle keeps its behavior. Both entry points share the unexported debugBundle with an upload switch, so the body stays where it was and merges cleanly with the MDM overlay change on main. Because the file variant leaves the zip to the caller and the upload variant only removes it after the upload finishes, a process killed in between leaves a zip behind in the cache. Remove stale bundles before generating a new one: RemoveStaleBundles deletes zips matching the generator's pattern that are older than an hour. Remote debug jobs write to the same directory, so younger files are treated as still in use. * Preserve network map for debug bundle on Android * [client] Keep exported Android debug bundles out of the stale cleanup DebugBundleFile hands the zip to the caller, but the file kept the netbird.debug.*.zip name that RemoveStaleBundles matches, so a later debug run could delete it once it was older than an hour. Rename the exported bundle to netbird.debug-file.*.zip after generation so the cleanup only ever touches bundles no caller owns. * [client] Warn when a stale debug bundle cannot be removed A failed removal means bundles pile up in the cache directory, so log it at Warn instead of Debug. A file that is already gone was removed by a concurrent cleanup and is skipped silently. * [client] Drop the outdated debugBundle comment The comment still said the file variant leaves the zip in place, but it is renamed by debug.ExportBundle since the stale-cleanup change. * [client] Test that the network map reaches the debug bundle Cover both halves of the path Android now relies on: the engine keeps the latest sync response once persistence is enabled, and the bundle generator writes it to network_map.json (anonymized or not) and omits the file when there is no sync response. * [client] Remove abandoned exported debug bundles after a day An exported bundle is owned by the caller, but if the app is killed before it copies and deletes the file, nothing ever removes it from the cache directory. Let RemoveStaleBundles also match exported bundles, with a 24 hour max age instead of the caller-provided one, so a bundle that is still being saved survives while an abandoned one goes. |
||
|
|
19c54b8226 |
[management] Refresh only affected peers on DNS zone and record changes (#8050)
Zone and record changes refreshed every peer in the account, and zone create/update passed the request context to the update goroutine, so it could be cancelled when the handler returned. They now compute affected peers from the zone's distribution groups inside the transaction and dispatch through ExpandAndUpdateAffected, which detaches the context. The resolver did not know about zones, so changing a group referenced only by a zone never pushed the zone to its added or removed members. It now folds the distribution groups of shipped zones on whole-group changes. |
||
|
|
0cd27ca14b | [management] Improve Base62 encoding/decoding performance and robustness (#3391) | ||
|
|
7b8fa29031 | [self-hosted] Replace "which" dependency by "command" from configure.sh script (#8007) | ||
|
|
88b26bb74f |
[infrastructure] Create the preflight artifacts directory before submitting (#7947)
* [infrastructure] Create the preflight artifacts directory before submitting * [infrastructure] Add a workflow to certify UBI images on demand Move the Red Hat certification job into redhat-certify.yml so it can be run by hand for any released version and component, or for all of them. release.yml calls it with component "all" on stable tags. Component IDs now come from REDHAT_CERT_ID_<COMPONENT> repository variables. With "all", components without a variable are skipped. * [infrastructure] Fail Red Hat certification on missing IDs or timeout Fail before certifying when a selected component's REDHAT_CERT_ID_* variable is missing, listing every missing variable. Filter the Pyxis poll by tag so older versions are found past the first page, and fail the job when both architectures are not certified within 10 minutes. |
||
|
|
3c4358dd36 |
[management] Require a private proxy cluster for cluster and direct upstream targets (#7984)
* [management] Require a private proxy cluster for cluster and direct upstream targets Cluster targets and direct upstream targets make the proxy dial the upstream from its own host network instead of through the embedded NetBird client. Only clusters running in private mode are meant to do that, but the service API accepted these targets on any cluster. Service create and update now reject such targets unless the service's proxy cluster reports the private capability. An unreported capability is treated as unsupported. * [management] Require every proxy in the cluster to be private The private capability is aggregated as any-true, so a cluster where only one proxy runs in private mode passed the check. The mapping is delivered to every proxy in the cluster, so the non-private ones would serve cluster and direct upstream targets from their host network too. Validate these targets against a unanimous aggregation instead. The existing any-true lookup stays as is for the dashboard flags and the agent network gateway. |
||
|
|
9f8ddc7131 |
[client] Discover interfaces lazily in stdnet instead of at construction (#7346)
* [client] Discover interfaces lazily in stdnet instead of at construction
stdnet.NewNet and NewNetWithDiscover ended with
return n, n.UpdateInterfaces()
handing back a non-nil *Net together with the discovery error. Three of the
five call sites (Engine.newWgIface, ice.NewAgent, SingleSocketUDPMux) logged
the error and kept using the instance, which is only safe as long as the
instance still works after a failed discovery.
That stopped being true when Interfaces() gained a lazily refreshed cache:
updateInterfaces sets lastUpdate only on success, so after a failed
construction the 30s cache guard never holds and Interfaces() returns an
error rather than the empty list it used to return. Feeding such an instance
to pion is worse than passing nothing at all - ice.NewAgent falls back to its
own stdnet when Net is nil, and the interface blacklist is applied separately
through AgentConfig.InterfaceFilter, so the fallback loses nothing. Instead,
a transient discovery failure (the Android bridge at boot, or an interface
disappearing between net.Interfaces() and Interface.Addrs()) turned into a
hard "error getting local interfaces" from ice.NewAgent, and aborted the STUN
and TURN probes, which never even need the interface list.
Since the accessors already refresh a stale cache on demand, the eager
discovery in the constructors is redundant: drop it, make both constructors
infallible, and let the discovery error surface at the call that actually
needs the interfaces. UpdateInterfaces had no callers left and is not part of
transport.Net, so it is removed along with it.
InterfaceByIndex and InterfaceByName read the cached slice directly and never
refreshed it, so they would have kept reporting ErrInterfaceNotFound forever
on an instance whose first discovery failed. They now go through the same
refresh path as Interfaces().
* [client] Warm the stdnet interface cache at construction
Moving discovery to first use regressed the privileged suites on the three
platforms that always build an ICE bind: Darwin, FreeBSD and Windows time out
in TestWGIface_UpdateAddr, TestRecreation, TestEngine_SSH and
TestEngine_MultiplePeers, while Linux stays green because a host with the
WireGuard kernel module takes the kernel-device branch and never drives the
mux that asks for interfaces.
interfaceFilter probes with wgctrl every interface the disallow list does not
already exclude. Discovering at construction ran that probe before the caller
had an overlay interface of its own; discovering at first use runs it after,
so on a userspace WireGuard platform the probe reaches the UAPI socket of the
same process. The tests reach it because they construct with a nil disallow
list, where the client passes DefaultInterfaceBlacklist and its own interface
is excluded by prefix.
Restore the original timing with an explicit warm-up. The constructors stay
infallible and the error is still reported by the accessor that needs the
interfaces, so the contract this branch is about is unchanged.
* Revert "[client] Warm the stdnet interface cache at construction"
This reverts commit
|
||
|
|
1b89880e30 |
[misc] Move the FreeBSD port test to release 15.1 (#7999)
* [misc] Move the FreeBSD port test to release 15.1 FreeBSD 15.0 reached end of life on 2026-09-30 and the ports tree marks it unsupported since freebsd/freebsd-ports@ed90b23fe9 (2026-10-01), so `make package` refuses to run on the 15.0 VM and the FreeBSD Port job fails on every PR. The pinned vmactions/freebsd-vm v1.4.8 ships a 15.1 image, so only the release needs to move. * [misc] Run the FreeBSD unit tests on release 15.1 too The job installs binary packages instead of building from the ports tree, so it kept passing on the EOL 15.0 image, but the client should be tested on the same supported release the port is built on, and the EOL image is only served from the archive mirror from now on. |
||
|
|
e2678d4e05 | [management] expose peer MAC addresses and make peers searchable by MAC (#6553) | ||
|
|
0712a5a5b9 | [management,proxy] Rename the OIDC session code query parameter (#7981) | ||
|
|
f400f4bee8 |
[proxy] Optionally refuse private addresses on direct-upstream dials (#7913)
Direct-upstream targets are dialled on the proxy host's network stack, outside the embedded client's LAN blocking. A proxy that serves untrusted accounts lets them reach the host's loopback, its LAN or cluster, and the cloud metadata service through such a target. NB_PROXY_DIRECT_UPSTREAM_BLOCK_PRIVATE adds a dialer control that refuses addresses that are not globally reachable. It checks each socket's resolved address just before connect, so hostnames and DNS rebinding are covered, and IPv4 embedded in IPv6 addresses is checked as IPv4. Refused dials are served as a 502. The setting defaults to off for private and self-hosted proxies; an unparsable value turns it on. |
||
|
|
5ceca6e500 |
[client] Report both peers' state when the connect test times out (#7944)
Test_ConnectPeers fails every few weeks on the Linux runner with a bare "waiting for peer handshake timeout after 30s". The failing logs show both kernel devices up and both peers configured within a second, then nothing for 30 s, which is six retries of the 5 s handshake retransmit and so a condition that lasted the whole window rather than a race. The failure cannot be reproduced locally and the log cannot tell whether initiations were sent, whether they arrived, or whether only one direction worked. On timeout the test now prints each device's view of its peer, the endpoint, the byte counters and the last handshake, so the next failure says which of those it is. The comment also states that the peers are kernel devices on the runner and that the first initiation of each side is always lost to the other side not knowing the peer yet. |
||
|
|
3906295446 |
[client] Add Homebrew cask e2e test (#7618)
* [client] Migrate macOS cask template to Homebrew install steps
Homebrew deprecated the postflight and uninstall_preflight cask stanzas
in favour of the declarative *_steps DSL, so every brew command that
evaluates netbirdio/tap now prints deprecation warnings. Once the
deprecation becomes a disable the generated cask stops loading and
netbird-ui can no longer be installed or upgraded through Homebrew.
The *_steps blocks take JSON-serialisable steps run in a sandbox rather
than arbitrary Ruby, so system_command is re-expressed as run/remove.
The two postflight blocks merge into one because a cask carries only a
single instance, preserving the original order. set_permissions moves
from a hardcoded /Applications to base: :appdir, matching what the
installer invocation already did. The launchctl fallbacks keep their
tolerant semantics through must_succeed: false, and remove is a no-op
when the plist is absent.
(cherry picked from commit
|
||
|
|
6425b04200 | [client] Only treat LocalSystem as a privileged identity by SID on Windows (#7889) | ||
|
|
82e5428c2f |
[management] Let usage_viewer read Agent Network access logs (#7750)
usage_viewer saw account-wide usage but only its own request logs, so the people reviewing cost could not drill into the requests behind it. The role now also holds Read on agent_network.logs, which makes the access-log and session endpoints return every caller's rows instead of self-scoping. Logs can contain captured prompts, so this widens what the role exposes; policies, guardrails, budgets and settings stay hidden. Co-authored-by: Misha Bragin <bangvalo@gmail.com> |
||
|
|
fd1a0203c7 |
[management] Clean up after account deletion (#7812)
Deleting an account left state behind that DeleteAccount's store associations don't reach. The Agent Network tables outlived the account, keeping its gateway domain claimed and its provider API keys stored. The proxies kept serving its gateway until they next resynced. Cloud-side state, such as managed proxy deployments, had no way to be cleaned up at all. Account deletion now runs registered hooks after the permission check and before any users or data are removed. A failing hook aborts the deletion. Agent Network registers one that tells the proxies to drop the account's gateway mappings. The account's settings, providers, policies, guardrails and budget rules are deleted in the account's transaction. Consumption counters, and the access logs of deleted accounts, are left to the background cleanup; usage records are kept. |
||
|
|
6c453a0f97 |
[client] Add debug cpu start and stop commands (#7749)
* Add debug cpu start and stop commands to profile the daemon without a restart * Restore test globals on every exit and stop the daemon in the cpu profile test * Add a no-updown flag to debug for * Enable sync response persistence with --no-updown and reset flags between debug test runs * Reset flags of every command between debug test runs * Reset slice flags with Replace in the debug test helper * Explain a running CPU profile in debug for and document cpu start and no-updown limits |
||
|
|
0dc729c4ea | [client] Cache the box shared key per remote peer in the Signal client (#7807) | ||
|
|
e72be6698f |
[client] Keep the advertised ICE session ID when following a remote restart (#7814)
A worker that saw a new remote session ID rebuilt its agent and also picked a new local ID. On the answer path nothing carries that ID back, so the next offer made the remote see a changed session, rebuild, and answer with yet another ID. Two peers kept tearing down working ICE connections on every offer and answer; nearly every answer in the affected logs carried a new remote session ID. Only a local restart changes the local ID now: a failed negotiation, as before, and an explicit Close, which previously kept the old ID and left the remote answering from a negotiation this side had abandoned. Following a remote restart keeps the ID the remote already knows, so the pair settles after one rebuild, also against peers that still pick a new ID when following a restart. |
||
|
|
39de33ceca | [management] handle db conn close on errpr (#7740) | ||
|
|
96bfcc3600 | [client] Classify Windows local accounts by NetBIOS name (#7628) | ||
|
|
8ab34fcf8b |
[proxy] Apply the upstream HTTP version before cloning transports (#7806)
createClientEntry and NewMultiTransport clone the secure transport into its insecure variant before newUpstreamTransport applies the configured HTTP version. http.Transport.Clone runs the source's one-time protocol setup, and at that point ForceAttemptHTTP2 is still false while a custom DialContext is set, so net/http disables HTTP/2 on the source for good. Setting ForceAttemptHTTP2 afterwards has no effect. As a result every TLS-verified upstream has been served over HTTP/1.1 since the upstream HTTP version became configurable, whatever NB_PROXY_UPSTREAM_HTTP_VERSION says, while skip-TLS-verify upstreams kept HTTP/2. gRPC upstreams break outright: unary calls get a 502 and streaming calls hang until the client gives up. Apply the version to the base transport before cloning it, and add a test that the direct and insecure transports both offer h2. |
||
|
|
8edc120370 | [client] Replace the eBPF WireGuard proxy with loopback endpoint addressing (#7316) | ||
|
|
30dd076b36 |
[management,proxy] Use single-use codes for OIDC session handoff (#7635)
* Generalize PKCE verifier store into SingleUseStore * Generalize PKCE verifier store into SingleUseStore * Extend single-use store to generate one-time retrieval codes * Hand off proxy OIDC session via one-time code instead of URL token * Use the single-use store in integration tests * Read active proxy versions by cluster * Detect proxy clusters that support session codes * Bind OIDC session handoff mode to signed state * Deprecate legacy OIDC session token handoff * Remove unrelated session code test stub * fix tests * fix merge * Fix session code compatibility detection * Isolate proxy session codes in shared cache * bump min session version |
||
|
|
7ff709f565 |
[client] Keep the delete-profile dialog open until the delete finishes (#7752)
* [client] Keep the delete-profile dialog open until the delete finishes Confirming a profile deletion closed the dialog straight away and left the daemon call running in the background. A slow delete then looked like nothing had happened: the dialog was gone, the profile was still listed, and the row only disappeared whenever the refresh landed. The confirm dialog now owns the action. It stays open with the confirm button spinning, closes once the call resolves, and surfaces a failure after it is gone rather than behind it. Nothing on the daemon path carries a deadline, so the wait is bounded in the dialog instead: Cancel comes back after five seconds and the wait is abandoned at thirty, which keeps a hung daemon from trapping the user in a modal that cannot be dismissed. A loading button keeps its own variant colours rather than the disabled skin, which dimmed the spinner to grey on the danger variant, and blocks input through aria-disabled and a click guard instead. * [client] Match the theme and anonymize pickers to the language switcher * [client] Settle a confirm dialog only from the run that opened it Cancelling at the stall point leaves the action running, and the provider is mounted once: take() read whichever settler the ref held when the stale run finally finished. Open another prompt in the meantime and that run answered it — a hung delete that later resolved confirmed a profile switch nobody accepted, and its timeout closed the new dialog with an unrelated error. Each run now remembers the settler it was dispatched for and settles only while the ref still points at it. A late arrival finds a stranger there and answers nothing. * [client] Add the shared Select the theme and anonymize pickers use The picker rework landed without the component both pickers import, so the branch did not compile. Add it, and name it for what it is: a select of a few options, with nothing settings-specific about it, so it sits with the other input controls rather than under a name that discourages reuse. * [client] Drop the Select header comment * [client] Fix cubic comments * [client] Name the Select trigger with the option it is showing * [client] Name the language trigger with the language it is showing * [client] Hold the confirm dialog for 15s before offering cancel |
||
|
|
10a04fcccb |
[management] Add a disabled state to the managed Agent Network proxy API (#7744)
A managed Agent Network gateway can be turned off by the platform. The derived state had no value for that, so a disabled deployment reported whatever the operator last saw, usually provisioning. Add `disabled` to the AgentNetworkManagedProxy state enum and say in the POST and GET descriptions that a disabled deployment answers with it and that provisioning again does not turn it back on. |
||
|
|
93cb226a5e | [management] fix login filter (#7739) | ||
|
|
082aca1556 |
[management] Refactor PKCE verifier store into a reusable single-use store (#7634)
* Generalize PKCE verifier store into SingleUseStore * Use the single-use store in integration tests * Remove unrelated session code test stub |
||
|
|
a2919e26dd | [management, proxy] Enforce strict base64url decoding for JWT validation (#7554) | ||
|
|
e830903a94 |
[management] Add a revocation guard hook to the proxy token API (#7732)
Let an embedding binary refuse DELETE /api/reverse-proxies/proxy-tokens/{id}
through an optional proxytoken.RevocationGuard passed to NewAPIHandler.
It is checked after the ownership check, so another account's token
still returns 404 without reaching the guard, and before the token is
revoked. A status error from the guard is written with util.WriteError;
any other error becomes a generic 500.
Nothing installs a guard here, so OSS behavior is unchanged.
|
||
|
|
782c943410 |
[management] extract shared db conn + data repository (#7649)
* extract shared db conn + data repository * extract repository interface * protect against nested transactions * fix mysql and db conn creation * fix context management * remove query warpper * remove withContext and withLock wrapper * remove context from function call * remove in memory mode * fix nested transaction handling * use db directly * remove leftover test * remove pool close on error during conn creation |
||
|
|
002755c412 | [infrastructure] Fix flow auth secret for external Relay migrations (#7731) | ||
|
|
164d92e78d |
[client] Stop dumping the whole device to clear one peer endpoint (#7632)
* [client] Stop dumping the whole device to clear one peer endpoint Clearing a peer's endpoint has to remove and re-add the peer, because neither the netlink API nor the wireguard-go UAPI can clear an endpoint in place. To keep the peer's allowed IPs across that dance, RemoveEndpointAddress read them back from the device: a full wgctrl.Device() dump on the kernel path, a full IpcGet plus text parse on the userspace one. Both cost a round trip proportional to the entire network map, both run under the interface lock, and both run on every relay and ICE transition. On a routing peer with ~15700 peers that is megabytes of netlink traffic per transition, at a measured 713 transitions per minute, with every other configuration operation queued behind it. RemoveAllowedIP paid the same price for the same reason. The allowed IPs cannot come from the caller: peer.Conn knows the peer's own overlay addresses, while the routed prefixes are attached separately by the route manager's refcounter, so a caller-supplied set would silently drop every route behind the peer. The configurer is the only writer of its device's peer set, so it can keep an authoritative mirror of what it configured and answer from memory instead. The mirror is fed by every operation that changes a peer's allowed IPs and reset by a device reconfiguration that replaces the peer set. A peer the mirror has not seen, which is what an out-of-band reconfiguration leaves behind, still falls back to reading the device and seeds the mirror from it. Prefixes are unmapped on the way in, so a v4-mapped address compares equal to the plain v4 prefix for the same network rather than registering as a second entry. Measured on a userspace device, allocations to clear one endpoint: peers 64 256 1024 4096 before 1452 - 21617 - after 91 91 91 91 * [client] Keep update-only allowed IP adds out of the peer mirror AddAllowedIP configures the device with update_only, which is a silent no-op when the peer does not exist, so its success says nothing about whether the device took the prefix. Recording it unconditionally let the mirror hold a peer the device had dropped, and RemoveEndpointAddress re-adds a peer without update_only: clearing the endpoint of such a peer recreated it, carrying allowed IPs the device never held. Allowed IPs are unique per device, so the recreated peer takes those prefixes away from the peer that legitimately holds them. This is not a theoretical window. Under lazy connections a routing peer's device entry is torn down and re-created on the idle transition, and a routed prefix re-added during that window is lost exactly because of update_only (#6863). Allowed IP adds now merge only onto a peer the store already knows, which mirrors the device: the operations that can create a peer record it, the update-only ones do not. A peer missing from the store still falls back to reading the device. * [client] Hand a prefix over to its new owner in the peer mirror An allowed IP belongs to exactly one peer: configuring a prefix on a peer takes it away from whichever peer held it before, and the configurer leaves that handover to the device rather than removing the prefix from the previous holder itself, which is what UpdatePeer's "wg will handle duplicated peer IP" refers to. The mirror recorded the prefix on the new peer while leaving it listed under the old one, so clearing the old peer's endpoint rewrote its allowed IPs from that stale list and took the prefix back from the peer that now owns it. Traffic for the routed prefix then went to the wrong peer. Reading the device before each write used to rule this out. The store now tracks the owner of each prefix and performs the same handover, so rewriting one peer's list cannot reclaim a prefix another peer holds. Prefixes are also masked on the way in. A device stores them masked, so a caller passing host bits would otherwise fail to match what a device fallback seeded and could never remove that prefix by value. Conversion back from the device now keys the v4-mapped decision on the mask width as well, so a genuine v6 prefix inside the mapped range stays v6 instead of being dropped as an invalid v4 prefix. * [client] Keep a mapped v6 prefix below /96 out of the v4 form normalizePrefix unmapped any v4-mapped address before masking it, keeping the original prefix length. For a genuine v6 prefix inside the mapped range, such as ::ffff:0:0/64, that pairs a v4 address with a v6 sized mask: netip.PrefixFrom returns an invalid prefix and Masked turns it into the zero prefix. The store then held a prefix whose Bits is -1, which cannot reproduce the allowed IP the device was given, so re-adding the peer after an endpoint removal could fail once the peer had already been removed. Masking now comes first, and it also decides the address family: only a prefix at least 96 bits long keeps the mapped marker through the mask, so anything shorter inside that range is v6 and stays v6. * [client] Record a peer created by a preshared key write Setting a preshared key without updateOnly creates the peer when it is absent, and Rosenpass applies a peer's first key exactly that way, since applyKeyLocked passes the peer's initialized flag. The store ignored that operation, so the peer could exist on the device while the store treated it as unknown. An update-only allowed IP add on such a peer then succeeded on the device, which moved the prefix away from its previous holder, while the store skipped the peer and left the previous holder still claiming it. Clearing that holder's endpoint rewrote it from the stale claim and took the prefix back, leaving the peer that owns the route with nothing. Every device operation that can create a peer now records it, which is the same rule the update-only operations already follow from the other side. * [client] Match a peer on the parsed key instead of its base64 form getPeer scanned the device comparing Key.String to the caller's key. wgtypes.Key is a 32 byte array, so it compares directly, while String base64 encodes it into a fresh allocation on every iteration. The scan therefore allocated once per peer on the device to find a single peer, and on a large network that is tens of thousands of allocations per lookup. The key is parsed once up front and the arrays are compared. Behaviour is unchanged: the callers already parse the same key before reaching here, so the new parse error is unreachable in practice and only guards the helper on its own. * [client] Normalize prefixes on their way to the device Prefixes were normalized when recorded but not when written, so a caller's raw prefix reached the device while a different form was kept for it. The conversion is also where a mapped prefix goes wrong: net.IPNet prints a v4-mapped address as v4 but takes the length from its 16 byte mask, so ::ffff:10.1.2.3/64 is handed to a userspace device as 10.1.2.3/0 — an allowed IP matching every v4 address, on a peer that was meant to carry one /64. prefixesToIPNets now normalizes, and the two hand-built conversions in AddAllowedIP go through it, so there is a single place where a prefix is turned into something a device is given and it cannot disagree with what is recorded for it. * [client] Parse the endpoint before configuring the peer The userspace UpdatePeer parsed the endpoint address after the device had already been configured, and returned on a parse failure. The device was then left holding a peer that neither the activity recorder nor the allowed IP store had been told about, so the peer was invisible to the wake path and the prefix handover for its allowed IPs never happened, leaving the previous holder still claiming them. The parse now happens before anything is written, so the only failure left after the device is touched is one the caller cannot cause. * [client] Keep the record when a peer removal fails The two configurers disagreed: the kernel one dropped its record only once the device had accepted the removal, the userspace one dropped it either way. Removing a peer is a single device write, so a failure leaves the peer exactly as it was, with the allowed IPs the record still describes. Dropping it there asserts nothing useful and only sends the next caller to read the whole device back for an answer it already had. The userspace one now follows the kernel and returns early on failure. * [client] Write down what the allowed IP store does not guarantee Two properties were relied on without being stated. The store's lock covers its map and not the device write beside it, so consistency between the two rests on callers being serialized, which WGIface does with its mutex; anyone removing that would have no way to learn it mattered. And the fallback to the device only covers a peer the store has never seen, so a peer first recorded from empty while the device already held prefixes keeps only what was recorded, and the next endpoint removal drops the rest. * [client] Key the allowed IP store on the parsed peer key The store keyed on the textual key, so a lookup compared 44 byte strings while the callers all held the parsed key already and the configurer had to carry both forms. wgtypes.Key is a 32 byte array and compares directly, which is what getPeer was changed to do for the same reason. The store and its helpers now take wgtypes.Key, the callers pass the key they parsed on entry, and the textual form survives only where something outside speaks it: parseStatus reports peers that way, so the userspace fallback converts once for its scan. * [client] Document the configurer methods the store changed The exported configurer methods now carry what the allowed IP store made true of them: when the mirror is reset, that a peer update merges its prefixes and takes them from their previous owner, that an update-only add on an absent peer does nothing, and what each side does with its record when a device write fails — where the two configurers differ, since the userspace one reports a prefix it does not have and the kernel one treats it as a no-op. mergeLocked states the lock its callers must already hold. Docstrings that only restated the name of a test are left out; the tests explain the scenario they set up in the body, where the explanation belongs. |
||
|
|
a8162aaac5 |
[client] Bump wails to v3.0.0-beta.25 (#7640)
* [client] Bump wails to v3.0.0-beta.25 * [client] Point wails to the merged integration commitv0.80.0-rc.1 |
||
|
|
5d92c89227 | [management] move rate limiter to shared package (#7727) | ||
|
|
979571a99f |
[management] Prevent deleting custom domains used by services (#7515)
Deleting a custom domain released its name while services still pointed at it, leaving them on a namespace the account no longer held. Deletion now refuses with 412 when a service in the same account uses the domain or a subdomain, including disabled ones. Service writes revalidate authorization inside their transaction and hold a shared lock on the matching registrations, so a delete racing a create cannot strand either. The dependency lookup is account-scoped: registrations are unique by name, so another account can hold team.example.com under example.com and its services are authorized by its own registration. |
||
|
|
c6aa6c232e |
[doc] Point bug reports at Discussions and add SUPPORT.md (#7647)
* add readme section and support file * [doc] Document anonymize levels in README and SUPPORT Reviewer feedback: mention both --anonymize-level values, not just -A. Wording follows the flag help in client/cmd/root.go. * Apply suggestion from @cubic-dev-ai[bot] * Update SUPPORT.md * adjust parameter descriptions * [doc] Fix duplicated sentence in SUPPORT.md The cubic suggestion replaced only part of the -U sentence, leaving the original line orphaned above it and dropping the blank line before the docs links. * [doc] Tighten anonymization claims in README and SUPPORT Strict mode keeps labels under netbird.io, so naming "the NetBird domains" overstated what it masks. Name the three peer domains instead. Anonymization is not full redaction: internal ranges survive at the default level and interface details are never masked, so say that rather than implying the bundle is safe to post unread. |
||
|
|
ba8bdfc5c7 | [infrastructure] Add proxy support to enterprise setup (#7651) | ||
|
|
2f27051439 |
[proxy] Add a release-wired UBI image variant (#7464)
Add a UBI-based reverse-proxy image for the internal Red Hat certification requirement, without changing the existing image or deployment defaults. Includes non-root execution, licensing and image metadata, plus an AMD64 GoReleaser entry with separate UBI tags. Preflight and TLS/overlay checks passed. An intermittent shutdown exit error remains deferred; this stays draft pending maintainer testing. |