* [client] Force interactive login when extending the auth session
A session extend must be answered from the account the peer is registered
under. With a silent PKCE flow (DisablePromptLogin or max_age=0) the IdP
answers from whatever session it already holds, which need not be the
peer's account when several are signed in; the token then fails the
user match in ExtendAuthSession with no way to pick another account.
Mark the PKCE flow request as a session extend so the management server
can force prompt=login for it, overriding the configured silent flow.
* [client] Reduce cognitive complexity of Server.Login
Login sat at cognitive complexity 27, over the 25 the linter allows.
Extract the interactive SSO branch into startSSOLogin, and split the
nested in-flight-flow reuse check out of it into reuseOAuthFlow, which
flattens the original if/else into early returns: it returns the cached
auth info when the previous flow targets the same client and still has
more than 90s left, otherwise cancels the stale wait and returns nil so
the caller requests a fresh flow.
The helpers take the contextState through a small statusSetter
interface, since internal.contextState is unexported and re-deriving it
with CtxGetState inside the helper would resolve against callerCtx
rather than rootCtx.
No behavior change: same ordering of state transitions, same mutex scope
around the oauthAuthFlow write, same error paths. Login is now at 21.
* [client] Respect DisablePromptLogin when extending the auth session
Forcing prompt=login on a session extend overrode DisablePromptLogin, which
is set for IdPs that break on it: Authentik triggers a double authentication
and social logins fail outright. Overriding it there trades a recoverable
extend for a login that cannot complete at all.
Keep the LoginFlag override, which only replaces max_age=0 or none with
prompt=login so the IdP honours login_hint, and leave DisablePromptLogin as
configured. Those deployments keep the silent flow, and with several accounts
signed in an extend answered from the wrong one still fails the user match.
* [client] Guard the shared OAuth flow state with the server mutex
reuseOAuthFlow read flow, expiresAt, waitCancel and info without holding
s.mutex, while startSSOLogin and WaitSSOLogin write them under it. Reading the
fields one at a time could also answer with auth info from a flow that was
already replaced, or cancel a wait that no longer belongs to the flow just
judged stale. Take one snapshot under the lock and decide from it.
WaitSSOLogin read oauthAuthFlow.flow twice outside the lock; both now use a
value snapshotted in the critical section that already installs actCancel.
Its stale waitCancel was read and called in a separate section from the one
installing the new one, so two racing calls could read the same predecessor and
leave one wait uncancelled. Swap the two in a single critical section. Both
cancels run after unlocking: the displaced wait takes s.mutex as it unwinds.
* [client] Verify the SSO login came back for the hinted account
login_hint is a suggestion the IdP may ignore: with a silent flow configured
(DisablePromptLogin or max_age=0) and a live IdP session for another account,
the login completes with that account's token. On a registered peer the
management server rejects it as a user mismatch, but on a fresh profile the
peer silently registers under the wrong account and the profile is then bound
to it — every later login follows the stored hint straight back.
After the token exchange, compare the ID token's email against the hint the
flow was sent with. On a mismatch, do not log in to management with the token;
run one more round asking the IdP to re-decide the account (prompt=login, via
ForceAccountPrompt — DisablePromptLogin still wins there). If the prompted
round also comes back different, proceed with a warning: the address may
legitimately have changed, and refusing forever would lock the user out of the
profile while the management server still rejects a token that does not own
the peer. A token or profile with no email to compare is not judged.
The retry differs per platform because of who opens the browser:
- CLI (netbird login foreground) and Android run the whole flow in one
process, so the mismatch retries automatically: the browser reopens with
the account prompt within the same login attempt.
- On desktop the login is split between the daemon and the GUI: Login hands
the authorize URL to the GUI, WaitSSOLogin blocks for the token, and only
the GUI can open a browser. A new URL cannot be handed out from inside
WaitSSOLogin (its response has no field for one, kept that way to avoid a
proto change), so the daemon arms forceAccountPrompt, fails the round with
"connect again to choose the account", and builds the next Login's flow
with the prompt — the user's next connect is the retry.
The flag and the flow annotations live in daemon memory only; SwitchProfile
drops them so the previous profile's hint cannot judge the next profile's
token. The device code flow has no prompt parameter (RFC 8628), so a prompted
round there runs as-is and a repeated mismatch is let through with the
warning rather than looping.
* [client] Address review comments on PKCE session extend flow
Fail the PKCE authorization flow test on request error instead of
continuing into a nil dereference, and make the godoc comments on the
touched exported symbols identifier-leading full sentences.
* [client] Match accounts only on the email claim of the ID token
The name-claim fallback in the ID token parsing is kept for the login
hint and display, but account matching now only considers a value that
came from the email claim, so a token without one no longer produces a
false account mismatch.
* [client] Drop the pending session extend on a profile switch
The profile-switch cleanup dropped the pending login flow and the
account-prompt flag, but left extendAuthSessionFlow untouched. Its device
code was issued by the previous profile's IdP client, so a
WaitExtendAuthSession still parked on the browser leg would submit the
resulting token against the new profile's engine.
* [client] Judge the SSO account against the flow that produced the token
WaitSSOLogin snapshotted the flow on entry but re-read the info, hint and
accountPrompted from the live s.oauthAuthFlow afterwards, in separate
critical sections. WaitToken blocks for the whole browser leg, so a
concurrent Login or RequestJWTAuth could replace the flow meanwhile and
the mismatch check would compare this wait's token against another flow's
account: either arming the prompt spuriously or letting a wrong-account
token through against an unrelated profile's hint. Take all of it in the
entry snapshot.
* [client] Keep the forced account prompt from being lost to flow reuse
startSSOLogin consumed forceAccountPrompt and applied the prompt to the
freshly built flow, but reuseOAuthFlow could then answer from a cached
flow for the same client — one built without prompt=login, e.g. by
RequestJWTAuth. The user got the same silent authorization URL that
produced the mismatch, with the flag already spent, so no later round
asked either. Rule reuse out when the prompt is forced, while still
cancelling the predecessor's wait.
RequestJWTAuth also wrote the flow fields one by one, leaving the previous
login's hint and accountPrompted behind for WaitSSOLogin to judge a later
token against. Both sites now replace the whole record.
* [client] Consume the forced account prompt after the retry
forceAccountPrompt was never cleared, so a flow that outlived the retry it
was armed for kept sending prompt=login on every later authorization
request and re-authenticated the user each time. RequestAuthInfo now takes
the flag as it builds the request.
* [client] Cancel the caller context in the SSO login tests
WaitSSOLogin parks a goroutine on the caller's context for the whole
browser leg. The tests passed context.Background(), which never cancels,
so each left one goroutine behind for the lifetime of the test binary.
* [client] Cancel the wait displaced by an OAuth flow replacement
Replacing the shared record with a whole struct value dropped the previous
flow's waitCancel, so an SSO browser wait still parked on it lost its
cancel: nothing could preempt it, and it could go on to run attemptLogin
or mutate the record behind the new flow. Both replacement sites now take
the displaced cancel over in the same critical section, via a shared
replaceOAuthFlow, and invoke it after the unlock.
* [client] Guard OAuth flow mutations by the flow that owns the wait
* [client] Arm the account prompt only from the wait that owns the flow
* [client] Adopt the three-value parseEmailFromIDToken in the device flow
The main merge brought in the device flow's email extraction from #7193,
which still used the two-value signature this branch replaced when account
matching was narrowed to the email claim. Git merged the files without a
textual conflict, so the branch stopped compiling.
Take the fromEmailClaim result and fill EmailClaim from it, the same way the
PKCE path does, so device-flow clients get the same account matching.
* [client] Populate the pending extend flow in the test server helper
SwitchProfile cancels and clears the pending session extend flow
unconditionally, the same way it clears the SSH JWT cache. New always
populates the field, but the hand-assembled test server did not, so
TestSwitchProfile_ClearsJWTCache panicked on a nil PendingFlow.
* return permission-denied error when setup key is expired or invalid
Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>
* fix tests, don't return permission-denied on internal errors
Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>
* return not found error when key fetch failed
Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>
* normalise errors returned from each call-site of GetSetupKeyBySecret
Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>
* return PermissionDenied after isValid() check failure
Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>
---------
Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io>
Account settings changes refreshed every peer, even an IPv6 group toggle
that re-addresses a few, and the refresh goroutine got the request context,
so it could be cancelled when the handler returned. Group paths that
reconcile IPv6 addresses only walked the changed group, so peers reaching a
re-addressed peer through its other groups missed the new address.
The IPv6 reconcile now returns the peers whose address changed and callers
pass them as changed peers. An IPv6-only settings change dispatches affected
peers, adding every IPv6 holder on a range change since the interface prefix
comes from the range. IPv4 range and account-wide changes keep the full
refresh with a detached context.
Zone and record changes refreshed every peer in the account, and zone
create/update passed the request context to the update goroutine, so it
could be cancelled when the handler returned. They now compute affected
peers from the zone's distribution groups inside the transaction and
dispatch through ExpandAndUpdateAffected, which detaches the context.
The resolver did not know about zones, so changing a group referenced only
by a zone never pushed the zone to its added or removed members. It now
folds the distribution groups of shipped zones on whole-group changes.
* [management] Require a private proxy cluster for cluster and direct upstream targets
Cluster targets and direct upstream targets make the proxy dial the
upstream from its own host network instead of through the embedded
NetBird client. Only clusters running in private mode are meant to do
that, but the service API accepted these targets on any cluster.
Service create and update now reject such targets unless the service's
proxy cluster reports the private capability. An unreported capability
is treated as unsupported.
* [management] Require every proxy in the cluster to be private
The private capability is aggregated as any-true, so a cluster where
only one proxy runs in private mode passed the check. The mapping is
delivered to every proxy in the cluster, so the non-private ones would
serve cluster and direct upstream targets from their host network too.
Validate these targets against a unanimous aggregation instead. The
existing any-true lookup stays as is for the dashboard flags and the
agent network gateway.
usage_viewer saw account-wide usage but only its own request logs, so
the people reviewing cost could not drill into the requests behind it.
The role now also holds Read on agent_network.logs, which makes the
access-log and session endpoints return every caller's rows instead of
self-scoping. Logs can contain captured prompts, so this widens what the
role exposes; policies, guardrails, budgets and settings stay hidden.
Co-authored-by: Misha Bragin <bangvalo@gmail.com>
Deleting an account left state behind that DeleteAccount's store
associations don't reach. The Agent Network tables outlived the account,
keeping its gateway domain claimed and its provider API keys stored. The
proxies kept serving its gateway until they next resynced. Cloud-side
state, such as managed proxy deployments, had no way to be cleaned up at
all.
Account deletion now runs registered hooks after the permission check
and before any users or data are removed. A failing hook aborts the
deletion. Agent Network registers one that tells the proxies to drop the
account's gateway mappings. The account's settings, providers, policies,
guardrails and budget rules are deleted in the account's transaction.
Consumption counters, and the access logs of deleted accounts, are left
to the background cleanup; usage records are kept.
* Generalize PKCE verifier store into SingleUseStore
* Generalize PKCE verifier store into SingleUseStore
* Extend single-use store to generate one-time retrieval codes
* Hand off proxy OIDC session via one-time code instead of URL token
* Use the single-use store in integration tests
* Read active proxy versions by cluster
* Detect proxy clusters that support session codes
* Bind OIDC session handoff mode to signed state
* Deprecate legacy OIDC session token handoff
* Remove unrelated session code test stub
* fix tests
* fix merge
* Fix session code compatibility detection
* Isolate proxy session codes in shared cache
* bump min session version
Let an embedding binary refuse DELETE /api/reverse-proxies/proxy-tokens/{id}
through an optional proxytoken.RevocationGuard passed to NewAPIHandler.
It is checked after the ownership check, so another account's token
still returns 404 without reaching the guard, and before the token is
revoked. A status error from the guard is written with util.WriteError;
any other error becomes a generic 500.
Nothing installs a guard here, so OSS behavior is unchanged.
* extract shared db conn + data repository
* extract repository interface
* protect against nested transactions
* fix mysql and db conn creation
* fix context management
* remove query warpper
* remove withContext and withLock wrapper
* remove context from function call
* remove in memory mode
* fix nested transaction handling
* use db directly
* remove leftover test
* remove pool close on error during conn creation
Deleting a custom domain released its name while services still pointed
at it, leaving them on a namespace the account no longer held.
Deletion now refuses with 412 when a service in the same account uses the
domain or a subdomain, including disabled ones. Service writes revalidate
authorization inside their transaction and hold a shared lock on the
matching registrations, so a delete racing a create cannot strand either.
The dependency lookup is account-scoped: registrations are unique by name,
so another account can hold team.example.com under example.com and its
services are authorized by its own registration.
* [client,management] Skip route firewall rule computation when no firewall
A peer that runs with the firewall disabled has no ACL manager and no
firewall to program, so nothing ever reads RoutesFirewallRules: the only
consumers are acl.Manager, which is reached solely when e.acl is set, and
the legacy-management probe in updateNetworkMap, which is guarded by a
non-nil firewall.
Building those rules is the most expensive part of a sync on a peer that
routes many network resources. On a 15k-peer deployment a debug bundle
showed getPeerNetworkResourceFirewallRules accounting for 62% of the
allocations of Calculate, and Calculate for effectively all of the
allocations of handleSync, which was taking 3.2s on average and holding
the engine lock for the duration.
Let the caller ask Calculate to leave the rules out. The client passes
its existing DisableFirewall setting; the management server keeps the
default and still produces them.
RoutesFirewallRulesIsEmpty is set from the resulting empty list, so a
receiver that would otherwise infer legacy management from an empty rule
set does not misread the skip.
* [client,management] Cover the skip flag through the envelope
Review feedback on #7624.
The components test compared only the length of the peer firewall rules, so
a change to their content would have passed while the message claimed they
came out unchanged. Compare the slices.
The skip path was also only exercised by setting the field directly on the
components, which bypasses the envelope conversion where
RoutesFirewallRulesIsEmpty is derived. That bit is what keeps the client from
reading skipped rules as a legacy management server, so it gets a test that
goes through EnvelopeToNetworkMap with the flag set.
* [management] Give the router a peer ACL so the rule comparison bites
Review feedback on #7624.
peer-router-1 appears in no peer ACL in the shared fixture, so its
FirewallRules came out empty and the equality assertion compared two empty
slices — it would have passed even if the peer rules were dropped entirely.
Add a policy covering the router and require the baseline to be non-empty
before comparing.
* Name the account owner in the pending approval error
A user refused because their account is pending approval had no way to
learn who could approve them. The refusal now carries the account
owner's address, masked, so a caller can name someone to contact without
being handed the address itself.
Resolving the owner is best effort: a lookup failure, or an account
predating the stored email, falls back to the refusal as it was.
* Name only the caller's own owner in the pending approval error
The refusal is raised before ValidateAccountAccess has established that
the caller belongs to the account the request asked about, and the user
is loaded by ID alone. Resolving the owner of the requested account
therefore disclosed that owner's address to a pending user with no claim
to it, reachable through any handler that takes an account ID from the
caller — DELETE /api/accounts/{accountId} passes one straight through.
The owner who can approve a pending user is the owner of their own
account, so resolve that one. The requested account is never read.
* Mask short local parts whole in MaskedEmail
Keeping the first two characters and the last hides nothing until the
local part is four long: at three or fewer they are the whole of it, so
"abc@example.com" masked to "ab****c@example.com" and a pending user
could recover the owner's address in full from what is meant to conceal
it. Short local parts are now replaced entirely.
* Name the owner from GetCurrentUserInfo instead of the permission gate
The gate could only read the stored user row, which carries no address
when an external IdP owns the identities — the usual case — so it named
no one in practice. It also had no way to reach the IdP without being
handed the account manager, which meant restoring bootstrap wiring that
a refactor had dropped.
GetCurrentUserInfo already holds that account manager, so it answers for
a pending user itself and reuses GetOwnerInfo, the same lookup /msp uses
to resolve an owner's address. The gate returns to exactly what it was,
and with it goes the risk of naming the owner of an account the caller
only asked about.
MaskedEmail becomes MaskEmail: with a UserInfo in hand there is no stored
row to hang it off.
* [management] Cover the pending approval refusal in GetCurrentUserInfo
The branch that names the owner had no coverage at the manager level, so
neither the named refusal nor the fallback for an owner without a resolvable
address was pinned down.
* [management] Cover the failed owner lookup in the pending approval refusal
The generic fallback has two ways in: no address on the resolved owner, and no
owner to resolve at all. Only the first was pinned down.
* [management] Pin the owner lookup to the caller's own account
A mismatched account claim must not steer which owner the refusal names, and
a blocked user is still answered before the claim is validated. Both are load
bearing and neither was covered.
The agent network gateway service is private: agents reach it over the WireGuard tunnel, authorised by peer identity, with the cluster as its only target. Only a reverse proxy cluster with private capabilities can serve that, reported per cluster as the `private` capability — the same `supports_private` flag the dashboard gates NetBird-only services on.
A bootstrap could pin an account to a cluster without private capabilities, leaving an immutable dead gateway. Both bootstrap paths now validate the picked cluster: one the account can see must have a connected proxy reporting the capability. Shared and account-owned clusters qualify alike. Known-ness comes from proxy rows, not heartbeat freshness, so a cluster without the capability stays refused while merely offline. A hostname no proxy has declared stays pinnable (address-first). Identity is compared case-insensitively over the account's cluster list.
An agent network bootstrap stores its cluster as proxy_address, which selects the proxy that serves the endpoint. An account-scoped proxy only receives its own account's mappings, so a pin onto a host another account's proxy declares can never be served, and the endpoint is immutable — a dead gateway until the account deletes its settings. Nothing refused that pin; the domain unique index only arbitrates between endpoints.
Both bootstrap paths now refuse, before the insert, a host that another account's proxy declares, a host another account has labeled pins beneath (self-addressed path), or a hostname that is another account's endpoint (labeled path).
Shared clusters are unaffected: shared proxies are never foreign, and labeled pins under one cluster are never asked about, so any number of accounts still pin beneath eu.proxy.netbird.io. Registration is deliberately unchanged — refusing a proxy for another account's pin would let a pin lock a tenant out after the reaper drops its rows.
Management / Unit (amd64, mysql) hit the 20 minute go test budget on #7516. The package was not hung: each of the 133 test store creations in management/server paid about 1.6s on MySQL for CREATE DATABASE, the pre-migrations, a 40-table AutoMigrate and the post-migrations, which puts the package at 10 minutes on a healthy runner and over the budget on a slow one.
The migration now runs once per test binary into a template database and each test database is cloned from it, with CREATE DATABASE ... TEMPLATE on Postgres and a replay of SHOW CREATE TABLE on MySQL. The MySQL test container also drops the binary log, doublewrite buffer and per-commit redo fsync. Two goroutine leaks in the test helpers are fixed.
tools/gotestsummary turns the go test -json stream into a readable log, and the Management unit and integration jobs now pipe through it, so a timeout names the tests still running. On MySQL, management/server went from 10m16s to 6m36s.
`validateDeleteGroup` already refuses to delete a group that is still used by routes, policies, nameservers, setup keys, users, network routers, reverse proxy services, and agent network policies. Account-level agent network budget rules also store group IDs in `TargetGroups`, but that check was missing.
Deleting such a group left a dangling ID on the budget rule. `budgetRuleApplies` then never matched callers by group, so the spend cap silently stopped applying.
This adds `isGroupLinkedToAgentNetworkBudgetRule` and uses it in `validateDeleteGroup`, matching the existing helpers.
Prevent unvalidated registrations from reserving domain names indefinitely.
Give pending registrations a 48-hour validation window and clean up expired entries at startup and every 60 minutes. Emit CustomDomainValidationExpired for each deletion and preserve registrations referenced by services.
Reject validation after expiry and prevent concurrent validation from recreating deleted registrations. Normalize domain names with the shared parser before registration.
Migrate existing pending registrations to receive a fresh 48-hour validation window.
Require validated custom domains when creating or updating reverse proxy services.
Propagate validation errors during updates and return HTTP 409 for duplicate domain claims.
Add regression tests for domain validation, ownership, and service creation and updates.
The management binary is embedded by downstream builds that override
server construction via SetNewServer, but the cobra command tree itself
was closed: rootCmd is unexported and fully assembled in init, with no
way to attach additional subcommands. Customize hands the built root
command to a caller-supplied function before Execute, so an embedding
binary can add its own commands next to — or under — the built-in ones,
such as extra administrative helpers beneath the existing admin group.
* Gather fresh system info on every management sync stream connect
The engine collected the peer meta once at start and reused the same
Info for every Sync stream reconnect, so a mobile network switch that
redials management kept reporting the old local network addresses.
The peer network range posture check was then evaluated against stale
data until the client restarted.
Sync now takes a gatherer that runs at each stream connect. The
gatherer is cheap: GetInfo plus the cached posture check file results,
kept in the new system.InfoSource, which the engine refreshes whenever
the checks list changes. No process enumeration runs on the reconnect
path.
Also fix the management mock server calling itself instead of SyncFunc.
* Evaluate the login response posture checks before the first sync connect
The engine starts with the checks the login response carried, and the
first sync stream request used to send their evaluated file results.
After moving the gather into InfoSource, the stream opened with an empty
cache and the first sync response did not refill it, because its checks
equal the ones the engine already holds. Desktop peers therefore never
reported process or file posture results.
Seed the cache once before the first connect, where the old gather ran,
so a timed out evaluation still falls through to the address-only info.
* Harden the sync info source against nil callbacks and shared slices
A nil getInfo opens the stream without metadata, as a nil sysInfo did
before. The cached posture results are a copy, so the Info returned by
Refresh cannot alias the snapshot later Current calls report. The
exclusion test asserts the remaining address count so it cannot pass
vacuously on a single-address host.
* Retry a posture check refresh that timed out or failed to sync
The checks list was recorded before the gather ran, so once the gather
timed out or SyncMeta failed, the next sync response carrying the same
list matched the recorded one and nothing retried. The peer kept
reporting the previous posture results until the list changed again.
Record the checks only after the meta reached management, so a failed
cycle is repeated on the next sync response.
* Log the skipped posture refresh, let the mock Sync return errors and deflake the reconnect test
* Drop the nil guard around the sync info callback
* Send the refreshed info on the first sync connect instead of gathering it twice