Account settings changes refreshed every peer, even an IPv6 group toggle
that re-addresses a few, and the refresh goroutine got the request context,
so it could be cancelled when the handler returned. Group paths that
reconcile IPv6 addresses only walked the changed group, so peers reaching a
re-addressed peer through its other groups missed the new address.
The IPv6 reconcile now returns the peers whose address changed and callers
pass them as changed peers. An IPv6-only settings change dispatches affected
peers, adding every IPv6 holder on a range change since the interface prefix
comes from the range. IPv4 range and account-wide changes keep the full
refresh with a detached context.
Zone and record changes refreshed every peer in the account, and zone
create/update passed the request context to the update goroutine, so it
could be cancelled when the handler returned. They now compute affected
peers from the zone's distribution groups inside the transaction and
dispatch through ExpandAndUpdateAffected, which detaches the context.
The resolver did not know about zones, so changing a group referenced only
by a zone never pushed the zone to its added or removed members. It now
folds the distribution groups of shipped zones on whole-group changes.
* [management] Require a private proxy cluster for cluster and direct upstream targets
Cluster targets and direct upstream targets make the proxy dial the
upstream from its own host network instead of through the embedded
NetBird client. Only clusters running in private mode are meant to do
that, but the service API accepted these targets on any cluster.
Service create and update now reject such targets unless the service's
proxy cluster reports the private capability. An unreported capability
is treated as unsupported.
* [management] Require every proxy in the cluster to be private
The private capability is aggregated as any-true, so a cluster where
only one proxy runs in private mode passed the check. The mapping is
delivered to every proxy in the cluster, so the non-private ones would
serve cluster and direct upstream targets from their host network too.
Validate these targets against a unanimous aggregation instead. The
existing any-true lookup stays as is for the dashboard flags and the
agent network gateway.
usage_viewer saw account-wide usage but only its own request logs, so
the people reviewing cost could not drill into the requests behind it.
The role now also holds Read on agent_network.logs, which makes the
access-log and session endpoints return every caller's rows instead of
self-scoping. Logs can contain captured prompts, so this widens what the
role exposes; policies, guardrails, budgets and settings stay hidden.
Co-authored-by: Misha Bragin <bangvalo@gmail.com>
Deleting an account left state behind that DeleteAccount's store
associations don't reach. The Agent Network tables outlived the account,
keeping its gateway domain claimed and its provider API keys stored. The
proxies kept serving its gateway until they next resynced. Cloud-side
state, such as managed proxy deployments, had no way to be cleaned up at
all.
Account deletion now runs registered hooks after the permission check
and before any users or data are removed. A failing hook aborts the
deletion. Agent Network registers one that tells the proxies to drop the
account's gateway mappings. The account's settings, providers, policies,
guardrails and budget rules are deleted in the account's transaction.
Consumption counters, and the access logs of deleted accounts, are left
to the background cleanup; usage records are kept.
* Generalize PKCE verifier store into SingleUseStore
* Generalize PKCE verifier store into SingleUseStore
* Extend single-use store to generate one-time retrieval codes
* Hand off proxy OIDC session via one-time code instead of URL token
* Use the single-use store in integration tests
* Read active proxy versions by cluster
* Detect proxy clusters that support session codes
* Bind OIDC session handoff mode to signed state
* Deprecate legacy OIDC session token handoff
* Remove unrelated session code test stub
* fix tests
* fix merge
* Fix session code compatibility detection
* Isolate proxy session codes in shared cache
* bump min session version
Let an embedding binary refuse DELETE /api/reverse-proxies/proxy-tokens/{id}
through an optional proxytoken.RevocationGuard passed to NewAPIHandler.
It is checked after the ownership check, so another account's token
still returns 404 without reaching the guard, and before the token is
revoked. A status error from the guard is written with util.WriteError;
any other error becomes a generic 500.
Nothing installs a guard here, so OSS behavior is unchanged.
* extract shared db conn + data repository
* extract repository interface
* protect against nested transactions
* fix mysql and db conn creation
* fix context management
* remove query warpper
* remove withContext and withLock wrapper
* remove context from function call
* remove in memory mode
* fix nested transaction handling
* use db directly
* remove leftover test
* remove pool close on error during conn creation
Deleting a custom domain released its name while services still pointed
at it, leaving them on a namespace the account no longer held.
Deletion now refuses with 412 when a service in the same account uses the
domain or a subdomain, including disabled ones. Service writes revalidate
authorization inside their transaction and hold a shared lock on the
matching registrations, so a delete racing a create cannot strand either.
The dependency lookup is account-scoped: registrations are unique by name,
so another account can hold team.example.com under example.com and its
services are authorized by its own registration.
* [client,management] Skip route firewall rule computation when no firewall
A peer that runs with the firewall disabled has no ACL manager and no
firewall to program, so nothing ever reads RoutesFirewallRules: the only
consumers are acl.Manager, which is reached solely when e.acl is set, and
the legacy-management probe in updateNetworkMap, which is guarded by a
non-nil firewall.
Building those rules is the most expensive part of a sync on a peer that
routes many network resources. On a 15k-peer deployment a debug bundle
showed getPeerNetworkResourceFirewallRules accounting for 62% of the
allocations of Calculate, and Calculate for effectively all of the
allocations of handleSync, which was taking 3.2s on average and holding
the engine lock for the duration.
Let the caller ask Calculate to leave the rules out. The client passes
its existing DisableFirewall setting; the management server keeps the
default and still produces them.
RoutesFirewallRulesIsEmpty is set from the resulting empty list, so a
receiver that would otherwise infer legacy management from an empty rule
set does not misread the skip.
* [client,management] Cover the skip flag through the envelope
Review feedback on #7624.
The components test compared only the length of the peer firewall rules, so
a change to their content would have passed while the message claimed they
came out unchanged. Compare the slices.
The skip path was also only exercised by setting the field directly on the
components, which bypasses the envelope conversion where
RoutesFirewallRulesIsEmpty is derived. That bit is what keeps the client from
reading skipped rules as a legacy management server, so it gets a test that
goes through EnvelopeToNetworkMap with the flag set.
* [management] Give the router a peer ACL so the rule comparison bites
Review feedback on #7624.
peer-router-1 appears in no peer ACL in the shared fixture, so its
FirewallRules came out empty and the equality assertion compared two empty
slices — it would have passed even if the peer rules were dropped entirely.
Add a policy covering the router and require the baseline to be non-empty
before comparing.
* Name the account owner in the pending approval error
A user refused because their account is pending approval had no way to
learn who could approve them. The refusal now carries the account
owner's address, masked, so a caller can name someone to contact without
being handed the address itself.
Resolving the owner is best effort: a lookup failure, or an account
predating the stored email, falls back to the refusal as it was.
* Name only the caller's own owner in the pending approval error
The refusal is raised before ValidateAccountAccess has established that
the caller belongs to the account the request asked about, and the user
is loaded by ID alone. Resolving the owner of the requested account
therefore disclosed that owner's address to a pending user with no claim
to it, reachable through any handler that takes an account ID from the
caller — DELETE /api/accounts/{accountId} passes one straight through.
The owner who can approve a pending user is the owner of their own
account, so resolve that one. The requested account is never read.
* Mask short local parts whole in MaskedEmail
Keeping the first two characters and the last hides nothing until the
local part is four long: at three or fewer they are the whole of it, so
"abc@example.com" masked to "ab****c@example.com" and a pending user
could recover the owner's address in full from what is meant to conceal
it. Short local parts are now replaced entirely.
* Name the owner from GetCurrentUserInfo instead of the permission gate
The gate could only read the stored user row, which carries no address
when an external IdP owns the identities — the usual case — so it named
no one in practice. It also had no way to reach the IdP without being
handed the account manager, which meant restoring bootstrap wiring that
a refactor had dropped.
GetCurrentUserInfo already holds that account manager, so it answers for
a pending user itself and reuses GetOwnerInfo, the same lookup /msp uses
to resolve an owner's address. The gate returns to exactly what it was,
and with it goes the risk of naming the owner of an account the caller
only asked about.
MaskedEmail becomes MaskEmail: with a UserInfo in hand there is no stored
row to hang it off.
* [management] Cover the pending approval refusal in GetCurrentUserInfo
The branch that names the owner had no coverage at the manager level, so
neither the named refusal nor the fallback for an owner without a resolvable
address was pinned down.
* [management] Cover the failed owner lookup in the pending approval refusal
The generic fallback has two ways in: no address on the resolved owner, and no
owner to resolve at all. Only the first was pinned down.
* [management] Pin the owner lookup to the caller's own account
A mismatched account claim must not steer which owner the refusal names, and
a blocked user is still answered before the claim is validated. Both are load
bearing and neither was covered.
The agent network gateway service is private: agents reach it over the WireGuard tunnel, authorised by peer identity, with the cluster as its only target. Only a reverse proxy cluster with private capabilities can serve that, reported per cluster as the `private` capability — the same `supports_private` flag the dashboard gates NetBird-only services on.
A bootstrap could pin an account to a cluster without private capabilities, leaving an immutable dead gateway. Both bootstrap paths now validate the picked cluster: one the account can see must have a connected proxy reporting the capability. Shared and account-owned clusters qualify alike. Known-ness comes from proxy rows, not heartbeat freshness, so a cluster without the capability stays refused while merely offline. A hostname no proxy has declared stays pinnable (address-first). Identity is compared case-insensitively over the account's cluster list.
An agent network bootstrap stores its cluster as proxy_address, which selects the proxy that serves the endpoint. An account-scoped proxy only receives its own account's mappings, so a pin onto a host another account's proxy declares can never be served, and the endpoint is immutable — a dead gateway until the account deletes its settings. Nothing refused that pin; the domain unique index only arbitrates between endpoints.
Both bootstrap paths now refuse, before the insert, a host that another account's proxy declares, a host another account has labeled pins beneath (self-addressed path), or a hostname that is another account's endpoint (labeled path).
Shared clusters are unaffected: shared proxies are never foreign, and labeled pins under one cluster are never asked about, so any number of accounts still pin beneath eu.proxy.netbird.io. Registration is deliberately unchanged — refusing a proxy for another account's pin would let a pin lock a tenant out after the reaper drops its rows.
Management / Unit (amd64, mysql) hit the 20 minute go test budget on #7516. The package was not hung: each of the 133 test store creations in management/server paid about 1.6s on MySQL for CREATE DATABASE, the pre-migrations, a 40-table AutoMigrate and the post-migrations, which puts the package at 10 minutes on a healthy runner and over the budget on a slow one.
The migration now runs once per test binary into a template database and each test database is cloned from it, with CREATE DATABASE ... TEMPLATE on Postgres and a replay of SHOW CREATE TABLE on MySQL. The MySQL test container also drops the binary log, doublewrite buffer and per-commit redo fsync. Two goroutine leaks in the test helpers are fixed.
tools/gotestsummary turns the go test -json stream into a readable log, and the Management unit and integration jobs now pipe through it, so a timeout names the tests still running. On MySQL, management/server went from 10m16s to 6m36s.
`validateDeleteGroup` already refuses to delete a group that is still used by routes, policies, nameservers, setup keys, users, network routers, reverse proxy services, and agent network policies. Account-level agent network budget rules also store group IDs in `TargetGroups`, but that check was missing.
Deleting such a group left a dangling ID on the budget rule. `budgetRuleApplies` then never matched callers by group, so the spend cap silently stopped applying.
This adds `isGroupLinkedToAgentNetworkBudgetRule` and uses it in `validateDeleteGroup`, matching the existing helpers.
Prevent unvalidated registrations from reserving domain names indefinitely.
Give pending registrations a 48-hour validation window and clean up expired entries at startup and every 60 minutes. Emit CustomDomainValidationExpired for each deletion and preserve registrations referenced by services.
Reject validation after expiry and prevent concurrent validation from recreating deleted registrations. Normalize domain names with the shared parser before registration.
Migrate existing pending registrations to receive a fresh 48-hour validation window.
Require validated custom domains when creating or updating reverse proxy services.
Propagate validation errors during updates and return HTTP 409 for duplicate domain claims.
Add regression tests for domain validation, ownership, and service creation and updates.
The management binary is embedded by downstream builds that override
server construction via SetNewServer, but the cobra command tree itself
was closed: rootCmd is unexported and fully assembled in init, with no
way to attach additional subcommands. Customize hands the built root
command to a caller-supplied function before Execute, so an embedding
binary can add its own commands next to — or under — the built-in ones,
such as extra administrative helpers beneath the existing admin group.