mirror of
https://github.com/netbirdio/netbird.git
synced 2026-10-10 23:49:09 +02:00
67c2041fb5534c19d7b2320617cfc3aaab8e7d3d
193
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
008cf47512 |
[client, management] Recover from stuck proof collections and keep re-tracked accounts on the challenge refresher (#7890)
Co-authored-by: Viktor Liu <viktor@netbird.io> |
||
|
|
f0a40e4395 |
[client, management] Harden the certificate posture client and keep challenge nonces fresh (#8052)
* implement certificate posture check * log signal address * add keychain and cert store support * read the console user's keychain through a user session helper A root daemon cannot reach a login keychain: securityd is per session and a key ACL needs a session to prompt in, so dropping uid is not enough. The daemon now answers certificate challenges from the System keychain itself, where MDM installs device identities, and launches "netbird posture cert-proof" into the console user's desktop session with launchctl asuser for the login keychain. Only the signature and the chain cross back, never the private key. The console user comes from SCDynamicStoreCopyConsoleUser, bound with purego like the keychain calls. The login window reports no user, root, or "loginwindow", and all three are treated as no keychain to read, so a Mac at the lock screen sends device proofs alone. Adds info logging across the path: the keychain search list, per class query status and item counts, the chain built per candidate, and the verification error for every rejected candidate. A run that sends nothing now says why. README.md documents the trust model, the console user limitation and how to read the logs. * read the signed-in user's certificate store on Windows A service reads LocalMachine\MY, where AD and Intune enrol device certificates. CurrentUser\MY lives in the signed-in user's registry hive with keys protected against their profile, and a service that opens it does not fail: "current user" resolves to HKU\S-1-5-18, so it silently reads the service account's own empty store. The service therefore reads the machine store itself and launches "netbird posture cert-proof" with the session token for the rest, mirroring the macOS console user helper. Windows lets a privileged service assume a user identity, so the token goes straight into the child process and no external tooling is involved. CREATE_NO_WINDOW keeps a console window from flashing on the desktop every sync. In-process impersonation would also work but is per OS thread while goroutines migrate, so the child process avoids that class of bug. Session selection prefers the physical console and falls back to any active session, so remote desktop and VDI hosts are covered. WTSQueryUserToken needs SE_TCB_NAME, so a user-run client skips the helper and reads the machine store alone. SystemStore takes a store location, gaining NewUserStore alongside NewSystemStore and the per candidate logging macOS already had. The request building and proof merging move to helper_spawn.go, shared by both platforms, and helperStore picks what the helper reads per platform. * start TPM support * split goreleaser to support pkcs11 and exclude on docker * update goreleaser * go mod tidy * add tpm pin to netbird config * split cert and key location and allow key lookup on tpm * add unsupported flag for mobile devices * Isolate the cert proof helper from the service environment and cap its output * Read the PKCS#11 token PIN from NB_TPM_PIN instead of the profile config * Bound certificate proof collection so a stuck token or keychain cannot hold the sync loop * Stop retrying a PKCS#11 PIN the token rejected * Log certificate posture details at debug level * Sign only nonces and peer keys of the size management issues * Skip certificate files whose key belongs to another certificate * Bound PKCS#11 driver sizes, pin template values, and log out only a login the session owns * Never pass NULL to CFRelease and skip unreadable keychain identities * Keep the macOS keychain code out of iOS and the PKCS#11 driver out of Android * Find a chain to each challenge's CAs through every intermediate the store holds * Require a token label whenever a PKCS#11 PIN is set * Read user certificates only from the session of the active profile's owner * Collect certificate proofs again when the owner's session changes and report lost proofs * Test the PKCS#11 build against SoftHSM in CI and warn once where the build has no driver * Document where an inline PKCS#11 PIN is stored and how it is protected * Refuse PKCS#11 URIs that this client cannot honour instead of widening the match * Trust certificate and key files only when no other user can write or redirect them * Explain a Windows certificate whose key only a legacy CryptoAPI provider holds * Use platform absolute module paths in tests and add a real owner session test for Windows * Match the Windows profile owner by name instead of resolving it through the domain controller * Keep the certificate stores and TPM library out of the WebAssembly build * [client] Read TSS2 key files on go-tpm, checked against the library it replaces The TSS2 parser was the only reason this repository depended on a crypto suite whose own build tooling it inherits. The replacement sits on go-tpm, which was already a direct dependency and is in fact what that suite calls underneath, so this removes a wrapper rather than porting onto a different library: the load, the derived storage root key and the signing commands are the same calls. Swapping a parser on the one path a customer actually runs is not something to assert, so the two are held side by side for this commit. One test feeds the replacement bytes the old library wrote and requires the same key type, empty auth flag, parent handle, blobs and decoded public key; the other feeds both the fixtures the tests are built on, so those are the shape the format calls for and not merely the shape the new parser reads. The scaffolding goes away with the dependency in the commit that follows. The encoder behind the fixtures is written out separately from the parser under test, so an encoder bug and a decoder bug cannot cancel each other out. * [client] Drop go.step.sm/crypto and the repo-wide upgrades it imposed The TSS2 parser was the only thing in the repository that used this module, and it brought 302 modules into the graph to do it — 35 of them linters, along with Google Cloud KMS and IAM, the AWS SDK and a terminal styling library. Those are the module's own development dependencies, which minimal version selection turns into floors in ours, and they are the whole reason gRPC, protobuf, the AWS SDK, OpenTelemetry, logrus and five x/ packages had moved. Management, signal, relay and proxy inherited every one of them for a feature none of them runs. Removing the import is not enough, because tidy never downgrades: the raised floors stay written in go.mod. Each one is pinned back to the version main had, then tidy is left to raise again whatever something still genuinely needs. It raised nothing: all 43 are back where they were, and go-tpm was already in the graph at the same version, so the certificate feature now costs no new module at all. The differential tests go with it. They existed to check the swap against the library while both were present, and there is nothing left to compare against. * [client] Clear the lint findings only the macOS and Windows runners see golangci-lint analyses one build at a time, so running it on Linux says nothing about the two platforms CI also lints. Against those builds the feature's packages reported eight findings, and the structural one is Config.dir: it is dead on macOS and Windows because neither reads a directory at all, their collectors take the configuration and discard it. Moving the method beside its only callers makes that visible in the layout instead of in a linter, and leaves the gap itself — no file or token store on those platforms — where it belongs, as something to decide rather than something to silence. An absent key file beside a certificate was reported as a nil signer with a nil error, which the caller then had to recognise by its nilness. It is a sentinel now, so the meaning is in the error rather than in the absence of one. The rest follow the standard library: the elliptic coordinates and the private scalar come from the encoding helpers rather than the deprecated big.Int fields, and an error string loses its trailing colon. Lint is clean on linux, darwin and windows; the hardware TPM path was exercised separately against a real device and passes. * Accept the TSS2 emptyAuth boolean OpenSSL writes and persistent parents on 32-bit builds * Count the certificates field in the peer meta store test * Check the store directory before listing it, refuse group-writable files, and reject a URI with two PIN sources * Share a PKCS#11 login between sessions and send each PIN at most once at a time * Collect certificate proofs again when the meta sync carrying them failed * Use no Windows user store when a domainless owner matches accounts of several domains * Use no user certificate store when the active profile's owner cannot be read * Document the PIN sources on CertPKCS11URI and keep the README PIN example off the command line * Test that the PKCS#11 URI stays out of the debug bundle and run the wrong-PIN test only on a disposable token * Refuse a TPM PSS signature request for the maximum salt length * Add the certificate fields to the network map golden data * Retry posture checks whose meta sync timed out instead of dropping them * Start no system info gathering while a timed-out one is still running * Guard the applied posture checks across goroutines and keep refreshing proofs while a pending update times out * Log what a successful certificate proof helper wrote to stderr * Send recollected certificate proofs to management only when the proven chains changed * Explain a macOS keychain key whose access list does not allow netbird * Kill the whole macOS certificate helper process group when it times out * End sudo option parsing before the macOS certificate helper binary * Hold off system info gathering only while a timed-out one is still running * Collect certificate proofs on the posture watcher instead of under the sync lock * Read the certificate store directory and PKCS#11 URI from the daemon environment, not the profile config * Install the RPM sysconfig file readable by root only and show the certificate posture variables * Move the certificate posture README into the package doc and the docs site * Name NB_CERT_PKCS11_URI in the PIN-without-token error * Keep the file check results of the latest-started system info refresh * Give the full import command for a keychain key netbird may not use, and correct the package doc * Restrict the service environment file to root on every package install * Search only the System keychain in the macOS daemon and only the login keychain in the user helper * Let the certificate proof helper read the PKCS#11 token from the environment on Linux * Ask a macOS user's keychain again only after an hour when it proved nothing * Clear the lint findings in certificate posture * Hold off the keychain helper only after a completed or timed-out run, independent of CA order * Keep free functions out of the method lists of PKCS11Store, URI and Challenger * Name the post-install permission helper in snake case and shorten the sysconfig certificate block * Drop the certificate store directory from certproof.Config, which only NB_CERT_STORE_DIR sets * [management] Renew certificate challenge nonces on quiet accounts A certificate challenge nonce is accepted for its own window and the one before it, and it only reaches a peer attached to a network map. An account where nothing changes sends no map, so after a day the peer re-sends the nonce it still holds, verification rejects its whole proof set, and the certificates stored for it are dropped. It fails the certificate check and loses every policy gated on it until some unrelated change happens to push a map. The outage repairs itself in seconds, which is what makes it expensive: it is intermittent, it only hits stable networks, and it is not reproducible on demand. Push the account's peers an update often enough that the nonce they hold is never close to expiring. Only accounts whose posture checks actually ask for a certificate are tracked, so a deployment without the feature does no extra work. The refresh runs from one goroutine over a map of accounts rather than a timer per account: the period is hours, so one pass every few minutes costs nothing next to it, and there is no timer to re-arm when an account that falls due sooner appears. Each account's first run is offset by a hash of its ID, because the challenge window is global and an instance restart would otherwise arm every account in the same moment. The push carries no administrative change, so it is counted as a refresh rather than an update and stays out of the figures that track what was edited. (cherry picked from commit |
||
|
|
53a14551c8 |
[client, management] implement certificate posture check (#7535)
Co-authored-by: mlsmaycon <mlsmaycon@gmail.com> |
||
|
|
515a01dd11 |
[client] Force interactive login when extending the auth session (#7216)
* [client] Force interactive login when extending the auth session A session extend must be answered from the account the peer is registered under. With a silent PKCE flow (DisablePromptLogin or max_age=0) the IdP answers from whatever session it already holds, which need not be the peer's account when several are signed in; the token then fails the user match in ExtendAuthSession with no way to pick another account. Mark the PKCE flow request as a session extend so the management server can force prompt=login for it, overriding the configured silent flow. * [client] Reduce cognitive complexity of Server.Login Login sat at cognitive complexity 27, over the 25 the linter allows. Extract the interactive SSO branch into startSSOLogin, and split the nested in-flight-flow reuse check out of it into reuseOAuthFlow, which flattens the original if/else into early returns: it returns the cached auth info when the previous flow targets the same client and still has more than 90s left, otherwise cancels the stale wait and returns nil so the caller requests a fresh flow. The helpers take the contextState through a small statusSetter interface, since internal.contextState is unexported and re-deriving it with CtxGetState inside the helper would resolve against callerCtx rather than rootCtx. No behavior change: same ordering of state transitions, same mutex scope around the oauthAuthFlow write, same error paths. Login is now at 21. * [client] Respect DisablePromptLogin when extending the auth session Forcing prompt=login on a session extend overrode DisablePromptLogin, which is set for IdPs that break on it: Authentik triggers a double authentication and social logins fail outright. Overriding it there trades a recoverable extend for a login that cannot complete at all. Keep the LoginFlag override, which only replaces max_age=0 or none with prompt=login so the IdP honours login_hint, and leave DisablePromptLogin as configured. Those deployments keep the silent flow, and with several accounts signed in an extend answered from the wrong one still fails the user match. * [client] Guard the shared OAuth flow state with the server mutex reuseOAuthFlow read flow, expiresAt, waitCancel and info without holding s.mutex, while startSSOLogin and WaitSSOLogin write them under it. Reading the fields one at a time could also answer with auth info from a flow that was already replaced, or cancel a wait that no longer belongs to the flow just judged stale. Take one snapshot under the lock and decide from it. WaitSSOLogin read oauthAuthFlow.flow twice outside the lock; both now use a value snapshotted in the critical section that already installs actCancel. Its stale waitCancel was read and called in a separate section from the one installing the new one, so two racing calls could read the same predecessor and leave one wait uncancelled. Swap the two in a single critical section. Both cancels run after unlocking: the displaced wait takes s.mutex as it unwinds. * [client] Verify the SSO login came back for the hinted account login_hint is a suggestion the IdP may ignore: with a silent flow configured (DisablePromptLogin or max_age=0) and a live IdP session for another account, the login completes with that account's token. On a registered peer the management server rejects it as a user mismatch, but on a fresh profile the peer silently registers under the wrong account and the profile is then bound to it — every later login follows the stored hint straight back. After the token exchange, compare the ID token's email against the hint the flow was sent with. On a mismatch, do not log in to management with the token; run one more round asking the IdP to re-decide the account (prompt=login, via ForceAccountPrompt — DisablePromptLogin still wins there). If the prompted round also comes back different, proceed with a warning: the address may legitimately have changed, and refusing forever would lock the user out of the profile while the management server still rejects a token that does not own the peer. A token or profile with no email to compare is not judged. The retry differs per platform because of who opens the browser: - CLI (netbird login foreground) and Android run the whole flow in one process, so the mismatch retries automatically: the browser reopens with the account prompt within the same login attempt. - On desktop the login is split between the daemon and the GUI: Login hands the authorize URL to the GUI, WaitSSOLogin blocks for the token, and only the GUI can open a browser. A new URL cannot be handed out from inside WaitSSOLogin (its response has no field for one, kept that way to avoid a proto change), so the daemon arms forceAccountPrompt, fails the round with "connect again to choose the account", and builds the next Login's flow with the prompt — the user's next connect is the retry. The flag and the flow annotations live in daemon memory only; SwitchProfile drops them so the previous profile's hint cannot judge the next profile's token. The device code flow has no prompt parameter (RFC 8628), so a prompted round there runs as-is and a repeated mismatch is let through with the warning rather than looping. * [client] Address review comments on PKCE session extend flow Fail the PKCE authorization flow test on request error instead of continuing into a nil dereference, and make the godoc comments on the touched exported symbols identifier-leading full sentences. * [client] Match accounts only on the email claim of the ID token The name-claim fallback in the ID token parsing is kept for the login hint and display, but account matching now only considers a value that came from the email claim, so a token without one no longer produces a false account mismatch. * [client] Drop the pending session extend on a profile switch The profile-switch cleanup dropped the pending login flow and the account-prompt flag, but left extendAuthSessionFlow untouched. Its device code was issued by the previous profile's IdP client, so a WaitExtendAuthSession still parked on the browser leg would submit the resulting token against the new profile's engine. * [client] Judge the SSO account against the flow that produced the token WaitSSOLogin snapshotted the flow on entry but re-read the info, hint and accountPrompted from the live s.oauthAuthFlow afterwards, in separate critical sections. WaitToken blocks for the whole browser leg, so a concurrent Login or RequestJWTAuth could replace the flow meanwhile and the mismatch check would compare this wait's token against another flow's account: either arming the prompt spuriously or letting a wrong-account token through against an unrelated profile's hint. Take all of it in the entry snapshot. * [client] Keep the forced account prompt from being lost to flow reuse startSSOLogin consumed forceAccountPrompt and applied the prompt to the freshly built flow, but reuseOAuthFlow could then answer from a cached flow for the same client — one built without prompt=login, e.g. by RequestJWTAuth. The user got the same silent authorization URL that produced the mismatch, with the flag already spent, so no later round asked either. Rule reuse out when the prompt is forced, while still cancelling the predecessor's wait. RequestJWTAuth also wrote the flow fields one by one, leaving the previous login's hint and accountPrompted behind for WaitSSOLogin to judge a later token against. Both sites now replace the whole record. * [client] Consume the forced account prompt after the retry forceAccountPrompt was never cleared, so a flow that outlived the retry it was armed for kept sending prompt=login on every later authorization request and re-authenticated the user each time. RequestAuthInfo now takes the flag as it builds the request. * [client] Cancel the caller context in the SSO login tests WaitSSOLogin parks a goroutine on the caller's context for the whole browser leg. The tests passed context.Background(), which never cancels, so each left one goroutine behind for the lifetime of the test binary. * [client] Cancel the wait displaced by an OAuth flow replacement Replacing the shared record with a whole struct value dropped the previous flow's waitCancel, so an SSO browser wait still parked on it lost its cancel: nothing could preempt it, and it could go on to run attemptLogin or mutate the record behind the new flow. Both replacement sites now take the displaced cancel over in the same critical section, via a shared replaceOAuthFlow, and invoke it after the unlock. * [client] Guard OAuth flow mutations by the flow that owns the wait * [client] Arm the account prompt only from the wait that owns the flow * [client] Adopt the three-value parseEmailFromIDToken in the device flow The main merge brought in the device flow's email extraction from #7193, which still used the two-value signature this branch replaced when account matching was narrowed to the email claim. Git merged the files without a textual conflict, so the branch stopped compiling. Take the fromEmailClaim result and fill EmailClaim from it, the same way the PKCE path does, so device-flow clients get the same account matching. * [client] Populate the pending extend flow in the test server helper SwitchProfile cancels and clears the pending session extend flow unconditionally, the same way it clears the SSH JWT cache. New always populates the field, but the hand-assembled test server did not, so TestSwitchProfile_ClearsJWTCache panicked on a nil PendingFlow. |
||
|
|
24cb7b75c2 |
[management] switch to libopenapi for managing of openapi-based api (#8056)
* use libopenapi OpenAPI generator Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * added a handler for v1alpha1/peers Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * wire up request validator Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * wired up spec-based validation Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * use sync validation; set base url to v1alpha1 Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * moved stuff around, extracted runtime libopenapu deps into runtime_tooling Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * testing /peers path params Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * fixed tests Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * added user schemas and paths Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * use api validation in tests Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * post-merge fixes Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * deleted tmp command used to test validator integration Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * require request body in post/put requests in order to force validation Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * making linter happy Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * switch to api/v1alpha1 Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * go mod tidy Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * keep the original order of middleware: metrics, cors, then auth Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * return 422 on validation errors Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * cleanups Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * more cleanups Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * v1alpha1 spec cleanups Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * no need for a double-pointer Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * more linter fixes Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * removed more double pointers Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * use di to inject api_v0 and api_v1 http routers Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * do not re-add middleware on repeated call to ApiHandler Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * calling ApiRouter() now also calls ApiV1Router() Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> --------- Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> |
||
|
|
5433e92183 |
[client] Stop one unreachable relay from churning the others (#8098)
* [client] Pass the home relay client into isForeignServer isForeignServer read m.relayClient directly, so it was only safe while the caller held relayClientMu, and that dependency was invisible at the call site. Taking the client as a parameter makes the requirement explicit and lets a caller narrow the window it holds the lock for. Pure refactor, no behavior change: the single caller passes m.relayClient while still holding the read lock. * [client] Release the relay client lock before dialing a foreign server OpenConn held relayClientMu for its whole body, dial included. Dialing a foreign relay takes as long as that server's connect timeout, and onServerDisconnected needs the same mutex in write mode; a pending writer also blocks later readers, so a single unreachable relay stalled every relay operation on the node for the duration of the dial. On a routing peer this stalled the disconnect handling of an unrelated relay for a full dial timeout. By the time the eviction ran, the track had already been rebuilt, so it dropped a working client and the next OpenConn built a duplicate connection to a server the peer was already connected to. Read the home client under the lock, then release it and dial unlocked. * [client] Keep a reconnected foreign relay client on a late disconnect notice evictForeignRelay deleted the track by server address without checking which client the entry held. The disconnect notice is delivered on its own goroutine and carries only the address, so it can arrive after the track has been rebuilt: the eviction then dropped a live client, and the next OpenConn built a second connection to a server the peer was already connected to. The relay answers a duplicate peer ID by closing the existing connection, which tears down every relayed channel on it, including the peers whose active path it is. Skip the eviction when the stored client is still connected. * [client] Drop the comments added with the previous commits Comment-only change, no behavior change. * [client] Cover the foreign relay stall and the late disconnect notice Two regression tests, both failing on the unfixed code: TestOpenConn_DoesNotHoldRelayClientLockAcrossDial dials a listener that accepts and never answers, then calls onServerDisconnected and fails if it does not return. On the unfixed code the writer waits out the whole dial. TestEvictForeignRelay_KeepsConnectedClient opens a real connection to a second relay and then replays a disconnect notice for that address. On the unfixed code the live client is dropped from the map. Both reuse the existing stallingRelayListener and relay server helpers. * [client] Keep a foreign relay track whose dial is still in progress openConnVia publishes the track before dialing and fills relayClient only once the dial finishes, so the readiness guard did not cover the dial-in-progress window: a disconnect notice arriving there deleted a track whose client was about to connect. That client then became unreachable through the map, so cleanUpUnusedRelays could not close it either, and the next OpenConn dialed a duplicate the relay answers by closing the first - the churn this branch is meant to remove. Skip the eviction while relayClient is still nil, the same signal cleanUpUnusedRelays already uses to leave an in-flight dial alone. Reported by cubic on PR #8098. |
||
|
|
2623feeb5b | [management] remove ingress ports (#8062) | ||
|
|
eb5a98c059 | [management] add tenant delete endpoint to openapi (#8054) | ||
|
|
e2678d4e05 | [management] expose peer MAC addresses and make peers searchable by MAC (#6553) | ||
|
|
0dc729c4ea | [client] Cache the box shared key per remote peer in the Signal client (#7807) | ||
|
|
30dd076b36 |
[management,proxy] Use single-use codes for OIDC session handoff (#7635)
* Generalize PKCE verifier store into SingleUseStore * Generalize PKCE verifier store into SingleUseStore * Extend single-use store to generate one-time retrieval codes * Hand off proxy OIDC session via one-time code instead of URL token * Use the single-use store in integration tests * Read active proxy versions by cluster * Detect proxy clusters that support session codes * Bind OIDC session handoff mode to signed state * Deprecate legacy OIDC session token handoff * Remove unrelated session code test stub * fix tests * fix merge * Fix session code compatibility detection * Isolate proxy session codes in shared cache * bump min session version |
||
|
|
10a04fcccb |
[management] Add a disabled state to the managed Agent Network proxy API (#7744)
A managed Agent Network gateway can be turned off by the platform. The derived state had no value for that, so a disabled deployment reported whatever the operator last saw, usually provisioning. Add `disabled` to the AgentNetworkManagedProxy state enum and say in the POST and GET descriptions that a disabled deployment answers with it and that provisioning again does not turn it back on. |
||
|
|
a2919e26dd | [management, proxy] Enforce strict base64url decoding for JWT validation (#7554) | ||
|
|
5d92c89227 | [management] move rate limiter to shared package (#7727) | ||
|
|
979571a99f |
[management] Prevent deleting custom domains used by services (#7515)
Deleting a custom domain released its name while services still pointed at it, leaving them on a namespace the account no longer held. Deletion now refuses with 412 when a service in the same account uses the domain or a subdomain, including disabled ones. Service writes revalidate authorization inside their transaction and hold a shared lock on the matching registrations, so a delete racing a create cannot strand either. The dependency lookup is account-scoped: registrations are unique by name, so another account can hold team.example.com under example.com and its services are authorized by its own registration. |
||
|
|
7009add7a9 | [management,signal,proxy] add pyroscope profiling (#7536) | ||
|
|
cb7ca8ef3f |
[client,management] Skip route firewall rule computation when no firewall (#7624)
* [client,management] Skip route firewall rule computation when no firewall A peer that runs with the firewall disabled has no ACL manager and no firewall to program, so nothing ever reads RoutesFirewallRules: the only consumers are acl.Manager, which is reached solely when e.acl is set, and the legacy-management probe in updateNetworkMap, which is guarded by a non-nil firewall. Building those rules is the most expensive part of a sync on a peer that routes many network resources. On a 15k-peer deployment a debug bundle showed getPeerNetworkResourceFirewallRules accounting for 62% of the allocations of Calculate, and Calculate for effectively all of the allocations of handleSync, which was taking 3.2s on average and holding the engine lock for the duration. Let the caller ask Calculate to leave the rules out. The client passes its existing DisableFirewall setting; the management server keeps the default and still produces them. RoutesFirewallRulesIsEmpty is set from the resulting empty list, so a receiver that would otherwise infer legacy management from an empty rule set does not misread the skip. * [client,management] Cover the skip flag through the envelope Review feedback on #7624. The components test compared only the length of the peer firewall rules, so a change to their content would have passed while the message claimed they came out unchanged. Compare the slices. The skip path was also only exercised by setting the field directly on the components, which bypasses the envelope conversion where RoutesFirewallRulesIsEmpty is derived. That bit is what keeps the client from reading skipped rules as a legacy management server, so it gets a test that goes through EnvelopeToNetworkMap with the flag set. * [management] Give the router a peer ACL so the rule comparison bites Review feedback on #7624. peer-router-1 appears in no peer ACL in the shared fixture, so its FirewallRules came out empty and the equality assertion compared two empty slices — it would have passed even if the peer rules were dropped entirely. Add a policy covering the router and require the baseline to be non-empty before comparing. |
||
|
|
bc0671fd21 |
[client] Fix peers not being notified when the relay connection drops (#7490)
* [relay] Signal relay disconnects through the conn context AddCloseListener deduplicated listeners by comparing reflect.ValueOf(callback).Pointer(). For a method value that pointer is the address of the compiler-generated wrapper, not an identity bound to the receiver, so every peer's w.onRelayClientDisconnected compared equal. All peers on the home relay register under the same connectionURL key, so only the first registration survived and the rest were silently dropped. On a relay disconnect those peers were never notified: statusRelay stayed connected and the reconnect guard never fired. The relayed net.Conn itself was closed by closeAllConns, so nothing leaked, but the peer state machine did not learn about it. Foreign relays had the same defect scoped to the peers sharing that server. Rather than fixing the deduplication, drop the peer-level listener registry entirely. A relayed Conn now exposes Context(), cancelled when the connection is torn down, with a cancellation cause naming the reason. This is the same shape quic-go uses for its Conn and Stream types, and it removes the whole class of problems around listener identity, lifetime and deregistration: the signal belongs to the resource instead of a side table. WorkerRelay watches that context in a goroutine whose lifetime matches the connection. A watcher that wakes up for a superseded connection compares the conn pointer against the current one and returns without touching the state machine, so a fast relay reconnect cannot have a stale watcher tear down the connection that replaced it. Client.SetOnDisconnectListener stays: it is server-level and drives the reconnect guard and foreign relay eviction, unrelated to peers. handleRelayReady also checks the conn context, closing the race where the relay dies between OpenConn and the readiness handoff and the peer would otherwise build a WireGuard endpoint over a dead connection. TestNotifierDoubleAdd covered the removed mechanism and is gone. TestForeignAutoClose asserted nothing (both branches logged); it now waits for the relay to leave the client map and fails if it does not. * [relay] Fix build: return the concrete conn from Client.OpenConn OpenConn now returns *Conn, but it still went through connContainer.netConn(), which widens to net.Conn. The helper had one caller and only existed to produce the interface value the signature no longer wants, so return container.conn directly and drop it. * [relay] Assert the local-close cancellation cause explicitly The local-close test only rejected ErrServerDisconnected, so it would also have passed for ErrPeerDisconnected or a bare context.Canceled. closeConn cancels with net.ErrClosed, so assert that. * [client] Ignore relay disconnects from superseded connections The relayed conn watcher compared the conn pointer under relayLock, released it, and only then tore the connection down. A new offer could install its replacement in that window, so a watcher that validated the old pointer went on to close the proxy of the connection that had already replaced it and report the peer as disconnected while it was up. Move the decision to where the teardown happens. Conn records which relayed connection the current proxy was built from, and onRelayDisconnected takes the connection the signal belongs to and drops it under conn.mu when it is no longer the current one. Check and effect are now in the same critical section, so the verdict cannot go stale before it is acted on. This also covers the proxy read loops, whose disconnect listener took no argument and had the same defect: it now names the connection it belongs to. The WG timeout path keeps passing nil, since it deliberately tears down whatever is current. * [client] Bind the relayed conn reference to the proxy swap relayedConnRef was set at the top of the readiness path, but wgProxyRelay only changes at the end, in setRelayedProxy. The two failure returns in between — newProxy and ConfigureWGEndpoint — left the reference pointing at a connection that never became active while the old proxy was still installed. A disconnect of that old, live relay would then be dismissed as belonging to a superseded connection and never cleaned up. Set the reference in setRelayedProxy, next to the proxy it belongs to. Both success paths go through it and neither failure path does, so no failure branch has to remember to roll anything back. |
||
|
|
314d88252d |
[management] Name the account owner in the pending approval error (#7533)
* Name the account owner in the pending approval error
A user refused because their account is pending approval had no way to
learn who could approve them. The refusal now carries the account
owner's address, masked, so a caller can name someone to contact without
being handed the address itself.
Resolving the owner is best effort: a lookup failure, or an account
predating the stored email, falls back to the refusal as it was.
* Name only the caller's own owner in the pending approval error
The refusal is raised before ValidateAccountAccess has established that
the caller belongs to the account the request asked about, and the user
is loaded by ID alone. Resolving the owner of the requested account
therefore disclosed that owner's address to a pending user with no claim
to it, reachable through any handler that takes an account ID from the
caller — DELETE /api/accounts/{accountId} passes one straight through.
The owner who can approve a pending user is the owner of their own
account, so resolve that one. The requested account is never read.
* Mask short local parts whole in MaskedEmail
Keeping the first two characters and the last hides nothing until the
local part is four long: at three or fewer they are the whole of it, so
"abc@example.com" masked to "ab****c@example.com" and a pending user
could recover the owner's address in full from what is meant to conceal
it. Short local parts are now replaced entirely.
* Name the owner from GetCurrentUserInfo instead of the permission gate
The gate could only read the stored user row, which carries no address
when an external IdP owns the identities — the usual case — so it named
no one in practice. It also had no way to reach the IdP without being
handed the account manager, which meant restoring bootstrap wiring that
a refactor had dropped.
GetCurrentUserInfo already holds that account manager, so it answers for
a pending user itself and reuses GetOwnerInfo, the same lookup /msp uses
to resolve an owner's address. The gate returns to exactly what it was,
and with it goes the risk of naming the owner of an account the caller
only asked about.
MaskedEmail becomes MaskEmail: with a UserInfo in hand there is no stored
row to hang it off.
* [management] Cover the pending approval refusal in GetCurrentUserInfo
The branch that names the owner had no coverage at the manager level, so
neither the named refusal nor the fallback for an owner without a resolvable
address was pinned down.
* [management] Cover the failed owner lookup in the pending approval refusal
The generic fallback has two ways in: no address on the resolved owner, and no
owner to resolve at all. Only the first was pinned down.
* [management] Pin the owner lookup to the caller's own account
A mismatched account claim must not steer which owner the refusal names, and
a blocked user is still answered before the claim is validated. Both are load
bearing and neither was covered.
|
||
|
|
794956a7a3 |
[client] Fix relay instance address race (#7498)
Read the relay instance URL and IP atomically to prevent reconnects from mixing values from different connections. Extend existing connection and offer/answer logs with relay URLs and IPs to help trace mismatched advertisements. |
||
|
|
76ea72237f |
[management] Add Agent Network managed proxy to the API spec (#7433)
Defines the cloud-side managed gateway provisioning surface (POST/GET /api/integrations/agent-network/managed-proxy) and its response objects so clients consume generated types instead of hand-written ones. POST is idempotent: 202 when the call starts (or restarts) provisioning, 200 when a deployment already exists; 409 names an already-assigned endpoint the managed flow does not own and 503 signals temporarily exhausted endpoint allocation. |
||
|
|
825389818c |
[client] Gather fresh system info on every management sync stream connect (#7409)
* Gather fresh system info on every management sync stream connect The engine collected the peer meta once at start and reused the same Info for every Sync stream reconnect, so a mobile network switch that redials management kept reporting the old local network addresses. The peer network range posture check was then evaluated against stale data until the client restarted. Sync now takes a gatherer that runs at each stream connect. The gatherer is cheap: GetInfo plus the cached posture check file results, kept in the new system.InfoSource, which the engine refreshes whenever the checks list changes. No process enumeration runs on the reconnect path. Also fix the management mock server calling itself instead of SyncFunc. * Evaluate the login response posture checks before the first sync connect The engine starts with the checks the login response carried, and the first sync stream request used to send their evaluated file results. After moving the gather into InfoSource, the stream opened with an empty cache and the first sync response did not refill it, because its checks equal the ones the engine already holds. Desktop peers therefore never reported process or file posture results. Seed the cache once before the first connect, where the old gather ran, so a timed out evaluation still falls through to the address-only info. * Harden the sync info source against nil callbacks and shared slices A nil getInfo opens the stream without metadata, as a nil sysInfo did before. The cached posture results are a copy, so the Info returned by Refresh cannot alias the snapshot later Current calls report. The exclusion test asserts the remaining address count so it cannot pass vacuously on a single-address host. * Retry a posture check refresh that timed out or failed to sync The checks list was recorded before the gather ran, so once the gather timed out or SyncMeta failed, the next sync response carrying the same list matched the recorded one and nothing retried. The peer kept reporting the previous posture results until the list changed again. Record the checks only after the meta reached management, so a failed cycle is repeated on the next sync response. * Log the skipped posture refresh, let the mock Sync return errors and deflake the reconnect test * Drop the nil guard around the sync info callback * Send the refreshed info on the first sync connect instead of gathering it twice |
||
|
|
c2b5d211d9 | [management] Enforce reverse proxy group access before minting and when honouring a session cookie (#7240) | ||
|
|
6aaeed744e |
[management] Check a provider's url and credential before saving it (#7301)
A bad upstream or key saved cleanly and surfaced minutes later as a failed request or an empty model picker, with nothing pointing back at the record. CreateProvider now spends the credential once against the vendor's model listing. UpdateProvider does the same when the upstream, the key, the catalog provider or the skip-TLS flag changed — only then, so renames and price edits neither wait on a vendor nor fail because one is down. Both run before the store write, so a rejected rotation leaves the working key where it was. What cannot be checked still saves: no listing endpoint, no derivable Bedrock control-plane host, a private upstream, a record skipping TLS verification. Everything else blocks, outages included — 5xx, 429 and timeouts leave the record unverified just as a refusal does. Refusals return 422 and carry no status code or echoed URL. Discovery now reads as a partial edit, so a retyped URL can be listed against without also rotating the credential. Entries with their own listing host (Bedrock) get their configured upstream resolved separately, since a successful listing said nothing about it. |
||
|
|
e3d6c3d0eb | [management] fix private services calc on new db path (#7383) | ||
|
|
ebc259e30b |
[management,client] Gate remote jobs behind an admin opt-in with MDM support (#7153)
This introduces a disabled-by-default allow-remote-jobs setting that controls whether the management server may run jobs (such as debug bundles) on a peer. The flag propagates end to end: through client configuration, the daemon SetConfig and Login requests, authentication, and system info, up to management, where it is stored on the peer and exposed on the peers API as remote_jobs_allowed. The client refuses any management-requested job unless the peer has opted in. Because enabling remote jobs crosses the user-to-root boundary, turning it on requires privilege, mirroring the SSH-server gate. Administrators can enforce the setting through MDM policy on both macOS and Windows, and MDM can also override the debug-bundle upload URL. The change ships policy documentation and generated profile templates, and adds configuration, conflict, and enforcement tests covering the opt-in, privilege, and MDM paths. |
||
|
|
3027130f0f |
[management] Add Agent Network access roles and self-service endpoints (#7221)
Delegating Agent Network today means handing out full account admin, and regular users cannot see their own usage or how to connect a local tool. Add two roles on top of the existing agent_network permission submodules. agent_network_admin owns the whole area (providers, policies, guardrails, budgets, usage, logs, settings) with read-only users, groups, peers, and account info needed to build policies, and nothing else in the account. usage_viewer is the regular User baseline plus read on the aggregated usage and cost overview: no provider configuration, no policies, no request-level logs, which can contain captured prompts. billing_admin gets a proper permission-map entry with the User baseline so role resolution stops failing with role-not-found; its plan and invoice permissions stay enforced cloud-side. Add the self-service endpoints behind the "My Agent Network" view, available to every authenticated user because both answers are scoped strictly to the caller. GET /api/agent-network/me/setup returns the account endpoint plus the providers and models the caller's own groups authorize, computed with the same rules the proxy enforces: policy filtering as in policy selection, model allowlist union intersected with declared models, orphan and disabled providers omitted. Not set up and no access are deliberately indistinguishable, and the response carries display metadata only. GET /api/agent-network/me/consumption returns the caller's own user-dimension counters. |
||
|
|
1081ca006d |
[management,client] Add anonymize level and upload URL to remote debug bundle jobs (#7147)
This extends the management-requested remote debug-bundle job with two new, optional parameters. anonymize_level selects how aggressively the bundle is scrubbed: "default" keeps internal (private) IP ranges readable, while "strict" also anonymizes private, CGNAT and link-local addresses; the value is trimmed and lowercased, and an unknown level is rejected at creation. upload_url lets an operator point the peer at a specific upload service instead of the default one; it must be a well-formed https URL with a host, and an empty value falls back to the default upload server. Both fields flow through the job workload API and are surfaced in the create-debug-job modal on the dashboard. Validation is shared so the client executor and the management boundary agree on what a valid upload URL is, preventing drift between the two checks. |
||
|
|
12e8874517 |
[client, relay, management] Bump go version to 1.26 and go-quic to v0.62.0 (#7359)
* Bump go version to 1.26 and go-quic to v0.62.0 * Replace deprecated ecdsa public key assembly and add tests for jwt * Update goversioninfo * Pin go toolchain to 1.26.7 |
||
|
|
353251d886 | [management] fix posture check evaluation for direct peers in policy definition (#7348) | ||
|
|
611a9291cd | [management] fix posture check flip evaluation for affected peers calc (#7347) | ||
|
|
e06c17cf59 |
[management] network map from nmap data type (#6919)
Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> Co-authored-by: Dmitri Dolguikh <dmitri.external@netbird.io> |
||
|
|
51095cb986 |
[client, management] Support per-peer lazy connection state and default proxy peers to lazy (#6762)
* Support per-peer lazy connection state and default proxy peers to lazy * Classify forward targets from incoming config in lazy exclusion * Set IsUserspaceBind mock so lazy manager starts in engine test * Skip lazy exclude reconciliation when the set is unchanged * Keep cached lazy flag when a sync carries no peer config |
||
|
|
15fff4c164 |
[client] Sweep connections on network loss via a shared netevents manager (#7254)
Losing the last network only flipped the availability state: the dead management, signal and relay sockets stayed silently connected until their own timeouts, so the client kept reporting Connected with no network at all. Introduce client/netevents with a Manager that ties the availability state, the connection sweeper and the status recorder together, and move the netstate and netsweep packages under it (netsweep renamed to sweep). SetNetworkAvailable(false) now also sweeps the registered connections so their owners redial and the listener reaches the NoNetwork state. The Android and iOS bindings own a Manager instance and inject it through the constructors; consumers hold the concrete *Manager whose nil zero value reports always-online and never sweeps, with interfaces kept only as parameter contracts. The relay guard settle wait moved into the Manager as WaitSettled, removing the netevents import from the relay package. |
||
|
|
f03853867b |
[proxy,management] Serve Bedrock model discovery from the control plane (#7250)
[proxy,management] Serve Bedrock model discovery from the control plane
A Bedrock provider could never answer a model-discovery request. The router
sent GET /inference-profiles to the record's upstream, which has to be
bedrock-runtime.<region> for InvokeModel to work, and that host does not
implement the operation. ListInferenceProfiles is a control-plane operation on
bedrock.<region>.amazonaws.com, and one provider record carries one upstream,
so the two hosts genuinely differ.
The route now carries a discovery host, taken from the catalog's declaration
with the region read back out of the configured upstream, and the listing — and
only the listing — goes there. Inference is untouched. A proxied or self-hosted
Bedrock endpoint gets no discovery host at all rather than a guessed one, since
inventing a host would send the operator's credential somewhere they never
configured.
Two things had to follow for the listing to be usable once it arrives. The
response filter only understood OpenAI's {"data":[{"id":…}]}, so a Bedrock
listing fell through it untouched, offering every profile in the account
whatever the policy said. And discoverableModels intersected by exact string,
so a record registering the raw profile id while a guardrail names the catalog
key intersected to nothing — bounding a working provider's listing down to
empty.
Normalisation is the third. The geography in front of a cross-region profile
was matched against a hardcoded list of four, so every profile issued under jp,
au, ca, sa or us-gov carried its prefix into the pricing key, matched no
catalog entry and metered at zero. It is now recognised by either the geography
or the vendor that follows it, so an id has to be new on both axes at once to
slip through — a live eu-central-1 listing returned "global.xai.grok-4.6" days
after the vendor list was first written.
|
||
|
|
5e88d3f87a |
[management] Offer a provider's live model list in the config form (#7246)
[management] Offer a provider's live model list in the config form Adds POST /api/agent-network/catalog/providers/models, which asks a vendor which models an operator's own credential can actually reach, so the provider form can offer a live list instead of only the compiled-in catalog. The catalog goes stale, and it cannot see an account: which OpenAI models an org is entitled to, which Bedrock inference profiles an account and region hold, which Vertex models a project has enabled. The endpoints, auth headers and response shapes come from probing the live APIs (#7244); each vendor invented its own envelope and none can be guessed from the request. Bedrock shaped the design: its listing lives on the control plane while inference must go to the runtime host, so Discovery carries its own host rather than reusing the record's upstream, and profile ids are taken verbatim because the region prefix is what AWS requires at invoke time. A caller supplies either the key they are typing or the id of a saved record whose stored credential is reused — never both, since accepting both would run an arbitrary credential under the identity of a record the caller may only be permitted to read. Gated on Create rather than Read, because this spends the operator's credential against a third party. Management has not made outbound calls on an operator's behalf before and it holds a credential for every provider, so every resolved address must be public — covering loopback, RFC1918, the cloud metadata address and NetBird's own 100.64/10 range — and redirects are not followed, since a redirect moves the request to a host the check never saw. The vendor is authoritative for the id; the catalog stays authoritative for pricing. A discovered model the shipped table cannot price returns pricing_known: false so the operator must set rates rather than being registered at a silent zero. |
||
|
|
766fcae3f8 |
[proxy,management] Conform the Agent Network endpoint to the LLM gateway protocol (#7154)
[proxy,management] Conform the Agent Network endpoint to the LLM gateway protocol Reviewed the proxy against Claude Code's published gateway contract. The transport layer already held up; fourteen gaps sat one layer up, in the model catalog and in the non-inference endpoints clients call. Two of them cost money. The catalog carried no claude-opus-5 or claude-sonnet-5, so an operator could not authorise the models coding agents default to — those requests denied as not-routable, or priced at zero where a catch-all carried them. And gateway records pin ParserID "openai" while the same record serves /v1/messages, so Anthropic responses were read with the OpenAI parser, which never looks at message_start where input tokens live: input metered as roughly zero on every stream and cost was skipped entirely. The rest fix requests refused for structural rather than policy reasons: model discovery denied for every account with a model allowlist, token counting denied on Bedrock and mis-parsed on Vertex, startup probes refused and written into the access log at every session start, and denials rendered in a shape no LLM client parses. Two changes are additive by design — the deny body keeps every field it had and adds the vendor's error object alongside, and body-level identity injection is now gated on the request's dialect so it stops sending OpenAI-shape fields into Anthropic bodies that reject them. The end-to-end work turned up one more: the discovery filter treated any slash in a model id as a gateway prefix, which would have dropped every self-hosted "Qwen/..." model from the picker. |
||
|
|
e206f8827d | [management] Suppress staticcheck warnings for deprecated proto fields (#7261) | ||
|
|
a144e8c144 |
[client, management] switch to go.uber.org/mock (#7253)
* switch to go.uber.org/mock/gomock Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * updated go:generate commands + regenerated mocks Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * update go:generate mockgen commands Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * removed duplicate import Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> * fix go:generate Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> --------- Signed-off-by: Dmitri Dolguikh <dmitri.external@netbird.io> |
||
|
|
070a0a7bf1 |
[client, android] Handle network changes without restarting the engine (#7144)
On network changes the client restarted the whole engine. That is heavy-handed and slow: it tears down working state to recover from a transition the engine could handle itself. This replaces the restart with proper network event handling. Suspend the retry loops while no network is available. Instead of burning through backoff intervals against an unreachable network, the reconnection loops park until the OS reports a usable network again. Reconnect immediately on a network switch. When the OS hands us a new network, connections bound to the old one are swept and re-dialed right away, rather than waiting for a timeout to notice they are dead. |
||
|
|
93e97f4bf1 |
[doc] Agent network docs update (#7020)
* [docs] Update agent-network docs for management-owned pricing The docs still described the retired proxy-side pricing: pricing.Loader, pricing_path, MiddlewareDataDir, embedded defaults_pricing.yaml, and the symlink-safe Unix loader. Rewrite them for the current design — management synthesizes the whole table and ships it in cost_meter's ConfigJSON, so the proxy carries no price list and has nothing to reload. |
||
|
|
58c09ead21 | [management] Document mutual exclusivity of policy rule ports and port_ranges (#7158) | ||
|
|
ebfdf7d7b8 |
[management] Rework Agent Network endpoint identity and settings bootstrap (#7085)
Store the per-account gateway endpoint as {domain, proxy_address} with a
global unique index on the full hostname; dedicated = (domain ==
proxy_address). Bootstrap becomes an explicit POST carrying exactly one
of proxy_address (server allocates an adjective-noun label beneath it)
or endpoint (claimed verbatim, address-first); provider create loses its
bootstrap side effect. PUT is a full replace with every field required —
the immutable identity fields must be echoed unchanged and a mismatch is
rejected with 422. A guarded DELETE releases the endpoint: refused with
412 while providers exist or a proxy is actively serving the endpoint
hostname (matched case-insensitively); re-creating bootstraps fresh. A
self-addressed pin excludes its address from the account's cluster allow
list, and the live mapping update path now addresses the serving proxy
from the synthesized service. Existing rows are migrated on all three
store engines.
|
||
|
|
e8671a811d | [client, relay] Migrate relay QUIC tracer to qlog and bump quic-go to 0.59.1 (#7124) | ||
|
|
5584f8ef0a | [client] Add strict anonymization level and MAC anonymization to debug bundles (#7102) | ||
|
|
19a6cedfff |
[relay] randomize the relay reconnect backoff (#7067)
## Describe your changes The relay client's reconnect backoff was constructed without a `RandomizationFactor` ([guard.go:156-165](https://github.com/netbirdio/netbird/blob/main/shared/relay/client/guard.go#L156-L165)), so it kept the zero value: every client that lost the same relay server retried on the identical 2/4/8/16/32/60, potentially in waves. It was the only exponential backoff in the codebase without a randomization factor. Use `backoff.DefaultRandomizationFactor`, as random factor. ## Issue ticket number and link No public issue. Found while reviewing the relay reconnect path for the client-metrics review: 22k relay reconnection events in 24h, and the shared transport's own retry schedule was identical across all clients. ## Stack <!-- branch-stack --> ### Checklist - [x] Is it a bug fix - [ ] Is a typo/documentation fix - [ ] Is a feature enhancement - [ ] It is a refactor - [x] Created tests that fail without the change (if possible) > By submitting this pull request, you confirm that you have read and agree to the terms of the [Contributor License Agreement](https://github.com/netbirdio/netbird/blob/main/CONTRIBUTOR_LICENSE_AGREEMENT.md). ## Documentation Select exactly one: - [ ] I added/updated documentation for this change - [x] Documentation is **not needed** for this change (explain why) Internal retry-timing change with no user-visible surface: no CLI flag, configuration option or API field is added or altered, and the mean reconnect delay is unchanged. ### Docs PR URL (required if "docs added" is checked) Paste the PR link from https://github.com/netbirdio/docs here: N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved reconnect timing to distribute repeated connection attempts more evenly and reduce synchronized retry spikes. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
bc7a15ab71 |
[management] Align agent-network API contracts for API clients (#7026)
## Describe your changes Work on the Terraform provider (terraform-provider-netbird #177–#183) surfaced places where the agent-network API broke its own contracts or deviated from the conventions the rest of the management API follows, forcing client-side workarounds. Settings reads now follow the settings-endpoint convention: GET always answers with a JSON object. Before bootstrap it returns the defaults with an empty cluster/subdomain/endpoint (previously 200 with a JSON `null` body, while the spec said 404). The settings PUT can bootstrap the account by carrying a `cluster` — previously the row could only come into existence through the first provider create, and a settings-first setup was impossible; a differing cluster on a bootstrapped account is rejected instead of silently ignored. PUT remains full-state. The provider PUT schema promised omit-preserves semantics for several operator-editable fields that the handler never delivered (it builds the row from the request, like every other update handler). The schema wording now matches the shipped full-state behavior; only the api_key (secret) and session keys stay preserved by the manager. Identity headers are always present in provider responses so an explicitly cleared value round-trips as an empty string. The Go REST client gains the full agent-network surface (catalog, providers, policies, guardrails, budget rules, settings), including a shim translating the legacy 200+`null` settings body from older servers into an `IsNotFound` error. Note for reviewers: the dashboard special-cased the `null` settings body; it needs a small follow-up for the new defaults response (in progress). |
||
|
|
075b319fb3 |
[client, android] Pull fresh TUN settings on Android rebuild (#6991)
## Describe your changes Pull fresh TUN settings on Android rebuild instead of push The Android TUN rebuild consumed state pushed through notifications and a Java-side snapshot, and both sources were unreliable. The DNS search-domain notifier fired OnNetworkChanged with an empty string, which the rebuild handler treated as the new route list, so any search domain change rebuilt the TUN with zero routes and cut all tunnel traffic. The rebuild also reused the search domains cached at the last establish, so search domain updates never reached the TUN at runtime. Make the notification a pure trigger and let the Java side pull a fresh snapshot instead. Expose GetTunSettings on the Android SDK client: it returns the current TUN route ranges, derived on demand by the route manager from the client routes, the exit-node selection and the fake IP blocks, together with the DNS search domains. The route notifier keeps only its last-announced baseline to suppress triggers for unchanged syncs; the TUN route state is owned by the route manager. SearchDomains now locks the DNS server mutex since the pull arrives from a Java thread. Requires the matching android-client change that switches recreateTUN to the pull API. ## Issue ticket number and link ## Stack <!-- branch-stack --> ### Checklist - [x] Is it a bug fix - [ ] Is a typo/documentation fix - [ ] Is a feature enhancement - [ ] It is a refactor - [ ] Created tests that fail without the change (if possible) - [ ] This change does **not** modify the public API, gRPC protocols, functionality behavior, CLI / service flags, or introduce a new feature — **OR** I have discussed it with the NetBird team beforehand (link the issue / Slack thread in the description). See [CONTRIBUTING.md](https://github.com/netbirdio/netbird/blob/main/CONTRIBUTING.md#discuss-changes-with-the-netbird-team-first). > By submitting this pull request, you confirm that you have read and agree to the terms of the [Contributor License Agreement](https://github.com/netbirdio/netbird/blob/main/CONTRIBUTOR_LICENSE_AGREEMENT.md). ## Documentation Select exactly one: - [ ] I added/updated documentation for this change - [x] Documentation is **not needed** for this change (explain why) ### Docs PR URL (required if "docs added" is checked) Paste the PR link from https://github.com/netbirdio/docs here: https://github.com/netbirdio/docs/pull/__ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added access to current TUN route ranges and DNS search domains. - TUN settings are returned in a mobile-friendly format for easier integration. - **Improvements** - Route changes are detected and synchronized more reliably. - Current routing information now reflects active routes, including supported fake-IP ranges. - Simplified network initialization for more consistent startup behavior. - **API Changes** - Removed the obsolete network-map retrieval method from the management client interface. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
f9b412228e | [management] fix handling of empty network map during decode and encode (#6987) | ||
|
|
0780a806f2 | [management, proxy] Management-owned LLM pricing: file-backed defaults + (#6965) |