BuildApiBlackBoxWithDBState[AndPeerChannel] built the account manager,
telemetry metrics, and API handler on context.Background() and registered
no cleanup. Every background loop those components start
(AccountRequestBuffer.processGetAccountRequests, the telemetry P95
flushers, PATUsageTracker.reportLoop, APIRateLimiter.cleanupLoop, proxy
service cleanup, cache janitors) exits only on ctx.Done(), so on a
never-cancelled context they ran forever and piled up across the package
— along with each server's sql.DB connection pool.
Over a package run that builds ~150 servers this exhausts DB connections
against the real Postgres/MySQL backends, so per-test store setup crawls
until the suite trips the 20m go-test timeout (seen as timeouts in
Management/Integration (postgres) and Management/Unit (mysql); the
in-process sqlite variants finish before it bites).
Give each helper a cancellable context tied to t.Cleanup(cancel) so the
manager and its goroutines/pools wind down when the test ends. Test-only
change; production already cancels the server context on shutdown.
Adds an e2e suite (e2e/remotejobs) that runs on the container harness and
exercises the two stacked PRs end-to-end against a live management server
and a real client:
- Remote-jobs opt-in (#7153): a peer that ran plain `netbird up` reports
remote_jobs_allowed=false via the peers API, and the client refuses a
streamed job ("remote jobs are not enabled on this peer"). After
`netbird up --allow-remote-jobs`, the flag flips to true on the API and
the same job is accepted for execution.
- Bundle job parameters (#7147): an unknown anonymize_level is rejected at
job creation, and a messy-but-valid value (' Strict ') is normalized to
'strict' in the stored job the API returns.
Adds a small harness helper, Client.Up(extraArgs...), to re-run
`netbird up` with flags so the opt-in can be toggled mid-test without
recreating the container.
TestSqlStore_SavePeer reflects over every field of PeerSystemMeta via
PopulateAll and asserts the count, so that a newly added metadata field
forces the author to confirm it round-trips through the store. This PR
added Flags.RemoteJobsAllowed (a value bool inside the JSON-serialized
Flags), taking the recursive leaf count from 32 to 33. The field does
round-trip via the existing Flags JSON serializer, so update the expected
count. Fixes the deterministic Management/Unit store failure on all
backends.
Populate DebugBundleUploadURL with a token-bearing value and assert the
rendered bundle contains neither the field name nor the token, in both
anonymize modes. The excluded-map entry only skips the missing-field
check; this guards against the value being serialized by a future change.
Two review follow-ups on the MDM debug-bundle upload override:
- applyMDMPolicy returned early on an empty policy before reaching the
clear-on-absent path, so a policy that dropped every key left a stale
upload target directing bundles on a reused Config. Resolve the override
unconditionally, up front, so an absent, empty, or invalid value fails
closed to "" (falling back to the management-supplied or default target).
- Extract the resolution into mdmDebugBundleUploadURL, dropping the outer
function's cognitive complexity from 26 to 20 (SonarCloud gate is 25).
Tests cover the empty-policy and invalid-URL clearing paths.
- Never log the MDM-provided upload URL (it can embed credentials or
signed query tokens): mark the key secret so it is redacted, and drop
the raw value from the invalid-URL warning.
- Clear DebugBundleUploadURL when a replacement policy no longer carries
the key, so a removed override can never keep directing uploads to a
previously-enforced host; covered by a policy-replacement test.
- macOS docs: state "https URL with a host" consistently, and make the
managed-plist helper fail closed on an invalid allowRemoteJobs value
(emit false rather than dropping the key).
TestAddConfig_AllFieldsCovered fails on any new Config field that is
neither rendered nor excluded. Render RemoteJobsAllowed alongside the
other collection toggles, and exclude DebugBundleUploadURL with a
justification: it is an MDM-provided URL that can carry credentials or
query tokens, so it stays out of the shared bundle.
Cover the opt-in defaulting off for both new and legacy configs (the key
difference from the SSH default), and that applyMDMPolicy enables the flag
from allowRemoteJobs and applies debugBundleUploadURL, rejecting a
non-https override.
Add allowRemoteJobs (bool) and debugBundleUploadURL (string) to the
remaining MDM policy schema artifacts so administrators can enforce them
through every supported channel: the Windows ADMX/ADML templates, the
macOS .mobileconfig and the bare-plist helper, and the .reg example. The
macOS managed-preferences plist was covered with the code change.
The dashboard needs to know which peers have opted out of remote jobs so
it can reflect that in the UI, the same way it surfaces the SSH server
flag. Report RemoteJobsAllowed as peer system-info: the client sets it on
the reported flags (like ServerSSHAllowed), the proto Flags message
carries it, and management decodes it onto the peer meta and exposes it
on the peers API as remote_jobs_allowed.
Kept out of the components/network-map path: unlike ServerSSHAllowed it
does not participate in firewall-rule calculation, so it only rides the
reporting flags, not ComponentPeer.
Remote jobs (debug bundles requested by the management server) run on the
peer with no local consent. This makes them an explicit opt-in, mirroring
the SSH-server opt-in: an --allow-remote-jobs flag persisted in the client
config, defaulting off. Enabling it off->on crosses the user-to-root
boundary and is refused for unprivileged IPC callers by the daemon gate,
the same way enabling the SSH server is. When disabled, the job-stream
handler refuses every job before doing any work.
Because the flag is admin-controlled, it is also MDM-managed: the
allowRemoteJobs policy key can enable or lock it, and a user SetConfig that
diverges from an enforced value is rejected like the other managed fields.
A second MDM key, debugBundleUploadURL, overrides the debug-bundle upload
service for remote jobs, taking precedence over the management-supplied
value (MDM > management > default). This lets an operator pin uploads to a
trusted host regardless of what management requests. The override is
validated as an https URL with a host, the same as the management value.
Defaulting the opt-in off is a behavior change: existing deployments that
rely on management-triggered debug bundles must opt in (flag or MDM) before
they work again.
Two review follow-ups.
The API validated anonymize_level after trimming and lowercasing but
persisted the value verbatim, so " default " passed as default yet reached
the client — which only lowercases — as an unrecognized value it resolves
to strict. Persist the normalized form so what was validated is what the
client parses.
The remote debug bundle job forwarded the management-supplied upload URL
to the uploader unchecked and logged it at info level, where it can leak a
host, credentials, or query tokens. Reject a malformed or non-https URL
before generating the bundle, and keep the URL out of the info-level line
while leaving the full parameters at debug. The accepted host is left
unrestricted for now, pending a decision on management-directed uploads.
The client resolves an unknown anonymization level to strict, a fail-safe
that is right for the wire but wrong for the API boundary: a caller that
misspells the level should be told so at job creation, not have a
different level than they asked for applied silently on the peer.
Reject any anonymize_level other than the known wire forms when building
a bundle job; an omitted or empty value still crosses the wire as empty
and defaults on the client. The accepted forms are taken from the client
anonymize package so the API and the consumer cannot drift.
PR #7102 added an anonymization level to debug bundles and the
anonymize_level proto field, but nothing on the management side ever set
it: the remote-job builder dropped the field and the REST schema never
exposed it, so a remotely triggered bundle always ran at the default
level regardless of what an operator asked for. The upload destination
for remote jobs was likewise fixed to the default upload server, with no
way to direct a bundle to a self-hosted one.
Expose anonymize_level and a new upload_url on the REST BundleParameters
and the management proto, and map both onto the job request streamed to
the client. Both are optional: an omitted value crosses the wire as the
empty string, which the client resolves to its own defaults — the
default anonymization level and the default upload server — matching how
the netbird CLI defaults the same inputs.
Prepares the repository for the release-branch process agreed internally:
one long-lived release-0.N branch per minor, with fixes backported by
cherry-pick and patch releases tagged from the branch.
Pushes to release-* branches now run the Release workflow and publish
immutable sha-* container images, the way pushes to main already do, so
a release branch can be tested before it is tagged. Release branches
never publish the floating "main" image tag. The push-to-main CI
workflows (Go tests on all platforms, frontend UI, install script,
mobile/wasm validation, infrastructure files, license check) also run
on release-* pushes; pull request triggers were already unfiltered, so
backport PRs were covered — this closes the post-merge gap.
Releases are no longer marked latest before signing: make_latest is
now false in all four goreleaser configs, so a release stays published
but not latest until the signing pipeline uploads the signed Windows
and macOS artifacts and marks it latest itself. Previously the release
became GitHub's "Latest release" at publish time, and the download
endpoints that resolve through the latest-release API could serve a
release whose signed installers did not exist yet. prerelease: auto
additionally labels rc tags as prereleases, so a release candidate can
never take the latest slot.
The trigger_sync_tag job is removed: it dispatched a downstream
image build on every v* tag (release candidates included), which would
race the deliberate release-branch build on every release. The android
and ios submodule bumps are unchanged.
Also sets perennial-regex = "^release-" so git-town never syncs or
ships a release branch into main.
The NSIS installer deleted HKLM/HKCU CurrentVersion\Run values it never
writes, which matches common AV heuristics for unwanted Run-key
manipulation and is suspected to contribute to Windows Defender and
third-party antivirus false positives on the installer.
Drop all autostart registry deletions from both the install and
uninstall sections so the installer only touches keys it creates
itself. Cleanup of the legacy machine-wide entry written by old
installers is left to documentation.
Extends the approach of the closed PR #6735, which only removed the
per-user deletion on uninstall.
e2e/harness documents itself as feature-agnostic, but three details
assumed the caller lives in this repo, so the terraform provider's
acceptance suite would otherwise carry a second harness for the same
product.
repoRoot took the first module root above the working directory as the
Docker build context, which from another module is the caller's own
root, with no combined/Dockerfile.multistage in it. It now requires that
ancestor to be this module, and otherwise asks the go tool for the
source: for a dependent, the extracted directory of the version it pins,
so the server matches the client library it was compiled against. That
lookup uses -mod=readonly, since automatic vendor mode otherwise reports
an empty Dir.
Geolocation was disabled unconditionally. Agent-network ingest does not
use it, but location-based posture checks need the database, and a rule
management cannot evaluate fails rather than passing.
StartClient pinned one network alias and set no hostname, so a second
agent could not start and a peer's name was arbitrary. Management
records that hostname, making it the peer's name in the API.
The client entrypoint is copied with an explicit mode: git tracks it
100755, but the module cache extracts 0444, so a dependent's build
produced a container exiting with "permission denied".
Adds CombinedOption, WithGeolocation, WithServerEnv, ClientOption and
WithClientName.
- Align default names and reuse same environment variables
- With the uploads now targeting the same stable/yum paths as the GTK4
packages, two packages named netbird-ui with the same version and arch
would collide in the repo indexes. Give the GTK3 variant its own
package name and mark the two as conflicting alternatives.
---------
Co-authored-by: Zoltan Papp <zoltan.pmail@gmail.com>
People who only ever reach private services through the reverse proxy were
invisible to activity accounting. Active users are counted from user.LastLogin
or from the LastSeen of a peer they own, and neither column was written on the
proxy paths — so a person signing in via SSO to a proxied service, or a peer
serving one over the mesh, never showed up in the 24 hour numbers.
Both writes now happen where the proxy already authenticates:
- GenerateSessionToken stamps LastLogin after the session token is signed,
the same column and the same way the dashboard and device login paths do.
- ValidateTunnelPeer stamps the calling peer's LastSeen, the column its owner
activates through.
The policy lives in a new reverseproxy/activity manager rather than in the gRPC
service, matching the module layout the other reverse proxy domains use. It
skips what can never count — service users, embedded proxy peers and WASM
clients — and throttles peer writes to once an hour, well inside the window
accounting asks about and far above the proxy's five minute tunnel cache.
The peer write is a single indexed UPDATE that touches only
peer_status_last_seen. Connected and SessionStartedAt are left alone so the
session-ownership fencing MarkPeerConnectedIfNewerSession relies on is never
disturbed, and the timestamp comes from the database clock rather than the
caller, for the same reason the other status writers take it from there. The
caller's cutoff travels into the statement's WHERE, so concurrent requests for
one peer collapse into a single write instead of each acting on its own stale
read, and a peer that was never seen — NULL last seen, since Status is an
embedded pointer — still records its first activity.
Nothing outside the reverse proxy changes behaviour: the only addition
elsewhere is the RefreshPeerLastSeen store method the manager calls.
At proxy connect time, the declared cluster address is validated for
shape and checked for availability (`IsClusterAddressAvailable`), and
from then on the declaration is what routes the cluster's mappings to
the connection. Deployments that embed management through the
integrations seam may need a policy on that claim — deciding which
credential is allowed to declare which address.
This adds an optional `ProxyConnectAuthorizer` hook on
`ProxyServiceServer`, following the pattern of the existing `Set*` seams
(`SetServiceManager`, `SetAgentNetworkSynthesizer`,
`SetAgentNetworkLimitsService`, `SetProxyController`):
- A nil-able interface field plus `SetProxyConnectAuthorizer`, guarded
by the existing mutex.
- One call at the end of `validateProxyConnect`, so both
`GetMappingUpdate` and `SyncMappings` are covered by a single site.
- **Nothing installs it by default** — with the hook unset (always, in
this repo), behavior is byte-for-byte unchanged, which the tests pin.
Design details:
- The authorizer runs **last** — after input validation and the
availability check — and **outside** the account-scoped branch, so
management-wide tokens and token-less connects are also presented to it
rather than bypassing policy.
- The authorizer receives the presented `*types.ProxyAccessToken` (nil
when none), the proxy ID, and the declared address. Everything it needs
is already in the request/context; no proto or schema change.
- A plain error from the authorizer surfaces as `PermissionDenied`,
keeping an authorization rejection distinguishable from the
`AlreadyExists` used for address conflicts in proxy logs. A status error
passes through unchanged so implementations can pick their own code.
Store the per-account gateway endpoint as {domain, proxy_address} with a
global unique index on the full hostname; dedicated = (domain ==
proxy_address). Bootstrap becomes an explicit POST carrying exactly one
of proxy_address (server allocates an adjective-noun label beneath it)
or endpoint (claimed verbatim, address-first); provider create loses its
bootstrap side effect. PUT is a full replace with every field required —
the immutable identity fields must be echoed unchanged and a mismatch is
rejected with 422. A guarded DELETE releases the endpoint: refused with
412 while providers exist or a proxy is actively serving the endpoint
hostname (matched case-insensitively); re-creating bootstraps fresh. A
self-addressed pin excludes its address from the account's cluster allow
list, and the live mapping update path now addresses the serving proxy
from the synthesized service. Existing rows are migrated on all three
store engines.
## Describe your changes
The native_webview2loader build tag embedded Microsoft's
WebView2Loader.dll via //go:embed. Those DLLs are gitignored in the fork
and go mod vendor resolves embed patterns regardless of build
constraints, so vendoring the module failed on the missing files when
packaging for openSUSE.
The fork now removes that branch along with the go-winloader dependency,
which drops out of the module graph here. The default GoWebView2Loader
path is unaffected.
## Issue ticket number and link
<!--
Required for anything that changes behavior. Link the issue (or the
validated
discussion it came from) that the NetBird team already agreed on. See
https://github.com/netbirdio/netbird/blob/main/CONTRIBUTING.md#ticket-first-pr-second
-->
## Stack
<!-- branch-stack -->
### Checklist
- [x] Is it a bug fix
- [ ] Is a typo/documentation fix
- [ ] Is a feature enhancement
- [ ] It is a refactor
- [ ] Created tests that fail without the change (if possible)
- [ ] I ran and tested this change locally — I did not rely on CI to
find out whether it works
- [ ] This PR has a single purpose (not a fix + refactor + feature in
one)
- [ ] This change is a trivial fix, **OR** it links an issue the NetBird
team agreed on beforehand. Changes to the public API, gRPC protocols,
functionality behavior, CLI / service flags, or new features always need
that agreement first. See
[CONTRIBUTING.md](https://github.com/netbirdio/netbird/blob/main/CONTRIBUTING.md#ticket-first-pr-second).
> By submitting this pull request, you confirm that you have read and
agree to the terms of the [Contributor License
Agreement](https://github.com/netbirdio/netbird/blob/main/CONTRIBUTOR_LICENSE_AGREEMENT.md).
## Documentation
Select exactly one:
- [ ] I added/updated documentation for this change
- [x] Documentation is **not needed** for this change (explain why)
### Docs PR URL (required if "docs added" is checked)
Paste the PR link from https://github.com/netbirdio/docs here:
https://github.com/netbirdio/docs/pull/__
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Chores**
* Updated the application framework revision.
* Removed an unused dependency.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Describe your changes
The UI shells out to xdg-open to launch the external browser for the SSO
verification page, which the embedded webview cannot open inline, and to
reveal the debug bundle in the file manager.
client/ui/build/linux/nfpm/nfpm.yaml lists xdg-utils for every package
format, but the released packages are built from the goreleaser configs,
where it was missing: the GTK and WebKitGTK dependencies carried over
and xdg-utils did not. Add it to all four nfpm dependency lists.
## Issue ticket number and link
<!--
Required for anything that changes behavior. Link the issue (or the
validated
discussion it came from) that the NetBird team already agreed on. See
https://github.com/netbirdio/netbird/blob/main/CONTRIBUTING.md#ticket-first-pr-second
-->
## Stack
<!-- branch-stack -->
### Checklist
- [x] Is it a bug fix
- [ ] Is a typo/documentation fix
- [ ] Is a feature enhancement
- [ ] It is a refactor
- [ ] Created tests that fail without the change (if possible)
- [ ] I ran and tested this change locally — I did not rely on CI to
find out whether it works
- [ ] This PR has a single purpose (not a fix + refactor + feature in
one)
- [ ] This change is a trivial fix, **OR** it links an issue the NetBird
team agreed on beforehand. Changes to the public API, gRPC protocols,
functionality behavior, CLI / service flags, or new features always need
that agreement first. See
[CONTRIBUTING.md](https://github.com/netbirdio/netbird/blob/main/CONTRIBUTING.md#ticket-first-pr-second).
> By submitting this pull request, you confirm that you have read and
agree to the terms of the [Contributor License
Agreement](https://github.com/netbirdio/netbird/blob/main/CONTRIBUTOR_LICENSE_AGREEMENT.md).
## Documentation
Select exactly one:
- [ ] I added/updated documentation for this change
- [x] Documentation is **not needed** for this change (explain why)
### Docs PR URL (required if "docs added" is checked)
Paste the PR link from https://github.com/netbirdio/docs here:
https://github.com/netbirdio/docs/pull/__
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Added required desktop integration support to Debian and RPM packages.
* Ensured GTK3 packages include the same runtime support for opening
links and files through the system.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
A user in the Pending Approval state could complete SSO and reach any
SSO-protected reverse proxy service distributed to a group they belong
to, including the All Users group. The reverse proxy authorization path
checked the session token signature, that the user exists, that the
user's account matches the service's account, and group membership —
never the user's account status. The REST API (`permissions/manager.go`)
and peer registration both gate on that state, but the proxy gRPC
service does not go through the permissions manager, so neither gate
applied. A pending user is persisted as blocked and pending approval, so
blocked users reached those services the same way.
`ValidateSession` now denies on account status, reporting
`pending_approval` or `user_blocked` so the proxy access log and the
denied page carry the cause rather than a generic refusal.
`GenerateSessionToken` refuses to mint a token for such a user at all,
so the browser never receives a session cookie and the OIDC callback can
tell the user why instead of showing "Service configuration error".
`ValidateUserGroupAccess` and `ValidateTunnelPeer` close the same gap;
for the tunnel path this covers a user blocked after their peer was
registered, since peer group membership alone kept mesh-origin access
open.
A single helper produces both the denied reason for the RPC responses
and the sentinel error for the error-returning callers, so the four
entry points cannot drift apart. A user the store cannot resolve is
denied rather than passed through.
One thing deliberately left out: session cookies are validated locally
by the proxy against the service public key with no management
round-trip, so a cookie issued before a user is blocked stays valid
until it expires (24h by default). That is a revocation-propagation
problem rather than this authorization gap, and every option for it
(per-request validation with a cache, short-lived tokens with refresh,
push-based revocation) changes the proxy hot path or the
proxy/management protocol. Worth its own ticket.
## Describe your changes
`recordConnectionMetrics` mapped only `conntype.Relay` to `relay` and
let a `default` branch
record everything else as `ice`. That silently included `ICETurn` — an
ICE connection through
a TURN server, which
[`conn.isRelayed`](https://github.com/netbirdio/netbird/blob/main/client/internal/peer/conn.go#L788-L795)
itself counts as relayed — and `None`, the transient state set when the
relay drops
([conn.go:632](https://github.com/netbirdio/netbird/blob/main/client/internal/peer/conn.go#L632))
or the peer state is reset
([conn.go:757](https://github.com/netbirdio/netbird/blob/main/client/internal/peer/conn.go#L757)).
Both were reported as direct peer-to-peer, so the `ice` share overstated
direct connections on
every platform.
The mapping now lists every priority explicitly and emits `ice_p2p`,
`ice_turn`, `relay` or
`unknown`. The new values deliberately do not reuse `ice` to avoid
ambuguity.
## Issue ticket number and link
No public issue. Found while reviewing the first production sample of
client metrics: 38% of iOS
connection events were tagged `ice` on a platform that forces relay by
default, which traced back
to the `default` branch at
[client/internal/peer/conn.go#L963-L968](https://github.com/netbirdio/netbird/blob/main/client/internal/peer/conn.go#L963-L968).
## Stack
<!-- branch-stack -->
### Checklist
- [x] Is it a bug fix
- [ ] Is a typo/documentation fix
- [ ] Is a feature enhancement
- [ ] It is a refactor
- [x] Created tests that fail without the change (if possible)
> By submitting this pull request, you confirm that you have read and
agree to the terms of the [Contributor License
Agreement](https://github.com/netbirdio/netbird/blob/main/CONTRIBUTOR_LICENSE_AGREEMENT.md).
## Documentation
Select exactly one:
- [] I added/updated documentation for this change
- [x] Documentation is **not needed** for this change (explain why)
Internal metrics documentation only, in
`client/internal/metrics/infra/README.md`: the four
`connection_type` values with their derivation, and a note that pre-fix
`ice` samples are not
comparable with `ice_p2p`. No public API, CLI or configuration change,
so no netbirdio/docs PR.
### Docs PR URL (required if "docs added" is checked)
Paste the PR link from https://github.com/netbirdio/docs here:
N/A
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Summary by CodeRabbit
* **New Features**
* Connection metrics now distinguish direct peer-to-peer, TURN-assisted,
relay, and unknown connection types.
* Metrics include clearer connection and peer identification details.
* **Documentation**
* Updated connection timing metric values, traffic semantics, priority
behavior, and historical data guidance.
* **Bug Fixes**
* Unset or unrecognized connection priorities are no longer incorrectly
classified as peer-to-peer.
* Unknown-transport metrics are skipped to prevent misleading connection
data.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
The migration wizard detected the community deployment only by the
Docker Hub image prefix (netbirdio/netbird-server), so deployments
installed from the ghcr.io mirror failed with "Could not find a service
running netbirdio/netbird-server*".
This broadens the server and dashboard detection to also accept
ghcr.io/netbirdio/... images. The regexes are anchored at the tag/digest
separator so Enterprise images (netbird-server-cloud, dashboard-cloud)
are still rejected, an already-migrated deployment must not be detected
as a community one. Error messages updated to mention both forms.
## Describe your changes
The Conn struct is reused across lazy-connection deactivate/activate.
Close
cancels the WireGuard watcher (via wgWatcherCancel, and ctxCancel also
tears
down its context) but left conn.wgWatcher pointing at the stopped
instance.
enableWgWatcherIfNeeded skips while conn.wgWatcher is non-nil, so the
next Open
never started a fresh watcher: once a lazy connection had idled and
woken, the
peer ran with no watcher at all — no WireGuard handshake-timeout
detection and
none of the escalation that depends on it.
Clear conn.wgWatcher and conn.wgWatcherCancel in Close so the next Open
re-arms
a fresh watcher.
## Issue ticket number and link
<!--
Required for anything that changes behavior. Link the issue (or the
validated
discussion it came from) that the NetBird team already agreed on. See
https://github.com/netbirdio/netbird/blob/main/CONTRIBUTING.md#ticket-first-pr-second
-->
## Stack
<!-- branch-stack -->
### Checklist
- [X] Is it a bug fix
- [ ] Is a typo/documentation fix
- [ ] Is a feature enhancement
- [ ] It is a refactor
- [ ] Created tests that fail without the change (if possible)
- [ ] I ran and tested this change locally — I did not rely on CI to
find out whether it works
- [ ] This PR has a single purpose (not a fix + refactor + feature in
one)
- [ ] This change is a trivial fix, **OR** it links an issue the NetBird
team agreed on beforehand. Changes to the public API, gRPC protocols,
functionality behavior, CLI / service flags, or new features always need
that agreement first. See
[CONTRIBUTING.md](https://github.com/netbirdio/netbird/blob/main/CONTRIBUTING.md#ticket-first-pr-second).
> By submitting this pull request, you confirm that you have read and
agree to the terms of the [Contributor License
Agreement](https://github.com/netbirdio/netbird/blob/main/CONTRIBUTOR_LICENSE_AGREEMENT.md).
## Documentation
Select exactly one:
- [ ] I added/updated documentation for this change
- [X] Documentation is **not needed** for this change (explain why)
### Docs PR URL (required if "docs added" is checked)
Paste the PR link from https://github.com/netbirdio/docs here:
https://github.com/netbirdio/docs/pull/__
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved connection cleanup by fully releasing WireGuard watcher
resources when a connection closes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
The community and enterprise bootstrap scripts now generate a dedicated
`server.auth.sessionCookieEncryptionKey` from 32 random bytes for fresh
installations.
The key is Base64-encoded, persisted independently from the datastore
encryption key, and reused from the generated configuration after
restarts.
The community script now also creates `config.yaml` with mode `0600`
before writing it, matching the existing enterprise behavior.
The server already supports the session cookie encryption key, but the
bootstrap scripts left it unset.
This adds defense-in-depth for newly generated deployments while
preserving the existing server-side nonce validation.
Because this only changes newly generated configuration, existing
installations and sessions are unchanged.
A focused regression test checks key presence, decoded length,
separation from the datastore key, YAML placement, and file mode for
both scripts.
It is included in the infrastructure workflow.
## Describe your changes
The relay client's reconnect backoff was constructed without a
`RandomizationFactor`
([guard.go:156-165](https://github.com/netbirdio/netbird/blob/main/shared/relay/client/guard.go#L156-L165)),
so it kept the zero value: every client that lost the same relay server
retried on the identical 2/4/8/16/32/60, potentially in waves.
It was the only exponential backoff in the codebase without a
randomization factor.
Use `backoff.DefaultRandomizationFactor`, as random factor.
## Issue ticket number and link
No public issue. Found while reviewing the relay reconnect path for the
client-metrics review:
22k relay reconnection events in 24h, and the shared transport's own
retry schedule was
identical across all clients.
## Stack
<!-- branch-stack -->
### Checklist
- [x] Is it a bug fix
- [ ] Is a typo/documentation fix
- [ ] Is a feature enhancement
- [ ] It is a refactor
- [x] Created tests that fail without the change (if possible)
> By submitting this pull request, you confirm that you have read and
agree to the terms of the [Contributor License
Agreement](https://github.com/netbirdio/netbird/blob/main/CONTRIBUTOR_LICENSE_AGREEMENT.md).
## Documentation
Select exactly one:
- [ ] I added/updated documentation for this change
- [x] Documentation is **not needed** for this change (explain why)
Internal retry-timing change with no user-visible surface: no CLI flag,
configuration option
or API field is added or altered, and the mean reconnect delay is
unchanged.
### Docs PR URL (required if "docs added" is checked)
Paste the PR link from https://github.com/netbirdio/docs here:
N/A
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved reconnect timing to distribute repeated connection attempts
more evenly and reduce synchronized retry spikes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->