* [management] Require a private proxy cluster for cluster and direct upstream targets
Cluster targets and direct upstream targets make the proxy dial the
upstream from its own host network instead of through the embedded
NetBird client. Only clusters running in private mode are meant to do
that, but the service API accepted these targets on any cluster.
Service create and update now reject such targets unless the service's
proxy cluster reports the private capability. An unreported capability
is treated as unsupported.
* [management] Require every proxy in the cluster to be private
The private capability is aggregated as any-true, so a cluster where
only one proxy runs in private mode passed the check. The mapping is
delivered to every proxy in the cluster, so the non-private ones would
serve cluster and direct upstream targets from their host network too.
Validate these targets against a unanimous aggregation instead. The
existing any-true lookup stays as is for the dashboard flags and the
agent network gateway.
* [client] Discover interfaces lazily in stdnet instead of at construction
stdnet.NewNet and NewNetWithDiscover ended with
return n, n.UpdateInterfaces()
handing back a non-nil *Net together with the discovery error. Three of the
five call sites (Engine.newWgIface, ice.NewAgent, SingleSocketUDPMux) logged
the error and kept using the instance, which is only safe as long as the
instance still works after a failed discovery.
That stopped being true when Interfaces() gained a lazily refreshed cache:
updateInterfaces sets lastUpdate only on success, so after a failed
construction the 30s cache guard never holds and Interfaces() returns an
error rather than the empty list it used to return. Feeding such an instance
to pion is worse than passing nothing at all - ice.NewAgent falls back to its
own stdnet when Net is nil, and the interface blacklist is applied separately
through AgentConfig.InterfaceFilter, so the fallback loses nothing. Instead,
a transient discovery failure (the Android bridge at boot, or an interface
disappearing between net.Interfaces() and Interface.Addrs()) turned into a
hard "error getting local interfaces" from ice.NewAgent, and aborted the STUN
and TURN probes, which never even need the interface list.
Since the accessors already refresh a stale cache on demand, the eager
discovery in the constructors is redundant: drop it, make both constructors
infallible, and let the discovery error surface at the call that actually
needs the interfaces. UpdateInterfaces had no callers left and is not part of
transport.Net, so it is removed along with it.
InterfaceByIndex and InterfaceByName read the cached slice directly and never
refreshed it, so they would have kept reporting ErrInterfaceNotFound forever
on an instance whose first discovery failed. They now go through the same
refresh path as Interfaces().
* [client] Warm the stdnet interface cache at construction
Moving discovery to first use regressed the privileged suites on the three
platforms that always build an ICE bind: Darwin, FreeBSD and Windows time out
in TestWGIface_UpdateAddr, TestRecreation, TestEngine_SSH and
TestEngine_MultiplePeers, while Linux stays green because a host with the
WireGuard kernel module takes the kernel-device branch and never drives the
mux that asks for interfaces.
interfaceFilter probes with wgctrl every interface the disallow list does not
already exclude. Discovering at construction ran that probe before the caller
had an overlay interface of its own; discovering at first use runs it after,
so on a userspace WireGuard platform the probe reaches the UAPI socket of the
same process. The tests reach it because they construct with a nil disallow
list, where the client passes DefaultInterfaceBlacklist and its own interface
is excluded by prefix.
Restore the original timing with an explicit warm-up. The constructors stay
infallible and the error is still reported by the accessor that needs the
interfaces, so the contract this branch is about is unchanged.
* Revert "[client] Warm the stdnet interface cache at construction"
This reverts commit 947e25288f.
* [client] Give the privileged tests the interface blacklist the client uses
The suites that create a WireGuard interface construct stdnet with a nil
disallow list, which the client never does: Engine passes
profilemanager.DefaultInterfaceBlacklist, whose "wt" and "utun" prefixes
exclude the overlay interface before the filter reaches its wgctrl probe.
With an empty list every interface reaches that probe, the one the test has
just created included, and on a userspace WireGuard platform the probe talks
to the UAPI socket of the same process. That is why Darwin, FreeBSD and
Windows timed out here while Linux, which takes the kernel-device branch on a
host with the module loaded, stayed green.
Pass the blacklist in both suites so they exercise the configuration the
client ships. client/iface declares the prefixes locally because
profilemanager imports it.
Also cover the constructors directly: the existing tests build the struct
literal, so nothing asserted that NewNet and NewNetWithDiscover leave the
cache cold.
* [client] Pass the blacklist in the remaining tests that build an interface
Same reason as the previous commit, four call sites it missed: engine_test,
the route manager and systemops suites, and the privileged DNS server suite
all construct stdnet with a nil disallow list and then create a WireGuard
interface. TestAddVPNRoute surfaced it on FreeBSD once the earlier two files
stopped timing out first.
client/internal/dns declares the prefixes locally; profilemanager imports
that package, so it cannot import profilemanager back.
* [misc] Move the FreeBSD port test to release 15.1
FreeBSD 15.0 reached end of life on 2026-09-30 and the ports tree marks
it unsupported since freebsd/freebsd-ports@ed90b23fe9 (2026-10-01), so
`make package` refuses to run on the 15.0 VM and the FreeBSD Port job
fails on every PR. The pinned vmactions/freebsd-vm v1.4.8 ships a 15.1
image, so only the release needs to move.
* [misc] Run the FreeBSD unit tests on release 15.1 too
The job installs binary packages instead of building from the ports
tree, so it kept passing on the EOL 15.0 image, but the client should
be tested on the same supported release the port is built on, and the
EOL image is only served from the archive mirror from now on.
Direct-upstream targets are dialled on the proxy host's network stack,
outside the embedded client's LAN blocking. A proxy that serves
untrusted accounts lets them reach the host's loopback, its LAN or
cluster, and the cloud metadata service through such a target.
NB_PROXY_DIRECT_UPSTREAM_BLOCK_PRIVATE adds a dialer control that
refuses addresses that are not globally reachable. It checks each
socket's resolved address just before connect, so hostnames and DNS
rebinding are covered, and IPv4 embedded in IPv6 addresses is checked
as IPv4. Refused dials are served as a 502. The setting defaults to
off for private and self-hosted proxies; an unparsable value turns it
on.
Test_ConnectPeers fails every few weeks on the Linux runner with a bare
"waiting for peer handshake timeout after 30s". The failing logs show
both kernel devices up and both peers configured within a second, then
nothing for 30 s, which is six retries of the 5 s handshake retransmit
and so a condition that lasted the whole window rather than a race.
The failure cannot be reproduced locally and the log cannot tell
whether initiations were sent, whether they arrived, or whether only
one direction worked.
On timeout the test now prints each device's view of its peer, the
endpoint, the byte counters and the last handshake, so the next
failure says which of those it is. The comment also states that the
peers are kernel devices on the runner and that the first initiation
of each side is always lost to the other side not knowing the peer
yet.
* [client] Migrate macOS cask template to Homebrew install steps
Homebrew deprecated the postflight and uninstall_preflight cask stanzas
in favour of the declarative *_steps DSL, so every brew command that
evaluates netbirdio/tap now prints deprecation warnings. Once the
deprecation becomes a disable the generated cask stops loading and
netbird-ui can no longer be installed or upgraded through Homebrew.
The *_steps blocks take JSON-serialisable steps run in a sandbox rather
than arbitrary Ruby, so system_command is re-expressed as run/remove.
The two postflight blocks merge into one because a cask carries only a
single instance, preserving the original order. set_permissions moves
from a hardcoded /Applications to base: :appdir, matching what the
installer invocation already did. The launchctl fallbacks keep their
tolerant semantics through must_succeed: false, and remove is a no-op
when the plist is absent.
(cherry picked from commit df3756151f)
* [client] Test the macOS Homebrew cask on a disposable runner
The cask template only runs on real macOS with Homebrew, sudo and
launchd, so changes to it have never been exercised before merge. This
job installs the rendered cask on a GitHub macOS runner, walks the
uninstall through a running, stopped and missing daemon, and reinstalls
over the tap's published legacy cask, which is the path every existing
user takes on their next upgrade.
The fixture is the published cask itself rather than a pinned version
and checksums, so the test follows each release instead of breaking at
the next one. The installer scripts inside the signed archives are not
under test, which is why their paths are left out of the trigger.
* [client] Address SonarCloud findings in the Homebrew cask test
Positional parameters move into local variables and the scenario switch
gains an explicit default, so an unknown scenario fails instead of
silently running the plain install and uninstall path.
* [client] Make the Homebrew cask test deterministic with a stub bundle
The released installer script opens the UI as root, which never returns
on a headless runner, so a test that installs the published archive
hangs until the job timeout. The cask itself never looks past two script
paths and a version argument, so the test now builds a stub bundle on
the runner, serves it from a local HTTP server and renders the template
against it. The scripts ship without the executable bit, which turns the
0755 check into proof that set_permissions ran, and the stub records the
version and uid it received. The published archives are still downloaded
to assert the two script paths exist, and the published cask still
supplies the legacy stanzas for the reinstall scenario.
* [client] Drop the launchctl stderr check from the Homebrew cask test
The test asserts what the cask template promises: install, uninstall and
no deprecation warnings. Whether the uninstall steps print launchctl
errors is a review remark on the template, not part of that contract.
* [client] Retry the daemon start in the Homebrew cask test stub
A reinstall runs the previous cask's bootout and the new postflight
within a second of each other. launchd is still tearing the old daemon
down at that point, so loading the same label again fails with EIO. The
stub now retries the start for up to fifteen seconds, and the test still
verifies afterwards that the daemon reached the running state.
---------
Co-authored-by: Daniele Casciani <d.casciani@genogra.com>
usage_viewer saw account-wide usage but only its own request logs, so
the people reviewing cost could not drill into the requests behind it.
The role now also holds Read on agent_network.logs, which makes the
access-log and session endpoints return every caller's rows instead of
self-scoping. Logs can contain captured prompts, so this widens what the
role exposes; policies, guardrails, budgets and settings stay hidden.
Co-authored-by: Misha Bragin <bangvalo@gmail.com>
Deleting an account left state behind that DeleteAccount's store
associations don't reach. The Agent Network tables outlived the account,
keeping its gateway domain claimed and its provider API keys stored. The
proxies kept serving its gateway until they next resynced. Cloud-side
state, such as managed proxy deployments, had no way to be cleaned up at
all.
Account deletion now runs registered hooks after the permission check
and before any users or data are removed. A failing hook aborts the
deletion. Agent Network registers one that tells the proxies to drop the
account's gateway mappings. The account's settings, providers, policies,
guardrails and budget rules are deleted in the account's transaction.
Consumption counters, and the access logs of deleted accounts, are left
to the background cleanup; usage records are kept.
* Add debug cpu start and stop commands to profile the daemon without a restart
* Restore test globals on every exit and stop the daemon in the cpu profile test
* Add a no-updown flag to debug for
* Enable sync response persistence with --no-updown and reset flags between debug test runs
* Reset flags of every command between debug test runs
* Reset slice flags with Replace in the debug test helper
* Explain a running CPU profile in debug for and document cpu start and no-updown limits
A worker that saw a new remote session ID rebuilt its agent and also
picked a new local ID. On the answer path nothing carries that ID back,
so the next offer made the remote see a changed session, rebuild, and
answer with yet another ID. Two peers kept tearing down working ICE
connections on every offer and answer; nearly every answer in the
affected logs carried a new remote session ID.
Only a local restart changes the local ID now: a failed negotiation, as
before, and an explicit Close, which previously kept the old ID and left
the remote answering from a negotiation this side had abandoned.
Following a remote restart keeps the ID the remote already knows, so the
pair settles after one rebuild, also against peers that still pick a
new ID when following a restart.
createClientEntry and NewMultiTransport clone the secure transport into
its insecure variant before newUpstreamTransport applies the configured
HTTP version. http.Transport.Clone runs the source's one-time protocol
setup, and at that point ForceAttemptHTTP2 is still false while a custom
DialContext is set, so net/http disables HTTP/2 on the source for good.
Setting ForceAttemptHTTP2 afterwards has no effect.
As a result every TLS-verified upstream has been served over HTTP/1.1
since the upstream HTTP version became configurable, whatever
NB_PROXY_UPSTREAM_HTTP_VERSION says, while skip-TLS-verify upstreams
kept HTTP/2. gRPC upstreams break outright: unary calls get a 502 and
streaming calls hang until the client gives up.
Apply the version to the base transport before cloning it, and add a
test that the direct and insecure transports both offer h2.
* Generalize PKCE verifier store into SingleUseStore
* Generalize PKCE verifier store into SingleUseStore
* Extend single-use store to generate one-time retrieval codes
* Hand off proxy OIDC session via one-time code instead of URL token
* Use the single-use store in integration tests
* Read active proxy versions by cluster
* Detect proxy clusters that support session codes
* Bind OIDC session handoff mode to signed state
* Deprecate legacy OIDC session token handoff
* Remove unrelated session code test stub
* fix tests
* fix merge
* Fix session code compatibility detection
* Isolate proxy session codes in shared cache
* bump min session version
* [client] Keep the delete-profile dialog open until the delete finishes
Confirming a profile deletion closed the dialog straight away and left the
daemon call running in the background. A slow delete then looked like nothing
had happened: the dialog was gone, the profile was still listed, and the row
only disappeared whenever the refresh landed.
The confirm dialog now owns the action. It stays open with the confirm button
spinning, closes once the call resolves, and surfaces a failure after it is
gone rather than behind it. Nothing on the daemon path carries a deadline, so
the wait is bounded in the dialog instead: Cancel comes back after five seconds
and the wait is abandoned at thirty, which keeps a hung daemon from trapping
the user in a modal that cannot be dismissed.
A loading button keeps its own variant colours rather than the disabled skin,
which dimmed the spinner to grey on the danger variant, and blocks input
through aria-disabled and a click guard instead.
* [client] Match the theme and anonymize pickers to the language switcher
* [client] Settle a confirm dialog only from the run that opened it
Cancelling at the stall point leaves the action running, and the provider is
mounted once: take() read whichever settler the ref held when the stale run
finally finished. Open another prompt in the meantime and that run answered it
— a hung delete that later resolved confirmed a profile switch nobody accepted,
and its timeout closed the new dialog with an unrelated error.
Each run now remembers the settler it was dispatched for and settles only while
the ref still points at it. A late arrival finds a stranger there and answers
nothing.
* [client] Add the shared Select the theme and anonymize pickers use
The picker rework landed without the component both pickers import, so the
branch did not compile. Add it, and name it for what it is: a select of a few
options, with nothing settings-specific about it, so it sits with the other
input controls rather than under a name that discourages reuse.
* [client] Drop the Select header comment
* [client] Fix cubic comments
* [client] Name the Select trigger with the option it is showing
* [client] Name the language trigger with the language it is showing
* [client] Hold the confirm dialog for 15s before offering cancel
A managed Agent Network gateway can be turned off by the platform. The
derived state had no value for that, so a disabled deployment reported
whatever the operator last saw, usually provisioning.
Add `disabled` to the AgentNetworkManagedProxy state enum and say in the
POST and GET descriptions that a disabled deployment answers with it and
that provisioning again does not turn it back on.
Let an embedding binary refuse DELETE /api/reverse-proxies/proxy-tokens/{id}
through an optional proxytoken.RevocationGuard passed to NewAPIHandler.
It is checked after the ownership check, so another account's token
still returns 404 without reaching the guard, and before the token is
revoked. A status error from the guard is written with util.WriteError;
any other error becomes a generic 500.
Nothing installs a guard here, so OSS behavior is unchanged.
* extract shared db conn + data repository
* extract repository interface
* protect against nested transactions
* fix mysql and db conn creation
* fix context management
* remove query warpper
* remove withContext and withLock wrapper
* remove context from function call
* remove in memory mode
* fix nested transaction handling
* use db directly
* remove leftover test
* remove pool close on error during conn creation
* [client] Stop dumping the whole device to clear one peer endpoint
Clearing a peer's endpoint has to remove and re-add the peer, because neither the
netlink API nor the wireguard-go UAPI can clear an endpoint in place. To keep the
peer's allowed IPs across that dance, RemoveEndpointAddress read them back from the
device: a full wgctrl.Device() dump on the kernel path, a full IpcGet plus text parse
on the userspace one. Both cost a round trip proportional to the entire network map,
both run under the interface lock, and both run on every relay and ICE transition.
On a routing peer with ~15700 peers that is megabytes of netlink traffic per
transition, at a measured 713 transitions per minute, with every other configuration
operation queued behind it. RemoveAllowedIP paid the same price for the same reason.
The allowed IPs cannot come from the caller: peer.Conn knows the peer's own overlay
addresses, while the routed prefixes are attached separately by the route manager's
refcounter, so a caller-supplied set would silently drop every route behind the peer.
The configurer is the only writer of its device's peer set, so it can keep an
authoritative mirror of what it configured and answer from memory instead. The mirror
is fed by every operation that changes a peer's allowed IPs and reset by a device
reconfiguration that replaces the peer set. A peer the mirror has not seen, which is
what an out-of-band reconfiguration leaves behind, still falls back to reading the
device and seeds the mirror from it.
Prefixes are unmapped on the way in, so a v4-mapped address compares equal to the
plain v4 prefix for the same network rather than registering as a second entry.
Measured on a userspace device, allocations to clear one endpoint:
peers 64 256 1024 4096
before 1452 - 21617 -
after 91 91 91 91
* [client] Keep update-only allowed IP adds out of the peer mirror
AddAllowedIP configures the device with update_only, which is a silent no-op when
the peer does not exist, so its success says nothing about whether the device took
the prefix. Recording it unconditionally let the mirror hold a peer the device had
dropped, and RemoveEndpointAddress re-adds a peer without update_only: clearing the
endpoint of such a peer recreated it, carrying allowed IPs the device never held.
Allowed IPs are unique per device, so the recreated peer takes those prefixes away
from the peer that legitimately holds them.
This is not a theoretical window. Under lazy connections a routing peer's device
entry is torn down and re-created on the idle transition, and a routed prefix
re-added during that window is lost exactly because of update_only (#6863).
Allowed IP adds now merge only onto a peer the store already knows, which mirrors
the device: the operations that can create a peer record it, the update-only ones
do not. A peer missing from the store still falls back to reading the device.
* [client] Hand a prefix over to its new owner in the peer mirror
An allowed IP belongs to exactly one peer: configuring a prefix on a peer takes it
away from whichever peer held it before, and the configurer leaves that handover to
the device rather than removing the prefix from the previous holder itself, which is
what UpdatePeer's "wg will handle duplicated peer IP" refers to. The mirror recorded
the prefix on the new peer while leaving it listed under the old one, so clearing the
old peer's endpoint rewrote its allowed IPs from that stale list and took the prefix
back from the peer that now owns it. Traffic for the routed prefix then went to the
wrong peer. Reading the device before each write used to rule this out.
The store now tracks the owner of each prefix and performs the same handover, so
rewriting one peer's list cannot reclaim a prefix another peer holds.
Prefixes are also masked on the way in. A device stores them masked, so a caller
passing host bits would otherwise fail to match what a device fallback seeded and
could never remove that prefix by value. Conversion back from the device now keys
the v4-mapped decision on the mask width as well, so a genuine v6 prefix inside the
mapped range stays v6 instead of being dropped as an invalid v4 prefix.
* [client] Keep a mapped v6 prefix below /96 out of the v4 form
normalizePrefix unmapped any v4-mapped address before masking it, keeping the
original prefix length. For a genuine v6 prefix inside the mapped range, such as
::ffff:0:0/64, that pairs a v4 address with a v6 sized mask: netip.PrefixFrom
returns an invalid prefix and Masked turns it into the zero prefix. The store then
held a prefix whose Bits is -1, which cannot reproduce the allowed IP the device
was given, so re-adding the peer after an endpoint removal could fail once the
peer had already been removed.
Masking now comes first, and it also decides the address family: only a prefix at
least 96 bits long keeps the mapped marker through the mask, so anything shorter
inside that range is v6 and stays v6.
* [client] Record a peer created by a preshared key write
Setting a preshared key without updateOnly creates the peer when it is absent, and
Rosenpass applies a peer's first key exactly that way, since applyKeyLocked passes
the peer's initialized flag. The store ignored that operation, so the peer could
exist on the device while the store treated it as unknown.
An update-only allowed IP add on such a peer then succeeded on the device, which
moved the prefix away from its previous holder, while the store skipped the peer
and left the previous holder still claiming it. Clearing that holder's endpoint
rewrote it from the stale claim and took the prefix back, leaving the peer that
owns the route with nothing.
Every device operation that can create a peer now records it, which is the same
rule the update-only operations already follow from the other side.
* [client] Match a peer on the parsed key instead of its base64 form
getPeer scanned the device comparing Key.String to the caller's key. wgtypes.Key
is a 32 byte array, so it compares directly, while String base64 encodes it into a
fresh allocation on every iteration. The scan therefore allocated once per peer on
the device to find a single peer, and on a large network that is tens of thousands
of allocations per lookup.
The key is parsed once up front and the arrays are compared. Behaviour is
unchanged: the callers already parse the same key before reaching here, so the new
parse error is unreachable in practice and only guards the helper on its own.
* [client] Normalize prefixes on their way to the device
Prefixes were normalized when recorded but not when written, so a caller's raw prefix
reached the device while a different form was kept for it. The conversion is also where
a mapped prefix goes wrong: net.IPNet prints a v4-mapped address as v4 but takes the
length from its 16 byte mask, so ::ffff:10.1.2.3/64 is handed to a userspace device as
10.1.2.3/0 — an allowed IP matching every v4 address, on a peer that was meant to carry
one /64.
prefixesToIPNets now normalizes, and the two hand-built conversions in AddAllowedIP go
through it, so there is a single place where a prefix is turned into something a device
is given and it cannot disagree with what is recorded for it.
* [client] Parse the endpoint before configuring the peer
The userspace UpdatePeer parsed the endpoint address after the device had already been
configured, and returned on a parse failure. The device was then left holding a peer
that neither the activity recorder nor the allowed IP store had been told about, so the
peer was invisible to the wake path and the prefix handover for its allowed IPs never
happened, leaving the previous holder still claiming them.
The parse now happens before anything is written, so the only failure left after the
device is touched is one the caller cannot cause.
* [client] Keep the record when a peer removal fails
The two configurers disagreed: the kernel one dropped its record only once the device
had accepted the removal, the userspace one dropped it either way. Removing a peer is a
single device write, so a failure leaves the peer exactly as it was, with the allowed IPs
the record still describes. Dropping it there asserts nothing useful and only sends the
next caller to read the whole device back for an answer it already had.
The userspace one now follows the kernel and returns early on failure.
* [client] Write down what the allowed IP store does not guarantee
Two properties were relied on without being stated. The store's lock covers its map and
not the device write beside it, so consistency between the two rests on callers being
serialized, which WGIface does with its mutex; anyone removing that would have no way to
learn it mattered. And the fallback to the device only covers a peer the store has never
seen, so a peer first recorded from empty while the device already held prefixes keeps
only what was recorded, and the next endpoint removal drops the rest.
* [client] Key the allowed IP store on the parsed peer key
The store keyed on the textual key, so a lookup compared 44 byte strings while the
callers all held the parsed key already and the configurer had to carry both forms.
wgtypes.Key is a 32 byte array and compares directly, which is what getPeer was changed
to do for the same reason.
The store and its helpers now take wgtypes.Key, the callers pass the key they parsed on
entry, and the textual form survives only where something outside speaks it: parseStatus
reports peers that way, so the userspace fallback converts once for its scan.
* [client] Document the configurer methods the store changed
The exported configurer methods now carry what the allowed IP store made true of them:
when the mirror is reset, that a peer update merges its prefixes and takes them from
their previous owner, that an update-only add on an absent peer does nothing, and what
each side does with its record when a device write fails — where the two configurers
differ, since the userspace one reports a prefix it does not have and the kernel one
treats it as a no-op. mergeLocked states the lock its callers must already hold.
Docstrings that only restated the name of a test are left out; the tests explain the
scenario they set up in the body, where the explanation belongs.
Deleting a custom domain released its name while services still pointed
at it, leaving them on a namespace the account no longer held.
Deletion now refuses with 412 when a service in the same account uses the
domain or a subdomain, including disabled ones. Service writes revalidate
authorization inside their transaction and hold a shared lock on the
matching registrations, so a delete racing a create cannot strand either.
The dependency lookup is account-scoped: registrations are unique by name,
so another account can hold team.example.com under example.com and its
services are authorized by its own registration.
* add readme section and support file
* [doc] Document anonymize levels in README and SUPPORT
Reviewer feedback: mention both --anonymize-level values, not just -A.
Wording follows the flag help in client/cmd/root.go.
* Apply suggestion from @cubic-dev-ai[bot]
* Update SUPPORT.md
* adjust parameter descriptions
* [doc] Fix duplicated sentence in SUPPORT.md
The cubic suggestion replaced only part of the -U sentence, leaving the
original line orphaned above it and dropping the blank line before the
docs links.
* [doc] Tighten anonymization claims in README and SUPPORT
Strict mode keeps labels under netbird.io, so naming "the NetBird domains"
overstated what it masks. Name the three peer domains instead.
Anonymization is not full redaction: internal ranges survive at the default
level and interface details are never masked, so say that rather than
implying the bundle is safe to post unread.
Add a UBI-based reverse-proxy image for the internal Red Hat certification requirement, without changing the existing image or deployment defaults. Includes non-root execution, licensing and image metadata, plus an AMD64 GoReleaser entry with separate UBI tags.
Preflight and TLS/overlay checks passed. An intermittent shutdown exit error remains deferred; this stays draft pending maintainer testing.
Inline the job in release.yml, gated on a stable vX.Y.Z tag on the upstream
repo so it stays out of main, release branches and pull requests. Its env and
contents:read permission move to the job, keeping them off the rest of the
workflow.
upload-server/Dockerfile only packaged the goreleaser-built binary, so
the image could not be built from a checkout. It is now a multi-stage
build on Chainguard static, running as the nonroot user (uid 65532),
with a VARIANT=debug build arg that swaps in busybox for a shell.
The goreleaser packaging file moves unchanged to Dockerfile.release and
.goreleaser.yaml points at it, so the published netbirdio/upload image
stays as it was: distroless and root.
The bases are pinned by digest, since Chainguard publishes only :latest
for free. A Dependabot docker entry for /upload-server moves them
weekly, leaving the release base alone and holding golang to patch
updates.
Connections to the daemon were left on gRPC's own defaults, which cap a received
message at 4 MB. A detailed status carries an entry per peer, so on a large
deployment the response outgrows that cap and the command fails outright:
netbird status -d
Error: status failed: grpc: received message larger than max (4287609 vs. 4194304)
The limit is raised where the daemon dial options are built, so every caller
inherits it: the CLI, the desktop UI, the JSON gateway, and the SSH client and
proxy. It is overridable through NB_DAEMON_GRPC_MAX_MSG_SIZE for a deployment
that outgrows the new default too, mirroring what the management client already
does with NB_MANAGEMENT_GRPC_MAX_MSG_SIZE, and reusing its 16 MB default.
Only the receive direction needs raising. Requests to the daemon are small, and
gRPC does not cap the send side by default, so the daemon could already send a
response the caller then refused to read.
Nothing in CI compiles the files behind //go:build android or //go:build ios.
The android bridge builds 7 of its 23 files on linux and skips client.go; the
iOS SDK is not built at all. The linter matrix picks a GOOS by picking a runner
OS, so it loads the same file set as the host build and never sees them either.
A type error in client/android/client.go therefore passes every check on its
PR, merges, and is discovered by netbirdio/android-client after sync-tag.yml
fires trigger_android_bump on the release tag.
The new Mobile workflow cross-compiles ./client/android/... for the GOARCH
values gomobile ships and ./client/ios/..., and vets the android bridge. The
new Android and iOS lint jobs run golangci-lint with GOOS/GOARCH in the job
env. No NDK, Xcode or gomobile is needed: these are library packages, so the
compiler type-checks them without a link step, and the dependency graph drags
in the android/ios-tagged files across client/iface, client/internal/dns and
client/internal/routemanager with them.
Linting those files for the first time surfaces one gosec G101 on the SSH
password-required marker. It is a sentinel string the Java side matches on,
not a credential, so it is suppressed at the declaration.
GitHub Actions runners no longer ship Node 20 for JavaScript actions, and the
ACTIONS_ALLOW_USE_UNSECURE_NODE_VERSION opt-out is gone, so any action whose own
action.yml declares `runs.using: node20` now fails to start. Eleven call sites
across three workflows were still on such actions: actions/setup-node (five),
pnpm/action-setup (four), actions/cache and actions/setup-go (one each). Every
target was verified by reading `runs.using` out of the pinned ref's action.yml
rather than inferred from its version number.
Pin style is preserved per call site: SHA-pinned refs stay SHA-pinned with a
corrected `# vX.Y.Z` comment, tag-pinned refs stay tag-pinned. cache and setup-go
go to v6 rather than the newest release so the stragglers join the versions the
rest of this repo already runs.
Breaking changes were checked and none apply. The setup-node v5/v6 automatic
dependency caching never triggers: it resolves package.json from the repo root,
which does not exist here, and the frontend caches the pnpm store itself. v7
drops the dummy NODE_AUTH_TOKEN export and adds cache outputs, neither of which
any workflow reads. setup-go v6 reworks toolchain selection, but the one
straggler passes the same go-version-file and `cache: false` as the 22 setup-go
v6.5.0 pins already in CI. The v5/v6 runner floor is met because every job runs
on GitHub-hosted runners.
pnpm/action-setup v4 added a hard error when the `version:` input disagrees with
package.json's `packageManager`. It stays dormant here only because the action
looks for package.json at the repo root and swallows the resulting ENOENT; all
four sites pass `version: 11` while client/ui/frontend/package.json says
pnpm@11.4.0. Left alone to keep this change to the runtime bump, but adding a
root package.json or setting `package_json_file` would make all four fail.
git-town/action is knowingly left on node20. Every release through the latest
v1.3.3, and main HEAD, still declares `runs.using: node20`, so there is nothing
to bump to. That job will break when the runtime is retired and needs its own
decision: drop it, fork the action onto node24, or file upstream.
node-version stays at 22. That is the Node toolchain used to build the frontend,
not an action runtime, so this retirement does not touch it, and the build cannot
be validated on this host because the binding generator needs Linux-only GTK4 and
WebKitGTK dev packages. It belongs in a separately verified change.
Goreleaser's RPM build is split into one nfpm entry per architecture, each pinned to a single-arch build, with the version substituted from the release job. The deb package, archives, and container images are unaffected.
Also fixes rpmlint: incoherent-version-in-changelog, found while investigating: nfpm writes the changelog title straight from semver and never appends the release, so entries read 0.79.0 against a 0.79.0-1 package.
* [client] Export the only-owner-writable path check from elevate
Pure refactor, no behavior change: the existing checkOnlyOwnerWritable gets a
thin exported wrapper so callers outside the elevation path can reuse it. No
call site changes here.
* [client] Validate the saved service parameters before applying them
The install reads <stateDir>/service.json and applies it to the service it then
registers: its arguments, its config path and its environment. The restricted
ACL that saveServiceParams puts on the state directory is applied when the file
is written, which is not necessarily before the file is first read, so the
install now checks the file rather than assuming it.
A file whose ownership or permissions are not the ones saveServiceParams
produces is treated as absent, and the install proceeds with its defaults. The
check covers the directories above the file as well, so what is checked is what
is read.
* [client] Restrict which environment variables the service is registered with
--service-env, and the service.json it persists to, accepted any name. A small
set of them decides how a process resolves the executables and libraries it
loads, and the daemon needs none of those: it now refuses them when they are
passed explicitly, and drops them with a warning when they come back from a
service.json written by an older version, so an upgrade does not fail over a
variable nobody needs.
* [client] Resolve netsh by absolute path
The lookup consulted PATH first and fell back to System32, in both the copy the
userspace firewall uses and the one that tears the interface down. It now asks
Windows for the system directory, so the resolution no longer depends on the
environment the service happens to be started with.
* [client] Move the System32 lookup into a package both callers share
Pure refactor, no behavior change: client/iface and client/firewall/uspfilter
carried a copy each of the same function, and neither imports the other, so the
body moves to client/internal/wincmd — alongside winregistry, which is where
the client's other Windows-only helper already lives. Both call sites now read
wincmd.System32("netsh").
* [client] Cover the System32 lookup with a test
Asserts what the previous commits changed: the lookup is absolute, and neither
PATH nor %SystemRoot% moves it.
* [client] Refuse the loader environment families by prefix
Review follow-up on the previous commit:
- LD_* and DYLD_* are now refused whole rather than name by name. Their members
differ per platform and libc and grow with new OS releases, so a list of them
is out of date as soon as it is written — DYLD_FALLBACK_LIBRARY_PATH and
DYLD_FALLBACK_FRAMEWORK_PATH were already missing from it.
- The names are folded to upper case only on Windows, where a variable is the
same one however it is spelled. Elsewhere the environment is case-sensitive,
so Path and PATH are two variables and only the exact spelling is the one that
is read; the fold refused the wrong one.
- TEMP and TMP stay in the denylist, but the rationale and the message now say
what they actually decide: where the service writes, not what it loads.