Commit Graph
3401 Commits
Author SHA1 Message Date
riccardom 278f2f3807 [management] Derive certificate challenge nonces from the stored encryption key
The nonce secret came from the server's WireGuard key, which is generated afresh
in every process and never persisted. A nonce carries no state, so the only
thing that lets one instance verify what another issued is deriving the same
secret — and that premise, written in the comment above the challenger, was not
met: every instance had its own key.

A peer reconnecting after a restart therefore presented a nonce minted under the
previous secret, verification failed with a mismatch, its whole proof set was
rejected and the certificates stored for it were dropped until it signed again.
Reproduced three times on the lab, each one logging "nonce was not issued to
this peer", which only a changed secret produces. On a single instance it costs
seconds of lost policy access per restart; across instances it is not transient
at all, because every reconnect that lands elsewhere is rejected the same way.

Derive from the data store encryption key instead: it is generated once, written
back to the configuration and read by every instance, so it survives restarts
and is shared. Where none is configured the secret falls back to the WireGuard
key with a warning — degraded but still unpredictable, which is the property
that matters most: a peer able to guess it could mint the nonces of future
windows, sign them while its key is present and keep passing after it is gone.

The challenger is now built once and passed to the two places that need it,
rather than re-derived per message.
2026-10-02 16:21:04 +02:00
riccardom 8d226074f3 [client] Drop go.step.sm/crypto and the repo-wide upgrades it imposed
The TSS2 parser was the only thing in the repository that used this module, and
it brought 302 modules into the graph to do it — 35 of them linters, along with
Google Cloud KMS and IAM, the AWS SDK and a terminal styling library. Those are
the module's own development dependencies, which minimal version selection turns
into floors in ours, and they are the whole reason gRPC, protobuf, the AWS SDK,
OpenTelemetry, logrus and five x/ packages had moved. Management, signal, relay
and proxy inherited every one of them for a feature none of them runs.

Removing the import is not enough, because tidy never downgrades: the raised
floors stay written in go.mod. Each one is pinned back to the version main had,
then tidy is left to raise again whatever something still genuinely needs. It
raised nothing: all 43 are back where they were, and go-tpm was already in the
graph at the same version, so the certificate feature now costs no new module at
all.

The differential tests go with it. They existed to check the swap against the
library while both were present, and there is nothing left to compare against.
2026-10-02 16:21:04 +02:00
riccardom a7f183c87c [client] Read TSS2 key files on go-tpm, checked against the library it replaces
The TSS2 parser was the only reason this repository depended on a crypto suite
whose own build tooling it inherits. The replacement sits on go-tpm, which was
already a direct dependency and is in fact what that suite calls underneath, so
this removes a wrapper rather than porting onto a different library: the load,
the derived storage root key and the signing commands are the same calls.

Swapping a parser on the one path a customer actually runs is not something to
assert, so the two are held side by side for this commit. One test feeds the
replacement bytes the old library wrote and requires the same key type, empty
auth flag, parent handle, blobs and decoded public key; the other feeds both the
fixtures the tests are built on, so those are the shape the format calls for and
not merely the shape the new parser reads. The scaffolding goes away with the
dependency in the commit that follows.

The encoder behind the fixtures is written out separately from the parser under
test, so an encoder bug and a decoder bug cannot cancel each other out.
2026-10-02 16:21:04 +02:00
riccardom dc4d0e0274 [management] Renew challenges only for the peers that answer one
The refresh pushed an update to every connected peer of the account, while only
the peers a certificate check applies to carry a nonce. On an account where a
handful of peers sit behind the check and the rest do not, everyone was woken
several times a day to be handed a map that changed nothing for them.

Push to the sources of the enabled policies whose posture checks include a
certificate check, which is exactly the set that is sent a challenge.

Resolving the set the other way round than the gRPC layer does is the risk here:
a peer the refresh forgets stops being renewed and falls out of its policies
silently, which is the failure this whole mechanism exists to prevent. So the
selection is held against processPeerPostureChecks, the per-peer rule that
decides who receives a challenge in the first place, by a test that asks both
the same question and requires the same answer.
2026-10-02 16:21:04 +02:00
riccardom dee68facfd [management] Pin the property that makes per-peer nonce state unnecessary
A nonce carries the window it was minted in, not the instant, and is accepted
for that window and the one before it. So a peer re-stamped at least once per
window can never be left holding one outside the accepted pair, whenever it was
last served and however much life its own nonce had left. That is the whole
reason management tracks nothing per peer, and it was resting on an argument
rather than a test.

The phases are part of the property, not decoration: accounts are deliberately
given a refresh phase of their own, so the guarantee has to hold off the window
boundary too. The negative case shows why that matters — a cadence of exactly
two windows lands inside the grace window when it is aligned to the boundary and
leaves a gap when it is not.
2026-10-02 16:21:04 +02:00
riccardom 0e38fcf409 [management] Make the certificate challenge window one knob to turn
Renewal was timed against the window in two different ways: the period derived
from it, the sweep interval did not. Shortening the window to watch a renewal in
an end-to-end run would have left the refresher still looking for due accounts
every quarter of an hour, so nothing would have been renewed in time and the
test would have reported the feature broken.

Derive the sweep from the period, within bounds that keep a very short window
from spinning and a normal one from checking less often than is useful, and
allow the window itself to be set through the environment so a run can take
seconds instead of half a day. A value that cannot be parsed or falls outside
the bounds keeps the default, because a window nobody intended is a security
property nobody chose, and an override is logged at warning level since it sets
how long a device keeps passing the check after its key is gone.

Every instance has to be given the same value: the window is part of the nonce,
so instances that disagree reject each other's.
2026-10-02 16:21:04 +02:00
riccardom 7ad4a0df37 [management] Renew certificate challenge nonces on quiet accounts
A certificate challenge nonce is accepted for its own window and the one before
it, and it only reaches a peer attached to a network map. An account where
nothing changes sends no map, so after a day the peer re-sends the nonce it
still holds, verification rejects its whole proof set, and the certificates
stored for it are dropped. It fails the certificate check and loses every policy
gated on it until some unrelated change happens to push a map. The outage
repairs itself in seconds, which is what makes it expensive: it is intermittent,
it only hits stable networks, and it is not reproducible on demand.

Push the account's peers an update often enough that the nonce they hold is
never close to expiring. Only accounts whose posture checks actually ask for a
certificate are tracked, so a deployment without the feature does no extra work.

The refresh runs from one goroutine over a map of accounts rather than a timer
per account: the period is hours, so one pass every few minutes costs nothing
next to it, and there is no timer to re-arm when an account that falls due
sooner appears. Each account's first run is offset by a hash of its ID, because
the challenge window is global and an instance restart would otherwise arm every
account in the same moment.

The push carries no administrative change, so it is counted as a refresh rather
than an update and stays out of the figures that track what was edited.
2026-10-02 16:21:04 +02:00
riccardom 3e360ebe0f [client] Drop the unreachable helper timeout from certificate proof collection
Both platform collectors wrapped the helper launch in a 30 second deadline, but
the context they wrapped was already capped at 10 seconds by the collector that
calls them, and CollectProofs has no other caller. The inner deadline could
never be reached, so it described a budget the code does not have.

Leave the collector as the single owner of the deadline rather than picking a
smaller inner value: on macOS one budget has to cover the System keychain read
and the helper launch that follows it, and how to divide it is a question about
the budget as a whole, not about the helper alone.
2026-10-02 16:21:04 +02:00
riccardom ec48a92065 [client] Cover the cert proof helper path the deadline actually cuts short
The existing case has the helper exit on its own, which reports ErrWaitDelay.
The daemon meets the other shape: the process it launches is a launcher, the
work happens in a grandchild, and the deadline kills the launcher while the
grandchild holds the inherited pipe. Wait then reports the kill and the wait
delay error is dropped, so the two paths through Cmd.Wait differ and only one
of them was exercised.
2026-10-02 16:21:04 +02:00
riccardom 8224e0f7bf [client] Bound the wait on a certificate proof helper that outlives its launcher
The collector that answers certificate posture challenges allows one collection
at a time, and the latch guarding that is released by the goroutine running the
collection, not by the caller that gave up waiting for it. The collection is
therefore assumed to end.

On macOS and Windows it reads the signed-in user's certificates through a helper
process, and there the assumption does not hold. Wait blocks on the output pipe
rather than on the process, and killing the process we launched does not close
the pipe its own children inherited: on macOS we launch launchctl, which launches
sudo, which launches the helper, so the deadline reaps launchctl and leaves the
other two holding the pipe open. A helper stuck on a keychain prompt is enough.

Wait then never returns, the latch stays set, and the peer sends no proof again
for the life of the daemon. It fails every certificate check and loses the
policies that carry one, logging a single line per sync and nothing else.

Set a wait delay so Wait closes the pipes itself once the process is gone. The
run fails rather than reporting proofs, which is what we want: the output was cut
short, so there is nothing trustworthy to report.
2026-10-02 16:21:04 +02:00
Viktor Liu 0ed3eb6139 Test the PKCS#11 build against SoftHSM in CI and warn once where the build has no driver 2026-10-02 12:44:17 +02:00
Viktor Liu 07cf08230c Collect certificate proofs again when the owner's session changes and report lost proofs 2026-10-02 12:37:51 +02:00
Viktor Liu a91fe94ada Read user certificates only from the session of the active profile's owner 2026-10-02 12:32:27 +02:00
Viktor Liu e0b6a38aa2 Require a token label whenever a PKCS#11 PIN is set 2026-10-02 12:27:00 +02:00
Viktor Liu 3cfff34d2c Find a chain to each challenge's CAs through every intermediate the store holds 2026-10-01 10:06:50 +02:00
Viktor Liu 21be4420d1 Keep the macOS keychain code out of iOS and the PKCS#11 driver out of Android 2026-10-01 09:56:22 +02:00
Viktor Liu 06bc8bfcad Never pass NULL to CFRelease and skip unreadable keychain identities 2026-10-01 09:50:31 +02:00
Viktor Liu b364af9206 Bound PKCS#11 driver sizes, pin template values, and log out only a login the session owns 2026-10-01 09:48:35 +02:00
Viktor Liu 59aeeb1c94 Skip certificate files whose key belongs to another certificate 2026-10-01 09:07:56 +02:00
Viktor Liu 5af909b111 Sign only nonces and peer keys of the size management issues 2026-10-01 09:06:14 +02:00
Viktor Liu 0129b90ae3 Log certificate posture details at debug level 2026-10-01 08:45:01 +02:00
Viktor Liu a230d39aa0 Stop retrying a PKCS#11 PIN the token rejected 2026-10-01 08:33:20 +02:00
Viktor Liu 83e7bf8ab4 Bound certificate proof collection so a stuck token or keychain cannot hold the sync loop 2026-10-01 08:21:13 +02:00
Viktor Liu ab2f8972be Read the PKCS#11 token PIN from NB_TPM_PIN instead of the profile config 2026-10-01 08:16:40 +02:00
Viktor Liu 0641f0d7bb Isolate the cert proof helper from the service environment and cap its output 2026-10-01 08:13:59 +02:00
pascal 574c29973a add unsupported flag for mobile devices 2026-09-29 23:11:41 +02:00
pascal 7fbde60b7c split cert and key location and allow key lookup on tpm 2026-09-22 17:04:50 +02:00
pascal d04aef6f14 add tpm pin to netbird config 2026-09-21 14:59:39 +02:00
pascal d3c5b6718a Merge branch 'main' into poc/certificate-posture 2026-09-21 11:22:36 +02:00
Eduard Gert 314d88252d [management] Name the account owner in the pending approval error (#7533)
* Name the account owner in the pending approval error

A user refused because their account is pending approval had no way to
learn who could approve them. The refusal now carries the account
owner's address, masked, so a caller can name someone to contact without
being handed the address itself.

Resolving the owner is best effort: a lookup failure, or an account
predating the stored email, falls back to the refusal as it was.

* Name only the caller's own owner in the pending approval error

The refusal is raised before ValidateAccountAccess has established that
the caller belongs to the account the request asked about, and the user
is loaded by ID alone. Resolving the owner of the requested account
therefore disclosed that owner's address to a pending user with no claim
to it, reachable through any handler that takes an account ID from the
caller — DELETE /api/accounts/{accountId} passes one straight through.

The owner who can approve a pending user is the owner of their own
account, so resolve that one. The requested account is never read.

* Mask short local parts whole in MaskedEmail

Keeping the first two characters and the last hides nothing until the
local part is four long: at three or fewer they are the whole of it, so
"abc@example.com" masked to "ab****c@example.com" and a pending user
could recover the owner's address in full from what is meant to conceal
it. Short local parts are now replaced entirely.

* Name the owner from GetCurrentUserInfo instead of the permission gate

The gate could only read the stored user row, which carries no address
when an external IdP owns the identities — the usual case — so it named
no one in practice. It also had no way to reach the IdP without being
handed the account manager, which meant restoring bootstrap wiring that
a refactor had dropped.

GetCurrentUserInfo already holds that account manager, so it answers for
a pending user itself and reuses GetOwnerInfo, the same lookup /msp uses
to resolve an owner's address. The gate returns to exactly what it was,
and with it goes the risk of naming the owner of an account the caller
only asked about.

MaskedEmail becomes MaskEmail: with a UserInfo in hand there is no stored
row to hang it off.

* [management] Cover the pending approval refusal in GetCurrentUserInfo

The branch that names the owner had no coverage at the manager level, so
neither the named refusal nor the fallback for an owner without a resolvable
address was pinned down.

* [management] Cover the failed owner lookup in the pending approval refusal

The generic fallback has two ways in: no address on the resolved owner, and no
owner to resolve at all. Only the first was pinned down.

* [management] Pin the owner lookup to the caller's own account

A mismatched account claim must not steer which owner the refusal names, and
a blocked user is still answered before the claim is validated. Both are load
bearing and neither was covered.
2026-09-21 10:34:04 +02:00
Maycon Santos 3073d18039 [proxy] Close the client connection on private service denials (#7590)
A client that hits a private service before its peer joins the overlay gets a 403 from the tunnel-peer check. After it connects to NetBird, the browser reuses the warm socket to the public listener, so the request never traverses the tunnel and keeps failing until the 120s idle timeout closes it.

Private service denials now set Connection: close and Cache-Control: no-store before the 403, both at the tunnel-peer check and at IP restriction denials on a private domain. Go's HTTP/1.1 server closes after the response; its HTTP/2 server turns the exact lowercase close token into a GOAWAY, which retires the stale connection for h2 clients. Public services and allowed private traffic keep their keep-alive behaviour.

Tests cover HTTP/1.1 and HTTP/2 denials over a real listener (retry lands on a new connection), public denials and allowed private requests (connection reused), and both IP restriction paths.
2026-09-20 20:24:01 +02:00
Bethuel Mmbaga 7d8f4fa31c [management] Handle empty trusted peer (#7589) 2026-09-18 18:21:34 +03:00
pascal af7b475552 go mod tidy 2026-09-17 17:52:43 +02:00
pascal a2e66a7d64 Merge branch 'main' into poc/certificate-posture
# Conflicts:
#	go.mod
#	go.sum
2026-09-17 17:52:19 +02:00
pascal 47318da65f update goreleaser 2026-09-17 17:45:32 +02:00
Bethuel Mmbaga f8c3e565f3 [management] Read X-Real-IP when extracting the peer connection IP (#7561) 2026-09-17 12:45:03 +03:00
Misha Bragin 85a3913331 [client] Fix - Add RPM metadata required for Red Hat software certification (#7562)
Declare the runtime dependencies, generate the changelog from git tags with
chglog at release time, and ship LICENSE, README.md and an example
/etc/sysconfig/netbird as %license, %doc and %config(noreplace). The unit
generated by "netbird service install" already reads that path via
EnvironmentFile, so post_install.sh is unchanged.
2026-09-17 09:24:55 +02:00
pascal 92d76acdd3 split goreleaser to support pkcs11 and exclude on docker 2026-09-17 00:17:37 +02:00
pascal 25078c4af0 Merge branch 'main' into poc/certificate-posture 2026-09-16 23:31:07 +02:00
pascal 24f5f2a7ec start TPM support 2026-09-16 23:00:44 +02:00
Maycon Santos eab510178a [misc] Load AGENTS.md every session and refuse attribution trailers (#7544)
AGENTS.md forbids attribution trailers, but a rule an agent has to go and read loses to the instruction it is handed every turn. CLAUDE.md now imports AGENTS.md so it is always in context; a commit-msg hook (via make setup-hooks) refuses the trailers at commit time; a CodeRabbit pre-merge check flags a PR whose description or commits carry them. The check reports rather than blocks, since the repository keeps CodeRabbit's request-changes workflow off; turning that on is a separate, repository-wide decision.
2026-09-15 22:48:49 +02:00
Riccardo Manfrin 08699a9e29 [client] Enforce HTTPS on install script downloads (#7545)
Every curl invocation in the install script that follows redirects now passes
`--proto` and `--proto-redir` set to https only, so neither the initial request
nor any hop in the redirect chain can drop to plaintext. This matters most for
the macOS .pkg download, whose URL is itself the result of a redirect
resolution, and for the release tarballs that get moved into the install dir as
root.

The protocol set lives in a single `PROTO_HTTPS` variable rather than being
repeated at each call site, and every expansion is quoted — the variable holds
one option value, not a list of flags.

The two call sites without `-L` (the release metadata lookups) are left alone:
they do not follow redirects and their URLs are https literals.

Verified against every URL the script fetches on curl 7.29.0 (CentOS 7),
7.68.0, 7.76.1, 7.88.1 and 8.14.1; both options have existed since curl 7.20.0.
Plaintext http:// is refused on all of them.
2026-09-15 14:20:15 +02:00
Zoltan Papp e70ec07320 Read the session deadline under the status read lock (#7550)
GetSessionExpiresAt took the exclusive lock for a plain field read, so
every caller queued behind writers and behind each other. The Android
SessionMonitor polls it from the main thread, and in the captured ANR
that is exactly where the main thread was blocked while hundreds of
peer-list callbacks held or waited on the same mutex.

d.mux is already an RWMutex and the other getters use RLock; this brings
the deadline read in line with them.
2026-09-15 12:05:57 +02:00
Riccardo Manfrin 58b5263c1a [client] Stage install script downloads in a private temp directory (#7534)
The install script downloaded both the macOS .pkg and the release tarballs
into /tmp under fixed, predictable names, then passed those same paths to the
privileged install steps (`installer -pkg`, `mv` into the install dir).

/tmp is shared, so those fixed names can collide with entries created there
beforehand, and the privileged steps consume whatever the path resolves to.

Stage every download in a directory from `mktemp -d` instead: unpredictable
name, mode 0700, owned by the caller, created atomically. Extraction now
targets that directory (`tar -C`, `unzip -d`) rather than relying on `cd /tmp`,
and an EXIT trap removes it, so a failed run no longer leaves the archive and
the unpacked LICENSE/README behind in /tmp either.
2026-09-15 10:30:33 +02:00
Maycon Santos f29249e7ef [management] Point the agent-config e2e providers at the mock upstream (#7542)
TestAgentConfigAllowlistOfDeclaredModels still pointed its providers at api.openai.com and bedrock-runtime with a dummy key, so every save has been refused with "the provider rejected the credential" and the Agent Network E2E has been red on main since — both subtests, every scheduled run.

Every other suite already uses the mock vLLM upstream (harness.StartVLLM), which answers both the OpenAI (/v1/models) and the Bedrock (/inference-profiles) listing; this test does the same. What it checks — the allowlist advertising the provider's declared ids on GET /api/agent-network/agent-config — never depended on the vendor.
2026-09-15 09:40:52 +02:00
Maycon Santos 7ce6a63dcb [management] Validate the proxy cluster an agent network bootstraps onto (#7402)
The agent network gateway service is private: agents reach it over the WireGuard tunnel, authorised by peer identity, with the cluster as its only target. Only a reverse proxy cluster with private capabilities can serve that, reported per cluster as the `private` capability — the same `supports_private` flag the dashboard gates NetBird-only services on.

A bootstrap could pin an account to a cluster without private capabilities, leaving an immutable dead gateway. Both bootstrap paths now validate the picked cluster: one the account can see must have a connected proxy reporting the capability. Shared and account-owned clusters qualify alike. Known-ness comes from proxy rows, not heartbeat freshness, so a cluster without the capability stays refused while merely offline. A hostname no proxy has declared stays pinnable (address-first). Identity is compared case-insensitively over the account's cluster list.
2026-09-15 08:42:21 +02:00
Maycon Santos ea294e1d46 [management] Refuse to pin an agent network gateway onto another account's host (#7519)
An agent network bootstrap stores its cluster as proxy_address, which selects the proxy that serves the endpoint. An account-scoped proxy only receives its own account's mappings, so a pin onto a host another account's proxy declares can never be served, and the endpoint is immutable — a dead gateway until the account deletes its settings. Nothing refused that pin; the domain unique index only arbitrates between endpoints.

Both bootstrap paths now refuse, before the insert, a host that another account's proxy declares, a host another account has labeled pins beneath (self-addressed path), or a hostname that is another account's endpoint (labeled path).

Shared clusters are unaffected: shared proxies are never foreign, and labeled pins under one cluster are never asked about, so any number of accounts still pin beneath eu.proxy.netbird.io. Registration is deliberately unchanged — refusing a proxy for another account's pin would let a pin lock a tenant out after the reaper drops its rows.
2026-09-14 22:20:25 +02:00
Maycon Santos ea216f8e73 [management] Speed up test store setup and summarize the unit test run (#7518)
Management / Unit (amd64, mysql) hit the 20 minute go test budget on #7516. The package was not hung: each of the 133 test store creations in management/server paid about 1.6s on MySQL for CREATE DATABASE, the pre-migrations, a 40-table AutoMigrate and the post-migrations, which puts the package at 10 minutes on a healthy runner and over the budget on a slow one.

The migration now runs once per test binary into a template database and each test database is cloned from it, with CREATE DATABASE ... TEMPLATE on Postgres and a replay of SHOW CREATE TABLE on MySQL. The MySQL test container also drops the binary log, doublewrite buffer and per-commit redo fsync. Two goroutine leaks in the test helpers are fixed.

tools/gotestsummary turns the go test -json stream into a readable log, and the Management unit and integration jobs now pipe through it, so a timeout names the tests still running. On MySQL, management/server went from 10m16s to 6m36s.
2026-09-14 19:42:02 +02:00
Viktor Liu a54d96cd72 [proxy] Make the upstream HTTP version configurable (#7410) v0.79.0-rc.1 2026-09-14 17:28:01 +02:00
pascal d89e9c86eb Merge remote-tracking branch 'origin/main' into poc/certificate-posture v0.80.0-canary.pr-7535.1 2026-09-14 15:49:55 +02:00