The receiver reported progress once per file, after the whole body had
been staged: spool.Write drains its reader before returning, so a large
file sat at nothing until it jumped to done. It now stages through a
reader that reports as the bytes land, the same way the sender already
did, and both sides thin their reports to one per percent and one per
200ms so a fast transfer cannot flood the history or the UI.
Progress also becomes an event of its own. It was left out because a
report per byte would have been unusable; thinned, it costs a handful of
events a second and spares every UI a poll. The daemon's event bridge
ignores the new kind, so nothing is published where a notification would
be noise.
Withdrawing consent mid-transfer did nothing. Cancel routed through
Decide, which only moves an offer out of Pending, so an accepted offer
kept its decision, kept its spool, and kept being uploaded into: the
receiver's list went quiet while the sender ran to completion. Revoke
takes an offer back whatever it has already answered, the staging reader
gives up as soon as consent is gone, and the sender stops rather than
retrying a refusal three times over. A transfer withdrawn this way reads
as declined on both ends, which is what it is, rather than as a failure.
Android re-establishes the VpnService interface on every route change, which
replaces tun0 with a fresh device. The file drop listeners are bound to the
overlay address of the interface being swapped out: the IPv4 one dies with
accept4: invalid argument and never comes back, so a peer dialing the overlay
IPv4 address gets an RST. Restart the receiver once the new device is in place.
Files move directly between peers over the overlay, with no server in the
path. The receiver listens on the WireGuard address only, so the port is
unreachable from outside the tunnel, and every offer is matched to a known
peer before anything is read.
Consent is the default: an offer carries metadata alone, and no payload
moves until the receiver accepts. Policy is per profile and device-local —
off, ask, or auto-accept, with per-sender exceptions on top.
Policy and history live in the profile's preferences, so removing a profile
takes its file drop state with it. Transfers interrupted by a restart are
settled on load; nothing survives to finish them, and left alone they would
sit in the log as permanently pending.
The Android bindings pull payload bytes through a chunk-returning stream:
gomobile copies a []byte argument into a fresh Java array and never copies
it back, so a fill-my-buffer method would hand back the right length with
no data.
Introduce a namespaced preference store owned by the profile manager,
persisted next to the profile config as <id>.prefs.json and deleted with
the profile. Sections are opaque JSON, so the profile manager stays free
of any feature schema, and writes reuse the state file's atomic path.
Migrate the Android SSH known-hosts store and session list onto it. Both
previously lived outside the profile lifecycle: known hosts in a
per-profile file under filesDir, the session list in Java
SharedPreferences, each needing its own sweep against the live profile
list to avoid outliving the profile they belonged to. A profile ID that
got reused would have inherited the trusted keys of a deleted profile.
Both now share the "ssh" and "ssh-sessions" namespaces of the profile's
preferences, so deleting a profile takes them along and the Java-side
pruning is gone.
The known-hosts entries keep the OpenSSH line format, only the container
changed, and host key verification keeps rejecting a changed key
outright. SetKnownHostsPath becomes SetKnownHostsStore, taking the
config dir and profile ID instead of a file path.
Existing known-hosts files are not migrated: hosts trusted before this
change prompt for confirmation once more, which errs towards safety.
GetOAuthFlow was the only flow factory without a hint parameter, which
forced its callers to apply the hint afterwards through a local setter
interface and a type assertion. Give it the same constructor-style hint
as NewOAuthFlow and set the hint on the concrete flows before they are
handed out as the interface, so a flow is always complete when built
and the caller-side ordering constraint disappears.
An empty hint is a valid value meaning the IdP chooses the account, so
the flows set it unconditionally.
Extract the shared RequestAuthInfo -> Open -> WaitToken sequence from the
login flow and the SSH JWT flow into runOAuthFlow. Open is now called
synchronously by both flows, matching iOS; openers must post their UI
work instead of blocking, which the app-side openers already do.
The SSH flow read its login hint via profilemanager.GetLoginHint, which
resolves desktop-layout files that the Android app never writes, so the
hint was always empty and the device-code flow could prompt for account
selection. Both flows now read the hint from the profile account file
via the config path, taken from authSnapshot so a concurrent profile
switch cannot pair one profile's config with another's hint.
The shared table used by the Android and wasm terminals was a strict
subset of the CLI one, leaving Ctrl+U, Ctrl+D, Ctrl+Z and friends
without explicit mappings. Export the full table from the ssh package
and derive both CLI variants from it; Windows adds its console-specific
modes to a copy so the shared map is never mutated.
Extract the dial-then-handshake sequence into nbssh.Handshake, which
applies the context deadline to the socket for the duration of the
handshake. Previously only the Android client did this; the CLI, wasm
and SSH proxy paths could block forever on a peer that accepts the TCP
connection and then goes silent, since ClientConfig.Timeout is not used
by NewClientConn.
Remove Engine.VerifySSHHostKey and keep GetPeerSSHKey as the only SSH
key API on the Engine. Verification now lives in the ssh package as a
PeerKeyLookup func type implementing HostKeyVerifier, shared by the
android and embed clients.
Align Android logout semantics with the desktop UI and CLI: logging out no
longer deletes the stored account email, so the next login passes it as the
OIDC login_hint and the IdP preselects the account. Removing the profile is
now the operation that deletes the email; previously RemoveProfile left the
account file behind, which the fixed-name default profile would have
inherited on recreation.
Extract the identical PTY session setup shared by the wasm and Android
terminal clients into ssh.StartPTYSession, and move the stored-key host
verification onto the engine so the embed client delegates and the
Android client passes the engine directly as HostKeyVerifier.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
English (en) is the source of truth for UI translation keys; the other
nine locales rely on runtime English fallback for any missing key, so a
gap never surfaces in CI. Add a dependency-free Node check that fails
when any locale declared in _index.json does not carry the exact same
key set as en (missing or orphaned keys), wired into a dedicated
UI Translations workflow that runs on locale changes.
Also close the one existing gap the check found: ja was missing
daemon.outdated.download ("Download Latest").
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Three pieces, each verified on Android across WiFi/cellular switches:
- A sweep-aware backoff wrapper: the first retry after a disconnect
that follows a recent network-change mark comes after 200ms instead
of the randomized [0..1.6s] interval. Any other failure keeps the
unchanged spread, so the clients of a restarted server still scatter
their reconnects.
- The retry sleep wakes on OS network availability transitions
(nbgrpc.Retry): a disconnect that precedes the offline flag by a few
milliseconds no longer sleeps blindly through the recovery - the
loop parks on the netstate gate and resumes the moment the network
returns.
- The connection state is re-checked after WaitForStateChange: a dial
settling in Ready proceeds immediately instead of burning another
backoff round on an already-usable channel.
Measured after a network switch: management and signal recover in
270-470ms deterministically, down from a 312-1593ms lottery.
Stamp every registered connection and in-flight dial with a network
generation, bumped by MarkNetworkChange, which replaces the immediate
full sweep with one delayed by a configurable 500ms. The sweep then
cuts only registrations older than the last change: subsystems that
redialed on their own hold fresh-generation connections and survive,
so the callers no longer need cancellation logic around the sweep.
A connection inherits its dial's generation, because the socket was
bound to the network that was default when the dial started.
This fixes the sweep being cancelled by the engine's management-level
reconnect while the relay was still down, and lets the mobile
notifiers shrink to plain forwarders.
Replace the fixed 1.5s pre-dial sleep with a wait driven by the OS
availability verdict: while online only a 200ms settle window is spent
(the offline flag lands a few milliseconds after the disconnect that
triggered the reconnect), and while offline the guard waits for the
network to return, bounded by the same 1.5s budget as before. Without
an injected netState the behavior is unchanged.
Measured on Android: relay recovery after a network switch drops from
1.7-1.9s to under 500ms after the cut.
Prepares the repository for the release-branch process agreed internally:
one long-lived release-0.N branch per minor, with fixes backported by
cherry-pick and patch releases tagged from the branch.
Pushes to release-* branches now run the Release workflow and publish
immutable sha-* container images, the way pushes to main already do, so
a release branch can be tested before it is tagged. Release branches
never publish the floating "main" image tag. The push-to-main CI
workflows (Go tests on all platforms, frontend UI, install script,
mobile/wasm validation, infrastructure files, license check) also run
on release-* pushes; pull request triggers were already unfiltered, so
backport PRs were covered — this closes the post-merge gap.
Releases are no longer marked latest before signing: make_latest is
now false in all four goreleaser configs, so a release stays published
but not latest until the signing pipeline uploads the signed Windows
and macOS artifacts and marks it latest itself. Previously the release
became GitHub's "Latest release" at publish time, and the download
endpoints that resolve through the latest-release API could serve a
release whose signed installers did not exist yet. prerelease: auto
additionally labels rc tags as prereleases, so a release candidate can
never take the latest slot.
The trigger_sync_tag job is removed: it dispatched a downstream
image build on every v* tag (release candidates included), which would
race the deliberate release-branch build on every release. The android
and ios submodule bumps are unchanged.
Also sets perennial-regex = "^release-" so git-town never syncs or
ships a release branch into main.
The NSIS installer deleted HKLM/HKCU CurrentVersion\Run values it never
writes, which matches common AV heuristics for unwanted Run-key
manipulation and is suspected to contribute to Windows Defender and
third-party antivirus false positives on the installer.
Drop all autostart registry deletions from both the install and
uninstall sections so the installer only touches keys it creates
itself. Cleanup of the legacy machine-wide entry written by old
installers is left to documentation.
Extends the approach of the closed PR #6735, which only removed the
per-user deletion on uninstall.
Rename ClientOption to Option in the management and signal client
packages so callers do not repeat the package name, and silence nilerr
on the offline-wait path where a cancelled context means shutdown
rather than a retryable failure.
e2e/harness documents itself as feature-agnostic, but three details
assumed the caller lives in this repo, so the terraform provider's
acceptance suite would otherwise carry a second harness for the same
product.
repoRoot took the first module root above the working directory as the
Docker build context, which from another module is the caller's own
root, with no combined/Dockerfile.multistage in it. It now requires that
ancestor to be this module, and otherwise asks the go tool for the
source: for a dependent, the extracted directory of the version it pins,
so the server matches the client library it was compiled against. That
lookup uses -mod=readonly, since automatic vendor mode otherwise reports
an empty Dir.
Geolocation was disabled unconditionally. Agent-network ingest does not
use it, but location-based posture checks need the database, and a rule
management cannot evaluate fails rather than passing.
StartClient pinned one network alias and set no hostname, so a second
agent could not start and a peer's name was arbitrary. Management
records that hostname, making it the peer's name in the API.
The client entrypoint is copied with an explicit mode: git tracks it
100755, but the module cache extracts 0444, so a dependent's build
produced a container exiting with "permission denied".
Adds CombinedOption, WithGeolocation, WithServerEnv, ClientOption and
WithClientName.
The effective state was computed under serverStateLock but published after
releasing it, so two transitions could reorder between compute and notify
and leave the listener on a state the notifier had already superseded,
e.g. NoNetwork surviving after the network came back.
Hold a publish lock across compute and notify on every publishing path.
- Align default names and reuse same environment variables
- With the uploads now targeting the same stable/yum paths as the GTK4
packages, two packages named netbird-ui with the same version and arch
would collide in the repo indexes. Give the GTK3 variant its own
package name and mark the two as conflicting alternatives.
---------
Co-authored-by: Zoltan Papp <zoltan.pmail@gmail.com>
The listener adapter introduced with the network state work turned a nil
listener into a non-nil peer.Listener holding a nil delegate, so the
notifier's nil check passed it through and setListener panicked on its
immediate OnAddressChanged callback. EngineRunner already forwards null,
so the path is reachable.
Drop the listener instead when it is nil, on Android and iOS alike.
pollUntil span a goroutine that looped forever when the condition never
held, which is exactly the path the test takes when it fails. Pass the
test context in and give up when it is done.
The connection and dial registries keyed on a bare uint64, which says
nothing about what the number identifies. Introduce sweepID so the maps,
the counter and the id fields state their intent. No behavior change.
Ticks taken while the OS reported no network were skipped, but they still
advanced the exponential backoff, so a peer could be tens of seconds from its
next attempt by the time connectivity returned. Recovery then depended on a
signal or relay event, which never arrives when both stayed up across the
outage, e.g. a short airplane mode toggle over Wi-Fi. A peer parked in ICE
hourly mode stayed there for the same reason.
React to the offline-to-online transition directly: re-arm the reconnect
ticker and reset the ICE retry state, so the peer retries at once. Add
netstate.Changed for callers that own a select loop and cannot block in Wait.
Move the Dial type and its Ctx/Release methods above the Sweeper in
netsweep.go, and relocate SetNetworkAvailable / NotifyNetworkChange below
GetTunSettings in the Android binding. Pure code moves, no behavior change.
WrapDialContext and WrapConn registered the dial and the connection
independently, so a sweep landing between the dial finishing and WrapConn
cancelled only the dial registration: the connection dialed on the old
network entered the fresh registry and survived the network change.
Replace the pair with a Dial handle. Sweep marks pending dials under the
sweeper mutex, and WrapConn decides under the same mutex: a swept dial's
connection is closed and ErrSwept returned, so the caller redials on the
new network; otherwise the connection transfers to the registry with no
window in between.
connPair closed the accepted connection right after the handshake, so the
reads in TestSweepClosesRegisteredConns failed on the peer's own close
rather than on the sweep. The test passed even with Sweep's close loop
removed. Hold the peer until cleanup so the read errors come from Sweep.
sweptConn embeds net.Conn, so it does not promote Protocol() from the
concrete relay connection. Asserting transportConn on the wrapper always
failed and left c.transport empty. Read the transport off the dialed
connection first, then wrap it.
Regular (non-NetBird) servers used InsecureIgnoreHostKey while also
offering the user's password, so an impersonating endpoint could collect
it. Replace that with a per-profile known-hosts store: an unknown host
returns a marker carrying the fingerprint so the client can show it and,
once confirmed, retry with the key trusted and persisted; a changed key
is rejected outright, as OpenSSH does. The confirmation is single-use and
cleared once the key is stored.
The server-type switch now handles the regular case explicitly and
rejects unknown types instead of routing them through the unverified
path. Java sets the store path (per profile, since an overlay IP is a
different host under a different profile) and can drop a host's key once
no session targets it.
DialContext limited only the TCP establishment, so a peer that accepted
the connection and then stayed silent left gossh.NewClientConn blocking
forever and the terminal stuck on "Connecting".
Set the socket deadline from the dial context before the handshake and
clear it on success, so the handshake shares the dial timeout instead of
being able to hang. Verified against a silent listener: the connect now
returns i/o timeout instead of blocking.
Any authentication failure on a regular server returned the
password-required marker, so against a server with password
authentication disabled the client asked again after every attempt and
reported each one as a wrong password.
gossh only lists a method under "attempted methods" when the server
offered it. When a supplied password never got attempted, surface the
real error instead of the marker, the same way the desktop client
reports it. A first connect without a password still prompts.