Files move directly between peers over the overlay, with no server in the
path. The receiver listens on the WireGuard address only, so the port is
unreachable from outside the tunnel, and every offer is matched to a known
peer before anything is read.
Consent is the default: an offer carries metadata alone, and no payload
moves until the receiver accepts. Policy is per profile and device-local —
off, ask, or auto-accept, with per-sender exceptions on top.
Policy and history live in the profile's preferences, so removing a profile
takes its file drop state with it. Transfers interrupted by a restart are
settled on load; nothing survives to finish them, and left alone they would
sit in the log as permanently pending.
The Android bindings pull payload bytes through a chunk-returning stream:
gomobile copies a []byte argument into a fresh Java array and never copies
it back, so a fill-my-buffer method would hand back the right length with
no data.
Introduce a namespaced preference store owned by the profile manager,
persisted next to the profile config as <id>.prefs.json and deleted with
the profile. Sections are opaque JSON, so the profile manager stays free
of any feature schema, and writes reuse the state file's atomic path.
Migrate the Android SSH known-hosts store and session list onto it. Both
previously lived outside the profile lifecycle: known hosts in a
per-profile file under filesDir, the session list in Java
SharedPreferences, each needing its own sweep against the live profile
list to avoid outliving the profile they belonged to. A profile ID that
got reused would have inherited the trusted keys of a deleted profile.
Both now share the "ssh" and "ssh-sessions" namespaces of the profile's
preferences, so deleting a profile takes them along and the Java-side
pruning is gone.
The known-hosts entries keep the OpenSSH line format, only the container
changed, and host key verification keeps rejecting a changed key
outright. SetKnownHostsPath becomes SetKnownHostsStore, taking the
config dir and profile ID instead of a file path.
Existing known-hosts files are not migrated: hosts trusted before this
change prompt for confirmation once more, which errs towards safety.
GetOAuthFlow was the only flow factory without a hint parameter, which
forced its callers to apply the hint afterwards through a local setter
interface and a type assertion. Give it the same constructor-style hint
as NewOAuthFlow and set the hint on the concrete flows before they are
handed out as the interface, so a flow is always complete when built
and the caller-side ordering constraint disappears.
An empty hint is a valid value meaning the IdP chooses the account, so
the flows set it unconditionally.
Extract the shared RequestAuthInfo -> Open -> WaitToken sequence from the
login flow and the SSH JWT flow into runOAuthFlow. Open is now called
synchronously by both flows, matching iOS; openers must post their UI
work instead of blocking, which the app-side openers already do.
The SSH flow read its login hint via profilemanager.GetLoginHint, which
resolves desktop-layout files that the Android app never writes, so the
hint was always empty and the device-code flow could prompt for account
selection. Both flows now read the hint from the profile account file
via the config path, taken from authSnapshot so a concurrent profile
switch cannot pair one profile's config with another's hint.
The shared table used by the Android and wasm terminals was a strict
subset of the CLI one, leaving Ctrl+U, Ctrl+D, Ctrl+Z and friends
without explicit mappings. Export the full table from the ssh package
and derive both CLI variants from it; Windows adds its console-specific
modes to a copy so the shared map is never mutated.
Extract the dial-then-handshake sequence into nbssh.Handshake, which
applies the context deadline to the socket for the duration of the
handshake. Previously only the Android client did this; the CLI, wasm
and SSH proxy paths could block forever on a peer that accepts the TCP
connection and then goes silent, since ClientConfig.Timeout is not used
by NewClientConn.
Remove Engine.VerifySSHHostKey and keep GetPeerSSHKey as the only SSH
key API on the Engine. Verification now lives in the ssh package as a
PeerKeyLookup func type implementing HostKeyVerifier, shared by the
android and embed clients.
Align Android logout semantics with the desktop UI and CLI: logging out no
longer deletes the stored account email, so the next login passes it as the
OIDC login_hint and the IdP preselects the account. Removing the profile is
now the operation that deletes the email; previously RemoveProfile left the
account file behind, which the fixed-name default profile would have
inherited on recreation.
Extract the identical PTY session setup shared by the wasm and Android
terminal clients into ssh.StartPTYSession, and move the stored-key host
verification onto the engine so the embed client delegates and the
Android client passes the engine directly as HostKeyVerifier.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
English (en) is the source of truth for UI translation keys; the other
nine locales rely on runtime English fallback for any missing key, so a
gap never surfaces in CI. Add a dependency-free Node check that fails
when any locale declared in _index.json does not carry the exact same
key set as en (missing or orphaned keys), wired into a dedicated
UI Translations workflow that runs on locale changes.
Also close the one existing gap the check found: ja was missing
daemon.outdated.download ("Download Latest").
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Three pieces, each verified on Android across WiFi/cellular switches:
- A sweep-aware backoff wrapper: the first retry after a disconnect
that follows a recent network-change mark comes after 200ms instead
of the randomized [0..1.6s] interval. Any other failure keeps the
unchanged spread, so the clients of a restarted server still scatter
their reconnects.
- The retry sleep wakes on OS network availability transitions
(nbgrpc.Retry): a disconnect that precedes the offline flag by a few
milliseconds no longer sleeps blindly through the recovery - the
loop parks on the netstate gate and resumes the moment the network
returns.
- The connection state is re-checked after WaitForStateChange: a dial
settling in Ready proceeds immediately instead of burning another
backoff round on an already-usable channel.
Measured after a network switch: management and signal recover in
270-470ms deterministically, down from a 312-1593ms lottery.
Stamp every registered connection and in-flight dial with a network
generation, bumped by MarkNetworkChange, which replaces the immediate
full sweep with one delayed by a configurable 500ms. The sweep then
cuts only registrations older than the last change: subsystems that
redialed on their own hold fresh-generation connections and survive,
so the callers no longer need cancellation logic around the sweep.
A connection inherits its dial's generation, because the socket was
bound to the network that was default when the dial started.
This fixes the sweep being cancelled by the engine's management-level
reconnect while the relay was still down, and lets the mobile
notifiers shrink to plain forwarders.
The NSIS installer deleted HKLM/HKCU CurrentVersion\Run values it never
writes, which matches common AV heuristics for unwanted Run-key
manipulation and is suspected to contribute to Windows Defender and
third-party antivirus false positives on the installer.
Drop all autostart registry deletions from both the install and
uninstall sections so the installer only touches keys it creates
itself. Cleanup of the legacy machine-wide entry written by old
installers is left to documentation.
Extends the approach of the closed PR #6735, which only removed the
per-user deletion on uninstall.
The effective state was computed under serverStateLock but published after
releasing it, so two transitions could reorder between compute and notify
and leave the listener on a state the notifier had already superseded,
e.g. NoNetwork surviving after the network came back.
Hold a publish lock across compute and notify on every publishing path.
The listener adapter introduced with the network state work turned a nil
listener into a non-nil peer.Listener holding a nil delegate, so the
notifier's nil check passed it through and setListener panicked on its
immediate OnAddressChanged callback. EngineRunner already forwards null,
so the path is reachable.
Drop the listener instead when it is nil, on Android and iOS alike.
pollUntil span a goroutine that looped forever when the condition never
held, which is exactly the path the test takes when it fails. Pass the
test context in and give up when it is done.
The connection and dial registries keyed on a bare uint64, which says
nothing about what the number identifies. Introduce sweepID so the maps,
the counter and the id fields state their intent. No behavior change.
Ticks taken while the OS reported no network were skipped, but they still
advanced the exponential backoff, so a peer could be tens of seconds from its
next attempt by the time connectivity returned. Recovery then depended on a
signal or relay event, which never arrives when both stayed up across the
outage, e.g. a short airplane mode toggle over Wi-Fi. A peer parked in ICE
hourly mode stayed there for the same reason.
React to the offline-to-online transition directly: re-arm the reconnect
ticker and reset the ICE retry state, so the peer retries at once. Add
netstate.Changed for callers that own a select loop and cannot block in Wait.
Move the Dial type and its Ctx/Release methods above the Sweeper in
netsweep.go, and relocate SetNetworkAvailable / NotifyNetworkChange below
GetTunSettings in the Android binding. Pure code moves, no behavior change.
WrapDialContext and WrapConn registered the dial and the connection
independently, so a sweep landing between the dial finishing and WrapConn
cancelled only the dial registration: the connection dialed on the old
network entered the fresh registry and survived the network change.
Replace the pair with a Dial handle. Sweep marks pending dials under the
sweeper mutex, and WrapConn decides under the same mutex: a swept dial's
connection is closed and ErrSwept returned, so the caller redials on the
new network; otherwise the connection transfers to the registry with no
window in between.
connPair closed the accepted connection right after the handshake, so the
reads in TestSweepClosesRegisteredConns failed on the peer's own close
rather than on the sweep. The test passed even with Sweep's close loop
removed. Hold the peer until cleanup so the read errors come from Sweep.
Regular (non-NetBird) servers used InsecureIgnoreHostKey while also
offering the user's password, so an impersonating endpoint could collect
it. Replace that with a per-profile known-hosts store: an unknown host
returns a marker carrying the fingerprint so the client can show it and,
once confirmed, retry with the key trusted and persisted; a changed key
is rejected outright, as OpenSSH does. The confirmation is single-use and
cleared once the key is stored.
The server-type switch now handles the regular case explicitly and
rejects unknown types instead of routing them through the unverified
path. Java sets the store path (per profile, since an overlay IP is a
different host under a different profile) and can drop a host's key once
no session targets it.
DialContext limited only the TCP establishment, so a peer that accepted
the connection and then stayed silent left gossh.NewClientConn blocking
forever and the terminal stuck on "Connecting".
Set the socket deadline from the dial context before the handshake and
clear it on success, so the handshake shares the dial timeout instead of
being able to hang. Verified against a silent listener: the connect now
returns i/o timeout instead of blocking.
Any authentication failure on a regular server returned the
password-required marker, so against a server with password
authentication disabled the client asked again after every attempt and
reported each one as a wrong password.
gossh only lists a method under "attempted methods" when the server
offered it. When a supplied password never got attempted, surface the
real error instead of the marker, the same way the desktop client
reports it. A first connect without a password still prompts.
The connect path logged the target host, port and username at info level,
which the guidelines reserve for debug and below.
Drop the two connect messages entirely rather than lowering them: both sat
directly in front of a return, so the same error already reaches the caller
and the terminal, and OnConnected reports the success. Keep the detected
server type, since it decides the auth path, but log it without the
endpoint.
The port arrives as an int because gomobile cannot carry uint16 across the
Java boundary, so nothing rejected a value outside the valid range. It
reached strconv.Itoa and only surfaced as a dial failure, after the server
detection had already spent its timeout.
Open and OnLoginSuccess were each started in their own goroutine, so they
raced. Open is what marks the surface as opened on the client side, and
OnLoginSuccess does nothing until it has, so a token that arrived quickly
left the browser sitting in front of the terminal — the dismissal was
dropped rather than delayed.
The login and session-extend flows do not hit this because their two calls
live in separate functions with a blocking wait between them. Here both
are in one function, so ordering has to come from calling them in turn.
Also groups the file's helpers with the code they serve.
The JWT device-code flow opened the verification URL through the URL
opener but never told it the round-trip had finished, so the Custom Tab
stayed in front of the terminal after the token had already been
collected and the user had to dismiss it by hand.
Call OnLoginSuccess once a non-empty token is in hand, which is what the
login and session-extend flows already do; the Android side reacts by
bringing its own activity forward.
On a network switch (e.g. cellular to WiFi) the management, signal and
relay sockets stay bound to the old network and look alive until the OS
tears them down — measured at 5 seconds of dead air on Android, while
the UI kept claiming Connected. The Android client papered over this
with a full engine restart, paying for it with a torn-down TUN device
and discarded peer state.
Introduce client/netsweep: connections register on dial and deregister
on close, and a sweep closes everything registered while aborting
in-flight dials through sweep-cancellable dial contexts. The aborted
dials matter: a relay dial started on the dying network would otherwise
hold the reconnect loop hostage for the QUIC handshake timeout. After a
sweep every failure surfaces as an ordinary read/write error and the
existing retry loops redial immediately on the new network.
The sweeper reaches the three long-lived connections through the same
options that carry the netstate gate: a gRPC dial option wraps the
management and signal transports (reconnects included), and the relay
client wraps its connection in one place for the picker, the guard and
foreign relays alike. Everything is nil-safe; platforms that inject no
sweeper are untouched.
Mobile clients expose the sweep as NotifyNetworkChange. Measured on
Android against the engine restart it replaces: recovery in 1.6s
instead of 3.2s, no Disconnected flash, and the TUN device, WireGuard
config and peer state survive.
Adding OnStateChanged to the gomobile interface forces every Swift
implementation to grow the method before the app builds again. Drop it
from the iOS binding for now — the adapter satisfies the internal
listener with a no-op and the legacy per-state callbacks keep firing —
so the app upgrades on its own schedule. The state constants stay
exported for that follow-up.
On mobile the client kept dialing management, signal, relay and peer
connections while the device had no usable network at all (airplane
mode), burning battery for attempts that cannot succeed. Stopping the
engine is not an option: tearing it down destroys the TUN device, and
traffic can leak outside the tunnel until it is rebuilt.
Add client/netstate, a small gate the platform feeds from its own
connectivity callbacks. Every reconnection loop waits on it instead of
retrying blindly, and resets its backoff when the network returns so
recovery is immediate. The state is injected through functional options
and consumers hold a *State that may be nil, so every platform that does
not report availability behaves exactly as before.
The relay quick-reconnect rechecks availability after its 1.5s wait: the
disconnect that triggers it is usually the first symptom of the network
going away, so the flag typically arrives while it sleeps.
Report the suspension to the UI as well. peer.Listener grows
OnStateChanged with a typed ClientState, re-exported across the gomobile
boundary as integer constants, and the notifier maps Connecting to a new
NoNetwork state while the OS reports no network, so mobile clients can
show "no network available" instead of a misleading "connecting".
Finally, exit the client retry loop cleanly when its context is
cancelled. backoff.WithContext surfaces the bare context error, which
callers could not distinguish from a real failure — on Android that
turned an engine restart into an unrecoverable error.
Connect() reports a password-required marker instead of a raw handshake
error when a regular SSH server turns down the NetBird key, so the caller
can prompt and retry as often as the user needs. NetBird servers are
excluded: they authenticate with a JWT or the NetBird key, so a failure
there is genuine. The marker is a string because gomobile flattens errors
to their message across the binding.
Errors that reach the terminal are unwrapped to their root cause, so a
dial failure reads "i/o timeout" rather than repeating every layer that
added context; the full chain still goes to the log. A normal shell exit
no longer surfaces as "EOF".
Reset() lets a closed client back a reconnect, which keeps the Java-side
session and its scrollback alive across a drop, and the JWT flow now
reports that it is waiting on the browser instead of blocking silently.