Riccardo ManfrinandZoltán Papp 3934e7c697 [client] Evaluate the session deadline against the wall clock (#7959)
* Skip session warnings that fire after their window

The warning timers run on the monotonic clock, which does not advance
while an Android device is suspended. A timer armed for T-10 or T-2 can
therefore fire long after the window it was armed for, delivering a
"session expires soon" notification once that window is already gone.

Gate both callbacks on the wall clock at fire time: the T-10 warning is
skipped once the final-warning window has been reached, and the final
warning is skipped once the deadline itself has passed. Both set their
edge guard before returning so a skipped warning cannot fire again for
the same deadline.

* Harden the late-warning guards

Clamp a non-positive final lead to zero in the T-10 guard so a disabled
final warning cannot move the cutoff past the deadline, matching how
armTimerLocked already treats it.

Strip the monotonic reading from both sides of the comparison so the
guard measures wall-clock time regardless of how the caller built the
deadline. The production deadline comes from a protobuf timestamp and
has no monotonic reading; this keeps the guard correct for callers that
derive one from time.Now.

* Log the deadline and lateness on skipped warnings

Include the deadline and how far past the cutoff the timer fired, so a
debug bundle shows how long the device was suspended.

* Inject the clock into the late-warning guard and cover it with tests

The guard read time.Now internally, so the skip paths were reachable
only through a deadline already in the past and the boundary depended
on real time. Extract the comparison into isLate and read the time
through a nowFn field, so tests can place a resume anywhere around the
deadline without sleeping.

* Send the final warning when the T-10 timer fires inside its window

A suspend between roughly eight and ten minutes long made the T-10
timer fire inside the final-warning window and the final timer fire
after the deadline, so both were skipped and a user who resumed with
time left got no warning at all. When the T-10 timer fires late but
before the deadline, send the final warning in its place and mark it
fired so the delayed final timer does not repeat it.

* Respect dismissal when promoting a late warning to the final one

fireFinal skips the final warning once the user dismissed the deadline,
but the promoted path did not, so a dismissed deadline could still get
a final warning. Check the dismissal first, and give each skip reason
its own log line so an already-fired final warning no longer logs a
negative lateness.

* Add a deadline-only mode to the session watcher

Android will schedule its own expiry warnings from the deadline, so
the engine must not arm the T-10 and T-2 timers there. NewDeadlineOnly
keeps the deadline validation, the recorder propagation and the
logging, and skips only the timers, so the status snapshot the app
reads stays correct and an out-of-range deadline is still rejected.

* Use the deadline-only watcher on Android and drop the warning callbacks

The warning timers run on the monotonic clock, which does not advance
while the device sleeps, so a warning armed for T-10 could fire long
after its window. The app now schedules the warnings itself with
WorkManager, anchored to the wall clock, from the deadline it reads
through SessionExpiresAtUnix on every OnStateChanged.

Wire the deadline-only watcher into the android build and remove the
event-driven path from the gomobile surface: OnSessionExpiring, the
event subscription behind it and DismissSessionWarning, which the app
never called.

* Describe the late-warning guard without naming Android

The guard stays for the desktop builds, where a timer can also stall
across a sleep. Android no longer arms the timers at all.

* [client] Evaluate the session deadline against the wall clock

The session deadline is an absolute instant published by management, but
the warnings for it were armed as relative timers. A relative timer runs
on the monotonic clock, which does not advance while a device is
suspended: it fires once that much awake time has passed, which can be
long after the deadline, and nothing re-evaluates when the device wakes
up. A device that suspends before the warning window and resumes with
minutes left is never told it can still extend. The same applies to a
device that boots before NTP has corrected its clock: the one-shot timer
has already fired by the time the correction lands.

Compare the tracked deadline against the wall clock on a ticker instead.
Every tick after a resume, or after a clock correction, sees the real
remaining time, so the warning is published whenever there is still time
to act on it and never once the deadline has passed. This supersedes the
late-callback guards: with no relative timer there is no late callback to
detect.

The warning windows keep their semantics. Each one publishes at most once
per deadline value, a dismissal still suppresses the final warning, and
inside the final window the interactive warning is skipped in favour of
the final one, since a device resuming there never saw it. Update
evaluates immediately so a deadline that already sits inside a window
does not wait for a tick, and Close stops the loop and waits for it.

Two trade-offs: a warning can land up to one tick (10s) after its lead,
which is noise against leads measured in minutes, and the loop runs from
the first tracked deadline until Close rather than being armed and
disarmed per deadline.

* [client] Poll whatever the clock reads when the deadline arrives

Update started the loop only for a deadline in the future, and read the
wall clock directly to decide. A device whose clock runs ahead before NTP
corrects it accepts the deadline as recent-past, starts nothing, and then
has no loop left to notice the correction: the next sync carries the same
deadline value, which Update treats as a no-op. The warning is lost for
exactly the case this mechanism exists to cover.

Start the loop for every accepted deadline. evaluate already ignores a
deadline that has genuinely passed, so the only cost is a ticker on an
expired session until the next deadline or Close.

Take the current time from nowFn there as well, so the sanity checks and
the evaluation agree on what time it is and a test can drive both.

* [client] Announce the deadline before any warning about it

Update released the lock before telling the recorder about the new
deadline, so a tick landing in that window could publish a warning for a
deadline consumers had not been told about yet, inverting the order the
function documents.

Gate publishing on the deadline the recorder has been told about: Update
records it after the recorder call, and an evaluation that finds the gate
shut leaves the warning for the next tick.

* [client] Describe the deadline-only mode in terms of the poll

The mode no longer skips arming timers, it skips the evaluation entirely,
and the doc comment said otherwise.

* [client] Publish warnings only from the evaluation loop

Update evaluated the new deadline on the caller's goroutine, so a warning
could still be on its way to the recorder after Close had returned: Close
waits for the evaluation loop, and that publish was not coming from it.

Hand the work to the loop instead. Update announces the deadline and
nudges a buffered wake channel, so a deadline that already sits inside a
warning window is still warned about at once, without this goroutine ever
touching the recorder. The loop is now the only caller of evaluate, which
makes waiting for it in Close enough.

* [client] Make a concurrent Close wait for the loop as well

The first Close stopped the evaluation loop and waited for it, but a
second concurrent Close saw the closed flag and returned straight away,
telling its caller the teardown was done while a warning was still on its
way to the recorder.

Keep the completion channel on the receiver after the teardown starts, so
the call that loses the race waits on the same loop.

* [client] Correct the poll start condition in the doc comment

Update starts the loop for every accepted deadline, including one that
already reads as expired, not only for a future one.

* [client] Unblock the parked publish on every test exit path

The concurrent-Close test parks the evaluation loop inside a publish and
releases it at the end. A t.Fatal before that line runs Goexit and skips
the release, stranding the loop and both Close calls for the rest of the
package run.

Release through a sync.Once registered with t.Cleanup, and close the
watcher there too, so an early failure reports itself instead of hanging.

* [client] Release the parked publish before closing in the test cleanup

t.Cleanup runs in reverse registration order, so the watcher was closed
before the release ran: Close waits for the evaluation loop, the loop was
still parked in the publish, and a t.Fatal hung the cleanup instead of
reporting the failure.

Register one cleanup that releases first and closes after.

---------

Co-authored-by: Zoltán Papp <zoltan.pmail@gmail.com>
2026-10-09 11:44:08 +02:00

Start using NetBird at netbird.io
See Documentation
Join our Slack channel or our Community forum


🚀 We are hiring! Join us at https://netbird.io/careers

🤖 NetBird Agent Network (Beta)

Identity-aware access control for AI agents — keyless access to LLM APIs and private resources over the encrypted NetBird tunnel. See agent-network/ or read the docs at netbird.ai.

NetBird combines a configuration-free peer-to-peer private network and a centralized access control system in a single platform, making it easy to create secure private networks for your organization or home.

Connect. NetBird creates a WireGuard-based overlay network that automatically connects your machines over an encrypted tunnel, leaving behind the hassle of opening ports, complex firewall rules, VPN gateways, and so forth.

Secure. NetBird enables secure remote access by applying granular access policies while allowing you to manage them intuitively from a single place. Works universally on any infrastructure.

https://github.com/user-attachments/assets/10cec749-bb56-4ab3-97af-4e38850108d2

Self-host NetBird (video)

Watch the video

Key features

Connectivity Management Security Automation Platforms
✓ Kernel WireGuard ✓ Admin Web UI ✓ SSO & MFA support ✓ Public API ✓ Linux
✓ Peer-to-peer connections ✓ Auto peer discovery and configuration ✓ Access control: groups & rules ✓ Setup keys for bulk provisioning ✓ macOS
✓ Connection relay fallback ✓ IdP integrations ✓ Activity logging ✓ Self-hosting quickstart script ✓ Windows
✓ Routes to external networks ✓ Private DNS ✓ Traffic events ✓ IdP groups sync with JWT ✓ Android
✓ Domain-based DNS routes ✓ Custom DNS zones ✓ Device posture checks ✓ Terraform provider ✓ Android TV
✓ Exit nodes ✓ Multiuser support ✓ Peer-to-peer encryption ✓ Ansible collection ✓ iOS
✓ IPv6 dual-stack overlay ✓ Multi-account profile switching ✓ SSH with central access policies ✓ Apple TV
✓ Browser SSH & RDP ✓ Quantum-resistance with Rosenpass ✓ FreeBSD
✓ Reverse proxy with auto-TLS ✓ Periodic re-authentication ✓ pfSense
✓ OPNsense
✓ MikroTik RouterOS
✓ OpenWRT
✓ Synology
✓ TrueNAS
✓ Proxmox
✓ Raspberry Pi
✓ Serverless
✓ Container

Quickstart with NetBird Cloud

Quickstart with self-hosted NetBird

This is the quickest way to try self-hosted NetBird. It should take around 5 minutes to get started if you already have a public domain and a VM. Follow the Advanced guide with a custom identity provider for installations with different IdPs.

Infrastructure requirements:

  • A Linux VM with at least 1 CPU and 2 GB of memory.
  • The VM should be publicly accessible on TCP ports 80 and 443 and UDP port 3478.
  • A public domain name pointing to the VM.

Software requirements:

Steps

  • Download and run the installation script:
export NETBIRD_DOMAIN=netbird.example.com; curl -fsSL https://github.com/netbirdio/netbird/releases/latest/download/getting-started.sh | bash

A bit on NetBird internals

  • Every machine in the network runs the NetBird agent, which manages WireGuard.
  • Every agent connects to the Management Service, which holds network state, manages peer IPs, and distributes updates to agents.
  • Agents use ICE (via pion/ice) to discover connection candidates for peer-to-peer connections.
  • Candidates are discovered with the help of STUN servers.
  • Agents negotiate a connection through the Signal Service, exchanging end-to-end encrypted messages with candidates.
  • When NAT traversal fails (e.g. mobile carrier-grade NAT) and a direct p2p connection isn't possible, the system falls back to a Relay Service and a secure WireGuard tunnel is established through it.

NetBird high-level architecture diagram

See a complete architecture overview for details.

Reporting bugs and requesting features

NetBird uses a discussion-first workflow. Bug reports and feature requests start in Discussions, not as issues.

What you want to do Where to go
Report a bug, regression, or unexpected behavior Issue Triage
Request a feature or share an idea Ideas & Feature Requests
Ask about setup, configuration, or self-hosting Q&A / Support
Report a security vulnerability Security policy, never a public thread

Our team and maintainers triage discussions, ask follow-up questions, check for duplicates, and reproduce bugs. Validated reports are promoted to issues. This keeps the issue tracker a clear answer to one question: what is the team working on.

Please search existing discussions and issues first, including closed ones. If something similar already exists, upvote it and add your details there instead of opening a duplicate.

For bug reports, include your NetBird version, operating system, deployment type (Cloud, self-hosted, Kubernetes, or Docker), reproduction steps, expected and actual behavior, and a debug bundle where relevant:

netbird version
netbird status -d -A
netbird debug for 1m -A -S -U

-U uploads the bundle and prints a file key you can paste instead of attaching the archive. -A anonymizes the output, which matters on a public thread. It masks most identifying details but is not full redaction, so read the bundle before posting it. Two levels are available:

Level How to select What it masks
default -A / --anonymize Public IP addresses, IPv6 ULA addresses, MAC addresses, and domains other than netbird.io, netbird.cloud, netbird.selfhosted, and netbird.stage. IPv4 private, CGNAT, and link-local ranges are kept
strict --anonymize-level strict (implies -A) The above, plus IPv4 private, CGNAT, and link-local ranges, peer names in front of netbird.cloud, netbird.selfhosted, and netbird.stage, and WireGuard public keys. Labels under netbird.io are kept, since it only hosts infrastructure

See collecting a debug bundle and the CLI reference for details.

See How to use Discussions, Issues, and Pull Requests for the full workflow, or SUPPORT.md for a shorter version.

Contributing

Contributions are welcome. Read CONTRIBUTING.md first. NetBird works ticket first, anything that changes behavior needs an issue the team has agreed on before you open a pull request.

Community projects

Note: The main branch may be in an unstable or even broken state during development. For stable versions, see releases.

Support acknowledgement

In November 2022, NetBird joined the StartUpSecure program sponsored by the Federal Ministry of Education and Research of the Federal Republic of Germany. Together with the CISPA Helmholtz Center for Information Security, NetBird brings security best practices and simplicity to private networking.

CISPA_Logo_BLACK_EN_RZ_RGB (1)

Acknowledgements

We build on open source technologies like WireGuard®, Pion ICE, and Rosenpass. We greatly appreciate the work these projects are doing, and we'd love it if you could support them too (e.g., by starring or contributing).

This repository is licensed under the BSD-3-Clause license, which applies to all parts of the repository except for the directories management/, signal/ and relay/. Those directories are licensed under the GNU Affero General Public License version 3.0 (AGPLv3). See the respective LICENSE files inside each directory.

WireGuard and the WireGuard logo are registered trademarks of Jason A. Donenfeld.

Languages
Go 94.6%
TypeScript 2.9%
Shell 1.5%
HTML 0.4%
Go Template 0.2%
Other 0.2%