Files
netbird/client/internal
Riccardo ManfrinandZoltán Papp 3934e7c697 [client] Evaluate the session deadline against the wall clock (#7959)
* Skip session warnings that fire after their window

The warning timers run on the monotonic clock, which does not advance
while an Android device is suspended. A timer armed for T-10 or T-2 can
therefore fire long after the window it was armed for, delivering a
"session expires soon" notification once that window is already gone.

Gate both callbacks on the wall clock at fire time: the T-10 warning is
skipped once the final-warning window has been reached, and the final
warning is skipped once the deadline itself has passed. Both set their
edge guard before returning so a skipped warning cannot fire again for
the same deadline.

* Harden the late-warning guards

Clamp a non-positive final lead to zero in the T-10 guard so a disabled
final warning cannot move the cutoff past the deadline, matching how
armTimerLocked already treats it.

Strip the monotonic reading from both sides of the comparison so the
guard measures wall-clock time regardless of how the caller built the
deadline. The production deadline comes from a protobuf timestamp and
has no monotonic reading; this keeps the guard correct for callers that
derive one from time.Now.

* Log the deadline and lateness on skipped warnings

Include the deadline and how far past the cutoff the timer fired, so a
debug bundle shows how long the device was suspended.

* Inject the clock into the late-warning guard and cover it with tests

The guard read time.Now internally, so the skip paths were reachable
only through a deadline already in the past and the boundary depended
on real time. Extract the comparison into isLate and read the time
through a nowFn field, so tests can place a resume anywhere around the
deadline without sleeping.

* Send the final warning when the T-10 timer fires inside its window

A suspend between roughly eight and ten minutes long made the T-10
timer fire inside the final-warning window and the final timer fire
after the deadline, so both were skipped and a user who resumed with
time left got no warning at all. When the T-10 timer fires late but
before the deadline, send the final warning in its place and mark it
fired so the delayed final timer does not repeat it.

* Respect dismissal when promoting a late warning to the final one

fireFinal skips the final warning once the user dismissed the deadline,
but the promoted path did not, so a dismissed deadline could still get
a final warning. Check the dismissal first, and give each skip reason
its own log line so an already-fired final warning no longer logs a
negative lateness.

* Add a deadline-only mode to the session watcher

Android will schedule its own expiry warnings from the deadline, so
the engine must not arm the T-10 and T-2 timers there. NewDeadlineOnly
keeps the deadline validation, the recorder propagation and the
logging, and skips only the timers, so the status snapshot the app
reads stays correct and an out-of-range deadline is still rejected.

* Use the deadline-only watcher on Android and drop the warning callbacks

The warning timers run on the monotonic clock, which does not advance
while the device sleeps, so a warning armed for T-10 could fire long
after its window. The app now schedules the warnings itself with
WorkManager, anchored to the wall clock, from the deadline it reads
through SessionExpiresAtUnix on every OnStateChanged.

Wire the deadline-only watcher into the android build and remove the
event-driven path from the gomobile surface: OnSessionExpiring, the
event subscription behind it and DismissSessionWarning, which the app
never called.

* Describe the late-warning guard without naming Android

The guard stays for the desktop builds, where a timer can also stall
across a sleep. Android no longer arms the timers at all.

* [client] Evaluate the session deadline against the wall clock

The session deadline is an absolute instant published by management, but
the warnings for it were armed as relative timers. A relative timer runs
on the monotonic clock, which does not advance while a device is
suspended: it fires once that much awake time has passed, which can be
long after the deadline, and nothing re-evaluates when the device wakes
up. A device that suspends before the warning window and resumes with
minutes left is never told it can still extend. The same applies to a
device that boots before NTP has corrected its clock: the one-shot timer
has already fired by the time the correction lands.

Compare the tracked deadline against the wall clock on a ticker instead.
Every tick after a resume, or after a clock correction, sees the real
remaining time, so the warning is published whenever there is still time
to act on it and never once the deadline has passed. This supersedes the
late-callback guards: with no relative timer there is no late callback to
detect.

The warning windows keep their semantics. Each one publishes at most once
per deadline value, a dismissal still suppresses the final warning, and
inside the final window the interactive warning is skipped in favour of
the final one, since a device resuming there never saw it. Update
evaluates immediately so a deadline that already sits inside a window
does not wait for a tick, and Close stops the loop and waits for it.

Two trade-offs: a warning can land up to one tick (10s) after its lead,
which is noise against leads measured in minutes, and the loop runs from
the first tracked deadline until Close rather than being armed and
disarmed per deadline.

* [client] Poll whatever the clock reads when the deadline arrives

Update started the loop only for a deadline in the future, and read the
wall clock directly to decide. A device whose clock runs ahead before NTP
corrects it accepts the deadline as recent-past, starts nothing, and then
has no loop left to notice the correction: the next sync carries the same
deadline value, which Update treats as a no-op. The warning is lost for
exactly the case this mechanism exists to cover.

Start the loop for every accepted deadline. evaluate already ignores a
deadline that has genuinely passed, so the only cost is a ticker on an
expired session until the next deadline or Close.

Take the current time from nowFn there as well, so the sanity checks and
the evaluation agree on what time it is and a test can drive both.

* [client] Announce the deadline before any warning about it

Update released the lock before telling the recorder about the new
deadline, so a tick landing in that window could publish a warning for a
deadline consumers had not been told about yet, inverting the order the
function documents.

Gate publishing on the deadline the recorder has been told about: Update
records it after the recorder call, and an evaluation that finds the gate
shut leaves the warning for the next tick.

* [client] Describe the deadline-only mode in terms of the poll

The mode no longer skips arming timers, it skips the evaluation entirely,
and the doc comment said otherwise.

* [client] Publish warnings only from the evaluation loop

Update evaluated the new deadline on the caller's goroutine, so a warning
could still be on its way to the recorder after Close had returned: Close
waits for the evaluation loop, and that publish was not coming from it.

Hand the work to the loop instead. Update announces the deadline and
nudges a buffered wake channel, so a deadline that already sits inside a
warning window is still warned about at once, without this goroutine ever
touching the recorder. The loop is now the only caller of evaluate, which
makes waiting for it in Close enough.

* [client] Make a concurrent Close wait for the loop as well

The first Close stopped the evaluation loop and waited for it, but a
second concurrent Close saw the closed flag and returned straight away,
telling its caller the teardown was done while a warning was still on its
way to the recorder.

Keep the completion channel on the receiver after the teardown starts, so
the call that loses the race waits on the same loop.

* [client] Correct the poll start condition in the doc comment

Update starts the loop for every accepted deadline, including one that
already reads as expired, not only for a future one.

* [client] Unblock the parked publish on every test exit path

The concurrent-Close test parks the evaluation loop inside a publish and
releases it at the end. A t.Fatal before that line runs Goexit and skips
the release, stranding the loop and both Close calls for the rest of the
package run.

Release through a sync.Once registered with t.Cleanup, and close the
watcher there too, so an early failure reports itself instead of hanging.

* [client] Release the parked publish before closing in the test cleanup

t.Cleanup runs in reverse registration order, so the watcher was closed
before the release ran: Close waits for the evaluation loop, the loop was
still parked in the publish, and a t.Fatal hung the cleanup instead of
reporting the failure.

Register one cleanup that releases first and closes after.

---------

Co-authored-by: Zoltán Papp <zoltan.pmail@gmail.com>
2026-10-09 11:44:08 +02:00
..