mirror of
https://github.com/netbirdio/netbird.git
synced 2026-10-09 23:19:11 +02:00
* Skip session warnings that fire after their window The warning timers run on the monotonic clock, which does not advance while an Android device is suspended. A timer armed for T-10 or T-2 can therefore fire long after the window it was armed for, delivering a "session expires soon" notification once that window is already gone. Gate both callbacks on the wall clock at fire time: the T-10 warning is skipped once the final-warning window has been reached, and the final warning is skipped once the deadline itself has passed. Both set their edge guard before returning so a skipped warning cannot fire again for the same deadline. * Harden the late-warning guards Clamp a non-positive final lead to zero in the T-10 guard so a disabled final warning cannot move the cutoff past the deadline, matching how armTimerLocked already treats it. Strip the monotonic reading from both sides of the comparison so the guard measures wall-clock time regardless of how the caller built the deadline. The production deadline comes from a protobuf timestamp and has no monotonic reading; this keeps the guard correct for callers that derive one from time.Now. * Log the deadline and lateness on skipped warnings Include the deadline and how far past the cutoff the timer fired, so a debug bundle shows how long the device was suspended. * Inject the clock into the late-warning guard and cover it with tests The guard read time.Now internally, so the skip paths were reachable only through a deadline already in the past and the boundary depended on real time. Extract the comparison into isLate and read the time through a nowFn field, so tests can place a resume anywhere around the deadline without sleeping. * Send the final warning when the T-10 timer fires inside its window A suspend between roughly eight and ten minutes long made the T-10 timer fire inside the final-warning window and the final timer fire after the deadline, so both were skipped and a user who resumed with time left got no warning at all. When the T-10 timer fires late but before the deadline, send the final warning in its place and mark it fired so the delayed final timer does not repeat it. * Respect dismissal when promoting a late warning to the final one fireFinal skips the final warning once the user dismissed the deadline, but the promoted path did not, so a dismissed deadline could still get a final warning. Check the dismissal first, and give each skip reason its own log line so an already-fired final warning no longer logs a negative lateness. * Add a deadline-only mode to the session watcher Android will schedule its own expiry warnings from the deadline, so the engine must not arm the T-10 and T-2 timers there. NewDeadlineOnly keeps the deadline validation, the recorder propagation and the logging, and skips only the timers, so the status snapshot the app reads stays correct and an out-of-range deadline is still rejected. * Use the deadline-only watcher on Android and drop the warning callbacks The warning timers run on the monotonic clock, which does not advance while the device sleeps, so a warning armed for T-10 could fire long after its window. The app now schedules the warnings itself with WorkManager, anchored to the wall clock, from the deadline it reads through SessionExpiresAtUnix on every OnStateChanged. Wire the deadline-only watcher into the android build and remove the event-driven path from the gomobile surface: OnSessionExpiring, the event subscription behind it and DismissSessionWarning, which the app never called. * Describe the late-warning guard without naming Android The guard stays for the desktop builds, where a timer can also stall across a sleep. Android no longer arms the timers at all. * [client] Evaluate the session deadline against the wall clock The session deadline is an absolute instant published by management, but the warnings for it were armed as relative timers. A relative timer runs on the monotonic clock, which does not advance while a device is suspended: it fires once that much awake time has passed, which can be long after the deadline, and nothing re-evaluates when the device wakes up. A device that suspends before the warning window and resumes with minutes left is never told it can still extend. The same applies to a device that boots before NTP has corrected its clock: the one-shot timer has already fired by the time the correction lands. Compare the tracked deadline against the wall clock on a ticker instead. Every tick after a resume, or after a clock correction, sees the real remaining time, so the warning is published whenever there is still time to act on it and never once the deadline has passed. This supersedes the late-callback guards: with no relative timer there is no late callback to detect. The warning windows keep their semantics. Each one publishes at most once per deadline value, a dismissal still suppresses the final warning, and inside the final window the interactive warning is skipped in favour of the final one, since a device resuming there never saw it. Update evaluates immediately so a deadline that already sits inside a window does not wait for a tick, and Close stops the loop and waits for it. Two trade-offs: a warning can land up to one tick (10s) after its lead, which is noise against leads measured in minutes, and the loop runs from the first tracked deadline until Close rather than being armed and disarmed per deadline. * [client] Poll whatever the clock reads when the deadline arrives Update started the loop only for a deadline in the future, and read the wall clock directly to decide. A device whose clock runs ahead before NTP corrects it accepts the deadline as recent-past, starts nothing, and then has no loop left to notice the correction: the next sync carries the same deadline value, which Update treats as a no-op. The warning is lost for exactly the case this mechanism exists to cover. Start the loop for every accepted deadline. evaluate already ignores a deadline that has genuinely passed, so the only cost is a ticker on an expired session until the next deadline or Close. Take the current time from nowFn there as well, so the sanity checks and the evaluation agree on what time it is and a test can drive both. * [client] Announce the deadline before any warning about it Update released the lock before telling the recorder about the new deadline, so a tick landing in that window could publish a warning for a deadline consumers had not been told about yet, inverting the order the function documents. Gate publishing on the deadline the recorder has been told about: Update records it after the recorder call, and an evaluation that finds the gate shut leaves the warning for the next tick. * [client] Describe the deadline-only mode in terms of the poll The mode no longer skips arming timers, it skips the evaluation entirely, and the doc comment said otherwise. * [client] Publish warnings only from the evaluation loop Update evaluated the new deadline on the caller's goroutine, so a warning could still be on its way to the recorder after Close had returned: Close waits for the evaluation loop, and that publish was not coming from it. Hand the work to the loop instead. Update announces the deadline and nudges a buffered wake channel, so a deadline that already sits inside a warning window is still warned about at once, without this goroutine ever touching the recorder. The loop is now the only caller of evaluate, which makes waiting for it in Close enough. * [client] Make a concurrent Close wait for the loop as well The first Close stopped the evaluation loop and waited for it, but a second concurrent Close saw the closed flag and returned straight away, telling its caller the teardown was done while a warning was still on its way to the recorder. Keep the completion channel on the receiver after the teardown starts, so the call that loses the race waits on the same loop. * [client] Correct the poll start condition in the doc comment Update starts the loop for every accepted deadline, including one that already reads as expired, not only for a future one. * [client] Unblock the parked publish on every test exit path The concurrent-Close test parks the evaluation loop inside a publish and releases it at the end. A t.Fatal before that line runs Goexit and skips the release, stranding the loop and both Close calls for the rest of the package run. Release through a sync.Once registered with t.Cleanup, and close the watcher there too, so an early failure reports itself instead of hanging. * [client] Release the parked publish before closing in the test cleanup t.Cleanup runs in reverse registration order, so the watcher was closed before the release ran: Close waits for the evaluation loop, the loop was still parked in the publish, and a t.Fatal hung the cleanup instead of reporting the failure. Register one cleanup that releases first and closes after. --------- Co-authored-by: Zoltán Papp <zoltan.pmail@gmail.com>