Files
netbird/client/internal/auth/sessionwatch/watcher.go
T
Riccardo ManfrinandZoltán Papp 3934e7c697 [client] Evaluate the session deadline against the wall clock (#7959)
* Skip session warnings that fire after their window

The warning timers run on the monotonic clock, which does not advance
while an Android device is suspended. A timer armed for T-10 or T-2 can
therefore fire long after the window it was armed for, delivering a
"session expires soon" notification once that window is already gone.

Gate both callbacks on the wall clock at fire time: the T-10 warning is
skipped once the final-warning window has been reached, and the final
warning is skipped once the deadline itself has passed. Both set their
edge guard before returning so a skipped warning cannot fire again for
the same deadline.

* Harden the late-warning guards

Clamp a non-positive final lead to zero in the T-10 guard so a disabled
final warning cannot move the cutoff past the deadline, matching how
armTimerLocked already treats it.

Strip the monotonic reading from both sides of the comparison so the
guard measures wall-clock time regardless of how the caller built the
deadline. The production deadline comes from a protobuf timestamp and
has no monotonic reading; this keeps the guard correct for callers that
derive one from time.Now.

* Log the deadline and lateness on skipped warnings

Include the deadline and how far past the cutoff the timer fired, so a
debug bundle shows how long the device was suspended.

* Inject the clock into the late-warning guard and cover it with tests

The guard read time.Now internally, so the skip paths were reachable
only through a deadline already in the past and the boundary depended
on real time. Extract the comparison into isLate and read the time
through a nowFn field, so tests can place a resume anywhere around the
deadline without sleeping.

* Send the final warning when the T-10 timer fires inside its window

A suspend between roughly eight and ten minutes long made the T-10
timer fire inside the final-warning window and the final timer fire
after the deadline, so both were skipped and a user who resumed with
time left got no warning at all. When the T-10 timer fires late but
before the deadline, send the final warning in its place and mark it
fired so the delayed final timer does not repeat it.

* Respect dismissal when promoting a late warning to the final one

fireFinal skips the final warning once the user dismissed the deadline,
but the promoted path did not, so a dismissed deadline could still get
a final warning. Check the dismissal first, and give each skip reason
its own log line so an already-fired final warning no longer logs a
negative lateness.

* Add a deadline-only mode to the session watcher

Android will schedule its own expiry warnings from the deadline, so
the engine must not arm the T-10 and T-2 timers there. NewDeadlineOnly
keeps the deadline validation, the recorder propagation and the
logging, and skips only the timers, so the status snapshot the app
reads stays correct and an out-of-range deadline is still rejected.

* Use the deadline-only watcher on Android and drop the warning callbacks

The warning timers run on the monotonic clock, which does not advance
while the device sleeps, so a warning armed for T-10 could fire long
after its window. The app now schedules the warnings itself with
WorkManager, anchored to the wall clock, from the deadline it reads
through SessionExpiresAtUnix on every OnStateChanged.

Wire the deadline-only watcher into the android build and remove the
event-driven path from the gomobile surface: OnSessionExpiring, the
event subscription behind it and DismissSessionWarning, which the app
never called.

* Describe the late-warning guard without naming Android

The guard stays for the desktop builds, where a timer can also stall
across a sleep. Android no longer arms the timers at all.

* [client] Evaluate the session deadline against the wall clock

The session deadline is an absolute instant published by management, but
the warnings for it were armed as relative timers. A relative timer runs
on the monotonic clock, which does not advance while a device is
suspended: it fires once that much awake time has passed, which can be
long after the deadline, and nothing re-evaluates when the device wakes
up. A device that suspends before the warning window and resumes with
minutes left is never told it can still extend. The same applies to a
device that boots before NTP has corrected its clock: the one-shot timer
has already fired by the time the correction lands.

Compare the tracked deadline against the wall clock on a ticker instead.
Every tick after a resume, or after a clock correction, sees the real
remaining time, so the warning is published whenever there is still time
to act on it and never once the deadline has passed. This supersedes the
late-callback guards: with no relative timer there is no late callback to
detect.

The warning windows keep their semantics. Each one publishes at most once
per deadline value, a dismissal still suppresses the final warning, and
inside the final window the interactive warning is skipped in favour of
the final one, since a device resuming there never saw it. Update
evaluates immediately so a deadline that already sits inside a window
does not wait for a tick, and Close stops the loop and waits for it.

Two trade-offs: a warning can land up to one tick (10s) after its lead,
which is noise against leads measured in minutes, and the loop runs from
the first tracked deadline until Close rather than being armed and
disarmed per deadline.

* [client] Poll whatever the clock reads when the deadline arrives

Update started the loop only for a deadline in the future, and read the
wall clock directly to decide. A device whose clock runs ahead before NTP
corrects it accepts the deadline as recent-past, starts nothing, and then
has no loop left to notice the correction: the next sync carries the same
deadline value, which Update treats as a no-op. The warning is lost for
exactly the case this mechanism exists to cover.

Start the loop for every accepted deadline. evaluate already ignores a
deadline that has genuinely passed, so the only cost is a ticker on an
expired session until the next deadline or Close.

Take the current time from nowFn there as well, so the sanity checks and
the evaluation agree on what time it is and a test can drive both.

* [client] Announce the deadline before any warning about it

Update released the lock before telling the recorder about the new
deadline, so a tick landing in that window could publish a warning for a
deadline consumers had not been told about yet, inverting the order the
function documents.

Gate publishing on the deadline the recorder has been told about: Update
records it after the recorder call, and an evaluation that finds the gate
shut leaves the warning for the next tick.

* [client] Describe the deadline-only mode in terms of the poll

The mode no longer skips arming timers, it skips the evaluation entirely,
and the doc comment said otherwise.

* [client] Publish warnings only from the evaluation loop

Update evaluated the new deadline on the caller's goroutine, so a warning
could still be on its way to the recorder after Close had returned: Close
waits for the evaluation loop, and that publish was not coming from it.

Hand the work to the loop instead. Update announces the deadline and
nudges a buffered wake channel, so a deadline that already sits inside a
warning window is still warned about at once, without this goroutine ever
touching the recorder. The loop is now the only caller of evaluate, which
makes waiting for it in Close enough.

* [client] Make a concurrent Close wait for the loop as well

The first Close stopped the evaluation loop and waited for it, but a
second concurrent Close saw the closed flag and returned straight away,
telling its caller the teardown was done while a warning was still on its
way to the recorder.

Keep the completion channel on the receiver after the teardown starts, so
the call that loses the race waits on the same loop.

* [client] Correct the poll start condition in the doc comment

Update starts the loop for every accepted deadline, including one that
already reads as expired, not only for a future one.

* [client] Unblock the parked publish on every test exit path

The concurrent-Close test parks the evaluation loop inside a publish and
releases it at the end. A t.Fatal before that line runs Goexit and skips
the release, stranding the loop and both Close calls for the rest of the
package run.

Release through a sync.Once registered with t.Cleanup, and close the
watcher there too, so an early failure reports itself instead of hanging.

* [client] Release the parked publish before closing in the test cleanup

t.Cleanup runs in reverse registration order, so the watcher was closed
before the release ran: Close waits for the evaluation loop, the loop was
still parked in the publish, and a t.Fatal hung the cleanup instead of
reporting the failure.

Register one cleanup that releases first and closes after.

---------

Co-authored-by: Zoltán Papp <zoltan.pmail@gmail.com>
2026-10-09 11:44:08 +02:00

488 lines
17 KiB
Go

// Package sessionwatch tracks the SSO session expiry deadline that the
// management server publishes via LoginResponse / SyncResponse and fires
// two warning events at fixed lead times before expiry: an interactive
// T-WarningLead notification and a dismiss-gated T-FinalWarningLead
// fallback dialog.
//
// The deadline is an absolute wall-clock instant, so the watcher compares
// it against the wall clock on a ticker rather than arming a relative
// timer for it. A relative timer runs on the monotonic clock, which does
// not advance while a device is suspended: it fires once that much awake
// time has passed, which can be long after the deadline, and nothing
// re-evaluates when the device wakes up. Polling makes every tick after a
// resume (or after an NTP correction) see the real remaining time.
//
// The watcher is idempotent: Update may be called as often as the network
// map snapshots arrive. Repeating the same deadline is a no-op; a new
// deadline starts a fresh warning cycle.
//
// Warning firing is edge-detected. Each unique deadline value publishes
// each warning at most once.
package sessionwatch
import (
"errors"
"fmt"
"sync"
"time"
log "github.com/sirupsen/logrus"
cProto "github.com/netbirdio/netbird/client/proto"
)
const (
maxPastHorizon = 30 * 24 * time.Hour
// maxDeadlineHorizon caps how far in the future an accepted deadline
// can sit. A timestamp beyond this is almost certainly a protocol
// glitch, and silently tracking a 100-year deadline would hide the bug.
maxDeadlineHorizon = 10 * 365 * 24 * time.Hour
// defaultEvalInterval is how often the tracked deadline is compared
// against the wall clock. The leads are minutes, so a coarse tick costs
// nothing in accuracy and keeps the wakeup cheap on battery-powered
// devices.
defaultEvalInterval = 10 * time.Second
// WarningLead is how far before expiry the first (interactive)
// warning fires. Drives the T-10 OS notification with
// Extend/Dismiss actions.
WarningLead = 10 * time.Minute
// FinalWarningLead is how far before expiry the fallback final
// warning fires. Drives the auto-opened SessionAboutToExpire dialog,
// but only when the user has not dismissed the T-WarningLead warning
// for the same deadline. Must be strictly less than WarningLead.
FinalWarningLead = 2 * time.Minute
)
var (
// ErrDeadlineBeforeEpoch is returned by Update when the supplied
// deadline pre-dates 1970-01-01.
ErrDeadlineBeforeEpoch = errors.New("session deadline before unix epoch")
// ErrDeadlineTooFarFuture is returned by Update when the supplied
// deadline is more than maxDeadlineHorizon in the future.
ErrDeadlineTooFarFuture = errors.New("session deadline too far in the future")
// ErrDeadlineInPast is returned by Update when the supplied deadline
// is more than maxPastHorizon in the past.
ErrDeadlineInPast = errors.New("session deadline in the past")
)
// StatusRecorder is the side-effect surface the watcher drives on every
// state transition. Production wires this to peer.Status (SetSessionExpiresAt
// for deadline change/clear, PublishEvent for the two warnings); tests pass
// a fake recorder so the same surface is observable without an engine.
//
// While the watcher runs, it owns the deadline propagated to the recorder:
// every set, clear and sanity-check rejection routes the value through
// SetSessionExpiresAt, so the SubscribeStatus snapshot the UI reads can
// never drift from the watcher's timer state. (SetSessionExpiresAt fans
// out its own state-change notification, so no separate notify is needed.)
// The recorder is server-scoped and outlives this engine-scoped watcher;
// Close deliberately leaves the recorder value in place so transient engine
// restarts don't blank it — the client run loop clears it on real teardown.
//
// PublishEvent's signature mirrors peer.Status.PublishEvent: the watcher
// composes the metadata internally so the wire format (MetaSession*) is
// owned by sessionwatch, not the caller.
type StatusRecorder interface {
SetSessionExpiresAt(deadline time.Time)
PublishEvent(
severity cProto.SystemEvent_Severity,
category cProto.SystemEvent_Category,
message string,
userMessage string,
metadata map[string]string,
)
}
// Watcher observes the latest session deadline and fires two warnings
// before it expires: the interactive T-WarningLead notification, and the
// fallback T-FinalWarningLead dialog (suppressed when the user dismissed
// the first one for the same deadline). Safe for concurrent use.
type Watcher struct {
lead time.Duration
finalLead time.Duration
interval time.Duration
deadlineOnly bool // record the deadline, leave the warnings to the caller
mu sync.Mutex
current time.Time
firedAt time.Time // deadline value the T-WarningLead warning last published for
finalFiredAt time.Time // deadline value the T-FinalWarningLead warning last published for
dismissedAt time.Time // deadline value the user dismissed via Dismiss(); gates the final warning
announcedAt time.Time // deadline value the recorder has been told about; gates publishing
closed bool
recorder StatusRecorder
nowFn func() time.Time
stop chan struct{} // closed to stop the evaluation loop; nil while it is not running
done chan struct{} // closed by the loop on its way out
wake chan struct{} // buffered nudge asking the loop to evaluate before its next tick
}
// New returns a watcher with the package defaults WarningLead and
// FinalWarningLead. Pass nil for recorder to silence side effects (handy
// in unit tests that exercise sanity checks without observing the publish
// path).
func New(recorder StatusRecorder) *Watcher {
return NewWithLeads(WarningLead, FinalWarningLead, recorder)
}
// NewWithLeads returns a watcher with custom lead times. Useful for tests.
// final must be strictly less than lead; otherwise the final warning takes
// over the whole warning window and the interactive notification never
// shows. A zero final lead disables the final warning entirely (see
// evaluate), leaving the interactive one as the only warning.
func NewWithLeads(lead, final time.Duration, recorder StatusRecorder) *Watcher {
return &Watcher{
lead: lead,
finalLead: final,
interval: defaultEvalInterval,
recorder: recorder,
nowFn: time.Now,
}
}
// NewDeadlineOnly returns a watcher that validates and records deadlines
// but publishes no warnings about them, and runs no evaluation loop to
// decide. Used where the deadline is handed on to something that schedules
// the warnings itself, such as the Android app.
func NewDeadlineOnly(recorder StatusRecorder) *Watcher {
w := New(recorder)
w.deadlineOnly = true
return w
}
// Update sets the latest deadline. Pass the zero time to clear (e.g. when
// a Sync push from the server omits the field because login expiration
// was disabled).
//
// Same-value updates are no-ops. A different non-zero value resets the
// "already fired" guards and starts a fresh warning cycle, evaluated
// immediately so a deadline that already sits inside a warning window
// warns without waiting for the next tick. A deadline already in the past
// (within maxPastHorizon) is recorded as-is and warns nothing: the session
// has expired and consumers render it that way.
//
// Returns one of the sentinel Err* values when the deadline fails the
// sanity checks (pre-epoch, far future, or past beyond maxPastHorizon).
// In every error case the watcher first clears its state so it stays
// consistent with what the caller will push into its other sinks (e.g.
// applySessionDeadline forces a zero deadline into the status recorder
// after a non-nil error).
func (w *Watcher) Update(deadline time.Time) error {
w.mu.Lock()
if w.closed {
w.mu.Unlock()
return nil
}
if deadline.IsZero() {
w.clearLocked()
return nil
}
now := w.nowFn()
switch {
case deadline.Before(time.Unix(0, 0)):
w.clearLocked()
return fmt.Errorf("%w: %v", ErrDeadlineBeforeEpoch, deadline)
case deadline.After(now.Add(maxDeadlineHorizon)):
w.clearLocked()
return fmt.Errorf("%w: %v", ErrDeadlineTooFarFuture, deadline)
case deadline.Before(now.Add(-maxPastHorizon)):
w.clearLocked()
return fmt.Errorf("%w: %v (now=%v)", ErrDeadlineInPast, deadline, now)
}
if deadline.Equal(w.current) {
w.mu.Unlock()
return nil
}
w.current = deadline
// Reset every per-deadline guard so a refreshed deadline starts a fresh
// warning cycle: both edge triggers and the user Dismiss decision
// (the user agreed to the old deadline expiring; a new deadline
// restarts the contract).
w.firedAt = time.Time{}
w.finalFiredAt = time.Time{}
w.dismissedAt = time.Time{}
w.announcedAt = time.Time{}
// Poll every accepted deadline, including one that reads as already
// expired: the clock may be running ahead of real time and get
// corrected later, and evaluate ignores a deadline that has genuinely
// passed anyway.
if !w.deadlineOnly {
w.startPollLocked()
}
recorder := w.recorder
w.mu.Unlock()
if recorder != nil {
recorder.SetSessionExpiresAt(deadline)
}
log.Infof("auth session deadline set to: %s (in %s)", deadline.Format(time.RFC3339), time.Until(deadline).Round(time.Second))
// Open the gate only once the recorder knows the new deadline, so a
// warning that refers to it can never reach consumers before the state
// change itself. A tick landing in between finds the gate shut.
w.mu.Lock()
if w.closed || !w.current.Equal(deadline) {
w.mu.Unlock()
return nil
}
w.announcedAt = deadline
wake := w.wake
w.mu.Unlock()
// Hand the evaluation to the loop rather than running it here, so a
// deadline that already sits inside a warning window is published at
// once without this goroutine ever touching the recorder.
if wake != nil {
select {
case wake <- struct{}{}:
default:
}
}
return nil
}
// Deadline returns the most recently observed deadline. Zero when no
// deadline is currently tracked.
func (w *Watcher) Deadline() time.Time {
w.mu.Lock()
defer w.mu.Unlock()
return w.current
}
// Dismiss records the user's "Dismiss" action against the current deadline
// and suppresses the final warning for that deadline. Idempotent: repeated
// calls are no-ops. A subsequent Update with a fresh deadline resets the
// dismissal so the final-warning cycle starts over.
//
// No-op when the watcher holds no deadline or has been closed.
func (w *Watcher) Dismiss() {
w.mu.Lock()
defer w.mu.Unlock()
if w.closed || w.current.IsZero() {
return
}
if w.dismissedAt.Equal(w.current) {
return
}
w.dismissedAt = w.current
log.Infof("auth session final-warning dismissed for deadline %s", w.current.Format(time.RFC3339))
}
// Close stops the evaluation loop and waits for it to exit. Update calls
// after Close are ignored. The recorder keeps its deadline: the watcher is
// engine-scoped and closes on every engine restart (network change,
// sleep/wake, stream errors) while the SSO deadline stays valid across
// those, so clearing here would blank the UI's "expires in" row on every
// transient reconnect. The client run loop clears the server-scoped
// recorder when it exits for real (Down, profile switch, permanent login
// failure).
func (w *Watcher) Close() {
w.mu.Lock()
if w.closed {
// A concurrent Close is already tearing down. w.done outlives it,
// so this caller waits for the same loop rather than returning
// while a warning is still on its way to the recorder.
done := w.done
w.mu.Unlock()
if done != nil {
<-done
}
return
}
w.closed = true
w.current = time.Time{}
w.firedAt = time.Time{}
w.finalFiredAt = time.Time{}
w.dismissedAt = time.Time{}
w.announcedAt = time.Time{}
// Copy the channels out before releasing the lock: the loop takes w.mu
// on every tick, so waiting for it while holding the lock would
// deadlock. w.done stays on the receiver for the branch above.
stop, done := w.stop, w.done
w.stop, w.wake = nil, nil
w.mu.Unlock()
if stop == nil {
return
}
close(stop)
<-done
}
// clearLocked drops the tracked deadline and notifies the recorder so
// downstream consumers (SubscribeStatus stream, UI) drop their anchor.
// The caller must hold w.mu; this helper releases it before invoking
// the recorder.
func (w *Watcher) clearLocked() {
if w.current.IsZero() {
w.mu.Unlock()
return
}
w.current = time.Time{}
w.firedAt = time.Time{}
w.finalFiredAt = time.Time{}
w.dismissedAt = time.Time{}
w.announcedAt = time.Time{}
recorder := w.recorder
w.mu.Unlock()
if recorder != nil {
recorder.SetSessionExpiresAt(time.Time{})
}
log.Infof("auth session deadline cleared")
}
// startPollLocked starts the evaluation loop unless it is already
// running. The loop starts lazily on the first accepted deadline, so a
// client whose server never publishes a session expiry never pays for a
// ticker, and it runs until Close: the watcher is engine-scoped, and a
// cleared deadline is normally followed by a fresh one on the next sync.
// Caller must hold w.mu.
func (w *Watcher) startPollLocked() {
if w.stop != nil {
return
}
stop := make(chan struct{})
done := make(chan struct{})
wake := make(chan struct{}, 1)
w.stop, w.done, w.wake = stop, done, wake
go w.poll(stop, done, wake, w.interval)
}
// poll re-evaluates the tracked deadline every interval, and as soon as a
// new deadline is announced. It is the only caller of evaluate, so Close
// waiting for it to exit is enough to know no warning is still on its way
// to the recorder. Its channels and interval are passed in rather than
// read off the receiver, so Close can clear them without racing it.
func (w *Watcher) poll(stop <-chan struct{}, done chan<- struct{}, wake <-chan struct{}, interval time.Duration) {
defer close(done)
ticker := time.NewTicker(interval)
defer ticker.Stop()
for {
select {
case <-stop:
return
case <-wake:
w.evaluate()
case <-ticker.C:
w.evaluate()
}
}
}
// evaluate compares the tracked deadline against the wall clock and
// publishes whichever warning the remaining time calls for. Inside the
// final-warning window the interactive warning is stale, so the final one
// is published in its place: a device that resumes there never saw the
// T-WarningLead notification.
func (w *Watcher) evaluate() {
w.mu.Lock()
if w.closed || w.deadlineOnly || w.current.IsZero() {
w.mu.Unlock()
return
}
deadline := w.current
if !w.announcedAt.Equal(deadline) {
// Update is still on its way to the recorder with this deadline.
w.mu.Unlock()
return
}
// Round(0) strips the monotonic reading so the comparison is wall
// clock on both sides, whether the deadline came off the wire or from
// a caller that derived it from time.Now.
remaining := deadline.Round(0).Sub(w.nowFn().Round(0))
switch {
case remaining <= 0:
// Already expired: the post-mortem SessionExpired flow owns it.
w.mu.Unlock()
case w.finalLead > 0 && remaining <= w.finalLead:
w.publishFinalLocked(deadline, remaining)
case remaining <= w.lead:
w.publishWarningLocked(deadline, remaining)
default:
w.mu.Unlock()
}
}
// publishWarningLocked emits the interactive T-WarningLead warning, at
// most once per deadline value. Caller must hold w.mu; this helper
// releases it.
func (w *Watcher) publishWarningLocked(deadline time.Time, remaining time.Duration) {
if w.firedAt.Equal(deadline) {
w.mu.Unlock()
return
}
w.firedAt = deadline
recorder := w.recorder
w.mu.Unlock()
if recorder == nil {
return
}
log.Infof("auth session expiry soon warning fired for deadline %s (in %s)",
deadline.Format(time.RFC3339), remaining.Round(time.Second))
publishWarning(recorder, deadline, false)
}
// publishFinalLocked emits the final warning, at most once per deadline
// value and never once the user dismissed that deadline. It marks the
// interactive warning as handled too: the final window is open, so a
// "expires in WarningLead minutes" notification would be wrong. Caller
// must hold w.mu; this helper releases it.
func (w *Watcher) publishFinalLocked(deadline time.Time, remaining time.Duration) {
if w.finalFiredAt.Equal(deadline) || w.dismissedAt.Equal(deadline) {
w.mu.Unlock()
return
}
w.firedAt = deadline
w.finalFiredAt = deadline
recorder := w.recorder
w.mu.Unlock()
if recorder == nil {
return
}
log.Infof("auth session final-warning fired for deadline %s (in %s)",
deadline.Format(time.RFC3339), remaining.Round(time.Second))
publishWarning(recorder, deadline, true)
}
// publishWarning composes the SystemEvent for a watcher-fired warning and
// pushes it through the recorder. Severity is CRITICAL on both — bypassing
// the user's Notifications toggle is deliberate: missing the warning
// window forces the post-mortem SessionExpired flow (tunnel torn down,
// lock icon, manual re-login), which is the UX we are trying to avoid.
func publishWarning(recorder StatusRecorder, deadline time.Time, final bool) {
lead := WarningLead
message := "session expiry warning"
meta := map[string]string{
MetaSessionWarning: "true",
MetaSessionExpiresAt: FormatExpiresAt(deadline),
}
if final {
lead = FinalWarningLead
message = "session expiry final warning"
meta[MetaSessionFinal] = "true"
}
meta[MetaSessionLeadMinutes] = FormatLeadMinutes(lead)
recorder.PublishEvent(
cProto.SystemEvent_CRITICAL,
cProto.SystemEvent_AUTHENTICATION,
message,
"",
meta,
)
}