The refresh pushed an update to every connected peer of the account, while only
the peers a certificate check applies to carry a nonce. On an account where a
handful of peers sit behind the check and the rest do not, everyone was woken
several times a day to be handed a map that changed nothing for them.
Push to the sources of the enabled policies whose posture checks include a
certificate check, which is exactly the set that is sent a challenge.
Resolving the set the other way round than the gRPC layer does is the risk here:
a peer the refresh forgets stops being renewed and falls out of its policies
silently, which is the failure this whole mechanism exists to prevent. So the
selection is held against processPeerPostureChecks, the per-peer rule that
decides who receives a challenge in the first place, by a test that asks both
the same question and requires the same answer.
Renewal was timed against the window in two different ways: the period derived
from it, the sweep interval did not. Shortening the window to watch a renewal in
an end-to-end run would have left the refresher still looking for due accounts
every quarter of an hour, so nothing would have been renewed in time and the
test would have reported the feature broken.
Derive the sweep from the period, within bounds that keep a very short window
from spinning and a normal one from checking less often than is useful, and
allow the window itself to be set through the environment so a run can take
seconds instead of half a day. A value that cannot be parsed or falls outside
the bounds keeps the default, because a window nobody intended is a security
property nobody chose, and an override is logged at warning level since it sets
how long a device keeps passing the check after its key is gone.
Every instance has to be given the same value: the window is part of the nonce,
so instances that disagree reject each other's.
A certificate challenge nonce is accepted for its own window and the one before
it, and it only reaches a peer attached to a network map. An account where
nothing changes sends no map, so after a day the peer re-sends the nonce it
still holds, verification rejects its whole proof set, and the certificates
stored for it are dropped. It fails the certificate check and loses every policy
gated on it until some unrelated change happens to push a map. The outage
repairs itself in seconds, which is what makes it expensive: it is intermittent,
it only hits stable networks, and it is not reproducible on demand.
Push the account's peers an update often enough that the nonce they hold is
never close to expiring. Only accounts whose posture checks actually ask for a
certificate are tracked, so a deployment without the feature does no extra work.
The refresh runs from one goroutine over a map of accounts rather than a timer
per account: the period is hours, so one pass every few minutes costs nothing
next to it, and there is no timer to re-arm when an account that falls due
sooner appears. Each account's first run is offset by a hash of its ID, because
the challenge window is global and an instance restart would otherwise arm every
account in the same moment.
The push carries no administrative change, so it is counted as a refresh rather
than an update and stays out of the figures that track what was edited.