backups: record that a config was checked, not only that it changed
The stale-backup card could not be built as designed, and the reason is more important than the card. Dedup means an unchanged configuration writes no revision, so collectedat moves only on a CHANGE. A machine stable for six months has a six-month-old newest revision and is perfectly healthy. Keying a staleness card on revision age would have flagged most of the fleet - exactly the noise that makes a board worth ignoring. Underneath that: ShopDB could not distinguish those cases at all. On a no-op the server returned "unchanged" and wrote nothing, so "we checked yesterday and it matched" was discarded. That fact is the one thing a backup system must be able to prove, and the only record of it was a line in a log file on the PC. lastseenat records the check rather than the change. Touched on every matching post including the no-op; set on creation, since a new revision has by definition just been seen; backfilled from collectedat or createdat so existing rows start from the last moment the config can be PROVEN current, rather than from now - claiming a check that never happened would be worse than silence. The card keys on it, one row per CHAIN rather than per asset: a machine with two part markers can have one still reporting while the other stopped, and a per-asset view would report the machine as fine. It stays deliberately silent about assets never backed up, because whether one SHOULD be is a question only the manifest can answer, and guessing would list a hundred healthy machines. The rule lives in services/staleness.py rather than the route, so it is testable without an auth layer in the way - the same split retention.py uses. Threshold is backups_staledays, default 3, and 0 disables the card.
This commit is contained in:
@@ -0,0 +1,43 @@
|
||||
"""backups: record when a configuration was last confirmed unchanged.
|
||||
|
||||
Dedup means an unchanged config writes NO revision, so `collectedat` moves only
|
||||
when something changes. A machine whose settings have been stable for six months
|
||||
therefore has a six-month-old newest revision and is entirely healthy - and
|
||||
ShopDB had no way to tell it apart from a machine whose backup stopped running
|
||||
six months ago. The evidence that a check happened existed only in a log file on
|
||||
the PC.
|
||||
|
||||
`lastseenat` records the check rather than the change. It is touched on every
|
||||
matching post, including the no-op that writes nothing else.
|
||||
|
||||
Backfilled from `collectedat` (falling back to `createdat`) so existing rows
|
||||
start from the last moment we can actually prove the config was current, rather
|
||||
than from now - claiming a fresh check that never happened would be worse than
|
||||
saying nothing.
|
||||
"""
|
||||
from alembic import op
|
||||
import sqlalchemy as sa
|
||||
|
||||
|
||||
# revision identifiers, used by Alembic.
|
||||
revision = 'backups0002lastseenat'
|
||||
down_revision = 'backups0001baseline'
|
||||
branch_labels = None
|
||||
depends_on = None
|
||||
|
||||
|
||||
def upgrade():
|
||||
columns = {c['name'] for c in
|
||||
sa.inspect(op.get_bind()).get_columns('backuprevisions')}
|
||||
if 'lastseenat' not in columns:
|
||||
op.add_column('backuprevisions',
|
||||
sa.Column('lastseenat', sa.DateTime(), nullable=True))
|
||||
op.execute('UPDATE backuprevisions '
|
||||
'SET lastseenat = COALESCE(collectedat, createdat)')
|
||||
|
||||
|
||||
def downgrade():
|
||||
columns = {c['name'] for c in
|
||||
sa.inspect(op.get_bind()).get_columns('backuprevisions')}
|
||||
if 'lastseenat' in columns:
|
||||
op.drop_column('backuprevisions', 'lastseenat')
|
||||
Reference in New Issue
Block a user