backups: a revision chain belongs to a PC, not just a machine
Dedup compared a posted config against the latest revision for the ASSET, which is only correct when a machine number means one PC. Several PCs share one here: the part markers on 0613, 0615 and WJPRT are separate devices, differing by COM port, filed under one machine number. Each marker's post therefore differed from whichever marker had posted last, nothing ever deduped, and the table grew by one row per PC per collection cycle. A chain is now (asset, kind, source hostname). An unchanged config is a no-op again, and each PC keeps its own history against the machine. NULL sources - rows written before the column was populated, and hand-loaded ones - form their own chain via IS NULL; `column == None` never matches in SQL, so without that those rows would have re-posted forever. Two consumers assumed the old key and are fixed with it. Retention pruned per asset, so a busy marker's revisions could evict a quiet marker's only backup; it now prunes each chain separately, protecting the newest and oldest of each. The revision diff compared against the previous revision on the machine, which across two markers reported one device's COM port as a change on the other; it now compares within the source's own chain. scripts/collapse_duplicate_backup_revisions.py cleans up what the old rule wrote. It removes only a revision whose hash repeats the one before it in the same chain - rows the fixed code would never have written - and keeps every genuine change, every chain's newest and oldest, and every source. Dry run by default. Its --report mode explains what grew each chain, which separates a legitimately shared machine number from two PCs wrongly carrying the same one, and from a value inside the config that changes on its own.
This commit is contained in:
203
scripts/collapse_duplicate_backup_revisions.py
Normal file
203
scripts/collapse_duplicate_backup_revisions.py
Normal file
@@ -0,0 +1,203 @@
|
||||
"""Collapse the duplicate backup revisions left by the shared-machine-number bug.
|
||||
|
||||
Dedup used to compare a posted config against the latest revision for the
|
||||
ASSET. That is correct only when a machine number means one PC. Several PCs
|
||||
legitimately share one here - the part markers on 0613, 0615 and WJPRT, whose
|
||||
configs differ by COM port - so each marker's post differed from whichever
|
||||
marker had posted last, nothing ever deduped, and the table grew by one row per
|
||||
PC per collection cycle. Dedup is now per source hostname; this cleans up what
|
||||
the old rule wrote.
|
||||
|
||||
A machine number that is NOT meant to be shared but shows two sources here is a
|
||||
different fault - two PCs carrying the same machine number - and this script
|
||||
will report it as two chains. Deal with that at the PC, not in the database.
|
||||
|
||||
WHAT IT REMOVES, and nothing else: within one chain (asset + kind + source
|
||||
hostname), a revision whose content hash equals the revision immediately before
|
||||
it from the same source. Those rows record no change and are exactly the rows
|
||||
the fixed code would not have created.
|
||||
|
||||
WHAT IT KEEPS: every genuine change, in order; the newest revision of every
|
||||
chain; the oldest revision of every chain; and every chain of every source, so
|
||||
a quiet PC's only backup is never touched. A config that reverts to an earlier
|
||||
state keeps both rows, because that is a real change and not a duplicate.
|
||||
|
||||
Share-backed kinds (partmarker) are metadata rows pointing at files on the
|
||||
share. Deleting a row does not touch the file.
|
||||
|
||||
Usage (on the target instance, in the app dir):
|
||||
venv\\Scripts\\python scripts\\collapse_duplicate_backup_revisions.py
|
||||
venv\\Scripts\\python scripts\\collapse_duplicate_backup_revisions.py --commit
|
||||
... --assetnumber 0615 # limit to one machine
|
||||
... --kind ntlars # limit to one kind
|
||||
|
||||
Dry run by default: it prints what it would delete and changes nothing. Take a
|
||||
database backup before --commit anyway; this deletes rows.
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import os
|
||||
import sys
|
||||
from collections import defaultdict
|
||||
|
||||
# Run from anywhere: put the repo root (parent of scripts/) on the path.
|
||||
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
|
||||
|
||||
from shopdb import create_app
|
||||
from shopdb.api import db
|
||||
from shopdb.core.models import Asset
|
||||
from plugins.backups.models import BackupRevision
|
||||
|
||||
|
||||
def _chains(assetnumber=None, kind=None):
|
||||
"""Every revision grouped into (assetid, kind, sourcehostname) chains."""
|
||||
query = db.session.query(BackupRevision)
|
||||
if assetnumber:
|
||||
asset = db.session.query(Asset).filter(
|
||||
Asset.assetnumber == assetnumber).first()
|
||||
if asset is None:
|
||||
raise SystemExit('no asset with assetnumber {!r}'.format(assetnumber))
|
||||
query = query.filter(BackupRevision.assetid == asset.assetid)
|
||||
if kind:
|
||||
query = query.filter(BackupRevision.backupkind == kind)
|
||||
|
||||
chains = defaultdict(list)
|
||||
for revision in query.order_by(BackupRevision.backuprevisionid.asc()).all():
|
||||
key = (revision.assetid, revision.backupkind, revision.sourcehostname)
|
||||
chains[key].append(revision)
|
||||
return chains
|
||||
|
||||
|
||||
def _redundant(revisions):
|
||||
"""Revisions in one chain that repeat the hash of the one before them.
|
||||
|
||||
The chain is in id order, so this walks it forwards and keeps the first
|
||||
appearance of each run. A later run of the same hash after a genuine change
|
||||
is kept: that is a revert, which is a real event.
|
||||
"""
|
||||
doomed = []
|
||||
previoushash = None
|
||||
for revision in revisions:
|
||||
if previoushash is not None and revision.contenthash == previoushash:
|
||||
doomed.append(revision)
|
||||
else:
|
||||
previoushash = revision.contenthash
|
||||
# Never drop the newest row of a chain, even if it repeats: it is the row
|
||||
# the asset panel and a restore both read.
|
||||
newestid = revisions[-1].backuprevisionid
|
||||
return [r for r in doomed if r.backuprevisionid != newestid]
|
||||
|
||||
|
||||
def _report(chains):
|
||||
"""Say what grew each chain, so a real cause can be told from the shared
|
||||
machine numbers.
|
||||
|
||||
Three findings, and they need different fixes:
|
||||
|
||||
several sources on one machine number - either a legitimately shared
|
||||
number (0613, 0615, WJPRT) or, on any other number, two PCs carrying the
|
||||
same machine number, which is a fault to fix on the PC.
|
||||
|
||||
one source, many revisions, few distinct hashes - the old asset-wide
|
||||
dedup, or a value flapping between two states.
|
||||
|
||||
one source, many revisions, all distinct hashes - something in the
|
||||
captured config changes on its own. The value names listed are the ones
|
||||
that changed most often between consecutive revisions; a timestamp or a
|
||||
session counter in the projection would show up here, and the fix is to
|
||||
exclude it from the hash, not to delete rows.
|
||||
"""
|
||||
from plugins.backups.api.routes import _diffprojections
|
||||
|
||||
bymachine = defaultdict(list)
|
||||
for key, revisions in chains.items():
|
||||
bymachine[key[0], key[1]].append((key[2], revisions))
|
||||
|
||||
for (assetid, kind), sources in sorted(bymachine.items()):
|
||||
asset = db.session.get(Asset, assetid)
|
||||
name = asset.assetnumber if asset else str(assetid)
|
||||
total = sum(len(r) for _, r in sources)
|
||||
print('{} / {}: {} revisions from {} source(s)'.format(
|
||||
name, kind, total, len(sources)))
|
||||
|
||||
for source, revisions in sorted(sources, key=lambda s: s[0] or ''):
|
||||
hashes = {r.contenthash for r in revisions}
|
||||
print(' {:<20} {:>5} revisions, {:>4} distinct'.format(
|
||||
source or '(no source)', len(revisions), len(hashes)))
|
||||
|
||||
if len(revisions) < 3 or kind != 'ntlars':
|
||||
continue
|
||||
churn = defaultdict(int)
|
||||
for older, newer in zip(revisions, revisions[1:]):
|
||||
if older.contenthash == newer.contenthash:
|
||||
continue
|
||||
for change in _diffprojections(older.payload, newer.payload):
|
||||
churn['{}\\{}'.format(change['keypath'],
|
||||
change['valuename'])] += 1
|
||||
for valuename, count in sorted(churn.items(),
|
||||
key=lambda kv: -kv[1])[:5]:
|
||||
print(' changed {:>4}x {}'.format(count, valuename))
|
||||
print()
|
||||
|
||||
|
||||
def main():
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument('--commit', action='store_true',
|
||||
help='apply the deletions (default is a dry run)')
|
||||
parser.add_argument('--report', action='store_true',
|
||||
help='explain what grew each chain and delete nothing')
|
||||
parser.add_argument('--assetnumber', help='limit to one machine number')
|
||||
parser.add_argument('--kind', help='limit to one backup kind')
|
||||
args = parser.parse_args()
|
||||
|
||||
app = create_app(os.environ.get('FLASK_ENV', 'production'))
|
||||
with app.app_context():
|
||||
chains = _chains(args.assetnumber, args.kind)
|
||||
if not chains:
|
||||
print('No revisions matched.')
|
||||
return
|
||||
|
||||
if args.report:
|
||||
_report(chains)
|
||||
return
|
||||
|
||||
assetnames = {}
|
||||
total = 0
|
||||
machines = defaultdict(int)
|
||||
|
||||
for (assetid, kind, source), revisions in sorted(chains.items(),
|
||||
key=lambda kv: kv[0]):
|
||||
doomed = _redundant(revisions)
|
||||
if not doomed:
|
||||
continue
|
||||
if assetid not in assetnames:
|
||||
asset = db.session.get(Asset, assetid)
|
||||
assetnames[assetid] = (asset.assetnumber if asset
|
||||
else str(assetid))
|
||||
name = assetnames[assetid]
|
||||
machines[name] += len(doomed)
|
||||
total += len(doomed)
|
||||
print('{:>10} {:<12} {:<20} {} of {} redundant'.format(
|
||||
name, kind, source or '(no source)', len(doomed),
|
||||
len(revisions)))
|
||||
for revision in doomed:
|
||||
db.session.delete(revision)
|
||||
|
||||
print()
|
||||
for name in sorted(machines):
|
||||
print('{:>10}: {} rows'.format(name, machines[name]))
|
||||
print('{} chains, {} redundant revisions'.format(len(chains), total))
|
||||
|
||||
if not total:
|
||||
print('Nothing to do.')
|
||||
return
|
||||
if args.commit:
|
||||
db.session.commit()
|
||||
print('DELETED {} revisions.'.format(total))
|
||||
else:
|
||||
db.session.rollback()
|
||||
print('Dry run - nothing changed. Re-run with --commit to apply.')
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
Reference in New Issue
Block a user