Files
shopdb-flask/docs/proposals/ge-enforce-plugin.md
cproudlock 4995456136 docs: take one site's name, hosts and paths off the public wiki
The publishability gate caught internal tooling names and developer paths but
nothing site-specific, so roughly sixty leaks reached the wiki: the site name in
ten documents, real fleet hostnames in the collector and GE-Enforce examples, an
internal database name through the whole import guide, imaging-share paths, and
a maintainer's username as the Deciders line of every ADR and inside a generated
curl example.

None of it is a security matter on an air-gapped fleet. It matters because these
pages are read by engineers at other plants, and a document that names one site
throughout reads as that site's notes rather than a product's documentation -
which is exactly what it then gets treated as.

Examples now use neutral hostnames, the site is "the reference site" where the
distinction carries meaning, and ADRs are decided by "ShopDB maintainers". The
gate carries all of these patterns, so the next one fails a build.

Two documents leave docs/ because they were never written for an outside reader.
PROJECT-REVIEW.md is an internal health memo pinned to a commit from July, whose
headline finding (an untracked playbook) has since been fixed - it is history,
and git holds it. PILOT-DEPLOY.md is one site's own cutover runbook, complete
with a "re-measure before publishing" placeholder; it moves next to the loader
it belongs to, in scripts/site_imports/wjf/.

ADR-015 is AMENDED rather than rewritten. Its enforcement section still said
report-only and its backlog still listed hardcodes that are now cleared, which
left the record contradicting itself. The amendment says what changed and why
the report-only period ended; the original text stays, because what the decision
looked like when it was taken is the part worth keeping.

Also corrects llms.txt's response envelope, which had errors at the top level
and pagination at meta.total. Both are nested one deeper, so anything written
against that description read undefined on every error it tried to handle.
2026-08-14 15:38:27 -04:00

649 lines
38 KiB
Markdown

# Proposal: GE-Enforce as a shopdb plugin
Status: ACCEPTED / built - see ADR-012 and plugins/geenforce/.
Author: planning session 2026-07-12.
## 1. What this is
Today GE-Enforce is a PowerShell manifest engine that reads per-PC-type
`manifest.json` files off an SMB share (`\\shopdb.example.net\
shared\dt\shopfloor\`). Each logon, a scheduled task running as SYSTEM mounts
the share, reads the manifest for the machine's PC type, and installs or
self-heals apps, files, drivers, registry values, and scripts. A parallel
`preinstall.json` runs the same schema once at imaging.
This proposal turns the *manifest* into shopdb data: the authoritative manifest
lives in the shopdb database, is edited through the shopdb UI (an expansion of
`/settings/pctypemapping`), and is served to clients over HTTP as JSON. The
*payloads* (MSI/EXE/PS1/config bytes) stay on SMB, on HTTP, or both, referenced
by URL/path from the manifest rows. GE-Enforce.ps1 changes from "read a file on
W:" to "GET a manifest from shopdb, then fetch each payload from wherever the
row says."
The result: managing imaging PC types, their apps, scripts, files, registry
rules, and version gates becomes a first-class shopdb feature instead of hand-
edited JSON on a file share.
## 2. Why it fits shopdb
- shopdb already models the fleet (the collector ingests every PC's hostname,
pctype, installed software, versions). Making shopdb *also* own what SHOULD be
installed closes the loop: desired-state (manifest) and observed-state
(collector) live in one system and can be diffed.
- `/settings/pctypemapping` already maps `gea-shopfloor-*` PC types to
`ComputerType`. That page becomes the entry point for full imaging-PC-type
management.
- The plugin contract (per-plugin models, migrations, API prefix, settings
cards, collector hooks) is exactly the shape this needs.
- ADR-004 (per-site instances) matches: each site's shopdb owns each site's
manifest. No multi-tenant complication.
## 3. Grounding: the real manifest schema
Source of truth for these field names (do not invent others):
- Schema: `<imaging-share>/SHOPDBHOST-v2/shared/dt/shopfloor/_meta/manifest-schema.json`
- Engine: `<imaging-share>/common/lib/Install-FromManifest.ps1`
- Dispatcher: `.../shopfloor/common/GE-Enforce.ps1`
- Architecture: `pxe/docs/ge-enforce-v2-architecture.md`
A manifest is `{ "Version": str, "_comment": str, "Applications": [entry, ...] }`.
Only `Name` and `Type` are required per entry.
### Per-entry fields (complete set)
Identity / action:
- `Name` (required, unique, also the status-key `<scope>/<Name>`)
- `Type` (required): one of `MSI EXE CMD BAT PS1 INF File Registry`
- `_comment` (documentation, heavily used in practice)
Type-specific payload references (sparse; depends on Type):
- MSI/EXE/CMD/BAT/INF: `Installer` (relative path) + `InstallArgs`
- PS1: `Script` (relative path, falls back to `Installer`) + `Args`
- File: `Source` (relative) + `Destination` (absolute on-PC path)
- Registry: `RegPath` + `RegName` + `RegValue` + `RegType`
(`RegType` in `String DWord QWord MultiString ExpandString Binary`)
- Optional `LogFile`, `WaitTimeoutSec` (EXE hang kill), `InUseCheck`
Detection (decides whether the action fires / self-heals):
- `DetectionMethod`: one of
`Registry File FileVersion Hash MarkerFile ValueMatches pnputil Always`
- `DetectionPath`, `DetectionName`, `DetectionValue`, `DetectionPattern`
- Note: `DetectionValue` is method-dependent - SHA256 for Hash, a 4-part
version for FileVersion, a registry value for Registry, ignored for
Always/File. Same column, different meaning per method.
- No `DetectionMethod` = always installs.
Targeting filters (all ANDed; each is multi-value):
- `PCTypes` (array; `"*"` = all; alias graph expands old<->new names)
- `PCSubTypes` / subtype via `<pctype>-<subtype>` values
- `TargetHostnames` (array; exact + `-like WJS-*` wildcards)
- `TargetMachineNumbers` (array; per-bay)
- `_CmmVersion` (scalar; per-entry PC-DMIS version gate, needs lib >= 2.6)
Nested:
- `InUseCheck`: `{ Behavior, Processes: [{Name, ExePath, GracefulCloseTimeoutSec}] }`
Behavior in `Defer CloseAndReopen ForceClose ScheduleForReboot`
Parsed-but-inert today (model them, mark inert):
- `ApplyMode` (`Nightly Immediate ImmediateReboot`), `UpdateWindow` (`HH:MM-HH:MM`)
Preinstall-only extras (phase discriminator):
- `PreEnrollment`, `KillAfterDetection`, `PCTypesStrict`, `_pcTypesNote`
### Load-bearing behaviors the model must preserve
1. **Array order IS execution order.** Config-restore entries are deliberately
placed AFTER their vendor installer so a mid-cycle overwrite heals the same
cycle (eMxInfo.txt after eDNC; udc_webserver_settings after UDC). We MUST
store an explicit per-scope `sortorder`, not a set.
2. **PCTypes alias graph** is many-to-many old<->new names resolved by set
intersection, with a `PCTypesStrict` escape hatch. Not a simple FK.
3. **Polymorphic entry by Type** - sparse column set per type. DECISION: one
wide `manifestentries` table with an `entrytype` discriminator column and
nullable per-type columns. NOT SQLAlchemy STI subclasses, NOT a JSON blob.
Justification (section 4): the whole fleet is ~64 entries, so sparse columns
cost nothing; real columns get validated, indexed, field-diffed, joined
against collector data, and read in plain SQL by an IT tech - a JSON blob
hides all of that, and class-per-type STI is expert ceremony for no gain. A
`validate()` that switches on `entrytype` (mirroring the engine's own
`switch ($App.Type)`) is ~40 obvious lines.
4. **Two manifest phases** - runtime (self-heal, per logon) and preinstall
(once at imaging) share the schema. One table with a `phase` discriminator.
## 4. Data model (new `geenforce` plugin)
Per-plugin Alembic chain (ADR-008). Tables (lowercase concatenated per naming
convention). Sizing that shapes every decision here: the real fleet is 10
runtime scopes = 43 entries, plus 1 preinstall manifest = 21 entries, so ~64
rows total. That smallness is why this stays deliberately low-tech (one wide
table, JSON-document snapshots, no row-mirroring) - the design target is an
average site IT tech maintaining it, not a specialist.
- `manifestscopes` - one row per imaging PC type / scope.
- `scopeid` PK
- `scopename` (e.g. `gea-shopfloor-cmm`)
- `phase` enum (`runtime` | `preinstall`)
- UNIQUE (`scopename`, `phase`), NOT `scopename` alone: `common` exists in
runtime, and a scope name can appear in both phases. Note the phases are
shaped differently - runtime is many per-pctype scopes (one manifest file
each), preinstall is ONE flat manifest gated internally by `PCTypes`, so
preinstall is modeled as a single `phase=preinstall` scope, not per-pctype
scopes.
- `computertypeid` FK -> `computertypes` (this REPLACES the thin
`pctypemap_<pxetype>` setting; the mapping becomes a column here).
Runtime-scope only; null for the preinstall scope.
- `measuringtooltypeid` FK -> `measuringtooltypes`, nullable (metrology
scopes: what device this scope implies; keeps imaging + collector agreed,
see section 11).
- `manifestversion` (string, mirrors manifest `Version`)
- `description`, `isactive`
- `iscommon` bool (the `common/` fleet-wide scope)
- `manifestentries` - one row per Applications[] entry (the working/draft copy).
- `entryid` PK, `scopeid` FK
- `sortorder` int (preserves array order; the ordering contract)
- `name`, `entrytype` (MSI/EXE/.../Registry), `comment`
- payload columns (nullable, per type): `installer`, `installargs`,
`scriptpath`, `scriptargs`, `sourcepath`, `destination`,
`regpath`, `regname`, `regvalue`, `regtype`
- `payloadsource` enum (`smb` | `http` | `inline`) + `payloadref`
(see section 5)
- `payloadsha256` - integrity hash of the payload bytes, INDEPENDENT of the
detection method. Mandatory for `http`/`inline` payloads; optional for
`smb`. Do NOT reuse `detectionvalue` for this - `detectionvalue` is a
SHA256 only when `detectionmethod = Hash`; an MSI with `Registry`/
`FileVersion` detection has no payload hash, so an HTTP fetch would
otherwise run unverified bytes (see section 5).
- `regvalue` stores the RAW JSON literal (`1` vs `"1"`) and is emitted
verbatim on export. `RegValue` is untyped in the manifest schema and real
entries carry numbers; the engine string-coerces for `ValueMatches` but
`Set-ItemProperty -Type DWord` cares, so preserve the literal.
- detection columns: `detectionmethod`, `detectionpath`, `detectionname`,
`detectionvalue`, `detectionpattern`
- gates: `cmmversion`, plus child tables for the multi-value filters
- control: `logfile`, `waittimeoutsec`, `applymode`, `updatewindow`
(`applymode`/`updatewindow` are parsed-but-INERT in the engine today; the
UI must label them "not yet enforced" so a tech does not trust a dead gate)
- preinstall flags: `preenrollment`, `killafterdetection`, `pctypesstrict`
- `isactive`
- `manifestpublishedversions` - immutable published snapshots, SIMPLIFIED to
freeze the rendered JSON DOCUMENT in a single `manifestjson` column (drop the
row-mirrored `manifestpublishedentries` family the earlier draft proposed).
The only consumer of a snapshot is the client, and it consumes exactly that
document, so freezing the text makes immutability structural (no UPDATE path),
rollback a one-flag `iscurrent` flip, serving a single-row read, and version
diffing a plain text diff - all things average IT can debug; row-mirroring
would add ~6 shadow tables and a copy routine that can drift. Columns:
`publishedversionid`, `scopeid`, `versionnumber` (1,2,3 per scope),
`manifestjson` (MEDIUMTEXT, verbatim), `publishedat`, `publishedby`,
`iscurrent`, `notes`. Editing `manifestentries` never affects the fleet;
"publish" freezes a new snapshot; the client is ALWAYS served the current
snapshot, never the live draft. Rollback = flip `iscurrent` to an older
version (the post-cutover safety net once the on-share JSON is retired).
Mirrors today's `_meta/history/<date>-<scope>.json` backups, but authoritative.
Revision history: every publish is a permanent, immutable revision kept
indefinitely (snapshots are small JSON text, ~10 scopes - storage is a
non-issue). An OPTIONAL retention policy (keep last M per scope, or prune
older than N months) can be added later; default is keep-everything, off.
- Draft-edit audit trail (field-level history BETWEEN publishes): drafts
(`manifestentries`) are not versioned - editing overwrites the working copy.
To answer "who changed this entry and when" in the window between two
published revisions, log every draft mutation through the EXISTING core audit
system (no new table): on create/update/delete of a scope, entry, or child
row, write an audit record with the actor, timestamp, entry name, and the
changed field(s). This gives per-edit provenance for free and shows up in the
same Audit Logs UI IT already uses; the published snapshots remain the
coarse-grained "what the fleet actually got" record.
- `manifestentrypctypes`, `manifestentryhostnames`, `manifestentrymachinenumbers`
- child rows for the ANDed multi-value filters (one value + a `sortorder` per
row, wildcards stored verbatim as patterns)
- `manifestinusechecks` + `manifestinusecheckprocesses`
- the nested InUseCheck object and its Processes[] child list (leave
`gracefulclosetimeoutsec` nullable; do not bake the engine's default of 10
into the row, emit it only when set)
- `manifestpayloads` - inline payload bytes for `payloadsource = inline`
(`entryid`, `filename`, `contenttype`, `payloadbytes` LONGBLOB, `payloadsha256`,
`uploadedat`). App-enforced size cap ~1 MB; the upload UI rejects larger with
"use SMB for this" so nobody pastes an MSI into the database. Can ship empty
and unused until P6.
- `pctypealiases` - a MIRROR of the old<->new name alias graph from
`Install-FromManifest.ps1:463-475`, for server-side resolve/validate only.
The engine lib stays the single source of truth (see section 10); shopdb
never becomes the authority the client depends on for aliases.
The JSON the client receives is REBUILT from a published snapshot in exact
array order. Parity with the current engine is proven by BEHAVIORAL equivalence,
not byte-identity (see section 9): re-serialized JSON will differ in key order
and whitespace, so the test is that both manifests parse to the same ordered
entry set with the same detection/targeting/action semantics.
## 5. Payloads: SMB and/or HTTP (both supported)
The user asked whether payloads can be SMB and/or HTTP. Yes - per entry:
- `payloadsource = smb`: `payloadref` is the current relative path
(`apps/eDNC_6-4-5.msi`); the client still mounts W: and resolves it against
the scope root exactly as today. The engine is unchanged for these rows (the
mount + scope-root resolution still happen; an HTTP-only site skips the mount
because it has no `smb` rows). This is the default and the migration target
for large binaries (MSIs are hundreds of MB; SMB streaming beats HTTP).
- `payloadsource = http`: `payloadref` is a URL (absolute, or relative to a
configured payload base). The client downloads to a local temp dir, verifies
the Hash/FileVersion detection value, then runs it. Good for small
config/script payloads and for sites with no SMB share.
- `payloadsource = inline`: for small text payloads (a `.ps1`, a config file, a
registry value), the bytes live in shopdb itself and are served in-band. No
external store at all. Best for scripts and File-type config drops.
Manifest generation emits, per entry, whatever the client needs to fetch the
bytes. The engine's existing "stage network EXE to local temp first" logic
(SYSTEM access-denied workaround) generalizes cleanly to HTTP download.
Payload integrity uses the dedicated `payloadsha256` column, NOT `DetectionValue`.
This is the correction to a subtle trap: `DetectionValue` is a SHA256 only when
`DetectionMethod = Hash`. Most binaries detect by `Registry` or `FileVersion`
and carry no payload hash at all, so relying on `DetectionValue` would let an
HTTP/inline-fetched MSI run unverified. Instead, publishing an `http`/`inline`
payload computes and stores `payloadsha256`, and the client verifies the fetched
bytes against it BEFORE running, independent of how the entry detects install
state. `smb` payloads may set it too (defense in depth) but the share ACL is
their primary trust boundary. Detection stays a separate concern: it decides
whether to act; the payload hash decides whether the bytes are trustworthy.
Transport security: the client fetches as SYSTEM, so the shopdb TLS cert must be
trusted machine-wide. Sites with a self-signed or air-gapped shopdb need the CA
in the machine trust store (provisioned by the same Azure DSC step that writes
the token). Plain HTTP is acceptable only inside a trusted segment, and even
then the `payloadsha256` check is what actually guarantees payload integrity.
## 6. API surface (`/api/geenforce/...`)
Two permissions via the plugin's `get_permissions()` hook (split so day-to-day
techs can edit but only a lead ships to the fleet):
- `geenforce.manage` - create/edit/reorder scopes, entries, drafts, payloads.
- `geenforce.publish` - publish, rollback, export-to-share (the fleet-affecting
actions).
Draft editing (`geenforce.manage`):
- `GET/POST /scopes`, `GET/PUT/DELETE /scopes/<id>` - imaging PC types
- `GET/POST /scopes/<id>/entries`, `PUT/DELETE /entries/<id>` - manifest entries
- `PUT /scopes/<id>/entries/reorder` - the ordering contract; Move Up/Down in the
UI (plain buttons + visible `sortorder`), not a drag-and-drop dependency
- `POST /entries/<id>/payload` - upload an inline/http payload (multipart),
compute + store its `payloadsha256` (the integrity hash; NOT `detectionvalue`)
- `GET /scopes/<id>/preview` - the draft JSON a client WOULD receive on next
publish; `GET /scopes/<id>/published` shows the currently-served snapshot
- `GET /scopes/<id>/simulate?pctype=&subtype=&hostname=&machinenumber=&cmmversion=`
- the "what would this PC get" simulator: runs the entry list through the same
filter logic the engine uses and returns which entries apply and why the rest
are filtered out. Reuses the P1 parity harness's filter engine, so it is
nearly free, and it is the single most IT-empowering endpoint - it answers
"why did/didn't app X install on PC Y" without reading a PowerShell log.
Publishing (`geenforce.publish`):
- `POST /scopes/<id>/publish` - freeze the current draft into a new immutable
`manifestpublishedversions` snapshot (this is what the fleet gets)
- `POST /scopes/<id>/rollback/<version>` - mark an older snapshot current
- `POST /scopes/<id>/export-share` (or a `flask geenforce export-share` CLI) -
write the current published JSON to `<shareroot>/<scope>/manifest.json` after
copying the existing file to `_meta/history/<date>-<scope>.json`. This is a
first-class feature, not a footnote: it is the Milestone 1 product (author in
shopdb, engine untouched) and the permanent break-glass path.
Client-facing (gated by a collector-style service token, `geenforce.fetch`
scope, reusing the PAT + `X-API-Key` machinery already built for the collector):
- `GET /manifest?pctype=<scope>&subtype=<s>&hostname=<h>&machinenumber=<n>`
Returns the latest PUBLISHED snapshot for that scope (never the live draft).
The server can pre-apply the PCTypes/hostname/machinenumber/cmmversion filters
(thin client) OR return the full scope and let the engine filter (fat client,
matches today). Start fat: return the scope manifest unchanged so the engine
logic is untouched. Include the snapshot version + an ETag so the client can
cache and no-op when unchanged.
- Payload fetch for `http`/`inline` rows: `GET /payload/<entryid>` streaming the
bytes; the client verifies them against `payloadsha256` from the manifest.
## 7. Frontend: expand `/settings/pctypemapping`
The current page (`PCTypeMappingSettings.vue`, "Collector PC Types") is a read-
only-ish table of `pxetype -> ComputerType` dropdowns. It grows into the imaging-
PC-type manager:
- **Scopes list**: add/rename/delete imaging PC types; each still carries its
`ComputerType` mapping (that column moves from a setting into `manifestscopes`).
A `phase` toggle (runtime vs preinstall). Common scope flagged.
- **Scope detail / manifest editor**: an ordered list of entries with Move
Up/Down buttons and a visible `sortorder` (the ordering contract made visible;
NOT drag-and-drop - a drag library is the kind of dependency that breaks
silently and average IT cannot fix; add drag later if wanted). Each entry is a
typed form - the visible fields switch on `entrytype` (MSI shows
Installer+InstallArgs; PS1 shows Script+Args; File shows Source+Destination;
Registry shows the Reg* quartet), one line of help per detection method.
Filter chips for PCTypes/hostnames/machine numbers. InUseCheck sub-editor.
Payload source selector (smb/http/inline) with upload for the latter two.
`applymode`/`updatewindow` sit behind an "Advanced (not yet enforced by the
engine)" disclosure. Ship the editor in three usable-alone increments: (a)
scope list + entry table, (b) the typed entry form, (c) publish + diff. That
keeps the biggest chunk of the build from ballooning.
- **Simulator ("what would this PC get")**: a small form (pctype, subtype,
hostname, machine number, CMM version) that calls `GET /scopes/<id>/simulate`
and lists which entries apply and why the rest are filtered. The single most
IT-empowering piece of the UI.
- **Draft, preview, publish**: editing changes only the draft; "publish" freezes
an immutable snapshot (see section 4) and is what the fleet then gets. Show the
draft-vs-published diff before publishing. Rollback republishes a prior
snapshot.
- **Desired vs observed (BUILT: observed-state reporting)**: rather than extend
the collector, the plugin has its own reporting path. Each enforcement cycle a
PC POSTs `POST /api/geenforce/report` (geenforce.report service token) with the
published version it applied, the installed/skipped/failed/filtered counts, and
per-entry outcomes. Stored in `manifestenforcementreports` (latest-per-host +
history) and `manifestenforcementresults` (per-entry). Two payoffs fall out:
RECEIVED - `receivedlatest` compares the applied version to the scope's current
published version, so the fleet view shows which PCs picked up an update; and
SELF-HEAL - each entry's action (installed = drift corrected, skipped = already
good, failed) with any warning/error message. Admin reads: `GET /reports`
(fleet compliance) and `GET /reports/<id>` (per-entry detail). This is the
observed half that makes the manifest a closed desired-vs-observed loop.
This is an ADR-010 settings card contributed by the geenforce plugin, so it only
appears when the plugin is enabled.
## 8. Client change (minimal, staged)
`GE-Enforce.ps1` today: mount W:, read `<scope>\manifest.json`, hand to
`Install-FromManifest`. New path: GET the manifest from shopdb, write it to the
same local location the engine reads, then run the engine unchanged. That is the
smallest possible client delta - the engine, detection logic, self-heal, and
SMB payload resolution all stay identical. Only the *source of the JSON* moves
from file to HTTP.
Payloads: `smb` rows need no client change. `http`/`inline` rows need a small
fetch-and-verify helper (download to temp, check SHA256, then the existing
installer action runs against the local copy). The engine already stages network
EXEs to temp, so this is an extension, not a rewrite.
Auth: the client already has SFLD credentials in
`HKLM:\SOFTWARE\GE\SFLD\Credentials`. Add a shopdb service token (a
`geenforce.fetch` PAT) provisioned the same way (Azure DSC writes it to
registry), sent as `X-API-Key`. If shopdb is unreachable, the client falls back
to the last-known-good manifest cached locally (fail-safe: never leave a PC
unmanaged because the web app is down). This mirrors today's "creds missing =
exit 0, retry next cycle" resilience.
## 9. Cutover strategy
The manifest is desired-state that runs as SYSTEM and installs software fleet-
wide. A bad cutover = a fleet-wide mis-install. Stage it:
1. **Import + parity.** Write a one-shot importer that reads the current
on-share manifests (common + every `gea-shopfloor-*` + preinstall.json;
skip `.bak` / `.pre-mtconnect.bak` variants) into the new tables. Then
generate JSON back out and prove BEHAVIORAL equivalence for every scope - do
NOT chase byte-identity. Re-serialized JSON will differ in key order,
whitespace, and `_comment` formatting, so a raw `diff` would never converge.
The correct test: parse both the original and the regenerated manifest,
normalize, and assert the same ordered entry list with identical
detection/targeting/action fields per entry (ideally a small harness that
mimics the engine's filter+detect decisions and confirms the same entries
would fire in the same order on representative machine profiles). That, not
byte equality, is what proves the model is lossless. (Same discipline as the
ADR-001 data migration.)
2. **Shadow mode.** shopdb serves the manifest at a new endpoint; a canary PC
fetches from shopdb but ALSO reads the share, and logs any diff. No install
behavior changes. Run across one of each PC type for a few cycles.
3. **Read cutover, payloads still SMB.** Flip GE-Enforce to source the JSON from
shopdb (payloads stay `smb`). The blast radius is only "where the JSON comes
from"; the bytes and engine are unchanged. Keep the share manifests as the
rollback (revert the dispatcher one-liner).
4. **Payload migration (optional, per entry).** Move small scripts/configs to
`inline`/`http` opportunistically. Leave big MSIs on SMB indefinitely - SMB
is the right transport for them.
5. **Author in shopdb.** Once read-cutover is stable, new manifest edits happen
in the shopdb UI and the on-share JSON is retired (or auto-exported as a
backup for break-glass).
Rollback during cutover (stages 2-4) is a one-line dispatcher revert, because
the engine and payload layout never stop working from the share. AFTER the share
JSON is retired (stage 5), that escape hatch is gone - post-cutover rollback is
republishing a prior `manifestpublishedversions` snapshot (section 4). Both
mechanisms must exist before stage 5, not just the dispatcher revert.
## 10. Risks / open questions
- **The engine is the contract.** Any drift between shopdb's generated JSON and
what `Install-FromManifest.ps1` expects is a fleet-wide install bug. The
byte-identical round-trip test (step 1) is non-negotiable, and the plugin must
pin which engine lib version it targets (>= 2.6 for `_CmmVersion`).
- **PCTypes alias graph** must be kept in sync with
`Install-FromManifest.ps1:463-475`. The engine lib stays the single source of
truth; shopdb only MIRRORS the map for server-side validation. Do NOT invert
this to have the engine fetch aliases from shopdb - that would add exactly the
availability coupling the next bullet warns against. When the lib's alias map
changes, update shopdb's mirror as part of shipping that lib version.
- **Availability coupling.** GE-Enforce currently depends only on SMB. Adding an
HTTP dependency on shopdb means shopdb downtime could stall enforcement -
hence the last-known-good local cache in section 8. Must be built in from day
one, not bolted on. This is also why alias resolution and payloads stay
independent of a live shopdb wherever possible.
- **Transport trust.** The client runs as SYSTEM, so shopdb's TLS cert must be
in the machine trust store (self-signed/air-gapped sites need the CA
provisioned via the same DSC step as the token). `payloadsha256` verification
is the real integrity guarantee and holds even over plain HTTP inside a
trusted segment (section 5).
- **Secrets in payloads.** Some config drops (site-config, credentials) may
contain secrets. `inline` payloads live in the shopdb DB - those must respect
the existing "secrets stay in .env, not the settings table" rule. Likely keep
any secret-bearing payload on SMB with ACLs, never inline.
- **Preinstall runner** is a separate consumer (`00-PreInstall-*` at imaging,
before enrollment). It may not have a shopdb token yet at that point in the
imaging sequence. Preinstall may need to stay share-sourced longer than
runtime, or fetch a bootstrap manifest anonymously over HTTP.
- **This is a big build.** Realistically phased: (P1) model + importer +
behavioral-parity test; (P2) admin API + CRUD + publish/snapshot/rollback;
(P3) frontend editor on /settings/pctypemapping; (P4) client fetch + shadow
mode; (P5) read cutover; (P6) payload migration. P1 is the gating de-risk - if
behavioral parity does not hold, stop. Snapshots (P2) must land before any
client points at shopdb (P4), since serving the live draft to the fleet is
unacceptable.
## 11. Relationship to existing work
- Replaces `plugins/computers/pctypemap.py` (the thin `pctypemap_<pxetype>`
settings) - the pctype -> ComputerType mapping becomes the `computertypeid`
column on `manifestscopes`. Two-source transition window: `pctype_mapping()`
must keep reading the settings until the geenforce plugin is enabled, then
fall back geenforce-table-first / settings-second, and only retire
`seed_pctype_settings` + the settings at Milestone 1 close. Also reconcile the
scope inventory: `pctypemap.py` lists `gea-shopfloor-display` but the share has
no such manifest dir, and the share has a `main/` legacy dir the model ignores
- the importer creates scopes only from what it finds (plus empty scopes for
mapped-but-absent pctypes), and the P1 gate review reconciles the list with
the floor team.
- Also folds in the metrology mapping now living in `pctypemap.py`
(`METROLOGY_TOOL_MAP`). The collector already auto-creates a MeasuringTool
asset and a directional PC->tool `controls` relationship when it sees a
metrology pctype (CMM / Keyence / Genspect / wax-and-trace); the PC stays a
shopfloor PC. A metrology scope in the manifest model should carry the
attached-measuring-tool type alongside its ComputerType so imaging and
collector agree on what device the scope implies.
- Reuses the collector's token machinery (PAT + `X-API-Key` + scopes) for the
client-facing endpoints.
- Reuses `get_permissions()` (contract 0.10.0) for `geenforce.manage` (edit
drafts) / `geenforce.publish` (publish, rollback, export) / `geenforce.fetch`
(the client service token).
- Pairs with the collector: desired-state (this plugin) + observed-state
(collector) enable a fleet compliance view.
## 12. Recommendation
Feasible and a strong architectural fit, but it is a multi-phase build with a
fleet-wide blast radius. The single most important gate is P1: import the real
manifests and prove BEHAVIORAL parity (same entries fire in the same order with
the same detection/targeting), not byte-identity. Do not build the UI or touch a
client until that parity holds. Three things separate a safe build from a
dangerous one and must not be cut: behavioral-parity import (P1), immutable
published snapshots with rollback before any client points at shopdb (P2/P4),
and a dedicated `payloadsha256` for every HTTP/inline payload (section 5). If and
when we proceed, this warrants a new ADR (ADR-012: GE-Enforce manifest
ownership) capturing the desired-state model, the published-snapshot contract,
the SMB/HTTP/inline payload + integrity model, and the fail-safe cache.
## 13. Execution plan (build order, gates, milestones)
Governing constraint: every step must be runnable and maintainable by average
site IT, not just the original developer. Where an earlier draft implied expert
machinery, this section simplifies it (and the model above already reflects
those simplifications: one wide table, JSON-document snapshots, no row-mirroring).
### Phases and gates
- **P0 - Scaffold (S, ~0.5-1 day).** `flask plugin new geenforce`, structure
copied from `plugins/measuringtools/`. Unlike bundled plugins' no-op migration
anchors, this NEW plugin's `0001_geenforce_baseline` actually creates the
tables and registers them in `PLUGIN_TABLE_OWNERS` (ADR-008). Deploy stays the
standard `flask db upgrade` + `flask plugin upgrade-all`. Manifest:
`api_prefix: /api/geenforce`, `default_enabled: false`, tight `core_version`.
- **P1 - Model + importer + parity harness (M, ~1-1.5 wk). THE GATE.** Order
inside: tables -> `flask geenforce import-share` (reads common + every
`gea-shopfloor-*` + preinstall.json, skips `.bak`, idempotent) -> exporter
(rebuilds each scope's JSON from rows in `sortorder`) -> the parity harness
(below). **GATE A:** `flask geenforce parity` prints PASS for all scopes. If
it cannot pass, STOP the project. No API/UI/client work before Gate A.
- **P2 - Publish/snapshot/rollback + admin API + export-to-share (M, ~1.5-2 wk).**
Publish freezes rendered JSON into `manifestpublishedversions`. CRUD per
section 6. Plus `flask geenforce export-share` + an "Export to share" button
that writes each scope's published JSON to the share after backing up the old
file to `_meta/history/`. Engine, dispatcher, share layout, payloads, PCs all
untouched. **GATE B = Milestone 1** (below).
- **P3 - Frontend editor (L, ~2-3 wk; parallel with P4 after P2 API freezes).**
Expand `PCTypeMappingSettings.vue` per section 7, in three shippable
increments; Move Up/Down not drag; the simulator.
- **P4 - Client fetch + shadow mode (M effort + soak time; needs P2, not P3).**
Week-1 spike: a ~20-line PS1 on ONE canary PC proves SYSTEM-context HTTP auth +
TLS trust before any real client change. Then `GE-Enforce.ps1` fetches JSON to
a local cache and hands the file to `Install-FromManifest.ps1` unchanged;
shadow mode installs from the share but logs any diff vs shopdb; ETag +
last-known-good cache from day one. **GATE C:** zero shadow diffs across one PC
of every pctype for >= 20 cycles.
- **P5 - Read cutover (S effort, M calendar).** Per-scope flip, canary first via
`TargetHostnames`. Payloads stay `smb`. Rollback = dispatcher revert; share
export continues as break-glass. **GATE D:** all scopes cut over.
- **P6 - Payload migration (S per entry, optional forever).** Small configs to
`inline` (verified by `payloadsha256`); MSIs stay on SMB. Each entry
independently revertible (flip `payloadsource`).
Hard ordering: P0 -> P1 -> P2 -> rest. **Snapshots (P2) MUST precede any client
pointing at shopdb (P4).** P3 and P4 parallelize. Preinstall stays share-sourced
through at least Milestone 1 (no token pre-enrollment; export writes
`preinstall.json` too, so it is authored-in-shopdb for free with no client risk).
### The P1 parity harness (concrete, IT-re-runnable)
`plugins/geenforce/parity.py` + a CLI, also wrapped as a CI test. Two checks per
scope, output one readable line per scope (`entries N/N identical profiles M/M
same-fire PASS`), exit 0/1, prints the first differing entry/field on fail:
1. **Lossless field check (order-preserving).** Canonicalize each entry to
exactly the fields the engine reads (Name, Type, the payload fields, all
Detection*, the filter arrays, `_CmmVersion`, InUseCheck, preinstall flags);
exclude `_comment` and key order (documentation, not behavior). Compare the
ordered lists position by position.
2. **Same-entries-fire-in-same-order.** Re-implement in ~120 lines of Python the
engine's four filter functions exactly as written in `Install-FromManifest.ps1`
(`Test-PCTypeMatches` incl. the alias groups at lines 463-475, `"*"`, and
`<Type>-<SubType>`; `Test-HostnameMatches` exact + `-like`;
`Test-MachineNumberMatches`; `Test-CmmVersionMatches`). For each machine-
profile fixture, run BOTH manifests through it and assert the identical
ordered list of entry names that pass all filters. Detection itself is not
executed - check 1 already proved detection fields identical, so identical
inputs to detection are guaranteed. This pair proves losslessness without
byte-diffing.
Fixtures (`plugins/geenforce/parityfixtures.json`, ~16-18 profiles): one per
pctype; CMM version variants `2016/2019/2026`/empty; collections machine-number
variants (a credentialed bay, an MTConnect bay, neither); legacy-alias profiles
(`Standard`+`Machine`, `CMM`) to exercise the alias graph both ways; a `WJS-*`
hostname-wildcard profile; preinstall profiles including one that hits
`PCTypesStrict`. Watch-items the harness must handle: empty `Applications: []`
scopes (4 exist), entries with NO `DetectionMethod` (fire every run), and the
`regvalue` literal typing.
### First slice: one vertical through `gea-shopfloor-cmm`
Only 4 entries but hits every hard part - MSI type, Registry detection with and
without a pinned value, nested InUseCheck with Processes[], and the `_CmmVersion`
gate. Tables: scopes, entries, entrypctypes, inusechecks + processes,
publishedversions, pctypealiases. `flask geenforce import-share --scope
gea-shopfloor-cmm`; `flask geenforce publish gea-shopfloor-cmm`; one endpoint
`GET /api/geenforce/manifest?pctype=gea-shopfloor-cmm` serving the published
snapshot (fat-client, ETag, collector-style `X-API-Key`/PAT auth reusing
`shopdb/core/api/collector.py`). **Done =** parity PASS for cmm; the endpoint's
JSON fed to `Install-FromManifest.ps1` on a bench CMM PC logs `4 skipped`
identically to the share manifest; editing a draft does NOT change the served
bytes but publishing does, and rollback restores the prior published bytes;
unauth = 401, wrong-scope = 401.
### Milestone 1 (the recommended first stop)
End of P2 plus the publish/scope-list slice of P3: **manifests are authored and
published in shopdb, exported to the share by a button, and the engine,
dispatcher, share layout, payloads, and every PC are completely unchanged.**
That delivers the real pain relief - validated editing instead of hand-edited
JSON, version history, one-click rollback (republish + re-export), desired-state
data sitting next to collector data - at ZERO client risk, with a rollback any
IT tech already knows (restore the `_meta/history` backup file). Natural point to
write ADR-012 with real experience behind it. P4/P5 (HTTP fetch, cutover) are a
separately green-lit second milestone.
### Ranked risks / fail-fast
1. **Generated-JSON vs engine drift (fleet-wide mis-install).** Parity harness
first; CI re-proves parity against checked-in real manifests on every
exporter change; pin lib >= 2.6.
2. **Serving a half-finished draft.** Structural: client reads only
`iscurrent` snapshots; test asserts a draft edit leaves served bytes
unchanged. Must exist before P4.
3. **Availability coupling.** Last-known-good local cache in the first client
prototype; shadow test blocks shopdb and confirms enforce-from-cache + WARN.
4. **SYSTEM HTTP auth + TLS trust.** The ~20-line canary spike in P4 week 1,
before the real client change. Hours of cost; if it fails, Milestone 1 still
delivers full value.
5. **Alias-graph drift.** Seed pins a lib version; harness legacy-name profiles
fail loudly on divergence; new-lib runbook includes "update the alias seed".
6. **Preinstall has no pre-enrollment token.** Keep share-sourced through
Milestone 1/2; decide later.
7. **Editor scope creep.** Three shippable increments; buttons over drag; reuse
JSON preview.
### IT operability (day-to-day runbook, proving the design is manageable)
All in Settings > Imaging PC Types. No PowerShell, no SQL, no share edits.
- **Add an app to a PC type:** open the PC type, Add Entry, pick Type (fields
adapt), fill installer + detection + targeting, Move Up/Down to order, Preview
(+ simulator), Publish with a note. PCs pick it up next 5-min cycle.
- **Bump a version:** drop the new MSI in the scope's `apps/` on the share,
update the entry's Installer + Detection value, Preview, Publish.
- **Roll back a bad publish:** History -> pick last-good version -> Roll Back
(during Milestone 1 also click Export to Share).
- **Canary a risky change:** add the one test PC under Target Hostnames, Publish;
when happy, remove the filter and Publish again.
- **Check "did PC Y get app X":** the simulator with that PC's type/machine
number/CMM version shows exactly which entries apply and why others are filtered.
- **See revision history / who changed what:** the PC type's History tab lists
every published version (date, author, note) with a Roll Back on each; the
Audit Logs page shows the finer-grained draft edits (who touched which entry
field, when) between publishes.