Files
shopdb-flask/docs/proposals/ge-enforce-plugin.md
cproudlock 035419fa51 ADR-015: stop shipping one site's values, and make the rule a gate
The scanner has been reporting the same count for weeks, which is what a rule
that only prints becomes. It now FAILS the build, and it looks where the leaks
actually were: PowerShell, the installer, the seeds, generated JSON, the
frontend - case-insensitively, across plugins, shopdb, scripts, deploy, tools.
A line that is deliberate declares itself with an ADR-015-OK marker and a
reason, so the claim is visible in review instead of tolerated in silence.

What it found, fixed here:

- The shadow client wrote one site's ShopDB URL into HKLM whenever the registry
  disagreed. At the site it was written for that reads as healing drift;
  anywhere else it overwrites the site's own address on every enforce cycle,
  and the site cannot win because the cycle repeats. The bay's value now wins,
  an explicit -BaseUrl seeds it, and with neither there is nothing honest to
  write, so it says so and skips.
- The kiosk dispatcher fell back to one plant's host when HKLM was unset, so a
  kiosk elsewhere quietly opened a server it has no business reaching. The
  fallback is now this site's site_base_url, baked in at seed time, and the
  dispatcher refuses rather than guessing when neither is set. Its legacy
  shortcut matcher derives the host from that URL instead of naming one.
- The OpenAPI generator hardcoded a production hostname into every spec it
  generated, which then published to a public wiki. The relative mount is the
  only server it can honestly name; a site passes its own by environment.
- Placeholders and examples in the UI and the client help offered real internal
  subnets and a real production URL. They now use documentation ranges.

Both publication gates - the export scrub and the docs publishability test -
carry the site patterns, which neither did. One plant's hostname, FQDN and
internal networks are out of the documentation and the generated specs.

Comments naming the reference site are reworded rather than deleted: the
reasoning is worth keeping, the plant name is not what makes it true.
2026-08-14 13:47:39 -04:00

38 KiB

Proposal: GE-Enforce as a shopdb plugin

Status: ACCEPTED / built - see ADR-012 and plugins/geenforce/. Author: planning session 2026-07-12.

1. What this is

Today GE-Enforce is a PowerShell manifest engine that reads per-PC-type manifest.json files off an SMB share (\\shopdb.example.net\ shared\dt\shopfloor\). Each logon, a scheduled task running as SYSTEM mounts the share, reads the manifest for the machine's PC type, and installs or self-heals apps, files, drivers, registry values, and scripts. A parallel preinstall.json runs the same schema once at imaging.

This proposal turns the manifest into shopdb data: the authoritative manifest lives in the shopdb database, is edited through the shopdb UI (an expansion of /settings/pctypemapping), and is served to clients over HTTP as JSON. The payloads (MSI/EXE/PS1/config bytes) stay on SMB, on HTTP, or both, referenced by URL/path from the manifest rows. GE-Enforce.ps1 changes from "read a file on W:" to "GET a manifest from shopdb, then fetch each payload from wherever the row says."

The result: managing imaging PC types, their apps, scripts, files, registry rules, and version gates becomes a first-class shopdb feature instead of hand- edited JSON on a file share.

2. Why it fits shopdb

  • shopdb already models the fleet (the collector ingests every PC's hostname, pctype, installed software, versions). Making shopdb also own what SHOULD be installed closes the loop: desired-state (manifest) and observed-state (collector) live in one system and can be diffed.
  • /settings/pctypemapping already maps gea-shopfloor-* PC types to ComputerType. That page becomes the entry point for full imaging-PC-type management.
  • The plugin contract (per-plugin models, migrations, API prefix, settings cards, collector hooks) is exactly the shape this needs.
  • ADR-004 (per-site instances) matches: each site's shopdb owns each site's manifest. No multi-tenant complication.

3. Grounding: the real manifest schema

Source of truth for these field names (do not invent others):

  • Schema: pxe-images/SHOPDBHOST-v2/shared/dt/shopfloor/_meta/manifest-schema.json
  • Engine: pxe-images/common/lib/Install-FromManifest.ps1
  • Dispatcher: .../shopfloor/common/GE-Enforce.ps1
  • Architecture: pxe/docs/ge-enforce-v2-architecture.md

A manifest is { "Version": str, "_comment": str, "Applications": [entry, ...] }. Only Name and Type are required per entry.

Per-entry fields (complete set)

Identity / action:

  • Name (required, unique, also the status-key <scope>/<Name>)
  • Type (required): one of MSI EXE CMD BAT PS1 INF File Registry
  • _comment (documentation, heavily used in practice)

Type-specific payload references (sparse; depends on Type):

  • MSI/EXE/CMD/BAT/INF: Installer (relative path) + InstallArgs
  • PS1: Script (relative path, falls back to Installer) + Args
  • File: Source (relative) + Destination (absolute on-PC path)
  • Registry: RegPath + RegName + RegValue + RegType (RegType in String DWord QWord MultiString ExpandString Binary)
  • Optional LogFile, WaitTimeoutSec (EXE hang kill), InUseCheck

Detection (decides whether the action fires / self-heals):

  • DetectionMethod: one of Registry File FileVersion Hash MarkerFile ValueMatches pnputil Always
  • DetectionPath, DetectionName, DetectionValue, DetectionPattern
  • Note: DetectionValue is method-dependent - SHA256 for Hash, a 4-part version for FileVersion, a registry value for Registry, ignored for Always/File. Same column, different meaning per method.
  • No DetectionMethod = always installs.

Targeting filters (all ANDed; each is multi-value):

  • PCTypes (array; "*" = all; alias graph expands old<->new names)
  • PCSubTypes / subtype via <pctype>-<subtype> values
  • TargetHostnames (array; exact + -like WJS-* wildcards)
  • TargetMachineNumbers (array; per-bay)
  • _CmmVersion (scalar; per-entry PC-DMIS version gate, needs lib >= 2.6)

Nested:

  • InUseCheck: { Behavior, Processes: [{Name, ExePath, GracefulCloseTimeoutSec}] } Behavior in Defer CloseAndReopen ForceClose ScheduleForReboot

Parsed-but-inert today (model them, mark inert):

  • ApplyMode (Nightly Immediate ImmediateReboot), UpdateWindow (HH:MM-HH:MM)

Preinstall-only extras (phase discriminator):

  • PreEnrollment, KillAfterDetection, PCTypesStrict, _pcTypesNote

Load-bearing behaviors the model must preserve

  1. Array order IS execution order. Config-restore entries are deliberately placed AFTER their vendor installer so a mid-cycle overwrite heals the same cycle (eMxInfo.txt after eDNC; udc_webserver_settings after UDC). We MUST store an explicit per-scope sortorder, not a set.
  2. PCTypes alias graph is many-to-many old<->new names resolved by set intersection, with a PCTypesStrict escape hatch. Not a simple FK.
  3. Polymorphic entry by Type - sparse column set per type. DECISION: one wide manifestentries table with an entrytype discriminator column and nullable per-type columns. NOT SQLAlchemy STI subclasses, NOT a JSON blob. Justification (section 4): the whole fleet is ~64 entries, so sparse columns cost nothing; real columns get validated, indexed, field-diffed, joined against collector data, and read in plain SQL by an IT tech - a JSON blob hides all of that, and class-per-type STI is expert ceremony for no gain. A validate() that switches on entrytype (mirroring the engine's own switch ($App.Type)) is ~40 obvious lines.
  4. Two manifest phases - runtime (self-heal, per logon) and preinstall (once at imaging) share the schema. One table with a phase discriminator.

4. Data model (new geenforce plugin)

Per-plugin Alembic chain (ADR-008). Tables (lowercase concatenated per naming convention). Sizing that shapes every decision here: the real fleet is 10 runtime scopes = 43 entries, plus 1 preinstall manifest = 21 entries, so ~64 rows total. That smallness is why this stays deliberately low-tech (one wide table, JSON-document snapshots, no row-mirroring) - the design target is an average site IT tech maintaining it, not a specialist.

  • manifestscopes - one row per imaging PC type / scope.

    • scopeid PK
    • scopename (e.g. gea-shopfloor-cmm)
    • phase enum (runtime | preinstall)
    • UNIQUE (scopename, phase), NOT scopename alone: common exists in runtime, and a scope name can appear in both phases. Note the phases are shaped differently - runtime is many per-pctype scopes (one manifest file each), preinstall is ONE flat manifest gated internally by PCTypes, so preinstall is modeled as a single phase=preinstall scope, not per-pctype scopes.
    • computertypeid FK -> computertypes (this REPLACES the thin pctypemap_<pxetype> setting; the mapping becomes a column here). Runtime-scope only; null for the preinstall scope.
    • measuringtooltypeid FK -> measuringtooltypes, nullable (metrology scopes: what device this scope implies; keeps imaging + collector agreed, see section 11).
    • manifestversion (string, mirrors manifest Version)
    • description, isactive
    • iscommon bool (the common/ fleet-wide scope)
  • manifestentries - one row per Applications[] entry (the working/draft copy).

    • entryid PK, scopeid FK
    • sortorder int (preserves array order; the ordering contract)
    • name, entrytype (MSI/EXE/.../Registry), comment
    • payload columns (nullable, per type): installer, installargs, scriptpath, scriptargs, sourcepath, destination, regpath, regname, regvalue, regtype
    • payloadsource enum (smb | http | inline) + payloadref (see section 5)
    • payloadsha256 - integrity hash of the payload bytes, INDEPENDENT of the detection method. Mandatory for http/inline payloads; optional for smb. Do NOT reuse detectionvalue for this - detectionvalue is a SHA256 only when detectionmethod = Hash; an MSI with Registry/ FileVersion detection has no payload hash, so an HTTP fetch would otherwise run unverified bytes (see section 5).
    • regvalue stores the RAW JSON literal (1 vs "1") and is emitted verbatim on export. RegValue is untyped in the manifest schema and real entries carry numbers; the engine string-coerces for ValueMatches but Set-ItemProperty -Type DWord cares, so preserve the literal.
    • detection columns: detectionmethod, detectionpath, detectionname, detectionvalue, detectionpattern
    • gates: cmmversion, plus child tables for the multi-value filters
    • control: logfile, waittimeoutsec, applymode, updatewindow (applymode/updatewindow are parsed-but-INERT in the engine today; the UI must label them "not yet enforced" so a tech does not trust a dead gate)
    • preinstall flags: preenrollment, killafterdetection, pctypesstrict
    • isactive
  • manifestpublishedversions - immutable published snapshots, SIMPLIFIED to freeze the rendered JSON DOCUMENT in a single manifestjson column (drop the row-mirrored manifestpublishedentries family the earlier draft proposed). The only consumer of a snapshot is the client, and it consumes exactly that document, so freezing the text makes immutability structural (no UPDATE path), rollback a one-flag iscurrent flip, serving a single-row read, and version diffing a plain text diff - all things average IT can debug; row-mirroring would add ~6 shadow tables and a copy routine that can drift. Columns: publishedversionid, scopeid, versionnumber (1,2,3 per scope), manifestjson (MEDIUMTEXT, verbatim), publishedat, publishedby, iscurrent, notes. Editing manifestentries never affects the fleet; "publish" freezes a new snapshot; the client is ALWAYS served the current snapshot, never the live draft. Rollback = flip iscurrent to an older version (the post-cutover safety net once the on-share JSON is retired). Mirrors today's _meta/history/<date>-<scope>.json backups, but authoritative. Revision history: every publish is a permanent, immutable revision kept indefinitely (snapshots are small JSON text, ~10 scopes - storage is a non-issue). An OPTIONAL retention policy (keep last M per scope, or prune older than N months) can be added later; default is keep-everything, off.

  • Draft-edit audit trail (field-level history BETWEEN publishes): drafts (manifestentries) are not versioned - editing overwrites the working copy. To answer "who changed this entry and when" in the window between two published revisions, log every draft mutation through the EXISTING core audit system (no new table): on create/update/delete of a scope, entry, or child row, write an audit record with the actor, timestamp, entry name, and the changed field(s). This gives per-edit provenance for free and shows up in the same Audit Logs UI IT already uses; the published snapshots remain the coarse-grained "what the fleet actually got" record.

  • manifestentrypctypes, manifestentryhostnames, manifestentrymachinenumbers

    • child rows for the ANDed multi-value filters (one value + a sortorder per row, wildcards stored verbatim as patterns)
  • manifestinusechecks + manifestinusecheckprocesses

    • the nested InUseCheck object and its Processes[] child list (leave gracefulclosetimeoutsec nullable; do not bake the engine's default of 10 into the row, emit it only when set)
  • manifestpayloads - inline payload bytes for payloadsource = inline (entryid, filename, contenttype, payloadbytes LONGBLOB, payloadsha256, uploadedat). App-enforced size cap ~1 MB; the upload UI rejects larger with "use SMB for this" so nobody pastes an MSI into the database. Can ship empty and unused until P6.

  • pctypealiases - a MIRROR of the old<->new name alias graph from Install-FromManifest.ps1:463-475, for server-side resolve/validate only. The engine lib stays the single source of truth (see section 10); shopdb never becomes the authority the client depends on for aliases.

The JSON the client receives is REBUILT from a published snapshot in exact array order. Parity with the current engine is proven by BEHAVIORAL equivalence, not byte-identity (see section 9): re-serialized JSON will differ in key order and whitespace, so the test is that both manifests parse to the same ordered entry set with the same detection/targeting/action semantics.

5. Payloads: SMB and/or HTTP (both supported)

The user asked whether payloads can be SMB and/or HTTP. Yes - per entry:

  • payloadsource = smb: payloadref is the current relative path (apps/eDNC_6-4-5.msi); the client still mounts W: and resolves it against the scope root exactly as today. The engine is unchanged for these rows (the mount + scope-root resolution still happen; an HTTP-only site skips the mount because it has no smb rows). This is the default and the migration target for large binaries (MSIs are hundreds of MB; SMB streaming beats HTTP).
  • payloadsource = http: payloadref is a URL (absolute, or relative to a configured payload base). The client downloads to a local temp dir, verifies the Hash/FileVersion detection value, then runs it. Good for small config/script payloads and for sites with no SMB share.
  • payloadsource = inline: for small text payloads (a .ps1, a config file, a registry value), the bytes live in shopdb itself and are served in-band. No external store at all. Best for scripts and File-type config drops.

Manifest generation emits, per entry, whatever the client needs to fetch the bytes. The engine's existing "stage network EXE to local temp first" logic (SYSTEM access-denied workaround) generalizes cleanly to HTTP download.

Payload integrity uses the dedicated payloadsha256 column, NOT DetectionValue. This is the correction to a subtle trap: DetectionValue is a SHA256 only when DetectionMethod = Hash. Most binaries detect by Registry or FileVersion and carry no payload hash at all, so relying on DetectionValue would let an HTTP/inline-fetched MSI run unverified. Instead, publishing an http/inline payload computes and stores payloadsha256, and the client verifies the fetched bytes against it BEFORE running, independent of how the entry detects install state. smb payloads may set it too (defense in depth) but the share ACL is their primary trust boundary. Detection stays a separate concern: it decides whether to act; the payload hash decides whether the bytes are trustworthy.

Transport security: the client fetches as SYSTEM, so the shopdb TLS cert must be trusted machine-wide. Sites with a self-signed or air-gapped shopdb need the CA in the machine trust store (provisioned by the same Azure DSC step that writes the token). Plain HTTP is acceptable only inside a trusted segment, and even then the payloadsha256 check is what actually guarantees payload integrity.

6. API surface (/api/geenforce/...)

Two permissions via the plugin's get_permissions() hook (split so day-to-day techs can edit but only a lead ships to the fleet):

  • geenforce.manage - create/edit/reorder scopes, entries, drafts, payloads.
  • geenforce.publish - publish, rollback, export-to-share (the fleet-affecting actions).

Draft editing (geenforce.manage):

  • GET/POST /scopes, GET/PUT/DELETE /scopes/<id> - imaging PC types
  • GET/POST /scopes/<id>/entries, PUT/DELETE /entries/<id> - manifest entries
  • PUT /scopes/<id>/entries/reorder - the ordering contract; Move Up/Down in the UI (plain buttons + visible sortorder), not a drag-and-drop dependency
  • POST /entries/<id>/payload - upload an inline/http payload (multipart), compute + store its payloadsha256 (the integrity hash; NOT detectionvalue)
  • GET /scopes/<id>/preview - the draft JSON a client WOULD receive on next publish; GET /scopes/<id>/published shows the currently-served snapshot
  • GET /scopes/<id>/simulate?pctype=&subtype=&hostname=&machinenumber=&cmmversion=
    • the "what would this PC get" simulator: runs the entry list through the same filter logic the engine uses and returns which entries apply and why the rest are filtered out. Reuses the P1 parity harness's filter engine, so it is nearly free, and it is the single most IT-empowering endpoint - it answers "why did/didn't app X install on PC Y" without reading a PowerShell log.

Publishing (geenforce.publish):

  • POST /scopes/<id>/publish - freeze the current draft into a new immutable manifestpublishedversions snapshot (this is what the fleet gets)
  • POST /scopes/<id>/rollback/<version> - mark an older snapshot current
  • POST /scopes/<id>/export-share (or a flask geenforce export-share CLI) - write the current published JSON to <shareroot>/<scope>/manifest.json after copying the existing file to _meta/history/<date>-<scope>.json. This is a first-class feature, not a footnote: it is the Milestone 1 product (author in shopdb, engine untouched) and the permanent break-glass path.

Client-facing (gated by a collector-style service token, geenforce.fetch scope, reusing the PAT + X-API-Key machinery already built for the collector):

  • GET /manifest?pctype=<scope>&subtype=<s>&hostname=<h>&machinenumber=<n> Returns the latest PUBLISHED snapshot for that scope (never the live draft). The server can pre-apply the PCTypes/hostname/machinenumber/cmmversion filters (thin client) OR return the full scope and let the engine filter (fat client, matches today). Start fat: return the scope manifest unchanged so the engine logic is untouched. Include the snapshot version + an ETag so the client can cache and no-op when unchanged.
  • Payload fetch for http/inline rows: GET /payload/<entryid> streaming the bytes; the client verifies them against payloadsha256 from the manifest.

7. Frontend: expand /settings/pctypemapping

The current page (PCTypeMappingSettings.vue, "Collector PC Types") is a read- only-ish table of pxetype -> ComputerType dropdowns. It grows into the imaging- PC-type manager:

  • Scopes list: add/rename/delete imaging PC types; each still carries its ComputerType mapping (that column moves from a setting into manifestscopes). A phase toggle (runtime vs preinstall). Common scope flagged.
  • Scope detail / manifest editor: an ordered list of entries with Move Up/Down buttons and a visible sortorder (the ordering contract made visible; NOT drag-and-drop - a drag library is the kind of dependency that breaks silently and average IT cannot fix; add drag later if wanted). Each entry is a typed form - the visible fields switch on entrytype (MSI shows Installer+InstallArgs; PS1 shows Script+Args; File shows Source+Destination; Registry shows the Reg* quartet), one line of help per detection method. Filter chips for PCTypes/hostnames/machine numbers. InUseCheck sub-editor. Payload source selector (smb/http/inline) with upload for the latter two. applymode/updatewindow sit behind an "Advanced (not yet enforced by the engine)" disclosure. Ship the editor in three usable-alone increments: (a) scope list + entry table, (b) the typed entry form, (c) publish + diff. That keeps the biggest chunk of the build from ballooning.
  • Simulator ("what would this PC get"): a small form (pctype, subtype, hostname, machine number, CMM version) that calls GET /scopes/<id>/simulate and lists which entries apply and why the rest are filtered. The single most IT-empowering piece of the UI.
  • Draft, preview, publish: editing changes only the draft; "publish" freezes an immutable snapshot (see section 4) and is what the fleet then gets. Show the draft-vs-published diff before publishing. Rollback republishes a prior snapshot.
  • Desired vs observed (BUILT: observed-state reporting): rather than extend the collector, the plugin has its own reporting path. Each enforcement cycle a PC POSTs POST /api/geenforce/report (geenforce.report service token) with the published version it applied, the installed/skipped/failed/filtered counts, and per-entry outcomes. Stored in manifestenforcementreports (latest-per-host + history) and manifestenforcementresults (per-entry). Two payoffs fall out: RECEIVED - receivedlatest compares the applied version to the scope's current published version, so the fleet view shows which PCs picked up an update; and SELF-HEAL - each entry's action (installed = drift corrected, skipped = already good, failed) with any warning/error message. Admin reads: GET /reports (fleet compliance) and GET /reports/<id> (per-entry detail). This is the observed half that makes the manifest a closed desired-vs-observed loop.

This is an ADR-010 settings card contributed by the geenforce plugin, so it only appears when the plugin is enabled.

8. Client change (minimal, staged)

GE-Enforce.ps1 today: mount W:, read <scope>\manifest.json, hand to Install-FromManifest. New path: GET the manifest from shopdb, write it to the same local location the engine reads, then run the engine unchanged. That is the smallest possible client delta - the engine, detection logic, self-heal, and SMB payload resolution all stay identical. Only the source of the JSON moves from file to HTTP.

Payloads: smb rows need no client change. http/inline rows need a small fetch-and-verify helper (download to temp, check SHA256, then the existing installer action runs against the local copy). The engine already stages network EXEs to temp, so this is an extension, not a rewrite.

Auth: the client already has SFLD credentials in HKLM:\SOFTWARE\GE\SFLD\Credentials. Add a shopdb service token (a geenforce.fetch PAT) provisioned the same way (Azure DSC writes it to registry), sent as X-API-Key. If shopdb is unreachable, the client falls back to the last-known-good manifest cached locally (fail-safe: never leave a PC unmanaged because the web app is down). This mirrors today's "creds missing = exit 0, retry next cycle" resilience.

9. Cutover strategy

The manifest is desired-state that runs as SYSTEM and installs software fleet- wide. A bad cutover = a fleet-wide mis-install. Stage it:

  1. Import + parity. Write a one-shot importer that reads the current on-share manifests (common + every gea-shopfloor-* + preinstall.json; skip .bak / .pre-mtconnect.bak variants) into the new tables. Then generate JSON back out and prove BEHAVIORAL equivalence for every scope - do NOT chase byte-identity. Re-serialized JSON will differ in key order, whitespace, and _comment formatting, so a raw diff would never converge. The correct test: parse both the original and the regenerated manifest, normalize, and assert the same ordered entry list with identical detection/targeting/action fields per entry (ideally a small harness that mimics the engine's filter+detect decisions and confirms the same entries would fire in the same order on representative machine profiles). That, not byte equality, is what proves the model is lossless. (Same discipline as the ADR-001 data migration.)
  2. Shadow mode. shopdb serves the manifest at a new endpoint; a canary PC fetches from shopdb but ALSO reads the share, and logs any diff. No install behavior changes. Run across one of each PC type for a few cycles.
  3. Read cutover, payloads still SMB. Flip GE-Enforce to source the JSON from shopdb (payloads stay smb). The blast radius is only "where the JSON comes from"; the bytes and engine are unchanged. Keep the share manifests as the rollback (revert the dispatcher one-liner).
  4. Payload migration (optional, per entry). Move small scripts/configs to inline/http opportunistically. Leave big MSIs on SMB indefinitely - SMB is the right transport for them.
  5. Author in shopdb. Once read-cutover is stable, new manifest edits happen in the shopdb UI and the on-share JSON is retired (or auto-exported as a backup for break-glass).

Rollback during cutover (stages 2-4) is a one-line dispatcher revert, because the engine and payload layout never stop working from the share. AFTER the share JSON is retired (stage 5), that escape hatch is gone - post-cutover rollback is republishing a prior manifestpublishedversions snapshot (section 4). Both mechanisms must exist before stage 5, not just the dispatcher revert.

10. Risks / open questions

  • The engine is the contract. Any drift between shopdb's generated JSON and what Install-FromManifest.ps1 expects is a fleet-wide install bug. The byte-identical round-trip test (step 1) is non-negotiable, and the plugin must pin which engine lib version it targets (>= 2.6 for _CmmVersion).
  • PCTypes alias graph must be kept in sync with Install-FromManifest.ps1:463-475. The engine lib stays the single source of truth; shopdb only MIRRORS the map for server-side validation. Do NOT invert this to have the engine fetch aliases from shopdb - that would add exactly the availability coupling the next bullet warns against. When the lib's alias map changes, update shopdb's mirror as part of shipping that lib version.
  • Availability coupling. GE-Enforce currently depends only on SMB. Adding an HTTP dependency on shopdb means shopdb downtime could stall enforcement - hence the last-known-good local cache in section 8. Must be built in from day one, not bolted on. This is also why alias resolution and payloads stay independent of a live shopdb wherever possible.
  • Transport trust. The client runs as SYSTEM, so shopdb's TLS cert must be in the machine trust store (self-signed/air-gapped sites need the CA provisioned via the same DSC step as the token). payloadsha256 verification is the real integrity guarantee and holds even over plain HTTP inside a trusted segment (section 5).
  • Secrets in payloads. Some config drops (site-config, credentials) may contain secrets. inline payloads live in the shopdb DB - those must respect the existing "secrets stay in .env, not the settings table" rule. Likely keep any secret-bearing payload on SMB with ACLs, never inline.
  • Preinstall runner is a separate consumer (00-PreInstall-* at imaging, before enrollment). It may not have a shopdb token yet at that point in the imaging sequence. Preinstall may need to stay share-sourced longer than runtime, or fetch a bootstrap manifest anonymously over HTTP.
  • This is a big build. Realistically phased: (P1) model + importer + behavioral-parity test; (P2) admin API + CRUD + publish/snapshot/rollback; (P3) frontend editor on /settings/pctypemapping; (P4) client fetch + shadow mode; (P5) read cutover; (P6) payload migration. P1 is the gating de-risk - if behavioral parity does not hold, stop. Snapshots (P2) must land before any client points at shopdb (P4), since serving the live draft to the fleet is unacceptable.

11. Relationship to existing work

  • Replaces plugins/computers/pctypemap.py (the thin pctypemap_<pxetype> settings) - the pctype -> ComputerType mapping becomes the computertypeid column on manifestscopes. Two-source transition window: pctype_mapping() must keep reading the settings until the geenforce plugin is enabled, then fall back geenforce-table-first / settings-second, and only retire seed_pctype_settings + the settings at Milestone 1 close. Also reconcile the scope inventory: pctypemap.py lists gea-shopfloor-display but the share has no such manifest dir, and the share has a main/ legacy dir the model ignores
    • the importer creates scopes only from what it finds (plus empty scopes for mapped-but-absent pctypes), and the P1 gate review reconciles the list with the floor team.
  • Also folds in the metrology mapping now living in pctypemap.py (METROLOGY_TOOL_MAP). The collector already auto-creates a MeasuringTool asset and a directional PC->tool controls relationship when it sees a metrology pctype (CMM / Keyence / Genspect / wax-and-trace); the PC stays a shopfloor PC. A metrology scope in the manifest model should carry the attached-measuring-tool type alongside its ComputerType so imaging and collector agree on what device the scope implies.
  • Reuses the collector's token machinery (PAT + X-API-Key + scopes) for the client-facing endpoints.
  • Reuses get_permissions() (contract 0.10.0) for geenforce.manage (edit drafts) / geenforce.publish (publish, rollback, export) / geenforce.fetch (the client service token).
  • Pairs with the collector: desired-state (this plugin) + observed-state (collector) enable a fleet compliance view.

12. Recommendation

Feasible and a strong architectural fit, but it is a multi-phase build with a fleet-wide blast radius. The single most important gate is P1: import the real manifests and prove BEHAVIORAL parity (same entries fire in the same order with the same detection/targeting), not byte-identity. Do not build the UI or touch a client until that parity holds. Three things separate a safe build from a dangerous one and must not be cut: behavioral-parity import (P1), immutable published snapshots with rollback before any client points at shopdb (P2/P4), and a dedicated payloadsha256 for every HTTP/inline payload (section 5). If and when we proceed, this warrants a new ADR (ADR-012: GE-Enforce manifest ownership) capturing the desired-state model, the published-snapshot contract, the SMB/HTTP/inline payload + integrity model, and the fail-safe cache.

13. Execution plan (build order, gates, milestones)

Governing constraint: every step must be runnable and maintainable by average site IT, not just the original developer. Where an earlier draft implied expert machinery, this section simplifies it (and the model above already reflects those simplifications: one wide table, JSON-document snapshots, no row-mirroring).

Phases and gates

  • P0 - Scaffold (S, ~0.5-1 day). flask plugin new geenforce, structure copied from plugins/measuringtools/. Unlike bundled plugins' no-op migration anchors, this NEW plugin's 0001_geenforce_baseline actually creates the tables and registers them in PLUGIN_TABLE_OWNERS (ADR-008). Deploy stays the standard flask db upgrade + flask plugin upgrade-all. Manifest: api_prefix: /api/geenforce, default_enabled: false, tight core_version.

  • P1 - Model + importer + parity harness (M, ~1-1.5 wk). THE GATE. Order inside: tables -> flask geenforce import-share (reads common + every gea-shopfloor-* + preinstall.json, skips .bak, idempotent) -> exporter (rebuilds each scope's JSON from rows in sortorder) -> the parity harness (below). GATE A: flask geenforce parity prints PASS for all scopes. If it cannot pass, STOP the project. No API/UI/client work before Gate A.

  • P2 - Publish/snapshot/rollback + admin API + export-to-share (M, ~1.5-2 wk). Publish freezes rendered JSON into manifestpublishedversions. CRUD per section 6. Plus flask geenforce export-share + an "Export to share" button that writes each scope's published JSON to the share after backing up the old file to _meta/history/. Engine, dispatcher, share layout, payloads, PCs all untouched. GATE B = Milestone 1 (below).

  • P3 - Frontend editor (L, ~2-3 wk; parallel with P4 after P2 API freezes). Expand PCTypeMappingSettings.vue per section 7, in three shippable increments; Move Up/Down not drag; the simulator.

  • P4 - Client fetch + shadow mode (M effort + soak time; needs P2, not P3). Week-1 spike: a ~20-line PS1 on ONE canary PC proves SYSTEM-context HTTP auth + TLS trust before any real client change. Then GE-Enforce.ps1 fetches JSON to a local cache and hands the file to Install-FromManifest.ps1 unchanged; shadow mode installs from the share but logs any diff vs shopdb; ETag + last-known-good cache from day one. GATE C: zero shadow diffs across one PC of every pctype for >= 20 cycles.

  • P5 - Read cutover (S effort, M calendar). Per-scope flip, canary first via TargetHostnames. Payloads stay smb. Rollback = dispatcher revert; share export continues as break-glass. GATE D: all scopes cut over.

  • P6 - Payload migration (S per entry, optional forever). Small configs to inline (verified by payloadsha256); MSIs stay on SMB. Each entry independently revertible (flip payloadsource).

Hard ordering: P0 -> P1 -> P2 -> rest. Snapshots (P2) MUST precede any client pointing at shopdb (P4). P3 and P4 parallelize. Preinstall stays share-sourced through at least Milestone 1 (no token pre-enrollment; export writes preinstall.json too, so it is authored-in-shopdb for free with no client risk).

The P1 parity harness (concrete, IT-re-runnable)

plugins/geenforce/parity.py + a CLI, also wrapped as a CI test. Two checks per scope, output one readable line per scope (entries N/N identical profiles M/M same-fire PASS), exit 0/1, prints the first differing entry/field on fail:

  1. Lossless field check (order-preserving). Canonicalize each entry to exactly the fields the engine reads (Name, Type, the payload fields, all Detection*, the filter arrays, _CmmVersion, InUseCheck, preinstall flags); exclude _comment and key order (documentation, not behavior). Compare the ordered lists position by position.
  2. Same-entries-fire-in-same-order. Re-implement in ~120 lines of Python the engine's four filter functions exactly as written in Install-FromManifest.ps1 (Test-PCTypeMatches incl. the alias groups at lines 463-475, "*", and <Type>-<SubType>; Test-HostnameMatches exact + -like; Test-MachineNumberMatches; Test-CmmVersionMatches). For each machine- profile fixture, run BOTH manifests through it and assert the identical ordered list of entry names that pass all filters. Detection itself is not executed - check 1 already proved detection fields identical, so identical inputs to detection are guaranteed. This pair proves losslessness without byte-diffing.

Fixtures (plugins/geenforce/parityfixtures.json, ~16-18 profiles): one per pctype; CMM version variants 2016/2019/2026/empty; collections machine-number variants (a credentialed bay, an MTConnect bay, neither); legacy-alias profiles (Standard+Machine, CMM) to exercise the alias graph both ways; a WJS-* hostname-wildcard profile; preinstall profiles including one that hits PCTypesStrict. Watch-items the harness must handle: empty Applications: [] scopes (4 exist), entries with NO DetectionMethod (fire every run), and the regvalue literal typing.

First slice: one vertical through gea-shopfloor-cmm

Only 4 entries but hits every hard part - MSI type, Registry detection with and without a pinned value, nested InUseCheck with Processes[], and the _CmmVersion gate. Tables: scopes, entries, entrypctypes, inusechecks + processes, publishedversions, pctypealiases. flask geenforce import-share --scope gea-shopfloor-cmm; flask geenforce publish gea-shopfloor-cmm; one endpoint GET /api/geenforce/manifest?pctype=gea-shopfloor-cmm serving the published snapshot (fat-client, ETag, collector-style X-API-Key/PAT auth reusing shopdb/core/api/collector.py). Done = parity PASS for cmm; the endpoint's JSON fed to Install-FromManifest.ps1 on a bench CMM PC logs 4 skipped identically to the share manifest; editing a draft does NOT change the served bytes but publishing does, and rollback restores the prior published bytes; unauth = 401, wrong-scope = 401.

End of P2 plus the publish/scope-list slice of P3: manifests are authored and published in shopdb, exported to the share by a button, and the engine, dispatcher, share layout, payloads, and every PC are completely unchanged. That delivers the real pain relief - validated editing instead of hand-edited JSON, version history, one-click rollback (republish + re-export), desired-state data sitting next to collector data - at ZERO client risk, with a rollback any IT tech already knows (restore the _meta/history backup file). Natural point to write ADR-012 with real experience behind it. P4/P5 (HTTP fetch, cutover) are a separately green-lit second milestone.

Ranked risks / fail-fast

  1. Generated-JSON vs engine drift (fleet-wide mis-install). Parity harness first; CI re-proves parity against checked-in real manifests on every exporter change; pin lib >= 2.6.
  2. Serving a half-finished draft. Structural: client reads only iscurrent snapshots; test asserts a draft edit leaves served bytes unchanged. Must exist before P4.
  3. Availability coupling. Last-known-good local cache in the first client prototype; shadow test blocks shopdb and confirms enforce-from-cache + WARN.
  4. SYSTEM HTTP auth + TLS trust. The ~20-line canary spike in P4 week 1, before the real client change. Hours of cost; if it fails, Milestone 1 still delivers full value.
  5. Alias-graph drift. Seed pins a lib version; harness legacy-name profiles fail loudly on divergence; new-lib runbook includes "update the alias seed".
  6. Preinstall has no pre-enrollment token. Keep share-sourced through Milestone 1/2; decide later.
  7. Editor scope creep. Three shippable increments; buttons over drag; reuse JSON preview.

IT operability (day-to-day runbook, proving the design is manageable)

All in Settings > Imaging PC Types. No PowerShell, no SQL, no share edits.

  • Add an app to a PC type: open the PC type, Add Entry, pick Type (fields adapt), fill installer + detection + targeting, Move Up/Down to order, Preview (+ simulator), Publish with a note. PCs pick it up next 5-min cycle.
  • Bump a version: drop the new MSI in the scope's apps/ on the share, update the entry's Installer + Detection value, Preview, Publish.
  • Roll back a bad publish: History -> pick last-good version -> Roll Back (during Milestone 1 also click Export to Share).
  • Canary a risky change: add the one test PC under Target Hostnames, Publish; when happy, remove the filter and Publish again.
  • Check "did PC Y get app X": the simulator with that PC's type/machine number/CMM version shows exactly which entries apply and why others are filtered.
  • See revision history / who changed what: the PC type's History tab lists every published version (date, author, note) with a Roll Back on each; the Audit Logs page shows the finer-grained draft edits (who touched which entry field, when) between publishes.