Collect what bays actually have, separately from what they are told to have
Some checks failed
CI / backend (push) Failing after 7m15s
CI / naming (push) Failing after 7m22s
CI / frontend (push) Failing after 7m14s
CI / migrations-mysql (push) Failing after 7m14s

ShopDB knew what a bay SHOULD have and nothing about what it DOES. Adding the
observed half makes a rollout a review instead of a typing exercise: the floor
reports itself in, you look, and you adopt.

The collection uses the mechanism that already exists rather than a new one.
POST /api/collector/printers dispatches to the printers plugin's
apply_collector_payload, the same ADR-006 hook the computers and backups plugins
implement. New client script, new plugin-owned table, no new transport and no new
credential.

OBSERVED AND ASSIGNED STAY APART, and that is the point rather than a detail. A
collector report can never write an assignment row: _reconcile_edges is the only
function that writes usesprinter/defaultprinter, it has two call sites, and both
are authenticated routes a human calls. If a drifted bay's own state were allowed
to become what it is told to install, every configuration error would become
permanent the next time that PC checked in.

Seeding an assignment from observed state is explicit -
POST /assignments/seed-from-observed - because a rollout adopts many machines at
once. It routes through the same _reconcile_edges as the editor, so there is one
write path with two doors, and a queue matching no known printer is REFUSED
rather than guessed into an assignment. That last rule is the lesson from the
measuring tools: adopting on a weak key produced 43 duplicate instruments.

Two fixes on top of what the agents built. The replace deleted a host's previous
rows by exact case-folded name while the read path treats a short name and its
FQDN as one machine, so a PC that changed spelling appeared to hold every queue
twice - which reads as drift that is not there. And the client sent 'reportedat'
where the declared schema said 'observedat'.

Also here: the legacy loader now imports machines.printerid, the classic system's
record of each machine's default printer, which it silently dropped - the
production import would have lost every one. And Set-ShopdbPrinters.ps1 finally
registers the per-user logon task, staging Apply-ShopdbDefaultPrinter.ps1 to
C:\ProgramData first because the share it lives on is mounted only during the
enforcement cycle and the task runs at logon when it is gone.

VALIDATED ON WINDOWS 11 (build 26200), not just on Linux pwsh, which parses these
scripts happily and executes none of the spooler branches.

The reporter: posts a correct payload with the X-API-Key header; resolves BaseUrl
and CollectorKey from HKLM when given no arguments; suppresses the virtual queues
by port; resolves port addresses; and reads the CONSOLE USER's default out of
HKU rather than SYSTEM's own, which is a different and usually wrong answer.

Two results matter more than the rest. With the spooler stopped, both the cmdlet
and the CIM path fail and the script posts NOTHING - verified against a capture
server that recorded zero requests, where an empty list would instead have
erased that host's observed rows and read as a bay that lost its printers. A
genuinely empty host still posts [], because that is a real and different fact.

The logon task registers as the Users group at Limited, and falls back to the
well-known SID S-1-5-32-545 when the group name will not resolve, as it will not
on localised Windows. It was then run with the source directory RENAMED AWAY, to
stand in for the share being unmounted, and it still moved the user's default -
which is the whole reason the script is staged to C:\ProgramData rather than run
from where it lives.

The guarantees against damage were re-checked rather than assumed: an empty
assignment changes nothing, an unreachable server changes nothing, -WhatIfOnly
leaves no queue, no task, no staged file and no registry value behind, and a
drifted queue is repointed IN PLACE with Set-Printer so whoever has it as their
default keeps it.

Not covered by any of this: the driver-staging path, which needs a real vendor
package rather than the class drivers a VM ships with.
This commit is contained in:
cproudlock
2026-08-19 14:45:39 -04:00
parent 1a5a1cd43d
commit 2d09fa3201
15 changed files with 2903 additions and 30 deletions

View File

@@ -8,6 +8,8 @@ a real caller (the GE-Enforce fleet agent) to it.
- Server code: `shopdb/core/api/collector.py`
- Computers schema + upsert: `plugins/computers/plugin.py`
(`get_collector_schema` / `apply_collector_payload`)
- Printers schema + replace (observed print queues): `plugins/printers/plugin.py`
(same two hooks)
- Contract rationale: `docs/adr/ADR-006-collector-contract.md`
---
@@ -321,6 +323,137 @@ classic reporter posts a full `networkInterfaces` array; the collector accepts
only one `ipaddress`, so pick the corp/routable NIC (see the corp-range gate in
the PowerShell below).
### Printers plugin field mapping (observed print queues)
`POST /api/collector/printers`. The other half of the printer story. ShopDB has
always known what a bay SHOULD have - `GET /api/printers/for-host/<hostname>`,
applied by `Set-ShopdbPrinters.ps1` - and has never known what it actually has.
This payload is that missing half. Same dispatcher, same auth, same audit row as
the computers collector; only the payload and the plugin differ.
- Schema + apply: `get_collector_schema` / `apply_collector_payload` in
`plugins/printers/plugin.py`.
- Client: `plugins/printers/client/Report-PrintersToShopDB.ps1` (SYSTEM context,
`Type=PS1` / DetectionMethod `Always` manifest entry, logs to
`C:\Logs\Shopfloor\report-printers-YYYYMMDD.log`, always exits 0).
- Storage: `printerobservedqueues`, one row per observed queue, owned by the
printers plugin. It is a separate table from the assignment on purpose - see
"Observed is not assigned" below.
The identity field is `hostname`. Unlike the computers payload this one is NOT
patch-style: `queues` is required, and the reported set replaces everything
previously recorded for that host.
| Payload field | Type | Server behaviour (`apply_collector_payload`) |
|---|---|---|
| `hostname` (required) | string | Identity of the report. `COMPUTERNAME` or its FQDN; matched case-insensitively. Rows are keyed by the NAME, so a bay that reports before its PC record exists still records everything - the report resolves to an asset the moment that record appears. An unknown hostname is a warning, never an error. |
| `queues` (required) | array of objects | Every real print queue on the host. This REPLACES the host's previous set. An empty array is a valid report meaning "this bay has no queues" and clears the host's rows. A missing key is rejected (400), because one client bug that dropped the field would otherwise erase the fleet's observed state host by host. |
| `queues[].queuename` (required) | string | Windows printer name. A queue with no name is skipped with a warning; a repeated name within one report is dropped with a warning (Windows cannot hold two queues of one name). |
| `queues[].drivername` | string | Driver name verbatim, as the INF spells it. Compared against the assignment's driver to detect drift. |
| `queues[].portname` | string | Windows port name. Also tried as a match key, since a port created outside the client is usually named after the host address. |
| `queues[].portaddress` | string | `PrinterHostAddress` of a TCP/IP port - an IP or FQDN. The PRIMARY key for matching a queue to a printer asset. Omit it for a non-TCP port (USB, WSD, redirected); such a queue is still worth reporting and matches on name alone. |
| `queues[].isdefault` | boolean | True on the one queue that is the user's default. If several arrive true, the first is kept and the rest are cleared with a warning - a bay has exactly one default, and keeping both would leave a seed picking one at random. |
| `queues[].isshared` | boolean | True when the queue is shared off this PC. Recorded, not acted on. |
| `observedat` | ISO-8601 datetime | Accepted and ignored, as is any other client timestamp. The server stamps `observedat` at ingest, one value for the whole report, so a bay with a wrong clock cannot report itself fresh or stale. |
Response `data` is the standard collector shape: `action` is always `updated`
(this endpoint replaces rows and creates no asset - calling an identical
re-report `noop` would hide that the bay is still checking in), `assetid` is the
resolved PC or `null`, `extra.queuecount` is how many rows were stored, and
`warnings` carries the soft problems above.
Schema source of truth: `get_collector_schema` in `plugins/printers/plugin.py`.
If you change the payload, change it there and re-check this table.
```
POST /api/collector/printers
X-API-Key: SECRET
Content-Type: application/json
{
"hostname": "WORKSTATION01",
"queues": [
{
"queuename": "Bay Label Printer",
"drivername": "Generic / Text Only",
"portname": "IP_192.0.2.40",
"portaddress": "192.0.2.40",
"isdefault": true
},
{
"queuename": "Office Laser",
"drivername": "HP Universal Printing PCL 6",
"portname": "IP_192.0.2.41",
"portaddress": "192.0.2.41",
"isdefault": false
}
]
}
```
#### The latest report replaces the previous one
This is CURRENT STATE, not history. Each report deletes every row previously
recorded for that hostname and writes the reported set in its place, inside the
dispatcher's transaction. Nothing accumulates, "what does this bay have" is
never a question about time, and a queue removed from a bay disappears from
ShopDB on the next cycle without anyone tidying up.
That has one sharp edge, and it belongs to the client: an empty `queues` array
is a legitimate report that wipes the host's rows. A client whose enumeration
FAILED must therefore send nothing at all rather than an empty list. Reporting
nothing loses one cycle; reporting `[]` after a WMI hiccup deletes real state
and reads as a bay that lost its printers. `Report-PrintersToShopDB.ps1` tracks
this with an `$enumerated` flag and exits without posting when both the cmdlet
and the WMI fallback failed.
#### Observed is not assigned
Stated plainly, because it is the whole point of keeping two tables:
**A collector report NEVER becomes an assignment.** Nothing on this path writes
a `usesprinter` or `defaultprinter` row. The moment a drifted bay's observed
state is treated as correct, enforcement stops meaning anything - a bay that
installed the wrong printer would make itself right simply by reporting it.
Observed state becomes assigned state only when a person asks for it, through
`POST /api/printers/assignments/seed-from-observed/<assetid>` (requires
`printers.edit`), after reviewing the comparison. That route refuses rather than
guesses: it will not seed from an ambiguous host, will not seed queues that
match no printer unless told to, and will not write an empty set.
#### Matching, and what "unknown" means
Nothing is matched at ingest - the raw observed strings are stored as reported,
and resolution to a printer asset happens at READ time, so a printer added to
ShopDB tomorrow matches yesterday's report without the bay reporting again.
At read time each queue is matched by PORT ADDRESS first (an IP or FQDN names
one device unambiguously), then by queue name. There is no fuzzier fallback: a
queue that matches nothing is reported as `unknown` rather than guessed, because
a wrong match seeds a wrong assignment, which is worse than no assignment.
#### Reading it back
| Endpoint | Purpose |
|---|---|
| `GET /api/printers/observed/<hostname>` | What the host last reported, every queue classified against the assignment it resolves to (`matching`, `drifted`, `extra`, `missing`, `unknown`), with `observedat`, a summary, and a `seedcandidate` preview. Read-only. Needs `printers.view`. |
| `POST /api/printers/assignments/seed-from-observed/<assetid>` | The one path from observed to assigned, human-triggered. Needs `printers.edit`. Normally posted against the MACHINE, so the assignment survives a reimage. |
#### Key delivery for the printers reporter
The reporter reads its server and its key from the registry the enforcement
client already provisions (`HKLM:\SOFTWARE\GE\ShopDB`, values `BaseUrl` and
`CollectorKey`), or from `-ApiUrl` / `-ApiKey` in the manifest entry's `Args`.
Nothing is baked into the script or the manifest JSON on the share - the same
rule as the computers reporter, for the same reason.
Server side, the key resolves per-plugin first: set `COLLECTOR_API_KEY_PRINTERS`
if you scope keys per collector, or rely on the shared `COLLECTOR_API_KEY`. A
site that has scoped the computers key per-plugin and set no printers key gets
401 on this endpoint until one of the two is in place. A `collector.ingest`
managed token works here exactly as it does for computers.
---
## 2. Breaking change: querystring `api_key` removed
@@ -703,8 +836,13 @@ fleet's key is scoped to the computers collector only.
| `No collector registered for plugin computers` | 404 | The computers plugin is disabled or not loaded on that instance. Enable it. |
| `hostname is required` / `No data provided` | 400 | Empty body, non-JSON body, or missing identity field. Check `-ContentType 'application/json'` and that `hostname` is set. |
| `Internal error processing collector payload` | 500 | Generic by design - the server does not leak the cause to the caller. The real error (DB, unexpected exception) is in the server log (`Collector upsert failed for computers`). Check there. |
| `No collector registered for plugin printers` | 404 | The printers plugin is disabled or not loaded on that instance. Enable it. A lean site without it simply does not collect observed queues. |
| `queues is required; send an empty array for a host with no queues` | 400 | The printers payload omitted `queues` entirely. Absent and empty are not the same thing here - see "The latest report replaces the previous one". |
| A bay's observed queues all vanished | 200 | It reported `queues: []`, which legitimately clears the host. Check the client log for `posting NOTHING` (a failed enumeration correctly sends nothing) versus `reporting an empty set`. If the bay really has queues, the enumeration on that host is the thing to fix. |
| Observed queues show as `unknown` | 200 | The queue matched no printer asset by port address or by name. Add the printer to ShopDB (with its IP) and re-read - matching happens at read time, so the bay does not need to report again. Never seeded automatically, by design. |
| `warnings` present but `action` is created/updated | 200 | Soft issues only; the row WAS written. Common: unmapped `pctype` (fix the pctypemap setting), unknown `osname` (add it to the `operatingsystems` vocab), unknown app name, `pcsubtype` not stored. No action needed unless the warning matters to you. |
Client-side log for the fleet reporter: `C:\Logs\Shopfloor\collector-YYYYMMDD.log`.
Client-side log for the fleet reporter: `C:\Logs\Shopfloor\collector-YYYYMMDD.log`;
for the printer-queue reporter: `C:\Logs\Shopfloor\report-printers-YYYYMMDD.log`.
Server-side: the Flask app log (the collector logs upsert failures and per-plugin
schema failures there).