Files
shopdb-flask/docs/adr/ADR-004-deployment-topology.md
cproudlock 96df19702e docs: one manual runbook, and ADR statuses that mean something
DEPLOY-WINDOWS-IIS was a second copy of the manual IIS procedure that had
diverged from the first: a different MySQL version (8.0, which reached end of
life in April), a different port, a different plugin list, and a profile file
that does not exist. Two runbooks for one procedure means a reader follows
whichever they found, and one of them was wrong. INSTALL-WINDOWS-IIS covers
everything it did plus a preflight step and the subpath method, so the one
section it uniquely had - redeploying a hand-built server - is folded in there,
with the plugin-chain step it was missing and a note to back up first, and the
duplicate is gone. Everything that pointed at it now points at the survivor.

Three ADR statuses said something untrue.

ADR-013 said PROPOSED while half of it had shipped and ADR-014 had been accepted
on top of it. A decision that has been implemented and depended upon is not
proposed, and leaving one that way devalues every other status in the index. The
catalog half is still unbuilt, which is the ordinary state of an accepted
decision: accepted means settled, not delivered.

ADR-016 said ACCEPTED for a design where nothing is built - the endpoint and
permissions it describes do not exist, so a reader goes looking for them. The
status stands, because the decision does; the header now says so plainly and
points at where today's credentials actually live.

ADR-003 and ADR-004 were ACCEPTED with their own Decision lines still opening
"**PROPOSED:**", which reads as though the decision was never taken.

And the dashboard proposal carried Status: ACCEPTED, which belongs to a decision
record. A proposal is a proposal; the contract it produced is the ADR.
2026-08-14 16:20:59 -04:00

5.3 KiB

ADR-004: Deployment topology (per-site instances)

  • Status: ACCEPTED
  • Date: 2026-05-08
  • Deciders: ShopDB maintainers

Context

shopdb-flask manages shop-floor inventory. If multiple GE Aerospace sites adopt it, the deployment can take one of two shapes:

Model How it works
Per-site instances Each site runs its own Flask + MySQL + Vue stack. Each site has its own DB, its own users, its own enabled-plugin list, its own deploy. Sites are isolated.
Multi-tenant single instance One central Flask + MySQL + Vue stack serves all sites. A siteid foreign key on every asset partitions data. Auth distinguishes which site a user belongs to.

The codebase today is single-tenant per deployment. There is no siteid column, no tenant filter, no cross-site auth model. Plugins can be enabled / disabled but only globally for the running instance.

Decision

Per-site instances. Each adopting site runs its own dedicated stack. The framework does not support multi-tenancy.

Each site:

  • Owns its database (own credentials, own backup policy, own retention). The database charset is part of the contract: it must be utf8mb4 (utf8mb4_unicode_ci). The migration chain creates every table utf8mb4, and the connection pins ?charset=utf8mb4. A site that creates the database with a different default charset (older MySQL defaults to latin1) gets a schema that silently diverges from every other site. See docs/DEPLOY.md.
  • Picks its own enabled plugins
  • Configures its own JWT secret, CORS allowlist, Zabbix integration, Active Directory binding
  • Deploys at its own cadence

The framework provides:

  • A Dockerfile and docker-compose.yml template suitable for a single-site deploy
  • A .env.example listing all required environment variables
  • A docs/DEPLOY.md walking through a fresh-site install

Consequences

Positive

  • Simpler code: no tenant filter on every query, no cross-tenant auth, no shared-state partitioning bugs.
  • Sites are independent. A schema change at one site does not affect another. A plugin crash at one site does not blast radius to other sites.
  • Clear ownership: each site's IT team owns their own stack and data. Compliance and audit boundaries match operational boundaries.
  • Aligns with how GE Aerospace sites already operate (independent IT, independent shop floors).

Negative / cost

  • No cross-site reporting out of the box. If GE corporate ever wants a fleet-wide view, it has to be built on top (e.g., a roll-up dashboard that queries each site's API). That layer is out of scope for the framework.
  • Each site administers its own stack. Higher operational overhead than a single central instance, but each site already runs its own infrastructure.
  • Updates require visiting each site's deploy. Fine for the current adoption model; revisit if dozens of sites adopt.

Neutral

  • No siteid column needed. The existence of one DB per site is the partition.

Alternatives considered

  1. Multi-tenant single instance. Lower operational overhead at scale, easier cross-site reporting, but adds significant code complexity and risk: every query needs a tenant filter, auth gets complex, schema migrations affect every site at once, and a bug at one site can leak data across sites. Rejected for v1; revisit if and only if more than five sites adopt and operational overhead becomes painful.
  2. Hybrid: per-site DB but central app server. Adds the operational complexity of multi-tenancy without isolating the failure domain (one app crash = all sites down). Rejected.

Migration strategy (resolved)

Deploys run two commands: flask db upgrade then flask plugin upgrade-all (lean/ADR-014 sites add an optional flask plugin prune-schema at initial provisioning). The core Alembic chain applied by flask db upgrade creates the full core AND bundled-plugin schema through the chain head (this includes migration 7c04_fold_plugin_schema). But every bundled plugin still carries its own Alembic chain per ADR-008: flask plugin upgrade-all stamps each plugin's own chain (the alembic_version_<plugin> tables) and applies any plugin-specific migrations added after the ownership cutover. The earlier Phase 7B state that folded everything into core with no per-plugin chains was superseded by ADR-008's per-plugin ownership. A fresh flask db upgrade reproduces the live core schema exactly (verified on a scratch DB).

External (out-of-tree) plugins per ADR-003 ship their own migrations too; the framework runs the same per-plugin chain mechanism (ADR-008) for both bundled and external plugins.

Open questions

  • Should the framework provide an optional read-only fleet roll-up mode where a "central" instance can pull aggregate metrics from each site's API? Defer. Out of scope for v1.
  • Backup strategy per site: framework recommendation, or each site decides? Framework should publish a recommended backup runbook (mysqldump + offsite copy) but not enforce.
  • Auth federation: each site has its own user table, or sites can share an LDAP / SSO? Recommend documenting the LDAP config knob in .env.example so sites can plug in their own auth without code change.

References

  • shopdb/config.py (currently single-tenant, no siteid)
  • ADR-001 (asset model is per-site, not cross-site)
  • ADR-003 (plugin distribution per site)