1
IMPORT ADOPTION
Cameron Proudlock edited this page 2026-07-13 19:51:21 -04:00

Importing a site's legacy data

Every adopting site has its own source database - it will not match another site's schema. So the import is split in two layers:

  1. The import API is the stable contract (docs/IMPORT-API.md). Whatever your source looks like, you create flask records through the same documented REST endpoints, authenticated with an admin PAT and the X-Import-Mode header (which preserves legacy timestamps). This layer is the product; it is schema-agnostic.
  2. A per-site loader is thin glue. It reads your source database and POSTs to those endpoints. Nobody runs another site's loader - you copy the pattern.

The West Jefferson loader in scripts/site_imports/wjf/ is reference implementation #1. Read it alongside this guide.

The shape of a loader

  • harness.py - builds the app against the target DATABASE_URL, mints an unscoped admin PAT in-process, and drives the real endpoints through the app test client with Authorization: Bearer <pat> + X-Import-Mode: true. This exercises the same routes/authz/validation an HTTP client would, no running server needed. It also holds read-only access to the source DB and a JSON IdMap of legacy-id -> new-id crosswalks.
  • run.py - ordered stage_* functions. Each reads a slice of the source, POSTs it, and records the crosswalk later stages resolve foreign keys against.

Stage order matters

Reference/lookup tables first (so foreign keys resolve), then the entity hub, then dependents, then links:

reference -> catalog -> assets (persist the source-id -> assetid crosswalk)
  -> dependents (installs, warranties, notifications, ...) -> relationships

The crosswalk is the keystone: capture every legacy id -> new id as you create rows, and resolve foreign keys through it in later stages. New autoincrement ids will not match the source's.

Producing the mapping

You do not have to hand-derive the source -> target mapping. Point the agent-assisted workflow at a source database plus this API contract and it emits a per-table mapping (source columns -> endpoint fields, transforms, what is importable vs out of scope) and a loader skeleton. That is the repeatable onboarding path.

Running (against a THROWAWAY import database)

  1. Build a fresh target: flask db upgrade + flask plugin upgrade-all + flask seed permissions/settings/reference-data. Enable every bundled plugin you need (some ship disabled; a plugin's routes only register when it is enabled at app start).
  2. Load your source dump into a scratch DB the loader can read.
  3. Run the loader stages in order, dry-running / spot-checking as you go.
  4. Verify: row-count + foreign-key-resolution audit against the source, then a UI spot-check (log in, eyeball the lists / map / a detail page).
  5. Only then point a real instance at the imported database.

What the WJ loader demonstrates

  • Fanning one legacy "machine" table out to the flask asset types (computer/machine/network/measuring-tool) by a routing rule, with the duplicate/placeholder/skip decisions applied.
  • Synthesizing a natural key when the source lacks one (printers -> PRN-{id}).
  • Folding a primary IP onto an asset, pairing a check-in/out event log into checkouts, deduping colliding names, reversing an inverse relationship type.
  • The handful of narrow gaps the API cannot cover (e.g. no bulk-communications endpoint) handled as documented direct-ORM writes.