Skip to content
Updated Sep 24, 2026 by Barča Dvořáková · Owner: analysisactivefeature Edit on GitHub

Feature: Tenant Provisioning ​

Purpose: documents the actual, currently-implemented workflow that takes a client from "just created, no tenant yet" through to a fully working tenant — schema created, migrated, seeded, and (for Porsenna platform admins) access converged. Written directly from the EM3 platform-backend code (core/tenant-registry/, core/platform-admin-sync/), not from any prior written spec — this is the gap the tenant entity page's "Provisioning" note flagged as undocumented. Covers both the periodic sweep/lease/retry mechanism and the platform-admin converge step it calls — both described in full below.

Overview ​

POST /v1/clients only ever creates the client row — it does not provision a tenant inline. A periodic sweep, running on the platform worker every 5 minutes, is what actually does the work: it finds every client with no tenant yet and every tenant stuck short of active, takes an exclusive lease on each one, and runs it through provisioning. A tenant that fails is retried by later sweep occurrences (up to a bounded attempt count) before landing in the terminal failed state; an operator can also force an immediate retry through a CLI, which goes through the exact same lease mechanism as the sweep so the two can never race each other onto the same tenant.

Lifecycle ​

                 claim (lease)                    all 9 steps succeed
  [no tenant] ──────────────────► provisioning ─────────────────────────► active ──► suspended ──► deleted
                                     │      ▲
        a step fails, attempts<max  │      │ claimRetry (lease expired, attempts<max)
                                     ▼      │
                              provisioning (lease expired, waiting for the next sweep)
                                     │
     attempts ≥ max (or a dead lease reaped past it)
                                     ▼
                                   failed ──────────────► provisioning   (operator: reset + re-claim)

See TenantStatus. failed is not a dead end — it just stops retrying automatically; a human resets it back to provisioning and it re-enters the same claim path.

The claim (lease) ​

Every provisioning attempt — whether picked up by the sweep or forced by the CLI — first takes a lease on the tenant row: provisioningStartedAt is stamped, provisioningAttempts is incremented, and only the caller holding that exact stamp is allowed to write the row's terminal outcome. This is what lets the sweep and a manual CLI retry run concurrently without corrupting each other's work — whichever claims first wins the row for that attempt; the other observes a no-op ("skipped") rather than a race. A lease has a fixed lifetime (15 minutes); if the process doing the provisioning dies mid-attempt, the lease simply expires and a later sweep reclaims the row — but the attempt already counted against provisioningAttempts, so a dead process still spends retry budget, not free ones.

A tenant with no lease yet (brand new client) is claimed by claimNewTenant; a tenant whose lease has expired without succeeding is claimed by claimRetry. A separate reap pass, run at the start of every sweep occurrence, moves any tenant whose lease has expired and whose attempts are already at the ceiling straight to failed, so a permanently-stuck lease can't sit forever consuming a "still in progress" appearance.

The provisioning steps ​

Once claimed, a tenant runs through the same nine steps whether the sweep or the CLI triggered it:

  1. Create the schema — an empty Postgres schema named tenant_<id without dashes> (the tenant's own id, no separate slug — see tenant.schemaName).
  2. Migrate — run the tenant-schema Liquibase changelog against the new schema, then grant the application's runtime role the privileges it needs on it.
  3. Seed the self-organisation — every tenant needs exactly one organisation row marked as itself (is_self), created here. Idempotent: re-running this step against a schema that already has one is a no-op, not a duplicate.
  4. Ensure the tenant's storage container exists (Azure Blob) — for anything the tenant later attaches (documents, exports).
  5. Seed the default sector — every tenant gets one seeded sector (code + name come from a fixed list; the first one is used).
  6. Seed the root node — the top of the tenant's building/organisation hierarchy, under the sector just seeded.
  7. Mark the tenant active — the terminal write for a successful attempt, guarded by the same lease stamp taken at claim time (if a later claimant has since reclaimed the row, this write is rejected rather than silently overwriting someone else's attempt). This write also clears lastProvisioningError, so a non-null error on an active tenant means exactly one thing: a converge failure recorded after this point, never a stale trace from an earlier attempt.
  8. Converge platform admins — see below. Runs only after step 7 succeeds, since it needs the tenant to already resolve as active.
  9. (Implicit — a step failing before #7 leaves status = provisioning, records the error, and leaves the tenant for the next sweep occurrence to retry, up to the attempt ceiling.)

Steps 1–2 run on a privileged, schema-owning database connection (the only steps that need DDL); every seed write from step 3 onward runs as an ordinary tenant-schema write under the same row-level security every other write in that schema goes through — provisioning does not get to bypass RLS just because the tenant isn't active yet.

Platform admin convergence (step 8) ​

This is the step most relevant to Users & Access (analysed on a separate, not-yet-merged branch): once a tenant is active, every current Porsenna platform admin (a holder of the platform all-access permissionSet) needs a permissionSetUserAssignment row scoped to this tenant — otherwise a brand-new tenant would be invisible to the very people who administer the platform. This is not bespoke to provisioning: it is the same converge/reconcile operation (SyncPlatformAdminAssignmentsUseCase) that also runs stand-alone (via the sync-platform-admin-assignments CLI) to fix up drift — a platform admin promoted or demoted after the fact, or a converge that failed here and needs a manual re-run, since re-running provisioning itself is a no-op once the tenant is active.

What it actually does, in one call:

  • Grants: inserts a permissionSetUserAssignment row for every (platform admin × active tenant) pair that doesn't already have one — this is what makes a brand-new tenant appear for every existing platform admin, and what would make a brand-new platform admin appear in every existing tenant if run without a tenant filter.
  • Revokes: soft-deletes any such row whose holder no longer holds the platform all-access set (only the rows this sync itself created — a genuine direct grant is never touched), and removes the revoked person from any tenant-side group membership the sync itself created for them.
  • Cache invalidation: every created or removed pair, plus (for a fully-revoked admin) their platform-wide cache key, are invalidated in the same call — grant, revoke, and reconcile share one code path precisely so none of the three can drift out of sync with the others.

This is worth calling out against the Users & Access analysis's "Group access propagation" design: it is the same shape — materialize the effective grant as its own row, keep it in sync with a reconcile step, invalidate the cache on the same write — already live in production, for the platform-admin case specifically. That analysis's own auto-provisioning back-fill (a newly-active tenant gets a group for every existing unrestricted organization-visibility permissionSet) is decided to run as a sibling step here, alongside step 8, right after a tenant reaches active — not folded into SyncPlatformAdminAssignmentsUseCase, since that use case reconciles platform-admin permissionSetUserAssignment rows specifically, a different table and a different concern from creating group / permissionSetGroupAssignment rows.

A converge failure at this step does not fail the whole attempt retroactively: the tenant is already active (step 7 already committed), so there is nothing left to retry from provisioning. It is recorded as its own durable trace (lastProvisioningError, tenant already active) and the working remedy is re-running the sync-platform-admin-assignments CLI, not re-running tenant provisioning (which sees already-active and exits without touching the converge step again).

Manual re-trigger ​

An operator can force an immediate attempt via the provision-tenant CLI, keyed by clientId. It goes through the exact same claim primitives as the sweep (claimNewTenant / claimRetry), so whichever side's claim commits first wins the row and the other safely no-ops — the CLI is not a separate, racier path, just an on-demand trigger of the same mechanism. Re-running it against an already-active tenant is a safe no-op (already-active); a failed tenant needs an explicit reset (clearing the attempt ceiling) before it becomes claimable again.

What this doesn't cover ​

  • Tenant suspension and deletion (suspended, deleted) — not part of this workflow; not yet documented anywhere in this catalog.
  • The client-side aggregated tenantStatus (none / provisioning / active / failed / suspended) shown on the Clients screen — a small, pure reduction over a client's own tenant(s) for display purposes, not a new lifecycle of its own.
  • Any UI for triggering, monitoring, or resetting a stuck tenant — today this is operator/CLI-only.