Skip to content

Onboarding & provisioning

The goal is a fully self-serve path: someone finds the landing page, signs up, pays, and is working in their own instance minutes later — the way a team adopts Slack or Teams, with no sales call required. Nothing about that flow is bespoke per customer. A new tenant is, mechanically, a fresh schema + one metadata deploy + a billing link, and everything below is how those three pieces come together automatically.

The load-bearing idea: almost none of this lives in a tenant’s own metadata. It lives in a control plane — a cross-tenant layer the kernel owns — plus a versioned base-org template that is itself just metadata. Provisioning is “stamp the template into a new schema and attach a subscription.”

Landing page ──▶ Sign up ──▶ Verify ──▶ Plan + pay ──▶ Provision ──▶ First login
(marketing) company & email / Stripe allocate admin lands
admin user card Checkout schema + in their org
base template
Step What happens Where it runs
1. Sign up Visitor submits company name, work email, desired subdomain, admin name. A provisional Org row is created in the control-plane registry with status pending. Control plane
2. Verify Email-ownership check (click-through link); optional work-domain check. Anti-abuse/velocity screen. Nothing is provisioned yet. Control plane
3. Plan + pay Visitor picks a plan and enters payment via the billing provider’s hosted checkout. A Billing Account is created and linked 1:1 to the Org; the subscription sets seats + plan. A valid card is itself a strong trust signal. Billing provider ↔ control plane
4. Provision On the “payment/trial confirmed” event, the provisioner allocates the tenant’s schema, deploys the base-org template into it, seeds the first admin, and wires routing. Seconds, unattended. Control plane → kernel provisioner
5. First login The admin lands at their-org.caos.example.com, already holding the Platform Admin capability bundle, in an org pre-populated with the standard objects, layouts, and permission bundles. Tenant

A free/developer tier is the same flow with step 3 as a card-optional trial and tighter trial limits; the contract, when there is one, is a digital click-through, not a step that blocks provisioning.

Control plane vs. tenant — what is kernel, what is metadata

Section titled “Control plane vs. tenant — what is kernel, what is metadata”

This is the crux of “what needs to be in the kernel vs. metadata.” Three distinct layers, and keeping them separate is what makes onboarding safe and repeatable.

Layer Holds Kind Scope
Control plane The tenant registry (every Org, its schema/connection, plan, status), the provisioner, the router (subdomain → Org → schema), auth/session, and billing-status enforcement Kernel code + a control-plane schema the kernel alone owns Cross-tenant — never visible to any tenant
Base-org template The standard objects, fields, layouts, list views, capability bundles, automation defaults, and AI settings a new org starts with; plan/entitlement definitions Metadata components (versioned, deploy-pipeline artifacts) A package, stamped into each new tenant
Tenant The customer’s own records, users, permission-set assignments, and any customization they make after signup Rows + per-tenant metadata in their schema One tenant — isolated by schema + RLS

The payoff of the middle row: what a new customer receives is not hard-coded — it is a metadata package run through the ordinary retrieve → diff → apply pipeline. Improving the new-customer starting point is a metadata change to the template, versioned and deployable, not a code release. The control plane (top row) is the only genuinely kernel-level new surface here; it is cross-tenant by nature, so it deliberately lives outside any tenant’s metadata.

The control plane needs somewhere to be operated from, and it cannot be a tenant surface. The registry, the base-org template catalogue, the plan and entitlement definitions, and the region map are cross-tenant objects — they exist above every schema and are named by none of them. The surface that edits them is the Operations console, and it is defined by its origin, its identity model, and its capability model.

Its own origin, for a mechanical reason. Tenants resolve under one wildcard — their-org.caos.example.com. The console does not: it is published on a host outside that wildcard’s DNS parent, and its session cookies are host-only. A cookie scoped to the tenant parent is presented to every tenant host on every request, so an operator session living inside that scope would be handed to front-end code that a tenant’s own administrator configures. A separate origin removes the possibility instead of defending against it. The rule runs backwards through the registry too: labels that could resolve to platform infrastructure — operations, admin, api, www, status, login — sit on a reserved list the subdomain allocator refuses, so no tenant can claim a host the platform answers on.

Operators are control-plane principals, not users. An operator has a row in the control-plane schema and no row in any tenant. They hold no permission-set assignment, appear in no tenant’s user list, and consume no seat. Authentication is federated against the platform’s own identity provider with a hardware-backed factor required, no password path, and short sessions — the platform’s operators are a small, known population, so the strongest available method is the only method rather than a policy someone opts into.

Capability comes from operator grants, which are not permission sets. A small fixed set, kernel-defined, additive and with no deny, assigned to operator principals directly or through operator groups:

Grant Opens
operator.registry.read The tenant registry, plan, status, and template version — no tenant data
operator.tenant.provision Creating a tenant, retrying a failed provision
operator.tenant.suspend Moving an Org to suspended and back
operator.tenant.deprovision Running offboarding, including rollback of a failed provision
operator.billing.manage Plan, seat, and entitlement changes outside the self-serve path
operator.template.publish Publishing a base-org template version to a region
operator.support.session Opening a support session inside a tenant (below)

They cannot be permission sets, for two independent reasons. A permission set is tenant metadata — a tenant’s own security administrator can edit one and assign it — so sourcing platform power from a tenant’s schema would let a tenant grant itself the registry or revoke the operator who is meant to suspend it. And a permission set is scoped to one schema: it can name that tenant’s objects, fields, and system permissions, and nothing else. The registry, the router, the template catalogue, and billing are not in any tenant’s schema, so there is no key a permission set could name to grant them. Operator grants are also not platform.root and do not imply it: root is a tenant’s own capstone grant, held by that tenant’s administrator, and holding every operator grant confers none of it.

Inside a tenant, an operator is a support session. The console reads the registry, not tenant data — there is no view in it that renders a customer’s records. Entering a tenant is a separate, deliberate act gated on operator.support.session, and it behaves like the tenant-side login-as it is modelled on: time-boxed, read-only unless a second grant is held, and requiring a typed reason before it starts.

The auditing consequence is the part that matters, and it follows from operators being control-plane principals rather than tenant users. A support session mints a session that resolves against the tenant’s own access planes and is stamped with the operator’s identity and the session id. Every entry it produces — metadata audit, data history, execution trace — names the operator as actor, in the tenant’s own streams, alongside the reason and the session id. Nothing an operator does inside a tenant is attributable to “system”, to the platform, or to a tenant user, because there is no code path that writes as one: the actor is a principal that only ever means one human. The tenant sees the session while it is open (a banner and a notification to its administrators) and afterwards (its audit stream), and by default a session requires the tenant’s approval before it opens — a tenant may pre-authorize a standing window, and the choice is theirs. Control-plane actions that are about a tenant rather than inside it — provision, suspend, deprovision, a plan change — land in a control-plane audit stream, and each tenant sees the entries concerning itself mirrored into its own stream, so “who suspended us, when, and why” is answerable by the customer without asking anyone.

“Their own instance” resolves against the decided multi-tenancy topology: single database, schema-per-tenant. Provisioning a standard tenant is fast and unattended:

  1. Allocate — insert the Org into the registry, create its Postgres schema, reserve its subdomain.
  2. Stamp the template — run the base-org metadata package into the new schema via the deploy pipeline (CREATE TABLE per standard object, layouts, list views, capability bundles, AI defaults). Same transactional apply used for any deploy.
  3. Seed identity — create the first admin user, assign the Platform Admin bundle, send a set-password / SSO invite.
  4. Wire — bind subdomain → Org → schema in the router and link the billing subscription.
  5. Activate — flip Org status to trialing or active; the tenant is open for business.

Because a standard tenant is a schema, this is seconds, not a build. Two isolation tiers exist, and the tier is a commercial choice, not a code fork — the provisioner just targets differently:

Tier Isolation Provision For
Standard (self-serve default) Own schema in the shared cluster, isolated by schema boundary + RLS Seconds, fully automated Most customers
Dedicated (enterprise upgrade) Own database or cluster Minutes, still automated, higher cost Customers who require physical isolation / residency

Offboarding is the inverse and equally scripted: on cancellation the Org moves to suspended (read-only, grace window), then deprovisioned — export the tenant’s data, drop the schema, release the subdomain. Nothing is entangled with another tenant, so teardown is clean by construction.

The registry carries exactly one status per Org, and it is the value everything else reads: the router refuses a host whose Org is not live, the kernel checks it at auth and at routing, and the entitlements the org runs under are resolved against it. Nine states, and three of them exist because provisioning is a multi-step operation that can stop in the middle. A status set with no failure states forces a half-provisioned tenant to be recorded as something it is not, and every recovery decision after that is made against a lie.

Status What it means Entered from Leaves to
pending Signup recorded. Nothing allocated — no schema, no subdomain binding, no subscription. (signup) provisioning, abandoned
provisioning The provisioner is running. A schema may exist; routing is not bound, so the org is unreachable. pending, provision_failed trialing, active, provision_failed
provision_failed A provisioning step failed and its retries are exhausted. Partial allocation exists and is unreachable. provisioning provisioning (retry), deprovisioned (rollback)
abandoned Signup never completed verification or checkout within the window. Nothing was allocated. Terminal. pending (terminal)
trialing Live, on trial entitlements. provisioning, suspended active, past_due, suspended
active Live, on paid entitlements. provisioning, trialing, suspended, past_due past_due, suspended
past_due Payment failed. Fully usable inside the grace window, with a banner. active, trialing active (recovered), suspended (dunning exhausted)
suspended Read-only or blocked. Data retained. past_due, active, trialing active, trialing (reinstated), deprovisioned
deprovisioned Exported, schema dropped, subdomain released. Terminal. suspended, provision_failed (terminal)

Two writers, and they never overlap. The provisioner owns pending, provisioning, provision_failed, and abandoned; the billing provider’s webhooks drive everything from trialing onward. A subscription event cannot move an Org that has not finished provisioning, and the provisioner cannot move one that has reached trialing. The one crossing point is the “payment/trial confirmed” event, which is the only thing that moves pending → provisioning.

A partial provision is unreachable, not half-usable. Routing is bound in the penultimate provisioning step and status flips to live in the last, so an Org in provisioning or provision_failed resolves to nothing at the router. Nobody lands in a broken org, and no first admin is invited into a schema that is missing half its template — the invitation is sent by the seed-identity step, which runs before wiring and after the stamp, so an invitation only exists if the metadata behind it does.

Retry is resumption, not restart. Each provisioning step records its completion in the registry and is idempotent, so a retry re-enters at the first incomplete step rather than re-stamping a schema that already holds the template. Three automatic attempts run with backoff; after that the Org parks in provision_failed and raises an operator alert. Nothing retries silently forever, because a provisioner that never stops trying is indistinguishable from one that is stuck.

Rollback from provision_failed is offboarding with nothing to export. It drops the schema if one was created, releases the reserved subdomain, and cancels the subscription. There is no customer data, so the export step is skipped rather than run empty — which is also why provision_failed → deprovisioned is safe to automate on an operator’s single action rather than requiring the grace window a live tenant gets.

abandoned is a sweep, and it is not a deletion. An Org that sits in pending for thirty days without completing verification or checkout is swept to abandoned; the reserved subdomain returns to the pool and the registry row is retained, because the velocity and fraud screening in the trust chain reads exactly this history. Nothing re-enters pending: someone who signs up again mints a new Org row, so an abandoned attempt can never become a live tenant by a route that skipped a check.

“How do I trust that” has a structural answer and a procedural one.

Structural — isolation makes a stranger safe. A new tenant is a schema, reachable only through the router, and every query is fenced by schema boundary + row-level security. A brand-new, unvetted signup cannot reach another tenant’s data by construction — so onboarding someone you have never spoken to does not risk existing customers. The blast radius of an abusive signup is their own empty schema.

Procedural — layered signals gate the flow:

  • Email verification — proves control of the address before anything provisions.
  • Payment method — a valid card (even on a trial) is a strong anti-fraud and identity signal, and it is the monetization path anyway.
  • Trial constraints — un-paid/developer orgs are provisioned with tighter entitlements: row caps, no production promote, no custom domain, a feature subset, and an expiry. These are enforced as the org’s plan entitlement (metadata), read by the kernel — not bolted on per org.
  • Velocity / fraud screening — rate-limit signups per IP/domain/card; flag disposable-email and known-abuse patterns before step 4.
  • Optional domain verification — for orgs that want SSO or a verified company identity, prove control of the email domain.

The trust chain is deliberately ordered so the expensive step (provisioning) happens only after the cheap checks (email, payment, screening) pass.

One Org, one Billing Account, linked 1:1 in the registry at signup. The subscription — plan + seat count — is the source of truth for what the Org may do, and the billing provider (Stripe is the reference implementation: hosted checkout, subscriptions, tax, dunning) drives it by webhook:

Subscription event Org status Kernel behavior
Trial started / payment succeeded trialing / active Full access at the plan’s entitlements
Payment failed past_due Grace window + in-app banner; still usable
Dunning exhausted suspended Read-only or blocked, data retained
Canceled deprovisioned (after export + grace) Offboarding runs

Billing status is checked by the kernel at auth and routing, so enforcement is uniform and immediate — a lapsed subscription changes behavior on the next request, the same way a metadata generation flip takes effect. Seats are the count of active users (permission-set assignments); exceeding the plan’s seat count blocks new assignments rather than silently over-billing. The plan → entitlement mapping is metadata, so introducing or repricing a plan is a template/entitlement change, not a code deploy.

A plan is a set of entitlements, and an entitlement is either a capability (this plan may promote to production; this plan may use a custom domain) or a meter. Capabilities are booleans and need no further explanation. Meters are the numbers an administrator watches, and there are seven of them. The set is closed: every limit the plan meters appears on the org’s usage surface, including the ones nowhere near their ceiling, and a limit that appears nowhere on that surface is not metered.

Meters come in two shapes, and the shape decides what happens at the ceiling.

  • A stock is a level that stays where it is put — seats, rows, bytes, provisioned environments. Consumption is against an allowance, and the allowance does not refill on its own, so reaching it blocks. Retrying cannot clear it; only deleting something or raising the allowance can.
  • A flow is consumption against a budget over a window that rolls — requests, job runs, assistant usage. Reaching the ceiling throttles, because the condition clears by itself as the window advances, and failing work that would succeed sixty seconds later is the wrong answer to a temporary condition.
Meter Shape Measured as At the ceiling Raise inside the plan?
Seats stock Active users holding at least one permission-set assignment New assignments and activations are blocked. Existing users are untouched. Yes — seats are a subscription quantity; adding them is a billing change, not a plan change
Records stock Rows across every object in the tenant schema Inserts blocked; updates and deletes continue, so the org can always work its way back under Yes, in purchasable increments up to the plan’s maximum
File storage stock Bytes of uploaded content. Renditions are a cache and are not metered Upload blocked; reads, downloads, and deletes continue Yes, in purchasable increments up to the plan’s maximum
Environments stock Concurrently provisioned non-production environments A deploy that would create one more fails with a limit error naming the meter; existing environments keep running No — the ceiling is a property of the isolation tier; more requires a plan change
API requests flow A per-minute burst allowance over a rolling 24-hour budget Throttled with a retry-after; requests are delayed, never dropped A time-boxed uplift can be granted for a migration window; a permanent increase is a plan change
Background jobs flow Concurrent run slots plus runs per 24 hours Runs stay queued rather than failing; at the daily budget, enqueue applies backpressure and the queue drains at the plan’s rate Same — uplift for a window, plan change for permanent capacity
Assistant usage flow Consumption against a monthly budget Bills as overage rather than blocking. A tenant may set a hard cap, which blocks instead — the choice of which failure it prefers is the tenant’s Yes — the budget is a number the administrator raises, because exceeding it costs money rather than capacity

Four rules hold across all seven.

The number is platform-owned; the request is not. An administrator can see every meter, see its ceiling, and request an increase. They cannot edit a ceiling, because the ceiling is the plan’s entitlement metadata and the plan is the commercial agreement. An increase inside a plan is a subscription quantity change applied by the same billing webhook that applies a seat change, so a raised limit lives in the entitlement metadata like every other limit and never as an exception recorded somewhere else.

Enforcement and display read the same counter. The value on the usage meter is the value the kernel compares against when it blocks or throttles. A usage page computed from a different source than the enforcement path eventually disagrees with it, and an administrator who has been told they have headroom and is then blocked has no way to tell which number was wrong.

Every meter names its consequence, and warns before it. Stocks raise a notification to the org’s administrators at 80% of allowance; flows raise one on the first throttle event in a window rather than at a percentage, because a flow that briefly touches its ceiling and recovers is normal and a flow that touches it repeatedly is not. Nobody discovers a ceiling by hitting it.

Trial entitlements are the same meters with different numbers. A trial org is not a different product with a separate cap system bolted on; it is an Org whose plan entitlement sets low values and switches some capabilities off. That is what makes “convert a trial to paid” a subscription event rather than a migration.

Most of this reuses what already exists — the deploy pipeline stamps the template, permission bundles seed access, RLS provides isolation, the generation flip is the enforcement model. The one net-new kernel surface is the control plane: the tenant registry, the provisioner, the subdomain router, and billing-status enforcement. It is small, it is cross-tenant, and it is the correct place for the answer to every “where does this live” question above — never inside a tenant’s own metadata.

Commercial parameters (set per business decision, not architecture)

Section titled “Commercial parameters (set per business decision, not architecture)”

These knobs are yours to set and can change without touching the mechanism:

  • Plan tiers, pricing, and seat definitions — how many plans, at what price, with what entitlements. The seven meters are the mechanism; their values per plan are a price-list decision.
  • What a failed provision returns — refund or credit when a charge was taken before a provision that rolled back. The rollback always cancels the subscription; whether the money comes back as a refund or a credit is a policy setting on the billing adapter.
  • Trial shape — card-required vs. card-optional, trial length, and the exact trial caps (rows, feature subset).
  • Contract terms — the click-through agreement presented at checkout.
  • Isolation-tier pricing — what a Dedicated instance costs and who qualifies.
  • Billing provider — Stripe is the reference implementation; the control plane treats it as a swappable adapter behind the subscription-event contract.