Skip to content

Environments

An environment is a named place an org’s metadata and records live and run — a tenant schema plus the serving that interprets it. It is not a running application in its own right, and it does not arrive with one. A freshly provisioned environment is bare kernel: the engine, a schema, and the only two surfaces the kernel bakes in — a sign-in screen and an empty-workspace landing — and nothing else a user would recognize as an app. The shell, the setup app, and every business app are package-delivered metadata, and the capabilities an org exposes are the capabilities its installed packages claim. An environment therefore governs where an org’s metadata and data live and where its logic runs; it is an infrastructure/metadata construct, not the application that runs inside it. This is the same engine/UI separation the whole platform is built on — see how the platform works.

Because it is infrastructure rather than business logic, an environment has no pure/effectful tier of its own; it is described instead by a point in a four-axis space: origin (what it was seeded from), lifecycle (how long it lives and how it dies), data-policy (masking, subsetting, and synthetic fill), and authority (who owns truth and how click-made changes flow back). There is one environment primitive configured along four orthogonal knobs, not four fixed environment products.

Provisioning, and what populates an environment

Section titled “Provisioning, and what populates an environment”

An environment starts empty and is filled by provisioning. Provisioning creates an org from an org shape — a starter bundle of packages plus the default metadata and starting configuration a fresh org of that kind begins with (object model) — and homes the org on a home instance in a region. The shape is what turns a bare schema into a working system: its packages install a shell, a setup app, and apps; its seed metadata and config populate the org’s objects. Until that install finishes, the environment shows the kernel’s empty-workspace landing — the honest “kernel online, nothing installed yet” state, not an error and not the finished product.

Two moments are worth keeping distinct, and both are ordinary metadata operations rather than anything special to environments:

  • Standing up the place — allocating the schema, wiring routing, resolving the home instance and region. This is the onboarding/provisioning path; an environment is where that stamp lands.
  • Stamping the shape — running the org shape’s package-and-metadata bundle into the new schema through the ordinary retrieve → diff → apply pipeline. What a new environment receives is not hard-coded; it is a versioned shape run through the same commit any deploy uses, so improving the starting point is a metadata change, not a code release.

The consequence is that “which apps and UI an environment has” is never a property of the environment primitive. It is a property of what has been installed into it — swap the shape, install a different shell, add or remove packages, and the same environment presents a different system with no change to the environment definition or the engine. The axes below govern the place; packages and the shape govern what runs there.

An environment is defined by four axes that vary independently. Salesforce collapses the same design space into a small set of fixed SKUs (Developer, Developer Pro, Partial Copy, Full sandbox; scratch org) each of which pins several axes at once; CAOS keeps them separable.

Axis What it controls Values
origin What the environment is seeded from at create time blank | from_repo | from_snapshot | from_live
lifecycle How long the environment lives and how it is torn down / kept current ephemeral (TTL, auto-deleted) | durable (persists, reconciled or refreshed)
data_policy How record data is protected, sampled, and filled mask (default mask-all) + subset (graph-coherent) + synthetic (fill)
authority Who owns configuration truth, and how in-env edits are captured back owner: repository + reverse_integration: continuous | confirmed | off

Two facts anchor the model:

  • The repository owns truth in every environment, including production. The metadata an environment runs is a projection of committed source, not the source. This inverts the Salesforce sandbox model, where every environment is copied from production and production is therefore the de-facto source of truth.
  • The safe data state is the default. A lower environment never receives raw PII unless its policy is explicitly downgraded. Masking is not an add-on that must be purchased, installed, and run; it is the default value of the mask knob.

Where the data lives, per origin. blank provisions schema with no records. from_repo builds records from committed seed/fixture data (the scratch-org analog). from_snapshot restores a captured, already-masked point-in-time image. from_live copies from another running environment — and is the only origin that can introduce production PII, so it is gated by the data-policy on the way in, not after the fact.

An environment is declared as a single metadata component with an origin, a lifecycle, a data_policy, and an authority block. The platform provisions the environment from that declaration; nothing about the environment is created by clicking except the record data the policy permits. (What runs inside the environment — the shell, setup, apps — arrives separately, as the org shape’s packages install; see Provisioning, and what populates an environment.)

Config vocabulary:

Key Required Meaning
origin yes One of blank, from_repo, from_snapshot, from_live. Determines seeding. from_snapshot/from_live take a source (env or snapshot id).
lifecycle yes { mode: "ephemeral", ttl: <duration> } or { mode: "durable", reconcile: "continuous" | "on_demand" }. Ephemeral envs are auto-deleted at TTL.
data_policy.mask yes Field-level masking policy. Default mask_all — every field classified as PII is masked unless individually exempted. none is a explicit, logged downgrade, not a default.
data_policy.subset when subsetting { seed_objects, sample, follow_graph: true, budget? } — a graph-coherent sample: start from seed_objects, take sample rows, then walk foreign keys to pull every referenced parent so the subset has no dangling references. Optional budget (a storage ceiling) constrains the seed sample, not the closure — the solver shrinks sample until the coherent set fits, never dropping required parents (see Semantics & evaluation).
data_policy.synthetic no { target_volume, distribution, epsilon? } — generate referentially-valid fake rows to reach a realistic size/shape without copying production data. distribution: match_source matches per-column marginals (not the joint distribution of real rows); optional epsilon sets the differential-privacy budget bounding how much any single source row can influence the output (see Semantics & evaluation).
authority.owner yes repository (the only supported value; the repo is authoritative in every env).
authority.reverse_integration yes continuous (lower envs — click-changes captured to source automatically), confirmed (production — captured only after review), or off.

The masking policy is authored per field via classification, not per environment: a field is tagged (e.g. pii: email), and each masking mode names the transform. Enumerable mask transforms: redact (replace with a fixed token), hash (deterministic pseudonymize — same input → same masked output within a run, preserving joins), shuffle (permute values within the column), synthetic (draw a format-valid fake value), and exempt (pass through — for non-sensitive fields only).

Worked example — a durable test environment seeded from a masked, graph-coherent 2% sample of production, topped up with synthetic rows, with click-changes captured continuously back to the repo:

{
"key": "test",
"label": "Test",
"type": "environment",
"body": {
"origin": { "mode": "from_live", "source": "prod" },
"lifecycle": { "mode": "durable", "reconcile": "continuous" },
"data_policy": {
"mask": "mask_all", // safe state is the default
"subset": {
"seed_objects": ["invoice"],
"sample": 0.02, // 2% of invoices …
"follow_graph": true // … plus every FK parent, kept intact
},
"synthetic": {
"target_volume": { "invoice": 5000 }, // fill to realistic scale
"distribution": "match_source" // shape without copying rows
}
},
"authority": {
"owner": "repository",
"reverse_integration": "continuous" // prod would be "confirmed"
}
}
}

Production is the same primitive with a stricter authority knob: origin: blank (built from the repo, never copied from elsewhere), lifecycle: durable, data_policy.mask: none (it is the real data), and authority.reverse_integration: confirmed.

When provisioning happens. Origin seeding runs once, at environment create. For from_live/from_snapshot, masking and subsetting are applied on the way in — a record that is masked in the target env is never written to the target in the clear, so there is no window in which raw PII lands and is later scrubbed. This is the structural difference from an install-and-run masking add-on, where the unmasked copy exists first.

Subset coherence. follow_graph: true makes subsetting a graph traversal, not a LIMIT. After sampling the seed rows, the sampler walks every outbound foreign key (including required lookups and master-detail parents) and pulls the referenced rows transitively, so the resulting set is closed under reference — no orphaned children, no dangling parents. Cycles, self-joins, and polymorphic references are handled by the traversal, not left to chance.

Subset budget. A subset must be both graph-coherent and within a storage budget, and coherence wins: the solver never drops a required parent to fit, because that reintroduces the orphans the traversal exists to prevent. Instead it treats the budget as a constraint on the seed sample, not on the closure. The build samples sample seed rows, computes the transitive FK closure, and measures its footprint; if the closure exceeds subset.budget, the seed fraction is reduced and the closure recomputed, converging on the largest coherent subset that fits. Only foreign keys the policy marks non-essential (nullable lookups flagged follow: false) are pruned from the walk; required and master-detail edges are always followed. If a single seed row’s mandatory closure alone exceeds the budget, the build fails loudly rather than emitting an incoherent set — an unsatisfiable budget is a configuration error, surfaced through the canonical error envelope, not silently truncated.

Masking determinism. hash masking is deterministic within a run: identical source values map to identical masked values, so a masked email used as a natural join key still joins. Determinism is scoped to the environment build, not global, so the same source value in two different environments need not share a masked image. Masking is applied field-by-field from classification; a field with no classification and no explicit exempt is treated as sensitive and masked by default (fail-closed).

Masking verification. Masked-by-default is a guarantee, not merely a default, because every non-production build ends with a verification pass that must succeed before the environment activates (before the single metadata-generation flip that makes it live). For every field classified as PII, the pass asserts two properties over the seeded rows: the value is not equal to the source value (for from_live/from_snapshot origins, proving the transform actually fired) and it conforms to the declared transform’s output domain (a redact field holds the fixed token, a synthetic email is format-valid but unlinkable, a hash field is the pseudonym, not the plaintext). The pass also closes the three leak vectors an add-on masker misses: free-text fields classified as possibly-PII must carry an explicit redact or synthetic transform or the build fails closed; computed/roll-up fields that derive from a masked source are recomputed from the masked inputs, never copied; and encoded-ID fields that embed PII (a customer number containing a tax ID, for instance) are classified as PII themselves and masked as a unit. A failed assertion aborts the build and reports the offending field through the canonical error envelope — the unmasked env never activates.

Reverse-integration timing and conflicts. A configuration change made by clicking inside an environment is captured back to the repository as a diff. In continuous mode (lower envs) the capture is automatic and immediate; in confirmed mode (production) the change is recorded as pending capture and merged into source only after review. Between the click and the merge, production and repo differ, and the system represents that gap explicitly rather than pretending truth is already synced.

Capture is a field-level three-way merge, never last-write-wins. The common ancestor is the source revision the environment was last reconciled to (its captured baseline for that component); the two sides are the in-env edit and any repo edit landed since. Changes to disjoint fields of the same component merge automatically; two edits to the same field of the same component are a genuine conflict. A conflict is never resolved by clobbering: in continuous mode the non-conflicting parts merge and the conflicting field is quarantined as pending capture for review; in confirmed mode every capture is review-gated regardless. This is why capture-back cannot silently eat work the way a sandbox refresh does — the only work ever held back is a field two actors changed at once, and it is surfaced, not discarded.

Synthetic-fill realism and privacy. synthetic fill reproduces production’s shape without reproducing its rows. It matches per-column marginal distributions plus any declared cross-column constraints (a ship_date after its order_date), and never interpolates between or copies real records. The privacy bound is differential privacy: generation runs under a configured ε budget, so no single source row measurably changes the output, which bounds linkage and membership-inference risk quantitatively rather than by inspection (differential privacy for synthetic data). Quasi-identifier columns are generated within that mechanism rather than jointly sampled from source, which defeats the demographic re-identification vector — the combination of five-digit ZIP, sex, and date of birth uniquely identifies an estimated 87% of the US population (Sweeney, Simple Demographics Often Identify People Uniquely), so joint fidelity on those columns is exactly what must not be preserved.

Determinism & reproducibility. An environment built with origin: from_repo (or blank) is reproducible from committed source: the same commit rebuilds the same schema and the same seed/synthetic data (synthetic generation is seeded), and — because the org shape is itself versioned metadata — the same shape reinstalls the same starting apps and config. from_live/from_snapshot environments are reproducible only relative to the captured source image, which is itself a moving target — a distinction the origin knob makes explicit.

An environment is where a tenant’s work runs; platform status is the platform’s own statement about whether that place is working. It is a first-class surface rather than a marketing page, and it is defined by four things: what it decomposes into, what states it can be in, how it is scoped, and where its data comes from. One rule sits above all four and is stated first because everything else bends around it: status is never authoritative over an error the tenant’s own request returned.

“Is the platform up” is not answerable, because the platform is not one thing. Status is stated per service — a named subsystem a tenant can reason about, each with its own health, each corresponding to a capability this guide defines elsewhere:

Service Covers A tenant notices when
identity Login, SSO, token issue and refresh, session policy Nobody can sign in, or sessions drop
data The record read and write path — queries, saves, the save order Records are slow to save or fail to load
metadata The component catalog, authoring surfaces, retrieve → diff → apply Setup is slow, or a deploy will not validate
deploy Validate, apply, promote, and the generation flip A deploy queues and does not start
jobs Background work — enqueue, claim, chunk, retry Queue depth climbs and nothing drains
files Upload, download, rendition, retention Attachments will not open
search Indexing and query for global and list search Results are stale or empty while records exist
notifications Delivery across every channel Alerts stop arriving
integrations Outbound callouts, the outbox, inbound endpoints External syncs back up
reports Report and dashboard execution, snapshots, subscriptions Dashboards time out while records load fine
ai The assistant, model routing, evals The assistant is unavailable while everything else works

Eleven services, and the decomposition is deliberately the same one the rest of this guide is organized by. A tenant reading “notifications degraded, data operational” knows exactly which of their own behaviors to expect to be wrong, which is the only thing a status page is for. A single global “all systems operational” light is not a smaller version of this — it is a different, less useful claim.

An incident is a record, not a banner. It carries an identity, a lifecycle, a scope, and a history, and it is the unit that status is composed from.

Two independent axes describe impact, and keeping them separate is what prevents “degraded” from meaning four things:

  • Impact kind — degradation (the service works and is slow, partial, or lossy) or disruption (the service does not work).
  • Severity — minor or major. Severity is about breadth and consequence, not about how long the fix takes.

Lifecycle is a fixed, forward-only sequence, and the state names are the ones an on-call engineer would use:

State Means Set by
investigating Impact is confirmed and the cause is not yet known. An operator, or automatically when a health probe crosses its threshold
identified The cause is known and a fix is in progress. An operator only
monitoring The fix is applied and the platform is watching for recurrence. An operator only
resolved Impact has ended. An operator only
postmortem Root cause and corrective actions are published against the resolved incident. An operator only

The asymmetry in that last column is the design. A machine signal may open an incident; only a person may close one. A probe that recovers proves the probe recovered, not that the tenant’s problem ended, and auto-resolution is how a status board learns to lie during a partial recovery. postmortem is a state rather than an attachment because an incident that never reaches it is visibly unfinished.

Maintenance is not an incident and never renders as one. A maintenance window is a separate record with scheduled → in_progress → complete, a planned start and end, the services and regions it touches, and an availability statement — fully_available, partially_available, or unavailable. Merging planned and unplanned into one stream trains readers to ignore both.

A service’s rolled-up state is derived, never typed: operational when it has no open incident; degraded or disrupted from the most severe open incident’s impact kind; maintenance while an unavailable or partially_available window is in progress; and unknown under the rule below. Derivation is one direction only — nobody edits a service’s colour, they open or close an incident.

Regional scoping, and what a tenant is shown

Section titled “Regional scoping, and what a tenant is shown”

Status is scoped on two dimensions, and a tenant is shown the intersection of them rather than the whole board.

Region. Every tenant schema lives in a named region, and services are operated per region. An incident carries the regions it affects; a tenant in one region does not see another region’s outage rendered as their own. affectsAll exists as an explicit flag for a genuinely global control-plane failure, and it is a claim an operator makes rather than a default.

Environment. A tenant’s production and its lower environments can run on separately-provisioned capacity, so an incident carries the environments it touches. “Test is disrupted, production is operational” is a common and important true statement, and a status model that cannot express it makes every test-environment problem look like an emergency.

The tenant-facing surface therefore lists one row per service per region-and-environment the tenant actually occupies, with a link to the full board for anyone who wants the platform-wide view. A tenant is never shown a red row for a place they do not run.

Status is a control-plane capability, not a tenant one. It is assembled from three inputs:

  1. Health probes, run per service per region by the control plane, measuring the same paths tenants use rather than a process liveness check. A probe crossing its threshold opens an incident in investigating.
  2. Operator-authored incident records, which are the only way an incident advances or closes.
  3. Maintenance components, which are declared in advance and deployed like any other component, so a planned window exists in source before it exists on a board.

Two constraints on that assembly matter more than the mechanics. The status store does not share a failure domain with the platform it reports on — it is separately provisioned, separately deployed, and reachable when the thing it describes is not, because a status page that goes down with its platform reports nothing at exactly the moment it is needed. And the status surface is readable without authentication for the public board and without any system permission for the in-app one: a person who cannot log in is the person most likely to need it, and gating status behind the login it is reporting on is a closed loop.

A status board is a summary of many tenants. An error envelope is evidence from one request. When they disagree, the envelope wins, and the surface says so rather than reconciling them away.

  • An envelope is never contradicted by a board. A request that returned class: internal with a correlationId failed. It failed whether or not any service is green, and the in-app status surface never renders a green service as a reason to disbelieve a specific failure the tenant just saw.
  • The tenant’s own recent error rate renders beside the board, drawn from their own envelopes — internal, limit, and integration classes, by service. A green board next to a nonzero own-error rate is displayed as a conflict, with both numbers and a way to report it. That combination is a real and frequent state: a fault affecting one tenant, one region shard, or one code path is invisible in an aggregate and total for the person hitting it.
  • The correlationId is the join key. An incident may list the correlation ids it explains, so a tenant holding an id can ask “is my failure this incident” and get an answer instead of a judgment call. Resolving an incident does not retroactively make the envelopes it explains disappear — they remain in the tenant’s own trace, now annotated with the incident.
  • The direction of causation is one-way. Envelopes feed incident detection; incidents never suppress, downgrade, or rewrite envelopes. There is no path by which declaring an incident resolved changes what a tenant’s request returned.

Unreachable reports unknown, never healthy

Section titled “Unreachable reports unknown, never healthy”

If the status service cannot be reached, or its most recent successful reading is older than its freshness budget, every service reads unknown — with the age of the last successful reading stated beside it, and the last known state shown as history rather than as current.

This is the one behavior that decides whether a status surface is worth reading. Fail-open — showing green when the reporter is unreachable — makes the healthy state and the blind state indistinguishable, which is precisely the moment they most need distinguishing. A cached green is never re-served as current: it is re-served, plainly labelled, as a reading from a stated time in the past.

The status surface’s own reachability is a separate fact from the platform’s, and it is rendered separately. “Status service unreachable” and “platform disrupted” are different sentences, and a reader must be able to tell which one they are being told.

Concern CAOS approach Salesforce’s exact limit / behavior (cited) The WHY behind the SF limit Does the CAOS stack still need an equivalent?
Refresh / reconcile cadence Repo-driven reconcile is continuous and unthrottled (it is cheap); a full from_live re-copy is rate-limited, the interval set by deployment policy against the infra budget rather than a fixed platform constant Refresh gated by type: Developer 1 day, Dev Pro 1 day, Partial Copy 5 days, Full 29 days Each refresh re-clones from prod; heavier copies (Full = entire prod) are throttled to bound copy-job cost Split — the reconcile half needs no gate; the re-copy half keeps one, but as a budget-driven policy knob, not a per-SKU constant
Storage per env Sized to the subset + synthetic fill the policy declares, on Postgres Developer 200 MB, Dev Pro 1 GB, Partial Copy 5 GB, Full = production data storage Fixed SKU tiers force “sample, don’t clone” at the Partial/Dev levels Yes in spirit — an env still has a storage footprint, but it is a function of the declared policy, not a SKU tier
Subset size Graph-coherent sample by fraction or count; closure over references guaranteed Partial Copy takes the first 10,000 records per selected object, and does not traverse to unrelated related records Bounds the sample cheaply; the cost is a referentially incoherent subset (orphans) No — coherence is the point of the follow_graph traversal; the CAOS cost moves to compute, not correctness
Data protection Masked by default; raw PII requires an explicit, logged downgrade Data Mask is a paid, separately-licensed add-on (a managed package); sandboxes are not masked by default Masking was built as an installable product, not a platform default No — masking is a default policy knob, not a purchased add-on; the equivalent “need” is that masking be correct, not that it be bought
Environment lifespan ephemeral with author-set TTL, or durable Scratch orgs: 30-day max, 7-day default, 1–30 selectable, then deleted Enforces ephemerality; prevents long-lived “pet” scratch orgs Yes — ephemeral CI envs still need a TTL and auto-teardown; the mechanism is the same, the numbers are policy
Concurrency / creation rate Governed by infra budget, not edition SKU Scratch orgs by Dev Hub edition: 3 active / 6 daily (Dev), 40 / 80 (Enterprise), 100 / 200 (Unlimited/Performance), rolling 24 h Free/edition tiers throttle CI fan-out and copy load Yes — env creation is a real resource cost and needs quotas; CAOS ties them to infra budget rather than a license edition
Loss of in-env work Reverse-integration captures click-changes to source (continuous / confirmed) A sandbox refresh overwrites anything not deployed back to prod; undeployed clicks are lost, and the env gets a new Org ID The copy-down model treats a sandbox as a lease on a snapshot, not a durable branch No — capture-back is the design goal, not an afterthought; the cost moves to conflict handling (a field-level three-way merge; see Semantics & evaluation)

The number of non-production environments an org may run concurrently is itself a metered plan entitlement, not a property of the environment primitive — see what the plan meters.

Salesforce splits the environment design space into two disjoint product families with different tooling.

  • Sandboxes are copies made from an existing org (ultimately production). Four types, differing only in record data:
    • Developer and Developer Pro copy all metadata but zero records — metadata-only shells (200 MB and 1 GB data storage; both refreshable once per day). Any data must be loaded or generated.
    • Partial Copy copies metadata plus a template-selected sample of up to 5 GB, capped at 10,000 records per selected object (the first 10,000), refreshable every 5 days. Because the sampler takes the first N rows and does not traverse to unrelated related records, the subset is referentially incoherent — orphaned children and dangling references unless the template is hand-curated.
    • Full copies everything — all records and attachments — at production storage size, refreshable every 29 days. It is production data, PII included.
  • Scratch orgs are built from source: a JSON definition file (config/project-scratch-def.json) declares edition, features, and settings, and source tracking syncs local files into the org. This is genuinely repo-as-truth — but scratch orgs are ephemeral (≤30 days, 7 default), data-empty, and not customer-facing, so they cover dev/CI only. Active/daily creation is throttled by Dev Hub edition (3/6 up to 100/200; partner PBOs 150/300).

Two structural weaknesses follow, and CAOS targets each with a named mechanism.

  • Data protection is opt-in and paid. Salesforce Data Mask is a managed package requiring a separate add-on license; masking jobs run only in sandboxes, after the copy. A Full or Partial sandbox therefore lands real production PII into a lower environment as-is, and it stays that way until someone buys, installs, configures, and runs Data Mask.
    • Genuinely better (mechanism): CAOS masks on the way in, from per-field classification, fail-closed by default — a lower env never holds the unmasked copy, and the safe state costs nothing extra.
    • Cost/risk: masking correctness is safety-critical and hard. Format-preserving masks can re-identify (a rare ZIP + birthdate), and free-text, computed, and encoded-ID fields are leak vectors. A masking bug is a data breach, not a cosmetic defect — so masked-by-default is enforced by a post-seed verification pass that aborts activation unless every PII-classified field is proven transformed, closing the free-text, computed, and encoded-ID vectors by construction (see Masking verification under Semantics & evaluation). The residual cost is that the guarantee is only as good as the field classification: an unclassified sensitive field is caught by fail-closed masking but a mis-classified one (tagged non-PII) is not, so classification review is the standing discipline.
  • Subsetting is incoherent, and truth is split-brain. Partial Copy’s first-10k-per-object sampling breaks referential integrity; and Salesforce runs two truth models at once — repo-truth for scratch orgs, org-truth for production and sandboxes — with a manual, lossy bridge between them (a click in production is not captured unless someone retrieves and commits it; a refresh silently discards undeployed sandbox work).
    • Genuinely better (mechanism): one env primitive with a graph-coherent subset (FK-closure traversal) and repository-owns-truth in every environment plus reverse-integration — continuous capture in lower envs, review-gated capture in production — which closes the manual-bridge gap by construction.
    • Mere parity: copy-from-a-live-env, build-from-repo, ephemeral and durable lifecycles, and deploy-from-repo to any env are table stakes; CAOS must match them, and does not claim novelty for them.
    • Cost/risk: conflict handling is the whole ballgame. A click-change in production and a repo edit to the same component is a merge conflict on live configuration; CAOS resolves it with a field-level three-way merge against the last-captured baseline (never last-write-wins), so disjoint edits merge and only a same-field collision is held back — capture-back cannot eat work the way a sandbox refresh does. Continuous capture still has a running cost — every lower-env click produces a diff to compute, store, and (eventually) reconcile.

Platform status is instance-shaped, and the tenant has to know their instance. Salesforce publishes status at status.salesforce.com, keyed on the instance an org happens to live on. An instance document carries a status ("OK" on every instance sampled), a location ("NA", "EMEA", "APAC"), an environment ("production"), a releaseVersion, a recurring maintenanceWindow as free text — "Saturdays 07:00 PM - 11:00 PM PST" — and arrays of Services, Products, Incidents, Maintenances, and Tags. The service decomposition is real and reasonably fine-grained: one sampled production instance lists fourteen services including coreService, search, analytics, Communities, liveAgent, and ServiceCloudVoice.

Incidents are records with a lifecycle. An incident carries status, type ("Degradation" or "Disruption"), instanceKeys and serviceKeys for scope, an affectsAll flag, IncidentImpacts with a severity of "minor" or "major", a message object with rootCause, actionPlan, and pathToResolution fields, and a timeline of impact_start / impact_end / event entries whose event types include update, resolved, and startTimeRevision. Maintenances are a separate record type with type: "release", status: "Confirmed", releaseType: "Major", planned start and end times, and a message carrying availability: "fullyAvailable" and eventStatus: "confirmed".

  • Genuinely better (mechanism): status is scoped to the tenant’s own regions and environments rather than to an instance identifier the tenant must first look up, so nobody has to know what instance they are on to know whether they are affected. The error envelope is authoritative over the board and a green-board-with-own-errors state renders as a stated conflict joined by correlationId, instead of leaving a tenant to argue with a status page. Detection may open an incident and only an operator may close one, so a recovering probe cannot mark a partial recovery resolved. And unreachable reports unknown, so the blind state is never drawn as the healthy state.
  • Mere parity: per-service decomposition, an incident record with a timeline and a severity, planned-maintenance windows separate from incidents, and a public unauthenticated board are table stakes. Salesforce does all four, and the status API’s shape is a reasonable model to follow.
  • Costs / risks (named): a separately-provisioned status store is real infrastructure that must be operated, patched, and paid for precisely so it can outlive the platform it describes — the value is entirely in the failure case, which makes it perpetually easy to under-resource. Operator-only closure means incidents stay open longer than the outage did, and a board that lags reality in the safe direction still erodes trust. Rendering a green-board-versus-own-errors conflict surfaces every single-tenant fault as a visible disagreement, which is honest and will generate support contacts that an aggregate board would have absorbed silently. And per-region, per-environment scoping multiplies the rows an operator must keep accurate; a decomposition nobody maintains degrades into eleven services that are always green.

An environment is one canonical component. All four axes are declared in its body; nothing about the environment is emergent across separate products or CLIs.

{
"key": "test",
"label": "Test",
"type": "environment",
"body": {
"origin": { "mode": "from_live", "source": "prod" },
"lifecycle": { "mode": "durable", "reconcile": "continuous" },
"data_policy": {
"mask": "mask_all",
"subset": { "seed_objects": ["invoice"], "sample": 0.02, "follow_graph": true },
"synthetic": { "target_volume": { "invoice": 5000 }, "distribution": "match_source" }
},
"authority": { "owner": "repository", "reverse_integration": "continuous" }
}
}

Retrieve → diff → deploy operates on this single object. A diff that changes data_policy.mask or subset re-provisions how the env is seeded on its next build; a diff that changes authority.reverse_integration changes capture behavior; the environment’s record data is never part of the component (it is data governed by the policy, not schema), and so is the UI that runs inside it — apps, the shell, and setup are package metadata installed into the environment, not fields of the environment component (see Packaging and The Exchange). Because the repo is authoritative, deploying the component is how an environment comes to exist or change — there is no out-of-band “create sandbox” click that the repo does not know about.

Salesforce Metadata API analog (for migration/parity mapping): there is no single environment component. A scratch org is described by the config/project-scratch-def.json definition file consumed by the CLI (sf org create scratch), not by the Metadata API. Sandboxes are created and refreshed through Setup or the Tooling API (SandboxInfo/SandboxProcess), separate from metadata deploys; masking is configured inside the installed Data Mask managed package; and moving config between environments uses change sets, the Metadata API (sf project deploy), packages, or DevOps Center — none of which is the environment definition itself.