Skip to content

Errors & diagnostics

CAOS does not pre-enumerate every possible failure. It enumerates the finite set of surfaces where code runs and places one normalizer boundary at each. A normalizer catches whatever happens at its surface and emits a single canonical envelope: a known failure becomes a typed code; an unknown failure becomes class: internal with a correlationId, and its raw text is captured to telemetry and never returned. The consequence is a construction guarantee, not a testing outcome: no raw Postgres, kernel, or runtime error can reach a user, ever — safety does not depend on having anticipated the specific error.

An error in CAOS is described on three orthogonal axes. Keeping them independent is what makes “it doesn’t matter where the error came from” literally true — a new surface plugs in without touching anything downstream.

  1. Origin / surface — WHERE it arose: the normalizer that caught it. The surfaces are a fixed set: authoring, deploy.apply, save.shape, save.adjust, save.validate, save.write.constraint, save.write.access, save.write.rollup, save.write.gate, save.effects, commit, post_commit, read.query, calc.sandbox, integration.callout, api.request, kernel. These map onto the fixed save order of execution plus the read, deploy, authoring, calc, integration, and API boundaries.
  2. Class — WHAT kind of failure it is: one of the small fixed taxonomy classes (below).
  3. Location — WHERE it surfaces to the user, which drives the UI target: field, record, toast, setup, deploy_result, api, async_status, or global.

The normalizer at an origin sets class, code, raw, and details; the location decides rendering. This separation is the design’s load-bearing property: a new origin only needs a normalizer that emits the envelope — nothing else changes. Routing is dynamic because the three axes vary independently; the same class (validation) can arise at several origins and surface at several locations without any of them knowing about the others.

Boot is a surface too. A failure while an organization boots — it cannot be identified, something installed in it fails its checks, or part of it does not load — reaches the error screen as the same envelope, from the boot path’s own normalizer, under catalogued codes: not_found.org, validation.bundle_component and internal.boot_degraded. The workspace normalizes the failures only it can see the same way: integration.engine_unreachable when the Engine cannot be reached, and internal.surface_unregistered when nothing can draw the surface the Engine chose. The screen renders the envelope’s fault and retriable flag rather than guessing them, so it offers to try again only when trying again can help.

Every error — at every surface, in every API, in the UI — is the same shape. The envelope has a machine face (stable, for monitoring/clients/tests), a human face (friendly, editable, localizable), and one dev-only field that is never serialized to an end user.

Field Face Purpose
code machine Stable, namespaced identifier — e.g. validation.required_field, computation.divide_by_zero, deploy.dependency_missing, internal.unexpected. Never changes.
class machine One of the fixed taxonomy classes (below).
message human Friendly, actionable, no jargon/cell-refs/raw text. Rendered from an editable, localizable template keyed by code.
origin machine The surface that raised it (axis 1).
location machine UI target + the offending thing: objectApi, recordId, fieldApi, componentKey, deployComponent. Drives axis 3.
severity both error | warning | info — blocks vs advisory. Lets warnings and info coexist with hard errors in one result.
retriable machine Whether the identical action can succeed on retry (deadlock/timeout = yes; validation = no).
fault both Who resolves it: user | admin | platform — lets the UI say “fix this” vs “ask your admin” vs “we’re on it (ref …)”.
correlationId both Trace id linking the friendly error to the full raw detail in telemetry. Shown for internal.
details machine Structured specifics keyed by api name: failing field, violated constraint, calc-function key + version, missing dependency.
docs machine Optional link to the relevant guide page.
raw dev-only The underlying original error, captured for telemetry. Never serialized to an end-user response.

The location locator is carried as data on one envelope, not expressed as two divergent error types — a field-level error is the same object as a record-level error with a fieldApi populated. Empty locator ⇒ record/page-level. (This mirrors Salesforce’s fields[] array design; see How Salesforce does it.)

Every error is exactly one class; codes are namespaced under the class. The set is small and fixed.

Class What it is Typical fault Typical location
validation A declared rule/constraint the user can fix (required, format, validation rule, uniqueness). user field / record
permission Action not allowed (object CRUD, FLS, Setup capability). admin record / setup
not_found The thing doesn’t exist or isn’t visible (see the security rule below). user record / global
conflict Concurrency/state conflict (stale version, lifecycle gate blocked, duplicate). user / admin record
computation A formula/roll-up/calc failed to produce a value (divide-by-zero, type-contract violation, calc raised, sandbox limit). user / admin field / record
limit A quota/resource/governor limit hit. user / admin toast / async_status
deploy A metadata validate/apply failure (dependency order, destructive-guard, migration, integrity/cycle). admin deploy_result
integration An external system failed (callout, sync, external auth). admin / platform record / async_status
internal The catch-all. Never the user’s fault. correlationId shown; raw hidden. The anti-“raw-DB-error” class. platform toast / global

The origins × locations map. An origin’s normalizer determines the class and code; the location determines what the user sees. The pairing is not fixed one-to-one — one origin can surface at different locations depending on the caller (UI vs API vs async job).

Origin (surface) Common classes raised Surfaces to (location)
authoring validation, computation, deploy setup (author-time, inline on the metadata editor)
deploy.apply deploy deploy_result
save.shape / save.adjust validation, computation field / record
save.validate validation field / record
save.write.constraint validation, conflict field / record
save.write.access permission, not_found record / global
save.write.rollup computation record
save.write.gate conflict record
save.effects computation, integration, internal record / async_status
commit conflict, internal toast / record
post_commit integration, internal async_status
read.query permission, not_found, limit record / global
calc.sandbox computation, limit field / record
integration.callout integration, limit record / async_status
api.request validation, permission, not_found api
kernel internal toast / global

Location decides the component, not the page. The location on the envelope is a UI target, and a fixed set of standard components renders each one — no surface writes its own error handling. field renders inline against the offending field, which keeps the typed value and takes focus back. record renders a ValidationSummary at the top of the owning section — the page-level “here is everything wrong” box — each entry linked to its field. toast is for an outcome with no field or record to attach to, and never for anything that needs a decision. setup renders inline on the metadata editor, deploy_result in the deploy report, api as the response body, async_status on the job, global as a page-level banner. The kernel decides which by emitting the envelope; the shell and these components decide how it looks. Those renderers are delivered by packages, not compiled into the engine — the envelope and the normalizer boundary are the kernel’s, while the components that paint each location arrive with the UI packages installed in the org. A validation rule that fails is therefore surfaced the same way whether the save came from a record page, a quick-create modal, an inline list edit, or a bulk import — the rule ran once in save.validate and returned one envelope, and only the renderer differs. A non-blocking warning or info on a row that did commit rides the same envelope and renders as an inline notice that does not block. The placement contract in UI terms lives in Notifications & feedback, a package-delivered surface; this is the origin side of it.

Partial success is first-class. A bulk write takes a switch — all-or-nothing (any failure rolls back the whole batch) or save-what-you-can (good rows commit; failed rows return with their own per-item errors), matching Salesforce’s allOrNone DML flag. Results carry per-item outcomes positionally aligned to the request, and the envelope’s severity lets warnings and info ride alongside hard errors — “row 3 failed on Name, row 7 warned on Amount, rows 1–2 committed” is expressible in one result. Beyond that, a row that did commit can still return a non-blocking warning/info on the same envelope (“saved — this value is unusually high”); Salesforce’s save result carries per-row errors but no advisory on a row that succeeded, so carrying severity on every envelope is a small step past it. A single top-level error field cannot represent any of this; the unit of the envelope is per-input-item.

The safety guarantee follows from the normalizer boundary, not from a complete error list.

  • Known failure → the normalizer maps it to a typed code with a curated friendly message.
  • Unknown failure → the normalizer emits class: internal, mints a correlationId, writes the full raw detail to telemetry, and returns only the id + a generic friendly message. The raw text is never in the response.

Because the unknown path is handled uniformly at every surface, the platform is safe on day one, before anyone has enumerated a single specific error. Anticipation is not a prerequisite for safety.

The telemetry promotion loop is how the catalog grows. An internal error is simultaneously safe (opaque id to the user) and logged (full raw detail server-side). An operator — or an automated monitoring agent — watches the internal-error stream; a frequently-recurring signature is promoted to a typed code with a friendly message and, where possible, a fault/retriable classification. This is additive — no big-bang enumeration, no migration. The catalog converges on the failures that actually happen, in frequency order.

“List every error” resolves into three tiers, and most of it is not hand-maintained.

  • Metadata-derived families — validation, computation, permission, deploy, plus constraint and lifecycle errors. Because validation rules, formulas, roll-ups, permissions, constraints, and lifecycle gates are all declared metadata, and the pure tier is statically analyzable, the set of possible errors for a given object/org is computable from the metadata graph. The platform generates these per org and surfaces likely failures at author time — inline, as a formula’s divide-by-zero risk or a validation rule’s field references or a deploy’s dependency order is written — not only reactively at runtime. For constraint and lifecycle errors the boundary is precise: the kernel catches the raw Postgres constraint violation or blocked gate (a fixed primitive), and metadata names it — mapping the violation back to the component that declared it and emitting the friendly, per-org code/message. A constraint a user declared surfaces as a metadata-derived code; a constraint with no metadata origin (a kernel invariant, a save-order mechanic) surfaces as a fixed kernel code.
  • A small fixed kernel catalog — conflict, limit, integration, internal, and the runtime primitives (input parsing, the save-order mechanics themselves). On the order of dozens of codes, defined once in the kernel, stable across orgs.
  • The catch-all — internal.unexpected makes everything not in the first two tiers safe from day one.

So the catalog is neither a giant hand-written enum nor absent: the large, org-specific part is generated and cannot drift from behavior (it is derived from the same metadata that produces the behavior), and the small, universal part is fixed in the kernel.

A starter set spanning every taxonomy class and a representative spread of surfaces. Metadata-derived rows (marked derived) are generated per org from the metadata graph; kernel rows (kernel) are fixed; new rows are appended as telemetry promotes recurring internal signatures. Columns: Code · Class · Origin · Surfaces to · Fault · Retriable · Friendly message (template).

Code Class Origin Surfaces to Fault Retriable Friendly message (template)
validation.required_field (derived) validation save.validate field user no “{fieldLabel} is required.”
validation.unique_violation (derived) validation save.write.constraint field user no “Another {objectLabel} already uses {fieldLabel} “{value}”.”
permission.no_edit (derived) permission save.write.access record admin no “You don’t have permission to edit {objectLabel}. Ask your administrator.”
not_found.record (kernel) not_found read.query record user no “That {objectLabel} doesn’t exist or isn’t available to you.”
conflict.stale_version (kernel) conflict commit record user yes “Someone else updated this {objectLabel} while you were working. Reload and try again.”
conflict.gate_blocked (derived) conflict save.write.gate record user no “This {objectLabel} can’t move to {stage} yet: {reason}.”
computation.divide_by_zero (derived) computation save.adjust field admin no “{fieldLabel} can’t be calculated: division by zero.”
computation.calc_raised (derived) computation calc.sandbox field admin no “{fieldLabel} couldn’t be computed. The calculation reported: {calcMessage}.”
limit.resource_exceeded (kernel) limit calc.sandbox toast admin no “This operation exceeded its resource budget ({limitName}: {observed}/{allowed}).”
deploy.dependency_missing (derived) deploy deploy.apply deploy_result admin no “{component} depends on {missing}, which isn’t in this deployment.”
integration.callout_failed (kernel) integration integration.callout async_status platform yes “Couldn’t reach {system}. We’ll retry; no action needed yet.”
internal.unexpected (kernel) internal kernel toast platform no “Something went wrong on our end. Reference {correlationId} — we’re looking into it.”

Salesforce’s error model is strong, and much of the CAOS design is parity with it, not novelty.

What Salesforce already has (parity — CAOS matches, does not reinvent):

  • A two-face envelope. Every API/DML error is {errorCode (machine), message (human), fields (locator)} — the same two-face shape, with the field/record target carried as a fields[] array of data, not as two divergent error types [1][5]. CAOS’s two-face envelope and data-carried locator are parity here.
  • A large, stable, enumerated code vocabulary shared across surfaces: StatusCode (DML/data) carries roughly 200 values, with a comparably large ExceptionCode (SOAP faults) set, and the REST/UI-API errorCode string carries that same vocabulary to HTTP clients. The exact count is version- and org-scoped — the authoritative machine-readable list lives in the org’s Enterprise/Partner WSDL, not a fixed doc table [1][2][5].
  • Partial-success bulk. Database.insert(records, allOrNone=false) returns a positional Database.SaveResult[], one entry per input row, each with isSuccess(), getId(), and getErrors() → Database.Error[]; a sibling rollback even has its own code, ALL_OR_NONE_OPERATION_ROLLED_BACK [3][4].
  • A wrapped catch-all for the failures it traps. When Salesforce catches an internal failure, it wraps it as UNKNOWN_EXCEPTION plus an opaque “Gack” Error ID (e.g. 87386591-78549) with an “include this Error ID if you contact support” message [6][7], and it normalizes known infrastructure conditions such as row-lock contention into named codes (UNABLE_TO_LOCK_ROW) rather than raw strings [1]. CAOS matches this wrapping — the mechanism is parity. Where the two part ways is coverage: Salesforce’s wrapping is not exhaustive (see the caution below), so an airtight catch-all is a CAOS advantage, not parity.

Where CAOS is genuinely better:

  • The catch-all is airtight by construction. Salesforce’s wrapping traps most internal failures but not all — the “Seven Dwarfs” have leaked raw ORA-##### / java.sql.SQLException text to end users (documented 2009, still reported as recently as 2021) [9][10]. CAOS makes that leak structurally impossible: every path to a user passes through exactly one normalizer, so an untrapped internal error becomes internal.unexpected + a correlationId, never a raw database string. The guarantee is a property of the boundary, not of having anticipated the error. Cost: it is only as strong as the discipline that no surface emits an un-normalized error — a backend that bypasses its normalizer reintroduces exactly the Seven-Dwarfs leak.
  • The correlation id is decodable by the org itself. Salesforce’s Gack/Error ID is opaque even to the tenant’s own admins — only Salesforce Support, holding the internal logs, can resolve it [6]. In CAOS every internal error’s full raw detail is written, keyed by correlationId, to an append-only trace store at a known location, retained 90 days hot then rolled up and archived (never hard-deleted). Decode (id → sanitized trace) is granted to the platform-operator/SRE role and to an automated monitoring identity that reads the internal-error stream on a schedule, triages each signature, and opens or updates a ticket with no human filing it — while end users still see only the opaque id. Cost/risk: the trace store must be gated as tightly as the errors it holds — leaking decode access to the wrong role recreates the exact information-disclosure risk the catch-all exists to prevent. This is the highest-value and highest-risk item; the security boundary is the whole point.
  • One code namespace + one envelope across every surface. Salesforce fragments the same conceptual error across StatusCode vs ExceptionCode vs REST errorCode vs UI-API output.errors, plus a deprecated getDmlStatusCode() [1][2]. CAOS emits one envelope identically from DB, API, and UI, removing an entire class of “which error format am I parsing?” work. Cost: every backend must funnel through one translation layer; any bypass reintroduces a dialect (this is the discipline behind the normalizer-boundary guarantee, not a config toggle).
  • Most of the catalog is metadata-derived and generated per org, and can surface proactively at author time. Salesforce’s code list is a fixed platform enum whose authoritative form is the per-org WSDL, surfaced only reactively at runtime [2]. CAOS derives the validation/computation/permission/deploy families from the same metadata that produces the behavior, so the catalog cannot drift, and static analysis of the pure tier can flag likely failures at edit time. Cost: the generator must be kept authoritative and versioned — a half-maintained generated catalog is worse than none, because people trust it.
  • Stable code / editable-localizable message split, and deterministic identity as an SRE signal. Salesforce lets validation-rule text be free-form, so one FIELD_CUSTOM_VALIDATION_EXCEPTION code ships wildly different human text per org [1]. CAOS separates the never-changing code from an editable, localizable message template, so text can be reworded or translated without breaking monitoring, tests, or client handling — and the deterministic code makes the internal rate an alertable health metric. Cost: it is a near-zero-cost win, with one discipline — the human message template must never carry machine-load-bearing detail (clients and tests key off code/details, never the prose), or the split silently collapses back into a single brittle string.

The limit-class contrast. Salesforce’s LimitException is uncatchable — “A governor limit has been exceeded. This exception can’t be caught” [8] — a deliberate multi-tenant isolation property: a limit breach kills the transaction and rolls back, and no catch can swallow it. CAOS surfaces limit-class breaches as a structured, observable envelope (code, which limit, observed/allowed) while keeping the hard stop server-side. The win is observability of the limit, not recoverability past it: making a limit catchable-and-continuable would break the isolation guarantee that made Salesforce make it uncatchable. Match the isolation; improve only the observability.