Errors & diagnostics
CAOS does not pre-enumerate every possible failure. It enumerates the finite set of surfaces where code runs and places one normalizer boundary at each. A normalizer catches whatever happens at its surface and emits a single canonical envelope: a known failure becomes a typed code; an unknown failure becomes class: internal with a correlationId, and its raw text is captured to telemetry and never returned. The consequence is a construction guarantee, not a testing outcome: no raw Postgres, kernel, or runtime error can reach a user, ever — safety does not depend on having anticipated the specific error.
The model
Section titled “The model”An error in CAOS is described on three orthogonal axes. Keeping them independent is what makes “it doesn’t matter where the error came from” literally true — a new surface plugs in without touching anything downstream.
- Origin / surface — WHERE it arose: the normalizer that caught it. The surfaces are a fixed set:
authoring,deploy.apply,save.shape,save.adjust,save.validate,save.write.constraint,save.write.access,save.write.rollup,save.write.gate,save.effects,commit,post_commit,read.query,calc.sandbox,integration.callout,api.request,kernel. These map onto the fixed save order of execution plus the read, deploy, authoring, calc, integration, and API boundaries. - Class — WHAT kind of failure it is: one of the small fixed taxonomy classes (below).
- Location — WHERE it surfaces to the user, which drives the UI target:
field,record,toast,setup,deploy_result,api,async_status, orglobal.
The normalizer at an origin sets class, code, raw, and details; the location decides rendering. This separation is the design’s load-bearing property: a new origin only needs a normalizer that emits the envelope — nothing else changes. Routing is dynamic because the three axes vary independently; the same class (validation) can arise at several origins and surface at several locations without any of them knowing about the others.
Boot is a surface too. A failure while an organization boots — it cannot be identified, something installed in it fails its checks, or part of it does not load — reaches the error screen as the same envelope, from the boot path’s own normalizer, under catalogued codes: not_found.org, validation.bundle_component and internal.boot_degraded. The workspace normalizes the failures only it can see the same way: integration.engine_unreachable when the Engine cannot be reached, and internal.surface_unregistered when nothing can draw the surface the Engine chose. The screen renders the envelope’s fault and retriable flag rather than guessing them, so it offers to try again only when trying again can help.
The envelope
Section titled “The envelope”Every error — at every surface, in every API, in the UI — is the same shape. The envelope has a machine face (stable, for monitoring/clients/tests), a human face (friendly, editable, localizable), and one dev-only field that is never serialized to an end user.
| Field | Face | Purpose |
|---|---|---|
code |
machine | Stable, namespaced identifier — e.g. validation.required_field, computation.divide_by_zero, deploy.dependency_missing, internal.unexpected. Never changes. |
class |
machine | One of the fixed taxonomy classes (below). |
message |
human | Friendly, actionable, no jargon/cell-refs/raw text. Rendered from an editable, localizable template keyed by code. |
origin |
machine | The surface that raised it (axis 1). |
location |
machine | UI target + the offending thing: objectApi, recordId, fieldApi, componentKey, deployComponent. Drives axis 3. |
severity |
both | error | warning | info — blocks vs advisory. Lets warnings and info coexist with hard errors in one result. |
retriable |
machine | Whether the identical action can succeed on retry (deadlock/timeout = yes; validation = no). |
fault |
both | Who resolves it: user | admin | platform — lets the UI say “fix this” vs “ask your admin” vs “we’re on it (ref …)”. |
correlationId |
both | Trace id linking the friendly error to the full raw detail in telemetry. Shown for internal. |
details |
machine | Structured specifics keyed by api name: failing field, violated constraint, calc-function key + version, missing dependency. |
docs |
machine | Optional link to the relevant guide page. |
raw |
dev-only | The underlying original error, captured for telemetry. Never serialized to an end-user response. |
The location locator is carried as data on one envelope, not expressed as two divergent error types — a field-level error is the same object as a record-level error with a fieldApi populated. Empty locator ⇒ record/page-level. (This mirrors Salesforce’s fields[] array design; see How Salesforce does it.)
The taxonomy
Section titled “The taxonomy”Every error is exactly one class; codes are namespaced under the class. The set is small and fixed.
| Class | What it is | Typical fault | Typical location |
|---|---|---|---|
validation |
A declared rule/constraint the user can fix (required, format, validation rule, uniqueness). | user |
field / record |
permission |
Action not allowed (object CRUD, FLS, Setup capability). | admin |
record / setup |
not_found |
The thing doesn’t exist or isn’t visible (see the security rule below). | user |
record / global |
conflict |
Concurrency/state conflict (stale version, lifecycle gate blocked, duplicate). | user / admin |
record |
computation |
A formula/roll-up/calc failed to produce a value (divide-by-zero, type-contract violation, calc raised, sandbox limit). | user / admin |
field / record |
limit |
A quota/resource/governor limit hit. | user / admin |
toast / async_status |
deploy |
A metadata validate/apply failure (dependency order, destructive-guard, migration, integrity/cycle). | admin |
deploy_result |
integration |
An external system failed (callout, sync, external auth). | admin / platform |
record / async_status |
internal |
The catch-all. Never the user’s fault. correlationId shown; raw hidden. The anti-“raw-DB-error” class. |
platform |
toast / global |
Where errors surface
Section titled “Where errors surface”The origins × locations map. An origin’s normalizer determines the class and code; the location determines what the user sees. The pairing is not fixed one-to-one — one origin can surface at different locations depending on the caller (UI vs API vs async job).
| Origin (surface) | Common classes raised | Surfaces to (location) |
|---|---|---|
authoring |
validation, computation, deploy |
setup (author-time, inline on the metadata editor) |
deploy.apply |
deploy |
deploy_result |
save.shape / save.adjust |
validation, computation |
field / record |
save.validate |
validation |
field / record |
save.write.constraint |
validation, conflict |
field / record |
save.write.access |
permission, not_found |
record / global |
save.write.rollup |
computation |
record |
save.write.gate |
conflict |
record |
save.effects |
computation, integration, internal |
record / async_status |
commit |
conflict, internal |
toast / record |
post_commit |
integration, internal |
async_status |
read.query |
permission, not_found, limit |
record / global |
calc.sandbox |
computation, limit |
field / record |
integration.callout |
integration, limit |
record / async_status |
api.request |
validation, permission, not_found |
api |
kernel |
internal |
toast / global |
Location decides the component, not the page. The location on the envelope is a UI target, and a fixed set of standard components renders each one — no surface writes its own error handling. field renders inline against the offending field, which keeps the typed value and takes focus back. record renders a ValidationSummary at the top of the owning section — the page-level “here is everything wrong” box — each entry linked to its field. toast is for an outcome with no field or record to attach to, and never for anything that needs a decision. setup renders inline on the metadata editor, deploy_result in the deploy report, api as the response body, async_status on the job, global as a page-level banner. The kernel decides which by emitting the envelope; the shell and these components decide how it looks. Those renderers are delivered by packages, not compiled into the engine — the envelope and the normalizer boundary are the kernel’s, while the components that paint each location arrive with the UI packages installed in the org. A validation rule that fails is therefore surfaced the same way whether the save came from a record page, a quick-create modal, an inline list edit, or a bulk import — the rule ran once in save.validate and returned one envelope, and only the renderer differs. A non-blocking warning or info on a row that did commit rides the same envelope and renders as an inline notice that does not block. The placement contract in UI terms lives in Notifications & feedback, a package-delivered surface; this is the origin side of it.
Partial success is first-class. A bulk write takes a switch — all-or-nothing (any failure rolls back the whole batch) or save-what-you-can (good rows commit; failed rows return with their own per-item errors), matching Salesforce’s allOrNone DML flag. Results carry per-item outcomes positionally aligned to the request, and the envelope’s severity lets warnings and info ride alongside hard errors — “row 3 failed on Name, row 7 warned on Amount, rows 1–2 committed” is expressible in one result. Beyond that, a row that did commit can still return a non-blocking warning/info on the same envelope (“saved — this value is unusually high”); Salesforce’s save result carries per-row errors but no advisory on a row that succeeded, so carrying severity on every envelope is a small step past it. A single top-level error field cannot represent any of this; the unit of the envelope is per-input-item.
Dynamic capture & the catch-all
Section titled “Dynamic capture & the catch-all”The safety guarantee follows from the normalizer boundary, not from a complete error list.
- Known failure → the normalizer maps it to a typed
codewith a curated friendlymessage. - Unknown failure → the normalizer emits
class: internal, mints acorrelationId, writes the fullrawdetail to telemetry, and returns only the id + a generic friendly message. The raw text is never in the response.
Because the unknown path is handled uniformly at every surface, the platform is safe on day one, before anyone has enumerated a single specific error. Anticipation is not a prerequisite for safety.
The telemetry promotion loop is how the catalog grows. An internal error is simultaneously safe (opaque id to the user) and logged (full raw detail server-side). An operator — or an automated monitoring agent — watches the internal-error stream; a frequently-recurring signature is promoted to a typed code with a friendly message and, where possible, a fault/retriable classification. This is additive — no big-bang enumeration, no migration. The catalog converges on the failures that actually happen, in frequency order.
Metadata-derived vs kernel-fixed catalog
Section titled “Metadata-derived vs kernel-fixed catalog”“List every error” resolves into three tiers, and most of it is not hand-maintained.
- Metadata-derived families —
validation,computation,permission,deploy, plus constraint and lifecycle errors. Because validation rules, formulas, roll-ups, permissions, constraints, and lifecycle gates are all declared metadata, and the pure tier is statically analyzable, the set of possible errors for a given object/org is computable from the metadata graph. The platform generates these per org and surfaces likely failures at author time — inline, as a formula’s divide-by-zero risk or a validation rule’s field references or a deploy’s dependency order is written — not only reactively at runtime. For constraint and lifecycle errors the boundary is precise: the kernel catches the raw Postgres constraint violation or blocked gate (a fixed primitive), and metadata names it — mapping the violation back to the component that declared it and emitting the friendly, per-orgcode/message. A constraint a user declared surfaces as a metadata-derived code; a constraint with no metadata origin (a kernel invariant, a save-order mechanic) surfaces as a fixed kernel code. - A small fixed kernel catalog —
conflict,limit,integration,internal, and the runtime primitives (input parsing, the save-order mechanics themselves). On the order of dozens of codes, defined once in the kernel, stable across orgs. - The catch-all —
internal.unexpectedmakes everything not in the first two tiers safe from day one.
So the catalog is neither a giant hand-written enum nor absent: the large, org-specific part is generated and cannot drift from behavior (it is derived from the same metadata that produces the behavior), and the small, universal part is fixed in the kernel.
The catalog table
Section titled “The catalog table”A starter set spanning every taxonomy class and a representative spread of surfaces. Metadata-derived rows (marked derived) are generated per org from the metadata graph; kernel rows (kernel) are fixed; new rows are appended as telemetry promotes recurring internal signatures. Columns: Code · Class · Origin · Surfaces to · Fault · Retriable · Friendly message (template).
| Code | Class | Origin | Surfaces to | Fault | Retriable | Friendly message (template) |
|---|---|---|---|---|---|---|
validation.required_field (derived) |
validation |
save.validate |
field |
user | no | “{fieldLabel} is required.” |
validation.unique_violation (derived) |
validation |
save.write.constraint |
field |
user | no | “Another {objectLabel} already uses {fieldLabel} “{value}”.” |
permission.no_edit (derived) |
permission |
save.write.access |
record |
admin | no | “You don’t have permission to edit {objectLabel}. Ask your administrator.” |
not_found.record (kernel) |
not_found |
read.query |
record |
user | no | “That {objectLabel} doesn’t exist or isn’t available to you.” |
conflict.stale_version (kernel) |
conflict |
commit |
record |
user | yes | “Someone else updated this {objectLabel} while you were working. Reload and try again.” |
conflict.gate_blocked (derived) |
conflict |
save.write.gate |
record |
user | no | “This {objectLabel} can’t move to {stage} yet: {reason}.” |
computation.divide_by_zero (derived) |
computation |
save.adjust |
field |
admin | no | “{fieldLabel} can’t be calculated: division by zero.” |
computation.calc_raised (derived) |
computation |
calc.sandbox |
field |
admin | no | “{fieldLabel} couldn’t be computed. The calculation reported: {calcMessage}.” |
limit.resource_exceeded (kernel) |
limit |
calc.sandbox |
toast |
admin | no | “This operation exceeded its resource budget ({limitName}: {observed}/{allowed}).” |
deploy.dependency_missing (derived) |
deploy |
deploy.apply |
deploy_result |
admin | no | “{component} depends on {missing}, which isn’t in this deployment.” |
integration.callout_failed (kernel) |
integration |
integration.callout |
async_status |
platform | yes | “Couldn’t reach {system}. We’ll retry; no action needed yet.” |
internal.unexpected (kernel) |
internal |
kernel |
toast |
platform | no | “Something went wrong on our end. Reference {correlationId} — we’re looking into it.” |
How Salesforce does it
Section titled “How Salesforce does it”Salesforce’s error model is strong, and much of the CAOS design is parity with it, not novelty.
What Salesforce already has (parity — CAOS matches, does not reinvent):
- A two-face envelope. Every API/DML error is
{errorCode (machine), message (human), fields (locator)}— the same two-face shape, with the field/record target carried as afields[]array of data, not as two divergent error types [1][5]. CAOS’s two-face envelope and data-carried locator are parity here. - A large, stable, enumerated code vocabulary shared across surfaces:
StatusCode(DML/data) carries roughly 200 values, with a comparably largeExceptionCode(SOAP faults) set, and the REST/UI-APIerrorCodestring carries that same vocabulary to HTTP clients. The exact count is version- and org-scoped — the authoritative machine-readable list lives in the org’s Enterprise/Partner WSDL, not a fixed doc table [1][2][5]. - Partial-success bulk.
Database.insert(records, allOrNone=false)returns a positionalDatabase.SaveResult[], one entry per input row, each withisSuccess(),getId(), andgetErrors()→Database.Error[]; a sibling rollback even has its own code,ALL_OR_NONE_OPERATION_ROLLED_BACK[3][4]. - A wrapped catch-all for the failures it traps. When Salesforce catches an internal failure, it wraps it as
UNKNOWN_EXCEPTIONplus an opaque “Gack” Error ID (e.g.87386591-78549) with an “include this Error ID if you contact support” message [6][7], and it normalizes known infrastructure conditions such as row-lock contention into named codes (UNABLE_TO_LOCK_ROW) rather than raw strings [1]. CAOS matches this wrapping — the mechanism is parity. Where the two part ways is coverage: Salesforce’s wrapping is not exhaustive (see the caution below), so an airtight catch-all is a CAOS advantage, not parity.
Where CAOS is genuinely better:
- The catch-all is airtight by construction. Salesforce’s wrapping traps most internal failures but not all — the “Seven Dwarfs” have leaked raw
ORA-#####/java.sql.SQLExceptiontext to end users (documented 2009, still reported as recently as 2021) [9][10]. CAOS makes that leak structurally impossible: every path to a user passes through exactly one normalizer, so an untrapped internal error becomesinternal.unexpected+ acorrelationId, never a raw database string. The guarantee is a property of the boundary, not of having anticipated the error. Cost: it is only as strong as the discipline that no surface emits an un-normalized error — a backend that bypasses its normalizer reintroduces exactly the Seven-Dwarfs leak. - The correlation id is decodable by the org itself. Salesforce’s Gack/Error ID is opaque even to the tenant’s own admins — only Salesforce Support, holding the internal logs, can resolve it [6]. In CAOS every
internalerror’s fullrawdetail is written, keyed bycorrelationId, to an append-only trace store at a known location, retained 90 days hot then rolled up and archived (never hard-deleted). Decode (id → sanitized trace) is granted to the platform-operator/SRE role and to an automated monitoring identity that reads theinternal-error stream on a schedule, triages each signature, and opens or updates a ticket with no human filing it — while end users still see only the opaque id. Cost/risk: the trace store must be gated as tightly as the errors it holds — leaking decode access to the wrong role recreates the exact information-disclosure risk the catch-all exists to prevent. This is the highest-value and highest-risk item; the security boundary is the whole point. - One code namespace + one envelope across every surface. Salesforce fragments the same conceptual error across
StatusCodevsExceptionCodevs RESTerrorCodevs UI-APIoutput.errors, plus a deprecatedgetDmlStatusCode()[1][2]. CAOS emits one envelope identically from DB, API, and UI, removing an entire class of “which error format am I parsing?” work. Cost: every backend must funnel through one translation layer; any bypass reintroduces a dialect (this is the discipline behind the normalizer-boundary guarantee, not a config toggle). - Most of the catalog is metadata-derived and generated per org, and can surface proactively at author time. Salesforce’s code list is a fixed platform enum whose authoritative form is the per-org WSDL, surfaced only reactively at runtime [2]. CAOS derives the
validation/computation/permission/deployfamilies from the same metadata that produces the behavior, so the catalog cannot drift, and static analysis of the pure tier can flag likely failures at edit time. Cost: the generator must be kept authoritative and versioned — a half-maintained generated catalog is worse than none, because people trust it. - Stable code / editable-localizable message split, and deterministic identity as an SRE signal. Salesforce lets validation-rule text be free-form, so one
FIELD_CUSTOM_VALIDATION_EXCEPTIONcode ships wildly different human text per org [1]. CAOS separates the never-changingcodefrom an editable, localizablemessagetemplate, so text can be reworded or translated without breaking monitoring, tests, or client handling — and the deterministic code makes theinternalrate an alertable health metric. Cost: it is a near-zero-cost win, with one discipline — the humanmessagetemplate must never carry machine-load-bearing detail (clients and tests key offcode/details, never the prose), or the split silently collapses back into a single brittle string.
The limit-class contrast. Salesforce’s LimitException is uncatchable — “A governor limit has been exceeded. This exception can’t be caught” [8] — a deliberate multi-tenant isolation property: a limit breach kills the transaction and rolls back, and no catch can swallow it. CAOS surfaces limit-class breaches as a structured, observable envelope (code, which limit, observed/allowed) while keeping the hard stop server-side. The win is observability of the limit, not recoverability past it: making a limit catchable-and-continuable would break the isolation guarantee that made Salesforce make it uncatchable. Match the isolation; improve only the observability.
Sources
Section titled “Sources”- SOAP Core Data Types Used in API Calls — the
Errortype, theStatusCodeenum,extendedErrorDetails,UNABLE_TO_LOCK_ROW. [1] - SOAP Error Handling — fault vs Error object;
ExceptionCode; the WSDL as authoritative code source. [2] - Bulk DML Exception Handling —
SaveResult[],allOrNone, partial success. [3] - Returned Database Errors —
Database.ErrorgetStatusCode/getMessage/getFields, per-row iteration. [4] - REST Status Codes and Error Responses — the JSON
{errorCode, message, fields[]}array shape;MALFORMED_ID. [5] - Xappex — Salesforce Error ID Lookup (secondary) — “Gack,” Error ID format, non-decodable by customers. [6]
- Gearset — UNKNOWN_EXCEPTION help (secondary) —
UNKNOWN_EXCEPTION+ “include this Error ID” wording. [7] - Exception Class and Built-In Exceptions —
LimitExceptionis uncatchable (verbatim). [8] - Salesforce, Oracle and the Seven Dwarfs — Salesforce Stack Exchange (community) — the seven internal error names; raw
ORA-/java.sql.SQLExceptiontext; “occasionally break out and appear to a user,” contact Support. [9] - Meaningful Error Messages #94 — The Silver Lining (2009) (secondary) — early public record of the Seven-Dwarfs errors surfacing to users. [10]
- Unable to Lock Row — Xappex (secondary) — Salesforce does not auto-retry
UNABLE_TO_LOCK_ROW; retry is the developer’s/integration’s responsibility. [11] - Serialization Failure Handling — PostgreSQL documentation — retry
40001/40P01in a bounded, backed-off loop; applications must be prepared to retry. [12]