Application logging & observability
CAOS is not blind. On every save the kernel already produces a structured execution trace — the save-order phases entered, the budget consumed, and, on failure, a canonical error envelope — and the audit & logs layer already writes three always-on streams keyed on a correlation id. What is missing is not capture; it is ergonomics. Today that signal is a return value and a telemetry sink, not something a developer browses. Application logging closes that gap: it persists the signal the kernel already has as first-class log records — with list views, record pages, tags, scenarios, and levels — and adds a small developer API to contribute application-level entries to the very same stream. It is an enhancement of the existing observability architecture, not a parallel logging system bolted on beside it.
This design is measured against Nebula Logger, the leading open-source Salesforce logging framework, because it is the best articulation of what developers want from logging on a metadata platform. We keep its ergonomics, improve on the parts the kernel can do better, and deliberately drop the parts that exist only to work around Apex and the Salesforce platform. See vs Nebula Logger for the full keep / improve / drop accounting.
What CAOS already captures (the honest baseline)
Section titled “What CAOS already captures (the honest baseline)”Before proposing anything new, here is what the kernel emits today, and why it is not yet browsable. This is the starting point the enhancement builds on — none of it is re-invented.
| Signal the kernel already has | Where it lives today | Why it isn’t yet a usable log |
|---|---|---|
| Save-order execution trace — every phase physically entered, in order, with a timestamp | Returned on the save result (trace), alongside a per-phase run counter (phaseRuns), retry attempts, and published-outbox ids |
It is an in-memory return value, discarded when the request ends. Nothing persists it, nothing lets you find the trace for a save that happened an hour ago. |
| Budget consumption — fuel, DML, rows scanned/returned, outward calls, value-cells against the per-transaction envelope | The save governor’s running counters; a breach raises a limit-class error |
Observable only inside the transaction. The final tally is never recorded next to the save it describes. |
Canonical error envelope — code, class, origin, location, severity, fault, retriable, and (on the catch-all path) a correlationId |
Emitted by the normalizer at every surface; raw detail written to a telemetry sink | The telemetry sink is in-memory and keyed only by the failure-path id. There is no durable, queryable error record and no id on the success path. |
| Three always-on streams — metadata audit, data history, execution summary | The audit & logs surface | These are platform-emitted — they answer “who changed config”, “what changed on this record”, “what did this transaction do”. None of them is a place a developer writes “reached the pricing branch with subtotal 8,646.” |
Three gaps follow directly, and they are the whole of the work:
- The correlation id is thin. It is minted only on the internal-error catch-all path, and it is not threaded end-to-end — there is no single id on the save request that flows through every phase, onto every outbox message, onto every error envelope, and into telemetry. Without it, the “walk one failure across all streams” promise of the audit surface cannot actually be kept for an arbitrary transaction. This is the foundational primitive, and everything else depends on it.
- The signal is ephemeral. The trace, the budget tally, and the error detail are computed and thrown away. They need to land in a durable place that a person can open, filter, and link from.
- There is no developer voice. A developer cannot say “log this, at INFO, tagged
pricing, in scenarioinvoice-recalc.” The kernel logs itself; the application has no way to log its own narrative into the same correlated stream.
The enhancement, in one sentence
Section titled “The enhancement, in one sentence”Thread one correlation id through the save, persist the signal the kernel already produces as first-class log-entry objects, and give developers a one-line API to add their own entries to the same stream — delivered rollback-safely by the mechanisms the kernel already owns.
Everything below is a consequence of that sentence.
The log-record model
Section titled “The log-record model”A log in CAOS is two ordinary metadata objects, not a bespoke store. Because an object is defined by the same schema whether it is an app object or a system object, log objects get physical tables, field-level security, and read-back through the query surface for free — the same way an Invoice or an Account does — and a standard package delivers list views and record pages over them as metadata. There is no separate “log database” to learn.
| Object | Grain | Answers | Key fields |
|---|---|---|---|
| Log | One transaction | “What happened in this unit of work?” | correlationId, actor, tenant, startedAt, durationMs, outcome (committed / rolled_back), scenario, origin, phaseRuns, budget (the governor snapshot), entryCount, highestLevel, retentionDate |
| Log Entry | One event within a transaction | “What did the code say, and in what context?” | log (parent lookup), level, message, loggedAt, phase (which save phase was executing), participantKey (which rule/effect), origin, relatedRecord, tags, scenario, budgetAtEntry, errorEnvelope (when the entry is an error) |
Two design choices carry most of the weight:
- The parent Log is the transaction. Its
correlationIdis the same id the error envelope carries and the audit streams join on — so a Log record is the pivot point the audit surface promises, made concrete and openable. ItsbudgetandphaseRunsare the governor snapshot and the phase counter the kernel already computed; persisting them is a write, not a new measurement. - Tags and scenarios are typed, not junction-table folklore. A scenario is a first-class field naming the unit of work (
invoice-recalc,nightly-sync); a tag is a typed, deploy-managed label. They are queryable and governable like any other field — no separate tag object to join through unless retention or reuse demands it.
Level is the standard ordered severity, so a Log’s highestLevel and a per-scenario threshold are both meaningful:
ERROR > WARN > INFO > DEBUG > TRACE
Five levels, not Nebula’s seven — see why we collapse FINE/FINER/FINEST.
How the pieces relate
Section titled “How the pieces relate”- Capture & the save order — how the kernel’s own trace, phase attribution, budget, and errors become log context automatically, so a developer inherits structured context instead of hand-assembling it.
- Rollback-safe delivery — why success-path logs ride the transactional outbox (atomic with the record) while failure-path logs use the out-of-band telemetry sink (survives the rollback) — and why CAOS therefore does not need Nebula’s platform-event indirection.
- The developer API — the one-line
log.info(...)surface, scenarios, tags, and levels, with nosaveLog()boilerplate and no save-method to choose. - Querying & retention — reading logs back under access control through the query surface, plus retention, archival, and purge.
- The CAOS Logger surface — the standard package developers install to browse logs: Logs, Log Entries, Scenarios, and Tags as tabs, the Log record page, and the correlation pivot — delivered as an ordinary package of metadata over the log objects, not a Setup screen or a bespoke console. The capture, delivery, and correlation engine underneath stays in the kernel.
- vs Nebula Logger — the keep / improve / drop verdict, grounded in Nebula’s actual architecture.
Why this is better than a bolt-on logger
Section titled “Why this is better than a bolt-on logger”The advantage is entirely a consequence of owning the kernel, and it is worth stating plainly because it is the reason this is an enhancement and not a port:
- The context is structured by construction, not by hand. A Nebula developer calls
.setRecord(r).addField(...).addTag(...)to attach context. A CAOS log entry already knows the phase it was raised in, the participant that raised it, the budget spent to that point, the caller’s access, and the correlation id — because the kernel is the thing running the code and it already tracks all of it. The developer writes the message; the platform supplies the context. - There is no delivery mechanism to reason about. Nebula makes the developer choose among
EVENT_BUS,QUEUEABLE,REST, andSYNCHRONOUS_DML, each with different rollback and ordering behavior. CAOS routes by outcome automatically: committed saves log atomically with the record, rolled-back saves log out of band. One API, correct by default. - Logs are just objects. List views, record pages, FLS, sharing, reports, and the query surface all apply with zero new plumbing — the same claim the audit surface makes, extended to developer logs.
Sources
Section titled “Sources”- Nebula Logger — project README (jongpie/NebulaLogger) — object model,
LogEntryEvent__e,SaveMethodoptions, levels, settings, the fluent API. - Nebula Logger — Logging in Apex (wiki) — the developer call pattern and buffering.
- Salesforce Ben — How to Debug Salesforce Flow, Apex, and LWC With Nebula Logger — why platform events are used so logs survive rollback.