Skip to content

Integrations & external services

A platform is only as useful as the systems it can reach and be reached by. CAOS treats every boundary crossing as a first-class, deployable thing rather than as code someone wrote: an inbound caller is an identity holding permission sets, an outbound call names a credential component instead of a URL and a secret, an event stream is a channel with a durable cursor, and a remote table can be projected as an object without copying it.

Two decisions carry the rest of the page.

There is no integration bypass. The inbound data API is generated from the object metadata, and it resolves through the same object/field permissions, the same record access, and the same system permissions as a person clicking in the UI — enforced in the query, not by a wrapper the caller could have been given a way around. Bulk ingest is not an exception: a million rows go through the same save order as one row. The only bypass in the platform is platform.root, and an integration principal that holds it is a decision an admin made explicitly and that the metadata audit recorded.

A callout never runs inside the write transaction. Not between BEGIN and COMMIT, under any phase, for any reason. An outbound call is enqueued to a transactional outbox during EFFECTS and executed after commit, or it is a standalone operation invoked before a save begins. Holding a Postgres transaction open across an unbounded network wait pins row locks, a pooled connection, and the vacuum horizon; and a call that has already reached the far side cannot be rolled back with the row that triggered it. Both failure modes are structural, so the rule is structural.

Seven boundary shapes, each a component or a generated surface:

Direction Shape What carries it
In, synchronous Data API over every object and field — record CRUD, batch, query Generated from object metadata; nothing authored
In, high volume Bulk ingest jobs — staged, then applied through the save order Generated; job resources
In, event Webhook endpoint — verify, acknowledge, dispatch webhook_endpoint
Out, synchronous Callout to a declared endpoint named_credential + external_credential
Out, typed Operations generated from an OpenAPI description external_service
Out, event Change events and platform events on durable channels event_channel (+ any object’s change feed)
Both, no copy Virtual object backed by a remote source external_data_source + an object with storage: "external"

Three invariants hold across all seven.

Every crossing carries an identity. Inbound, that identity is a user — usually a dedicated integration user — assigned permission sets like any other. Outbound, it is a principal on an external credential, and which users may use that principal is itself a permission-set grant. There is no unauthenticated write path and no “runs as the platform” mode.

A secret is never a field value. Secret material lives in an encrypted credential store addressed by (external credential, principal). It is not a custom field, not a setting record, not a component body, and not readable by any API — including the metadata retrieve pipeline. The only operation on a secret is use it, and only the callout executor can perform it.

Failures resolve to the ordinary error envelope. A refused webhook signature, an exhausted rate-limit bucket, an open circuit breaker, and a remote 500 all emit class: integration or class: limit with a code, a correlationId, and a retriable flag, at the integration.callout or api.request origin. An integration client parses one error shape for everything, including errors that originated three systems away.

An external credential describes how to authenticate. A named credential describes where to call and with which authentication. Splitting them means one authentication setup serves many endpoints, and an endpoint can be repointed at a sandbox without touching the auth configuration.

{
"key": "ec_erp_oauth",
"label": "ERP — OAuth 2.0 client credentials",
"type": "external_credential",
"body": {
"scheme": "oauth2_client_credentials",
"tokenUrl": "https://auth.erp.example.com/oauth2/token",
"scopes": ["orders.read", "orders.write"],
"principals": [
{ "key": "svc_integration", "mode": "named", "grantedTo": ["ps_erp_integration"] },
{ "key": "per_user", "mode": "per_user", "grantedTo": ["ps_erp_users"] }
]
// No client secret here. Secrets are written to the credential store
// through a separate write-only operation and never appear in a body.
}
}
{
"key": "nc_erp_orders",
"label": "ERP — Orders API",
"type": "named_credential",
"body": {
"url": "https://api.erp.example.com/v2",
"externalCredential": "ec_erp_oauth",
"principal": "svc_integration",
"timeoutMs": 5000,
"retry": {
"attempts": 4,
"backoff": "exponential",
"baseMs": 250,
"jitter": 0.2,
"retryOn": ["timeout", "connection", "http_429", "http_5xx"]
},
"circuitBreaker": { "failureThreshold": 10, "windowSec": 60, "openSec": 30 },
"budget": { "concurrency": 8, "requestsPerMinute": 600 }
}
}

Supported schemes: oauth2_client_credentials, oauth2_authorization_code (for per_user principals), oauth2_jwt_bearer, mtls, aws_sigv4, api_key, basic, and custom (a declared header/parameter set whose values come from the credential store). A callout addresses the endpoint by component key — callout("nc_erp_orders", …) — so no URL, host, token, or key is ever written in an expression, a body, or a deployed artifact.

An external_service imports an OpenAPI 3.x description and generates typed, callable operations. The generated surface is metadata, not code: request and response schemas become platform types, each operation becomes an invocable with a checked signature, and a schema change on re-import produces a diff a reviewer reads.

{
"key": "es_erp_orders",
"label": "ERP Orders",
"type": "external_service",
"body": {
"namedCredential": "nc_erp_orders",
"spec": { "format": "openapi_3_1", "source": "specs/erp-orders.yaml", "hash": "sha256:9c1f…" },
"operations": ["getOrder", "createOrder", "listOrders"], // omit to generate all
"onSpecDrift": "fail" // re-import that changes a generated signature fails the deploy
}
}

onSpecDrift: "fail" is the default and the reason the hash is stored: a vendor editing their published spec cannot silently change the shape of an operation an automation already calls. The alternative, "warn", regenerates and records the signature change in the metadata audit.

A webhook endpoint declares the three things a receiver must get right — how a message is verified, how a repeat is recognized, and where a verified message goes — so none of them is re-implemented per integration.

{
"key": "wh_payments_events",
"label": "Payments — inbound events",
"type": "webhook_endpoint",
"body": {
"path": "/hooks/payments",
"runAs": "user_svc_payments", // ordinary user; permission sets decide what it may write
"verification": {
"scheme": "hmac_sha256",
"externalCredential": "ec_payments_signing", // the shared secret lives here
"signatureHeader": "X-Payments-Signature",
"signedPayload": "{timestamp}.{rawBody}",
"timestampHeader": "X-Payments-Timestamp",
"toleranceSec": 300
},
"dedupe": { "idHeader": "X-Payments-Event-Id", "windowHours": 72 },
"dispatch": { "publishTo": "ch_payment_events" } // or: automation key, or an upsert mapping
}
}

A data source supplies the adapter and its capability declaration; an object with storage: "external" binds a remote entity to typed platform fields. The pushdown block is a contract — what the adapter claims it can translate is what the platform will let an author write.

{
"key": "eds_erp_odata",
"label": "ERP — OData v4",
"type": "external_data_source",
"body": {
"adapter": "odata_v4",
"namedCredential": "nc_erp_odata",
"pushdown": { "filter": true, "sort": true, "paging": "server", "aggregate": false },
"pageSize": 200,
"listViewDeadlineMs": 2000
}
}
{
"key": "obj_erp_shipment",
"label": "Shipment (ERP)",
"type": "object",
"body": {
"storage": "external",
"dataSource": "eds_erp_odata",
"remoteEntity": "Shipments",
"externalId": "ShipmentNumber",
"ownerColumn": "AssignedToEmail", // optional; enables a pushed record-access predicate
"fields": [ /* mapped, typed, and subject to field permissions like any other object */ ]
}
}

One generated REST surface, versioned in the path:

Resource Behavior
GET /data/v3/{object}/{id} One record. Fields projected through field permissions; an invisible column comes back null, not omitted-and-guessable. A row the caller cannot see is not_found, never forbidden.
POST /data/v3/{object} Insert. Runs the full save order.
PATCH /data/v3/{object}/{id} Update. Optimistic concurrency via If-Match on the record version; a mismatch is conflict.stale_version.
DELETE /data/v3/{object}/{id} Routes through safe-delete, exactly as a Setup delete does.
POST /data/v3/{object}/batch Up to 200 records, allOrNone switch, per-item outcomes positionally aligned to the request — the partial-success model the platform uses everywhere.
POST /data/v3/query Structured query: object, projected fields, a filter, order, limit, cursor.
POST /data/v3/jobs/ingest Bulk. Create job → upload parts → close → poll → fetch per-row results.

The query interface is structured JSON, not a query string. The filter is the same pure boolean grammar used by validation rules and sharing-rule coverage:

{
"object": "invoice",
"fields": ["id", "name", "total_price", "account.name"],
"filter": "record.status == \"Sent\" && record.total_price > 50000",
"orderBy": [{ "field": "created_at", "dir": "desc" }],
"limit": 200,
"cursor": "eyJrIjoiMjAyNi0wNy0yMlQxNDoxMjowMFoiLCJpZCI6IjhmM2EifQ"
}

A string DSL would be simpler to type and worse in three ways: it cannot be type-checked against the metadata graph before it runs, it is not analyzable by the where-used and blast-radius machinery every other expression participates in, and it invites string concatenation at the client. The structured form compiles to the same access-checked plan the UI issues, so a query cannot express a projection the caller has no field permission for. Cursors are opaque, keyset-based, and stable under concurrent inserts.

Bulk ingest does not bypass anything. Uploaded parts land in a per-job staging table by COPY; the applier then drains the staging table in batches of 200 through the ordinary save order, so formulas, roll-ups, validation rules, automation, and the access planes all run. The cost is real and is stated plainly in Limits: this is slower than a raw load. It is the right trade, because a bulk path that skipped validation would make “the data in this object satisfies its rules” false, and every downstream consumer of that guarantee — roll-ups, formulas, reports — would silently inherit the lie.

An integration has to survive both a platform release and its own org’s schema edits. Those are two independent axes, and conflating them is the mistake to avoid.

The protocol version — the v3 in the path — pins the shape of the conversation: resource paths, envelope structure, parameter names, pagination style, the error envelope’s fields. It changes rarely. Each protocol version is supported for a minimum of three years after its successor ships, deprecation is announced at least a year before support ends, and a call to a retired version returns 410 with an error envelope naming the successor version and the migration guide. That floor matches Salesforce’s published policy deliberately; anything shorter is a promise integrators cannot plan around.

The org’s schema is not versioned, and does not need to be, because three mechanisms make a deploy non-breaking by construction:

  1. Additive changes never break a caller. A new field appears in the metadata graph; it appears in a record read only if the caller projected it or asked for the full record. Bulk and query require an explicit fields list, so a new column cannot widen an existing job’s output.
  2. A rename leaves a permanent alias. Field and object API names are stable keys; renaming registers the old name as an alias that keeps resolving on every surface — read, write, query, filter — until someone explicitly retires it, which is itself a destructive change subject to rule 3.
  3. A destructive change is blocked when a live client is using the field. Removals route through safe-delete, and the deploy pipeline consults the execution trace: if any registered client read or wrote that field within the deprecation window (default 30 days), the deploy fails with deploy.field_in_use, naming the clients and their last-use timestamps. It proceeds only with an explicit breaking: true acknowledgement on the component, which the metadata audit records against the person who made it.

So a “version pin” pins shape, not semantics. Pinning v3 does not freeze validation rules, permissions, formulas, or business behavior — those follow the org, and they must, because an integration that could opt out of the org’s validation rules is a hole in the rules. Any platform that lets a pinned client keep old behavior has to run and support every historical behavior forever; the cost of that shows up as version-conditional branches in the vendor’s own runtime.

Position Allowed? Why
SHAPE, VALIDATE No Pure tier. A network call is neither deterministic nor side-effect-free, and the compiler rejects it.
ADJUST (before_write) No Bounded to field assignments on the triggering row; no outward reach.
EFFECTS (after_write) Enqueue only Inside the transaction. outbox.publish(…) writes a row that commits with the record.
COMMIT No —
POST-COMMIT (post_commit) Yes Outside the transaction, own budget, retried and dead-lettered by the worker.
Standalone invocable operation Yes Not part of a save. Its own transaction, its own budget.

The pattern that needs a remote value before a save — a credit check, an address normalization, a live rate — is served by the standalone invocable: the caller (a UI action, an automation not currently inside a save, or an API client) runs the callout, gets a typed result, and submits the save with that value as an ordinary field. This is exactly as capable as calling out from a before-trigger, and it cannot produce the “uncommitted work pending” class of failure, because there is no open transaction to be uncommitted.

The complement is the transactional outbox. An effect enqueues; the row commits atomically with the record; a worker drains it after commit. A message therefore cannot describe a record that was rolled back, and a rollback cannot retract a message that was already on the wire — the two failure modes that make in-transaction callouts unsound.

How long an outbox row lives. The outbox is a queue, not a log, and the two have opposite retention shapes:

  • Delivered. The row is removed at acknowledgement. It is not archived, because the drainer claims against this table and a table that accumulates its own history makes every claim scan it. What survives is the delivery itself — endpoint, attempt count, status, latency, correlation id — on the execution trace, where it lives out that stream’s 90-day window.
  • Dead-lettered and unresolved. Retained in full and indefinitely: the payload, the resolved endpoint, every attempt’s error envelope, and the correlation id of the transaction that produced it. The clock starts at resolution, not at failure — the same rule a dead-lettered job run follows, and for the same reason. An undelivered message is work the far system is still owed, and expiring it on a timer discards the only record of what was owed.
  • Discarded, or dead-lettered and then resolved. Retained 90 days from the resolution, matching the execution-trace window so the message and the trace of its attempts expire together, then purged. The decision outlives the data: the discard, its actor, and its required reason are metadata-audit entries with that stream’s two-year retention, so “who decided this invoice would never be posted” is answerable long after the invoice body is gone.

A retained outbox row is a second copy of record values, so it is governed like one. Sensitive field values carry the placeholder treatment rather than the value, and an erasure request redacts the payloads of retained rows that reference the subject with the same salted hash it applies to data history — an undelivered message must not be the thing that survives an erasure.

Timeouts and retries are per named credential, not per call site, so the operational policy for a system lives in one reviewable place. The defaults above — 5 s timeout, four attempts, exponential backoff with jitter — are overridable per component and cap at 60 s for a synchronous invocable and 300 s for a post-commit worker. Retries are automatic for idempotent methods (GET, HEAD, PUT, DELETE); a POST is retried only when the operation declares an idempotency key, because retrying a non-idempotent create is how one order becomes four. Each named credential carries a circuit breaker and a budget (concurrency cap and requests-per-minute), so a remote system that has become slow degrades to fast failures instead of consuming the worker pool that every other integration shares.

Two producers, one substrate:

  • Change events — automatic per-object change capture, opt-in per object. Emitted through the same outbox as any other effect, so a change event exists if and only if the write committed. The payload carries the changed fields, the prior values for changed fields, the operation, and the committing user.
  • Platform events — author-declared event objects published explicitly from automation or an API call. Ordinary typed fields; no record storage.

Both land in a partitioned event log and are read through one subscribe API.

Property Behavior
Cursor A bigint sequence, assigned at publish time by the single drainer for a partition. Monotonic and gap-free within a partition — a consumer that has seen n knows it has seen everything up to n.
Ordering Total within a partition key. Change events partition on record id, so all changes to one record are ordered; platform events partition on an author-declared key. There is no global order, and a consumer must not assume one.
Delivery At-least-once. Every event carries eventId (uuid), sequence, and commitTs. A consumer must be idempotent — dedupe on eventId, or make the handler naturally repeatable.
Replay Subscribe from latest, earliest, or an explicit sequence. Server-side cursors are stored per registered subscription, so a consumer that crashes without checkpointing resumes where the platform last acknowledged; a consumer may also supply its own cursor and ignore the stored one.
Retention Per channel, default 7 days, maximum 30. Beyond retention, earliest starts at the oldest retained event and the response says so rather than silently starting mid-stream.
Oversize An event exceeding the channel’s maxBytes (default 1 MB) is rejected at publish, synchronously, with a limit-class error naming the field that blew the budget. It is not accepted and then quietly replaced downstream.
Fan-out Any number of consumers may read the same partition. Reads are range scans on a table the platform already wrote; there is no separate delivery allocation, and adding a second consumer does not halve the first one’s headroom.

What a consumer must tolerate, stated so it appears in an integrator’s design rather than in their incident review: duplicates, out-of-order across partitions, a gap after an outage longer than retention (surfaced explicitly, never silently), and schema evolution — a change event’s payload gains fields as the object does, so a consumer must ignore unknown fields rather than fail on them.

An object with storage: "external" has no rows in the tenant schema. Reads become adapter calls. What that costs is specific and is not negotiable by configuration:

  • No roll-ups. A roll-up over virtual children would require a remote scan on every parent write. Declaring one is a compile error, not a runtime surprise.
  • No storage-enforced constraints. No unique index, no foreign key pointing at a virtual row, no check constraint. A local object may hold a lookup to a virtual object, but the reference is validated on read, not enforced on write.
  • Filtering is bounded by pushdown. A predicate the adapter can translate is pushed to the source. One it cannot — a function the source has no analogue for, a cross-source join — is rejected at author time with the reason. The platform never quietly fetches a page and filters in memory while presenting the result as complete.
  • Paging belongs to the remote. Cursors are the source’s. Page n is not guaranteed stable if the source changed between pages, and the adapter reports when it detects that rather than presenting a shifted page as consistent.
  • Record access degrades to the object plane. Record access needs an owner_id on a row the platform stores. For a virtual object, ownership exists only if the source exposes a column that can be mapped (ownerColumn), in which case an owner predicate is pushed down; without one, the record-access plane collapses to “the object permission decides,” and the object’s Setup page says so in those words. Field permissions still apply in full — they are a projection concern and are enforced after the fetch.
  • Latency is the source’s latency. A list view is a callout. Each page beyond the first is another callout. The listViewDeadlineMs on the data source is a hard budget; exceeding it returns the partial page with a limit-class warning rather than hanging the view.

Sync instead when any of three things is true: aggregation is needed (roll-ups, report grouping, dashboards over the remote set); record-level sharing is needed and the source has no owner column; or the source’s p95 exceeds the interactive budget — take 400 ms for a list view as the working threshold. The replacement is a scheduled ingest into an ordinary object through the bulk API, which restores every capability above at the cost of staleness bounded by the sync interval. Choosing between them is a latency-versus-capability decision, and the platform’s job is to make the trade explicit rather than to pretend the virtual object is a local one.

Question Answer
Where do they live? An encrypted store outside the tenant schema, keyed by (external credential, principal), with a per-tenant data-encryption key wrapped by a key-management service.
Who may read them? Nobody. There is no read operation. The callout executor resolves material in-process, and it is never rendered, logged, or returned.
Who may write them? Holders of the integration.manage_credentials system permission. The operation is write-only, and the metadata audit records the fact of the write, the credential, and the actor — never the value.
How does rotation work? A credential holds current and next. Writing next stages; promoting swaps them and keeps the prior value valid for a grace window (default 24 h) so in-flight tokens do not fail mid-rotation. Both transitions are audited.
Can a secret be a field value? No. There is no field type that would hold one, and the credential store is the only sink. A platform where a secret can be a field value is one FLS mistake away from an exfiltration.
Do secrets deploy? No. retrieve emits credential components with secret slots absent; apply creates the component and leaves it unpopulated until a write. A deploy that would need a secret to succeed fails with a deploy error naming the credential — environment promotion never carries production secrets into a sandbox or the reverse.

Rate limits are per client, not per org. Every integration authenticates through a registered client component carrying a token bucket (sustained requests/second plus a burst allowance) and a concurrency cap. Exceeding it returns 429 with Retry-After and a limit-class envelope. The org has a ceiling, and each client’s bucket is a declared share of it, so a runaway integration throttles itself and nothing else — the containment property that matters, since the failure being guarded against is one team’s retry loop taking down everyone’s.

Callout budgets are per named credential, as above: concurrency, requests per minute, and a circuit breaker. A dependency that has become slow cannot consume the shared worker pool.

Everything is observable. Per-client request counts, throttle events, callout latency histograms, circuit-breaker transitions, event lag per subscription, and webhook verification failures are streams in the execution trace with the same envelope and correlation id as everything else. Alerting is on the same data an admin reads, not a separate paid telemetry product.

Concern CAOS Salesforce’s documented limit Why their limit exists
API requests Per-client token bucket carved from an org ceiling; 429 + Retry-After Org-wide daily allocation: 15,000 (Developer), 100,000 + 1,000/license (Enterprise), 100,000 + 5,000/license (Unlimited), 5,000,000 (Full sandbox). Soft at first, then “all subsequent API calls blocked,” 403 REQUEST_LIMIT_EXCEEDED Protects a shared multitenant substrate. Scoped to the org because there is no per-client bucket to scope it to
Concurrent long requests Per-credential and per-client concurrency caps 25 concurrent requests running longer than 20 seconds (production/sandbox), 5 (Developer/trial) Long-running requests hold shared app-server threads
Batch write 200 records per batch, allOrNone switch, per-item results 25 subrequests per Composite call, of which at most 5 may be sObject-collection or query operations Bounds work per request on shared infrastructure
Bulk ingest Parts up to 100 MB, job up to 5 GB, applied in 200-row batches through the full save order 150 MB per Bulk 2.0 job; 150,000,000 records per 24-hour rolling period; 15,000 batches per 24 h; results retained 7 days Bounds asynchronous queue depth
Callouts per save Zero inside the transaction; unbounded after commit, subject to per-credential budgets 100 callouts per transaction; 120 s cumulative callout timeout; 10 s default and 120 s maximum per callout; 6 MB request/response synchronous, 12 MB asynchronous Callouts run inside the Apex transaction, so they must be capped to bound transaction duration
Callout after a write Not applicable — the outbox is the only path CalloutException: “You have uncommitted work pending. Please commit or rollback before calling out” Callouts and DML share one transaction that cannot be held open across a network wait
Event retention 7 days default, 30 max, per channel 72 hours for high-volume events; 24 hours for PushTopic/generic/standard-volume. “Salesforce doesn’t guarantee the storage of events beyond the retention period of 72 hours” Retention on a bespoke event bus is capacity a tenant does not pay for directly
Event delivery No delivery allocation; fan-out is a range scan 50,000 (Performance/Unlimited), 25,000 (Enterprise), 10,000 (Developer) delivered events per 24 h, counted per subscriber Delivery is metered because each external subscriber costs bus capacity
Event publishing Bounded by the org’s write budget, same as any write 250,000/hour (Enterprise, Performance, Unlimited), 50,000/hour (Developer); 1 MB max event size Bounds bus ingest
Change capture scope Any object, opt-in per object 5 objects selectable for change notifications without an add-on license Each captured object multiplies bus volume
Replay cursor bigint, monotonic and gap-free within a partition An opaque ReplayId, documented as not guaranteed contiguous and recommended to be stored as bytes An opaque id lets the implementation change without a contract break — at the cost of a consumer being unable to reason about gaps
Virtual object reads Adapter pushdown declared per source; unpushable predicates rejected at author time queryMore() costs one web service callout per page; default batch 500 rows; server-driven paging caps at 2,000 rows per page; external data changing between queryMore() calls “may receive unexpected query results” External data has no local index and no snapshot, so paging is the remote’s problem

Inbound. Salesforce generates REST, SOAP, Bulk, and Composite APIs over every object and field. Composite batches up to 25 subrequests, chains them by @{referenceId.Field}, honors an allOrNone flag, and counts as a single call toward the org’s API allocation. Bulk API 2.0 handles volume, at 150 MB per job and 150,000,000 records per rolling 24 hours. The allocation itself is org-wide, and when it is exhausted the platform first allows an overage, then enforces “a system protection limit … and all subsequent API calls blocked,” returning 403 REQUEST_LIMIT_EXCEEDED. The containment tool an admin is given is an API usage notification: an email when usage crosses a percentage threshold in an interval. There is no per-client bucket, so a single misbehaving integration exhausts the allocation for every other integration in the org, and the admin learns about it by email.

Custom inbound endpoints are hand-written Apex REST classes annotated @RestResource. These are the platform’s one genuine bypass surface: Salesforce’s own documentation notes that “sharing declarations don’t enforce object-level access or field-level security,” so an Apex REST endpoint enforces CRUD and FLS only to the extent the developer wrote code to enforce them.

API versioning. Every request names a version. Salesforce commits to “supporting each API version for a minimum of 3 years from the date of first release” and to notifying customers “at least 1 year before support for the version ends.” Retired versions return 410:GONE for REST. Versions 7.0–20.0 became unavailable in Summer ’22; 21.0–30.0 in Summer ’25 — after the retirement was postponed from Summer ’23 “following consultation with the community and our partners … to ensure a smooth transition,” which is the clearest available evidence of how much pain the original schedule caused. Version settings run deeper than the wire protocol: an Apex class or trigger is “stored with the version settings for a specific Salesforce API version,” so that as the platform evolves “a class or trigger is still bound to versions with specific, known behavior,” and a package version setting “determines the exposed interface of any Apex code in the installed package.” That is genuine behavior pinning, and it obliges Salesforce to keep every historically-pinned behavior alive.

Outbound. Named Credentials and External Credentials split endpoint from authentication. A callout writes callout:My_Named_Credential/some_path and “Salesforce manages all authentication for Apex callouts that specify a named credential as the callout endpoint so that your code doesn’t have to.” External credentials support OAuth 2.0 in several variants (including client credentials with a JWT assertion), JWT, AWS Signature v4, Basic, and Custom, with principals as the unit that permission sets grant access to. Secrets are excluded from metadata: the Metadata API “can’t fully expose the definition of a credential and render sensitive information like tokens in plain text,” so a packaged credential arrives unpopulated and an admin fills it in after install.

Callout limits are per Apex transaction: 100 callouts, a 120-second cumulative timeout, a 10-second default per callout with a 1 ms–120,000 ms range, and a 6 MB (synchronous) / 12 MB (asynchronous) request-and-response ceiling. And because callouts and DML share one transaction, a callout after a write raises System.CalloutException: You have uncommitted work pending. Please commit or rollback before calling out — one of the most-searched errors in the ecosystem, worked around by reordering operations or moving the callout to @future/Queueable.

External Services imports an OpenAPI description and generates typed operations — “objects and operations defined in the external service’s registered API specification become Apex classes and methods in the ExternalService namespace,” so a developer makes “a type safe callout … without needing to use the Http class or perform transforms on JSON strings.” Spring ’22 added OpenAPI 3.0 JSON support and the composition/polymorphism keywords allOf, anyOf, oneOf, and discriminator; media types outside the specification must be hand-mapped at registration.

Events out. Platform Events and Change Data Capture publish to a shared bus consumed over the gRPC/HTTP-2 Pub/Sub API (which also provides “flow control that lets you specify how many events to receive in a subscribe call”) or the older CometD Streaming API. Each event carries an opaque ReplayId; the durability documentation is explicit that the id is opaque, is not guaranteed contiguous across maintenance, and should be stored as bytes, and that “Salesforce doesn’t guarantee the storage of events beyond the retention period of 72 hours.” Subscribers replay with -1 (new events only) or -2 (everything within retention). Change events larger than 1 MB “aren’t delivered and are replaced by a gap event message.”

The allocation model is where practitioners hit a wall. Publishing runs at 250,000 events/hour on Enterprise and above, while delivery to external subscribers is 25,000 (Enterprise) or 50,000 (Performance/Unlimited) per 24 hours — and the delivery counter is per subscriber, so, as one practitioner writes, “if you are planning on delivering events to 10 different systems you can only publish 15k events per rolling 24 hours.” When the delivery allocation is exhausted, “all cometD subscriptions start failing” with 403::Organization total events daily limit exceeded — the integration is simply down. The same account names the absence of centralized event logging as the reason overages are hard to diagnose.

The older outbound message mechanism is candid about its guarantees, and they are the guarantees any event consumer has to design for: messages “stay in the queue until sent successfully, or until they’re 24 hours old,” retry with exponential backoff “up to a maximum of two hours between retries,” and — verbatim — “while each message is usually delivered one time, it can sometimes be delivered more than one time” and “messages are retried independent of their order in the queue. As a result, messages can be delivered out of order.”

Webhooks in. There is no first-class inbound webhook receiver. The established pattern is an @RestResource Apex class exposed through a Salesforce Site, with HMAC verification, replay protection, deduplication, and fast acknowledgement all written by hand; outbound messages, for their part, carry no native HMAC signature.

External data. Salesforce Connect projects OData, an Apex-written custom adapter, or another org’s objects as external objects. The documented mechanics: queryMore() “results in a Web service callout” per page against a default 500-row batch, server-driven paging caps at 2,000 rows per page, queryAll() “behaves the same as query()” because Salesforce doesn’t track external changes, and if the external data changes between queryMore() calls “you may receive unexpected query results.” Apex triggers and Apex-managed sharing are not supported on external objects.

Where CAOS is genuinely better:

  • Per-client rate buckets instead of one org allocation. Each registered client holds a token bucket and concurrency cap carved from the org ceiling, so a runaway retry loop returns 429 to itself rather than blocking every other integration and notifying an admin by email after the fact.
  • A callout cannot sit inside a write transaction. The outbox.publish / post-commit split removes the “uncommitted work pending” failure mode by making the illegal arrangement inexpressible, and removes the transaction-duration risk that forces a 100-callout, 120-second cap in the first place.
  • A monotonic, gap-free replay cursor per partition. A bigint sequence assigned by a single per-partition drainer lets a consumer prove it has processed everything up to n. An opaque, possibly-non-contiguous id cannot support that claim, which is why consumers of it end up maintaining their own dedupe stores anyway.
  • No delivery allocation. Events are rows in a partitioned table; fan-out is a range scan. Adding a tenth consumer does not divide the ninth’s headroom, and the publish/deliver asymmetry that takes integrations offline does not exist to be exhausted.
  • Oversize events are rejected at publish. A payload over the channel budget fails synchronously with the offending field named, rather than being accepted and replaced downstream by a gap event the publisher never sees.
  • Destructive schema changes are blocked by observed usage. The deploy pipeline reads the execution trace and refuses to remove a field a registered client used inside the deprecation window. Version pinning cannot provide this, because a pin protects the client that pinned and says nothing about the client that did not.
  • Webhook receipt is a component, not a hand-written endpoint. Signature scheme, tolerance window, deduplication, acknowledgement deadline, and dispatch target are declared and reviewable, so the security-critical parts are not re-implemented per integration with per-integration bugs.
  • Bulk goes through the save order. Volume ingest is subject to the same validation, automation, and access planes as a single record, so “the rows in this object satisfy its rules” stays true regardless of how they arrived.
  • Pushdown is declared and enforced at author time. A filter the adapter cannot translate fails when it is written, with the reason, instead of degrading silently into a partial scan presented as a complete result.

Parity: the endpoint/authentication split with grantable principals; secrets excluded from metadata retrieve and populated post-install; typed operations generated from an OpenAPI description; a generated REST surface over every object; a bulk pipeline with staged upload and per-row results; partial-success semantics with per-item outcomes; at-least-once event delivery with replay from a cursor; a three-year minimum support window with a year of deprecation notice. These are the right answers and CAOS copies them.

Costs and risks:

  • Bulk through the save order is slower than a raw load, by a wide margin on large jobs. The guarantee is worth the throughput, but “load 50 million rows in an hour” is not a thing this platform will do, and any migration plan has to be sized against the applier’s real rate rather than against COPY.
  • A single drainer per partition is a throughput ceiling and a failure domain. Gap-free ordering requires one writer assigning sequences per partition. Partition count is the parallelism knob, and a partition whose drainer is stuck stops its own stream — a stall that must be alarmed on, because a silent one looks exactly like quiet.
  • Retention costs storage that a bespoke bus does not. Seven days of every change event on every captured object is a real table with real bytes. Per-channel retention and capture opt-in are the levers, and an org that captures everything at 30 days will notice.
  • The credential store is a concentrated target. One encrypted store holding every tenant’s outbound secrets is worth more to an attacker than the same secrets scattered. Per-tenant key wrapping, a write-only interface, no read path, and an audited management permission are what make it defensible, and they are load-bearing rather than hardening to be added later.
  • Permanent aliases accumulate. Never breaking a renamed field means carrying its old name indefinitely; over years an org acquires alias debt that makes its own metadata harder to read. Retirement is available and is a destructive change; nothing forces anyone to use it.
  • Usage-based deploy blocking depends on trace retention. A client that runs monthly and a 30-day window are one clock skew apart from a false negative. The window is configurable per client for exactly this reason, and a client that cannot be characterized should be pinned to a longer window rather than trusted to the default.
  • Virtual objects will still disappoint someone. No roll-ups, no constraints, degraded record access, and remote-bound latency are inherent to not owning the data. Declaring the limits at author time makes the disappointment early rather than absent.
  • Owning the event substrate means owning its operations. Retention policies, partition rebalancing, dead-letter handling, consumer-lag alerting, and back-pressure are infrastructure a hosted bus provides. On Postgres, they are on the platform’s roadmap and its on-call rotation.
Component type Body
External credential external_credential scheme, protocol parameters, principals[] with grantedTo permission sets. No secret material.
Named credential named_credential url, externalCredential, principal, timeoutMs, retry, circuitBreaker, budget
External service external_service namedCredential, spec (format, source, hash), operations[], onSpecDrift
Webhook endpoint webhook_endpoint path, runAs, verification, dedupe, dispatch
External data source external_data_source adapter, namedCredential, pushdown, pageSize, listViewDeadlineMs
Virtual object object storage: "external", dataSource, remoteEntity, externalId, ownerColumn, fields[]
Event channel event_channel partitionKey, retentionDays, maxBytes, schema (platform events) or object + captured fields (change events)
Registered client integration_client authScheme, runAs, rateLimit (sustained, burst, concurrency), deprecationWindowDays

Subscriptions, stored cursors, credential values, and issued tokens are runtime state, never components — the same split the permission model draws between a permission set and its assignments. A retrieve of an external_credential returns its scheme and principals with the secret slots absent; an apply creates the component unpopulated; the environment’s own operator writes the values. This is what makes environment promotion safe by default: a sandbox that received a deploy from production holds production’s endpoints and shapes and none of its credentials, so it cannot accidentally call a production system with production authority.

Everything else follows the standard pipeline — retrieve → diff → validate → apply, transactionally, under the generation flip. Validate is where the integration-specific checks run: an OpenAPI spec whose hash moved and whose generated signatures changed fails under onSpecDrift: "fail"; a virtual object declaring a roll-up fails; a filter expression the declared adapter cannot push down fails; and a field removal fails when the execution trace shows a registered client using it.

Salesforce Metadata API analogs, for migration mapping: NamedCredential, ExternalCredential, ExternalServiceRegistration, ExternalDataSource, CustomObject with an external data source, PlatformEventChannel and PlatformEventChannelMember, ConnectedApp, ApiUsageNotification, WorkflowOutboundMessage, and ApexClass carrying @RestResource.