Skip to content

Background work & scheduling

Everything that runs outside a user’s request is a job. Post-commit effects handed off from a save, a one-off task queued from a button, a nightly cron, a batch pass over eleven million rows, an overnight import — these are one runtime primitive with different enqueue paths and different chunking, not five subsystems with five sets of limits.

Two decisions carry this page.

A job runs as a real user. It has an owner, a runAs identity that resolves to an actual user record, and it clears the same three access planes every interactive request clears. There is no privileged automation identity, because there is no system mode anywhere in the platform — that is a kernel invariant, and background work is the place platforms usually break it. Salesforce breaks it: a platform-event trigger “runs as the Automated Process system user” by default (Platform Events Developer Guide), a synthetic entity that community guidance describes as running “in System Context without Sharing, which is the highest amount of access” and which “doesn’t show up in places like Profiles and Permission Sets so it’s hard to assign specific permissions to it” (UnofficialSF).

Delivery is at-least-once, and idempotency is the platform’s job to make achievable. Exactly-once execution across a process boundary is not purchasable at any price. So every run is designed to be safely repeatable, and the platform supplies the primitives — a run-scoped idempotency key, a once() barrier, upsert-by-key writes, and callout keys — rather than leaving each author to invent deduplication.

A job definition is metadata. A job run is record data. The definition says what the work is and how it must behave; the run is one attempt-bearing instance of it moving through states.

Concept Kind Identity Lives in
Job definition Metadata component (job) key — e.g. invoice.sync_to_erp The repository, deployed
Schedule Metadata component (schedule) key, referencing one job The repository, deployed
Job run Record data run_id (uuid) The tenant’s job_run table
Correlation id Field on every run Inherited or minted Threaded into all three audit streams

Every run carries five identifiers, and each answers a different question:

  • run_id — this attempt-bearing instance. Returned synchronously by every enqueue, so the caller always has something to poll, link to, or cancel. There is no enqueue path that returns nothing.
  • job_key — which definition. What the Jobs list groups by.
  • correlation_id — the transaction that caused this work. If the run was handed off from a save’s post-commit phase, it is the save’s correlation id, not a new one. That single inheritance is what makes “a user pressed Submit, and forty minutes later an invoice was wrong” a query rather than an investigation.
  • idempotency_key — the key under which this run’s effects are deduplicated. Stable across retries of the same run.
  • generation — the metadata generation the run resolves against. Assigned at claim, not at enqueue. See Interaction with deploys.

Two distinct people, deliberately kept separate:

  • owner — the human accountable for the definition. Failures route to them. A dead-lettered run is somebody’s problem by construction.
  • runAs — the user identity the run’s queries and writes execute under. It resolves to a real user row with real permission-set assignments, a real role, and real record access.

runAs takes one of three forms, and the default is the honest one:

runAs Resolves to Used for
enqueuer (default) The user whose transaction enqueued the run Post-commit effects, queued one-offs — the work is a continuation of what that user did
owner_of_record The owner_id of the triggering record Work whose natural authority is the record’s owner rather than whoever touched it
service:<key> A named service user — an ordinary user record, provisioned with permission sets like anyone else Scheduled and batch work, which has no originating human

A service user is not a special class of principal. It appears in the user list, holds permission sets, sits in the role hierarchy, owns records that show up in sharing predicates, and can be audited, restricted, and revoked. It cannot log in interactively, and that is its only distinguishing property. Granting it platform.root is possible, loudly audited, and never the default — the platform ships no job that requires it.

A run whose runAs user is deactivated does not silently fall back to a wider identity. It fails with a non-retriable forbidden-class error naming the deactivated user, and the definition’s owner is notified. Widening privilege to keep a job alive is the failure mode this rule exists to prevent.

Five shapes, one primitive underneath. What differs is the enqueue path and the chunking, never the execution engine, the retry semantics, the permission model, or the observability surface.

Kind How a run appears Chunking Notes
Post-commit effect Enqueued to the transactional outbox during the save’s EFFECTS phase, published after COMMIT One run per effect Inherits the save’s correlation id and runAs: enqueuer
Queued one-off jobs.enqueue("key", payload) from any effectful context One run Returns a run_id immediately
Scheduled / recurring Materialized from a schedule component One run per occurrence Occurrence identity is (schedule_key, scheduled_for) — the deduplication anchor
Batch jobs.enqueue on a definition that declares a chunk block One run per chunk, plus a parent run Chunks claim independently; the parent tracks progress
Long-running import / recalculation A batch whose source is an uploaded file or a full-object scan Chunked, checkpointed, resumable Distinguished only by having a resumable cursor and a progress surface

A batch is not a separate mechanism. It is a job definition with a chunk block that names a cursor expression — a keyset over an indexed column, never OFFSET — and a chunk size. The kernel plans the cursor, enqueues chunk runs, and maintains the parent’s progress counters. Chunk failure is isolated: a chunk that exhausts its retries dead-letters as a chunk, and the parent completes with a partial outcome that names exactly which key ranges failed, rather than the whole pass being lost or, worse, silently reported as complete.

queued ──claim──▶ running ──┬──▶ succeeded
▲ │
│ ├──▶ failed ──retry?──▶ queued
│ │ └─exhausted─▶ dead_letter
scheduled ├──▶ cancelled (terminal, retained, surfaced)
│ │
(occurrence due) └──▶ paused ──resume──▶ queued

dead_letter is terminal but not an ending — the run is retained in full, with its payload, its last error envelope, its attempt history, and its correlation id. Nothing is deleted to make a queue look healthy.

A job definition is a canonical component like any other: {key, label, type, body}. Its run body is ordinary code in the one typed language, in the effectful tier — this is a place side effects are authorized, so callouts are legal here and nowhere in the pure tier.

{
"key": "invoice.sync_to_erp",
"label": "Sync invoice to ERP",
"type": "job",
"body": {
"runAs": "service:integration_erp",
"owner": "user:integrations_lead",
"priority": "default", // interactive | default | bulk
"maxAttempts": 8,
"backoff": { "strategy": "exponential", "base": "10s", "cap": "1h", "jitter": "full" },
"timeout": "5m", // per attempt, wall clock
"concurrencyKey": "record.id", // serializes runs touching the same invoice
"idempotency": "record.id + ':' + record.version",
"callouts": { "maxPerRun": 4, "totalTimeout": "60s" },
"enqueues": ["invoice.post_sync_notify"], // static chain edges, checked at deploy
"run": "…"
}
}

A batch adds one block, and the presence of that block is what makes it a batch:

{
"key": "invoice.recalculate_all",
"label": "Recalculate all invoices",
"type": "job",
"body": {
"runAs": "service:pricing_engine",
"owner": "user:pricing_lead",
"priority": "bulk",
"timeout": "10m",
"chunk": {
"over": "invoice",
"where": "record.status != \"Archived\"", // pure expression, same grammar as everywhere
"cursor": "id", // keyset, indexed; never OFFSET
"size": 2000,
"checkpoint": true // resumable; survives pause and metadata change
},
"run": "…"
}
}

A schedule is its own component, so the same job can carry several — a frequent light pass and a nightly heavy one — without duplicating the definition:

{
"key": "sched.nightly_erp_reconcile",
"label": "Nightly ERP reconcile",
"type": "schedule",
"body": {
"job": "invoice.reconcile_with_erp",
"cron": "0 15 2 * * *", // sec min hour dom mon dow
"timeZone": "America/Chicago", // required — an IANA zone, never an offset
"onMissed": "runOnce", // skip | runOnce | runAll
"catchUpWindow": "6h", // bounds runOnce/runAll after an outage
"overlap": "skip", // skip | queue | allow
"enabled": true
}
}

timeZone has no default. A schedule that omits it fails to deploy, with an error that says so. Every hard-won class of “the report ran an hour early for three weeks in March” bug traces to a scheduler that guessed a zone, and the cheapest place to refuse to guess is at compile time.

Enqueuing from code is one verb, and it participates in the enclosing transaction:

// Inside an after_write effect. Enqueue is transactional with the save:
// if the save rolls back, the run was never enqueued.
const run = jobs.enqueue("invoice.sync_to_erp", { invoiceId: record.id });
// run.id is available immediately — link it, poll it, cancel it.

Inside a run body, three platform affordances do the idempotency work:

// 1. A once-barrier: the block executes at most once per key, forever.
once(`erp-post:${payload.invoiceId}:${payload.version}`, () => {
erp.postInvoice(payload);
});
// 2. Writes are upsert-by-key, not blind insert.
upsert("erp_sync_log", { externalKey: payload.invoiceId }, { syncedAt: now() });
// 3. Outbound calls carry the run's idempotency key automatically as a request
// header, so a well-behaved remote can deduplicate a retried delivery.
http.post(url, body); // Idempotency-Key: <run.idempotency_key>

The queue is Postgres, and the claim is FOR UPDATE SKIP LOCKED

Section titled “The queue is Postgres, and the claim is FOR UPDATE SKIP LOCKED”

There is no external broker. The queue is a partitioned job_run table in the tenant’s schema, and workers claim with the standard Postgres queue primitive. The PostgreSQL documentation names this use case directly: with SKIP LOCKED, “any selected rows that cannot be immediately locked are skipped … this is not suitable for general purpose work, but can be used to avoid lock contention with multiple consumers accessing a queue-like table” (SELECT — The Locking Clause).

That is the design, and the reason is not convenience. Enqueue must be transactional with the save that caused it. A record write and the job that follows from it commit together or not at all; there is no window in which a message is published for a transaction that rolled back, and none in which a committed record’s follow-up work was lost in a broker handshake. Salesforce documents the failure this avoids in its own future-method behavior: “Future jobs queued by a transaction aren’t processed if the transaction rolls back” (Apex Developer Guide) — the correct semantic, achieved there by special-casing one mechanism. Here it falls out of the queue being a table in the same database.

The claim is not a naive ORDER BY priority, run_at. That ordering starves tenants. The claim is fairness-first, priority-within:

-- One worker slot, one claim. Fairness is the outer decision.
WITH eligible_tenant AS (
SELECT t.tenant
FROM tenant_queue_state t
WHERE t.runnable_runs > 0
AND t.active_runs < t.concurrency_budget
ORDER BY t.last_served_at ASC -- least-recently-served wins
LIMIT 1
)
SELECT r.*
FROM job_run r
JOIN eligible_tenant e ON r.tenant = e.tenant
WHERE r.state = 'queued'
AND r.run_at <= now()
AND (r.concurrency_key IS NULL OR NOT EXISTS (
SELECT 1 FROM job_run x
WHERE x.tenant = r.tenant
AND x.concurrency_key = r.concurrency_key
AND x.state = 'running'))
ORDER BY r.priority, r.run_at -- priority applies inside a tenant
LIMIT 1
FOR UPDATE SKIP LOCKED;

Three properties follow from that shape:

  • A tenant cannot starve another. Tenant selection happens before priority is consulted, so a tenant enqueueing a million interactive runs consumes its own budget and nothing else’s. Priority is a within-tenant ordering, never a cross-tenant one. This is the single most important structural property of the queue and it is why priority is deliberately weak: making priority global would hand any tenant a starvation lever.
  • Priority is three bands, not a number. interactive (a user is watching a spinner), default, bulk (a batch chunk). Bands are ordered and non-preemptive; a running bulk chunk is never interrupted by an arriving interactive run, it simply loses the next slot. Integer priorities invite an arms race in which every job is priority 1.
  • concurrencyKey serializes without locking. Two runs that would fight over the same record — same invoice, same external account — are prevented from running concurrently by the claim predicate itself, so the second one is left queued rather than deadlocking against the first. Salesforce’s future methods have exactly this hazard and no such guard: “It’s possible that two future methods could run concurrently, which could result in record locking if the two methods were updating the same record” (Trailhead).

A claim is a lease, not a checkout. The claiming update stamps leased_until = now() + timeout. A worker that dies without committing leaves a run whose lease expires, and a reaper returns it to queued. There is no state in which a crashed worker’s runs are lost, and no state in which they are invisible.

Attempts are counted before execution, not after

Section titled “Attempts are counted before execution, not after”

The attempt counter increments in the claim transaction, which commits before the run body starts. A run that hard-crashes its worker — an out-of-memory, a segfault in a dependency, a pathological payload — therefore burns an attempt on every try and dead-letters like any other failure. Incrementing after the body would let a crash-looping poison message occupy a worker slot forever while its attempt count stayed at zero. This is the poison-message defence, and it is why the counter is where it is.

Retry eligibility comes from the error envelope, not from the worker’s guesswork. Every error carries a retriable flag and a class; the worker retries when and only when the envelope says to.

Error class Retried? Why
conflict, unavailable, timeout, rate_limited Yes Transient by definition
internal Yes, with a lower attempt cap Might be transient; might be a bug — bounded either way
validation, forbidden, not_found, precondition No Retrying a rejected payload produces the same rejection at a cost

Backoff is exponential with full jitter — the delay is drawn uniformly from [0, min(cap, base × 2^attempt)). Full jitter rather than a fixed multiplier because the failure that triggers retries is usually a shared dependency, and a fleet of workers backing off on identical schedules re-synchronizes into a thundering herd precisely when the dependency is trying to recover. Both base and cap are per-definition, and a rate_limited error’s Retry-After overrides the computed delay when the envelope carries one.

When attempts are exhausted, the run moves to dead_letter. It is retained, not deleted, and it is loud:

  • The full payload, every attempt’s error envelope, the correlation id, and the pinned generation are kept for the definition’s retention window.
  • The definition’s owner is notified. A dead letter with no addressee is a dead letter nobody reads, so owner is a required field on every job component and a deploy without it fails.
  • The Jobs surface shows a Dead letters count that does not clear itself. Clearing it requires a human decision: replay (re-enqueue with the same idempotency key, so already-completed effects do not repeat), discard (an audited action with a required reason), or replay with an edited payload (a new key, audited as such).
  • A definition whose dead-letter rate crosses a threshold within a window is auto-paused, so a broken integration stops burning the tenant’s throughput budget while its owner is being paged. Pausing a definition never touches runs already in flight.

A job run is record data, so it lives under a retention policy like any other object — one component covering the three stages, shipped as a default on job_run that a tenant edits like any other. The windows and the reason behind each number:

Run outcome Hot Archived Purged
succeeded, cancelled 90 days to 13 months at 13 months
failed, then retried to success On the successful run, same window to 13 months at 13 months
dead_letter, unresolved Indefinitely — the clock has not started — —
dead_letter, resolved by replay or discard 90 days from the resolution to 13 months from it at 13 months from it

90 days hot is the execution-summary window. A run’s value is mostly its trace: the payload says what was asked, and the trace says what happened. Retaining the run longer than the trace it deep-links to produces a stub with a correlation id that resolves to nothing, which reads as data loss rather than as expiry. Thirteen months to purge covers an annual audit that lands weeks after the period it examines, plus a year-over-year comparison of the same month.

A dead letter’s retention clock starts when a human resolves it, not when it failed. An unresolved dead letter is the one run that is definitionally still someone’s open work, and a window that expired underneath it would delete the evidence needed to decide what to do — and would let the Dead letters counter clear itself by attrition, which is the one thing that counter must never do. Replay, discard, and replay-with-an-edited-payload each start the clock, and each is already an audited action with a required reason.

Attempt history is complete, not sampled. Every attempt on a run keeps its full error envelope, and the count is bounded by the definition’s own maxAttempts rather than by a row cap, so there is no run whose earlier failures were dropped to make room for its later ones — the first attempt’s error is frequently the diagnostic one, and a policy that truncated from the front would discard exactly it. A batch parent keeps a per-chunk summary — key range, outcome, attempt count, last error code — after its chunk runs are purged, so the parent stays a complete account of the pass once the chunks have aged out.

One floor bounds what a tenant may configure downward: a run is never purged while it is the most recent run of its definition, because a definition that runs quarterly would otherwise show an empty Last run column between passes and read as broken. Lengthening the windows is cheap for the queue and not for storage — purge of aged runs is a partition drop on the time-partitioned job_run table rather than a row-by-row delete, so a longer window costs bytes and never claim latency.

At-least-once is the delivery contract. Every run must therefore be safely re-runnable, and the platform makes that achievable rather than aspirational:

  • A stable idempotency key per run, derived from the definition’s idempotency expression (a pure expression over the payload) or, when unspecified, from run_id. Retries of a run reuse the key; a replay from the dead-letter surface reuses it too. That is what makes replay safe by default.
  • once(key, body) — a barrier backed by a unique index on (tenant, key) written inside the run’s transaction. The insert and the body’s writes commit together, so the barrier cannot be recorded for work that rolled back. A second execution under the same key skips the body and returns the first execution’s recorded result.
  • Upsert-by-key writes. The write verbs available in a job body are keyed upserts, not blind inserts, so a repeated run converges instead of duplicating.
  • Outbound idempotency keys. Every callout carries the run’s key as an Idempotency-Key header. This cannot make a remote system idempotent — it makes CAOS the party that has done its half, and no more than that.
  • Occurrence identity for schedules. A scheduled run’s key includes (schedule_key, scheduled_for), so two workers that both believe an occurrence is due produce one run, not two.

The honest boundary: exactly-once delivery across a process boundary is impossible, and no amount of platform machinery changes that. What is achievable is exactly-once effect, and it is achieved by deduplicating at the sink — the once barrier for local effects, the idempotency key for remote ones.

A run may enqueue another run. The enqueue participates in the child-producing run’s transaction, exactly as it does inside a save.

Static edges are declared and checked. A definition lists the definitions it enqueues in enqueues[]. That list is a dependency graph the compiler owns:

  • A cycle in the static graph is a deploy error, reported with the full path (a → b → c → a). Not a runtime detection, not a depth counter that eventually trips — a compile failure, the same treatment a cycle among after edges gets in the save order.
  • Enqueuing a definition not in enqueues[] is a runtime error. The graph is not advisory.

One back-edge is legal, and it is bounded. A definition may enqueue itself — the batch-continuation pattern, where a chunk enqueues the next chunk. Self-enqueue is exempt from the cycle check and governed instead by two budgets carried on the run and decremented as it propagates: chainDepth (default 10,000 for a chunk-bearing definition, 20 otherwise) and chainRuns (total runs traceable to one root). Exhausting either fails the run with a non-retriable error naming the root — a runaway chain is a bug that surfaces with a diagnosis, not a queue that quietly grows.

Ordering. Runs within a chain are ordered by causation: a child cannot be claimed before its parent commits, because it did not exist until then. Between unrelated runs there is no ordering guarantee and none is implied — a job that needs “B after A” declares it, and jobs.enqueue(..., { after: runId }) holds the child queued until the named run reaches succeeded. If the dependency dead-letters, the dependent is cancelled with a reason naming it, rather than waiting forever.

A schedule is a cron expression plus a required IANA time zone. The materializer converts occurrences to absolute UTC instants ahead of time and writes them as scheduled runs, so the queue — not a clock-reading loop — is the thing workers consult.

Daylight saving is resolved by policy, not by luck. Converting a local wall-clock time in a named zone to an instant is ambiguous twice a year, and both cases have a stated answer:

DST case What happens locally Platform policy
Spring forward — the scheduled local time does not exist (02:15 in a zone that jumps 02:00 → 03:00) Naive schedulers skip the day entirely, silently The occurrence fires at the first instant after the gap (03:00:00 local). It runs, once, on the correct day.
Fall back — the scheduled local time occurs twice Naive schedulers either fire twice or fire at a shifted local hour The occurrence fires once, at the first occurrence of the local time. The second is suppressed by occurrence identity.
Zone rule change (a government moves or abolishes DST) Previously materialized instants become wrong Materialization horizon is short (7 days) and re-materializes against the current tz database, so already-shipped rules do not fossilize
timeZone: "UTC" No transitions Legal and correct for machine-facing work. It is a choice, never a default that hides the question.

This is not a hypothetical. Salesforce’s scheduled flows store times against a GMT anchor, and the community-documented consequence is that a job scheduled for 02:00 local has no valid moment on the spring-forward date and simply does not run — with, per that write-up, no error raised, “because the platform interprets this as the moment simply never occurring rather than a failure” (Salesforce Ben). A silent skip is the worst available outcome: the work did not happen and nothing said so.

Missed runs after downtime. If the platform, a tenant, or a worker pool was unavailable when an occurrence came due, the schedule’s onMissed decides:

  • skip (default) — the occurrence is recorded as missed and abandoned. Correct for the common case, where the work is a refresh whose next run supersedes it. It is recorded, so “did the 02:00 job run last Tuesday” has an answer.
  • runOnce — all occurrences missed within catchUpWindow collapse into one catch-up run, flagged catchUp: true with the count and the range of skipped occurrences in its payload. Correct for reconciliations, where running once over the whole gap is right and running twelve times is wasteful.
  • runAll — every missed occurrence within the window runs, in order. Correct for period-stamped work where each occurrence produces a distinct artifact.

catchUpWindow is mandatory whenever onMissed is not skip. A week-long outage must not produce a thundering herd of ten thousand catch-up runs the moment capacity returns.

Overlap. When an occurrence comes due and the previous run is still going, overlap decides: skip (default — record the occurrence as skipped-for-overlap and move on), queue (hold it, with a bounded backlog depth; exceeding the depth reverts to skipping and raises an alert), or allow (concurrent runs, legal only when the definition declares no concurrencyKey conflict). A schedule also holds a short occurrence lease, so two workers that both notice an occurrence is due cannot both start it.

Background work is the part of a platform most likely to fail invisibly, so it gets a first-class surface rather than a query an operator must already know to write. The job-monitor surfaces below are package-delivered UI; the durable queue, the fairness claim, and the run records they read stay in the kernel, and everything they show is a plain table anyone can query without them.

  • Jobs — the definitions list: key, owner, runAs, schedule, last run, success rate over a window, dead-letter count, current state (active / paused / auto-paused).
  • Job detail — the run history for one definition: timeline of runs, duration distribution, attempts, failure clustering by error code.
  • Run detail — one run: payload, runAs, state transitions with timestamps, every attempt with its full error envelope, the pinned generation, budgets consumed, and a deep link to the run’s execution trace.
  • Queue health — depth by tenant and band, oldest queued run, claim latency, worker utilization, lease expiries (a rising lease-expiry rate means workers are dying, which nothing else reports as clearly).

Everything on those surfaces is a table, so a CLI view over it would be a plain SQL query. Verified 2026-09-10 — the kernel serves both the read and the control half of that view; the CLI carries the read half.

The kernel now serves two read routes over caos.job_run: jobs/list-runs, a page of runs newest-first, filtered by job key, run state or correlation id and paged by an opaque cursor; and jobs/get-run, one run in full. Both are gated jobs.view — a permission this catalogue had reserved since it was written and which no route consumed until CAOS-1140. Their command-line form ships in the CLI as caos jobs runs and caos jobs show <runId>: the Run detail surface above, in a terminal.

The kernel also serves the operator writes, each a thin route over a primitive the engine already had. jobs/replay-run re-queues a dead letter under its original idempotency key, or under a new one when an edited payload is given, because edited work is different work; jobs/discard-run gives it up with a required reason. jobs/pause, also with a required reason, stops new claims of a job key without touching runs in flight, and jobs/resume lifts a pause of either kind; pausing a key the platform already auto-paused keeps the automatic pause and its reason. jobs/run enqueues one run now, joined by its correlation id to the request that forced it. Replay, discard, pause and resume are gated jobs.manage; forcing a run is gated jobs.enqueue. Replaying or discarding a run that is not dead-lettered is refused naming the state it is actually in, under a row lock, so two operators resolving the same run at once get one success and one refusal.

Still absent: cancelling a live run — there is no cancel primitive, and cancelling a run that holds a lease means deciding what happens to the worker mid-flight — and listing a schedule’s occurrences, since nothing in production materializes a schedule into runs today. The Jobs definitions list (key, owner, runAs, schedule, last run, success rate, dead-letter count, state) reads metadata rather than run records and has no CLI form yet either. Queue health remains UI-only — it is an aggregate over the same table, not a row listing.

Correlation is the whole point. A run’s correlation_id is inherited from the transaction that enqueued it, so one id joins: the error a user saw → the execution trace of the save → the run that save handed off → that run’s own trace → the field changes it wrote in data history → the metadata audit entry for the deploy that changed its behavior. A job’s writes appear in data history with the runAs user as actor and the originating correlation id, so “why did this field change at 3 a.m. with nobody logged in” resolves to a named job, a named run, and a named originating request.

Limits exist. A platform without them is one tenant away from an outage, and Salesforce’s async limits — however much they are complained about — are protecting a real constraint. The difference is where the limits come from: CAOS budgets are per-tenant plan configuration, visible in the run detail and previewable before a job ships, not constants baked into a shared kernel.

Concern Salesforce’s limit (cited) Why it exists there CAOS approach
Total async executions 250,000 or user licenses × 200 per 24 h, whichever is greater — shared across batch, future, Queueable, and scheduled Apex [1] One shared async pool across all tenants Per-tenant throughput budget (runs/hour and runs/day), by plan; visible on Queue health
Concurrent batch jobs 5 queued or active per org [2] Bounds concurrent long-running work on shared infrastructure Per-tenant concurrency budget on slots, not on job kinds; a batch consumes slots like anything else
Overflow queue 100 batch jobs in Holding in the Apex flex queue; FIFO unless an admin reorders [3] A second queue exists because the first one is only 5 deep One queue. Depth is bounded by the tenant’s storage and throughput budget, not by a 100-slot antechamber
Scheduled jobs 100 at one time (5 in Developer Edition) [4] A fixed per-org scheduler capacity No fixed count. Schedules are components; the bound is the throughput budget the occurrences consume
Chained async depth Queueable chains one child per parent; depth 5 in Developer Edition and Trial orgs, unbounded in production [5] Prevents a self-replicating job explosion Default chain depth 50, raisable per run to a hard ceiling of 500 — beats the incumbent’s too-tight default of 5 (which it had to make configurable) without inheriting production’s “unbounded.” Enforced by chainDepth + chainRuns budgets on the run, and a compile-time cycle check on the static graph makes an infinite self-enqueue impossible, not merely capped.
Async resource budget SOQL 200, DML 150, heap 12 MB, CPU 60,000 ms, callouts 100, total execution 10 minutes [1] Protects the shared multitenant substrate Per-run timeout (wall clock, enforced by lease), statement_timeout (CPU), query/DML counters, callout budget — all per-definition within a plan ceiling
Batch chunk size 200 default; max 2,000 with a QueryLocator; QueryLocator returns at most 50,000,000 records [2] Bounds per-transaction work and cursor size chunk.size per definition; the cursor is a keyset over an indexed column, so total record count is not itself a limit
Bulk data loads 150,000,000 records per rolling 24 h; 150 MB per job; 10,000 query jobs per 24 h; results retained 7 days [6] Bounds ingestion against shared storage and processing Import jobs are ordinary batch jobs against the tenant’s throughput budget; results retained per the tenant’s log retention policy, not a fixed 7 days
Event-trigger retries Up to 10 runs (initial + 9 retries); afterwards the trigger “moves to the error state and stops processing new events,” and events sent during that state “aren’t resent” [7] Bounds retry cost on a shared bus maxAttempts per definition; exhaustion dead-letters that run — it never stops the definition from processing unrelated work, and nothing is dropped un-recorded
May May not
Outbound HTTP callouts (post-commit context only), within a per-run callout budget Run past its timeout — the lease expires and the run is terminated
Write records, subject to the runAs user’s full access planes Bypass FLS, record access, or object permissions; there is no system mode
Enqueue declared child jobs, and itself for chunk continuation Enqueue an undeclared definition, or form a static cycle
Read and write large volumes through chunked cursors Hold an unbounded loop — statement_timeout and the lease cap it
Raise a retriable error to request its own retry Extend its own budget, raise its own priority, or opt out of fairness
Log, trace, and emit metrics Assume it runs exactly once

Unbounded loops are not detected — they are bounded. Halting-problem detection is not on offer; a wall-clock lease and a statement timeout are, and they cut the same problem off at a known cost. A run that exceeds either is terminated, marked timeout (retriable, so a genuinely slow-but-progressing job gets another pass), and its trace is retained with the budget line that ran out.

Four escalating controls, each audited:

  1. Cancel a run. Sets a cancellation flag the run body observes at its next checkpoint (chunk boundary, callout boundary, or explicit checkpoint()), letting it stop cleanly and commit partial progress.
  2. Hard-kill. If a cooperative cancel does not land within a grace period, the worker’s backend is terminated. The run’s transaction rolls back; the lease reaper marks it cancelled.
  3. Pause a definition. New runs stop being claimed; in-flight runs finish. This is what auto-pause does when a dead-letter threshold trips.
  4. Drain a tenant. All claiming for one tenant stops. Nothing is lost — the queue simply stops being served, and depth grows visibly on Queue health. Reserved for incidents.

Salesforce splits background work across six mechanisms, each with its own limits, its own monitoring, and its own rules about what may call what.

Future methods (@future) are the oldest. They “can only return a void type,” their parameters “must be primitive data types, arrays of primitive data types, or collections of primitive data types,” they “must be static methods,” and “a future method can’t invoke another future method” (Apex Developer Guide). They return no job id, so a caller cannot track what it started. They are “not guaranteed to execute in the same order as they are called,” and “it’s possible that two future methods could run concurrently, which could result in record locking if the two methods were updating the same record” (Trailhead). Salesforce’s own guidance is to use Queueable instead.

Queueable Apex fixed the worst of that — job ids, non-primitive arguments, chaining — but capped the chain hard: “when chaining jobs, you can add only one job from an executing job with System.enqueueJob,” and “for Developer Edition and Trial orgs, the maximum stack depth for chained jobs is 5, which means that you can chain jobs four times” (Trailhead). Up to 50 may be enqueued from a synchronous transaction; from an async context the governor is 1 (Execution Governors and Limits). The one-child rule prevents a self-replicating job explosion, and it also prevents ordinary fan-out — a job that legitimately needs to start three children cannot.

Batch Apex chunks a QueryLocator (up to 50 million records) into transactions of 200 by default, up to 2,000 with a QueryLocator source, and each execute is a discrete transaction — “if the first transaction succeeds but the second fails, the database updates made in the first transaction aren’t rolled back” (Apex Developer Guide). Chaining is possible only from finish, and only since API 26.0. Concurrency is the sharp edge: “up to 5 batch jobs can be queued or active concurrently.”

The Apex flex queue exists because five is not many. “You can place up to 100 batch jobs in a holding status for future execution”; “up to five queued or active jobs can be processed simultaneously for each org”; a job moved out of the flex queue “changes from Holding to Queued”; and “otherwise, jobs are processed ‘first-in, first-out’” unless an admin reorders them by hand (Salesforce Help). It is a second queue built to hold the overflow of the first.

Scheduled Apex is capped at a flat count: “you can only have 100 scheduled Apex jobs at one time,” the cron format is Seconds Minutes Hours Day_of_month Month Day_of_week Optional_year, “the System.schedule method uses the user’s time zone as the basis of all schedules,” and “actual execution can be delayed based on service availability” (Apex Scheduler). The 100-job ceiling is a widely-documented operational problem — one write-up notes it “can sneak up on you when your org scales” and describes the standard workaround of building a master job that runs every five minutes and dispatches everything else from a custom setting (DZone), and ISVs publish support articles for the resulting error, “You have exceeded the maximum number (100) of Apex scheduled jobs” (Ortoo). A platform capability that customers routinely replace with a hand-built scheduler is a capability that did not fit.

Platform-event-triggered Apex and flows are the modern async-decoupling path, and they carry the identity problem. “By default, the trigger runs as the Automated Process system user with a batch size of 2,000 event messages,” and PlatformEventSubscriberConfig exists to override that so the trigger can “send emails, properly populate OwnerId fields, and create debug logs under that user’s identity” (Platform Events Developer Guide) — an admission that the default identity produces wrong data. Retry is bounded and the exhaustion behavior is severe: a trigger may throw EventBus.RetryableException and “run up to 10 times when it’s retried (the initial run plus 9 retries)”; after that “it moves to the error state and stops processing new events,” and “events sent after the trigger moves to the error state and before it returns to the running state aren’t resent” (Retry Event Triggers). One poison message stops the subscriber, and the events that arrive meanwhile are gone.

Bulk API 2.0 handles large ingestion as its own job system with its own state machine (Open, UploadComplete, InProgress, JobComplete, Failed, Aborted), its own limits — 150,000,000 records per rolling 24 hours, 150 MB per job, 10,000 query jobs per 24 hours — and its own result retention of 7 days (Bulk API limits). It shares nothing with Apex jobs: not the queue, not the monitoring page, not the retry model.

Across all six, a shared 24-hour ceiling applies — “250,000 or the number of applicable user licenses multiplied by 200, whichever is greater” — along with the async governor set: 200 SOQL, 150 DML, 12 MB heap, 60,000 ms CPU, 100 callouts, and a 10-minute total execution cap (Execution Governors and Limits).

And deploys collide with jobs. Changing an Apex class with work pending produces “This Apex class has batch or future jobs pending or in progress,” because “the class is scheduled to run later” and “editing it or its dependencies could cause the class to behave differently from when it was originally scheduled.” The remedy is either to abort the jobs or to enable “Allow deployments of components when corresponding Apex jobs are pending or in progress” (Salesforce Help) — which is to say: block the deploy, or turn the safety off entirely. There is no third option in which the job knows which version it is running.

Where CAOS is genuinely better:

  • One primitive, six shapes. Post-commit effects, one-offs, schedules, batches, and imports are one definition type, one queue, one retry model, one monitoring surface. The mechanism is that chunking and scheduling are fields on a job definition rather than separate subsystems, so a capability added to jobs (idempotency keys, dead-letter replay, correlation) is added to all of them at once.
  • Jobs run as a real user. runAs resolves to a user record with permission sets, a role, and record access, enforced in the query like every other request. The mechanism is the absence of a system mode in the kernel — there is no privileged identity to fall back to, so there is no Automated Process to explain away or override per-subscriber.
  • Per-tenant fairness is in the claim query. Tenant selection precedes priority ordering, so no tenant’s backlog can consume another’s capacity. The mechanism is the outer least-recently-served selection in the SKIP LOCKED claim, not a downstream throttle.
  • Attempts increment before execution. A crash-looping poison message dead-letters instead of occupying a worker forever. The mechanism is counting in the claim transaction rather than in the completion path.
  • A poison message never stops the pipeline. Exhausted retries dead-letter one run; the definition keeps processing unrelated work, and nothing is dropped un-recorded. The mechanism is per-run state rather than per-subscriber state.
  • Cycles are a deploy error. The static enqueue graph is compiled and checked, so a → b → c → a fails at deploy with the path named, instead of being bounded at runtime by a depth counter. The mechanism is enqueues[] being declared metadata rather than an emergent property of code.
  • Time zones are mandatory and DST is policy. Every schedule names an IANA zone; skipped local times fire after the gap and ambiguous ones fire once. The mechanism is materializing occurrences into UTC instants against the tz database with a stated resolution rule, rather than anchoring to a fixed offset and letting the local hour drift.
  • Deploys are never blocked by in-flight jobs, and jobs never straddle a generation. A run pins the generation active when it was claimed; a long batch checkpoints at a generation change. The mechanism is claim-time generation pinning plus resumable chunk checkpoints — see below.

Parity: an asynchronous execution path decoupled from the user’s transaction; cron-shaped recurring schedules; chunked processing of large record sets with per-chunk transactions; retry with backoff; a bulk ingestion path; job monitoring with per-run status and error detail; and hard resource ceilings on async work. All of it is table stakes. Salesforce has every one of these somewhere, and matching them is the floor, not the achievement.

Costs and risks:

  • Postgres as a queue is real operational work. A high-churn job_run table produces dead tuples faster than default autovacuum settings handle. This is answered with time-partitioned tables (completed runs age into partitions that are dropped rather than deleted row-by-row), a partial index on (tenant, state, priority, run_at) WHERE state = 'queued' so the claim never scans history, and per-table autovacuum tuning. It is answered, not avoided, and it must be load-tested at target throughput rather than assumed.
  • The claim query is a hot path with a fairness join. Two-stage claiming costs more than a naive ORDER BY … SKIP LOCKED. The tenant-state table is small and cached, but claim latency under high worker counts is the metric that decides whether this design holds, and it is on Queue health for exactly that reason.
  • Fairness caps peak throughput for a single tenant. A tenant with capacity available elsewhere in the fleet still cannot exceed its own concurrency budget. That is the point, and it is also a real cost: bursty legitimate work runs slower than an unfair scheduler would run it.
  • At-least-once pushes correctness onto authors. The platform supplies once, keyed upserts, and callout keys, but an author who ignores all three will double-post an invoice on the first retry. Mitigations are a compile-time warning when a run body performs a callout with no once barrier and no idempotency expression, and a dry-run replay in the job editor that executes a run twice against a scratch environment and diffs the resulting writes. Neither is proof.
  • Dead letters accumulate into an ignored counter. Every queue system eventually grows a dead-letter list nobody triages. Auto-pause on a dead-letter rate threshold and a required owner make ignoring it cost something, but the discipline is organizational and no design forces it.
  • Cooperative cancellation is only as good as the checkpoints. A run body with a tight loop and no checkpoint can only be hard-killed, losing partial progress. Chunk boundaries give batches natural checkpoints; a hand-written long-running body may have none.
  • DST policy will surprise someone. Firing a skipped 02:15 occurrence at 03:00 is the right default, and it will still be wrong for a job that must not run after 02:30. That is why the resolution rule is documented as a rule rather than left emergent — a stated policy can be designed around; a silent skip cannot.
  • Owning the whole stack means owning the pager. Salesforce operates its async infrastructure for its customers. Worker fleet health, lease reaping, partition maintenance, and backlog alerting are all CAOS’s to run.

A job holding a stale metadata generation is a genuine correctness hazard: a run that resolves fields, formulas, validation rules, and permissions against one generation while the tenant has moved to another will write values the current rules would have rejected. The resolution is precise and has three parts.

1. Runs pin at claim, not at enqueue. A run resolves all metadata against the generation active at the moment a worker claims it, and holds that generation for the whole attempt. Pinning at enqueue was rejected: a run sitting in the queue for four hours would execute against demonstrably stale metadata, and a scheduled definition would be pinned to whatever generation was active when the schedule was authored — it would never see a new field. Claim-time pinning gives a run exactly the same rule a fresh user request gets: you see the generation that is active when your work begins.

2. A run never straddles a generation. Once claimed and pinned, an attempt runs to completion under generation N even if N+1 activates mid-run. This is the same guarantee the save order gives an editing user, extended to background work. The superseded generation stays readable until the last run pinned to it drains — the metadata-retention obligation already stated for interactive sessions, now covering job runs too. A run whose pinned generation has aged out is failed with a retriable error and re-claimed under the current generation; it is never silently upgraded mid-flight.

3. Long batches checkpoint at a generation change. This is where the hazard actually lives — a forty-minute batch is long enough for a deploy to land underneath it. Each chunk claims and pins independently, so chunks naturally track forward; the parent records the generation it started under, and the kernel compares each chunk’s claim-time generation to it:

Change between the batch’s start generation and a chunk’s claim generation Batch behavior
No change Continue
Class A — catalog-only (formulas, validation rules, layouts, permission sets, added automation) that does not touch the objects or fields the batch reads or writes Continue under the new generation; the change is recorded on the parent run
Class A touching objects or fields in the batch’s own read/write set Suspend at the last checkpoint, state paused, reason metadata_changed, naming the components that changed
Class B — physical schema change (column add, type change, constraint, index) affecting the batch’s read/write set Suspend at the last checkpoint, same treatment

A suspended batch is resumable: its cursor position, its accumulated state, and its completed key ranges are all on the parent run, and Resume re-plans against the current generation and continues from the checkpoint. If re-planning is impossible — the cursor column was dropped, the source object retired — resume reports that, and the operator’s choice is restart or cancel. Either way the decision is made by a human who can see what changed, rather than by a job that continued blindly against a schema it was not compiled for.

Deploys are never blocked by in-flight jobs. The Class A / Class B distinction is exactly the one the deploy pipeline already draws, so no new classification machinery is needed — the deploy publishes what it changed, and the batch decides. This is the deliberate inversion of Salesforce’s answer, which blocks the deploy (“This Apex class has batch or future jobs pending or in progress”) until either the jobs are aborted or the safety is switched off org-wide (Salesforce Help). Blocking a deploy on a running batch makes the release process hostage to the job schedule; disabling the check makes stale-metadata execution silent. Generation pinning plus checkpointed suspension is the third option: the deploy lands immediately, and the job knows precisely which version it is running.

Schedules materialize under the current generation. Occurrences are materialized on a short (7-day) horizon, so a deploy that changes a schedule’s cron, zone, or overlap policy takes effect for occurrences not yet materialized, and already-materialized occurrences beyond the change are re-materialized at the flip. Retiring a job definition through safe-delete checks for queued and scheduled runs first and reports them as blockers with their run ids, rather than orphaning work whose definition no longer exists.

Component type Body
Job definition job runAs, owner, priority, maxAttempts, backoff, timeout, concurrencyKey, idempotency, callouts, enqueues[], chunk?, run
Schedule schedule job, cron, timeZone, onMissed, catchUpWindow, overlap, enabled
Service user service_user label, permissionSets[], role? — the identity, not the assignment

Job runs are record data and are never components. They belong to an environment, not to the repository — the same split permission sets draw between a set and its assignments, and sharing draws between a rule and a share. A deploy carries definitions and schedules; it never carries queue contents. Promoting from a sandbox does not promote a backlog.

Compile-time checks that run on every deploy touching a job:

  • owner present and resolvable, runAs present and resolvable.
  • enqueues[] acyclic; every referenced definition exists.
  • idempotency, concurrencyKey, and chunk.where are pure expressions, type-checked against the bound object and recorded in the dependency graph — so a field rename knows which jobs read it, and a blast-radius query includes background work.
  • timeZone present on every schedule and a valid IANA zone name.
  • catchUpWindow present whenever onMissed is not skip.
  • Budgets within the tenant plan’s ceilings; a definition requesting a 6-hour timeout under a 1-hour plan ceiling fails at deploy with the ceiling named, rather than being truncated at runtime.

Salesforce Metadata API analogs, for migration mapping: ApexClass implementing Queueable / Database.Batchable / Schedulable, @future methods inside an ApexClass, Flow of type ScheduledFlow and PlatformEventTriggeredFlow, PlatformEventSubscriberConfig (running user and batch size), and CronTrigger / CronJobDetail for scheduled-job instances — noting that CronTrigger rows are runtime records rather than deployable metadata, which is why a scheduled job commonly has to be aborted and re-created around a deploy.