AI platform
Every app built on Cloud Atlantis OS gets AI without asking for it, for the same reason every app gets validation, audit, and sharing: the capability lives in the kernel, above the tenant’s schema, and is configured rather than integrated. There is no AI feature to bolt on, no per-app model client, no second copy of the schema for the model to read.
Three decisions carry this page.
Grounding is the metadata. The platform already holds a complete, typed, machine-readable description of every object, field, relationship, picklist, validation message, permission grant, and page layout in the org — as canonical components, the same ones the runtime executes. That description is what makes a useful assistant possible with nothing trained, fine-tuned, indexed overnight, or synced. A field created one second ago is groundable on the next request.
Permission-awareness is structural, not procedural. The model never issues a query. It emits a typed tool call; the kernel executes that call on the asking user’s session, through the same three access planes as a list view. There is no AI service role, no integration user, no system mode. A row or column the user could not read on screen is not in the model’s context because the query that would have produced it returned nothing.
A write is a proposal until a human accepts it. AI-originated changes are patches rendered for review, and acceptance is an ordinary save performed by the accepting user — so validation rules, automation, sharing, and save order all apply unchanged. The AI path is not a privileged path; it is the same path with a different author of the draft.
The model
Section titled “The model”The AI kernel is one service with four stages. Apps do not talk to model providers; they call capabilities.
| Stage | What happens | Who controls it |
|---|---|---|
| Ground | Assemble a typed context: a schema projection, tool results, and retrieved document chunks — all filtered by the asking user’s effective access | The kernel, from the component store and the user’s session |
| Invoke | Send prompt + context to a version-pinned model on a configured provider | Org configuration, overridable per capability |
| Constrain | Require a structured result — a tool call or a schema-validated object — and reject anything that does not parse, typecheck, or compile | The capability’s declared output contract |
| Propose | Render the result as an answer, or as a reviewable patch that a human accepts or discards | The user, gated by ai.use / ai.author / ai.approve |
Nothing in this pipeline is optional or per-app. An app that wants a summarize button binds a capability; it does not acquire a model client, a prompt string, a key, or a budget of its own.
Grounding: what the metadata actually supplies
Section titled “Grounding: what the metadata actually supplies”Three grounding inputs, assembled per request, each with a different trust character.
1 — The schema projection. A compact typed rendering of the part of the org the request concerns, generated directly from the component store. It carries, for each object in scope:
- object key, label, and plural label;
- every field’s key, label, type with its declared precision and scale, required-ness, and default;
- picklist fields’ active values only — retired values are not offered, so the model cannot propose a value the platform would reject;
- relationships: target object, cardinality, and whether the relationship is master-detail;
- help text and the human-readable messages attached to validation rules, which are the closest thing the org has to a written statement of its own business rules;
- the sections and field order of the page layout the asking user actually sees — a layout a package delivers as a surface, which the projection reads from the component store rather than from anything compiled in;
- the asking user’s effective permission on each field —
none,read, oredit.
Fields resolving to none are absent from the projection entirely, not marked as forbidden. The model cannot name a column it was never shown, and cannot infer one from a gap.
The projection is scoped, not dumped. A request in the context of a record projects that object, its parents and children one hop out, and any object whose label the request names — resolved against a small label index — subject to a size budget. It is also generation-pinned: the projection is built at one metadata generation, and its hash is written into the audit entry, so “what did the model know” is answerable exactly, later, by anyone with diagnostics.view.
This is the whole reason no training step exists. The assistant does not learn the schema; it reads it, per request, from the authority. There is no drift window between a deploy and the assistant’s understanding of it, because there is no second copy to fall out of date.
2 — Record grounding. Actual values, returned by tool calls the kernel executed, typed. Each value arrives labelled with its field key and type, so the model is not inferring that 4.50 is a currency at scale 2 or that 2026-03-01 is a date rather than a string. Nothing is scraped from a rendered page.
3 — Corpus grounding. Chunks of attachments, notes, and long text, retrieved by vector search over embeddings stored in the tenant’s own schema. Every chunk row carries its tenant, source record id, and owner, and lives under the same row-security predicate as the record it came from — so retrieval is a query the asking user is running, not a lookup into a shared index. Chunks are re-checked against the user’s access at assembly time, after retrieval, because an embedding index is exactly the kind of derived structure that goes stale relative to a share being revoked.
Embeddings exist only for unstructured content. Structured questions are answered by structured queries; there is no reason to approximate a WHERE clause with cosine distance when the platform owns a real one.
Permission-aware by construction
Section titled “Permission-aware by construction”This is the load-bearing guarantee, and it is worth being exact about why it holds, because “the AI respects your permissions” is a sentence every vendor writes and most of them mean as a policy.
The AI service holds no credential of its own. It is a caller like the UI and the API. Every query it makes runs on the session of the user who asked, which means the row predicate from record access and the column projection from FLS apply in the query, in Postgres, on the same single enforcement path everything else uses. There is no elevated mode to reach for and no integration user to misconfigure, because neither exists in the kernel.
The model emits tool calls, never SQL. The tool surface is kernel-owned and small — query records, read one record, list related records, read a file, run a saved report. Every parameter is an object or field key, validated against the schema projection before execution. Since the projection already excludes anything the user cannot read, a request for a forbidden column fails at parameter validation, before any query is built. Two independent checks therefore have to fail before an unreadable value could reach a prompt: the projection filter and the query itself.
Absence, not refusal. When a tool call returns nothing because the rows were filtered, the model receives an empty result — not-found beats forbidden, the same rule the rest of the platform follows. A model told “you are not allowed to see the 4 matching Opportunities” has been handed the count. A model handed an empty set has been handed nothing.
Writes do not get a shortcut. A proposal accepted by a user is that user’s save. It passes object and field write permission, record access, validation rules, and every phase of the save order. An AI-proposed value that violates a validation rule fails exactly as a typed one would, and the failure is returned as an ordinary validation error on the proposal review surface.
The ai domain’s four system permissions. ai.use gates the assistant entry point, ai.author gates letting AI create or modify metadata on the user’s behalf, ai.approve gates approving AI-proposed changes before they deploy, and ai.manage gates configuring the platform itself — providers, models, capabilities, budgets, and evals. They are ordinary boolean grants in a permission set, unioned like any other, with no invisible baseline. An org that wants AI to read and never write assigns ai.use and withholds the other two; that is a complete configuration, not a partial one.
Provenance is queryable, not promised. Because the grounding hash, the tool calls, and the results are all written to the execution trace, “did the model see this customer’s phone number” is a query someone runs. It is not an assurance in a trust document.
Authoring
Section titled “Authoring”Five component types, all ordinary canonical components that retrieve, diff, validate, and apply like everything else.
A provider is a connection: endpoint, credential reference, region, and the retention terms that apply to it. Credentials are references to the secret store, never literals in the body.
{ "key": "ai_provider_primary", "label": "Primary model provider", "type": "ai_provider", "body": { "endpoint": "https://api.example-provider.com/v1", "credential": "secret://ai/provider_primary", "region": "us-east", "retention": "zero", "dataProcessingAgreement": "dpa-2026-01" }}A model is a pin. It names one provider, one model identifier, one immutable version, and the sampling parameters. There is no floating alias — no latest, no unversioned family name — because a model whose behavior can change without a deploy makes every eval result and every audit entry unreproducible.
{ "key": "ai_model_reasoning_v1", "label": "Reasoning model — v1", "type": "ai_model", "body": { "provider": "ai_provider_primary", "model": "example-reasoning-2026-02-11", "temperature": 0, "maxOutputTokens": 4096, "priceInput": "0.000003", "priceOutput": "0.000015", "currency": "USD" }}A prompt is a set of published versions plus a pointer at the one in force. Each version is instructions plus a typed input signature plus an output contract — never a string. Inputs are declared with types, so a prompt that expects an Invoice record cannot be bound to an Account.
Those four things live on the VERSION, not on the component, because they are exactly what a version freezes. Publishing is appending an entry; activating or rolling back is moving activeVersion. Both are ordinary deploys.
{ "key": "prompt_invoice_summary", "label": "Invoice summary", "type": "prompt", "body": { "versions": [ // …versions 1–3, still here and still frozen… { "version": 4, "inputs": [{ "name": "record", "type": "invoice" }], "grounding": { // `schema` names object and field keys, and must resolve at deploy. "schema": ["invoice", "invoice_line", "account"], // `records` names record PATHS rooted at the input names above — resolved // per-caller at invoke time, not against the component store. "records": ["record", "record.lines", "record.account"] }, "output": { "kind": "object", "schema": { "summary": "string", "risks": "string[]", "citations": "fieldRef[]" } }, "instructions": "Summarize this invoice for a sales manager. Cite every number you state by field. If a value is absent from the grounding, say so rather than guessing." } ], "activeVersion": 4 }}A published version is immutable: the deploy refuses a changeset that edits version 4’s contents, drops it, or slots a new version in beneath it. To change the wording, append version 5 and point activeVersion at it.
A capability binds a prompt to a model, a budget, an autonomy level, and the permission required to invoke it. This is the component an app actually calls, and the indirection is what lets an org repoint a capability at a different model without touching an app.
{ "key": "ai_capability_invoice_summary", "label": "Summarize invoice", "type": "ai_capability", "body": { "prompt": "prompt_invoice_summary", "model": "ai_model_reasoning_v1", "requires": "ai.use", "autonomy": "read_only", "budget": { "perInvocation": "0.05", "perUserPerDay": "2.00", "currency": "USD" }, "evals": "ai_eval_invoice_summary" }}An eval suite is the test set that gates the prompt. It is covered under evaluation.
Org-level defaults — the default provider, the default model per capability class, the org budget, and whether unclassified fields may be sent to an external provider — are configuration data, edited on the AI settings page in Setup and deployed like any other configuration.
Semantics & evaluation
Section titled “Semantics & evaluation”The request lifecycle, in order, with the failure at each step named:
- Authorize. The caller’s
ai.use(and, for an authoring capability,ai.author) is checked. Missing →permission. - Budget pre-flight. Estimated tokens × the pinned model’s declared price is checked against the per-invocation, per-user-per-day, and org budgets. Over →
limit.ai_budget, raised before any spend. - Ground. The schema projection is built at the current generation and filtered by effective access; declared record grounding is fetched by tool calls on the user’s session; corpus chunks are retrieved and re-checked. The projection hash is recorded.
- Invoke. One call to the pinned model, with the capability’s declared timeout. Provider failure →
integration, withretriableset from the provider’s response. - Constrain. The result must satisfy the declared output contract. An object output is validated against its schema; a tool call is validated against the tool’s signature and the projection; generated logic is compiled. Failure → one bounded repair attempt carrying the exact compiler or schema error, then refusal.
- Propose or answer. A read capability returns its answer. A write capability returns a proposal record; nothing is written yet.
- Record. Capability, prompt key and version, model key and pinned version, grounding hash, token counts, actual cost, latency, and outcome are written to the execution trace under the request’s
correlationId.
Structured output is the default, prose is the exception. Four of the five shipped capabilities return objects or tool calls, not text. This is a containment mechanism before it is an ergonomic one: a value that arrives as {"field": "close_date", "value": "2026-04-30"} can be typechecked against the field, range-checked, and rejected. The same value inside a paragraph can only be believed.
Generated logic is compiled before it is offered. A capability that produces a formula, validation rule, roll-up input, or automation body emits the one typed language — the same TypeScript-syntax source a person writes, with record and prior bound the same way. There is no separate AI dialect and no natural-language “rule” that gets interpreted at runtime. Before the draft is ever rendered:
- it is parsed and typechecked against the target field’s declared type;
- it is checked against the surface’s tier contract — a validation slot demands pure and boolean, a formula slot demands pure and the field’s type, an automation body permits effects;
- its dependencies are extracted and cycle-checked against the org’s existing dependency graph.
A draft that fails any of these is not shown. The compiler’s error is fed back for one bounded repair attempt; if the second attempt also fails, the capability refuses and says which contract it could not satisfy. The user is never asked to review code the platform already knows it would reject — and, because generation targets the same language everything else is written in, a generated formula is indistinguishable from a hand-written one the moment it lands.
Reproducibility. A pinned model version, a pinned sampling temperature, a hashed grounding snapshot, and a pinned metadata generation together mean an invocation can be replayed. That is what makes an eval suite meaningful and what makes an audit entry more than a receipt.
Refusal is a first-class outcome. Every capability declares its scope. Out-of-scope requests return a typed refusal with a code and a reason, not an improvisation. So does insufficient grounding: "the grounding did not contain a value for this field" is a valid, expected result, and the capability contract has a slot for it. A model that cannot say “I don’t have that” will invent it.
Capabilities the kernel ships
Section titled “Capabilities the kernel ships”Five, each a component binding, each usable from any app.
| Capability | Input | Output contract | Default autonomy |
|---|---|---|---|
| Ask — natural-language record query | An utterance, plus the current context | A query spec (object, filters, fields, sort, limit), executed by the kernel | Read-only |
| Generate logic — field, formula, validation rule, roll-up, automation | An intent plus the target surface | Compiled source in the one typed language | Propose |
| Summarize — a record and its related lists | A record reference | Prose plus required fieldRef citations |
Read-only |
| Extract — document or file into typed fields | A file plus a target object | A typed field patch, per-field confidence, unmapped list | Propose |
| Assistant — in-app conversational surface | A conversation | Composes the above; proposes actions | Propose |
Ask does not answer from the model’s reading of data. It converts the utterance into a query spec, the kernel runs the spec, and the result set is the answer — rendered as a real list, with the spec shown above it and openable as a saved list view. The model’s contribution is the translation, which is checkable; the data is the platform’s, which is authoritative. A user who disagrees with the answer can inspect and edit the query rather than re-phrasing a sentence and hoping.
Extract is where types earn their keep. The output patch must satisfy every target field’s contract: a picklist value must be one of the field’s active values, a currency must fit the column’s declared scale, a lookup must resolve to a record the user can read. Values that cannot be mapped are returned in an explicit unmapped list with the source span, and are never coerced into the nearest legal value. Per-field confidence rides along so the review surface can sort the uncertain to the top.
Summarize is the one prose surface, and it is the one that requires citations. Every number or claim in the summary carries a fieldRef to the field it came from; the review surface renders those as links back into the record. An uncited assertion is treated as a contract violation by the output validator, not as a stylistic lapse.
Assistant composes the others inside one conversation, and its proposals go through the same review as any other. It is not a separate agent runtime with its own permissions.
Actions and consent
Section titled “Actions and consent”A proposal is a record, not a message. When a capability produces a change, the kernel writes an ai_proposal holding:
| Field | What it carries |
|---|---|
capability, promptVersion, modelVersion |
Exactly what produced it |
groundingHash, generationId |
Exactly what it saw, and at which metadata generation |
patch |
The change as a typed patch — a record field patch, or a metadata component diff |
rationale |
The model’s stated reason, for the reviewer, never load-bearing |
basisVersion |
The record version (or generation) the patch was computed against |
correlationId |
Joins it to the invocation, the trace, and any resulting write |
Review is the platform’s existing diff surface. A data patch renders as field-level before → after, the same rendering as data history. A metadata patch renders as the same component diff a deploy shows, because it is one.
Acceptance is an ordinary save by the accepting human. The kernel applies the patch as that user’s write. The actor on the resulting audit entry is the person; proposedBy carries the capability key and the proposal id. Consequences that follow from this and not from policy:
- the write is subject to the accepting user’s object, field, and record access — a user cannot accept a proposal that edits a field they may not edit;
- validation rules, roll-ups, automation, and lifecycle gates all run;
- a metadata proposal deploys through validate → apply, so it can fail deploy validation like any other change.
Stale proposals are refused, not merged. If the record’s version or the metadata generation has moved since basisVersion, acceptance fails with a conflict and the proposal must be re-derived. Accepting a diff computed against data that no longer exists is the single most likely way an assistant quietly destroys a value.
Reversibility. Every accepted action is undoable by construction, because the patch is typed and the prior state is recorded:
- a data write reverses from data history — the inverse patch is derivable, and the review surface offers it for the reversibility window;
- a metadata write reverses by redeploying the prior generation, the ordinary rollback path;
- an effect that left the platform (an email, a callout) is not reversible, which is why capabilities that produce outbound effects cannot hold
auto_applyautonomy at all.
The reversibility window is 30 days from acceptance. The bound is not storage — data history holds the prior values for years, and the inverse patch is derived from it. The bound is the safety of applying an inverse. A month after a value was accepted, the record has usually moved on: other people have edited it, automation has run on it, roll-ups have carried it upward. Beyond that point a one-click inverse stops being an undo and becomes a blind overwrite of everything done since. Inside the window, undo is offered on the proposal and on the record’s history entry; outside it, the change is reversed the ordinary way — a person edits the field, informed by history, through the save order. Undo is itself a save: it runs validation, automation, and the access planes, and it is refused with a conflict if the record’s version has moved since acceptance, under the same basisVersion rule that governs acceptance. A metadata proposal is not bound by this window at all, because its reversal is redeploying the prior generation and its horizon is the metadata-generation retention. auto_apply gets the same 30 days as everything else, which is one of the reasons it is confined to reversible patches.
Proposals are retained on their own schedule, by outcome:
| Proposal state | Retained | Why that long |
|---|---|---|
| Open — neither accepted nor discarded | Expires 30 days after it was produced | A proposal is computed against a basisVersion, so an old one is usually already unacceptable; expiry is a state with a reason, not a silent deletion, and the surface says the proposal went stale rather than showing an accept button that would fail |
| Discarded | 90 days, then purged | What a reviewer rejected is the highest-value evaluation signal the platform produces, and 90 days is long enough to harvest a release’s worth of rejections into eval cases. It matches the execution-trace window, so a discarded proposal and the invocation that produced it expire together |
| Accepted | 13 months, then purged | It is the provenance record for a write while that write is still recent enough to be questioned. After it, the answer comes from the surviving data-history entry, which carries proposedBy — the capability key and the proposal id — for the history stream’s own retention. The patch body was never the authority; the resulting history entry is |
The invocation record outlives all three. Capability, prompt version, model version, grounding hash, token counts, cost, latency, and outcome are execution-trace entries and expire on that stream’s schedule, so spend and usage analysis does not depend on keeping patch bodies around. And because a proposal body can contain record values, purging one is a genuine reduction in exposure rather than a housekeeping saving.
Autonomy has three settings, and the middle one is the default.
| Level | Meaning | Configurable by |
|---|---|---|
read_only |
Never produces a patch | metadata.author |
propose |
Produces patches; a human accepts each one | metadata.author |
auto_apply |
Applies without review, restricted to an explicit allowlist of (capability, object, field) triples, reversible-only, never outbound effects, always within a spend budget | ai.approve |
auto_apply exists because refusing to ship it would push orgs toward scripting around the platform, which is worse. It is deliberately narrow: an allowlist rather than a switch, no effects, and every application still written to the audit streams with the capability as proposedBy and the platform as actor.
Everything is audited. AI activity produces entries in all three always-on streams under one correlationId: the execution trace records the invocation (prompt version, model version, grounding hash, tokens, cost, outcome, refusal reason); data history records field changes from accepted patches; the metadata audit records component diffs from accepted metadata proposals. The prompt text itself is stored on the execution entry — and because that entry can contain record values, reading it honors field-level security: a user who cannot read margin_pct does not see it in an archived prompt either. The audit stream is not an FLS bypass, and AI does not make it one.
Prompt and instruction management as metadata
Section titled “Prompt and instruction management as metadata”A prompt is a deployable component with versions. It is never a string literal in application code, and there is no runtime path that accepts an arbitrary system prompt from a caller.
-
Versions are immutable once published. A published version’s instructions, output contract, and grounding declaration are frozen; changes create a new version. The component carries an
activeVersionpointer, and moving that pointer is a deploy. -
Promotion is a deploy. Prompt changes ride the same retrieve → diff → validate → apply pipeline as a field or a layout, through the same environments. A prompt is tested in a lower environment and promoted; it is not edited in production and hoped over.
-
Prompts participate in the dependency graph. A prompt names object and field keys in the
grounding.schemaof its ACTIVE version, so where-used and blast-radius analysis cover it. Renaming a field knows which prompts read it, and a deploy that would orphan a reference fails validation rather than producing a prompt that silently grounds on nothing.Two things are deliberately NOT edges, and both are worth knowing before you rely on this. Instructions are prose, not references — merge-style mentions inside the instruction text are not scanned, because matching English against component keys makes ordinary words like
summaryortoneinto pointers and permanently blocks deleting whatever they collide with. Only the active version’s grounding counts — an archived version can name an object the org has since removed without blocking its deletion, because a published version can never be edited and enforcing it forever would leave no legal repair. Integrity is re-checked when the pointer moves instead: rollingactiveVersionback onto a version whose grounding no longer resolves is refused by the deploy that moves it. -
Retirement routes through safe-delete. A prompt with a live capability binding cannot be deleted, only retired.
-
Localization is a variant, not a fork. Instructions carry per-locale variants under one component key, so a translated prompt cannot drift out of structural sync with the original.
The practical consequence: “why did the assistant answer that way in March” is answerable, because the March prompt version, the March model pin, and the March grounding are all still on disk and all still joinable by correlationId.
Evaluation and safety
Section titled “Evaluation and safety”An eval suite is a component, and it gates the deploy. ai_eval holds cases — fixed inputs with a frozen grounding fixture — and assertions of three kinds:
| Assertion | Checks | Gating |
|---|---|---|
| Structural | The emitted tool call, its target object, and its arguments match expectation | Hard — any failure blocks |
| Typed | The output satisfies the declared schema; generated logic compiles under its tier contract | Hard — any failure blocks |
| Judged | A rubric scored by a pinned judge model, aggregated across cases | Threshold — the suite declares its floor |
Publishing a prompt version runs its suite; a failed suite blocks the deploy exactly as a failed validation blocks a save. Structural and typed assertions are deterministic and therefore fully gating. Judged assertions are not deterministic, so they gate on an aggregate score with a declared floor, and the judge model is pinned so that “the score moved” means the prompt moved.
A model change is a deploy with tests. Because a model component pins an immutable version, a provider shipping a new model does not change org behavior. Repointing to a new version is a metadata change that re-runs every eval suite bound to that model, in a lower environment, before promotion. Silent vendor upgrades — the most common cause of “it worked last week” in an AI feature — cannot occur, and the cost of that guarantee is that upgrades are deliberate work someone has to schedule.
Hallucination containment is mechanical. Four mechanisms, in order of how much they carry:
- Structured tool calls over free text. The model chooses among typed operations; it does not narrate data. A tool call is validated before it runs.
- Closed vocabularies. Picklist values, object keys, and field keys come from the projection, so the legal set is finite and enumerated. An out-of-set value fails validation rather than reaching a record.
- Citations bound to identifiers. Prose capabilities must attach
fieldRefcitations; the validator rejects uncited assertions, and the reviewer can click through to the source. - “No answer” as a supported output. Insufficient grounding returns a typed refusal. The contract has a slot for not knowing, which is the cheapest hallucination prevention available and the one most often omitted.
Refusal behavior. A capability refuses when the request falls outside its declared scope, when grounding is insufficient, when its output cannot be made to satisfy its contract in one repair attempt, or when a budget is exhausted. Each refusal is a typed error envelope with a code and a fault — a scope refusal is fault: user, a budget refusal is fault: admin, a provider failure is fault: platform — so the UI can say the right thing without the model composing an apology.
PII and what leaves the tenant
Section titled “PII and what leaves the tenant”Stated plainly, per capability class:
| Capability class | What leaves the tenant | What never leaves |
|---|---|---|
| Metadata authoring (generate logic, build an object) | The schema projection: object, field, and type names, labels, help text, picklist values | Record values — these capabilities do not ground on data at all |
| Record capabilities (ask, summarize, extract, assistant) | The record values in the grounding, as the asking user could read them | Values in fields the user cannot read; values in fields classified aiVisibility: never |
| Embeddings | Text sent once to an embedding model, if the embedding model is external | The vectors and chunks themselves, which are stored in the tenant’s own schema |
Three controls, and they are the honest ones:
- Field classification. The same classification that drives environment masking drives AI exposure. A field tagged
aiVisibility: neveris excluded from every projection and every tool result, regardless of who is asking — a plane above FLS, because it restricts a value the user may read from being sent outward. The org sets a default for unclassified fields, and that default is a deliberate configuration choice, not an accident. - Provider selection and residency. The provider component carries region and retention terms, and a capability can be pinned to a provider — so a capability handling regulated data can be routed to an in-region or in-tenant endpoint while the rest of the org uses a general one.
- In-tenant models. The provider abstraction accepts a self-hosted endpoint. In that configuration nothing leaves at all, at the cost of running the model. This is the only configuration in which “no data leaves the tenant” is a true statement, and it is stated as such rather than approximated by masking.
Masking is deliberately not the answer here. Redacting entities out of a prompt and re-inserting them afterward degrades exactly the capability being paid for — the vendor with the most-marketed masking layer disabled it for its own agents in order to recover accuracy, and documents a smaller usable context window when it is on (see below). Classification, provider choice, and residency are structural; masking is a heuristic that costs quality and provides a guarantee only as good as its entity detector. Where an org still wants it, it is a per-provider transform, off by default, with its accuracy cost stated in the setup surface rather than in a footnote.
Cost governance
Section titled “Cost governance”Budgets are denominated in currency, not in a platform-defined unit. A model component declares its per-token prices; a capability declares its budgets in the org’s currency; the kernel projects the cost before the call and records the actual after it. An abstract “credit” whose exchange rate the vendor sets and revises is not a budget an organization can plan against, and the industry has an instructive example of that (see below).
| Control | Scope | Enforcement |
|---|---|---|
perInvocation |
One call | Pre-flight cost check; over → limit.ai_budget, nothing spent |
perUserPerDay |
One user, rolling day | Pre-flight against recorded actuals |
| Org budget | Whole tenant, per period | Warn at a declared threshold; hard stop at the ceiling |
| Rate limit | Per user and per capability | Concurrency and requests-per-minute ceilings, independent of spend |
Spend is not a separate dashboard. Cost per invocation lands on the execution trace, which is a queryable table, so spend by capability, by user, by object, by day, or by individual record is an ordinary SQL query and an ordinary report. A hard stop returns a typed limit error with fault: admin and a message naming which budget was exhausted, so the user is told what happened rather than watching a feature quietly stop working.
Limits, and the reasons behind them
Section titled “Limits, and the reasons behind them”| Concern | CAOS approach | Salesforce | Why theirs is shaped that way |
|---|---|---|---|
| Grounding source | The live component store, filtered by effective access, generation-pinned | Prompt Builder merge fields, Flow/Apex data providers, and Data Cloud retrievers over a vector search index | Grounding was built as a retrieval product on top of a separate data platform, not as a projection of the schema |
| Access enforcement | Every grounding query runs on the asking user’s session; no service role exists | Documented as respecting permissions and sharing per feature and per grounding path | Enforcement is per-integration-point rather than a single query path |
| Masking | Not used as the primary control; classification, provider selection, and in-tenant models are | Pattern-, ML-, and field-based masking — disabled for Agentforce agents to protect accuracy; 65,536-token context ceiling when enabled via the Models API | Masking was the promised control, and it conflicts with the accuracy of multi-step agents |
| Model version | Immutable pin per model component; no floating aliases | Model chosen per prompt template version (primaryModel); BYO LLM via LLM Open Connector |
Model choice is versioned with the template, which is close to right |
| Prompt versioning | Immutable published versions, activeVersion pointer, deploy-gated by evals |
GenAiPromptTemplate + GenAiPromptTemplateVersion, status Draft/Published, activeVersionIdentifier |
Genuinely good; this is table stakes and is copied |
| Generated logic | Must compile under a tier contract before it is offered | Generated Apex/formula is reviewed by a human; no purity or tier contract exists to check it against | There is no analyzable tier in the incumbent to compile into |
| Write consent | Proposal → human accept → ordinary save; auto_apply allowlisted and reversible-only |
Agent actions execute; Agent Script exposes a confirmation property on reasoning actions | Actions were designed to run, with confirmation as an option |
| Cost unit | Currency, with per-token prices declared on the model component | Flex Credits — $500 per 100,000, 20 credits (~$0.10) per action | A fungible credit spans agents, data operations, and BYO-LLM prompts, which is flexible for the vendor and hard to forecast for the buyer |
| Budget enforcement | Pre-flight cost check, hard stop with a typed error | Consumption billed; ceilings are commercial, not runtime | Consumption is a revenue model before it is a control |
How Salesforce does it
Section titled “How Salesforce does it”Salesforce’s AI stack is four products stacked: the Einstein Trust Layer as a gateway, Prompt Builder for templates, Agentforce for agents, and Data Cloud for grounding — with Model Builder / BYO LLM underneath for model choice and the Agentforce Testing Center beside it for evaluation.
The Trust Layer is a gateway every generative request passes through. Its published stages are secure data retrieval, dynamic grounding, data masking, prompt defense, a zero-retention agreement with model providers, toxicity detection, and an audit trail. Salesforce’s own engineering write-up describes client-side grounding “when a prompt is being selected in the context of a record page” and server-side grounding “when a response is being generated behind the scenes”; masking that replaces each detected entity with “a combination of its type and a sequential number. For example, the first detected name becomes PERSON_0”; prompt defense via “instruction defense” and “post-prompting instructions”; and toxicity scoring across seven categories producing “an overall safety score from 0 (least safe, most toxic) to 1 (most safe)” (Inside the Einstein Trust Layer). The audit trail logs timestamped metadata including the original prompt, safety scores, and user feedback.
Masking is the stage that did not survive contact with agents. Salesforce’s own developer documentation for the Models API states that masking covers “selected personally identifiable information (PII) and payment card industry (PCI) data,” requires the caller to “specify the correct locale in your API request,” acknowledges that “no model can guarantee 100% accuracy,” notes that “cross-region and multi-country use cases can affect the ability to detect specific data patterns,” and — the hard constraint — that when data masking is enabled, “all models are currently limited to a context size of 65,536 tokens” (Data Masking — Agentforce Developer Guide). Salesforce Help goes further for agents specifically: pattern-based and field-based data masking for LLMs is disabled for agents, stated as being done to improve the performance and accuracy of agents (Data Masking Limitations in Agentforce). The most-marketed privacy control is off in the flagship product.
Prompt Builder is the strongest piece, and it is metadata. GenAiPromptTemplate carries masterLabel, type, relatedEntity, relatedField, visibility, and activeVersionIdentifier; GenAiPromptTemplateVersion carries versionIdentifier, versionNumber, status (Draft or Published), content, primaryModel, inputs, templateDataProviders, and generationTemplateConfigs (Metadata API — GenAiPromptTemplate). Versioned, deployable, model-per-version: correct, and CAOS copies the shape.
Agentforce organizes an agent into topics and actions, where an action calls a Flow, a prompt template, or an Apex class, and the Atlas reasoning engine plans across them. Agent Script declares tools in a subagent’s reasoning.actions block — “executable functions that the LLM can choose to call, based on the tool’s description and the current context” — binds parameters with with, assigns outputs with set, supports available when to “deterministically specify when the tool is available,” and exposes a requiresConfirmation property on reasoning actions (Tools (Reasoning Actions) — Agentforce Developer Guide). Confirmation is available; it is a property on an action rather than the shape of the write path.
Data Cloud grounding works by building a search index that chunks source data, embeds it, and stores it in a Vector Data Model Object, then defining a retriever over that index which a prompt template invokes; a retriever version must be activated before templates can use it (Retrieving Content with Vector Search). This is a competent RAG pipeline — and it is a pipeline, with indexes to build, activate, and keep fresh, sitting on a separately licensed data platform.
Model choice runs through Model Builder / BYO LLM. External models from OpenAI, Azure OpenAI, Amazon Bedrock (Anthropic Claude), and Google Vertex AI can be connected, and Salesforce states that “all inference requests from your external model are routed through the LLM Gateway and Einstein Trust Layer before surfacing content to your users” (Bring Your Own LLM in Einstein 1 Studio).
Evaluation is the Agentforce Testing Center: batch tests from a CSV of utterances with expected topic, expected actions, and expected response, scored on accuracy plus completeness, coherence, conciseness, latency, and instruction adherence (Agentforce Testing Center). It is a good tool. It is not a deploy gate — testing is something a team does, not something publication requires.
Pricing moved twice. Agentforce launched at $2 per conversation, then in May 2025 moved to Flex Credits at $500 per 100,000 credits, with “one Agentforce action consumes 20 Flex Credits ($0.10 USD),” plus 100,000 credits included at no cost for Enterprise Edition customers with Salesforce Foundations (Constellation Research). Credits are fungible across agent actions, Data Cloud operations, and BYO-LLM prompts, which is what makes forecasting hard: a single user request can consume several actions, and the unit is the vendor’s, not the buyer’s.
The accuracy record, from Salesforce’s own research. CRMArena-Pro, published by Salesforce researchers, evaluates LLM agents across nineteen expert-validated CRM tasks. Leading agents achieve “around 58% single-turn success on CRMArena-Pro, with performance dropping significantly to approximately 35% in multi-turn settings,” workflow execution is the tractable case at over 83% single-turn success, and — the finding that matters most for a platform claiming permission-awareness — “agents exhibit near-zero inherent confidentiality awareness; though targeted prompting can improve this, it often compromises task performance” (CRMArena-Pro, arXiv:2505.18878). An agent asked politely to respect confidentiality gets worse at its job. That is the empirical case for enforcing access in the query rather than in the instructions.
Where CAOS is genuinely better:
- Grounding is a projection of the live component store, not an index to build. No search index to create, activate, and refresh, and no window in which the assistant’s view of the schema lags a deploy — the projection is generated at a pinned metadata generation on every request.
- Enforcement is the query, not the integration point. Every grounding read runs on the asking user’s session under the same row predicate and column projection as a list view, because the AI service holds no credential of its own. Confidentiality is not something the model is asked to observe.
- The model cannot name what it was not shown. Fields resolving to
noneare absent from the projection rather than marked forbidden, so there is no gap to reason about and no denial to leak a fact. - Generated logic is compiler-checked before it is offered. Output targets the one typed language under a tier contract, so a draft that would not typecheck never reaches a human. There is no analyzable tier in the incumbent to check generated Apex against.
- Writes are proposals executed by a human, so they inherit every guarantee. Acceptance is that user’s save — validation, automation, sharing, and save order all apply, and
basisVersionrefuses a diff computed against state that has moved. - Evals gate publication. A prompt version cannot be published with a failing suite, and repointing a model re-runs every suite bound to it. Testing Center is a tool beside the pipeline; this is in it.
- Models are pinned, so behavior is reproducible. No floating aliases means no silent vendor upgrade, and an audit entry from March replays.
- Budgets are in currency. Per-token prices on the model component, pre-flight cost checks, hard stops with typed errors, and spend as a queryable column — rather than a fungible credit whose conversion the vendor controls.
- Nothing is a separate SKU. Grounding does not require buying a data platform, and audit of AI activity is part of the always-on log streams.
Parity: a gateway all model traffic passes through; prompts as versioned, deployable metadata with a model bound per version; grounding that mixes schema, record data, and retrieved documents; connectable external models from multiple providers; zero-retention agreements with providers; toxicity and prompt-injection defenses; batch evaluation of prompts and agents; an audit trail carrying the prompt, the model, and the outcome. Salesforce got the shape of most of this right, and copying it is the correct move.
Costs and risks:
- The projection is a real token cost on every request. Scoping, a size budget, and a label index keep it bounded, but a wide org with hundreds of fields per object will spend meaningfully on context before the question is even asked. This must be measured per capability, not assumed.
- Pinning models means someone owns upgrades. Immutable versions buy reproducibility and cost the work of scheduling migrations, re-running suites, and eventually being forced off a deprecated version by a provider on the provider’s timetable.
- Judged evals are a soft gate. Structural and typed assertions are deterministic; rubric scores are not. A pinned judge model reduces drift but does not remove it, and a threshold that is too tight blocks good changes while one that is loose blocks nothing.
- The provider boundary is not solved, only bounded. Classification, residency, and in-tenant endpoints are the honest controls, and the strongest of them — running the model in-tenant — carries real infrastructure cost. Any other configuration means readable values leave.
auto_applyis a genuine risk surface. Narrow allowlists, reversibility, and no outbound effects contain it; they do not eliminate the case where a reversible-but-unnoticed write propagates through automation before anyone reviews it.- Corpus grounding has a staleness seam. Embeddings are derived data, and revoking access does not un-embed a chunk. Post-retrieval re-checking closes the read path, and re-indexing on access change is still work the platform has to do and to monitor.
- Structured output constrains what the assistant can do. Requiring tool calls and citations makes some genuinely useful open-ended answers unavailable. That is the trade being made deliberately, and users will notice it.
- Refusal has a cost in perceived quality. A capability that says “the grounding did not contain that” is more trustworthy and less impressive than one that guesses. Expect the comparison to be made against products that guess.
Metadata & deploy representation
Section titled “Metadata & deploy representation”| Component | type |
Body |
|---|---|---|
| Provider connection | ai_provider |
endpoint, credential (secret reference), region, retention, dataProcessingAgreement |
| Model pin | ai_model |
provider, model, immutable version, sampling params, per-token prices, currency |
| Prompt | prompt |
versions[] (each frozen once published: inputs[], grounding, output contract, instructions, locale variants) and activeVersion |
| Capability binding | ai_capability |
prompt, model, requires, autonomy, budget, evals |
| Eval suite | ai_eval |
cases[] with frozen grounding fixtures, assertions[], thresholds |
Org defaults — default provider, default model per capability class, org budget, and the exposure default for unclassified fields — are configuration data rather than components, because they are per-environment values rather than per-environment shape. A proposal (ai_proposal) is record data, never a component: it is the output of an invocation, not a declaration.
Deploys behave like any other. A prompt or capability change is metadata-only and takes effect at the next generation flip. Publishing a prompt version runs its eval suite as part of deploy validation, so a failing suite surfaces as a deploy error with the failing cases in details. Retiring a prompt or model that a capability still binds routes through safe-delete and fails validation rather than leaving a dangling reference.
Salesforce Metadata API analogs, for migration mapping: GenAiPromptTemplate and GenAiPromptTemplateVersion map to prompt; GenAiPlannerBundle / agent and topic definitions map to ai_capability bindings; connected model configurations from Model Builder map to ai_provider + ai_model; Testing Center test definitions map to ai_eval.
Sources
Section titled “Sources”- Inside the Einstein Trust Layer — Salesforce Developers — trust-layer stages, client- vs server-side grounding,
PERSON_0masking placeholders, instruction defense, the 0–1 safety score across seven toxicity categories, audit-trail contents. - Data Masking — Agentforce Developer Guide (Models API) — PII/PCI scope, locale requirement, “no model can guarantee 100% accuracy,” cross-region detection caveats, and the 65,536-token context ceiling when masking is enabled.
- Data Masking Limitations in Agentforce — Salesforce Help — pattern-based and field-based masking disabled for agents, to improve agent performance and accuracy.
- GenAiPromptTemplate — Metadata API Developer Guide — template and version fields, Draft/Published status,
activeVersionIdentifier,primaryModel,templateDataProviders. - Tools (Reasoning Actions) — Agentforce Developer Guide —
reasoning.actions, tool descriptions driving LLM selection,with/setbinding,available when,requiresConfirmation. - Retrieving Content with Vector Search — Salesforce Help — search index, chunking and embedding into a Vector Data Model Object, retrievers, and activation before use in prompt templates.
- Bring Your Own Large Language Model in Einstein 1 Studio — Salesforce Developers — connectable providers and the statement that external-model inference routes through the LLM Gateway and Trust Layer.
- Agentforce Testing Center — Salesforce Help — batch testing of agents against expected topics, actions, and responses.
- CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions — arXiv:2505.18878 — ~58% single-turn and ~35% multi-turn success, >83% on workflow execution, and near-zero inherent confidentiality awareness that targeted prompting improves only at a cost to task performance.
- Salesforce revamps Agentforce pricing with Flex Credits — Constellation Research — $2 per conversation previously; $500 per 100,000 Flex Credits; 20 credits (~$0.10) per action; 100,000 credits free with Salesforce Foundations. (Analyst coverage, May 2025.)
- Agentforce Flex Credits Explained — Nebula Consulting — what consumes credits, multi-step conversations summing several actions, and unmetered per-user pricing as the predictability alternative. (Secondary.)
- The Doomed Evolution of Salesforce’s Agentforce Pricing — Monetizely — documented buyer complaints about undefined units, forecasting difficulty, and cost as an adoption barrier. (Secondary, opinion.)