Skip to content

Search & indexing

Search is the surface where a permission model usually leaks. The classic architecture copies record content into a separate search tier, ranks it there, returns the top n, and then filters the survivors against the user’s access. That order is backwards, and every symptom follows from it: a result cap consumed by rows the user was never allowed to see, a count that reveals how many records matched, a snippet drawn from a column the user cannot read, and a second datastore whose freshness is somebody else’s problem.

CAOS does not have a search tier. Search is a query over the same tables, through the same row-security policies and the same field-permission resolution as every other read. The index is a set of Postgres indexes attached to the tenant’s own relations, so access filtering happens inside the plan, as predicates the planner uses to choose an access path — before ranking, before limiting, before any row becomes a result.

The load-bearing decision on this page: a searchable value is stored in a column whose visibility is already decided. Field-level security is not applied to search output; it decides which indexed columns the query is allowed to touch. A user cannot match on a value they cannot read, so there is no title, no snippet, no rank contribution, and no count that could betray one.

Layer What it does Where it lives
Search document The tokenized form of a record’s searchable content tsvector columns on the object’s own table, GENERATED ALWAYS AS … STORED
Visibility band A partition of an object’s fields by identical field-permission profile One search-document column per band
Derived content Text that is not on the record — extracted file text, parent labels A side table keyed by record id, joined through the base table
Lexical index Word matching, stemming, phrase, rank GIN on each band’s tsvector
Fuzzy index Substring, typo, prefix pg_trgm on identifier and name columns
Query UNION ALL over the objects the user may read, blended and ordered Generated per request from metadata

There is one mechanism here that does not exist elsewhere and carries the whole security argument, so it is worth stating plainly before anything else.

A record does not get one search document. It gets one per band, where a band is an equivalence class of the object’s fields under field-level security: two fields are in the same band when the set of permission sets granting Read on them is identical.

Bands are derived, not authored. At deploy time the kernel reads the fieldPermissions rows across the whole permission catalog, groups the object’s searchable fields by their grant signature, and emits one generated tsvector column per group. Real objects collapse to very few — a public band, one or two restricted bands, sometimes a band of one for a genuinely singular column like margin.

At query time the effective-field-permission resolver says which bands the user can read, and only those columns appear in the query:

-- a user who holds Read on the public band and the pricing band, and nothing else
WHERE (invoice.search_b0 @@ q OR invoice.search_b2 @@ q)

A lexeme from a band the user cannot read is not in the query, so it cannot match, cannot rank, cannot be counted, and cannot be excerpted. The guarantee is structural rather than procedural: there is no filtering step to forget, no result to strip, no “protected fields” list with a ceiling on it. Contrast this with the incumbent, which applies field-level security to a bounded number of fields per object and leaves the remainder able to influence results (see How Salesforce does it).

Record access needs no new mechanism at all. The search columns are columns on the object’s own row, so the object’s row-security predicate governs them exactly as it governs every other column. A record the predicate excludes never enters the candidate set.

Global search is a UNION ALL over one subquery per participating object. An object participates when it is marked searchable and the user holds object Read on it — nothing else. Each branch is independently row-filtered and band-filtered, produces a uniform (object, record_id, title, rank, matched_field) shape, and the outer query blends and orders them.

Per-object search — the box above a list view — is the same generated subquery with the list’s own filter, sort, and scope ANDed in. It is not a separate code path and not a client-side filter over a fetched page: it searches the whole object within the list’s criteria, so a record on page 40 of a list is findable from page 1.

Lookup search is the same subquery again, scoped to the reference field’s target object with the reference’s own filter ANDed as an ordinary predicate. See Lookup search.

Searchability is a property of the field, decided by its storage primitive and overridable per field.

Primitive Default How it is indexed
text searchable Full lexical: stemmed, stop-worded, phrase-capable. Weight from the field’s role
enum searchable The active label, unstemmed
reference searchable The target record’s title, in the reference field’s band
number searchable Canonical text form as a single unstemmed token — 13315 matches, 133 does not, except through the fuzzy path
temporal searchable ISO form as a single token, plus the tenant-locale rendering
boolean not indexed A term matching half the table is noise, and a boolean is a filter, not a search
json not indexed Opt in per JSON path; indexing a whole document indexes its keys as well as its values
geo never indexed Spatial predicates are a different query shape, not a text match

Two consequences worth naming, both of which are places the incumbent stops:

  • Computed fields are searchable. A formula field and a roll-up summary are stored computed columns, so they are indexable like any other column with no special handling. A record findable by its calculated tag is findable by search.
  • Numbers and dates are searchable. They are the identifiers people actually type. Indexing them as exact tokens gives precise matching without polluting the lexical corpus with digit fragments.

Opting out is a field property: searchable: false. It removes the field from its band’s generated column, which is a physical change to the table, so it deploys as an expand → migrate → contract migration and takes effect at the generation flip, not before.

Sensitive fields are excluded, not banded. A field marked sensitive is absent from every search document. A tsvector is a decomposition of a value, and a value that must not leak must not be decomposed — even into a band nobody currently reads, because band membership is derived from a permission catalog that can change.

An object’s search behavior is one canonical component. Bands are not in it — they are derived from field permissions, and an admin who could hand-assign them could hand-assign a leak.

{
"key": "search_invoice",
"label": "Invoice — search",
"type": "search_settings",
"body": {
"object": "invoice",
"searchable": true,
"title": "invoice.invoice_number",
"subtitle": ["invoice.account_id", "invoice.status"],
"fields": [
{ "field": "invoice.invoice_number", "weight": "A", "fuzzy": true },
{ "field": "invoice.name", "weight": "A", "fuzzy": true },
{ "field": "invoice.account_id", "weight": "B" },
{ "field": "invoice.scope_notes", "weight": "C" },
{ "field": "invoice.margin_pct", "searchable": false }
],
"derived": [
{ "source": "attachment_text", "weight": "D" }
],
"typeWeight": 1.4,
"recencyHalfLife": "P30D"
}
}

weight maps onto the four tsvector weight labels — A for titles and identifiers, B for short descriptive text, C for long prose, D for derived content — which is what makes a match on an invoice number outrank a match buried in a scope note. fuzzy: true adds a trigram index for typo and substring matching on that column; it is opt-in per field because a trigram index on a long text column is expensive and rarely worth it.

A saved search is record data, not metadata — it belongs to the user who saved it and is shared through the ordinary record-access plane. It becomes a component only when an admin promotes it, at which point it is a list view; a saved search and a list view are the same object at two points in its life, and modeling them as two things would mean two query builders.

{
"key": "listview_invoice_open_west",
"label": "Open invoices — West",
"type": "list_view",
"body": {
"object": "invoice",
"filter": "record.status != \"Paid\" && record.region == \"West\"",
"columns": ["invoice.invoice_number", "invoice.account_id", "invoice.total_price"],
"searchable": true
}
}

Three planes, all applied in the plan, in this order:

  1. Object permission decides whether the object contributes a branch to the UNION ALL at all. No Read, no branch — the object is not searched, and its absence is indistinguishable from having no matches.
  2. Record access is the object’s row-security predicate, evaluated as the branch’s USING clause. Cheapest clauses first, short-circuiting, exactly as described in Roles & record sharing.
  3. Field access decides which band columns the @@ operator is applied to.

Ranking runs after all three, on the surviving rows only. Limits apply after ranking. So the ordering is filter → rank → limit, and every user-visible number — the result count, the per-object count, the “showing 25 of 340” — is computed over the filtered set. There is no arithmetic a user can do on a search page that yields a fact about a record they cannot read.

platform.root short-circuits all three, so root searches every band of every object with nothing to enumerate.

Not-found beats forbidden, consistent with the error model: a filtered record is absent from results, never present-but-refused. A search page that said “3 results you may not view” would be a disclosure surface with a polite label.

Explaining a match. Because the query names which band column matched, a result can report why it matched — the field, and the excerpt from it. The same machinery answers the inverse question in Setup: given a term and a user, which objects were searched, which bands were readable, and which predicate excluded a record. Search that cannot explain itself is indistinguishable from search that is broken.

A record’s own fields are indexed transactionally. The search documents are GENERATED ALWAYS AS (to_tsvector('english', …)) STORED columns, so Postgres recomputes them inside the statement that writes the row and the GIN indexes are updated in the same transaction. The write and its index entry commit together or not at all. The latency between a save and its searchability is zero, and there is no window in which a committed record is unfindable.

This is the decision, and it is deliberate. Eventual consistency here is not a performance optimization with an acceptable cost — it is a correctness hazard that shows up as a support ticket. A user who saves a record, searches for it, does not find it, and creates it again has produced a duplicate that the platform will carry forever. Paying the index-maintenance cost on the write path buys that away.

Two implementation constraints come with the choice and are worth stating rather than discovering:

  • A generated column’s expression must be immutable, so the text-search configuration is named as a literal ('english') rather than resolved from a setting. Changing an object’s search configuration is therefore a metadata change with a column rewrite behind it, run as an online migration.
  • fastupdate is off on search indexes. The default is ON, which batches new entries into a pending list; leaving it on trades deterministic read latency for insert throughput. Search is a read-latency surface, so the trade goes the other way, and bulk loads that need the throughput drop and rebuild the index instead.

Derived content is eventually consistent, and only derived content. Extracted text from an attached document, and denormalized parent titles, are enqueued through a transactional outbox in the writing transaction and applied afterwards. The budget is p95 under two seconds, p99 under ten, measured per tenant and reported on the Setup search-health surface. Missing derived content never produces a wrong result — only a narrower one — and the record itself is already findable by every one of its own fields. A derived payload that has not landed is visible as such on the health surface rather than inferred from a user complaint.

Band recomposition is the one index-maintenance event with real cost. Changing field permissions can change the grant signature of a field, which changes band membership, which rewrites generated columns. It runs as an expand → migrate → contract migration: the new generation’s columns are built and indexed alongside the old, then the generation pointer flips atomically. Both generations are internally consistent throughout, so no query ever sees a band whose membership disagrees with the grants being enforced against it. On a large object this is a genuine online migration and is scheduled like one.

The score is a weighted blend, and every weight is metadata:

Signal Source Default behavior
Lexical relevance ts_rank_cd over the matched bands Cover-density ranking, so proximity of the matched terms counts, with length normalization on
Field weight setweight labels A–D Postgres’s defaults {0.1, 0.2, 0.4, 1.0} for {D, C, B, A} — a title match is ten times a derived-content match
Exact-identifier boost Term equals a whole token in an A-weighted identifier column Dominates; typing an invoice number returns that invoice first
Recency Exponential decay on last-modified, half-life per object 30 days by default; short for transactional objects, long for reference data
Record type weight typeWeight on the object’s search settings Uniform until an admin says otherwise
User affinity The user’s own recent-items and ownership Boost for records the user owns, recently viewed, or recently edited

Affinity is computed at query time from the user’s own record data, and no behavioral profile is stored. The signal is the user’s recent items — which are records, subject to the record-access plane like anything else — so a user who loses access to a record loses its influence on their ranking in the same instant they lose the record.

Ranking weights are visible and editable in Setup. Relevance tuning that requires a platform release is relevance tuning that never happens, and the loudest complaint about search on any CRM is that the admin can see it is wrong and cannot reach it.

Top results is a blended cross-object list, capped at three per object so one high-volume object cannot monopolize the page. A View all link opens that object’s group. The cap is a presentation rule applied after ranking; it never changes what was searched.

Per-object groups follow, each with its own count and its own pagination. Counts are exact up to 500 and rendered 500+ beyond, because an exact count of a large result set costs a full scan of the filtered candidate set to answer a question nobody is asking precisely. Every count, exact or capped, is over the access-filtered set.

Snippets are generated by ts_headline over only the band columns the user can read, and the result names which field produced the excerpt. Two rules govern them:

  • A snippet is computed only for the page of results actually returned. ts_headline reads the original text rather than the tsvector summary and the Postgres documentation warns it “can be slow”; computing it for a full candidate set would make the cost of a broad search quadratic in the wrong dimension.
  • A snippet can only contain text from a readable band, because a snippet is generated from the same columns the match came from. There is no separate “safe snippet” check to get wrong.

The results page is a list-filtered surface delivered by a package — the same template as any faceted list, with object, owner, and date facets in the filter region — and the search box itself is package-rendered; the query engine, the index, and the access planes below it stay in the kernel. Typeahead is a popover-menu overlay anchored to the search box. Search is not a bespoke design language; it is two ordinary surfaces a package renders over an engine that is pure.

They are different queries with different budgets, and conflating them is how a search box becomes slow.

Typeahead Full search
Trigger 2 characters, 150 ms debounce Explicit submit
Latency budget p95 ≤ 120 ms p95 ≤ 800 ms
Bands searched A-weighted identifier and title columns only All readable bands
Matching Prefix (tsquery :*) and trigram Full lexical, phrase, plus trigram fallback
Snippets None Yes, for the returned page
Counts None Yes
Candidate cap Hard cap, per object Cost-budgeted
Seeded with The user’s recent items before any keystroke —

The two-character minimum is not politeness; a one-character prefix is non-selective against a trigram index and turns a keystroke into a scan. Typeahead deliberately does not search long text: a suggestion list is for going somewhere, and a match in paragraph nine of a scope note is never the thing the user was navigating to.

A reference field’s picker runs the same generated subquery against the target object, with the reference’s declared filter ANDed in as an ordinary predicate. That means the filter narrows the candidate set in the plan, alongside access — a filtered-out record cannot be matched, ranked, or counted, for exactly the same reason an inaccessible one cannot.

Three behaviors follow:

  • The whole searchable set is searched, not just the title field. An invoice identified by its number, a part identified by its stock code, and an account identified by its city are all reachable from the picker. Restricting a picker to the title field is a limitation, not a safety feature.
  • Recent items seed the empty state. Clicking into a picker before typing shows the user’s recent records of that object, which is the correct answer a large fraction of the time.
  • A filtered-out record is explained, not hidden. If a user searches a picker for a record that exists and is readable but fails the reference’s filter, the picker says the filter excluded it and names the filter. This is the one place where “not found” is unhelpful, because the record’s existence is not a secret — the user can see it on its own record page — and silence sends them to an admin.

Recent items are per-user record data: a bounded ring of the last 200 records the user opened or edited, per object. They seed typeahead, they feed the affinity ranking signal, and they are read through the record-access plane, so the list self-heals when access changes.

Saved searches persist a term plus its facets plus its object scope. A saved search is owned, shareable, and — when an admin promotes it — becomes a list view component. Search terms in a saved search are re-executed, never cached: a saved search is a question, not an answer, so it reflects current data and current access every time it runs.

Concern CAOS approach Why
Visibility bands per object Hard cap of 8 Each band is a stored column plus a GIN index maintained on every write. An object needing more than eight distinct field-permission profiles has a permission-model problem that a ninth index will not fix
Fields protected by FLS in the index All of them Protection is band membership, not a list with a ceiling
Searchable fields per object No count cap; bounded by the object’s write-amplification budget The real constraint is index maintenance cost per write, which is what a budget should measure
Result set Bounded page (25 per object group, 50 in top results), keyset pagination The cap is applied after filtering, so it is never consumed by rows the user cannot see
Exact counts Exact to 500, then 500+ Counting past the cap costs a full scan of the filtered set and changes no decision the user is making
Minimum term length 2 characters Below that, no index is selective
Non-selective queries Rejected before execution with a narrowing prompt See below
Query concurrency Per-tenant limiter on the search path Schema-per-tenant isolates data, not CPU
Index size Reported per object on the search-health surface; long-text and trigram indexing opt-in A GIN index over long text is often larger than the text

What makes a query non-selective, and how a pathological one is contained. A search is non-selective when its terms are estimated to match a large fraction of the object’s rows: a single common lexeme, a stop-word-only query after normalization, or a leading-wildcard trigram probe on a large table. Four defenses, in order:

  1. Planning. The query’s cost is estimated before it is run and compared against the object’s search budget. Over budget, it is not executed.
  2. Narrowing, not failing. An over-budget query returns a narrowing surface — add a term, pick an object, add a facet — rather than an error. It is a UI outcome, not an error envelope, because the user did nothing wrong.
  3. statement_timeout on the search path, always, as the backstop for a plan that was wrong.
  4. Per-tenant concurrency limit, so a script hammering the search endpoint degrades its own tenant’s search and nothing else in the schema.

Every rejection is recorded in the execution trace with the term shape (never the term itself) and the planned cost that triggered it, so “search is slow” is answerable from telemetry rather than reproduction.

Salesforce runs a separate search tier. When a record is created or updated, “the search engine comes along, makes a copy of the data, and breaks up the content into smaller pieces called tokens,” which are stored in a search index with links back to the records (Trailhead — Choose the Right Search Solution). The index adds spell correction, nickname recognition, lemmatization, and synonym matching. SOSL queries that index; SOQL queries the database.

The order of operations is the architectural difference. SOSL’s own reference states that “the search engine looks for matches to the search term across a maximum of 2,000 records,” and then — in the same document — that “Admins (users with the View All Data permission) see the full set of results returned,” while other users see a set filtered by their access (SOSL Limits on Search Results). Access filtering is applied to the output of the 2,000-record scan, not to its input. Records the user cannot see consume scan slots before they are removed.

Salesforce documents the user-visible consequence under the name search crowding — Help lists it among the reasons a record cannot be found, describing “a large number of records match a search term, pushing some matching records out of the visible results” (Unable to find records in global search). The same article lists two more structural causes: a custom object’s Allow Search checkbox, and tab visibility — “If the Tab setting is set to Hidden for any object (standard or custom), Global Search will not return records for that object.” Searchability is coupled to navigation.

Field-level security in the index has a documented ceiling. Salesforce Help states that “Search can only protect up to 100 searchable custom fields per object with field-level security, even if you set field-level security for more than 100 custom fields,” that beyond that “the extra fields are unprotected,” and that “Search Manager shows how many objects have unprotected fields.” Custom picklists are always protected and do not count toward the 100. Since Summer ’24 an admin can choose which 100 are protected, with the guidance to “move fields with less sensitive information to the Not Protected list” (Field-Level Security for Custom Fields in Search). Unprotected fields still influence which records match.

Indexing is asynchronous, which is a class of incident rather than a footnote: Salesforce ships a proactive alert for it, which “triggers when indexing for one or more of the entities is delayed, which may lead to a degraded search experience” (Proactive Alert Monitoring: Search Indexing Delay).

Not every field type is searchable. Salesforce’s own list-view search limitations name the excluded types outright: “Lookup fields, Derived fields, Formula fields, Non-text fields (such as Integer and Currency)” (Limitations of list view search). The same article documents that list-view search “only searches records that are included in the list view. Because list views return a maximum of 2,000 records, records outside that limit cannot be found through the search,” that partial search terms are supported only for Record Name fields, and that a term must be at least two characters.

Result limits are shaped by the scan, not by the user’s page. SOSL returns “a maximum of 250 records” when querying one object, and for multiple objects “each object returns up to the minimum number between 2,000/n … and 250” — ten objects means 200 each. In Apex the governor limits are 20 SOSL queries per transaction and 2,000 records retrieved by a single SOSL query (Execution Governors and Limits). Field scope is narrowed with IN SearchGroup — ALL FIELDS (the default), NAME FIELDS, EMAIL FIELDS, PHONE FIELDS, SIDEBAR FIELDS (IN SearchGroup).

Lookup search reads one field. “Search uses only the record name field when looking for matches, unless the name is an auto-number. For auto-number names, search uses the secondary field to match results under certain circumstances.” The dropdown blends “recent items, items from your most frequently used objects, and items from the current object” (Lookup Searches).

Einstein Search adds personalization and conceptual queries on top. Salesforce’s engineering write-up describes a two-stage ranker — “a first algorithm leverages the user profile to rank objects based on their likelihood to match the user’s query intent. Next, algorithms at the object level rank matching records, boosting records with field values matching the user profile” — trained on clicked results at global scale, and notes that “user search profiles are currently constructed on the fly at query time and never stored to preserve data privacy” (How We Built Personalization and Natural Language into CRM Search).

Where CAOS is genuinely better:

  • Access filtering is a predicate in the plan, not a post-pass. The result cap is applied to the filtered set, so it cannot be consumed by rows the user may not see, and no user-visible count is derived from an unfiltered scan.
  • Field-level security in the index is total, by construction. Bands are derived from the grant signature of every searchable field, so there is no protection ceiling, no protected/not-protected triage, and no set of fields that silently influence matching.
  • Search is transactional with the write. Generated tsvector columns are maintained by the statement that writes the row, so the indexing-delay class of incident does not exist for a record’s own fields. Only derived content is asynchronous, and its staleness is reported rather than inferred.
  • Every field type is searchable, including computed ones. A formula field is a stored column; indexing it needs no special case. Numbers and dates index as exact tokens.
  • Searchability is independent of navigation. An object is searchable because it is marked searchable and the user holds Read — not because a tab is visible.
  • Lookup search reads the object’s whole searchable set with the reference’s filter applied in the plan, and explains a filter exclusion instead of returning nothing.
  • Ranking weights are metadata. Field weight, recency half-life, and record-type weight are editable in Setup, so relevance is an admin lever rather than a platform release.

Parity: tokenization, stemming, stop words, synonyms and spell tolerance, a top-results blend over per-object groups, typeahead seeded with recent items, and behavioral personalization computed at query time and never stored. That last one is a good decision and CAOS copies it deliberately. The two-character minimum term is also copied — the underlying reason is identical.

Costs and risks:

  • Write amplification is the price of transactional indexing. Every write recomputes up to eight tsvector columns and updates up to eight GIN indexes, with fastupdate off. Bulk load paths must drop and rebuild rather than insert through the indexes, and that path has to exist before the first large migration, not after it.
  • Band recomposition is an online migration. An FLS change that regroups fields rewrites generated columns on the whole object. It is correct and atomic, but on a large table it is scheduled work, and an admin toggling field permissions casually will be surprised by that.
  • Band count is a real ceiling. Eight is a design choice that will occasionally be wrong for a genuinely complex object, and the answer is to simplify the permission model — which is the right answer and an unwelcome one.
  • Relevance quality is bounded by Postgres. ts_rank_cd plus a weighted blend is a good classical ranker, not a learned one. There is no cross-tenant training corpus, and per-tenant term statistics are noisy on small datasets. The escape hatch — an external search engine — is closed on purpose, because it would reintroduce exactly the post-filter security architecture this design exists to avoid.
  • ts_headline cost is real. Restricting it to the returned page keeps it bounded, but a page of long-text results is measurably slower than a page of short ones, and the budget must be enforced rather than assumed.
  • Index storage multiplies by tenant. Schema-per-tenant means every tenant carries its own GIN indexes and its own autovacuum load. That is excellent isolation and a genuine capacity-planning burden on the control plane, where a single shared search tier would have amortized it.
  • Non-selective queries are a denial-of-service surface. Planning, narrowing, timeouts, and a concurrency limiter are mandatory hardening on the search path, not optional polish.
Component type Body
Object search settings search_settings object, searchable, title, subtitle[], fields[] (weight, fuzzy, searchable), derived[], typeWeight, recencyHalfLife
List view / promoted saved search list_view object, filter, columns[], searchable
Search ranking defaults search_ranking Blend weights for lexical, recency, affinity, type; count and page caps

Visibility bands are never a component. They are derived at deploy from the field-permission rows in the permission catalog, so an object’s band layout is a consequence of its security model rather than a parallel declaration that could disagree with it. A deploy diff reports band changes as a derived effect with the migration they imply, the same way it reports an index rebuild.

Saved searches and recent items are record data and never appear in a deploy — the same split the permission model draws between a permission set and its assignments.

A deploy that changes only ranking weights is metadata-only and takes effect at the generation flip with nothing to rebuild. A deploy that changes which fields are searchable, or that changes field permissions in a way that regroups bands, carries a column rewrite and runs expand → migrate → contract.

Salesforce Metadata API analogs, for migration mapping: SearchSettings (org-level search configuration), the searchLayouts element on CustomObject (which fields appear in search results), enableSearch on CustomObject (the “Allow Search” checkbox), and ListView.