Skip to content

Files & content

Every application accumulates binary content: a signed contract, a stamped drawing, a photograph from the field, a spreadsheet someone emailed in. None of it fits a metadata model that thinks in columns, and the usual response is to bolt on a separate subsystem with its own storage, its own sharing rules, its own limits, and its own idea of what deleting means.

The load-bearing decision on this page: a file is a record. file is an object in the same sense account is an object — a real Postgres table with typed columns, an owner, a record-access predicate, field-level security, custom fields an admin can add, page layouts, list views, validation rules, automation, history, and audit. The only thing that distinguishes it is that one of its properties is a pointer to bytes held outside the database. Everything else about it is the platform working normally.

That decision decides the rest of the page. Because a file is a record, it does not need a bespoke sharing model — it uses the record-access plane. It does not need a bespoke permission surface — it uses permission sets and FLS. It does not need a private delete path, a private audit trail, or a private search index. A file that were not a record would need all six, and every one of them would be a place where access could resolve differently than it does everywhere else.

Three tables, and the split between them is deliberate.

Table Holds Mutability
file The durable identity of a document — name, owner, classification, custom fields, a pointer to the current version Mutable; this is the record users see, link, and edit
file_version One immutable snapshot — content hash, byte count, media type, author, reason for change Append-only; a version is never edited or overwritten
file_link The junction between a file and a record it is attached to Mutable; rows come and go as attachments change

file is a base-org object with a fixed set of platform fields and room for as many custom fields as the org wants:

Field Type Notes
name text The display title, independent of the uploaded filename
extension text Derived from the current version
media_type text The verified IANA media type, sniffed from content, not trusted from the client
size_bytes number The current version’s size
content_hash text SHA-256 of the current version’s bytes
owner_id lookup(user) Ownership in the ordinary sense — the record plane reads this column
current_version_id lookup(file_version) Which version is current
version_count rollup Maintained the same way any roll-up is
state picklist draft, scanning, available, quarantined, failed
access_mode picklist inherited or explicit — see access resolution
retention_class picklist The input to retention policy, an ordinary picklist an org defines

A custom field on file is a real column on the file table, exactly as on any other object — contract_effective_date, drawing_revision, confidentiality. Those fields get FLS, appear on the file’s page layout, are filterable in list views — the layout and the list views being package-delivered surfaces over the kernel’s stored record, not compiled-in UI — and are readable by formula fields and roll-up summaries on the records the file is linked to. An invoice can carry a roll-up of “how many drawings are attached” because a link is a real foreign key and a file is a real row.

A file_version is immutable once written. It carries version_number (monotonic per file), content_hash, size_bytes, media_type, created_by, created_at, and reason — a free-text note the uploader supplies. Uploading new content for an existing file inserts a version and repoints file.current_version_id; it never mutates an existing row.

Because the blob is addressed by its hash rather than by its version id, restoring a prior version copies nothing. Restore inserts version N+1 whose content_hash equals version K’s, and the storage layer already holds those bytes. History is append-only in both directions: a restore is a new event, not an erasure of the events after K.

file_link is a junction with exactly four meaningful columns — file_id, target_object, target_id, created_by — and no access semantics whatsoever. No share type, no visibility flag, no per-link permission. A link says one thing: this file is attached to that record.

That absence is the decision. The moment a link carries access semantics, the same file has as many different answers to “who can read this” as it has links, and the answer depends on which link the reader arrived through. Access is a property of the file, resolved once, and a link is a relationship.

Three shapes fall out of this without any of them being a special case:

  • Attached to one record — one link row. The common case.
  • Shared to several records — several link rows, one file, one set of bytes, one version history. Editing the file changes what every linked record shows, because there is one file. Nothing is copied and nothing diverges.
  • Attached to nothing — zero link rows, which is a perfectly ordinary state, not an orphan. A library asset, a template, a photo someone uploaded before deciding where it belongs. It resolves on its own owner and grants, it appears in the Files list view, and no cleanup job comes for it.

Bytes live in object storage. The database holds metadata and a pointer, and never the blob.

The reason is not squeamishness about large columns; it is that the two have incompatible operational profiles. Blobs are large, immutable, and read by streaming; database rows are small, mutable, transactional, indexed, backed up, replicated, and restored as a unit. Putting terabytes of immutable bytes inside the transactional store makes every backup, every restore, every replica, and every point-in-time recovery pay for content that never changes. Object storage is priced, replicated, and lifecycle-tiered for exactly that shape.

The storage key is the content hash: tenant/<tenant-id>/blob/<sha256>. This gives three properties for free.

  • Deduplication within the tenant. The same PDF attached to forty invoices stores once. Ten users uploading the same specification store once. Every version row still exists — history is unaffected — but the bytes behind identical versions are one object, refcounted by the number of file_version rows that name the hash.
  • Integrity by construction. The name of the object is a claim about its content, and the claim is verified on write and re-verifiable at any time. A corrupted or substituted object is detectable without a separate checksum column to trust.
  • Free restore and free copy. Restoring a version, cloning a record with its attachments, or deploying a sandbox seed all reference existing hashes rather than duplicating bytes.

Deduplication is per tenant and never across tenants. Cross-tenant dedup would make storage a side channel: an attacker who can upload a file and observe whether it deduplicated learns whether another tenant already holds that exact byte sequence. That is a real disclosure for contracts, offer letters, and standard documents. The schema-per-tenant boundary is a boundary for blobs too, and each tenant’s objects are encrypted at rest under a per-tenant key.

File records are data. What an admin declares as metadata is the policy around them.

Whether an object accepts attachments, and how access flows across the link, is a block on the object component:

{
"kind": "object",
"key": "invoice",
"files": {
"enabled": true,
"access": "inherit",
"maxFileBytes": 5368709120,
"allow": ["application/pdf", "image/*", "application/vnd.openxmlformats-officedocument.*"]
}
}
  • access: "inherit" — a user who can read the invoice can read files linked to it; a user who can edit the invoice can edit them. This is the default because it is what people mean by “attach a file to a record.”
  • access: "independent" — a link grants nothing. The file resolves on its own owner and grants alone. Set this on objects whose readership is deliberately wider than their attachments: a public knowledge article, a portal-visible order.
  • allow — an allow-list of media types, matched against the sniffed type, not the declared one. Absent, any type is accepted subject to the platform’s refused types.

Retention and legal hold are canonical components, each carrying a pure predicate over the file record in the one typed language:

{
"key": "retention_invoice_attachments",
"label": "Invoice attachments — 7 years",
"type": "file_retention_policy",
"body": {
"covers": "record.linkedTo(\"invoice\") && record.retention_class != \"permanent\"",
"retainFor": "P7Y",
"from": "lastLinkRemoved",
"then": "purge"
}
}
{
"key": "hold_matter_2941",
"label": "Legal hold — matter 2941",
"type": "legal_hold",
"body": {
"covers": "record.linkedTo(\"account\", \"a3f9c1d2-…\") || record.created_at >= date(\"2026-01-01\")",
"reason": "Litigation hold, matter 2941",
"releaseRequires": "files.legal_hold"
}
}

Rendition specs — which previews the platform generates, at what sizes, from which media types — are components too, so a tenant that needs a 1600-pixel drawing preview declares one rather than filing a request.

Because covers is an ordinary pure expression, retention and hold participate in the dependency graph like everything else: renaming retention_class knows which policies read it, and a deploy that would orphan a hold predicate fails validation rather than silently ceasing to hold anything.

A file clears the same three planes as any record, in the same query, with no system mode and no separate file-permission subsystem:

  1. Object and field permissions on file — whether the user may read or edit files at all, and which of the file’s own fields they may see. A confidentiality field readable only by Legal is an ordinary FLS grant.
  2. Record access — which files, resolved as a predicate.
  3. System permissions — files.share_external to mint an external link, files.legal_hold to place or release a hold, files.purge to hard-delete content ahead of policy. Three permissions, added because three surfaces exist that need gating.

The record predicate is the standard one plus a single additional clause for link inheritance:

USING (
is_root(current_user_id())
OR owner_id = current_user_id()
OR owner_id = ANY (subordinate_user_ids(current_user_id()))
OR <org-wide-default clause for `file`>
OR EXISTS (SELECT 1 FROM file_share s
WHERE s.file_id = file.id
AND s.grantee = ANY (current_user_groups()))
OR (access_mode = 'inherited'
AND EXISTS (SELECT 1 FROM file_link l
WHERE l.file_id = file.id
AND target_readable(l.target_object, l.target_id)))
)

target_readable is a SECURITY DEFINER helper that dispatches to the linked object’s own policy — the same cycle-break the sharing model uses, for the same reason: a policy that consults another policy must not re-enter row security.

Four properties follow, and they are the whole access story:

  • Access is a union, and it is additive. Adding a link can only widen access; removing one can only narrow it. There is no deny link, because there is no deny anywhere on the platform.
  • access_mode = 'explicit' suppresses the inheritance clause for one file. A confidential contract attached to an account half the company can read resolves on its own owner and explicit shares only. This is a property of the file, set once, visible on the file’s record page, and identical no matter which record a reader arrived from.
  • Edit follows the same shape. The write predicate substitutes target_editable, so “can edit the invoice” implies “can replace the drawing on it” under inherit, and implies nothing under independent.
  • A file the user cannot see is absent, not refused — consistent with the error model. A related list renders the files the reader may see and gives no count of the ones they may not.

caos access explain --object file --record 8f3a… --user jane names the clause that admitted her — ownership, hierarchy, an explicit share, or a specific link to a specific record — and names that none did when the answer is no.

Download, and why a signed URL is not a permission

Section titled “Download, and why a signed URL is not a permission”

The dangerous shape in every file system is the download URL, because it is a bearer token that outlives the check that produced it. If a URL is minted at the moment access is granted and remains valid afterward, then revoking access does nothing to the URL, and the URL is now a permission the access model cannot see.

Downloads therefore resolve in two hops, and the check happens at the second:

  1. The client requests /files/{id}/download. The kernel re-resolves all three planes for the current user, right now. On success it issues a storage pre-signed URL scoped to one object key, GET only, no listing, no range beyond the requested one, valid 60 seconds, and redirects to it.
  2. The browser follows the redirect and streams the bytes directly from object storage. The application tier never proxies the payload.

The kernel-side URL is stable, shareable, and worthless on its own: pasting it into another user’s browser re-runs step 1 against that user and fails. The storage-side URL is unstable, unguessable, and expires before it can be circulated. Revoking a permission takes effect on the next request rather than at the end of some cached grant, because there is no cached grant — the check is at issue, every time.

The same path serves previews and thumbnails. A rendition is not separately shareable and carries no access of its own; it resolves against its source file. A thumbnail of a document nobody may read is a document nobody may read.

Sending a file to someone who has no login is a real requirement, and pretending otherwise pushes people to email attachments. It is served by an explicit file_external_link record — a first-class object with its own record page, owner, and audit trail — and creating one requires files.share_external.

Every external link has, without exception:

  • An expiry. Required, defaulting to 14 days, capped by a tenant policy. There is no non-expiring external link.
  • A revocation switch that takes effect immediately, because redemption re-resolves the link record rather than trusting the token.
  • A scope: one file, one version — pinned, so a later version is not silently disclosed — and one of view or view + download.
  • An access log: every redemption is an audit entry with timestamp, IP, and user agent, correlated to the link record.

A password is optional and on by default; turning it off is a deliberate, audited edit rather than the absence of a feature. An external link is not a hole in the access model — it is a grant to an anonymous principal, recorded as such, with an expiry and an owner.

Upload, and the states a file passes through

Section titled “Upload, and the states a file passes through”

Bytes go directly to object storage. The application tier issues a scoped upload ticket and never touches the payload, which is what makes a multi-gigabyte upload the same cost to the platform as a small one.

draft ──▶ uploading ──▶ scanning ──┬──▶ available
└──▶ quarantined
└──▶ failed (abandoned, hash mismatch, type refused)
  • draft — the file and file_version rows exist; no bytes yet. The record is visible only to its creator.
  • uploading — an upload session is open against a per-tenant staging prefix. Uploads are chunked and resumable: the session records which parts have landed, so a dropped connection resumes at the boundary rather than restarting. Sessions expire after 24 hours and abandoned parts are collected as background work.
  • scanning — on completion the kernel verifies the byte count and the SHA-256 against what the client declared, sniffs the media type from the content’s magic bytes, and refuses the upload if the sniffed type contradicts the declared one or falls outside the object’s allow list. A type disagreement is a rejection, not a warning, because “declared as an image, actually an executable” is the entire shape of the attack. Then a malware scan runs as background work.
  • available — the scan cleared. Only now is the file linkable, previewable, downloadable, searchable, and visible in a related list.
  • quarantined — the scan found something. The bytes are retained for investigation, the record is visible to security admins and to the uploader, and every read path refuses. Quarantine is a state, not a deletion, because the security team needs the sample.

Nothing becomes visible before it is scanned, on any upload path. There is no API that trades the scan for latency, and there is no size above which scanning is skipped — a large file simply stays in scanning longer. The cost is honest and named below.

A rendition is derived content: a thumbnail, a page image, a PDF of an office document, or a plain-text extraction used for search and for AI extraction. Generation is background work, queued when a version reaches available, and the queue is the ordinary one with the ordinary retry, budget, and trace behavior.

Renditions are keyed by (content_hash, spec), not by file id. The same PDF attached to forty invoices renders once. A restored version needs no rendering at all, because its hash already has renditions.

Source Renditions
Images (jpeg, png, webp, gif, tiff, heic) Thumbnails, a normalized web-safe display image
PDF Per-page images, thumbnails, extracted text
Office documents (docx, xlsx, pptx, and their legacy forms) A PDF rendition, then per-page images and text from it
Plain text, Markdown, CSV, source code Syntax-aware inline rendering; no image generated
Audio and video A poster frame and duration; playback streams the original
CAD, archives, and unrecognized binaries None — the UI says the type has no preview, immediately, rather than showing a spinner that never resolves

Page images are generated lazily by range and cached, so a 3,000-page document previews its first page as fast as a one-page document, and page 2,900 renders when someone asks for it. There is no page ceiling past which preview stops working.

Renditions are a cache, not history. They are safe to evict, regenerate on demand, and — the decision that matters commercially — they do not count against the tenant’s storage meter. A tenant is metered on the content it uploaded, not on the platform’s choice of how many thumbnail sizes to keep.

Pulling typed values out of an uploaded document — a PO number from a scanned purchase order, line items from a supplier invoice — runs on the AI platform and is not redefined here. Three properties are guaranteed by this layer:

  • Extraction reads the text rendition, which means it reads only what a user with access to the file could read, and it runs under a principal whose access is resolved by the same three planes.
  • Its output is a proposal, not a write. Accepted values travel through the ordinary save order — validation, automation, and audit all see them — so an extracted value is indistinguishable downstream from a typed one, and equally reversible.
  • The extraction, its confidence, and the version it read are recorded, so “where did this number come from” resolves to a specific version of a specific file.

Deleting a record does not delete its files. It deletes its links. This follows from the model rather than being a policy choice: a file linked to three records is one file, and deleting one of the three cannot be allowed to destroy content the other two depend on. A file whose last link is removed becomes an unlinked file, owned by its owner, visible in the Files list — not an orphan and not a candidate for immediate collection.

Deletion of the file is the platform’s ordinary three-stage path, matching safe-delete:

  1. Soft delete — the record is removed from views and remains restorable for 15 days.
  2. Purge — the row and its versions are removed. Gated on files.purge.
  3. Blob collection — the underlying object is deleted only when its refcount reaches zero, and only after a 30-day grace window past that point. Because storage is content-addressed, a hash reachable from any surviving version anywhere in the tenant is never collected, which makes “we purged a record and broke an unrelated file” structurally impossible.

Legal hold beats everything. A legal_hold component’s covers predicate is evaluated on every retention and purge path. A held file cannot be soft-deleted into purge, its versions cannot be removed, blob collection skips its hashes, and retention jobs pass over it. Placing and releasing a hold are audited events, the hold itself is a deployable component with a diff history, and releasing one requires files.legal_hold — so “who lifted the hold, when, and why” is answerable from the metadata audit stream without asking anybody.

Retention policies express the ordinary case: keep invoice attachments seven years from the removal of their last link, then purge. Bulk lifecycle operations — an archive sweep, a mass reclassification — run through data management rather than being a separate file-only tool.

Concern CAOS Salesforce Why theirs is shaped that way
Per-file size One limit, every path. Tenant policy, default 10 GB, ceiling set by the object store’s single-object limit Eleven different limits by path: 10 GB UI/libraries/related lists, 2 GB Chatter posts, 2 GB REST, 150 MB Bulk API 2.0, 38 MB SOAP, 10 MB Bulk API, 10 MB Visualforce, 25 MB Classic attachments and email, 5 MB Documents tab / Knowledge / chat transfer Each path was built at a different time against a different transport; the limit is a property of the pipe, not of the file
Non-multipart API upload Not applicable — bytes never travel as encoded field data 50 MB text or 37.5 MB base64 Base64 blobs inside a SOAP/REST envelope must be buffered and decoded in the app tier
Links per file Unbounded; a link is a junction row 2,000 shares per file, records and people and groups combined A share is an access-bearing row, so the count is bounded by the sharing engine
Versions per file Unbounded; versions are rows, identical content stores once Unbounded rows, but each version’s bytes count against file storage No content addressing, so a re-upload of identical bytes is charged again
Preview page depth Unbounded; pages render lazily by range and cache First 500 pages (reported) Eager whole-document rendering has to stop somewhere
Malware scan coverage Every file, every path, before visibility. No size ceiling Files ≤ 100 MB only; UI uploads blocked on detection, API uploads allowed and scanned asynchronously; pre-existing files scanned on first download, and that first download proceeds Synchronous scanning on the API path would change the API’s latency contract
Storage metering Unique stored bytes per tenant. Renditions, thumbnails, and text extractions are free Per-org base 10 GB file storage (Developer and Personal: 20 MB) plus per licensed user 2 GB (Enterprise, Performance, Unlimited) or 612 MB (Contact Manager, Group, Professional) File storage is a sold unit, so it is allocated by license rather than by usage shape
Over-quota behavior Soft threshold warns; hard threshold blocks new uploads only. Reads, downloads, previews, and deletes always work Uploads fail at the limit —

Storage usage is a first-class surface, not a Setup page nobody visits: a storage_usage object with per-object, per-owner, and per-retention-class breakdowns, current and projected, queryable in a list view and reportable like any other data.

Salesforce carries two file systems and has for over a decade.

The legacy one is Attachment — a child row hanging off a single ParentId, one parent, no versions, 25 MB in Classic. It still functions and Salesforce’s own knowledge base states that “Attachments are deprecated in Lightning Experience and are not indexed for search within the Lightning interface” (Attachments from Classic are not searchable in Lightning). The recommended remedy — a Setup toggle that uploads future attachments as Files — applies only going forward; existing attachments stay unsearchable until migrated. Migration is a project, not a switch: attachment rows must be rewritten as three objects with their parentage, ownership, and timestamps preserved, and open-source conversion tooling exists precisely because the platform does not ship it (sfdc-convert-attachments-to-chatter-files).

The modern one is Salesforce Files, a triple:

  • ContentDocument — the file’s identity, pointing at LatestPublishedVersionId.
  • ContentVersion — an immutable version row holding VersionData, VersionNumber, IsLatest, ContentSize, Checksum, and ReasonForChange. A new version is an insert against the same ContentDocumentId (ContentVersion).
  • ContentDocumentLink — the join to a record, a user, a group, or a library (ContentDocumentLink).

The version model is right, and CAOS copies it. The link model is where it goes wrong. ContentDocumentLink carries ShareType — V for Viewer, C for Collaborator, I for Inferred, meaning “whatever the linked record says” — and Visibility — AllUsers, InternalUsers, or SharedUsers. Access semantics therefore live on the edge, so one file reached through two links can grant two different levels, and the answer to “who can edit this file” requires enumerating every link. The I value’s behavior is not obvious from its name, and developers routinely add an after-insert trigger on ContentDocumentLink to force the share type the platform did not set for them. Inserting a link is also an access-bearing write, which is why guest and site users hit INSUFFICIENT_ACCESS_ON_CROSS_REFERENCE_ENTITY attaching a file to a record they can otherwise write to (community report).

Querying compounds it: SOQL against ContentDocumentLink must filter on ContentDocumentId or LinkedEntityId, so “list every record this file is attached to” and “list every file on these records” are shaped by the query restriction rather than by the question.

Limits are per-path rather than per-file. The same document is 10 GB through Files Home, 2 GB in a Chatter post, 38 MB over SOAP, 10 MB through Bulk API, and 5 MB in the Documents tab (File Size and Sharing Limits); a non-multipart REST insert is capped at 50 MB of text or 37.5 MB of base64 (Insert or Update Blob Data). A file may be shared at most 2,000 times across records, people, and groups combined.

Malware scanning now exists and is on by default, with a shape worth reading closely: a UI upload is scanned and blocked, but “when a user uploads a file via the API, Salesforce allows the upload and scans the file asynchronously”; files are “scanned only when they’re 100 MB or smaller”; and a file that predates the feature is scanned on first download, with that download allowed to proceed (Malware Scanning for Salesforce Files).

Renditions are PDF, THUMB120BY90, THUMB240BY180, and THUMB720BY480. For shared files “renditions process asynchronously after upload”; for private files they “process when the first file preview is requested, and aren’t available immediately after the file is uploaded” (File Rendition).

External sharing splits into two mechanisms with different security properties. Content deliveries support expiration and password protection; public links support neither, and “external users can access it without needing a Salesforce login” (Content Deliveries vs Public Links). Both are ContentDistribution rows carrying ExpiryDate, Password, PreferencesPasswordRequired, and ViewCount (ContentDistribution).

Files Connect surfaces external repositories — SharePoint Online, OneDrive, Google Drive — inside Salesforce, with documented constraints: the Google Drive Recent list is limited to the 24 most recently accessed documents from the last 30 days, Name queries support a single trailing % wildcard matching prefixes only, and SharePoint folders containing # or % appear in listings but their contents are inaccessible (Files Connect Implementation Guide).

Storage is a metered, sold resource: 10 GB per org for most editions, plus 2 GB per licensed user on Enterprise, Performance, and Unlimited, or 612 MB on Contact Manager, Group, and Professional (Salesforce Files Storage Allocations). Running out is a well-documented operational event, and the standard remedies — mass deletion, retention sweeps, and third-party offloading — are all responses to storage being priced per gigabyte rather than to it being technically scarce.

Where CAOS is genuinely better:

  • One file system, not two. There is no legacy attachment object, so there is no migration, no split search index, and no “which related list is this document actually in.”
  • Access lives on the file, never on the link. One file, one resolution, one answer — instead of ShareType and Visibility per edge producing as many answers as there are edges.
  • One size limit across every path. The maximum is a property of the file and the tenant, not of the transport that carried the bytes, because bytes never travel through the app tier as encoded field data.
  • Content-addressed storage. Deduplication, integrity verification, zero-copy version restore, and safe refcounted collection all fall out of naming the object by its hash.
  • Scan before visible, everywhere. No API path trades the scan for latency and no size ceiling switches scanning off, so “uploaded” and “reachable” are never the same instant.
  • Signed URLs are not grants. Every download re-resolves access at redemption; the storage-side URL lives 60 seconds and is scoped to one key. Revocation is immediate because nothing was cached.
  • External links always expire and are always revocable. There is no equivalent of a public link with no password and no expiry.
  • Files are ordinary records. Custom fields, layouts, list views, validation rules, automation, roll-ups, history, and audit apply without a file-specific version of any of them.
  • Renditions are free. Derived content is a cache and is not metered, so the platform’s rendering choices are not billed to the customer.

Parity: an immutable version chain with a pointer to the current version, a junction that lets one file attach to many records, previews generated asynchronously off the request path, expiring password-protected external delivery, and metered storage with a visible usage breakdown. Salesforce’s version model in particular is correct, and CAOS reproduces it rather than inventing an alternative.

Costs and risks:

  • Scan-before-visible adds latency to every upload, and it is most visible exactly where it is least welcome — large files on slow scanners. The mitigation is a real progress state in the UI and a queue with capacity headroom, not an escape hatch, because an escape hatch is the whole hole.
  • Two storage systems must stay consistent. A row can exist without its blob (a crashed upload) or a blob without its row (a rolled-back transaction). Reconciliation is a standing background job with alerting, and the grace window before blob collection exists so a reconciliation bug loses nothing.
  • The inheritance clause is a subquery over another object’s policy. It is the most expensive clause in the file predicate, and it must be indexed on file_link (file_id) and (target_object, target_id) and benchmarked at realistic link counts. A file linked to thousands of records is a legitimate shape and a slow one.
  • Union inheritance can over-grant quietly. Linking a confidential file to a widely readable record widens access with no prompt. access_mode = 'explicit' is the answer, but it is opt-in, and a file’s access surfacing on its record page is a requirement rather than a nicety.
  • Content addressing means bytes outlive rows. A hash referenced anywhere survives deletion everywhere else, which is correct for integrity and awkward for “delete every copy of this document.” Purge-by-content is therefore an explicit, audited, hold-aware operation and not a side effect of deleting a record.
  • Unmetered renditions are a real cost the platform absorbs. Lazy generation, eviction of cold renditions, and a per-tenant rendering budget are required controls, not optimizations.
  • Object storage is a second failure domain. Its availability, region, and lifecycle policy now bound the application’s, and backup/restore must be coordinated across two systems so a database restore does not land on collected blobs.
Component type Body
Object file policy (a files block on the object component) enabled, access, maxFileBytes, allow[]
Retention policy file_retention_policy covers, retainFor, from, then
Legal hold legal_hold covers, reason, releaseRequires
Rendition spec rendition_spec from[], kind, dimensions, format
Custom fields on file field Ordinary field components, no special casing

file, file_version, file_link, file_share, and file_external_link rows are data, never components — the same split the permission model draws between a permission set and its assignments, and the sharing model draws between a sharing rule and a share row.

A deploy that changes an object’s access from inherit to independent is metadata-only: it recompiles the file predicate and takes effect at the next query under the ordinary generation flip. No rows move. A deploy that lowers maxFileBytes does not retroactively invalidate existing files; it applies to the next upload, and the validator says so.

Salesforce Metadata API analogs, for migration mapping: there is no metadata type for the file objects themselves — ContentDocument, ContentVersion, ContentDocumentLink, ContentDistribution, and Attachment are all data. ContentAsset (asset files) and ContentDeliveryPolicy-adjacent org settings are the closest metadata surfaces, and the storage allocation is a licensing artifact rather than a deployable one.