Masking a Salesforce sandbox usually means scrambling names, emails, and phone numbers in structured fields. That work is necessary but incomplete. Every org also carries an ocean of unstructured content sitting in Files, ContentVersion, and the legacy Attachment object, and none of it gets touched by a standard field-level masking pass. A driver's license scan attached to a Case record stays exactly as it was uploaded, full resolution, full name, full address, sitting in a sandbox that twenty contractors and three offshore QA vendors can open.
Why field masking never touches binary files
Masking tools work by reading a field's value, running it through a transformation, and writing the fake value back. That model assumes the PII lives inside a queryable text or number field. Salesforce Files and Attachments are stored as binary blobs referenced by metadata records, ContentVersion for Files, Attachment for the older object type. The masking engine can see the file name and the owner, but it has no reason to open the blob itself.
This isn't a flaw in the masking logic. It's a scope decision, and most vendors never advertise it clearly. If your masking tool's documentation doesn't mention files anywhere, assume it does nothing to them. I have yet to see a default Salesforce sandbox masking configuration that opens a PDF and redacts a name printed inside it, because doing that reliably requires OCR and document parsing, not string substitution.
The result is a two-tier sandbox. Structured data looks clean and defensible. Unstructured data, often the most sensitive kind, ships through untouched. An auditor who checks only the Contact object will sign off. An auditor who opens a few attached files will find the gap in about ninety seconds.
Where PII actually hides in Salesforce files
Attachments accumulate PII in places admins rarely audit. Sales reps attach signed contracts to Opportunities. Support agents attach screenshots of customer emails to Cases. HR-adjacent orgs attach ID documents, medical notes, or background check results to custom objects nobody reviews during a data protection audit.
- Signed contracts and order forms with full legal names, addresses, and signatures
- Scanned government IDs and passports attached during KYC or verification workflows
- Exported spreadsheets and CSV reports attached to Chatter posts or Cases for reference
- Email threads saved as .eml or .msg files, often containing entire conversation histories
- Screenshots of customer support tickets from other systems, pasted in as images
None of these are edge cases. They're routine business behavior in any org that has been live for more than a year. And because Files can be linked to almost any object through ContentDocumentLink, they follow the sandbox refresh right along with the records they're attached to, no extra step required to bring them along.
Why this matters more under GDPR than field masking does
GDPR defines personal data broadly. Article 4 doesn't care whether the data sits in a queryable field or a scanned image. A PDF of a passport is just as much personal data as a Contact record with the same person's name in it, arguably more sensitive given the document types involved. Regulators assessing a breach or an audit finding won't accept "we masked the fields" as a defense if the sandbox still contains readable scanned IDs.
The risk profile is also different in kind. A masked field that somehow leaks reveals one data point. A leaked attachment can reveal a signature, a date of birth, a national ID number, and a photo, all in one file. That concentration of sensitive data in a single object makes attachments a higher-value target for anyone probing a sandbox for exploitable information, and a bigger liability if that sandbox gets exposed through a misconfigured guest user profile or an over-shared community.
What actually needs to happen with files before a refresh
Masking the file's content is rarely practical at scale. Trying to OCR every PDF and attachment across a multi-gig sandbox, identify PII within it, and redact it in place is slow, error-prone, and still leaves metadata like file names intact. A more workable approach treats files as a category to exclude or replace rather than transform.
| Approach | What it does | Trade-off |
|---|---|---|
| Strip files during refresh | Excludes ContentVersion and Attachment records from the sandbox copy entirely | Breaks workflows that depend on test files being present |
| Replace with dummy files | Swaps real attachments for placeholder documents of the same file type | Preserves record counts and file-type testing without exposing content |
| Restrict by object type | Strips files only from high-risk objects like Case, Contact, or custom ID-verification objects | Requires ongoing maintenance as new objects gain file attachments |
| Leave untouched | Default behavior for most masking tools | Full PII exposure in every refreshed sandbox |
My preference, and the one MaskEzee applies by default, is dummy-file replacement scoped by object sensitivity. Developers testing a file-upload trigger still get a file of the right type and size class to work with. Nobody gets a readable passport scan in a dev sandbox that half the org can log into.
Auditing your org's exposure before you fix it
Before choosing a strategy, run a query against ContentDocumentLink and Attachment to see what you're actually carrying. Group by LinkedEntityId's object type and you'll usually find a handful of objects account for most of the volume, Cases and Opportunities being the usual suspects. Check file extensions too. A sandbox full of .pdf, .jpg, and .msg files attached to Contact or Case records is a strong signal that ID documents and email exports are sitting there unmasked.
It's also worth checking who can see these files once the sandbox exists. Files inherit sharing from their parent record in most configurations, but Chatter-attached files and Files uploaded through Experience Cloud can have broader visibility than admins expect. A file quietly shared org-wide because of a sharing rule inherited from its parent record turns a masking gap into an access control problem too.
Document what you find, even informally. When someone asks why the sandbox masking policy excludes certain file types, having the audit numbers ready beats explaining the reasoning from scratch every time.
Building files into your masking policy properly
Treat files as a named category in whatever masking policy document your org maintains, not an afterthought bolted onto the field-masking rules. Specify which objects get files stripped, which get dummy replacement, and which are low-risk enough to leave alone, if any genuinely are. Revisit the list every time a new object type starts accepting attachments, because that happens more often than most admins track.
Tie the policy to your refresh cadence, not to a one-time cleanup project. A file-masking rule that only runs when someone remembers to trigger it manually will lapse within two release cycles. Build it into the same automated pipeline that handles field masking, so every sandbox refresh applies the same file-handling rules without anyone needing to remember a manual step.
Full Copy sandboxes carry the largest file volume since they mirror production storage entirely, but Partial Copy sandboxes inherit files tied to whatever records the copy includes, so the exposure isn't limited to your biggest environment. Scope the policy to cover every sandbox type that gets refreshed with real data, not just the one everyone assumes is the risky one.
Frequently Asked Questions
Does Salesforce's standard sandbox masking feature handle Files and Attachments?
No. Salesforce's built-in Data Mask feature operates on field values within standard and custom objects, not on binary content stored in ContentVersion or the Attachment object. Files and attachments carry over into refreshed sandboxes exactly as they existed in production unless a separate process strips or replaces them.
What's the difference between the Attachment object and Salesforce Files for masking purposes?
Attachment is the legacy object tied directly to a single parent record, while Files use ContentVersion and ContentDocumentLink to allow sharing across multiple records. Both store binary content that field-level masking tools cannot read or transform, so both need to be addressed separately in any masking policy.
Can OCR-based tools redact PII inside PDFs and scanned documents automatically?
Some specialized document-redaction tools can OCR a file and black out detected PII, but accuracy varies with scan quality and layout, and file names and metadata still need separate handling. For most sandbox refresh workflows, replacing sensitive files with placeholder documents is faster and more reliable than attempting in-place redaction at scale.
Which Salesforce objects typically carry the most sensitive attached files?
Case and Contact records tend to accumulate the highest volume of sensitive files, including scanned IDs, signed forms, and support-related screenshots. Opportunity and any custom objects built for verification or onboarding workflows are close behind and often get overlooked because they're not the first place admins think to check.
Should dummy files preserve the original file type and size?
Yes, preserving file type and an approximate size class matters if developers or testers rely on file-upload triggers, validation rules, or storage limit testing. A placeholder PDF of similar size to the original lets those tests run normally without exposing any real personal data inside the document.