On September 4, 2025, security researcher Oren Yomtov reported a vulnerability through HackerOne that allowed one Cloudflare customer to read another customer's data. The flaw was in Cloudflare's Containers service, where storage blocks weren't properly zeroed between tenant uses. By September 19, Cloudflare had deployed fixes across its infrastructure. Although no customer data was exposed in production, the incident highlights how multi-tenant isolation can fail in ways your monitoring might not catch.
What Happened
Cloudflare's Containers service allocates storage blocks to customer workloads. When a container finishes using a block, it should be zeroed before being reassigned. This zeroing step failed, leaving residual data from one tenant accessible to the next tenant using that block.
Yomtov's testing found residual material on 18 of 24 container placements and across 20 of 22 underlying nodes. The vulnerability was exploitable through normal container operations, not through any sophisticated attack technique.
Timeline
- September 4: Researcher submits report via HackerOne
- September 19: Cloudflare completes all mitigation actions across infrastructure
- 15 days total: From disclosure to full remediation
This timeline is fast for infrastructure-level changes that affect multiple data centers. Cloudflare's automated deployment systems enabled them to push fixes without manual intervention on each node.
Which Controls Failed
Three specific controls either failed or were missing:
Data sanitization between tenants. The service didn't enforce cryptographic erasure or verified zeroing of storage blocks. ISO 27001 Annex A.8.10 requires deletion of information from media when no longer required. This applies to logical deletion in multi-tenant systems, not just physical media disposal.
Tenant boundary testing. Pre-production testing didn't validate that storage blocks were clean between allocations. NIST 800-53 Rev 5 control SC-4 (Information in Shared System Resources) specifically requires preventing unauthorized information transfer through shared system resources. Your test cases should explicitly verify that tenant A cannot read tenant B's residual data.
Runtime isolation verification. The platform lacked continuous validation that isolation was working. You can't assume isolation controls stay effective just because they passed initial testing. Memory and storage can leak across boundaries due to race conditions, incomplete cleanup routines, or optimization logic that skips expensive zeroing operations.
What Standards Require
If you're running multi-tenant infrastructure, several requirements apply directly:
NIST 800-53 Rev 5 SC-4 states: "The information system prevents unauthorized and unintended information transfer via shared system resources." The control enhancement SC-4(2) adds: "The information system prevents unauthorized information transfer via shared resources in accordance with organization-defined procedures when system processing explicitly switches between different information classification levels or security categories."
Your implementation needs to handle both the switch (when a block moves from tenant A to tenant B) and the verification (proving the switch was clean).
SOC 2 Type II Common Criteria CC6.6 requires logical access controls that restrict access to sensitive data. If your SOC 2 scope includes multi-tenant systems, your auditor will want evidence that tenant isolation is both designed correctly and operating effectively. That means automated tests in your CI/CD pipeline and runtime monitoring, not just architecture diagrams.
PCI DSS v4.0.1 Requirement 3.2.1 requires that cardholder data be rendered unrecoverable when no longer needed. In a multi-tenant context, "no longer needed" includes the moment a storage block is released by one tenant. If you process payment data in containers, you need cryptographic erasure or verified multi-pass wiping, not just a zeroing routine that might be skipped.
Lessons and Action Items
Here's what you can implement this quarter:
Build tenant boundary tests into your CI/CD pipeline. Create test cases where tenant A writes known patterns to storage, releases the resource, then tenant B attempts to read. Fail the build if tenant B sees anything but zeros or random data. Run these tests on every deployment, not just during security reviews.
Implement verified erasure. Don't trust that a zeroing operation completed. Read back the block after zeroing and verify it's clean. Yes, this adds latency. The alternative is hoping your cleanup code never has a bug.
Monitor for cross-tenant data leakage in production. Instrument your allocation system to periodically sample released resources and verify they're clean. If you find residual data, trigger an incident. Don't wait for a researcher to report it through HackerOne.
Review your bug bounty scope. Cloudflare's HackerOne program caught this before any customer reported it. If you don't have a bug bounty program, you're relying on customers to notice isolation failures and report them privately instead of publishing. That's optimistic.
Document your tenant isolation architecture. When an auditor asks how you prevent cross-tenant data exposure, you need more than "we use containers." Map out every shared resource (CPU cache, storage blocks, network buffers, memory pages) and document the isolation mechanism for each. Then test that the mechanism works.
Automate your remediation pipeline. Cloudflare fixed this across 22 nodes in 15 days because they could push changes automatically. If you need manual intervention to patch each server, your exposure window is measured in weeks or months. Build the automation now, before you need it during an incident.
The Cloudflare incident had a good outcome because a researcher reported it responsibly and the company responded quickly. Your incident might not be reported at all. It might just leak customer data until someone notices. The controls above help you catch isolation failures before they become breaches.





