Skip to main content
AI Found 23,000 Bugs. Humans Reviewed 1,900.Incident
3 min readFor Compliance Teams

AI Found 23,000 Bugs. Humans Reviewed 1,900.

The Problem with AI-Driven Vulnerability Detection

Anthropic's Claude Mythos Preview scanned 281 open-source projects and identified 23,019 potential vulnerabilities. External security firms reviewed only 1,900 of these, resulting in 88 published advisories. Of these, just 27 were assigned CVEs. This leaves 21,119 findings unreviewed, their severity unknown. This backlog highlights the challenge of relying solely on AI for vulnerability detection without adequate human oversight.

The Gap in Detection and Review

Details on the timeline of Anthropic's analysis and reviews are sparse. Here's what we know:

  • 23,019 vulnerabilities were identified across 281 projects.
  • 1,900 findings were reviewed by external firms.
  • 88 advisories were published.
  • 27 CVEs were assigned.
  • Echo's survey of over 80 senior US security leaders highlights detection overload as a significant issue.

The disparity between detected vulnerabilities and those reviewed underscores a critical failure in the process.

Missing Controls and Failures

Inadequate Vulnerability Assessment Scoping. Running an AI scanner without the capacity to review the results is ineffective. It's like ordering 23,000 X-rays with only three radiologists available.

Lack of Severity Calibration. The AI's severity ratings often didn't match human assessments. AI can identify patterns but lacks the context to evaluate business impact or exploitability.

Mismatch in Resource Planning. Echo's survey found that 37% of security leaders cite detection overload as a barrier to improving security. This isn't unique to Anthropic; it's an industry-wide issue where detection outpaces review capability.

No Plan for Unreviewed Findings. There's no strategy for the 21,119 unreviewed findings. Without a plan, these findings are effectively ignored.

Compliance Standards and Requirements

PCI DSS v4.0.1 Requirement 6.3.2 mandates identifying vulnerabilities using recognized sources and assigning risk rankings. AI-generated severity ratings without human validation don't meet this requirement.

ISO/IEC 27001:2022 Annex A.8.8 requires timely information on vulnerabilities and appropriate action. Generating unreviewed alerts doesn't equate to actionable intelligence.

NIST 800-53 Rev 5 Control RA-5 requires remediation based on risk assessments. This involves human analysis, not just data collection.

SOC 2 Type II CC7.1 requires monitoring systems for anomalies. If your team can't review all findings, you're not effectively monitoring.

Actionable Steps for Your Team

Align Detection with Review Capacity. Before deploying an AI scanner, determine how many findings your team can review weekly. Adjust the scope or add reviewers to match this capacity.

Calibrate AI Severity Ratings. Compare AI-flagged findings with human assessments. If discrepancies are consistent, document a translation layer to align AI and human severity ratings.

Prioritize Before Scanning. Focus AI analysis on high-risk projects or dependencies. Not all code requires equal scrutiny.

Set a Review SLA. Commit to a timeline for reviewing critical findings. If you can't meet this, adjust your detection strategy.

Track Your Backlog. Log unreviewed findings as technical debt. Reporting them in your risk register can prompt leadership to allocate resources or adjust detection scope.

Understand Coverage vs. Security. Scanning many projects doesn't mean they're secure. Document your review rate for compliance purposes.

AI can enhance detection, but it can't replace human judgment. If findings outpace your ability to act, you're not improving security; you're just documenting risk. Focus on actionable insights, not just data collection.

Topics:Incident

You Might Also Like