Yesterday, Google's Mandiant team published the architecture of a tool they've been running internally for ten months. They call it the Agentic Vulnerability Discovery Harness — AVDH. During a recent incident response engagement involving stolen corporate repositories, the system found over 100 true-positive critical vulnerabilities in two days.
Not potential issues. Not false positives to triage. A hundred real, exploitable, high-severity bugs in 48 hours. Across tens of millions of lines of code.
Mandiant said they're publishing the blueprint so defenders can build similar tools. That's genuinely good. It's also the clearest public confirmation yet that AI-assisted vulnerability discovery at machine speed is real, it works, and it's accessible to anyone who builds it.
How AVDH Works
AVDH is a coordinated pipeline of specialized agents. An Explorer agent reads a target codebase, figures out what it does, and assigns focused specialists to look at authentication, routing, authorization, and domain-specific attack surfaces. Those specialists produce vulnerability hypotheses. Then multiple Validation agents independently assess each hypothesis. A Synthesis agent compares their reasoning and produces a final, prioritized finding.
The design is specifically built to kill noise. Traditional static analysis scanners flag thousands of patterns that resemble known bugs. Most of those flags are wrong, and security teams spend most of their time sorting through them. AVDH agents challenge each other's conclusions and check findings against methodologies written by Mandiant's own consultants. The result is something closer to what a skilled human would find — except it runs around the clock across codebases no human team could cover.
Across ten months of internal use, AVDH has produced tens of thousands of findings, resulting in 12 assigned CVEs from open-source projects and widely used web extensions, with a dozen more currently in active disclosure.
Why This Changes the Threat Model
The old security math went like this: attackers find bugs slowly, manually, opportunistically. Defenders have weeks to months between a vulnerability existing and widespread exploitation. Patch windows were designed around this assumption.
That assumption is being dismantled.
What AVDH represents on the defensive side is a parallel capability that's been building on the offensive side — AI-assisted scanning of codebases and running infrastructure to find exploitable conditions before defenders catch them. The difference is that defenders are transparent about their tools and attackers aren't.
You can see the effect in the recent exploitation timelines. CVE-2025-62593, the critical remote code execution bug in the Ray AI framework, was actively being exploited by the RondoDox botnet two days before it was publicly disclosed last November — meaning attackers found it first, or found it simultaneously. A cryptomining campaign called ShadowRay 2.0 has been running on unpatched Ray clusters with NVIDIA GPUs ever since. (We covered that specifically yesterday.) That pattern — exploitation before or alongside disclosure — is getting more common as scanning tools get faster.
The patch window is compressing. What used to be months is becoming days. CISA added Ray's bug to the Known Exploited Vulnerabilities catalog Monday and gave federal agencies until today to patch. Three days. That's not bureaucracy — that's a realistic reflection of how fast the active exploitation window closes.
What This Means If You're Not Google
Here's the uncomfortable part: the capability Mandiant just described isn't exclusive to large security firms. The agent pipeline they've published is architecturally similar to what anyone with access to a capable model API and some orchestration can build. Security researchers and red teams will implement versions of this. Some already have. Automated scanning of open-source projects, vendor products, and customer environments is going to accelerate.
Small orgs are exposed in a specific way here. You're running software — web apps, internal tools, AI workflow stacks — that was written before this scanning speed was possible. The codebase you shipped two years ago may have bugs in it that a well-prompted agent chain would find in an afternoon. You don't have a security team running continuous code review. The attacker equivalent of AVDH doesn't need to target you specifically to find you; it just needs to sweep the category of software you use.
What to Actually Do
Know what's running. This is less obvious than it sounds. AI tooling has proliferated across dev teams in the last two years. Ray, LangFlow, LiteLLM, Ollama, CrewAI — each of these has an API surface. Do you know which of these your team is running? Which have network-accessible endpoints? Which are current on patches?
Set automated dependency alerts. GitHub Dependabot, Renovate, or pip-audit on a scheduled job will surface CVE additions before they hit the three-day CISA deadline. The ShadowRay campaign ran for nine months on unpatched clusters because no one noticed. Automated alerts close that gap.
Limit exposure of internal tools. Dashboards and admin APIs that are meant to be internal shouldn't be reachable externally, even when the default setup makes it easy to leave them open. Ray's dashboard runs on port 8265 with no authentication by design. LangFlow, Flowise, and similar tools have similar defaults. These are designed for development — not internet exposure.
Think about what RCE on a dev machine actually means. If an attacker gets code execution on a developer's laptop or a CI runner, what can they reach? Production cloud credentials? Shared secrets in environment variables? Access to git history with sensitive strings? The machine itself is usually not the prize — it's what the machine can touch.
Consider what AI-assisted scanning could find in your own codebase. The AVDH architecture Mandiant published is public. Running something like it defensively — auditing your own code before attackers do — is now feasible without a Mandiant-scale team. The tools exist to build a basic version of this.
The Shift Worth Naming
Mandiant published this architecture for defenders. But the more useful framing is: they've confirmed that this level of automated vulnerability discovery has been production-proven at scale. Security is now, in part, a race between automated scanning pipelines. Defenders who don't automate are working at human speed against tools that don't sleep.
Small orgs can't match the scale of a Mandiant red team operation. But they can stop leaving the low-hanging fruit: unpatched frameworks, exposed internal APIs, dev machines with production access. That's where AI-assisted scanning starts its sweep anyway.
This is work we help small teams get ahead of — not as a governance exercise, but as an actual audit of what's running, what's exposed, and what needs to close before the next CVE hits the KEV catalog.