Why We Built Sherlock Audit Engine

The orchestration layer for AI-native security review. Frontier LLMs, purpose-built AI auditors, and AI-native security researchers run against your codebase in one coordinated engagement, with every validated finding consolidated into one final audit result.

The orchestration layer for AI-native security review

For months, we kept Audit Engine quiet while we tested one ambitious idea against real, high-stakes code: bring the strongest AI auditing approaches we could find into the same review and see what they could uncover together. The goal was to build the ultimate engine for finding vulnerabilities and assemble it with the strongest AI security approaches we could find. Leading AI auditors. Frontier models. World-class security researchers using their own AI workflows and tools. Everything went into the same fixed review. What came out was one validated, deduplicated audit outcome. We learned where AI is already extraordinary. We learned where it still falls short. We learned that different systems fail differently, and that coordinating those differences produces a stronger review.

Those lessons became the foundation of Audit Engine.

Sherlock Audit Engine is live.

Security intelligence is multiplying

AI has created a new abundance of security intelligence.Frontier LLMs can reason through unfamiliar code at remarkable speed. Purpose-built AI auditors are wrapping those models in specialized harnesses and security workflows. Researchers are using AI to understand systems faster, probe more attack paths, and push their own methodology further.The field is expanding further every week, and that capability is available to attackers too.Black hats can now use increasingly capable AI to search for vulnerabilities faster and at far greater scale. For teams responsible for high-stakes code, the question is increasingly who finds the vulnerability first: the defenders, or an attacker. That changes the standard for defense. Finding 99 of 100 vulnerabilities may still not be enough if the one you miss is the one an attacker finds first. For high-value systems, billions of dollars can depend on closing that last gap. At the same time, protocol teams face a practical challenge: every additional AI system brings more setup, context sharing, dashboards, duplicate findings, false positives, and output for engineers to evaluate.

Audit Engine is designed to solve both problems. It brings the strongest vulnerability-finding approaches into one coordinated audit, designs the right mix of tools and participants around each scope, and aims to maximize real vulnerabilities found per dollar spent.

Maximum vulnerabilities per dollar spent, minus the false positives

Audit Engine puts the entire field of vulnerability-finding approaches in one room.

Image

Frontier LLMs. The strongest (and newest) general-purpose models available, optimized for relentless vulnerability-finding.

Purpose-built AI auditors. Specialized systems designed to outperform frontier LLMs, each with its own architecture, methodology, and strengths.

AI-native security researchers. World-leading security researchers combining years of expertise with their own AI workflows, custom tools, and original investigative approaches.

Every layer works independently, but their output is brought together inside one coordinated review.

Sherlock orchestrates the review around them, bringing every submission into one system for judging, validation, clustering, deduplication, and synthesis.

That combination is where Audit Engine becomes more powerful than any individual approach. Different systems reach different parts of the attack surface, and Sherlock combines their collective output into a higher-intensity security review.

Signal without the pileup

AI can generate an enormous amount of security findings. Volume alone does very little for the team responsible for shipping code.

The real challenge is turning that volume into consolidated and valuable security signal.

Which findings are valid? Which submissions describe the same underlying issue? Which severities hold up under review? Where is the noise coming from? And which individual AI auditing approaches are actually performing best against the code?

Audit Engine handles that burden inside the engagement.

As findings arrive, they move through Sherlock’s proprietary AI judging system, with multiple layers of validation and two iterative improvement cycles before reaching you.

Audit Engine also uses an “Invalid but Interesting” category to preserve useful analysis that falls short of a validated finding, while filtering noise and consolidating duplicates.

That standard matters to us. Sherlock has always optimized to maximize valid vulnerability discovery while keeping noise and low-quality spam out of protocol teams' way.

Sherlock has years of experience across two distinct review models: Collaborative Audits deliver focused, private reviews with hand-selected researchers.

Audit Contests

bring a much broader field of security researchers into a competitive review, with Sherlock judging and validating submissions at scale.

Audit Engine draws on Sherlock’s experience across both models: concentrated expertise where it matters, broad participation where it adds value, and Sherlock’s judging infrastructure tying the output together.

Every engagement leaves a map

One of Audit Engine's core benefits is that it shows how every participant performs across 20+ metrics that matter to your review.

Image

Coverage (share of validated findings discovered), precision (share of submissions that were valid), speed, focus, noise, overlap, unique findings, and incremental contribution are all measured throughout the engagement.This makes it clear which approaches performed best on your codebase, where they added value, and how much each contributed to the final result. That matters because our data shows security performance is codebase-specific. One AI auditor may deliver exceptional breadth on one architecture and add little on another. Another may produce fewer findings with much higher precision. A researcher may surface an issue that every automated system misses. By the end of the engagement, you receive an audit report with every deduplicated, validated finding, plus a benchmarking report showing how each participant performed on your code across 20+ metrics. That intelligence can shape future security spend, internal tooling, participant selection, and the composition of the next review.

Polygon became the proving ground

Audit Engine started with an ambitious idea, and we weren’t sure it would work in practice. To see whether that model could actually hold up against real, high-stakes code, we needed a serious proving ground. @0xPolygon’s Heimdall V2 review became exactly that.

Image

To see whether that model could actually hold up against real, high-stakes code, we needed a serious proving ground. @0xPolygon’s Heimdall V2 review became exactly that.

The engagement brought more than 200 participants and systems across every security layer into one review: AI auditors, frontier LLMs, specialized security skills, and top security researchers.

What we saw was intriguing.

AI auditors emerged among the strongest performers for overall coverage, showing how capable the best systems have already become. But no individual approach came close to capturing the full collective issue set. Different systems and researchers consistently performed better on different parts of the attack surface.

The Polygon engagement made the blind spots of any single security approach impossible to ignore.

Following Polygon, we ran further engagements on codebases like the entire @SkyEcosystem, and the results have only strengthened this conviction.

There is no single approach capable of finding every vulnerability today. Audit Engine is built around that reality.

When you need Audit Engine

Audit Engine is built for codebases where a critical vulnerability is catastrophic.

A major protocol release. A high-stakes upgrade. Core infrastructure. A system responsible for significant economic value. Any moment when a team absolutely needs the best possible security for their money.

The depth of the review scales with the stakes. Some engagements lean heavily on automated AI approaches for speed. Others bring in deeper human investigation for novel, high-stakes architecture.

Audit Engine adapts the mix around the scope, timeline, budget, and level of scrutiny required. It sits alongside Sherlock’s Collaborative Audits, Audit Contests, and Blackthorn reviews as a dedicated AI-native review format.

Audit Engine future-proofs your protocol’s security

The state of the art in AI security is moving fast enough that today’s best approach may be outdated within months.

Frontier models will get stronger. New AI auditing companies will emerge. Researchers will develop new workflows. Better harnesses, agents, and methodologies will push vulnerability discovery further. The same progress will reach attackers.

Because of this, Audit Engine is constantly improving.

Stronger systems are immediately incorporated into new Audit Engine engagements. As existing systems improve, their contribution can be benchmarked against real code. As the field changes, the composition of an Audit Engine engagement changes with it.

As the frontier advances, Audit Engine advances with it.

Audit Engine brings the strongest available security approaches into one environment, benchmarks what each contributes, and combines their collective capability into a higher-intensity audit.

That future is already here. If you're building high-stakes code that absolutely needs to be secure, Audit Engine is for you.

Sherlock Audit Engine is live

Contact our team

to scope an Audit Engine engagement for your codebase.Note: Audit Engine engagements are currently backlogged, but we are opening up more capacity as fast as we can.