
CodeRabbit launched CodeRabbit Security a few months back. It is powered by AI reasoning with the ability to find deep, complex vulnerabilities in your code. We tested CodeRabbit Security and three other systems on their ability to find known vulnerabilities in open-source repositories.
At its core, CodeRabbit Security works in stages: map, hunt, verify and (optional) fix. It maps the codebase, investigates potential attack paths, independently verifies candidate findings against the source and drafts a fix for supported findings. For the benchmark, we measure whether that process recovers the known vulnerability. To earn detection credit, a finding must identify the vulnerability and explain how it can be exploited, accounting for the application’s defenses along the way.
The results
The following results come from our internal comparative development slice:
Vulnerability detection: CodeRabbit 86.7%, Devin 73.3%, Codex Security Plugin 62.0%, Claude Security 40.0%. Internal comparative development slice.
The benchmark also lets us examine something more specific than a final percentage. When a vulnerability is missed, we can trace exactly where the analysis broke down. An entry point can go unseen, a useful hypothesis can get lost between stages, or a valid finding can get rejected during verification. Each failure points to concrete, fixable work that our pipeline can iterate over.
Building the evaluation from real vulnerabilities
Our evaluation cases come from the GitHub Advisory Database and the OSV database.
Each system received the vulnerable source code, with the advisory and patch withheld. We chose known vulnerabilities in real repositories so the tests would retain the surrounding code and configuration needed to understand each flaw. Whether a vulnerability can be exploited may depend on how the framework behaves, how data moves between files, or whether an earlier authorization check blocks the attack.
The test bed covers the following:
Test bed: 100 vulnerabilities, 94 repositories, 11 languages, 11 vulnerability families. Severity: 12 critical, 50 high, 38 medium.
The cases span authentication, authorization, injection, deserialization, SSRF, XSS, cryptography, data exposure, path handling, protocol integrity, and denial of service. For each case, we identify the commit that fixed the vulnerability and its direct parent. We then check that the parent contains the vulnerable code and that the case meets our evaluation requirements.
What the security agent receives
The agent receives a snapshot of the vulnerable repository. We remove Git history and remotes, and the scan runs without network access. We also withhold the advisory, GHSA or CVE identifier, CWE classification, severity, repository identity, vulnerable file, target lines, fixing commit, and patched source from the benchmark instructions. That information remains in the evaluator’s answer key.
These controls limit the clues available during a scan. Prior exposure during model training remains a separate concern because the vulnerabilities and repositories are public. We treat the setup as a controlled test of vulnerability recovery, with that limitation in mind.
A closer examination of CodeRabbit Security
The benchmark measures vulnerability detection. Fix quality is outside the reported score.
Map
AI agents explore the repository inside a secure sandbox and build a system-level map of the application. They identify entry points, trust boundaries, authentication and authorization checks, and the relationships between them. They also build a reachability graph connecting externally accessible inputs to the code and resources they can affect. This focuses the investigation on attack paths that fit the application’s architecture.
Hunt
The map narrows the field. Specialized agents investigate different risk areas in parallel, including authorization, injection, business logic, data exposure, and AI-specific threats. They trace attacker-controlled input from entry point to sink, examine the trust boundaries and controls along the way, and gather supporting code for each candidate vulnerability.
Verify
Before CodeRabbit publishes a finding, it has to hold up under independent verification.
The verifier reopens the cited paths and checks whether the code is reachable, whether safeguards exist elsewhere in the application, whether the conditions required for exploitation can occur, and whether the evidence supports the claimed impact.
Duplicates are removed. Candidates are also rejected when they depend on test-only or unreachable code, overlook existing protections, or rest on unsupported assumptions. And when the evidence doesn't support a conclusion either way, CodeRabbit says so rather than treating incomplete analysis as proof that a vulnerability is exploitable or that the repository is secure.
(Optional) Fix
For eligible findings, Fix with AI uses the confirmed attack path and application context to draft a scoped remediation on a security branch and open a reviewable pull request or merge request. The PR preserves the technical account behind the finding, from entry point and sink through the reachability analysis, exploitation conditions, and potential impact.
Developers inspect the patch, review the reasoning, and run their existing tests and checks before deciding whether to merge. The finding, its evidence, the proposed fix, and the review history stay together in the development workflow instead of scattering across a scanner dashboard, a ticketing system, and a separate remediation project.
The code paths the benchmark asks scanners to follow
The interactive graphic below explores four examples. Click a case to see the vulnerability, what it tests, and supporting evidence.
Inside the evaluation: deepstream, python-statemachine, Fission, and pymonocypher. CVSS scores: 9.9 (v3.1), 9.3 (v4.0), 9.9 (v3.1), and 5.1 (v4.0). See the case descriptions below for evidence.
Across these examples, recognizing a sensitive operation is only the beginning. The scanner has to explain how the vulnerable behavior arises in that application and identify the source that supports the finding.
Scoring methodology
A scan earns a detection credit when a reported finding matches the labeled vulnerability’s underlying mechanism at an accepted source location. A different vulnerability in the same file receives no credit toward that target. A matching CWE classification is also insufficient if the report describes a different attack path. Duplicate reports count once, and we track infrastructure failures separately from model misses.
This scoring method measures how often a system detects the known target. Additional findings require separate assessment.
Why mixing models wins
Our experiments compared mixed-model configurations with single-model setups. We evaluated how each combination affects the number of known vulnerabilities recovered and how well findings survive investigation, verification, and reporting.
Model selection also involves cost tradeoffs. Using a less expensive model at one stage can lower execution cost, but it can also reduce the number of correct findings that reach the final report. We therefore evaluated each change across the whole pipeline, including how one stage’s output affects the work that follows.
A strong harness is what lets a good model reach its full potential. Pairing a good model with a weak harness still only delivers average results. A weak model, however, can’t be rescued by even the best harness. The two compound and is the reason why CodeRabbit invests as heavily in the harness as in model selection.
We continue tuning model selection and routing across the pipeline. As new models become available, we test where each fits best and which combinations improve detection.
Stay tuned for more updates to CodeRabbit Security!
Watch the CodeRabbit Security launch video below to see it in action.






