MagAudit vs CodeRabbit
Written by the people who make one of the two. Read it with that in mind.
Where CodeRabbit is ahead, and it is not close
CodeRabbit is far bigger and far better funded than we are, and their bot has
commented on more than 6.5 million public pull requests — a figure
anyone can reproduce with a GitHub search for commenter:coderabbitai[bot],
which is how we got it on 18 September 2026. They review reasoning and design, not just
patterns. If you want a second opinion on how the code is written, they do that
and we do not.
We removed three other figures that used to sit in this paragraph — customer count, revenue, weekly volume. They came from reporting we could not re-verify today, and a number we cannot check has no business on a page comparing us to someone else.
We put both of them on the same pull request
On 18 September 2026 we installed CodeRabbit on our own public repository and opened a pull request containing a file we had made defective on purpose. Both reviewers saw the same diff at the same moment. CodeRabbit ran on its own defaults throughout — it says so in its comment — because a comparison where we configure the other side is worth nothing. We published our own expected result before opening the pull request, as a SHA-256 hash committed twelve seconds earlier, so it could not be edited afterwards.
| Planted defect | MagAudit | CodeRabbit |
|---|---|---|
| Hardcoded credential | Found | Missed |
pickle.loads on outside data | Found | Missed |
shell=True with an interpolated value | Found | Missed |
| TLS verification disabled | Found | Found |
| SQL built from a variable | Found | Found |
| Double refund — check-then-write race | Missed | Found |
| Fee truncated to zero by integer division | Missed | Missed |
| Three pieces of correct code that look wrong | 0 false alarms | 0 false alarms |
Where they beat us, they beat us on the thing we cannot do. CodeRabbit found that a refund function reads a balance, checks it, then writes it with no transaction and no row lock, so the same entry can be refunded twice. Every line of that function is valid Python and valid SQL; there is no pattern to match. It also followed a value out of one function and into another, and ran the query in a sandbox to show the effect. That is a class of bug we do not detect and do not claim to.
Where we were ahead, it was on severity: it did not raise the committed credential, the
pickle.loads or the shell=True. The last two are textbook remote code
execution. And the integer-division bug that quietly zeroes small fees was missed by both of
us, which we would rather write down than leave out.
One asymmetry we have to declare, because it favours us. CodeRabbit could see our findings and we could not see its: our check run publishes annotations on the pull request, and its comment quotes them. It also ran Ruff and OpenGrep. We ran on the diff alone. Its three findings were reached with our five already visible, and it still did not echo them — but the two columns are not strictly comparable, and saying so is cheaper than being caught not saying it.
The first attempt at this comparison was ours to lose and we lost it: the fixture documented itself, CodeRabbit read the comment saying it was a deliberate test and correctly reported nothing, and we fired anyway because we do not read context. That round is published too, including why it did not count. Both rounds, the answer key and the hash are here — one file, one run each, so read it as an illustration of where two tools differ in kind, not as a score.
What their own market says goes wrong
These are not our numbers, and we could not re-verify their source today, so read them as reported rather than as measured. A third-party audit of 28 pull requests was reported to find that 15% of CodeRabbit's comments were noise and 21% were nitpicking. We keep the figure because it is the kind of number we ask to be judged by ourselves — but we will not present someone else's reporting as if we had measured it.
What we can say from our own side: we publish our own false-positive rate, derived from the tests, at this page. They do not publish theirs. That asymmetry is the argument; the 15% is not.
The three differences that follow from that
| CodeRabbit | MagAudit | |
|---|---|---|
| Configuration before it is useful | A .coderabbit.yaml to curate |
None required. An optional .magaudit.yml exists if you want to tune it |
| If it is wrong about your code | Config file, or @coderabbitai ignore |
Reply @magaudit ignore <id>, a config file, or the marker your existing scanner already uses — pragma: allowlist secret, nosec, nosemgrep, gitleaks:allow |
| Pricing | From $24 per developer per month | Flat per organisation. Ten people: €69, not $240 |
| Known vulnerabilities in the dependencies you add | Add-on, usage-based pricing | Included. Checked against the GitHub Advisory Database, and only reported when the line pins an exact version |
| Dockerfiles, workflows, Terraform, Kubernetes | Via linters you configure | Included. Including pull_request_target with checkout of the contributor's branch — the one that hands your secrets to anyone who opens a pull request |
| One-click fixes | Generated by a model — powerful, and occasionally wrong | Derived, not invented. Offered only where the correction follows from the rule; where it does not, no suggestion at all |
| Run it locally before you push | CLI, on paid plans | The same detector, offline. If it is clean here it is clean there |
| Its own false-positive rate | Not published | Published, and derived from the tests |
| Reads the code already in the repo | Reviews the diff | Yes, at install |
| README badge | — | Yes |
What we are not claiming
We are not saying we are quieter than CodeRabbit. We have not measured them, and the 15% above is someone else's reporting that we could not re-verify today — so we mark it as reported and we do not restate it as ours. Until 18 September this paragraph said we linked that audit. We did not, and there was no link on the page. That was wrong and it is corrected here rather than quietly removed. What we can say is what our rate was, because we measured it and published it: the first automated run produced 150 findings over 249 public pull requests, 79 of them critical, and almost none were real. Here is what they were.
The two cases that cost us a public apology
We were preparing security notices for public repositories, and we check every one by hand before sending. Both of these were wrong, and both are now pinned by a regression test written from the exact line:
1. A variable forced empty is the fix, not the bug
...(command === "build"
? { "import.meta.env.VITE_OPENAI_API_KEY": '""' }
: {}),
That line removes the key from every production build. A six-line
comment above it explains that Vite inlines VITE_* into the
bundle. We would have told a team that understands this better than our
scanner does that they have the bug they had already fixed.
2. Reading a variable in Node is not publishing it
module.exports = {
openAIApiKey: process.env.VITE_OPENAI_API_KEY,
}
A translation tool's configuration. It runs in Node at build time. The
VITE_ prefix misleads: the variable is never referenced by client
code, so nothing reaches the browser. Any scanner that greps for the prefix
will flag this, and be wrong.
What we measure, and publish
| Measurement | Result |
|---|---|
| First automated run over 249 public pull requests | 150 findings, 79 CRITICAL — almost none real |
| Same population after four rounds of fixes | 39 findings, 0–1 CRITICAL |
| Dependencies resolved live against PyPI and npm | the live count |
| AI-hallucinated packages confirmed | 0 |
That last row works against us: we built this around invented package
names, measured it, and found the attack is rarer than the industry suggests.
A package that does not exist makes pip install fail and CI go
red, for free. We publish it because a vendor you can check is worth more than
a vendor you have to believe.
Last reviewed: 21 September 2026. Every figure about us is our own measurement and reproducible against our public endpoint and our error record. Figures about CodeRabbit are either verifiable by you in one step — their price, read from their own pricing page; their public comment count, from a GitHub search — or explicitly marked as reported by someone else. Until today this line said every figure here was our own measurement, which was not true of four of them. If you find anything else on this page that does not hold up, tell us.