AI code review that follows your coding rules: how to test it
A reviewer that loads your rules file is not a reviewer that enforces it. Two instruction-following studies give you three tests that tell the two apart.
Category · 16 of 37 articles
How to run an evaluation, a trial or a rollout without wasting a quarter on it.
A reviewer that loads your rules file is not a reviewer that enforces it. Two instruction-following studies give you three tests that tell the two apart.
Two vendors posted "#1 on Martian's Code Review Bench." Both still improved their scores. Both lost first place. How to read the live leaderboard without getting burned.
Self-hosting an AI code reviewer moves the data, not the risk. Here is what to verify in the harness before you point it at a real repo, with the receipts.
Teams reviewing a growing volume of AI-generated code need to score what a reviewer fails to notice, not only what it flags. A measurable protocol.
A canary test, block tests, and independent judging to prove an AI code reviewer actually follows your team's coding standards, not just claims to.
Ask the right question: AI review tools cut how long PRs WAIT, not much how long they take to READ. Plus a two-week PR-slice protocol.
Teams drowning in AI-generated code often let an LLM review the LLM's own patches. Amazon's judge-correlation work shows why that misses real defects.
Atlassian says Rovo cut PR cycle time 45%. The number is real but self-attested. Here's how to measure whether AI actually reduces your review time.
The volume of AI-generated code is rising faster than review capacity. The fix starts in evaluation design: don't let the model that wrote the code also judge it.
AI is producing more code than teams can review. First-party data on why the old loop breaks (Salesforce, DORA) and what actually scales.
AI made PRs smaller but much more numerous. Reviewing the volume isn't a per-PR speed problem, it's a routing problem. Here's how teams actually triage AI-generated code.
The failure mode that matters in AI review is not missing a bug, it is fluent output that is structurally wrong and easy to trust. How to benchmark for it.
How Martian's Code Review Bench separates reproducible fixed-dataset evals from streaming real-world evals, and the tradeoffs hidden in each.
A practical playbook for how to evaluate AI code review tools: a 9-standard scoring rubric, red flags, a 2-week trial protocol, and vendor questions.
Open source AI code review tools compared: Kodus (AGPL), PR-Agent (MIT), and more — real licenses, BYOK costs, and how they stack up against closed SaaS.
Self-hosted AI code review explained: full-stack vs BYOK vs on-prem runners, verified vendor options, and what deployment really costs in 2026.