# Agent-performed code review finds real issues (`verify.agent_code_review`)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/criterion/verify.agent_code_review

Area: [Checking and finishing](https://feedbackbench.com/criteria/checking.md)

**Definition.** A review mode or review agent finds real defects in code or PRs, without re-flagging code that already passed or producing noise.

**Boundary.** Not this: see [Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md) for fixing a known symptom. Not this: see [Caves to or argues with the user's judgement](https://feedbackbench.com/criteria/work.sycophancy_pushback.md) for reversing findings under pressure.

Rated author-weeks, all agents: 570. Complaint share: 27%.

## The brief

Written by Claude Opus 5.5 from 57 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**Reviews catch real bugs when a different model does the reviewing.**

TL;DR:

- OpenAI Codex earns praise mostly as the second reviewer auditing Claude Code output.
- Claude Code reviewing its own work draws complaints about missed defects and endless fix loops.
- Diff-only review misses runtime bugs. Users want independent reviewers and severity-filtered findings.

In plain terms: Users increasingly run one agent to write and another to review. Cross-family checks surface real gaps. Self-review glosses over defects, defers to code comments, and can spiral into rounds where every fix spawns fresh findings.

### How it breaks

- **Reviews gloss over real defects** ([Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md)). The worst reports are not noise but silence: review passes that see a problem and never flag it.
  Users describe reviewers that skip code that is unused or half-wired without a mention. One post says a review trusted a comment claiming a risky shortcut was intentional and let it through. Another found a reviewing sub-agent missed unapproved changes it was explicitly told to hunt for. The shared pattern is a reviewer that reads intent from the code's own narration instead of checking behaviour.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-24: “yeah i've seen some of these comments about it being lazy and spend a couple hours with it. it is absolutely lazy. had 5.6 terra code review and it literally picked up some stuff that it just decided to completely gloss over. like, didn't even flag that it was there or not being used etc.” [source](https://www.reddit.com/r/codex/comments/1wp5x1a/gpt6_sol_is_not_good/pbsxeuk/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-15: “you can actually review every line of code. if you're not, you've no idea what you're shipping. claude's not infallible. i'm right now going through the process of cleaning up a backend that was vibe-coded by a product guy with claude, and it trusts input from jwt's without verifying its signature. claude's "review" didn't pick it up, i assume because there's a comment above the decode that states it's on purpose and planned for a future iteration. if you're not reading the code you're shipping, you're asking for trouble.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wgm4si/engineers_who_write_all_their_code_with_claude/p9wv1ic/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-02: “yesterday, when i was writing a spec, i asked a sub-agent to review it specifically for any decisions or changes that had been made without my approval, and even the sub-agent missed them. things have been pretty bad lately, so i've been reviewing the spec documents very carefully myself. that's how i found several things in there that i never instructed it to add in the first place.” [source](https://www.reddit.com/r/codex/comments/1w50rhx/codex_ignored_an_explicit_uuid_v7_requirement_in/p7bs40p/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-01: “when i highlighted that the code base had grown too large and will soon run into context issues, i was chastised because it was “hiding the wins of engineering”. also, i mentioned that there’s areas to start spitting into more manageable chunks so there’s smaller code reviews. it was praised but then told it was out of my scope to comment and i shouldn’t comment as it distracts from delivery. there are a lot of “let’s pick one lane” moments i’ve been having with ai and code reviews. what’s more disturbing is when it’s easy to spot the problem looking at the code and no review caught it. it was very much also “this is the problem” from ai code reviews but dismissed.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w3nc4h/the_i_dont_know_claude_wrote_this_pandemic/p738fhd/)

- **Review loops that never converge** ([Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md)). Review rounds can generate more work than they remove, through inflated severity, flood-level findings, and fix cycles that breed new bugs.
  One user reports an overnight run where the agent went 40 rounds against its own review without settling. Others say a reviewer inflates severity and proposes band-aid fixes, or that review skills misfire and turn simple tasks into a review ceremony. Users ask for severity-filtered findings and a cap on non-convergent automated review loops.
  Evidence:
  - Complaint, GitHub Copilot, r/cscareerquestions, 2026-09-04: “lgtm. but can you please look at the 100 issues copilot reports? thanks.” [source](https://www.reddit.com/r/cscareerquestions/comments/1w6sz0g/rant_how_do_you_deal_with_ai_slop_from_more/p7uamfp/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-14: “seems to depend on workload i guess. there's just as many people who like codex more. for me, cc opus and even fable can't pass a /code-review to save its life. every review round begets more bugs, i've had opus go a horrifying 40 rounds overnight against its own review and git integrated codex. meanwhile astra has been one-shotting double reviews left and right, including a zero downtime ecs cloudformation to eks terraform migration. i subscribed to 20x codex for the first time and must've merged 200 prs this weekend. i've been building this app since august last year and had only done 1600 total before.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wftqej/trying_codex_for_the_first_time_very_frustrating/p9p8763/)
  - Complaint, OpenAI Codex, r/ClaudeCode, 2026-09-25: “using it as my daily driver with fable advising (/advisor). will explicitly ask for a fable review if the work is complicated. i do have codex reviewing in github with opus validating the issues and fixing them. astra is a great reviewer but it tends to inflate severity and recommend overly complex, bandaid fixes” [source](https://www.reddit.com/r/ClaudeCode/comments/1wq1pjh/opus_55_for_everything_or_mixing_models_across/pc0bh68/)
  - Complaint, OpenCode, r/opencode, 2026-09-11: “do you have any skills that miss fire? it was the case for me. i had some code review, change review, and other review skills that were constantly miss-firing, turning every simple coding task into a review ceremony” [source](https://www.reddit.com/r/opencode/comments/1wd7aka/anyone_else_spending_10x_more_time_on_reviewdebug/p93tdpe/)

- **Same model family shares blind spots** ([Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md)). Users trust findings far more when the reviewer comes from a different model family and never sees the writer's session.
  The dominant workflow in these posts pairs a writer from one vendor with a read-only reviewer from another. Users argue an agent grading its own session confirms its own assumptions, and that diff-only, fresh-context review is the only version they trust. An independent reviewer model separate from the writer is the most requested change on this page. A minority counterpoint says same-model review in a new context still finds correct issues.
  Evidence:
  - Complaint, Devin, @cognition, 2026-09-23: “@dewyashtwts @supercodeai @devinai @cognition self-review is the part i'd push back on. an agent grading its own session just confirms its own blind spots. fresh reviewer, zero access to that chat, diff only, that's the only way i trust it.” [source](https://twitter.com/1415688330/status/2102647905425465627)
  - Praise, OpenAI Codex, r/cursor, 2026-09-17: “on the same project i keep claude as the one that patches. cursor and codex run as read-only reviewers on the plan, different model families, and neither sees the other's notes. same-family loops just share blind spots. i only concede a finding after i can point at the file in the repo.” [source](https://www.reddit.com/r/cursor/comments/1wih7y1/people_who_own_both_cursor_and_claudecodex_plans/pab6n7k/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-26: “i also prefer to have different models review test suites in new contexts. i wouldn’t trust one person (including myself) to write and test something thoroughly end-to-end so i don’t let my models do it either.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqfne5/whats_the_strategy_to_understand_what_your_app_is/pc423fg/)
  - Complaint, Cursor, @cursor_ai, 2026-09-26: “@cursor_ai self-verification on cursorbench is a model skill, not a permission. an agent that checks its own diffs can still ship the wrong change if the only reviewer is the same loop that wrote it.” [source](https://twitter.com/1288646414394896389/status/2103764156361150570)

- **Diff review cannot see runtime** ([Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md)). Static review of the diff passes bugs that only show up when the app actually runs.
  One user's worst bug had a clean diff, passed two review bots, and was caught only by a real browser check after the agent had marked the task done. Praise goes to verifiers that test the live app against the spec and report how many spec items pass, not just a silent pass. Spec-check plugins draw similar credit for catching drift.
  Evidence:
  - Complaint, Cursor, @cursor_ai, 2026-09-22: “@kentcdodds @cursor_ai @devinai @coderabbitai bugbot and coderabbit review the diff. they don't run the app. my worst bug, a chart line drawn twice, had a clean diff and passed static review; only a real browser check caught it, because the agent had already marked the task done.” [source](https://twitter.com/1748692144213405696/status/2102277944152576079)
  - Praise, Warp, @warpdotdev, 2026-09-08: “perfect example of llm-as-judge in the runtime path from @warpdotdev. instead of grading the coding agent’s code on a rubric (boring), it evals the application and blocks the flow unless it passes. - implementation agent codes a feature - verification agent evals the live app against the spec using a computer-use model - sends it back around the loop if needed astra and its off-the-charts computer-use capability will make this pattern more common at runtime.” [source](https://twitter.com/65392279/status/2097386394347807178)
  - Praise, Warp, @warpdotdev, 2026-09-08: “@josharosen @warpdotdev what makes this work: the verifier gets the spec, not the implementation's summary of itself. the run that wrote the feature is its worst judge. one addition: the verifier prints a denominator. 'passed' is silence; '7 of 9 spec items pass' is evidence.” [source](https://twitter.com/97718205/status/2097459317465293144)
  - Praise, Amp, @AmpCode, 2026-09-18: “@adamwathan wrote an @ampcode plugin that checks if an llm’s code output follows the spec i give it. it checks against specific verification criteria and claims. it’s already caught some stuff for me” [source](https://twitter.com/33135576/status/2101006859314561178)

### Who stands out

- **OpenAI Codex (stronger)**. Codex wins here largely as the outside auditor that checks another agent's work and surfaces gaps the writer missed.
  Posts describe Codex reviewing Claude Code output as a standing habit that cuts cleanup work sharply. Users call it a second set of eyes and say the two model families complement each other. Complaints exist: one user calls a recent model lazy at review, and another says a separate chat-based review finds issues Codex missed.
  Evidence:
  - Praise, OpenAI Codex, r/ClaudeCode, 2026-09-15: “my codex reviewed what my claude code did for quite a while now.. reduces the slop i have to fix by a metric fuckton..” [source](https://www.reddit.com/r/ClaudeCode/comments/1wga9u4/ai_is_using_ai_now/p9zqkdd/)
  - Praise, OpenAI Codex, X search: OpenAI Codex, Codex CLI, Codex app, 2026-09-26: “i am using it in reverse 🔁 i have claude code create something and have it audited by codex cli. today, i posted 2 exam preparation posts after passing them through codex ✅ when i show it to a model from another company, gaps that i wouldn't notice myself come out 👀🐢 <strict_link>” [source](https://twitter.com/2037901657799794688/status/2103826697850360135)
  - Praise, OpenAI Codex, r/vibecoding, 2026-08-31: “i built [<strict_link> entirely with claude models, currently i use codex for code review but still claude as my main, i would advice you to try claude in claude code as a start, you are missing a lot because atm claude and gpt models complement each other.” [source](https://www.reddit.com/r/vibecoding/comments/1w2zfh5/already_have_2_plus_codex_subs_and_looking_at/p6wglda/)
  - Praise, OpenAI Codex, r/codex, 2026-09-09: “mostly in codex i use it as a second set of eyes on complex problems when i build with claude.” [source](https://www.reddit.com/r/codex/comments/1wbaote/how_are_you_guys_using_gpt_astra_right_now/p8ojws1/)

- **Claude Code (weaker)**. Claude Code gets heavy praise as a writer and for review skills, but self-review is where complaints cluster.
  Users say Claude reviewing Claude shares the same blind spots, and report review rounds that keep producing new bugs instead of closing out. Others push back: Opus reviewing Opus in a fresh context usually returns correct findings, and the review skills draw explicit praise. The split tracks whether a different model sits in the reviewer seat.
  Evidence:
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-24: “this biggest problem with any claude instance reviewing your “claude” code is that they come with from same derivative. imo, this is where adversarial supports compounds value” [source](https://www.reddit.com/r/ClaudeCode/comments/1wowilt/i_still_dont_understand_this_agentic_workflow/pbqr654/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-14: “seems to depend on workload i guess. there's just as many people who like codex more. for me, cc opus and even fable can't pass a /code-review to save its life. every review round begets more bugs, i've had opus go a horrifying 40 rounds overnight against its own review and git integrated codex. meanwhile astra has been one-shotting double reviews left and right, including a zero downtime ecs cloudformation to eks terraform migration. i subscribed to 20x codex for the first time and must've merged 200 prs this weekend. i've been building this app since august last year and had only done 1600 total before.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wftqej/trying_codex_for_the_first_time_very_frustrating/p9p8763/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-15: “you can actually review every line of code. if you're not, you've no idea what you're shipping. claude's not infallible. i'm right now going through the process of cleaning up a backend that was vibe-coded by a product guy with claude, and it trusts input from jwt's without verifying its signature. claude's "review" didn't pick it up, i assume because there's a comment above the decode that states it's on purpose and planned for a future iteration. if you're not reading the code you're shipping, you're asking for trouble.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wgm4si/engineers_who_write_all_their_code_with_claude/p9wv1ic/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-26: “well they definitely are because i routinely use opus to review work done by another opus, and it almost always produces findings - sometimes minor and sometimes major, and typically correct. different tasks activate different things in the model.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wpz6ms/do_models_actually_get_nerfed/pc4iee1/)

- **Cursor (mixed)**. Cursor's review agent catches real errors and cuts repeat flags, but loses head-to-head comparisons on depth.
  One user reports the review agent found two real errors in a late-night PR. Another says the bigger win is no fresh flags on code that already passed review. On the other side, a team found Devin's reviews consistently surfaced more bugs than Bugbot on the same PR, and an auto-review classifier drew complaints about blanket rejections.
  Evidence:
  - Praise, Cursor, @cursor_ai, 2026-09-24: “@vibecoderofek @cursor_ai too early for a clean number. i just switched it on. the bigger win so far is noise: no more fresh flags on code that already passed review. i'll share the numbers once i have a full week to compare” [source](https://twitter.com/2201379657/status/2103054480199541159)
  - Praise, Cursor, @cursor_ai, 2026-09-10: “just this morning i had to say to a team member, 'don't feed me this human slop' after being sent a pr at 1:30am. i ran it through the @cursor_ai review agent and two obvious real errors were presented. i've been insisting that they use the tools given them to do the work and if they won't, to at least use them to check your work. not my employee so just going to keep pointing it out in chat.” [source](https://twitter.com/1521677378073866240/status/2098020208543687159)
  - Complaint, Cursor, @cursor_ai, 2026-09-10: “in particular, devin's reviews have been far more detailed than cursor. with the same pr and codebase on github and cursor origin so that the two do not conflict, devin regularly found more bugs and issues than cursor's bugbot. devin's deepwiki is a gamechanger for teams to get new employees up and running much faster, and it's always up-to-date with the codebase. even for myself, i have found it rather useful to reference how specific parts of the codebase work.” [source](https://twitter.com/2070908287978246144/status/2098100357478150466)
  - Complaint, Cursor, @cursor_ai, 2026-09-07: “poteto poteto poteto, is there any issue going on with the auto-review classifier in @cursor_ai? i'm seeing lots of rejections today. even for the simple `gh` cmds like `issue view`. @poteto” [source](https://twitter.com/1109410339/status/2097017413703479494)

- **Google Antigravity (weaker)**. Thin evidence, but posts lean negative, with verification failing to catch cut corners.
  Users say they still need at least one manual review-and-fix round because verification lets shortcuts through. Praise is narrow: occasional unprompted bug hunting, and value as one half of a two-agent cross-check alongside Claude.
  Evidence:
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-23: “that's sad. i hoped that boost could detect cut corners at verification phase. usually i need at least 1 round of a review and fixes because of that” [source](https://www.reddit.com/r/google_antigravity/comments/1wo8qye/what_is_your_experience_with_boost_deep_reasoning/pbltash/)
  - Praise, Google Antigravity, r/google_antigravity, 2026-09-09: “claude sometimes finds bugs in antigravity code, and vice versa. think the best setup is not to rely on a single agent but use two.” [source](https://www.reddit.com/r/google_antigravity/comments/1waream/antigravity_is_incredible_goodbye_claude/p8r3r84/)
  - Praise, Google Antigravity, r/google_antigravity, 2026-09-07: “also i have found like it automatically hunt thosse problems which i have not mentioned sometimes. thats really cool tbh” [source](https://www.reddit.com/r/google_antigravity/comments/1w9vwgv/tbh_flash_38_is_really_good_when_it_comes_in/p8djqld/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-24: “sarcasm is on that take, they have a fast and decent model but dont use it for reviewing commands.” [source](https://www.reddit.com/r/google_antigravity/comments/1w69qmt/finally_someone_popular_bringing_attention_to/pbq07cn/)

### Fine print

- Many praise posts credit a reviewer while discussing a different writer agent, so credit and blame can blur across agents.
- Most agents here have too few posts to separate from the field. Only Claude Code, OpenAI Codex and Cursor carry real volume.

## Top requests

What users ask to add or change, most asked first. 61 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Author-weeks | Posts | Agents (author-weeks) |
|---|---|---|---|---|
| 1 | Independent reviewer model separate from writer | 9 | 9 | Claude Code 4, Cursor 2, OpenAI Codex 1, GitHub Copilot 1, Devin 1 |
| 2 | Richer benchmark reports beyond scores | 8 | 8 | Claude Code 3, Cursor 2, Devin 2, Zed 1 |
| 3 | Automated PR and merge request review | 7 | 7 | Zed 2, Amp 1, Google Antigravity 1, Claude Code 1, OpenAI Codex 1, Cursor 1 |
| 4 | Less noisy, severity-filtered review findings | 6 | 6 | Claude Code 3, Cursor 2, OpenAI Codex 1 |
| 5 | More realistic, frequently updated eval benchmarks | 6 | 6 | Factory 2, Pi 2, Claude Code 1, Cursor 1 |
| 6 | Post-deploy monitoring and deployment acceptance | 4 | 4 | Cursor 4 |
| 7 | Review of plans and subagent output | 3 | 3 | Claude Code 1, OpenAI Codex 1, Pi 1 |
| 8 | Security review with OWASP and fix guidance | 3 | 3 | GitHub Copilot 1, Cursor 1, Factory 1 |
| 9 | Structured, trackable review findings | 2 | 3 | Cursor 2 |
| 10 | Cap non-convergent automated review loops | 2 | 2 | Claude Code 1, OpenAI Codex 1 |
| 11 | Merge gates requiring signoff for risky changes | 2 | 2 | Claude Code 1, Cursor 1 |
| 12 | Review findings linked to tickets and scope | 2 | 2 | Claude Code 1, Cursor 1 |

### 1. Independent reviewer model separate from writer

- Claude Code, 2026-09-26, @ClaudeDevs (X): “@claudedevs this solve a great issue. also devs need automotion for "review" and "fix" on a single sesstion with low token useges with deffrent models so we get some saving on letest and greatest models.” [source](https://twitter.com/1424656235018653697/status/2103672127430000746)
- Cursor, 2026-09-24, r/cursor (Reddit): “i dont, only cause its the same model that does work and checks itself; at least that was the case when i first used it... if it had a model write code, then a completely separate reviewer critique it - id be all over it.” [source](https://www.reddit.com/r/cursor/comments/1wow14z/does_anyone_use_goal/pbqi974/)
- Devin, 2026-09-23, @cognition (X): “@dewyashtwts @supercodeai @devinai @cognition self-review is the part i'd push back on. an agent grading its own session just confirms its own blind spots. fresh reviewer, zero access to that chat, diff only, that's the only way i trust it.” [source](https://twitter.com/1415688330/status/2102647905425465627)

### 2. Richer benchmark reports beyond scores

- Zed, 2026-09-25, @zeddotdev (X): “@zeddotdev a harness report gets much more informative with reruns, task-level traces, and recovery data: retries, human interventions, and rollbacks. cost per accepted change would make the comparison even more useful for teams.” [source](https://twitter.com/2052583923918503944/status/2103567067530318310)
- Cursor, 2026-09-23, @cursor_ai (X): “@cursor_ai cursorbench this type of chart is very suitable for comparing "how much each task costs," but for real projects, we also need to consider the pass rate and manual wrap-up time. if the model is 40% cheaper but makes developers spend an extra 20 minutes checking changes, the perceived cost may not decrease. it would be best to disclose the success rate, rollback times, and total time together.” [source](https://twitter.com/1800749035135074304/status/2102577017371668616)
- Cursor, 2026-09-23, @cursor_ai (X): “@cursor_ai for game prototypes, i care less about a single benchmark score than whether an agent can preserve scene state through a bunch of edits. any plans to show a longer end-to-end task trace alongside cursorbench?” [source](https://twitter.com/1863169058428137472/status/2102549168443236584)

### 3. Automated PR and merge request review

- Claude Code, 2026-09-26, r/ClaudeCode (Reddit): “\+ link another tool to auto review mr's, and have the main ai response to those comments automatically 1hr after submitting to account for reviewer ai's delay.” [source](https://www.reddit.com/r/ClaudeCode/comments/1v776fe/instead_of_make_no_mistakes_what_do_you_genuinely/pc59kxi/)
- OpenAI Codex, 2026-09-18, X search: OpenAI Codex, Codex CLI, Codex app (X): “@thsottiaux when i was delegated to the luna reserve usage, the sandbox does not allow git operations, making it super tough to use. and also would love to see "autofix ci &amp; comments" option in the @chatgpt codex app. that would really be a claude code killer” [source](https://twitter.com/1027581358217080833/status/2100890296964026622)
- Amp, 2026-09-17, @AmpCode (X): “@ampcode @sqs possible to get @typesafeai for review &amp; judgements within amp?” [source](https://twitter.com/107126704/status/2100486464203268153)

### 4. Less noisy, severity-filtered review findings

- Cursor, 2026-09-25, @cursor_ai (X): “@cursor_ai would be interesting to see how it behaves when the evidence is genuinely ambiguous, like a latency bump that's real but belongs to a different change. an agent that says inconclusive correctly might end up more trusted than one that always picks a suspect” [source](https://twitter.com/737610404264706048/status/2103336974816096525)
- Claude Code, 2026-09-05, r/ClaudeCode (Reddit): “i find this very frustrating for pr reviews, i've asked fable to review prs a few times and no matter how bullet proof the code is it'll still pick up on the most obscure inconsequential nitpicky "bugs" possible, it almost never just gives a pr a green light” [source](https://www.reddit.com/r/ClaudeCode/comments/1w7vicj/fable_51_not_getting_shit_done/p828n21/)
- Claude Code, 2026-09-05, r/ClaudeCode (Reddit): “if anything, it would be ideal if it found less bugs - the ones that actually matter.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w7y24r/my_experience_with_opusfable_vs_astra/p7yibdo/)

### 5. More realistic, frequently updated eval benchmarks

- Cursor, 2026-09-24, @cursor_ai (X): “@cursor_ai benchmark says 57.8%, production says “why is this diff 4,000 files”. ship the boring evals too.” [source](https://twitter.com/2315892500/status/2102966746462466395)
- Factory, 2026-09-23, @FactoryAI (X): “i really think @factoryai should have something similar to @cursor_ai cursorbench. i place great trust in factoryai and i'm willing to determine my workflow based on them. <strict_link>” [source](https://twitter.com/1989809437213896704/status/2102743881058254945)
- Pi, 2026-09-17, r/PiCodingAgent (Reddit): “would be great to add evals to see if this actually improves the resulting code or not” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wimfhg/piwarden_a_jevpowered_second_pair_of_eyes_for_pi/pad20ef/)

### 6. Post-deploy monitoring and deployment acceptance

- Cursor, 2026-09-26, @cursor_ai (X): “@cursor_ai rollouts moved "discovering regressions after deployment" one step forward, making monitoring plans and validation more practical than simply letting the agent generate code. if it can also show the basis for each judgment and the rollback path when failures occur, small teams will be more willing to integrate it into production.” [source](https://twitter.com/1800749035135074304/status/2103651575403282524)
- Cursor, 2026-09-25, @cursor_ai (X): “@cursor_ai good instinct -- catching regressions as they deploy beats a scan that runs once a quarter. the gap we keep seeing in live apps is one step further out: something reviews the code change, but nothing watches what happens after it ships. error tracking still gets skipped.” [source](https://twitter.com/2066258931853488128/status/2103320309516611675)
- Cursor, 2026-09-24, @cursor_ai (X): “@cursor_ai this is very practical to put the monitoring plan in the deployment process. for the acceptance of the enterprise agent go-live, in addition to the demo green light, will the runtime evidence (error rate/delay/key write-back and rollback) and regression thresholds be solidified together?” [source](https://twitter.com/1235054222275538946/status/2102940137860776137)

### 7. Review of plans and subagent output

- OpenAI Codex, 2026-09-26, r/codex (Reddit): “astra should review the code that subagents created. otherwise there is a spaghetti code.” [source](https://www.reddit.com/r/codex/comments/1wql2gl/is_agent_routing_still_meta/pc6xqg1/)
- Pi, 2026-09-12, r/PiCodingAgent (Reddit): “ppl just fire and forget... i think u should analyze the whole pre walk, plan and, then, outcome of that task. u should, also, have hooks to keep not only task done but also code quality high” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wafm5s/how_do_you_guys_analyze_chatsinteractions_with_ai/p9ax3nu/)
- Claude Code, 2026-09-05, r/ClaudeCode (Reddit): “tried the three compartment approach and, indeed, the reviewer caught one critical (valid and confirmed by implementor), and a dozen of mediums and minors. thanks! how about a review pass on the planner's output itself? just an idea, but maybe a cleaner and more coherent plan might reduce the surface for the implementer's defects?” [source](https://www.reddit.com/r/ClaudeCode/comments/1w4zdwu/when_should_you_use_clear/p7yd73o/)

### 8. Security review with OWASP and fix guidance

- Cursor, 2026-09-24, @cursor_ai (X): “@cursor_ai faster review is useful, but latency is the easy metric; the dangerous misses are secrets in generated diffs and tenant boundaries that only fail under a real permission set. i’d gate merge on explicit auth-path coverage, not just a green scan.” [source](https://twitter.com/2053606955571105792/status/2103105910294077545)
- Factory, 2026-09-01, @FactoryAI (X): “@chahvivi @enoreyes @factoryai speed is great until the review loop becomes the bottleneck. the useful version of low-friction security feedback is one that catches the secret and tells the builder what to fix next.” [source](https://twitter.com/1671319169793654785/status/2094668248843444733)
- GitHub Copilot, 2026-08-31, r/GithubCopilot (Reddit): “add owasp standards.” [source](https://www.reddit.com/r/GithubCopilot/comments/1w2bke4/suggestions_for_agentsmd_to_make_gpt_56_write/p6xioaq/)

### 9. Structured, trackable review findings

- Cursor, 2026-09-24, @cursor_ai (X): “@cursor_ai the real metric is not review latency alone. track escaped findings and merge-blocker precision by change class, then tune the reviewer against that eval set.” [source](https://twitter.com/2099871292480421888/status/2103142332195594722)
- Cursor, 2026-09-17, @cursor_ai (X): “i have a love-hate relationship with @cursor_ai bugbot, leaning hate whenever it flags something i'm 99% sure is fine. what i really want is a "not now, track it" button. until then, i created a gh action to open an issue, link both ways and resolve the thread. <strict_link>” [source](https://twitter.com/1980521913861730304/status/2100616602689687962)

### 10. Cap non-convergent automated review loops

- Claude Code, 2026-09-26, r/ClaudeCode (Reddit): “if you use it as a reviewer, you’ll end up in a never ending loop with no convergence. that’s a huge problem with the opus 5 family. i hope anthropic can fix it at some point.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wq1pjh/opus_55_for_everything_or_mixing_models_across/pc4k299/)
- OpenAI Codex, 2026-09-09, r/codex (Reddit): “**frontier-simplify** — stops codex from inventing process around its own work, and puts a hard stop on automated code review. the review half is the part i use daily. automated review never terminates on its own: fix three findings, push, get three new ones. so the runner carries the previous round's findings into the next review, reuses the cached attempt when inputs are identical instead of re-running the model, and stops after **3 automatic a” [source](https://www.reddit.com/r/codex/comments/1wavxwy/show_us_all_what_youve_been_building_with_codex/p8oql2w/)

### 11. Merge gates requiring signoff for risky changes

- Cursor, 2026-09-18, r/cursor (Reddit): “honestly the market answer is no. vibe coders who want to understand the code already slow down and read the diff, and the ones who don't care won't sit through your questions. the people who'd pay are teams with review requirements, so build it as a pr gate that flags risky ai changes and forces a signoff, not a learning tool. that's where an actual budget exists.” [source](https://www.reddit.com/r/cursor/comments/1wjz3xe/do_you_need_this_in_your_vibecoding_life/pancs69/)
- Claude Code, 2026-09-08, @ClaudeDevs (X): “@fipetru @claudeai @claudedevs @anthropicai @bcherny @dickson_tsai @amorriscode @trq212 same gap i hit. cloud agents that open the pr still need a human if there is no review loop. a harness with allow lists plus eval checks before the pr step is what makes it zero babysitting.” [source](https://twitter.com/2009223361969442816/status/2097118084502782456)

### 12. Review findings linked to tickets and scope

- Cursor, 2026-09-24, @cursor_ai (X): “@cursor_ai nice improvement—cutting the review loop from 4.8m to 3.8m makes this much easier to fit into a deploy gate. for enterprise rollouts, can the finding + verification result be written back to the ticket/change record with an auditable link?” [source](https://twitter.com/1235054222275538946/status/2102941737350258866)
- Claude Code, 2026-09-16, r/ClaudeCode (Reddit): “put an agent code reviewer in that is a product check. it checks the code against the scope of the ticket and fails if it’s less than 100% coverage or more than 100%” [source](https://www.reddit.com/r/ClaudeCode/comments/1whj547/how_do_you_stop_fable_51_from_scope_creeping_and/pa3vln9/)

## Every agent

| Agent | Overall rank | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|---|
| [OpenAI Codex](https://feedbackbench.com/agents/codex.md) | 2 | Better than peers | 0.542 | 0.510–0.572 | 201 | 159 | 42 |
| [Cursor](https://feedbackbench.com/agents/cursor.md) | 4 | Typical | 0.515 | 0.491–0.542 | 36 | 30 | 6 |
| [Claude Code](https://feedbackbench.com/agents/claude-code.md) | 1 | Worse than peers | 0.472 | 0.448–0.499 | 258 | 176 | 82 |
| [GitHub Copilot](https://feedbackbench.com/agents/copilot.md) | =8 | Too few posts | – | – | 17 | 10 | 7 |
| [OpenCode](https://feedbackbench.com/agents/opencode.md) | 3 | Too few posts | – | – | 16 | 13 | 3 |
| [Google Antigravity](https://feedbackbench.com/agents/antigravity.md) | =5 | Too few posts | – | – | 11 | 2 | 9 |
| [Devin](https://feedbackbench.com/agents/devin.md) | =5 | Too few posts | – | – | 10 | 8 | 2 |
| [Pi](https://feedbackbench.com/agents/pi.md) | 7 | Too few posts | – | – | 7 | 7 | 0 |
| [Warp](https://feedbackbench.com/agents/warp.md) | 15 | Too few posts | – | – | 4 | 3 | 1 |
| [Cline](https://feedbackbench.com/agents/cline.md) | =8 | Too few posts | – | – | 2 | 2 | 0 |
| [Factory](https://feedbackbench.com/agents/factory.md) | =11 | Too few posts | – | – | 2 | 2 | 0 |
| [Amp](https://feedbackbench.com/agents/amp.md) | =11 | Too few posts | – | – | 2 | 2 | 0 |
| [Grok Build](https://feedbackbench.com/agents/grok-build.md) | 16 | Too few posts | – | – | 2 | 2 | 0 |
| [Augment Code](https://feedbackbench.com/agents/augment.md) | 17 | Too few posts | – | – | 2 | 2 | 0 |
| [Zed](https://feedbackbench.com/agents/zed.md) | =8 | Too few posts | – | – | 0 | 0 | 0 |
| [Kiro](https://feedbackbench.com/agents/kiro.md) | 13 | Too few posts | – | – | 0 | 0 | 0 |
| [Conductor](https://feedbackbench.com/agents/conductor.md) | 14 | Too few posts | – | – | 0 | 0 | 0 |

## Posts

Receipts rule: The 5 most recent praise and complaint posts per area (first 700 characters) and 3 per criterion (first 450 characters).

### OpenAI Codex

- Praise, 2026-09-27, X search: OpenAI Codex, Codex CLI, Codex app (X): “cancelled my coderabbit subscription this morning spent 2 hours writing a custom script that uses codex cli + gpt-6 luna at max reasoning effort to do the exact same thing. reviews prs, leaves comments, catches issues and i also sync my review rules from notion so it actually follows my standards and it's basically free. runs off my existing codex sub, and luna at max reasoning is so token-efficient it barely registers as usage meanwhile coderabb” [source](https://twitter.com/1895398810299318272/status/2104152229334896683)
- Praise, 2026-09-27, X search: OpenAI Codex, Codex CLI, Codex app (X): “@__roycohen yeah i mean, i love the codex app, i've loved 5.6 sol and was completely out of anthropic. but there's no denying that if you give the same task to astra and to opus 5.5 right now, opus 5.5 feels significantly more magical. astra is a great reviewer of opus though.” [source](https://twitter.com/174970722/status/2104223206823649358)
- Praise, 2026-09-27, r/ChatGPTPro (Reddit): “i have claude max then then it calls codex and gemini as reviewers. been working well for me.” [source](https://www.reddit.com/r/ChatGPTPro/comments/1wr7nqy/if_you_were_paying_which_one_would_you_go_with/pcgptpu/)
- Complaint, 2026-09-26, r/codex (Reddit): “never had those issues but something happened last night when codex crashed. what i found out using chatgpt for code review is better for me than codex and it’s free. use sol model or 6 pro if you have the pro plan. i normally tell chatgpt exactly what to audit and to read documentation about it. for example i’ve built a 3d11x renderer for an old game (built a new client) and i’ve told chat got precisely to follow directx, nvidia, and amd documen” [source](https://www.reddit.com/r/codex/comments/1wqo1s6/confused_how_to_use_codex/pc5il4y/)
- Complaint, 2026-09-26, r/codex (Reddit): “first thing i got a new model is i make it check a repo for bugs and inefficiencies and opus 5.5 surprised me by finding stuff astra and prior models just didnt.” [source](https://www.reddit.com/r/codex/comments/1wqtd1g/openai_gpu_are_really_cooling_down_theo_just/pc6vyh5/)
- Complaint, 2026-09-25, r/ClaudeCode (Reddit): “using it as my daily driver with fable advising (/advisor). will explicitly ask for a fable review if the work is complicated. i do have codex reviewing in github with opus validating the issues and fixing them. astra is a great reviewer but it tends to inflate severity and recommend overly complex, bandaid fixes” [source](https://www.reddit.com/r/ClaudeCode/comments/1wq1pjh/opus_55_for_everything_or_mixing_models_across/pc0bh68/)

### Cursor

- Praise, 2026-09-24, r/cursor (Reddit): “risk tiers helped, but the bigger win was a second model that only reviews. different family from the one that wrote the patch, read only, never sees the writer's notes. claude only concedes a finding after it greps the actual file. that kills a lot of the hallucinated-import chase before i open the diff.” [source](https://www.reddit.com/r/cursor/comments/1woujtl/i_timed_agent_diff_review_for_5_days_generation/pbqig27/)
- Praise, 2026-09-24, r/cursor (Reddit): “that separation is a useful guardrail. i have been treating the reviewer as a verifier of the actual diff and test result, not as another source of implementation context, which makes it easier to reject findings that do not survive a file check. i may try your rule of requiring a grep before accepting a finding.” [source](https://www.reddit.com/r/cursor/comments/1woujtl/i_timed_agent_diff_review_for_5_days_generation/pbv6kcp/)
- Praise, 2026-09-24, r/cursor (Reddit): “that is a useful refinement. i like turning risky paths into a decision boundary rather than asking the reviewer to reread everything. i did not measure review time by task yet, but your comparison suggests a good next cut would be tracking time to sign-off alongside pass rate and token cost.” [source](https://www.reddit.com/r/cursor/comments/1woujtl/i_timed_agent_diff_review_for_5_days_generation/pbv76dh/)
- Complaint, 2026-09-26, @cursor_ai (X): “@cursor_ai self-verification on cursorbench is a model skill, not a permission. an agent that checks its own diffs can still ship the wrong change if the only reviewer is the same loop that wrote it.” [source](https://twitter.com/1288646414394896389/status/2103764156361150570)
- Complaint, 2026-09-24, r/cursor (Reddit): “i dont, only cause its the same model that does work and checks itself; at least that was the case when i first used it... if it had a model write code, then a completely separate reviewer critique it - id be all over it.” [source](https://www.reddit.com/r/cursor/comments/1wow14z/does_anyone_use_goal/pbqi974/)
- Complaint, 2026-09-22, @cursor_ai (X): “@kentcdodds @cursor_ai @devinai @coderabbitai bugbot and coderabbit review the diff. they don't run the app. my worst bug, a chart line drawn twice, had a clean diff and passed static review; only a real browser check caught it, because the agent had already marked the task done.” [source](https://twitter.com/1748692144213405696/status/2102277944152576079)

### Claude Code

- Praise, 2026-09-27, r/ClaudeCode (Reddit): “if you have 3 accounts with 20$ plan then yes, it's enough for heavy coding. /code-review consumes a lot, and it's essential to find bugs , that people always miss” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrf0ld/is_claude_pro_actually_worth_20_just_for_one/pcbysrl/)
- Praise, 2026-09-27, @ClaudeDevs (X): “@claudedevs if you're wondering what to do with these: i use them to have an isolated claude tear my plans apart before i build. open-sourced it: <strict_link>” [source](https://twitter.com/14996888/status/2104055969881665716)
- Praise, 2026-09-26, r/ClaudeCode (Reddit): “at least it actually finds real issues, astra just breaks everything whilst you're thinking it's fixed things.” [source](https://www.reddit.com/r/ClaudeCode/comments/1woe3lo/is_opus_55_really_better_than_fable_in_your/pc3cssl/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “the mistaken bug report is worth keeping in the write-up because it shows how a reference can point to real code and still get the explanation wrong. publishing that example makes it much clearer why a human reviewer needs to check whether the code actually supports the claim. that includes checking why the compatibility branch exists” [source](https://www.reddit.com/r/ClaudeCode/comments/1wri3sa/i_let_my_claude_code_plugin_describe_flasks/pce9k9p/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “the problem with reviews is, that claude will always find something. had a talk with my colleagues about that and we all agreed on that.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wropen/i_keep_hitting_usage_limits_on_20x_plan_have/pcei3f2/)
- Complaint, 2026-09-26, r/ClaudeCode (Reddit): “if you use it as a reviewer, you’ll end up in a never ending loop with no convergence. that’s a huge problem with the opus 5 family. i hope anthropic can fix it at some point.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wq1pjh/opus_55_for_everything_or_mixing_models_across/pc4k299/)

### GitHub Copilot

- Praise, 2026-09-27, r/GithubCopilot (Reddit): “get a claude or chatgpt pro sub then proxy it into copilot via byok to use the harness. i keep a copilot sub also for adversarial review (like pair gpt with a claude to argue), copilot code review on prs, and overages.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wr5io2/best_cheaper_alternative/pc9tpgy/)
- Praise, 2026-09-26, r/ExperiencedDevs (Reddit): “i don't know about you, but my ape brain couldn't for the life of me review any pr with even close to the depth and thoroughness of the team of claude, codex, copilot and deepseek agents that i use for my work projects. so the question is really: do you want code quality or do you want kabuki theater and the warm human feeling of "being in control"? coding is ~~largely~~ solved.” [source](https://www.reddit.com/r/ExperiencedDevs/comments/1wql3g2/interviewed_candidates_for_ai_engineer_roles_this/pc5d5jy/)
- Praise, 2026-09-23, r/cscareerquestions (Reddit): “creative ai usage what are some creative ways you use ai to complete work? in visual studio copilot i have an agent file where i add mistakes i make which were pointed out in pr comments. i find it's a big help to code review my work before making pull requests” [source](https://www.reddit.com/r/cscareerquestions/comments/1woinjn/creative_ai_usage/)
- Complaint, 2026-09-25, @GitHubCopilot (X): “alright, looking for a replacement for @githubcopilot reviews. what's the best ai pr review software nowadays? @coderabbitai , @greptile , @cubic_dev_ , @claudeai pr reviews / or @cursor_ai bugbot?” [source](https://twitter.com/4279716508/status/2103294237525848323)
- Complaint, 2026-09-23, r/ClaudeCode (Reddit): “something similar can happen with code reviews, too. after enough cycles, a reviewer agent (whether that's copilot running remotely on github pull requests or a claude agent) will start to find problems like: _if a user submits an upload at 2:46am on the first tuesday in a calendar month with two full moons during the year 2246, then this list will contain only one entry, but the internal logs unconditionally use the plural "entries"._ then clau” [source](https://www.reddit.com/r/ClaudeCode/comments/1wohqnu/endless_slicing/pbnen3c/)
- Complaint, 2026-09-18, r/ClaudeCode (Reddit): “i have tried claude code, coderabbit and gh copilot in the past for code review but they are always doing the review "in their own way". also, most of them don't show me how they review my code and instead just gives the comment once it's done reviewing. having worked as an engineer for years, i have accumulated some rules that i always follow when i do pr (especially while vibecoding).what i did instead is to create my own agentic persona that k” [source](https://www.reddit.com/r/ClaudeCode/comments/1wjblf7/dont_pay_other_services_for_agentic_code_review/)

### OpenCode

- Praise, 2026-09-22, r/opencode (Reddit): “i am using it to review ds 4.1 flash code and the results are quite good as a reviewer.” [source](https://www.reddit.com/r/opencode/comments/1wn191t/mimo_26_flash_is_a_beast_and_seems_to_use_very/pbbmn1v/)
- Praise, 2026-09-20, r/opencode (Reddit): “i definitely agree with you. i’m having the total opposite experience with 1.3 contributor. i think it’s amazing. though i depend totally on my project/global level skills, pre-commit hooks, github actions and my automated reviewer” [source](https://www.reddit.com/r/opencode/comments/1wh9l72/i_lost_faith_in_artificial_analysis_muse_spark_13/pazd22d/)
- Praise, 2026-09-20, @opencode (X): “@aapakari got amazing code review from spark under @opencode” [source](https://twitter.com/260941463/status/2101816702992335157)
- Complaint, 2026-09-18, @opencode (X): “@btsouth @commandcodeai @ollama @opencode it's my default too, but at some point, you need at least sol to review the work” [source](https://twitter.com/801763720544317440/status/2100776067216662698)
- Complaint, 2026-09-14, @opencode (X): “using @opencode for the day in my app to make the integration smooth and i really wish it had codex's auto review sandbox mode, if not as an implementation at least as a contract so i could wire the app to it and provide the implementation through a plugin” [source](https://twitter.com/1931997694286770176/status/2099408425448780037)
- Complaint, 2026-09-11, r/opencode (Reddit): “do you have any skills that miss fire? it was the case for me. i had some code review, change review, and other review skills that were constantly miss-firing, turning every simple coding task into a review ceremony” [source](https://www.reddit.com/r/opencode/comments/1wd7aka/anyone_else_spending_10x_more_time_on_reviewdebug/p93tdpe/)

### Google Antigravity

- Praise, 2026-09-13, r/google_antigravity (Reddit): “completely agree. they can check each other which is awesome.” [source](https://www.reddit.com/r/google_antigravity/comments/1waream/antigravity_is_incredible_goodbye_claude/p9gfoxi/)
- Praise, 2026-09-09, r/google_antigravity (Reddit): “claude sometimes finds bugs in antigravity code, and vice versa. think the best setup is not to rely on a single agent but use two.” [source](https://www.reddit.com/r/google_antigravity/comments/1waream/antigravity_is_incredible_goodbye_claude/p8r3r84/)
- Praise, 2026-09-07, r/google_antigravity (Reddit): “also i have found like it automatically hunt thosse problems which i have not mentioned sometimes. thats really cool tbh” [source](https://www.reddit.com/r/google_antigravity/comments/1w9vwgv/tbh_flash_38_is_really_good_when_it_comes_in/p8djqld/)
- Complaint, 2026-09-24, r/google_antigravity (Reddit): “sarcasm is on that take, they have a fast and decent model but dont use it for reviewing commands.” [source](https://www.reddit.com/r/google_antigravity/comments/1w69qmt/finally_someone_popular_bringing_attention_to/pbq07cn/)
- Complaint, 2026-09-23, r/google_antigravity (Reddit): “that's sad. i hoped that boost could detect cut corners at verification phase. usually i need at least 1 round of a review and fixes because of that” [source](https://www.reddit.com/r/google_antigravity/comments/1wo8qye/what_is_your_experience_with_boost_deep_reasoning/pbltash/)
- Complaint, 2026-09-22, r/codex (Reddit): “when i use claude code, i have codex and gemini (antigravity) do the audit and when i use codex, i audit with the other two. gemini sometimes catches bugs but mostly, i have to say codex is the best auditor - always comes up with something useful. i think it may be in your case, codex is not doing well debugging because it’s the same model and it’s debugging itself. but in my case, as of the moment, codex is the best performer.” [source](https://www.reddit.com/r/codex/comments/1wnfc94/is_gemini_38_flash_actually_outperforming_gpt_56/pbevcm7/)

### Devin

- Praise, 2026-09-25, @cognition (X): “@cognition i was a hater, but i'm now using deepwiki and devin reviews on ci a lot. so congrats.” [source](https://twitter.com/1458111452397674503/status/2103515665168470122)
- Praise, 2026-09-23, @cognition (X): “ok @cognition devin is the code reviewer you want checking everything. and great value.” [source](https://twitter.com/1440778796727091206/status/2102831105569378664)
- Praise, 2026-09-23, @cognition (X): “@lordofafew @cognition devin reviewing code is a strong vote of confidence” [source](https://twitter.com/1949872909872254977/status/2102832167717834927)
- Complaint, 2026-09-23, @cognition (X): “@dewyashtwts @supercodeai @devinai @cognition self-review is the part i'd push back on. an agent grading its own session just confirms its own blind spots. fresh reviewer, zero access to that chat, diff only, that's the only way i trust it.” [source](https://twitter.com/1415688330/status/2102647905425465627)
- Complaint, 2026-09-17, @cognition (X): “@cognition a repo-wide scan is useful only if the evidence threshold is higher than “open a pr.” the ui can find dead code across acme/webapp; the hard part is proving the deletion is safe without turning every uncertain finding into review load.” [source](https://twitter.com/2046212002859610112/status/2100635161058525285)

### Pi

- Praise, 2026-09-25, r/PiCodingAgent (Reddit): “yes in my experience, its a work horse on medium and also good at reviews, catching multiple bugs in both astra and sols work” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wpcis6/opus_55_vs_gpt6_sol_luna_in_piagent_results_on_my/pbze4yj/)
- Praise, 2026-09-22, r/PiCodingAgent (Reddit): “for me pi-lens has been a game changer in adding context-level confidence to the code written by pi. it lets me use smaller models like deepseek-4.1-flash without worrying about whether it writes clean code. sometimes i just tell it “fix the lens warnings” and it’ll do it. pretty amazing tool for fighting ai slop” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wml882/what_extensions_do_you_think_are_essential_and/pbaxf03/)
- Praise, 2026-09-22, r/PiCodingAgent (Reddit): “switched my code-review flow from using lang-graph to using pi through the sdk. using astra from my codex sub. worked like a charm, and the review quality is even better. how do i add pi extensions into here? i would like to add subagents/goal functionality” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wm2g7w/piagentpythonsdk_use_your_installed_pi_coding/pbbv0b3/)

### Warp

- Praise, 2026-09-18, @warpdotdev (X): “@warpdotdev grading past agent sessions on efficiency is how you catch the first-run mess before it becomes the factory default” [source](https://twitter.com/1462653589617360896/status/2100947643077673054)
- Praise, 2026-09-08, @warpdotdev (X): “perfect example of llm-as-judge in the runtime path from @warpdotdev. instead of grading the coding agent’s code on a rubric (boring), it evals the application and blocks the flow unless it passes. - implementation agent codes a feature - verification agent evals the live app against the spec using a computer-use model - sends it back around the loop if needed astra and its off-the-charts computer-use capability will make this pattern more common” [source](https://twitter.com/65392279/status/2097386394347807178)
- Praise, 2026-09-08, @warpdotdev (X): “@josharosen @warpdotdev what makes this work: the verifier gets the spec, not the implementation's summary of itself. the run that wrote the feature is its worst judge. one addition: the verifier prints a denominator. 'passed' is silence; '7 of 9 spec items pass' is evidence.” [source](https://twitter.com/97718205/status/2097459317465293144)
- Complaint, 2026-09-08, @warpdotdev (X): “@josharosen @warpdotdev now it’s just a matter of ensuring that the agent knows what to evaluate and doesn’t just give itself a pat on the back 😅” [source](https://twitter.com/2028622577690947584/status/2097458250098884711)

### Cline

- Praise, 2026-09-15, @cline (X): “@howdevelop @cline scheduled reviews become valuable when they test the system's promises against the implementation. catching the corrupted-file overwrite risk shows a good task contract: compare docs, inspect failure paths, and return evidence with the finding.” [source](https://twitter.com/1764321378507931648/status/2099742319184691245)
- Praise, 2026-09-14, @cline (X): “i got early access to the new @cline open source desktop app! i wanted to see how it would fit into my everyday development workflow, so i tested it with the free glm-5.3 flash model across coding, scheduled reviews, and conversation handoffs. for the coding task, i asked cline to build a small node.js task-list cli with persistence and tests. after resolving an initial environment issue and fixing how the cli handled corrupted json, it successfu” [source](https://twitter.com/1071875988122951682/status/2099545544532144129)

### Factory

- Praise, 2026-09-24, @FactoryAI (X): “@factoryai @fireworksai_hq excellent that legacy-bench measures more than just whether the code compiles. in payroll, erp, and closures, a plausible but incorrect output can alter withholdings or reconciliations. evaluating edge cases and traceable evidence, with final human review, is key.” [source](https://twitter.com/1569177389959192578/status/2103233553891025320)
- Praise, 2026-09-02, G2 (G2): “q: what problems is the product solving and how is that benefiting you? a: having used factory ai in our engineering workflows, the standout feature for me is its autonomous agents, which they call droids. the biggest problem it solves for us is developer fatigue from multi-file refactoring, ongoing maintenance, and pull requests. most coding ai tools just sit inside your code editor and offer line-by-line autocomplete, which still leaves the man” [source](https://www.g2.com/products/factory-ai/reviews/factory-ai-review-13397954)

### Amp

- Praise, 2026-09-18, @AmpCode (X): “@adamwathan wrote an @ampcode plugin that checks if an llm’s code output follows the spec i give it. it checks against specific verification criteria and claims. it’s already caught some stuff for me” [source](https://twitter.com/33135576/status/2101006859314561178)
- Praise, 2026-09-02, @AmpCode (X): “i simply don't review every diff like that anymore. i rely on deep planning ahead of time and a multi-panel agent review against various criteria like security performance and adherence to the specs that i provide. when an implementation is ready, i probe it with pointed questions about how something works until i'm satisfied.” [source](https://twitter.com/33135576/status/2095241802362273842)

### Grok Build

- Praise, 2026-09-22, r/ClaudeCode (Reddit): “i’m sure there’s a more elegant way, i had astra leading, calling fable but it works fine in the reverse also. i created a ‘delegate’ skill and prompt the orchestrator agent to use the delegate skill to bring in whatever model(s) i specify. using claude -p when delegating to a claude model. i actually have an antigravity sub, a grok super heavy sub (which gives me quota via grok build and cursor ultra), and the new $50 muse sub. i always use cla” [source](https://www.reddit.com/r/ClaudeCode/comments/1wneic4/saw_this_today/pbgume6/)
- Praise, 2026-09-11, r/ClaudeCode (Reddit): “i use fable as my orchestrator, sonnet recon and headless grok build. will shift to sonnet/opus build once the heavy grok discount runs out. it is slower, a lot slower but token efficient and all the models add a layer of checking the others work. less baby sitting and cleaner code.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wd9eio/fable_as_orchestrator_and_opussonnet_as_executers/p9574gb/)

### Augment Code

- Praise, 2026-09-23, r/ExperiencedDevs (Reddit): “i would definitely start getting comfortable with it, you don't need to let it be an agent and do everything for you. it can be fun to figure out where that boundary is. i work in a small team that owns and maintains several software systems, from vendor based to integration layers and some full stack software with both internal and customer users, so knowing everything about everything is effectively impossible. another case is i have it integra” [source](https://www.reddit.com/r/ExperiencedDevs/comments/1wo4d8p/job_requiresuses_very_little_ai_sinking_ship_or/pblbka0/)
- Praise, 2026-09-10, @augmentcode (X): “line-by-line code review will soon disappear. the future is a risk-gated handoff between humans and agents, where people get pulled in only for the reviews that need judgment. @augmentcode published a schematic of how that works and it's a stellar blueprint for this new architecture. 𝐑𝐢𝐬𝐤 𝐫𝐨𝐮𝐭𝐢𝐧𝐠 𝐛𝐞𝐟𝐨𝐫𝐞 𝐭𝐡𝐞 𝐪𝐮𝐞𝐮𝐞: every pr gets classified first. docs and config auto-approve with a written justification. everything else gets tagged with the dime” [source](https://twitter.com/771267202762670081/status/2098097592001548319)
