# Breaks existing code or reintroduces bugs (`work.regressions_introduced`)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/criterion/work.regressions_introduced

Area: [Doing the work](https://feedbackbench.com/criteria/work.md)

**Definition.** Edits break previously working code, reintroduce fixed bugs, or cycle between introducing and fixing bugs.

**Boundary.** Not this: see [Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md) for repetition without code damage. Not this: see [Games checks instead of fixing the problem](https://feedbackbench.com/criteria/work.reward_hacking.md) for deliberately gaming checks.

Rated author-weeks, all agents: 577. Complaint share: 91%.

## The brief

Written by Claude Opus 5.5 from 60 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**Agents still break working code, and model updates make it worse.**

TL;DR:

- Complaints swamp praise for every agent with real volume. Regressions are a category-wide problem.
- OpenAI Codex draws the most regression complaints and rates worse than peers, despite vocal defenders.
- Model swaps drive much of the pain. Users report fixed bugs returning after new versions ship.

In plain terms: You ask for a small change and something unrelated stops working. You ask for a fix and an old bug returns. Many users treat git rollback as routine, and burn quota repairing the agent's own damage.

### How it breaks

- **New model versions bring old bugs back** ([Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md)). The sharpest complaints follow a model upgrade, when previously fixed bugs resurface and users roll back to an older model to regain stability.
  Posts tie regressions to specific releases more than to the tools themselves. One founder running a live platform says old bugs keep returning since a new model shipped, and that an older version feels more reliable. Others report a new release breaking a project in its first session. Users also ask vendors to let them pin the previous model. The pattern hits Claude Code, Cursor and OpenAI Codex users alike.
  Evidence:
  - Complaint, Cursor, @cursor_ai, 2026-09-05: “@elonmusk bring back @grok 4.5 to @cursor_ai !! it is way better than 4.6. grok 4.6 doesn’t listen and breaks everything thinking it’s helping. you can’t code with it. 4.5 was amazing. at least give us the option to keep using 4.5. grok 4.6 is not worth paying money for.” [source](https://twitter.com/863062011386527748/status/2096077784506568741)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-14: “hey all, i’ve built a booking platform with claude code (think airbnb style, different industry). it’s live with real hosts/customers and doing a couple grand a month. i could turn the traffic up pretty easily, but i’m honestly scared to scale it because of the bugs. everything was going really well initially, but since opus 5 came out it feels like things have gone backwards. old bugs we already fixed keep coming back, new bugs are appearing, and claude sometimes completely misunderstands what i’m asking it to do. i’ve tried fable 5.1 which seems better, but it absolutely eats my usage. i’ve now gone back to opus 4.8 and weirdly it seems more reliable, although i’m unsure about using an older model. anyone running an actual production platform with claude code have advice? how are you keeping things stable as you scale?” [source](https://www.reddit.com/r/ClaudeCode/comments/1wfy7fa/looking_for_bug_advice/)
  - Complaint, Claude Code, @ClaudeDevs, 2026-09-01: “@mrtacticalx can confirm it’s a turd 💩 1 fable 5.1 agent 4 opus5 subs same prompt over and over, fable 5 would hit 75-80% 5.1 chewed my 4hr window in 30mins, broke my project, completed zero work. its token usage is worse than 5 @claudedevs <strict_link>” [source](https://twitter.com/1525673840919351296/status/2094935532166066378)
  - Praise, OpenAI Codex, r/ClaudeCode, 2026-09-12: “i have to use verify and simplify in every pr and even double agent reviews. the number of bugs produced by claude opus is substantially higher than codex and it's not even close and this is just happening recently with opus 5. the situation was completely opposite until 4.6 and maybe even 4.8 was ok to an extent.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wdw386/i_have_never_had_to_use_verify_so_much_in_cc_its/p9asg4d/)

- **Each new feature breaks an old one** ([Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md)). Users describe projects decaying feature by feature, where every addition quietly breaks something that already worked.
  The complaint is cumulative damage, not one bad edit. A Google Antigravity user says every new feature breaks something else and the project has become a buggy mess. A Pi user on local models sees a small follow-up change break half of what worked. An OpenAI Codex user finds a working GUI left unusable after a request for extra error checks. The top request here is blunt: preserve working code and manual edits.
  Evidence:
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-18: “i got antigravity because it’s basically the only frontier model that lets me code a lot these days without spending a fortune, but i’ve noticed that every time i add new features, something else breaks. it’s been like this over and over for the past month, and my project has turned into a buggy mess. i genuinely don’t know if it’s me or the model at this point. is anyone else dealing with this?” [source](https://www.reddit.com/r/google_antigravity/comments/1wk1p0l/feeling_like_antigravity_is_slowly_destroying_my/)
  - Complaint, Pi, r/PiCodingAgent, 2026-09-22: “i have an asus ascent gx10 with 128gb vram hosting models via ollama. yet, i'm really struggling to get much going with opencode or pi. even single python files or small html projects are a struggle. i am currently using one of the qwen coder models with pi. sometimes it will get a request mostly right. then when i ask for a small change, it breaks half of what was working before. i'm coming from 2 years on cursor, so maybe my expectations are too high. i would truly love to be able to rely on open source agentic coding with local llm, but my mileage so far has been underwhelming.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wnmsho/i_want_to_believe_in_local_llms_for_coding_but/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-14: “as an example, if i upload a working gui and want some extra error checks, i am surprised that the first iteration doesnt leave the gui in a workable state after its done.” [source](https://www.reddit.com/r/codex/comments/1wg5ydh/what_am_i_doing_wrong/p9syv1r/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-17: “you have to have the agents think ahead or they will always introduce new problems.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wi5gcz/ive_been_doing_an_experiment_with_a_project_where/paaxayi/)

- **Fix-break cycles that eat quota** ([Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md)). Agents fix one bug, introduce another, then repeat, and users pay for every lap in usage.
  Users frame this as a billing problem as much as a quality one. A Cursor user says fixing the same bot-introduced UI bugs again and again is not a workflow, and asks that visual changes require proof before the agent declares done. A Devin post mocks a harness where half the work produces bugs and the other half fixes them. Experienced users advise hard-capping rounds and rolling back the moment a pass looks worse.
  Evidence:
  - Complaint, Cursor, @cursor_ai, 2026-09-05: “burning ai usage fixing the same bot-introduced ui bugs again and again isn’t a workflow. visual changes should require build + on-device proof before “done.” this needs to be a product-level reliability issue, not user babysitting. @bot @cursor_ai” [source](https://twitter.com/1630452238593437696/status/2096139574087196770)
  - Complaint, Devin, @DevinAI, 2026-09-23: “@hraness @devinai @zeddotdev so this is what a 50% produce the bugs, the other 50% fixes on repeat harness looks like 🔥” [source](https://twitter.com/1099277230260264962/status/2102715926315761838)
  - Complaint, Claude Code, r/ClaudeCode, 2026-08-31: “i understand and feel the same way. but i’ll say, i didn’t like gpt before i tried pro, and that’s didn’t really change. but also it is fully able to do what claude does just in different quality. i was getting super annoyed by claude using all of its usage and claiming to fix bugs when it just wasn’t fixing the bugs and was creating new bugs, and with gpt i kind of had the same experience but worse in a couple areas. havjng both is a way to pretty much double usage without going all the way to 5x (veiw it as the $40 plan) if there’s no free trial and you’re really getting stopped by usage maybe you can try both for a month? maybe you can let claude expire, try gpt plus, and if you don’t like it more try it with both, then you shouldn’t have usage issues. although i will repeat that in my experience gpt $20 plan is like clauses $20 plan but worse. but maybe that’s my workload.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w3diti/gpt_vs_claude/p6z7m44/)
  - Complaint, Cursor, r/cursor, 2026-09-23: “i'd stop using auto for anything that can touch ci. silent routing means you don't notice when it drops you onto something shallow mid-loop. pin the model and effort yourself, and hard-cap the rounds. if the next pass of test tweaks looks worse than the last one, the patches are injecting bugs, so stop and roll back instead of letting it spiral.” [source](https://www.reddit.com/r/cursor/comments/1wno9df/something_has_changed/pbigzbs/)

- **Collateral damage outside the task** ([Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md)). Edits spill past the requested change and break tooling, history or schema that the user never asked the agent to touch.
  Posts describe damage to the surrounding environment, not just the target file. Google Antigravity users report its edits breaking linting, live debugging and the git index. An OpenCode user says a feature landed at the cost of rewriting existing migrations. An OpenAI Codex user says a simple update turned into a detour that broke session history and auth before the original job got done.
  Evidence:
  - Complaint, Google Antigravity, r/GoogleAntigravityIDE, 2026-09-21: “when the extension proposes changes in the code, it breaks the linting and my live-debug session because both old and new codes are in the codebase. i had to switch back to the ide.” [source](https://www.reddit.com/r/GoogleAntigravityIDE/comments/1vyskmd/recently_they_announced_the_antigravity_ide/pb8kdny/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-10: “same here, vs code + antigravity, its touchs code and breaks git index” [source](https://www.reddit.com/r/google_antigravity/comments/1w996ck/bug_report_gitindex_corruption_race_condition_on/p912g7i/)
  - Complaint, OpenCode, r/opencode, 2026-09-03: “i started a working on a feature with shopify mcp without inbetween context switches. it implemented the feature at the cost of breaking migrations. migrations should be forward and not modify the existing ones. if it lacks basic common sense, then it's bad imo that too at xhigh effort. and ignored the frontend changes and said it applied it but didn't. glm's flash with default thinking & deepseek flash model with high thinking. no context switch mid conversation. all worked on same branch forks on separate worktrees. the only situation where something similar happened was with openai's luna model at medium. where it repeatedly does the same.” [source](https://www.reddit.com/r/opencode/comments/1w5tuwx/so_whats_the_early_consensus_about_the_new_muse/p7jhmpn/)
  - Complaint, OpenAI Codex, X search: OpenAI Codex, Codex CLI, Codex app, 2026-09-23: “i asked it to update codex cli + openclaw and switch the bot to gpt 6 sol xhigh. instead it went down a rabbit hole of plugin/runtime migrations, broke session history/auth, rolled things back, and still didn’t update openclaw. i had to keep steering it and suggest the openclaw update again before it finally worked. wasted ~1 hour on what should’ve been a 5 min task for 5.6. happy to share the full story on dm and hope they fix these issues in the future, but i think this is gpt 6 terra, just renamed to sol.” [source](https://twitter.com/507855158/status/2102906703461150877)

### Who stands out

- **OpenAI Codex (weaker)**. The largest pile of regression complaints, with users describing broken code, failed routine tasks and refunds, even as a minority swear it never breaks anything.
  Complaints center on specific model tiers inside Codex. Users say some tiers break routine weekly tasks or apply changes in the wrong place, and one says the damage was bad enough to win a partial refund. Defenders exist and are loud: several say Codex verifies before delivering and produces hardly any regressions compared with Claude. Opinion splits by model tier more than by product.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-19: “got a 110$ refund out of 200$... it broke so much of my code/sql - its actually unbelievable - i probably won't use/trust open ai for many years.” [source](https://www.reddit.com/r/codex/comments/1wk4tdp/im_exhausted/paodbo7/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-05: “theres only one usable model in codex, sol. astra burns 8x more quota than sol, terra is practically unusable (stops in the middle of a task, or sometimes marks a ticket as complete without even starting it), and luna is too slow to be useful and messes up big time (applied the schema changes to an entirely different database of a completely different project). i used codex for a month after 6 months of cc and im going back as soon as my sub expires.” [source](https://www.reddit.com/r/codex/comments/1w86ix6/is_codex_pro_5x_worth_it_coming_from_claude_max_5x/p80iwfv/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-13: “android development. a year ago it was bad, you needed to be a programmer to work along side it. now even as a vipe coder you could probably make stuff. even now, chatgpt sol is making a lot of mistakes. i spend more time telling it to fix than actually coming up with new stuff. opus 5 is not as bad. it occasionally make mistakes but it doesn't break stuff. astra is much better than sol in term of breaking stuff. i haven't tried fable but i am 100% sure it would make less mistakes than opus. and astra on extra high is much much better than sol. they are almost on completely different level. note: i am working on a mihon fork, which is a massive app with a lot of layers on every front. it should be different on simpler stuff.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wes79a/ive_been_vibecoding_seriously_for_about_a_year/p9gel0m/)
  - Praise, OpenAI Codex, r/codex, 2026-09-10: “it's only cheaper in the sense that it makes less errors, and hardly any regression. so far that has been my experience.” [source](https://www.reddit.com/r/codex/comments/1w6k3w1/gpt6_astra_credit_usage_i_didnt_expect_this/p8xgn9o/)

- **Claude Code (mixed)**. Users split sharply, with some reporting near-bug-free shipping and fewer new bugs than rivals, and others blaming a new model for returning bugs.
  Praise posts say recent Claude Code produces fewer new bugs and lasts longer than Codex on the same work. Complaint posts say a newer model reintroduced fixed bugs on a production app and broke projects outright. Two of the top requests come from Claude Code users asking Anthropic to fix regressions in recent releases.
  Evidence:
  - Praise, Claude Code, r/ClaudeCode, 2026-09-14: “i strongly disagree here, mate, because i've been using this product myself and have also shipped it to a bunch of my people, and i've never had any feature break or anything. the tool was shipped without any bugs. even though there is a minor bug, it can be easily fixed with a single prompt. back in 2024 and 2025, claude wasn't able to do this, but now ai coding is almost negligible with bugs.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wgcfst/cost_of_building_with_ai_vs_human_teams/p9tt17r/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-26: “i have both x20 claude/codex i just did a test on 10 different repo parallel tasks. for a simple task. astra high to xhigh. codex burned around 15% of the weekly. too afraid to use sol 6 after so many mistakes since its release. same task on opus 5.5 high/xhigh, it burned 4% and did it a lot faster than astra. hands down claude is the winner right now. also on my usual workflow codex can last 1-2 days while being token costs efficient. claude is 3-4 days, most results were a lot better on claude, less error and less new bugs.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqjb89/i_ran_out_of_codex_on_chatgpt_pro_with_4_days/pc4ygoo/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-14: “hey all, i’ve built a booking platform with claude code (think airbnb style, different industry). it’s live with real hosts/customers and doing a couple grand a month. i could turn the traffic up pretty easily, but i’m honestly scared to scale it because of the bugs. everything was going really well initially, but since opus 5 came out it feels like things have gone backwards. old bugs we already fixed keep coming back, new bugs are appearing, and claude sometimes completely misunderstands what i’m asking it to do. i’ve tried fable 5.1 which seems better, but it absolutely eats my usage. i’ve now gone back to opus 4.8 and weirdly it seems more reliable, although i’m unsure about using an older model. anyone running an actual production platform with claude code have advice? how are you keeping things stable as you scale?” [source](https://www.reddit.com/r/ClaudeCode/comments/1wfy7fa/looking_for_bug_advice/)
  - Complaint, Claude Code, @ClaudeDevs, 2026-09-01: “@mrtacticalx can confirm it’s a turd 💩 1 fable 5.1 agent 4 opus5 subs same prompt over and over, fable 5 would hit 75-80% 5.1 chewed my 4hr window in 30mins, broke my project, completed zero work. its token usage is worse than 5 @claudedevs <strict_link>” [source](https://twitter.com/1525673840919351296/status/2094935532166066378)

- **Cursor (mixed)**. Cursor earns the most praise here, but mostly for tooling that catches regressions, not for writing fewer of them.
  The praise cluster responds to a feature that catches regressions before users see them. Users call the regression-catching system the real moat. Complaints still describe the agent reintroducing the same UI bugs and a model update that breaks everything it touches. The product gets credit for the safety net while its agents still need one.
  Evidence:
  - Praise, Cursor, @cursor_ai, 2026-09-25: “@cursor_ai catching regressions in production automatically before users report them is a massive gain for solo dev speed.” [source](https://twitter.com/2075291394189541376/status/2103520875211341873)
  - Praise, Cursor, @cursor_ai, 2026-09-24: “@cursor_ai catching regressions before users see them is such a needed feature, glad to see this!” [source](https://twitter.com/1730179397750030336/status/2103116762460143835)
  - Praise, Cursor, @cursor_ai, 2026-09-25: “@cursor_ai harness work is the unsexy moat. models rotate. the system that catches regressions is what keeps teams shipping.” [source](https://twitter.com/1606668181137166337/status/2103388034851098969)
  - Complaint, Cursor, @cursor_ai, 2026-09-05: “burning ai usage fixing the same bot-introduced ui bugs again and again isn’t a workflow. visual changes should require build + on-device proof before “done.” this needs to be a product-level reliability issue, not user babysitting. @bot @cursor_ai” [source](https://twitter.com/1630452238593437696/status/2096139574087196770)

- **Cline (stronger)**. Every Cline post on this criterion is praise, driven by a harness change that users say sharply cut agent mistakes without swapping models.
  Users highlight that Cline kept the same models and changed only the harness, and mistake rates dropped steeply. Commenters single out the instrumentation as the detail most developers miss. The sample is small and comes from one announcement thread, so treat this as a signal, not a verdict.
  Evidence:
  - Praise, Cline, @cline, 2026-09-03: “@cline quiet harness swaps are the best kind of ship. we did one last month and nobody pinged us either. that 6.34% → 0.62% mistake drop is the changelog line i'd actually lead with.” [source](https://twitter.com/1888453273679740928/status/2095321474156388649)
  - Praise, Cline, @cline, 2026-09-04: “@cline interesting report. you kept the models and only changed the harness, and mistake_limit_reached dropped from 6.34% to 0.62%. wow! those are the little details a lot of devs are not aware of. smart to instrument the legacy harness first; otherwise, there is no baseline for that.” [source](https://twitter.com/3448284313/status/2095904557394018581)

### Fine print

- Many posts blame a specific model rather than the agent, so scores mix harness quality with model choice.
- Cline, Devin, Pi, Kiro and others have too few posts to rank with confidence here.
- Cursor praise clusters around one feature announcement thread, which may overstate sentiment.

## Top requests

What users ask to add or change, most asked first. 16 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Author-weeks | Posts | Agents (author-weeks) |
|---|---|---|---|---|
| 1 | Preserve working code and manual edits | 4 | 4 | Claude Code 2, Google Antigravity 1, OpenCode 1 |
| 2 | Fix regressions in recent releases | 3 | 3 | Claude Code 2, Cursor 1 |
| 3 | Stop reintroducing previously fixed bugs | 2 | 2 | Google Antigravity 1, OpenAI Codex 1 |
| 4 | Stronger pass/fail gates against regressions | 2 | 2 | Google Antigravity 1, Claude Code 1 |

### 1. Preserve working code and manual edits

- Google Antigravity, 2026-09-09, r/google_antigravity (Reddit): “do not remove a change just because it wasn't made by you. the agent almost always discards my manual code changes because it thinks they were excidently made by it lmao” [source](https://www.reddit.com/r/google_antigravity/comments/1wbcu5v/what_is_the_most_repeated_instruction_that_you/p8pcx7f/)
- Claude Code, 2026-08-31, r/ClaudeCode (Reddit): “the worst is if you have claude fix bugs in the codebase; it'll destroy any existing documentation you have in the codebase with it's own comments about the bug its fixing. the only way ive been able to fix this is to strip the comments and have an agent with no context of the bug add docs to the entire class after undertstanding the codebase.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w3auw5/ive_removed_all_inline_documentation_from_my/p6z9n6d/)
- OpenCode, 2026-08-31, r/opencodeCLI (Reddit): “muse fucks up my code base on the regular with unnecessary edits and rewrites. the fucker constantly rewrites my convex auth to use email provider and manual fetch calls to the resend api, when i have already perfectly setup the resend provider with the resend sdk. anytime it encounters a error that is remotely auth related, it just goes welp time to rewrite the email setup again despite explicit warnings not to.” [source](https://www.reddit.com/r/opencodeCLI/comments/1w2vu9m/top_10_best_opencode_go_models_by_capabilities/p6xsb4r/)

### 2. Fix regressions in recent releases

- Claude Code, 2026-09-10, @ClaudeDevs (X): “@claudedevs how about fixing the bugs before talking more crap on. this is really cool though.” [source](https://twitter.com/1973879601324584960/status/2098173616403702131)
- Cursor, 2026-09-04, r/cursor (Reddit): “true hopefully they fix that or atleast hit us with new composer .” [source](https://www.reddit.com/r/cursor/comments/1w6hn2i/i_guess_all_good_things_have_to_end/p7u7doq/)
- Claude Code, 2026-09-04, r/ClaudeCode (Reddit): “i've been experiencing serious reliability issues with claude opus 5, and they're significant enough that i feel compelled to speak up. the model consistently loses track of where it has edited files, leaving me to hunt down changes it made and then forgot about. it fails to clean up after itself — scratch files and temporary artifacts are left scattered across my project instead of being removed when they're no longer needed. worse, it creates r” [source](https://www.reddit.com/r/ClaudeCode/comments/1w6u8ka/the_performance_of_claude_opus_5_has_gotten/)

### 3. Stop reintroducing previously fixed bugs

- OpenAI Codex, 2026-09-15, r/codex (Reddit): “moved off claude partly because weekend work and the 5-hour window suck. sol high to plan, terra high to execute. docs supposedly migrated, but "change a and b" only does a, then fixing b reverts a. last time terra flipped a back while you were fixing b, how long did that ping-pong run, and what did you end up shipping vs leaving broken?” [source](https://www.reddit.com/r/codex/comments/1wgyi2b/i_just_migrated_from_claude_what_am_i_doing_wrong/p9yboov/)
- Google Antigravity, 2026-09-03, r/google_antigravity (Reddit): “the extension in vscode is unusable. file state is messy, impossible to know if the changes made by the agent have been applied or if you have to accept the changes first. i can see bugs being introduced and the agent remaking the same modifications. i wish i could use it but it's simply impossible.” [source](https://www.reddit.com/r/google_antigravity/comments/1w637hk/antigravity_20_release_v2120/p7mnu82/)

### 4. Stronger pass/fail gates against regressions

- Google Antigravity, 2026-09-21, r/GoogleAntigravityIDE (Reddit): “when the extension proposes changes in the code, it breaks the linting and my live-debug session because both old and new codes are in the codebase. i had to switch back to the ide.” [source](https://www.reddit.com/r/GoogleAntigravityIDE/comments/1vyskmd/recently_they_announced_the_antigravity_ide/pb8kdny/)
- Claude Code, 2026-09-10, r/ClaudeCode (Reddit): “it’s going to sound weird but i’ve found that it reflects your energy back at you. if you’re kind, build it up, and treat it as a colleague that you’re working with to solve a problem then it genuinely seems to do better. i unironically like to mix in words of encouragement occasionally. it sounds dumb but it really seems to work. for specific bugs where it breaks something to fix something else, it needs better pass/fail gates to prevent regres” [source](https://www.reddit.com/r/ClaudeCode/comments/1wcfu1t/how_do_you_make_your_cc_not_keep_making_mistakes/p8zagfr/)

## Every agent

| Agent | Overall rank | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|---|
| [Cursor](https://feedbackbench.com/agents/cursor.md) | 4 | Typical | 0.510 | 0.482–0.538 | 50 | 11 | 39 |
| [Google Antigravity](https://feedbackbench.com/agents/antigravity.md) | =5 | Typical | 0.507 | 0.459–0.558 | 53 | 6 | 47 |
| [Claude Code](https://feedbackbench.com/agents/claude-code.md) | 1 | Typical | 0.505 | 0.446–0.556 | 153 | 12 | 141 |
| [OpenCode](https://feedbackbench.com/agents/opencode.md) | 3 | Typical | 0.487 | 0.457–0.524 | 31 | 2 | 29 |
| [OpenAI Codex](https://feedbackbench.com/agents/codex.md) | 2 | Worse than peers | 0.448 | 0.390–0.498 | 255 | 12 | 243 |
| [Devin](https://feedbackbench.com/agents/devin.md) | =5 | Too few posts | – | – | 10 | 1 | 9 |
| [Cline](https://feedbackbench.com/agents/cline.md) | =8 | Too few posts | – | – | 7 | 7 | 0 |
| [GitHub Copilot](https://feedbackbench.com/agents/copilot.md) | =8 | Too few posts | – | – | 6 | 0 | 6 |
| [Pi](https://feedbackbench.com/agents/pi.md) | 7 | Too few posts | – | – | 5 | 1 | 4 |
| [Kiro](https://feedbackbench.com/agents/kiro.md) | 13 | Too few posts | – | – | 5 | 0 | 5 |
| [Zed](https://feedbackbench.com/agents/zed.md) | =8 | Too few posts | – | – | 1 | 0 | 1 |
| [Amp](https://feedbackbench.com/agents/amp.md) | =11 | Too few posts | – | – | 1 | 0 | 1 |
| [Factory](https://feedbackbench.com/agents/factory.md) | =11 | Too few posts | – | – | 0 | 0 | 0 |
| [Conductor](https://feedbackbench.com/agents/conductor.md) | 14 | Too few posts | – | – | 0 | 0 | 0 |
| [Warp](https://feedbackbench.com/agents/warp.md) | 15 | Too few posts | – | – | 0 | 0 | 0 |
| [Grok Build](https://feedbackbench.com/agents/grok-build.md) | 16 | Too few posts | – | – | 0 | 0 | 0 |
| [Augment Code](https://feedbackbench.com/agents/augment.md) | 17 | Too few posts | – | – | 0 | 0 | 0 |

## Posts

Receipts rule: The 5 most recent praise and complaint posts per area (first 700 characters) and 3 per criterion (first 450 characters).

### Cursor

- Praise, 2026-09-25, @cursor_ai (X): “@shivanisis89840 @cursor_ai glad the stronger self-verification and 256k context in grok 4.7 helped cut agent refactoring errors in cursor. holding more of the codebase and checking results carefully makes multi-step changes far more reliable. thanks for sharing the details.” [source](https://twitter.com/1720665183188922368/status/2103328453269283207)
- Praise, 2026-09-25, @cursor_ai (X): “@cursor_ai harness work is the unsexy moat. models rotate. the system that catches regressions is what keeps teams shipping.” [source](https://twitter.com/1606668181137166337/status/2103388034851098969)
- Praise, 2026-09-25, @cursor_ai (X): “@cursor_ai the useful part is that deployment becomes an observable loop instead of a one-time handoff. catching regressions before users report them could make agent-written changes much easier to trust in production.” [source](https://twitter.com/2095173354613260290/status/2103448969980309874)
- Complaint, 2026-09-26, r/cursor (Reddit): “tired of cursor fixing 1 line and breaking 3 other features? we built an open-source 23-protocol agent constitution + mcp server to fix it. [removed]” [source](https://www.reddit.com/r/cursor/comments/1wqhwi5/tired_of_cursor_fixing_1_line_and_breaking_3/)
- Complaint, 2026-09-25, @cursor_ai (X): “@cursor_ai agent writes code quickly, the difficulty lies in blocking regressions before going live. this kind of automatic monitoring plan is quite appealing.” [source](https://twitter.com/2024314068967026688/status/2103400686377730060)
- Complaint, 2026-09-25, @cursor_ai (X): “truly insane how quickly i went from almost getting an annual @cursor_ai subscription to canceling outright despite a 50% "loyalty" offer. where was that when grok bot burned an extra $200 doing nothing but break stuff for 8 hours?” [source](https://twitter.com/1595034974255718400/status/2103494681258860846)

### Google Antigravity

- Praise, 2026-09-27, r/ClaudeCode (Reddit): “short answer no. gemini pro subscription with flash 3.8 in high mode inside of antigravity will provide more bang for the buck on same code quality. no constant rewriting if you use it directly inside of vs code. only inline def changes more bang for the buck.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrf0ld/is_claude_pro_actually_worth_20_just_for_one/pcci0wy/)
- Praise, 2026-09-17, r/google_antigravity (Reddit): “that's the same reason i am happy about antigravity and 3.8 on high now - because it stopped breaking stuff and started to understand context in which it makes differences. especially with custom skills it is now much more useful. it follows commands much better and checks for blast radius of it's actions, still not intelligent for big work, but with right skills, first time antigravity is really usefull. and now, especially with this much of us” [source](https://www.reddit.com/r/google_antigravity/comments/1wi8z7u/did_they_change_the_model_or_what/pabbxud/)
- Praise, 2026-09-14, r/google_antigravity (Reddit): “yep. it did all that until i got opus to write me a gemini.md to fix it. people keep telling me to stop posting ai slop, so if you want it, send me a message. i'll explain more. i have had a very good experience, building out a messageboard frontend for a compacted forum archive. i modified gemini.md a few times along the way, but the final iteration is doing about as well as opus 4.6, on far fewer tokens. (or far higher quota) most of my prompts” [source](https://www.reddit.com/r/google_antigravity/comments/1wft9p8/gemini_cherrypicks_easy_tasks_and_falsely_reports/p9q857k/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “hello! i've been using antigravity on a medium-sized unity game development project with a mcp, and lately i've been having three problems with the ai agents. i'm not sure whether this is a problem with antigravity itself (limitations and intelligence) or with the way i'm using it. 1. **loops** since version 3.8 was released, i've noticed that the agent gets stuck in loops quite often. a request to fix or add something simple can make it loop for” [source](https://www.reddit.com/r/google_antigravity/comments/1wr740g/antigravity_agents_getting_stuck_in_loops_and/)
- Complaint, 2026-09-27, r/GoogleAntigravityIDE (Reddit): “hello! i've been using antigravity on a medium-sized unity game development project with a mcp, and lately i've been having three problems with the ai agents. i'm not sure whether this is a problem with antigravity itself (limitations and intelligence) or with the way i'm using it. 1. **loops** since version 3.8 was released, i've noticed that the agent gets stuck in loops quite often. a request to fix or add something simple can make it loop for” [source](https://www.reddit.com/r/GoogleAntigravityIDE/comments/1wrf9ly/antigravity_agents_getting_stuck_in_loops_and/)
- Complaint, 2026-09-24, @antigravity (X): “@jointhebnc @jamesor @antigravity @googleaistudio @googlecloud bro i am trying to fiy so many bugs on my apps, it just creates more bugs. its not a good model to build complex apps and games.” [source](https://twitter.com/2001613318633693184/status/2103013688076861652)

### Claude Code

- Praise, 2026-09-27, r/ClaudeCode (Reddit): “this behavior is reduced by 80% after i install ponytail, but again my project is simple web app, i’m happy though” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrib98/do_you_feel_that_claude_code_unit_tests_are/pcf2db6/)
- Praise, 2026-09-26, r/ClaudeCode (Reddit): “i have both x20 claude/codex i just did a test on 10 different repo parallel tasks. for a simple task. astra high to xhigh. codex burned around 15% of the weekly. too afraid to use sol 6 after so many mistakes since its release. same task on opus 5.5 high/xhigh, it burned 4% and did it a lot faster than astra. hands down claude is the winner right now. also on my usual workflow codex can last 1-2 days while being token costs efficient. claude” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqjb89/i_ran_out_of_codex_on_chatgpt_pro_with_4_days/pc4ygoo/)
- Praise, 2026-09-26, r/ClaudeCode (Reddit): “tbh who cares about token to solution as long as it's within the subscription limit. what's important is human time to a _properly implemented_ solution and amount of bugs/regressions. i don't use superpowers to save on upfront tokens, i use it for ease of mind knowing most stuff there won't cause troubles the moment it hit production. maybe it is outdated and vanilla harness is exactly as good - but that's a totally different consideration.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wpczby/superpowers_skill_is_so_bad_now/pc50vpi/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “i had opus 5.5 make mistakes and break things right after release, so i'm not sure "broke a couple things" is enough evidence. llms can make mistakes regardless of quantization. what you should do is have fable replay the last 20 or so prompts from before the degradation occurred with the now possibly weaker model in the exact same spot in your commit history, and have it judge the results. then come back with it's evaluation. your analysis is ju” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr67vx/i_agree_with_dario_regulating_llms/pca1k7u/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “yes. i also find it over confident with its own decisions and doubling down, do a half ass job, then the problem i caught earlier came back biting its' ass.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqnjkp/be_careful_with_opus_55s_confidence/pcbaxvx/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “this is the sad truth. opus makes me wait for 5 hour window resets, but i would spend even more time with sonnet, just because after every task there was another task scheduled to fix mistakes of the previous one” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr5lzf/practically_speaking_what_tasks_fit_into/pcbqt41/)

### OpenCode

- Praise, 2026-09-27, r/opencode (Reddit): “this has helped me a lot to not break anything <strict_link> i only mark opencode” [source](https://www.reddit.com/r/opencode/comments/1wrbwg8/la_base_de_datos_de_opencode_paso_a_14gb_como_la/pcbcp9x/)
- Praise, 2026-09-08, r/opencode (Reddit): “i have not had any regression issues with spark 1.3 and for the last few days it has been writing a non-traditional compiler in an invented general purpose language and honestly doing really well. i'm not using the opencode version since it was hitting limits, etc. i just use the api in the muse cli -extremely cheap labor. got good 'ol sol reviewing and crackin' the whip, so that probably helps. i also have a project specific harness that i run” [source](https://www.reddit.com/r/opencode/comments/1waq3e5/muse_spark_13_free_is_ass/p8l31oq/)
- Complaint, 2026-09-27, @opencode (X): “@seridarivus13 @opencode @claude been there, done that. this rate seems to be higher in open code compared to claude” [source](https://twitter.com/1216432139610144768/status/2104149753927967057)
- Complaint, 2026-09-26, r/opencode (Reddit): “yup. it's good at spotting problems. but sometimes just doesn't fix them, or introduces new ones. i think that has something to do with the "context-bug" i mentioned. something doesn't seem right.” [source](https://www.reddit.com/r/opencode/comments/1wqi4a8/my_honest_opinion_about_spacebunny/pc493wb/)
- Complaint, 2026-09-26, r/opencodeCLI (Reddit): “"confirmed it's worthless to use with caution", the model itself has told me that they have put "deepseek v4 flash", and removed glm 4.6, what a shame.. bye bye cucumber.., it's fine sometimes but it gets out of control, it doesn't plan even less than glm4.6 xd, but glm 4.6 was a genius at programming, this substitute only has moments of clarity, fixes one and breaks 3..., it works for us, does not attend, does not focus, and does not follow inst” [source](https://www.reddit.com/r/opencodeCLI/comments/1t1yrs0/big_pickle_is_unusable_now/pc8ltrn/)

### OpenAI Codex

- Praise, 2026-09-27, r/codex (Reddit): “what i find is that opus can even break the code and leave you with **unusable** software, while astra will never break anything, and will verify that things work –in its own way– but that they work before delivering results.” [source](https://www.reddit.com/r/codex/comments/1wpveoe/astra_vs_opus_55_my_impressions_on_hard_project/pcbtxos/)
- Praise, 2026-09-23, r/codex (Reddit): “i just find it less bug ridden when using codex, and astra tends to daydream less. may be my bias, i originally worked on work spaces, maybe its just more experience. i use work space only for documents and drafts for non code applications” [source](https://www.reddit.com/r/codex/comments/1wnnv36/plus_users_hows_your_experience_so_far/pbi47sn/)
- Praise, 2026-09-21, r/codex (Reddit): “same experience, also a hands-on cto, etc., etc. the key is multiple layers of tests and gold copy database backups (assuming your software interacts with a database) used to run before/after db state verification (end-to-end/integration testing.) all changes need to be guarded with tests that link back to the detailed pr notes or docs folder md implementation plans with the business rules the fix/feature does spelled out. essentially, big-org r” [source](https://www.reddit.com/r/codex/comments/1wlvf0j/state_of_agentic_coding/pb3jyac/)
- Complaint, 2026-09-27, r/codex (Reddit): “not astra-specific in my experience. i've had agents garble utf-8 on several setups, and the cause was usually the path the edit took rather than the model: the agent rewrites the file through a shell command (powershell redirection, a script running under a non-utf-8 locale) instead of its own edit tool. that's also why it's intermittent. it only happens in sessions where the agent picks that route. a global-prompt rule helped less than a check” [source](https://www.reddit.com/r/codex/comments/1wree2f/astra_medium_corrupts_encodings_sometimes/pcbwxdo/)
- Complaint, 2026-09-27, r/codex (Reddit): “6 sol is actively bad. not like "hey let's save some tokens and i'll just have to watch it more closely" bad. like "do not let this thing touch your code" bad. it needs to be yoinked out of the lineup before people fuck their shit up with it.” [source](https://www.reddit.com/r/codex/comments/1wrftcs/gpt6_sol_is_massive_downgrade/pccqtgj/)
- Complaint, 2026-09-27, r/codex (Reddit): “astra fucked my code base so hard it took all of my opus allotment just to fix it. just random, stupid bugs introduced into months old processes” [source](https://www.reddit.com/r/codex/comments/1wrizyf/months_of_throttled_codex_usage_then_openai/pcctnpx/)

### Devin

- Praise, 2026-09-12, @cognition (X): “swe-2 from @cognition just decomposed every file over 1,200 lines in the <strict_link> codebase. i ran agent swarms in 6 phases. many files as big as 7k lines of code were decomposed by 90%. all files passed verifications, security checks and had zero breaks. the best part is it costed me $0 because swe-2 is free on any devin paid plan till end of october 2026. a plan starts from $20 a month. i have also used their newly launched fusion model” [source](https://twitter.com/1878743506094895104/status/2098688776029757842)
- Praise, 2026-09-12, @cognition (X): “@_colton_harris @cognition decomposition of monolithic files happened successfully no vandalism.” [source](https://twitter.com/1878743506094895104/status/2098703668711490019)
- Complaint, 2026-09-26, @cognition (X): “@cognition a run rate annualizes a snapshot. how much of the code devin merged six months ago still sits in main, or did the engineers reviewing it quietly rewrite it?” [source](https://twitter.com/1603553957204627456/status/2103640350900552120)
- Complaint, 2026-09-23, @DevinAI (X): “@hraness @devinai @zeddotdev so this is what a 50% produce the bugs, the other 50% fixes on repeat harness looks like 🔥” [source](https://twitter.com/1099277230260264962/status/2102715926315761838)
- Complaint, 2026-09-23, @DevinAI (X): “@hraness @devinai @zeddotdev 99% of the time, this ai fixes bugs it has created, while creating 2x more bugs…” [source](https://twitter.com/83444822/status/2102889437093056804)

### Cline

- Praise, 2026-09-14, @cline (X): “@cline i surprised you guys vooked this time it seems non ai slop mot buggy another harness congrats and thanks” [source](https://twitter.com/2070816272967979008/status/2099550871411728781)
- Praise, 2026-09-04, @cline (X): “@cline interesting report. you kept the models and only changed the harness, and mistake_limit_reached dropped from 6.34% to 0.62%. wow! those are the little details a lot of devs are not aware of. smart to instrument the legacy harness first; otherwise, there is no baseline for that.” [source](https://twitter.com/3448284313/status/2095904557394018581)
- Praise, 2026-09-04, @cline (X): “@cline the 10x reduction in tasks hitting the mistake limit is a pretty impressive result, especially with such a careful rollout.” [source](https://twitter.com/1488761257092005892/status/2095910397467611387)

### GitHub Copilot

- Complaint, 2026-09-26, r/ExperiencedDevs (Reddit): “mate don’t. we had very similar. a new rewrite front and backend of a client web application for a small subset of what is now considered legacy. they did the entire front end with claude first, the backend claude as a copilot assistant. first iteration saw maybe 30-40 bugs raised, and who knows how many more post the push from testenv to preprod. it took like 3 months. and then the quarterly management stat report came in about unplanned work an” [source](https://www.reddit.com/r/ExperiencedDevs/comments/1wql3g2/interviewed_candidates_for_ai_engineer_roles_this/pc7cdzh/)
- Complaint, 2026-09-17, r/ClaudeCode (Reddit): “at least it doesn't mangle files and fuck up character encoding like the powershell commands any model writes when you run it in copilot...” [source](https://www.reddit.com/r/ClaudeCode/comments/1whvvp8/why_does_claude_write_python_scripts_to_change/pab3f3b/)
- Complaint, 2026-09-14, r/AI_Agents (Reddit): “five ai coding tools, five production incidents, and the filter i wish i had first i use vibecoding to build software for a living, and here is what i found after using different ai tools. before you pick one, you should know my filter: * it must integrate into existing workflows without forcing a rewrite of how we develop. * it must not silently introduce logic changes (safe defaults, test awareness). * it should respect existing architecture an” [source](https://www.reddit.com/r/AI_Agents/comments/1wgfrah/five_ai_coding_tools_five_production_incidents/)

### Pi

- Praise, 2026-09-12, @pidotdev (X): “@johan_vd_meer @thsottiaux yes, as per looking at <strict_link> it was clearly this. just switched to the @pidotdev harness, no more fuckups.” [source](https://twitter.com/15266830/status/2098887000426189273)
- Complaint, 2026-09-23, @pidotdev (X): “@pidotdev it keeps mangling text files, it is working through the planned phases though. so i think i'll let it be, if it starts struggling too much i'll switch it out for deepseek v4.1 and see if it runs smoother... working on this btw: <strict_link> you can see it in realtime” [source](https://twitter.com/432486005/status/2102622029694533984)
- Complaint, 2026-09-22, r/PiCodingAgent (Reddit): “i have an asus ascent gx10 with 128gb vram hosting models via ollama. yet, i'm really struggling to get much going with opencode or pi. even single python files or small html projects are a struggle. i am currently using one of the qwen coder models with pi. sometimes it will get a request mostly right. then when i ask for a small change, it breaks half of what was working before. i'm coming from 2 years on cursor, so maybe my expectations are to” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wnmsho/i_want_to_believe_in_local_llms_for_coding_but/)
- Complaint, 2026-09-20, @pidotdev (X): “@pidotdev @mitsuhiko spent four hours undoing what took it thirty seconds to break.” [source](https://twitter.com/1910204476973096960/status/2101684433061298632)

### Kiro

- Complaint, 2026-09-18, r/codex (Reddit): “i only had this when i was using kiro with claude. never with codex directly. when ai was using gpt 5.5, i sometimes trashed its work and restored from git.” [source](https://www.reddit.com/r/codex/comments/1wjyeps/astra_just_worked_on_a_feature_for_an_hour_then/pamjn49/)
- Complaint, 2026-09-16, r/kiroIDE (Reddit): “i had multiple sub agents launch and ultimately throw errors. i contacted support for a refund for something clearly wrong with their service at the time. they refused to credit 20$ in tokens. that’s the last time i’ve used kiro.” [source](https://www.reddit.com/r/kiroIDE/comments/1wgutoi/so_let_me_get_this_straight_the_kiro_team_goes/pa2qy26/)
- Complaint, 2026-09-15, r/kiroIDE (Reddit): “same here, one small task on a card in a specific page, sol used almost 700 credits, 1.5h.. entire page was broken at the end... and without working api in the page” [source](https://www.reddit.com/r/kiroIDE/comments/1wgutoi/so_let_me_get_this_straight_the_kiro_team_goes/p9xwn7s/)

### Zed

- Complaint, 2026-09-08, r/ZedEditor (Reddit): “it's been like that on every. single. app. those are the consequences of everyone using ai to code everything these days. i'm not ai hating here: i'm using it on my job too - and on my personal projects - but everyone is doing that while still trying to figure out a quality control process for this new generation, and basically nobody has yet. claude, cursor, hermes, zed - every single app i use that ships codes to users has been carrying some b” [source](https://www.reddit.com/r/ZedEditor/comments/1wagx6o/lately_so_many_subtle_annoying_bugs/p8kp0jv/)

### Amp

- Complaint, 2026-09-24, @AmpCode (X): “mainly when i throw in a new idea, hand it a design mockup, or ask it to refactor something big, it tends to make a mess. breaks things, ignores instructions. these are problems most harness solved earlier this year, but amp's harness hasn't caught up. also, no plan mode. for larger tasks the model still needs to plan before it acts. amp has oracle but it wasn't enough, i ended up writing my own planning skill to compensate. cc handles this nativ” [source](https://twitter.com/1592160489965948933/status/2103099981473382778)
