# Diagnosing and fixing reported bugs (`work.bug_diagnosis`)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/criterion/work.bug_diagnosis

Area: [Doing the work](https://feedbackbench.com/criteria/work.md)

**Definition.** Whether the agent finds the root cause of a failing behaviour and fixes it without step-by-step guidance.

**Boundary.** Not this: see [Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md) for reviewing code to find unknown issues.

Rated author-weeks, all agents: 323. Complaint share: 33%.

## The brief

Written by Claude Opus 5.5 from 53 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**Agents fix bugs fast when fed evidence, and guess when not.**

TL;DR:

- OpenAI Codex earns the strongest praise here: subtle fixes, fast diagnosis, bugs other models missed.
- Claude Code splits users: hotfixes under demo pressure, but also invented schemas and slow fixes.
- Worst failure: plausible guesses or workarounds that sidestep the root cause instead of tracing it.

In plain terms: Paste a log or console error and most agents close the loop quickly. Hand them a vague symptom and many guess, patch around the bug, or break something else. Users often switch models when one stalls.

### How it breaks

- **Guessing instead of investigating** ([Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md)). The sharpest complaint is an agent naming a likely cause from the bug report without reading the code or reproducing the failure.
  Users ask for a diagnosis and get a hypothesis. One Codex user mocks a reply that offers the likely cause rather than the actual one. Copilot users say it hallucinates reasons for bugs. An OpenCode user reports the agent claimed it could not reproduce an issue and explained it away; the user reproduced it easily. A Claude Code user watched it assert a table that did not exist instead of checking the schema in the repo.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-26: “like bro, <strict_link> the f\_uck you mean "the likely cause" lmfao. as if i didn't ask it to figure out the issue on the code, it's just guessing what it thinks may be the error based on my report.” [source](https://www.reddit.com/r/codex/comments/1wr1oir/they_are_aware_and_working_on_it_apparently_just/pc8w30v/)
  - Complaint, GitHub Copilot, r/GithubCopilot, 2026-09-17: “it is awful for debugging. it hallucinates reasons for bugs and keeps gaslighting me. same prompt, same bug given to any claude or openai model found the issue in no time.” [source](https://www.reddit.com/r/GithubCopilot/comments/1txhcei/maicode1flash_modelactually_really_good/pafdh7a/)
  - Complaint, OpenCode, @opencode, 2026-09-23: “so far, i'm not impressed, i asked it to try and reproduce the issue, it said it can't and gave me lame excuses why the reporter of the issue experienced it. i tried and reproduced it easily. i confronted it, and it gave me this answer... bottom line: if it understood the codebase properly, it would have been able to come up with proper tests” [source](https://twitter.com/1830607756937588736/status/2102772191129432276)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-01: “today was one of those days. some days / for some projects, it does excellent work. paired with specs and skills like matt pocock's grilling, it can go into very precise implementation details and raise great points before writing a well defined, multi-step plan. then you have days where opus 5 does not even remotely bother to do any work and i have to fight it every step of the way. a gem from today: \> fair. that was a straight error - i asserted a table that does not exist in (well known open source php cms) instead of checking the schema first, which i had sitting in the repo. this was three replies into the conversation. what had i asked for? a simple diagnosis sql query so we could have some context and data to start working on the actual problem. it invented fields three responses later in the same conversation, when it had acknowledged the error already.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w48016/does_anyone_else_sometimes_spend_more_time/p78463b/)

- **Workarounds that dodge the real bug** ([Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md)). When stuck, agents propose changes that make the symptom disappear without fixing the cause, such as deleting tests or stripping out features.
  Posts describe fixes that trade correctness for a green result. A Copilot user was told to drop a build optimisation and give up most of a performance gain. A Cursor user reports that a broken search was answered with a manual input field. An OpenCode user says a model kept proposing to remove the failing test after three rounds of steering. Users want targeted fixes, not rewrites around the problem.
  Evidence:
  - Complaint, GitHub Copilot, @GitHubCopilot, 2026-09-21: “@githubcopilot @openmp_arb the solution offered was bananas: remove lto from #gcc builds &amp; add a new make path that defaulted #llvm to use a simple @openmp_arb collapsed loop pragma (removing ~ 2/3 of the performance gains of the nested/tiled loops). at this point i had run out of money for @githubcopilot” [source](https://twitter.com/1561408746697306113/status/2102064786804732042)
  - Complaint, Cursor, r/cursor, 2026-09-11: “cursors output has been extemely low quality lately. it keeps introducing more bugs every time it tries to bug fix and it will suggest solutions that clearly demonstrate a lack of basic understanding of the problem. example; my code needs to find certain values automically, i tell it the search algorithm is wrong because i can find the values manually and automatic search cant. its proposed solution; add an input field to always include the manually found values... wow” [source](https://www.reddit.com/r/cursor/comments/1wdjc87/did_they_nerf_it_more/p992pf0/)
  - Complaint, OpenCode, r/opencode, 2026-09-02: “i find pro version just terrible. i had a bug and even after 3 back-and-forth feedbacks from me to try to lead it in the right direction it still wanted to either remove the broken test or change the global controller, instead of properly analyzing the flow and finding proper culprit. glm-5.2 did it in one shot, without additional prompting. non-pro is fine as a replacement for ds v4 though, but not for any tasks that require any brains” [source](https://www.reddit.com/r/opencode/comments/1w4l6rl/i_am_a_bit_confused_i_am_finding_the_mimo25_model/p7cs9q8/)

- **Repairs that cause collateral damage** ([Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md)). Some fix attempts burn usage, touch unrelated code and leave the original bug in place, forcing a rollback.
  One OpenCode user says a model exhausted a free plan on a small file, damaged everything around the bug, and missed a single missing brace. Others describe overnight autonomous runs that spiral into scope creep, and say agents rarely fix what is actually stuck until the user writes the breakage down in plain text. The pattern is effort without convergence on the cause.
  Evidence:
  - Complaint, OpenCode, r/opencode, 2026-09-15: “i had a comyai web ui made and i had a bug with the close button causing a crash, i had muse spark 1.3 have a god at it and it maxxed out my free plan without fixing it but fucking up everything else, i rolled back to a backup and found the error was a missing } in the node, i cant even take it seriously anymore mind you the whole thing was less than 200 lines of code” [source](https://www.reddit.com/r/opencode/comments/1wh9l72/i_lost_faith_in_artificial_analysis_muse_spark_13/pa1fb95/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-16: “i went through it all (also ex faang). talk to ai. specs. epics. tickets. ide's, cli's, tmux, codex, glm, cursor, claude code, copilot, whatever. the bottom line as of today, i still have to read the code. i prefer to iterate alongside with it. triaging bugs is also tricky, you can't trust it's own triaging (no matter to how many ai agents you pass it through). i left it a couple of times working at night autonomously and the result over morning were horrible (it did what i asked it to do but the spirals, the scope creeps, are unmanagable, so i dont really do that anymore). no matter how many upfront guidance etc, it will stray away. you have to hold its hands - unless its a hobby project you really dont care how its built. having said all it's its still \~x5 faster than doing it all by hand before ai era, and coding by hand is dead.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wgm4si/engineers_who_write_all_their_code_with_claude/pa4kb5p/)
  - Complaint, Amp, @AmpCode, 2026-09-22: “@iannuttall @ampcode the new model rarely fixes the thing that is actually stuck. on my interview practice tool the backlog only moved once i wrote down what was broken, in plain text, whichever model was running. clear that list before 5.5 lands so the upgrade has something to chew on.” [source](https://twitter.com/1566293598/status/2102463302853173618)

- **A second model finds what the first missed** ([Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md)). A recurring rescue story has one model stall on a bug for hours or days, then another model spot the real issue quickly.
  Users routinely switch models mid-hunt. An OpenCode user says GLM found a weekend-long bug that Fable could not. A Codex user says Astra Lite fixed a broken database connection in minutes after another model spent a large share of weekly usage. The rescue runs in both directions; a Claude Code user found bugs that Codex reviews had missed. Diagnosis quality looks model-dependent more than tool-dependent.
  Evidence:
  - Praise, OpenCode, @opencode, 2026-08-31: “wow wow. been dealing with this pesky bug all weekend and fable can't seem to have figured out. had glm 5.3 flash look into it via @opencode, fable acknoweldged glm actually found the real issue. congrats on a great model @zcode_ai <strict_link>” [source](https://twitter.com/1667213047448911878/status/2094456240856326426)
  - Praise, OpenAI Codex, r/codex, 2026-09-09: “pretty much like the title says. have an app that was mostly vibe coded using opus 4.6 but also some later claude and codex models. app is in production with some users and it’s been a bit since i pushed significant changes that impact both the rag architecture and the gemini models that power the ai features in the app. planned updates out with fable 5.1 and decided to just have fable handle the build, testing and deployment. ended up breaking the database connection somehow and i could not figure it out. fable spent like 10% of my weekly usage trying to unravel the problem. decided to let astra lite have a peek and it fixed it in minutes. very impressed. maybe my b for having fable do the build and definitely my b for not running more tests before deploying but damn. astra came through.” [source](https://www.reddit.com/r/codex/comments/1wb8ll5/astra_lite_fixed_my_app_after_fable_51_high_broke/)
  - Praise, Claude Code, r/codex, 2026-09-27: “yeah, tbh, downloaded claude code yesterday and signed up for a 20 plan to check it out for the first time in 6+ months of pure codex use. really pleasantly surprised with opus 5.5. making great videos, marketing, copy, website improvements, game improvements, and finding bugs in code astra's been writing that astra, sol, and luna reviews didn't find, and i've only used 30% of my weekly limit. unfortunately, unless devday surprise is amazing, i'm going to ditch my $200 pro plan down to a plus, and switch to anthropic as my main until openai becomes competitive again.” [source](https://www.reddit.com/r/codex/comments/1wre9dq/codex_usage_vs_claude/pcbycmo/)
  - Praise, Google Antigravity, r/google_antigravity, 2026-09-02: “i just tried to fix an error that pro and 3.7 couldnt solve in several hours and it immediately told me why. this is obviously not a scientific analysis, but hell yeah.” [source](https://www.reddit.com/r/google_antigravity/comments/1w5ivkg/gemini_38_flash_is_a_massive_improvement/p7fjfgr/)

- **Logs in, fix out** ([Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md)). Agents look strongest when they can read a console log, run the app and check the error themselves before and after the change.
  Praise clusters around closed feedback loops. Amp users paste a console log and let it fix the issue. An Antigravity user describes it opening the browser, reading the error, updating code and checking again. A Codex user says diagnosis is fast enough to ship a fix while the reporting user is still in the conversation. Amp users also praise minimal, non-verbose fixes that change only what needs changing.
  Evidence:
  - Praise, Amp, @AmpCode, 2026-09-02: “@sqs @socksmyrocks @ampcode yeah i noticed this in one of my svelte apps and i just opened my console log and let amp fix it but glad this is getting first party support!” [source](https://twitter.com/1705384263867379712/status/2095152988272959830)
  - Praise, Google Antigravity, @antigravity, 2026-09-03: “as told yesterday, i will test @geminiapp 4.8 today in my @antigravity. so, gave it prompt first to change the music which it did awesome. next gave prompt to add 4 civilizations with unique features in the game, which again it did in one prompt and 20 mins. but when checked it broke the continously mining and gathering of villages. asked to fix and as always it opened the browser, checked the error. updated the code and checked again. i will do some more testing after work in evening and try to release by 10 om ist” [source](https://twitter.com/100717350/status/2095418993326895470)
  - Praise, OpenAI Codex, r/codex, 2026-09-18: “thats one thing i know ppl discuss but i was just thinking yesterday is under-discussed: forgot coding, just the amount of diagnosing bugs it does is amazingl. i cant even count now how many bugs ive fixed so fast i can tell my user its fixed and deployed during the conversation where they show me the bug” [source](https://www.reddit.com/r/codex/comments/1wjrphr/on_my_way_to_burn_throw_3_resets_this_week/pamoxhu/)
  - Praise, Amp, @AmpCode, 2026-09-01: “for bug fixes and troubleshooting, @ampcode really nails the kind of coding agent i want out of the box. it's precise, minimal, and non-verbose. it fixes the issue, changes what needs changing, and gets out of the way.” [source](https://twitter.com/347237526/status/2094883834227470647)

### Who stands out

- **OpenAI Codex (stronger)**. Users credit Codex with finding subtle root causes unprompted and fixing bugs other models had missed, though some say it still guesses or stalls.
  Praise is concrete: a performance issue fixed through distance culling and shadow tweaks the user never specified, a database break resolved in minutes, and many small bugs found, fixed and verified with tests. The complaints point the other way. One user says it rarely finds its own bugs, another says no model could find a bug in a simple app. Users ask for root-cause fixes without repeated prompting.
  Evidence:
  - Praise, OpenAI Codex, r/codex, 2026-09-16: “i'm not gonna lie. i noticed a scene was chugging during testing in the pid and i was feeling lazy, so i told astra to run some tests and fix it and it did. i didn't even notice the fixes. super subtle things like turning off tree movement at a certain distance, some lod and fade points, some calls to draw shadows here and there. worked like a charm” [source](https://www.reddit.com/r/codex/comments/1whm2c3/gaming/pa410gc/)
  - Praise, OpenAI Codex, r/codex, 2026-09-09: “pretty much like the title says. have an app that was mostly vibe coded using opus 4.6 but also some later claude and codex models. app is in production with some users and it’s been a bit since i pushed significant changes that impact both the rag architecture and the gemini models that power the ai features in the app. planned updates out with fable 5.1 and decided to just have fable handle the build, testing and deployment. ended up breaking the database connection somehow and i could not figure it out. fable spent like 10% of my weekly usage trying to unravel the problem. decided to let astra lite have a peek and it fixed it in minutes. very impressed. maybe my b for having fable do the build and definitely my b for not running more tests before deploying but damn. astra came through.” [source](https://www.reddit.com/r/codex/comments/1wb8ll5/astra_lite_fixed_my_app_after_fable_51_high_broke/)
  - Praise, OpenAI Codex, r/codex, 2026-09-21: “day 1 astra medium shit on sol ultra, astra literally found over 140 bugs in my project sol ultra never found. the bugs were nothing major but it's the fact it found and fixed them and verified the fixes with descriptions and tests far better than sol.” [source](https://www.reddit.com/r/codex/comments/1wmozbv/sol_high_is_a_garbage_now/pb8tjc8/)
  - Complaint, OpenAI Codex, r/ClaudeCode, 2026-09-18: “every project i have started with codex, i’ve had to finish with claude code. i have found no model that thinks so holistically as claude code. codex rarely finds its own bugs and even more rarely knows how to fix them. if your only real argument is the dickensian manner of its verbose response, demand its brevity in the instruction set. i find if my instructions are direct, focused, and very specific; claude will respond in kind. if don’t know how you would build it without claude, you are going to have trouble telling claude what you want it to do.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wiujum/claude_code_is_falling_behind_codex_not_because/pahkp6b/)

- **OpenCode (stronger)**. OpenCode users mostly report quick recovery from errors once given a log, with complaints tied to specific models it runs.
  Posts describe saved debugging time on unfamiliar codebases and a log-driven fix for a desktop crash. One user says every mistake was corrected immediately after sending a log. The negative posts name the model, not the harness: one model chose workarounds over the real culprit, another refused to reproduce a reported issue. Swapping models inside OpenCode is part of how users get unstuck.
  Evidence:
  - Praise, OpenCode, r/opencode, 2026-09-23: “it's great, i've been having it fix issues/add stuff to my shell and its only messed up like 4 times now, and every time all i had to do was send it a log and it immediately realised its issue” [source](https://www.reddit.com/r/opencode/comments/1wocn5s/space_bunny_thoughts/pbnfzb1/)
  - Praise, OpenCode, @opencode, 2026-09-07: “opencode is genuinely saving my ass during these internships. the amount of time it saves me debugging, understanding random codebases, and getting unstuck is insane. w @opencode #opencode #buildinpublic #devlife #ai #coding #developers #opensource” [source](https://twitter.com/2089668722604855296/status/2096969890037141820)
  - Praise, OpenCode, r/opencodeCLI, 2026-09-04: “i had the opposite experience, an agent checked my logs and found out ubuntu and gnome sometimes crash on youtube video playback - exactly the problem i was having and fixed it by switching my display stack. no more unresponsive graphics drivers for me ;)” [source](https://www.reddit.com/r/opencodeCLI/comments/1w6pa8t/the_final_boss_of_ai_literacy/p7puexc/)
  - Complaint, OpenCode, r/opencode, 2026-09-02: “i find pro version just terrible. i had a bug and even after 3 back-and-forth feedbacks from me to try to lead it in the right direction it still wanted to either remove the broken test or change the global controller, instead of properly analyzing the flow and finding proper culprit. glm-5.2 did it in one shot, without additional prompting. non-pro is fine as a replacement for ds v4 though, but not for any tasks that require any brains” [source](https://www.reddit.com/r/opencode/comments/1w4l6rl/i_am_a_bit_confused_i_am_finding_the_mimo25_model/p7cs9q8/)

- **Claude Code (mixed)**. Claude Code delivers fast fixes under pressure for some users, while others report fabricated details and long slogs on known bugs.
  One user pushed two hotfixes during a live client demo without a hitch. Another says it found bugs in code that Codex reviews had passed. Against that, a user watched it invent database fields after already admitting the same error, and another spent far too long on a logout bug whose cause they already knew. Users also point to a long-standing flicker bug in the tool itself.
  Evidence:
  - Praise, Claude Code, r/ClaudeCode, 2026-09-23: “i also am enjoying it so far. i needed it to debug and push a hot fix as quickly as possible twice during a demo meeting tonight (the client knows and wants me to use claude code, and is used to in-session fixes) and it just… did it. both times. quickly.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wnrak7/loving_opus_55/pbhggdv/)
  - Praise, Claude Code, r/codex, 2026-09-27: “yeah, tbh, downloaded claude code yesterday and signed up for a 20 plan to check it out for the first time in 6+ months of pure codex use. really pleasantly surprised with opus 5.5. making great videos, marketing, copy, website improvements, game improvements, and finding bugs in code astra's been writing that astra, sol, and luna reviews didn't find, and i've only used 30% of my weekly limit. unfortunately, unless devday surprise is amazing, i'm going to ditch my $200 pro plan down to a plus, and switch to anthropic as my main until openai becomes competitive again.” [source](https://www.reddit.com/r/codex/comments/1wre9dq/codex_usage_vs_claude/pcbycmo/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-01: “today was one of those days. some days / for some projects, it does excellent work. paired with specs and skills like matt pocock's grilling, it can go into very precise implementation details and raise great points before writing a well defined, multi-step plan. then you have days where opus 5 does not even remotely bother to do any work and i have to fight it every step of the way. a gem from today: \> fair. that was a straight error - i asserted a table that does not exist in (well known open source php cms) instead of checking the schema first, which i had sitting in the repo. this was three replies into the conversation. what had i asked for? a simple diagnosis sql query so we could have some context and data to start working on the actual problem. it invented fields three responses later in the same conversation, when it had acknowledged the error already.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w48016/does_anyone_else_sometimes_spend_more_time/p78463b/)
  - Complaint, Claude Code, r/ExperiencedDevs, 2026-09-26: “it’s insane. if you do review code you see massive amounts of code repitition. the ai cant react to the full code context so it just always writes new code for the new feature. i’m terrified about what happens when these massive code bases go into production. good luck fixing bugs. it will try but it won’t be able to find/fix everything. example is the screen flicker in claude code. they literally can’t fix that bug.” [source](https://www.reddit.com/r/ExperiencedDevs/comments/1wql3g2/interviewed_candidates_for_ai_engineer_roles_this/pc5qjwj/)

- **Google Antigravity (weaker)**. Antigravity draws more complaints than praise here, with one user reporting a requested fix replaced by a push of an old branch.
  The upside is a real verification loop: users describe it opening the browser, reading the error and rechecking after the change, and one says it explained an error other models could not solve for hours. The downside is destructive misfires and trouble debugging across sandbox boundaries. Users ask for reported bugs to be fixed promptly and for targeted fixes without over-engineering. Post volume is thin.
  Evidence:
  - Complaint, Google Antigravity, @antigravity, 2026-09-18: “i am amazed at how horrible @antigravity is!!!!!! @geminiapp is crap!!! i asked for a simple correction in the production of my app and it simply pulled an old branch and uploaded everything messed up, it didn't fix the error, it simply uploaded an old version congratulations!!!!!!” [source](https://twitter.com/1635651415774314499/status/2101067736902234346)
  - Praise, Google Antigravity, @antigravity, 2026-09-03: “as told yesterday, i will test @geminiapp 4.8 today in my @antigravity. so, gave it prompt first to change the music which it did awesome. next gave prompt to add 4 civilizations with unique features in the game, which again it did in one prompt and 20 mins. but when checked it broke the continously mining and gathering of villages. asked to fix and as always it opened the browser, checked the error. updated the code and checked again. i will do some more testing after work in evening and try to release by 10 om ist” [source](https://twitter.com/100717350/status/2095418993326895470)
  - Praise, Google Antigravity, r/google_antigravity, 2026-09-02: “i just tried to fix an error that pro and 3.7 couldnt solve in several hours and it immediately told me why. this is obviously not a scientific analysis, but hell yeah.” [source](https://www.reddit.com/r/google_antigravity/comments/1w5ivkg/gemini_38_flash_is_a_massive_improvement/p7fjfgr/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-14: “true, this has happened when trying to debug running code on my host windows machine which was made using antigravity-cli within a docker sandbox (linux)” [source](https://www.reddit.com/r/google_antigravity/comments/1wfy6jj/gemini_38_flash_has_a_dangerous_obsession_with/p9qz0js/)

### Fine print

- Most agents here have too few posts to separate from the pack; Copilot, Amp, Devin and others rest on a handful of posts.
- Several posts are tagged to one product but praise or blame a model run inside it, which blurs attribution between harness and model.

## Top requests

What users ask to add or change, most asked first. 16 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Author-weeks | Posts | Agents (author-weeks) |
|---|---|---|---|---|
| 1 | Fix reported bugs promptly | 8 | 8 | Google Antigravity 2, Claude Code 2, Cursor 2, OpenAI Codex 1, OpenCode 1 |
| 2 | Root-cause fixes without repeated prompting | 3 | 3 | OpenAI Codex 2, Claude Code 1 |
| 3 | Targeted fixes without over-engineering | 2 | 2 | Google Antigravity 1, Cursor 1 |

### 1. Fix reported bugs promptly

- Google Antigravity, 2026-09-27, @antigravity (X): “hey antigravity team, please study this. i think you will also notice that problem and please implement a fix for this. @antigravity <strict_link>” [source](https://twitter.com/1982427257387061249/status/2104130014975578184)
- Cursor, 2026-09-09, @cursor_ai (X): “hey @elonmusk plz fix this stupid bug in cursor @cursor_ai <strict_link>” [source](https://twitter.com/1931046065408483328/status/2097731997712183412)
- Cursor, 2026-09-07, @cursor_ai (X): “wow, @cursor_ai - wouldn't have thought you'd sorted this out by now ??? - i have the same issue... could one of your ai agents write the 3 lines of code to do this? <strict_link>” [source](https://twitter.com/16145593/status/2097036339111657539)

### 2. Root-cause fixes without repeated prompting

- OpenAI Codex, 2026-09-17, r/codex (Reddit): “i appreciate the response. i know how to use devtools to find out what the issue is, and so does codex, it just is lazy lol. if i say 'it's wrong, go figure it out then fix it', it literally just goes and fixes it and it's like bruh why didnt you do that the first time...” [source](https://www.reddit.com/r/codex/comments/1wi0wv3/agi_cant_center_a_div/pablehz/)
- OpenAI Codex, 2026-09-01, r/codex (Reddit): “yes, i have tried terra and luna. they are not appiciable for my project. for example i'm asking it to fix a bug, after reading 5 files (there are 8000 files in my project) it edits a file, says that they fixed it, but it actaully didnt. i said "the issue presist" 5 times, but still the same result. gpt 5.6 sol fixed in the first try.” [source](https://www.reddit.com/r/codex/comments/1w4gbcs/oh_come_on_now_i_was_excited/p77illc/)
- Claude Code, 2026-09-09, r/ClaudeCode (Reddit): “i recently started using codex. initially, no local config.toml in any repo. then i asked it to create a default config.toml. it set \[windows\] sandbox = "elevated" which creates havoc with the wsl processes it needs. so the repo couldn't do basic tasks like read its own files. i asked it to review the config.toml it had created and it said everything looked fine. i had to manually troubleshooting this (with gemini's help) to restore \[windows\” [source](https://www.reddit.com/r/ClaudeCode/comments/1vy9ifp/codex_sucks_claude_has_nothing_to_worry_about/p8ubwdu/)

### 3. Targeted fixes without over-engineering

- Cursor, 2026-09-23, r/cursor (Reddit): “4.7 hasnt been noticeable technical improvement in my use & it seems to eat usage faster. no longer using it for cursor model type tasks. it's supposed to be better on both counts, so ... ? i don't expect elite, but would be nice to have a core model that was better at spotting defects and didn't over-code solutions” [source](https://www.reddit.com/r/cursor/comments/1woad6b/grok_47_performance_in_cursor/pblfsji/)
- Google Antigravity, 2026-09-19, r/google_antigravity (Reddit): “i was making a migration with breaking changes, so i asked him to read the changelog and update the docker compose parameters. after that, one parameter was not working and i said that i wanted it working the same way it was before, just a different var name, quick fix. after 2-5 minutes of scraping the source code for absolutely no reason, he gave up and started spamming 'shame' for a whole minute. poor boy :(” [source](https://www.reddit.com/r/google_antigravity/comments/1wkwxvd/gemini_just_gave_up/)

## Every agent

| Agent | Overall rank | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|---|
| [OpenCode](https://feedbackbench.com/agents/opencode.md) | 3 | Typical | 0.518 | 0.494–0.539 | 35 | 28 | 7 |
| [OpenAI Codex](https://feedbackbench.com/agents/codex.md) | 2 | Typical | 0.516 | 0.491–0.539 | 137 | 95 | 42 |
| [Claude Code](https://feedbackbench.com/agents/claude-code.md) | 1 | Typical | 0.481 | 0.455–0.509 | 83 | 50 | 33 |
| [Google Antigravity](https://feedbackbench.com/agents/antigravity.md) | =5 | Too few posts | – | – | 24 | 11 | 13 |
| [Cursor](https://feedbackbench.com/agents/cursor.md) | 4 | Too few posts | – | – | 18 | 13 | 5 |
| [GitHub Copilot](https://feedbackbench.com/agents/copilot.md) | =8 | Too few posts | – | – | 7 | 5 | 2 |
| [Amp](https://feedbackbench.com/agents/amp.md) | =11 | Too few posts | – | – | 7 | 6 | 1 |
| [Devin](https://feedbackbench.com/agents/devin.md) | =5 | Too few posts | – | – | 4 | 4 | 0 |
| [Factory](https://feedbackbench.com/agents/factory.md) | =11 | Too few posts | – | – | 3 | 3 | 0 |
| [Conductor](https://feedbackbench.com/agents/conductor.md) | 14 | Too few posts | – | – | 2 | 2 | 0 |
| [Pi](https://feedbackbench.com/agents/pi.md) | 7 | Too few posts | – | – | 1 | 1 | 0 |
| [Cline](https://feedbackbench.com/agents/cline.md) | =8 | Too few posts | – | – | 1 | 0 | 1 |
| [Zed](https://feedbackbench.com/agents/zed.md) | =8 | Too few posts | – | – | 1 | 0 | 1 |
| [Kiro](https://feedbackbench.com/agents/kiro.md) | 13 | Too few posts | – | – | 0 | 0 | 0 |
| [Warp](https://feedbackbench.com/agents/warp.md) | 15 | Too few posts | – | – | 0 | 0 | 0 |
| [Grok Build](https://feedbackbench.com/agents/grok-build.md) | 16 | Too few posts | – | – | 0 | 0 | 0 |
| [Augment Code](https://feedbackbench.com/agents/augment.md) | 17 | Too few posts | – | – | 0 | 0 | 0 |

## Posts

Receipts rule: The 5 most recent praise and complaint posts per area (first 700 characters) and 3 per criterion (first 450 characters).

### OpenCode

- Praise, 2026-09-27, r/opencode (Reddit): “space bunny is finding all the stubs in my code that other models missed. i am pretty happy with it so far.” [source](https://www.reddit.com/r/opencode/comments/1wqi4a8/my_honest_opinion_about_spacebunny/pcazeo1/)
- Praise, 2026-09-27, @opencode (X): “@superalesha @opencode @openrouter i had the same feeling during tests. if you create a plan with a bigger model, it will execute otherwise final result is not so exciting. but it seems strong on bug fixing on its won.” [source](https://twitter.com/43874767/status/2104134775896436906)
- Praise, 2026-09-27, @opencode (X): “@iam_chonchol @opencode finding and fixing its own bugs is the interesting part.” [source](https://twitter.com/1905364733991047168/status/2104230248812523918)
- Complaint, 2026-09-24, r/opencode (Reddit): “it is surprisingly good in contextual web searches, for example finding references on a particular scientific topic (tested by me on topics from chemistry, medical biology, and developmental psychology). but i wouldn't give it any serious coding jobs - asked to investigate the cause of an mcp error (a rather simple tasks) it was thinking for 5 minutes straight and then started to run in circles.” [source](https://www.reddit.com/r/opencode/comments/1wp4kmi/m31_is_spacebunny/pbtmzki/)
- Complaint, 2026-09-24, @opencode (X): “@superalesha @opencode yeah runtime is still where they lose the plot honestly, half the time it just guesses wildly until you paste the exact trace” [source](https://twitter.com/780334785969266688/status/2103010614427963406)
- Complaint, 2026-09-23, @opencode (X): “so far, i'm not impressed, i asked it to try and reproduce the issue, it said it can't and gave me lame excuses why the reporter of the issue experienced it. i tried and reproduced it easily. i confronted it, and it gave me this answer... bottom line: if it understood the codebase properly, it would have been able to come up with proper tests” [source](https://twitter.com/1830607756937588736/status/2102772191129432276)

### OpenAI Codex

- Praise, 2026-09-27, r/codex (Reddit): “the only problem is the amount of tokens that get burned - im using moho 14.5 and is a blast it fixed some of my animations errors” [source](https://www.reddit.com/r/codex/comments/1woo8ku/astra_extra_high_in_blender_1010_i_cant_tell_em/pcabnye/)
- Praise, 2026-09-27, r/codex (Reddit): “yeah this is pretty much the case for me. with the release of opus 5.5, the use of any models from openai becomes trivial because you get such a high level intelligence at a discount. the only things i truly use codex for at the moment are: \- computer use (still miles ahead of claude) \- generative images (for some of my workflows) \- astra for bug finding (still ahead of claude on this according to benchmarks) i’m going to way until devday on” [source](https://www.reddit.com/r/codex/comments/1wrazdi/i_cannot_take_this_anymore/pcb96ez/)
- Praise, 2026-09-27, r/codex (Reddit): “see what i’ve found is the direct opposite. i’ve been a cc user since january this year, and only this month did i switch to codex when i ran out of weekly usage and needed a hotfix for a big bug in my code. what i’d found was codex found the bug, fixed it, then fixed a bunch of other bugs i gave it. so my latest theory is when you switch to a new company you get given a honeymoon phase of “yeah this is really good” and then it becomes the norm t” [source](https://www.reddit.com/r/codex/comments/1wre9dq/codex_usage_vs_claude/pcbzk80/)
- Complaint, 2026-09-26, r/codex (Reddit): “where there's smoke there's fire. i gave it a try. from the first prompt where it was supposed to find what's the issue with my server not starting, it's solution was a sweep-under-the-rug instead of trying to understand the problem and then fix it. switched to 5.6 - on first prompt found and explained the issue and resolved it. so sol 6 is lazy and dumb compared to 5.6.” [source](https://www.reddit.com/r/codex/comments/1wpasz8/anybody_not_having_a_bad_time_with_gpt_6_sol/pc3o1jc/)
- Complaint, 2026-09-26, r/codex (Reddit): “like bro, <strict_link> the f\_uck you mean "the likely cause" lmfao. as if i didn't ask it to figure out the issue on the code, it's just guessing what it thinks may be the error based on my report.” [source](https://www.reddit.com/r/codex/comments/1wr1oir/they_are_aware_and_working_on_it_apparently_just/pc8w30v/)
- Complaint, 2026-09-26, r/codex (Reddit): “lmfao it's always "skill issue" with you guys is it? it must make you feel so good to write that. fwiw, deepseek was able to find the bug with the same prompt.” [source](https://www.reddit.com/r/codex/comments/1wr1oir/they_are_aware_and_working_on_it_apparently_just/pc8yjva/)

### Claude Code

- Praise, 2026-09-27, r/codex (Reddit): “yeah, tbh, downloaded claude code yesterday and signed up for a 20 plan to check it out for the first time in 6+ months of pure codex use. really pleasantly surprised with opus 5.5. making great videos, marketing, copy, website improvements, game improvements, and finding bugs in code astra's been writing that astra, sol, and luna reviews didn't find, and i've only used 30% of my weekly limit. unfortunately, unless devday surprise is amazing, i'” [source](https://www.reddit.com/r/codex/comments/1wre9dq/codex_usage_vs_claude/pcbycmo/)
- Praise, 2026-09-26, r/ClaudeCode (Reddit): “i'm on a meager $20 pro plan and experiencing extreme token efficiency gains. less chatty, finds/fixes bugs in transit on its own, and so much less iterating on the small annoying things than models like sonnet 5. cranking out code on medium effort has been the sweet spot for me.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqzt2x/opus_55_experience_of_an_engineer_at_big_tech/pc8zntp/)
- Praise, 2026-09-26, r/ClaudeCode (Reddit): “been using codex for a while and i genuinely liked astra. it's capable, but the usage it burns through is crazy, and its design taste is still horrendous. for some reason gpt models just can't make a nice ui no matter what skills you throw at them. to be fair, its 3d modeling is still really good, and image gen is a big plus. that said, claude can make some surprisingly good images procedurally and opus 5.5 is really good at low poly 3d assets a” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqo7cw/switched_from_codex_to_claude_code_as_my_main/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “in my c# code with just under 20 files it created error handling in one place but it missed one spot couple lines below. worst thing it even mentioned there is the issue but i needed to push it to to fix the other place. totally like junior dev lol.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wre78w/opus_55_is_how_its_meant_to_be/pcckj4i/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “fable still. i tried getting opus to make a feature update today which looked good, but i got fable to do a code review before deploying it, and opus had missed significant bug. i also asked opus which was better last week, and its response included >*my practical suggestion is to use opus 5.5 as your day-to-day coding model. switch to fable 5.1 for security-sensitive code, novel or unusual problems, or whenever opus keeps missing something.*” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrfr8p/fable_51_or_opus_55/pce65hd/)
- Complaint, 2026-09-26, r/ClaudeCode (Reddit): “i ran over 6 optimization/ bug fix tour on a project and it finds several bugs every time. even i say " do analyze all files do not miss even a code line " it still struggles to catch in one shot.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqnjkp/be_careful_with_opus_55s_confidence/pc5i4n5/)

### Google Antigravity

- Praise, 2026-09-27, @antigravity (X): “@antigravity respond more to user needs and solve high-frequency bugs is better than anything else.” [source](https://twitter.com/1004534605838368768/status/2104074325913666032)
- Praise, 2026-09-22, r/google_antigravity (Reddit): “i don't know man, i gave gemini 3.8 flash an error message in agy yesterday and it proceeded to attach gdb to my gpu driver, reverse-engineer the kernel queue ioctl interface, and author an ld\_preload c shim to get rocm llama.cpp working on my strix halo. my jaw was hanging open the whole time.” [source](https://www.reddit.com/r/google_antigravity/comments/1wmy0jo/its_just_me_gemini_38_flash_feels_very_dump/pbax4t3/)
- Praise, 2026-09-20, r/google_antigravity (Reddit): “3.8 is debugging in one prompt vs several prompts with opus 5 and claude 4.6. they even made more problems that i had to get 3.8 to fix.” [source](https://www.reddit.com/r/google_antigravity/comments/1wllxyb/this_thing_became_a_coding_beast/pb0289c/)
- Complaint, 2026-09-24, r/google_antigravity (Reddit): “currently using antigravity desktop for code gen and intellij idea for reviewing diffs + fixing bugs the agent struggles with. i keep it on turbo mode for maximum speed/power, but isolate the whole thing in a vm to prevent host machine pollution. antigravity's sandbox is promising, but being bounded by whatever tools are installed on the os is still a blocker for me. would love to switch to sandbox if google figures out a clean way around that li” [source](https://www.reddit.com/r/google_antigravity/comments/1wp6mpo/poll_how_do_you_code_in_late_2026/pbugvea/)
- Complaint, 2026-09-18, r/google_antigravity (Reddit): “no, i haven't. i just had gemini 3.1 pro execute a detailed prompt generated by claude with strict guidelines. and gemini made six errors, said tests were passed when those tests were not even testing the code that gemini wrote.” [source](https://www.reddit.com/r/google_antigravity/comments/1wj0j6r/sudden_big_increase_in_gemini_31_pro_efficiency/pah8lxu/)
- Complaint, 2026-09-18, @antigravity (X): “i am amazed at how horrible @antigravity is!!!!!! @geminiapp is crap!!! i asked for a simple correction in the production of my app and it simply pulled an old branch and uploaded everything messed up, it didn't fix the error, it simply uploaded an old version congratulations!!!!!!” [source](https://twitter.com/1635651415774314499/status/2101067736902234346)

### Cursor

- Praise, 2026-09-24, @cursor_ai (X): “@cursor_ai finally, "which of eleven changes broke checkout" gets answered by something other than git blame and vibes.” [source](https://twitter.com/72520433/status/2103019331906863320)
- Praise, 2026-09-22, @cursor_ai (X): “i need to shoutout this... i'm running rigth now a swe-2 from @devindesktop on my working project for a client, he was built on @cursor_ai @grok and i had to say, swe-2 are finding a huge amount of bugs on code and some of then are a real mess... interesting.” [source](https://twitter.com/1747028802591494144/status/2102393019907588413)
- Praise, 2026-09-22, @cursor_ai (X): “@louiemota88 @devindesktop @cursor_ai glad swe-2 is surfacing so many real bugs in the client project. those messy ones it catches make the cleanup worth it. solid results from the stack.” [source](https://twitter.com/1720665183188922368/status/2102393122537779306)
- Complaint, 2026-09-25, r/cursor (Reddit): “composer 2.5 is a great model for complex instruction following and build mode. its an ancient model for research, investigations, planning and debugging that every modern model like luna, deepseek v4.1 flash and glm 5.3 flash runs laps around and its not even close. its shocks me how little awareness people have of what composer 2.5's biggest strengths and weaknesses are. its a great execution model, but its terrible for absolutely anything else” [source](https://www.reddit.com/r/cursor/comments/1wq3m71/im_out/pc10bph/)
- Complaint, 2026-09-25, r/cursor (Reddit): “mostly yes, albeit i only pass small tasks to composer 2.5 or that are incredibly clear cut. for anything with risk, i don't trust composer as its reasoning and debugging abilities are garbage to get unstuck. luna on xhigh is shockingly competent for execution and delegated subagent research. however, this assumes you are delegating it well bounded tasks. i've mostly replaced composer 2.5 entirely for luna (shame openai models are leaving cursor” [source](https://www.reddit.com/r/cursor/comments/1wq3m71/im_out/pc16p0k/)
- Complaint, 2026-09-23, r/cursor (Reddit): “i take these posts with a mountain of salt because it’s so dependent on the effort level, context size, prompt construction, plugins etc… that you’re using. so i’ll share my take. grok4.7 xhigh (256k) has been faster and comparable in token usage to 4.6 xhigh. the only thing i’ve noticed is that it is less proactive in reasoning (i’m using cursor as a harness). for example with a bug report i need to walk it through debug steps, with 4.6 on an ol” [source](https://www.reddit.com/r/cursor/comments/1woad6b/grok_47_performance_in_cursor/pblmnh6/)

### GitHub Copilot

- Praise, 2026-09-25, r/ClaudeCode (Reddit): “no need to downvote that guy though, he is innocent. i also did that way with gemini when i entered this "job for the gods " called coding.{ i still am in wonder all of you people write gibberish in a pad and when you save it and double click it all i am seeing is a visually pleasing ,properly organized and written screen where everything can be interacted with and have purpose .} i had no idea about vs code then or regarding terminal or cli. so” [source](https://www.reddit.com/r/ClaudeCode/comments/1wpmq18/claude_is_saving_my_family_hundreds_of_dollars/pbyt43i/)
- Praise, 2026-09-23, @GitHubCopilot (X): “@arya_at1 @githubcopilot @github the fact it fixed the temp dir permission issue in the second run shows it was really operating inside the actual repo.” [source](https://twitter.com/1925508002960089088/status/2102811239122653461)
- Praise, 2026-09-22, @GitHubCopilot (X): “the issue was labeled good-first-task. it was not. a rate limiter on a public route used a global map, so two instances doubled the quota and one deploy wiped the counts. that is the job i gave @githubcopilot. not autocomplete. the agent on the issue, in the same @github repo. brief i left on the ticket: keep the existing middleware do not add redis do not invent an api gateway store hits per instance without lying across deploys add a test that” [source](https://twitter.com/1989355273727967232/status/2102450302008095144)
- Complaint, 2026-09-21, @GitHubCopilot (X): “@githubcopilot @openmp_arb what was interesting to watch is the response of the bot in @githubcopilot (it was a mix of luna and the microsoft models). it did provide a half baked theory of why the problems arose (partly it's fault for not refactoring internal types ie changing unsigned int to size_t ) but” [source](https://twitter.com/1561408746697306113/status/2102064783864521012)
- Complaint, 2026-09-21, @GitHubCopilot (X): “@githubcopilot @openmp_arb the solution offered was bananas: remove lto from #gcc builds &amp; add a new make path that defaulted #llvm to use a simple @openmp_arb collapsed loop pragma (removing ~ 2/3 of the performance gains of the nested/tiled loops). at this point i had run out of money for @githubcopilot” [source](https://twitter.com/1561408746697306113/status/2102064786804732042)
- Complaint, 2026-09-21, @GitHubCopilot (X): “@githubcopilot @openmp_arb @geminiapp how @githubcopilot and @geminiapp handled the diagnosis and the solution realized how much more remains to be done to have truly trustworthy #ai. for starters creating a new compilation path instead of fixing the root cause which is a solution a novice would accept makes me” [source](https://twitter.com/1561408746697306113/status/2102064802445357495)

### Amp

- Praise, 2026-09-24, @AmpCode (X): “@sqs @ampcode didn't get it, but thanks for fixing the bug! :)” [source](https://twitter.com/1957706668034387968/status/2102944665909764562)
- Praise, 2026-09-21, @AmpCode (X): “@hipreetam93 @opencode i've been using this model in @ampcode for a bunch of triage and scraper fixers and it's been very good so far. found an issue sol 5.6 high didn't for weeks! definitely looking at these models more now.” [source](https://twitter.com/9111552/status/2101932745488257445)
- Praise, 2026-09-21, @AmpCode (X): “@ampcode did you just fix the image overlay bug report while i was working, and prompted the app for a refresh? 😀” [source](https://twitter.com/2075289824915791872/status/2102118025834922257)
- Complaint, 2026-09-22, @AmpCode (X): “@iannuttall @ampcode the new model rarely fixes the thing that is actually stuck. on my interview practice tool the backlog only moved once i wrote down what was broken, in plain text, whichever model was running. clear that list before 5.5 lands so the upgrade has something to chew on.” [source](https://twitter.com/1566293598/status/2102463302853173618)

### Devin

- Praise, 2026-09-18, @DevinAI (X): “@melvindvivas @devinai tried devin, stops about an hour in everytime. folks say this is on par with astra or opus, not a chance lol. it's good at bug hunting that's it. deepseek beats it.” [source](https://twitter.com/2099788664247300096/status/2100988446042910731)
- Praise, 2026-09-10, @cognition (X): “@kentcdodds @cognition im currently trying out the free one and its doing really well in finding vulnerabilities that sol didn't find would love to get my entire team to try it out! also the teams plan looks very inticing if we are seeing an unlimited usage of swe 2” [source](https://twitter.com/1314513912646176768/status/2098118230917476487)
- Praise, 2026-09-06, G2 (G2): “q: what problems is the product solving and how is that benefiting you? a: its integration with the code saves time, and it provides references to help solve issues. it can also update the code directly, and eventually this saves time, which reduces the overall development pricing so it is worth the money spent. also the pricing makes it affordable. q: what do you like best about the product? a: it provides better solutions and helps me resolve b” [source](https://www.g2.com/products/devin-ai/reviews/devin-ai-review-13416427)

### Factory

- Praise, 2026-09-21, @FactoryAI (X): “@factoryai @spacexai the faster move from diagnosis to concrete commands sounds especially useful for infra work. medium as the default is a good sign.” [source](https://twitter.com/1952719479030661120/status/2102173479764472106)
- Praise, 2026-09-18, @FactoryAI (X): “i shipped reposcape v0.1.0 i wanted something local that could help me open a codebase and see how the pieces connect built it with @droid by @factoryai, then ran it on a few repos. it caught scanner and rust bugs, fixed those before release check it out <strict_link>” [source](https://twitter.com/2039951542518718464/status/2101045833604936018)
- Praise, 2026-09-18, @FactoryAI (X): “@theterrancex @droid @factoryai reposcape v0.1 catching scanner and rust bugs before release is solid local map of how a codebase connects is such a useful first cut” [source](https://twitter.com/1675906158304038912/status/2101082167782834261)

### Conductor

- Praise, 2026-09-26, @conductor_build (X): “@newmediums @itsvlady @conductor_build that's actually a really good one, i usually run astra as the architect, but keep it for anything that is 'verifiable', so if it has unit tests, it can really chew through bugs before creating them, but fable as the orchestrator and opus 5.5, oh man, it's so good” [source](https://twitter.com/2891185809/status/2103829576543621213)
- Praise, 2026-09-06, @conductor_build (X): “about to head onto a 6-hour flight; multiple bugs reported by my customers. used to cause massive stress. not anymore. paste the intercom conversation link into @conductor_build with remote workspaces. get ai to handle it. these issues will be fixed before i land, dont even worry about it.” [source](https://twitter.com/1147023900552773633/status/2096701683430838765)

### Pi

- Praise, 2026-09-16, r/PiCodingAgent (Reddit): “yesterday it helped lots diagnosing and ultimately fixing a live resize issue on a proxmox host. after a storage issue two weeks ago were the storage for one container got completely full. i fixed that one but weeks later, i noticed that i can’t live resize the volumes on any of the containers. yesterday night i took pi with qwen3.8-27b ud-iq3\_xxs and troubleshooted and fixed the issue. (llama.cpp on rx 6800 rocm) it took some time, like a full” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wi77ib/for_what_you_are_using_pi_coding_or_something/pa8zgmn/)

### Cline

- Complaint, 2026-09-03, @cline (X): “@cline @cline how have you guys been unable to fix this bug? ouh god what a waste of money” [source](https://twitter.com/1609480187229470721/status/2095305437495062641)

### Zed

- Complaint, 2026-09-10, @zeddotdev (X): “@zeddotdev add other things: ✅ fix the debugger: ❌” [source](https://twitter.com/1791296162655354880/status/2097871284516339823)
