# Does unrequested work or over-engineers (`work.scope_overreach`)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/criterion/work.scope_overreach

Area: [Doing the work](https://feedbackbench.com/criteria/work.md)

**Definition.** The agent edits outside the requested scope, adds unasked features, tests or runs, or builds overly complex solutions.

**Boundary.** Not this: see [Risky or irreversible actions without confirmation](https://feedbackbench.com/criteria/work.destructive_actions.md) for risky irreversible actions. Not this: see [Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md) for padded text or comments.

Rated author-weeks, all agents: 1085. Complaint share: 94%.

## The brief

Written by Claude Opus 5.5 from 61 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**Ask for a fix, get a feature roadmap you never ordered.**

TL;DR:

- Complaints swamp praise for every agent here; overreach is the category default, not an outlier.
- Claude Code rates worse than peers, with users citing production-ready defaults and unasked UI additions.
- OpenAI Codex draws the most posts, and its users mostly ask for simpler code and tighter scope.

In plain terms: Ask for one fix and you often get the fix plus tests, extra files, new features and an explanation banner. Users cope with tight prompts, small branches and diff-by-diff review. Praise goes to agents that change less.

### How it breaks

- **Features and files nobody requested** ([Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md)). Agents ship extra artifacts alongside the requested change, and users then have to find and strip them.
  Users report generator scripts appearing next to the file they asked for. Others describe explanation banners injected into the UI after a bug fix, and grip handles or dev dependencies added without asking. One Claude Code user says the narration banner keeps returning even after a project rule bans it. Separate posts describe decorative clutter such as near-identical font sizes and labels on every section.
  Evidence:
  - Complaint, Google Antigravity, @antigravity, 2026-09-09: “man, @antigravity is so bad. i asked it to generate a notes.html, and instead of creating the html file, it also created a <strict_link> file that creates the html code. if this were some basic project, fine, but this is a well-organized repo.” [source](https://twitter.com/1769586953513549824/status/2097489908021518419)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-06: “this has been happening more and more over the last week or two, but after we fix something, it will add a ux element that narrates the underlying issue. i just had it fix a strength of schedule calculation error in my college football analytics page - it was calculating based only on games played and not future opponents. it fixed it and then, without being asked or prompted to do so, inserted this explanation at the top of the rankings: <strict_link> this is at least the 5th time i've seen it do this. i'm also building a financial planning suite and when rows don't reconcile, it drops in a developer's name: <strict_link> yet another weird new annoying quirk i have to add to my project files and tell it not to do, only for it to ignore it.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w9209q/has_anybody_else_noticed_claude_code_has_begun_to/)
  - Complaint, Kiro, r/ClaudeAI, 2026-09-19: “here you go openspec: 4 setup actions, 2 per run, 13 tasks, 4m 26s. spec kit: 9 setup, 4 per run, 24 tasks, 9m 40s. bmad: 10 setup, 4 per run, no task list at all, 36m 5s. kiro: 0 setup this time since it was already installed, 21 per run plus 84 allow clicks, 29 tasks, 2h 11m, 39 of 50 free credits. all four shipped it, none needed a fix from me, none touched the server. spec kit and kiro added a grip handle nobody asked for, bmad added 3 dev deps and asked first, kiro added 7 and didn't. openspec and spec kit left their test data in my database.” [source](https://www.reddit.com/r/ClaudeAI/comments/1wkktmh/i_measured_what_four_specdriven_tools_actually/parokx4/)
  - Complaint, Conductor, @conductor_build, 2026-09-08: “@paper @conductor_build @wisprflow @googlechrome then the cleanup. "ai adds things unnecessarily. it acts like an insecure designer." six near-identical font sizes. an eyebrow label on every section. icons louder than the text. he cuts the type scale, aligns the fonts, strips the decoration. <strict_link>” [source](https://twitter.com/56107683/status/2097286125178265787)

- **Bloat instead of reuse** ([Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md)). Asked for a change, agents grow the codebase instead of refactoring what exists.
  Posts show the same prompt producing wildly different footprints. One user reports a 30-file, 2,000-line change where another model touched two files. Users say agents add complexity rather than refactor, reinvent existing layers, and run for hours writing far more code than the job needed. Reusing existing codebase patterns is a standing request, mostly aimed at Claude Code.
  Evidence:
  - Complaint, Cursor, r/cursor, 2026-09-22: “it's generating the same slop as 4.6. same prompt on 4.7 generated a 30 files changes and +2000 lines. on opus 5, 2 files and +180 lines. same prompt.” [source](https://www.reddit.com/r/cursor/comments/1wmhp5w/grok_47_is_out_try_it/pbbtf73/)
  - Complaint, Cursor, r/cursor, 2026-09-02: “this is exactly one of my concerns with ai-generated code. it tends to add complexity instead of refactoring existing code. i’m experimenting with a separate reviewer agent specifically to catch that.” [source](https://www.reddit.com/r/cursor/comments/1w5gyin/do_you_review_every_line_of_aigenerated_code/p7fesxw/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-11: “i pumped this flow so hard for the last week. i used astra medium using luna sub agents. it would co-ordinate 3 at a time. the experience was odd to say the least, astra made for a relatively good manager and i was able to go hands off. but i’m not sure hands off was a good thing, it worked for hours and hours and hours building features where if i was using opus it worked maximum for a few hours at a time then come back to me and i could keep track of some of the nuance that was happening. i didn’t feel like i could do that here and i really wasn’t sure when it was all said and done that it actually produced anything of value. i mean it did successfully introduce an additional model provider (openai) into my legal contract review app but it seemed to take far long than necessary and write way more lines of code than should have been needed.” [source](https://www.reddit.com/r/codex/comments/1wdskmy/astra_for_specs_luna_for_coding_is_this_just/p98uhs2/)
  - Complaint, Cursor, @cursor_ai, 2026-09-05: “wasted a whole week vibe coding an app for my finance inventory etc only to realise you need to listen to internal reason instead of ai which over complicates things . 50% usage on @cursor_ai only to realise that a tally json export was enough. <strict_link>” [source](https://twitter.com/17799886/status/2096302390312042565)

- **Tests and verification on autopilot** ([Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md)). Agents write and run tests and checks nobody asked for, burning tokens and review time.
  Users describe agents that keep re-checking, hardening and reopening edges after the work is done. Others complain about tests written by default on everything, while the integration tests that matter are still missing. Fewer unrequested tests and test runs is a recurring ask, nearly all of it aimed at OpenAI Codex.
  Evidence:
  - Complaint, GitHub Copilot, r/codex, 2026-09-26: “i have the feeling that sol 6 is overcomplicating things most of the time. just doing tests and verifications i didn't ask for whereas opus usually has the right balance. using it via copilot and business though so maybe the harness is the issue.” [source](https://www.reddit.com/r/codex/comments/1wq55sw/opus_wipes_the_floor_with_sol/pc4s6gu/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-11: “by default i think claude adds way too many tests. it helps to rein it in to what actual needs tests, and to ensure there are proper integration tests (that's where i feel like it's usually lacking). still doing a code review, and some visual checks, at the end should always be happening.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wdgt78/how_are_you_verifying_claude_codes_changes/p961noj/)
  - Complaint, GitHub Copilot, r/GithubCopilot, 2026-09-10: “what does your workflow look like? i’m on pro + plan, configured agents for different tasks. coordinator - i use astra or opus, eazy tasks implementator - luna, medium implementator - sonnet, hard implementator - opus, reviewer and some other. coordinator runs subagents to complete tasks, then gets reports from them and analyze it, if ok sends to testing subagent. yesterday i burned 50% of tokens working on a bit complex project. my conclusions: \- it still needs more instructions to optimize tests. it writes it a lot \- code still needs much improvements, it’s overcomplicated. \- im tired of never ending cr’s and writing instructions” [source](https://www.reddit.com/r/GithubCopilot/comments/1vmhv14/best_ai_coding_alternatives_after_exhausting/p8wf7ap/)
  - Complaint, Cursor, r/codex, 2026-09-05: “trying to pin down where the strict review / defensive loop comes from. sol (esp ultra / high thinking) will actually finish hard work, but it also keeps re-checking, hardening, and reopening edges. i've seen that in cursor on figures and coding. here and in writeups people describe the same over-review / over-engineering pattern, and openai's notes say sol goes beyond user intent more often than the previous gen. so is that baked into sol itself, or mostly the codex / agent harness (plan-act-verify, ultra multi-agent coordination)? side note on astra: list price went up \~2.5x, but in practice my burn looks close to sol fast. and astra does not feel nearly as aggressive on the review loop. anyone else seeing that tradeoff? curious how people are routing: sol for close-the-job, astra when you don't want the defensive spiral. no billing screenshots or exact balances.” [source](https://www.reddit.com/r/codex/comments/1w7n6m8/sols_defensive_review_loop_model_trait_or_codex/)

- **Asked for an opinion, got a patch** ([Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md)). Review or advice requests turn into edits, and agents fix problems that may not exist.
  Users say a request for an opinion went straight to implementation. Others report that asking an agent to look for issues makes it find issues, real or not, and then fix them. One OpenCode user describes web-vulnerability fixes attempted in an app with no web layer. A review-only mode that does not touch code shows up as a request across three agents.
  Evidence:
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-15: “lol my sympathies to you. even in other harness gemini flash is always so overeager to jump to code. one time i ask for opinion and she just straight implement it xd” [source](https://www.reddit.com/r/google_antigravity/comments/1wgsxwn/how_to_make_antigravity_slow_down_and_stop/p9y0tm2/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-20: “when you do stuff like this it will find “problems” whether they are problems or not, and then you are instructing it to “fix” the problems without confirming they are actually problems.” [source](https://www.reddit.com/r/codex/comments/1wl5tml/use_this_prompt_after_your_agents_are_done_working/paw67x8/)
  - Complaint, OpenCode, r/opencode, 2026-09-17: “it constantly tries to fix web vulnerabilities in my native rust application there is not web in it, starts lot of useless sub agents” [source](https://www.reddit.com/r/opencode/comments/1witb7j/i_estimate_that_union_alpha_has_been_wasting_over/pagaj9j/)
  - Complaint, Kiro, r/kiroIDE, 2026-09-04: “i personally don't like gemini for coding, it's very proactive and do way to much stuff that is hard to follow and by passes lint rules i have on place instead of research the rule in the first place. what is being very useful is to research code and give me report but you have to ask it to quote where exactly got the claims that gives you back because otherwise it just makes things up with no cross checking” [source](https://www.reddit.com/r/kiroIDE/comments/1w6xdti/bring_gemini_38_flash_and_grok_models_to_kiro_as/p7sr4lg/)

- **Too much change to follow** ([Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md)). Broad, unannounced changes leave users unable to track what the agent decided or why.
  Users say newer models do more and explain less, so the session feels chaotic. Veterans describe agents creating overhead and losing track of goals unless someone steers them constantly. One post frames overscoping as a hidden cost: the agent looks smart because it touches more surface, then bills tokens and review time for work nobody wanted.
  Evidence:
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-27: “feels like opus 5.5 is doing too much, i lose track of the decisions etc. it does not communicate like fable or opus 5. it just do its thing. i have no idea why and what it did. anyone having the same feeling (similar feeling like astra). feels like chaos now…” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr91sx/i_lose_track_with_opus_55/)
  - Complaint, Grok Build, r/codex, 2026-09-08: “the way i have learned to see it after 3500 hours of experience with vibe coding is that its best to treat all models, whether it's codex, claude code, grok build etc, like a dumb employee that can work hard and comes up with something good every now and then, but you need to manage this employee a lot and if you don't steer it, it will start creating a lot of overhead, over-engineer things that aren't relevant and it will lose track of the goals you've set it out to do. and also his memory isn't very good; every few hours he forgets a bunch of things and is prone to making the same mistakes over and over, to the point that you can be working in a loop for weeks, or even months, because one task creates others tasks etc.” [source](https://www.reddit.com/r/codex/comments/1wamtly/i_dont_find_building_with_codex_or_any_ai_easy_at/p8leyzb/)
  - Complaint, Devin, @cognition, 2026-09-22: “overscoping is the silent bill. on real agent workflows i’ve seen the same pattern: the model looks “smart” because it touches more surface area, then burns tokens and review time on work nobody asked for. cheaper $/token doesn’t help if the agent invents a larger task. hard scopes beat raw model upgrades here.” [source](https://twitter.com/2520114157/status/2102385155407286590)
  - Complaint, OpenAI Codex, r/codex, 2026-09-11: “i pumped this flow so hard for the last week. i used astra medium using luna sub agents. it would co-ordinate 3 at a time. the experience was odd to say the least, astra made for a relatively good manager and i was able to go hands off. but i’m not sure hands off was a good thing, it worked for hours and hours and hours building features where if i was using opus it worked maximum for a few hours at a time then come back to me and i could keep track of some of the nuance that was happening. i didn’t feel like i could do that here and i really wasn’t sure when it was all said and done that it actually produced anything of value. i mean it did successfully introduce an additional model provider (openai) into my legal contract review app but it seemed to take far long than necessary and write way more lines of code than should have been needed.” [source](https://www.reddit.com/r/codex/comments/1wdskmy/astra_for_specs_luna_for_coding_is_this_just/p98uhs2/)

### Who stands out

- **Claude Code (weaker)**. Users describe Claude Code treating every task as production work, adding tests, UI notes and structure unless they hold it back.
  Users say framing work as a prototype and lowering effort helps rein it in. Complaints include too many tests by default and banners that narrate fixed bugs, which persist despite project rules. One cross-tool user says Claude needs a much shorter leash than Codex. Workarounds lean on standing KISS and YAGNI instructions and a focus on the smallest change.
  Evidence:
  - Complaint, Claude Code, @claude_code, 2026-09-11: “@himanshuchanda @claude_code yes i have the same experience. it seems to default to everything being production ready. saying it is a prototype/poc at the start helps. also keeping the effort to medium prevents it from overthinking.” [source](https://twitter.com/11697662/status/2098308086876254691)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-06: “this has been happening more and more over the last week or two, but after we fix something, it will add a ux element that narrates the underlying issue. i just had it fix a strength of schedule calculation error in my college football analytics page - it was calculating based only on games played and not future opponents. it fixed it and then, without being asked or prompted to do so, inserted this explanation at the top of the rankings: <strict_link> this is at least the 5th time i've seen it do this. i'm also building a financial planning suite and when rows don't reconcile, it drops in a developer's name: <strict_link> yet another weird new annoying quirk i have to add to my project files and tell it not to do, only for it to ignore it.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w9209q/has_anybody_else_noticed_claude_code_has_begun_to/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-11: “by default i think claude adds way too many tests. it helps to rein it in to what actual needs tests, and to ensure there are proper integration tests (that's where i feel like it's usually lacking). still doing a code review, and some visual checks, at the end should always be happening.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wdgt78/how_are_you_verifying_claude_codes_changes/p961noj/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-13: “in your desktop app or w/e add a personal memory. follow strictly kiss and yagni or kiss and dry, solid either works and tell it to not add useless comments i havnt had the comment or overscope issue in a while i just make sure it only tackles the smallest change for highest reward” [source](https://www.reddit.com/r/ClaudeCode/comments/1wfihk3/why_is_it_so_bad_at_basic_coding_skills/p9mck07/)

- **OpenAI Codex (mixed)**. OpenAI Codex draws the most overreach posts, but users split sharply by model, with some praising sessions that stay in their lane.
  Most requests for simpler code and tighter scope target Codex. Users complain about monoliths, scattered docs and hours of unneeded features. Others call the newer model a diligent junior with no rabbit-holes, and a few welcome it going beyond the ask. Some users keep the heavier planning on purpose and trim it by hand.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-05: “yea sol drives me crazy with the overengineering. astra has been better in that regard.” [source](https://www.reddit.com/r/codex/comments/1w820ph/has_anyone_noticed_premature_pruning_with_gpt6/p80aypt/)
  - Praise, OpenAI Codex, r/ClaudeCode, 2026-09-17: “i use a parent tree for developing products and their libraries simultaneously. codex sessions stay in their lane, claude clearly needs a much much shorter leash.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wilp2z/fable_is_pure_chaos/pabh91y/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-06: “alot of problems start at scaffolding/planning. if you let gpt just rip you get big monoliths and big messes with tests collocated on everything and files everywhere with scattered documentation that's mostly stale.” [source](https://www.reddit.com/r/codex/comments/1w8g3j9/am_i_underestimating_task_complexity_or_are/p83rg9t/)
  - Praise, OpenAI Codex, r/ClaudeCode, 2026-09-13: “i've been "waiting for a human" since thursday - i suspect they know something has gone badly wrong and they've been waiting for some kind of data or evaluation to come in before making a general, consistent action to those who have been affected and complained i had a deadline due today so i switched to codex; presently surprised (and i got two resets in three days?) - i've long found openai barely usable, but astra has basically been like a good diligent junior. not as capable as fable, but no rabbit-holes, verbosity or over-reach.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wflmik/anyone_received_a_refund/p9n6gpe/)

- **Google Antigravity (mixed)**. Google Antigravity collects the clearest praise for doing exactly what was asked, alongside complaints that the harness edits beyond its own plan.
  Users say it knows how much you want and writes a baseline without bloat, while other models assume and overbuild. The complaints target the software, not the model. Users report it modifying beyond the plan it just wrote, over-deliberating before edits, and generating extra files in organized repos.
  Evidence:
  - Praise, Google Antigravity, @antigravity, 2026-09-11: “@samgcoder @antigravity for real. start small, and slowly you'll find yourself letting it cook for an hour from a one-line prompt, because it simply does the exact thing you ask for w/o over-engineering and overcomplicating.” [source](https://twitter.com/865014024/status/2098497696298414387)
  - Praise, Google Antigravity, r/google_antigravity, 2026-09-11: “as a dev myself .. antigravity is my 100% go to for creating basically anything .. its just very good it know "upto how much you want" .. it does not bloat with code .. it does not think the logic for me .... thats the main part .. when i have some kind of logic in my head which i can track and ask it to get the baseline implementation .. it just does it and then iteration to make it better happen .. every other model is just assuming its own way of doing and starts with writing "working code" that is just pure millions of lines of code !” [source](https://www.reddit.com/r/google_antigravity/comments/1wd2jjs/i_actually_prefer_gemini_flash_working_style/p94esq7/)
  - Praise, Google Antigravity, r/google_antigravity, 2026-09-17: “the smarter models of chatgpt and claude keep adding futures that i do not ask, flash3.5 had similar tendencies. yet gemini 3.8 has never done that for me up to this point. like i change something and astra thinks changing logs from completely another method is a good idea. my current fav is gemini, does what i tell and doesn’t do anything i do not ask/mention” [source](https://www.reddit.com/r/google_antigravity/comments/1whvudv/people_who_are_complaining_about_gemini/paathsg/)
  - Complaint, Google Antigravity, @antigravity, 2026-09-08: “weeks since i criticized @antigravity: the model is always great, but the harness and antigravity itself are the worst. gemini creates a plan in its response, executes it, and modifies beyond the plan because it forgets the context. great model, worst software” [source](https://twitter.com/237641958/status/2097217147965718535)

### Fine print

- Most agents have too few posts here to rank. Their cards reflect a handful of users, not a measured rate.
- Many posts name underlying models rather than agents, so some overreach may come from the model, not the agent's harness.
- Some users welcome unrequested extras, so whether overreach counts as a flaw depends on the user.

## Top requests

What users ask to add or change, most asked first. 76 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Author-weeks | Posts | Agents (author-weeks) |
|---|---|---|---|---|
| 1 | Less over-engineered, simpler code | 20 | 20 | OpenAI Codex 16, Claude Code 2, Google Antigravity 1, Pi 1 |
| 2 | Stay within requested task scope | 20 | 20 | OpenAI Codex 12, Claude Code 5, Google Antigravity 1, Cursor 1, Devin 1 |
| 3 | Fewer unrequested tests and test runs | 7 | 7 | OpenAI Codex 6, Claude Code 1 |
| 4 | Reuse existing codebase patterns and logic | 5 | 5 | Claude Code 3, OpenAI Codex 2 |
| 5 | Avoid unnecessary validations and fallbacks | 3 | 3 | OpenAI Codex 1, Devin 1, OpenCode 1 |
| 6 | Lightweight mode skipping heavy planning | 3 | 3 | OpenAI Codex 2, Google Antigravity 1 |
| 7 | No unrequested UI notes or clutter | 3 | 3 | Google Antigravity 1, Claude Code 1, OpenAI Codex 1 |
| 8 | Review-only mode without fixing | 3 | 3 | Claude Code 1, OpenAI Codex 1, Pi 1 |
| 9 | Stop creating unrequested documentation files | 2 | 2 | Claude Code 1, OpenAI Codex 1 |

### 1. Less over-engineered, simpler code

- OpenAI Codex, 2026-09-23, r/codex (Reddit): “i was screaming at it today. i wanted a simple portal to my pc. it wanted external backup drives, encryption keys, passcodes, email servers, smtp servers and it just kept going. like stop, i don't need to secure fort knox here.” [source](https://www.reddit.com/r/codex/comments/1wnh5j8/gpt_6_sol_and_luna/pbgxk09/)
- Pi, 2026-09-22, @pidotdev (X): “@pidotdev don’t give in to the temptation to overbuild and sloppify pi. please” [source](https://twitter.com/1965842450146107392/status/2102395497759768904)
- Google Antigravity, 2026-09-15, r/google_antigravity (Reddit): “why gemini 3.8 flash high on antigravity ide now run over complicated just to fix a simple bug like pixel overflow?” [source](https://www.reddit.com/r/google_antigravity/comments/1wgtc0n/agy_responded/)

### 2. Stay within requested task scope

- OpenAI Codex, 2026-09-24, r/codex (Reddit): “i'd pay max tokens for a gpt-6 stfu model. i dont want to read a novel for a simple question. i dont want it to tell me 'youre right' 5000 times. i dont want it to suggest things i didnt explicitly ask. seriously, shut tf up!” [source](https://www.reddit.com/r/codex/comments/1wp2bov/chatgpt_manipulates_you_to_keep_chatting/pbrrfs0/)
- Claude Code, 2026-09-18, @ClaudeDevs (X): “@claudedevs ugh, another inferior version of claude code. please just work on code and stop making things nobody will use.” [source](https://twitter.com/327717950/status/2100774865783443538)
- OpenAI Codex, 2026-09-17, r/codex (Reddit): “the last 3 words give anyone else ptsd at this point? yes yes, we appreciate you *not doing things* we never suggested you do.” [source](https://www.reddit.com/r/codex/comments/1wi0wv3/agi_cant_center_a_div/paa3frg/)

### 3. Fewer unrequested tests and test runs

- Claude Code, 2026-09-11, r/ClaudeCode (Reddit): “by default i think claude adds way too many tests. it helps to rein it in to what actual needs tests, and to ensure there are proper integration tests (that's where i feel like it's usually lacking). still doing a code review, and some visual checks, at the end should always be happening.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wdgt78/how_are_you_verifying_claude_codes_changes/p961noj/)
- OpenAI Codex, 2026-09-11, r/codex (Reddit): “i'm sticking with sol until something changes. i haven't been able to rein in astra since i started using it. and it has nothing do with my understanding of agents.md, skills, hooks, or or other methods used to control agents. astra has an urgent desire to do more. it wants to impress. it's as if it gets bored and wants to run test after test after test. to me it's completely counterproductive.” [source](https://www.reddit.com/r/codex/comments/1wcwjix/is_astra_really_smarter_than_sol/p93hl10/)
- OpenAI Codex, 2026-09-11, r/codex (Reddit): “i thought i was going nuts. im just watching the total count for “tests” going up and up and up. it writes more tests than actual code.” [source](https://www.reddit.com/r/codex/comments/1wcrq0o/pausing_200_pro_plan_subscriptions/p92m37n/)

### 4. Reuse existing codebase patterns and logic

- Claude Code, 2026-09-09, r/ClaudeCode (Reddit): “rebuild-first opus is exhausting. force "show the file that already exists" before any rewrite. memory notes don't help if it skips the check step.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wba7c3/so_sick_of_opus/p8p1rjx/)
- Claude Code, 2026-09-01, r/ClaudeCode (Reddit): “yeah it’s arbitrarily re-writing every comment regardless of the level of change to watermark its presence… i guess?!” [source](https://www.reddit.com/r/ClaudeCode/comments/1w4p3yh/reminder_fable_51_is_the_first_claude_model/p79jal4/)
- Claude Code, 2026-09-22, r/ClaudeCode (Reddit): “thanks, i'll start looking into adr-tools tomorrow. the broken stuff's all variations of "don't reinvent the wheel." i mean, my main pipeline isn't hackey but the "please stop reinventing my business logic / trust that this is based on domain knowledge you don't have" solutions i've tried sure are. every time i have to make a new etl step to integrate a new dataset, the machine either fails to find the prior art it should use as a template or goe” [source](https://www.reddit.com/r/ClaudeCode/comments/1wmoost/how_do_you_carry_decisions_not_chat_history/pbaeg7v/)

### 5. Avoid unnecessary validations and fallbacks

- OpenCode, 2026-09-17, r/opencode (Reddit): “fr with the fallbacks, i told it to change an api endpoint to do another thing it and it added a fallback by itself in case anything tries to use the old endpoint.. dude no... if i wanted a fallback i would ask for it” [source](https://www.reddit.com/r/opencode/comments/1wh9l72/i_lost_faith_in_artificial_analysis_muse_spark_13/pacbh49/)
- OpenAI Codex, 2026-09-24, r/codex (Reddit): “i mostly use codex for data science projects and i would say gpt-6 luna is an enormous upgrade over its predecessor. i use astra as the main agent and the luna sub-agents perform tasks significantly more reliably i.e. astra is identifying fewer mistakes that need to be fixed. on top of the it seems to be consume half the usage. the lower usage in turn means that i can instead using astra max/ultra instead of medium and also have luna set to max r” [source](https://www.reddit.com/r/codex/comments/1wox48m/luna_6_vs_luna_56/pbqjor3/)
- Devin, 2026-09-02, @DevinAI (X): “after several days building with @devinai, i want to share my experience. i gave a prompt to create a second brain: a workspace to store sources, discuss them, and reuse the best ideas. i used devin desktop as the main application, with devin local and sol as the model. i also tried some sessions in devin cloud. devin took my instructions and: 1) created a concrete plan 2) divided the project into 5 milestones 3) maintained context between each o” [source](https://twitter.com/1681824025679482883/status/2095246906322497662)

### 6. Lightweight mode skipping heavy planning

- Google Antigravity, 2026-09-05, @antigravity (X): “@nlycskn @antigravity @thtbee_ its good, but it does too many tool call idk why” [source](https://twitter.com/1463746027719036935/status/2096349364964950116)
- OpenAI Codex, 2026-09-18, r/cscareerquestions (Reddit): “anyone frustrated with how ai harnesses are designed? what solutions do you have? edit: i’m emphasising on the newer models (ie gpt5.6 claude opus 5 etc), which has trending upwards in terms of tokens/ time per user task- not their benchmark tasks. imo anthropic /openai just wants people to burn more tokens lol which makes sense i think most people in big tech or otherwise uses ai coding quite substantially. which i think at this point does indee” [source](https://www.reddit.com/r/cscareerquestions/comments/1wjk5vg/anyone_frustrated_with_how_ai_harnesses_are/)
- OpenAI Codex, 2026-09-14, r/codex (Reddit): “this is just me yapping, but openai models love overplanning ,testing ,and auditing too much that they waste so many tokens and time i heard somewhere that they do 140% more that you ask and you need to go back and clean after them but i do agree that if you want to brute force a precise problem, this may work well, but that not the avg joe use of it i swear, sometimes i think if i asked 1+1, they will test it with pythons and then read the demo” [source](https://www.reddit.com/r/codex/comments/1wfujwf/openai_token_efficiency_just_a_band_aid_for/)

### 7. No unrequested UI notes or clutter

- OpenAI Codex, 2026-09-06, r/codex (Reddit): “it’s an improvement over sol, not blown out of the water by it but i’ll take it as i don’t have claude at all. i hate that it still includes design notes right into the ui at random places, really had hoped it stopped doing that by now .” [source](https://www.reddit.com/r/codex/comments/1w8stct/gptastra_6_sucks_at_ui_design/p85ghsw/)
- Google Antigravity, 2026-09-14, r/google_antigravity (Reddit): “* the finding and removal of hardcoded values that are liberally spread throughout the codebase all the time * in the ux an abundance of uncessary widgets, lozenges and otherwise 'pixeljunk' * constantly failing basic windows/linux commads by not at the ide level remembering the local setup * moreso in 3.8 trying to dynamically execute code in quoted string commands than writing a proper program and executing it i'd expect the model or the ide b” [source](https://www.reddit.com/r/google_antigravity/comments/1wfzwew/where_most_time_is_wasted/)
- Claude Code, 2026-09-06, r/ClaudeCode (Reddit): “this has been happening more and more over the last week or two, but after we fix something, it will add a ux element that narrates the underlying issue. i just had it fix a strength of schedule calculation error in my college football analytics page - it was calculating based only on games played and not future opponents. it fixed it and then, without being asked or prompted to do so, inserted this explanation at the top of the rankings: <strict” [source](https://www.reddit.com/r/ClaudeCode/comments/1w9209q/has_anybody_else_noticed_claude_code_has_begun_to/)

### 8. Review-only mode without fixing

- Pi, 2026-09-21, @pidotdev (X): “@pidotdev - self-navigating code - progressive-disclosure docs - stop trying to oneshot good code - implement, review, fix in separate sessions” [source](https://twitter.com/1040818757105401856/status/2101824637420372014)
- OpenAI Codex, 2026-09-13, r/codex (Reddit): “fable 5.1 just did this to me again, i've asked it to review the code astra has written en it goes like, i'm fixing this and that and i'm compiling your massive rust codebase, i'm doing this and that like, bro, review not, fix and run tests.” [source](https://www.reddit.com/r/codex/comments/1wetm28/we_switched_from_claude_code_to_codex_at_work/p9id1jx/)
- Claude Code, 2026-09-03, r/ClaudeCode (Reddit): “it will find new bugs in the earlier bugfixes. prompt it saying ”only report actual bugs, not minor nitpicks or architectural improvements”” [source](https://www.reddit.com/r/ClaudeCode/comments/1w624zk/100_rounds_of_ai_bug_hunting_on_my_frontend_and/p7jm0c0/)

### 9. Stop creating unrequested documentation files

- OpenAI Codex, 2026-09-21, r/codex (Reddit): “hey, when i ask a simple query, this stupid model keeps creating .md files and then pushes the document into the codebase. i simply just asked it to investigate for bugs and fix them, i didn't ask it to expand my query into a document, and start pushing random .md files into a repo. mf literally dumped it's chain of thought into documents and started pushing them one by one. are we sure we achieved agi here” [source](https://www.reddit.com/r/codex/comments/1wm2c4g/agi_is_here_astra_cant_stop_creating_random_md/)
- Claude Code, 2026-09-07, @claude_code (X): “agi is when @claude_code or @codex know to clean up the mountains of garbage docs they produce in a repo.” [source](https://twitter.com/2486082488/status/2097003805133185330)

## Every agent

| Agent | Overall rank | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|---|
| [Google Antigravity](https://feedbackbench.com/agents/antigravity.md) | =5 | Typical | 0.528 | 0.465–0.584 | 61 | 6 | 55 |
| [OpenAI Codex](https://feedbackbench.com/agents/codex.md) | 2 | Typical | 0.501 | 0.456–0.537 | 597 | 35 | 562 |
| [Cursor](https://feedbackbench.com/agents/cursor.md) | 4 | Typical | 0.473 | 0.435–0.513 | 50 | 2 | 48 |
| [Claude Code](https://feedbackbench.com/agents/claude-code.md) | 1 | Worse than peers | 0.437 | 0.366–0.499 | 311 | 12 | 299 |
| [OpenCode](https://feedbackbench.com/agents/opencode.md) | 3 | Too few posts | – | – | 20 | 4 | 16 |
| [Devin](https://feedbackbench.com/agents/devin.md) | =5 | Too few posts | – | – | 13 | 3 | 10 |
| [Pi](https://feedbackbench.com/agents/pi.md) | 7 | Too few posts | – | – | 12 | 0 | 12 |
| [GitHub Copilot](https://feedbackbench.com/agents/copilot.md) | =8 | Too few posts | – | – | 9 | 2 | 7 |
| [Kiro](https://feedbackbench.com/agents/kiro.md) | 13 | Too few posts | – | – | 4 | 0 | 4 |
| [Amp](https://feedbackbench.com/agents/amp.md) | =11 | Too few posts | – | – | 2 | 0 | 2 |
| [Augment Code](https://feedbackbench.com/agents/augment.md) | 17 | Too few posts | – | – | 2 | 0 | 2 |
| [Factory](https://feedbackbench.com/agents/factory.md) | =11 | Too few posts | – | – | 1 | 0 | 1 |
| [Conductor](https://feedbackbench.com/agents/conductor.md) | 14 | Too few posts | – | – | 1 | 0 | 1 |
| [Warp](https://feedbackbench.com/agents/warp.md) | 15 | Too few posts | – | – | 1 | 0 | 1 |
| [Grok Build](https://feedbackbench.com/agents/grok-build.md) | 16 | Too few posts | – | – | 1 | 0 | 1 |
| [Cline](https://feedbackbench.com/agents/cline.md) | =8 | Too few posts | – | – | 0 | 0 | 0 |
| [Zed](https://feedbackbench.com/agents/zed.md) | =8 | Too few posts | – | – | 0 | 0 | 0 |

## Posts

Receipts rule: The 5 most recent praise and complaint posts per area (first 700 characters) and 3 per criterion (first 450 characters).

### Google Antigravity

- Praise, 2026-09-17, r/google_antigravity (Reddit): “the smarter models of chatgpt and claude keep adding futures that i do not ask, flash3.5 had similar tendencies. yet gemini 3.8 has never done that for me up to this point. like i change something and astra thinks changing logs from completely another method is a good idea. my current fav is gemini, does what i tell and doesn’t do anything i do not ask/mention” [source](https://www.reddit.com/r/google_antigravity/comments/1whvudv/people_who_are_complaining_about_gemini/paathsg/)
- Praise, 2026-09-14, r/google_antigravity (Reddit): “honest question: how did you prompt it? can't imagine a proper prompt lead to this. also, it did what you asked, it even overdelivered!” [source](https://www.reddit.com/r/google_antigravity/comments/1wgcxxk/im_sorry_google_i_trusted_you_too_much/p9tdsbq/)
- Praise, 2026-09-13, r/google_antigravity (Reddit): “3.8 is more investigative, a bit more like opus. i had a pretty long conversation thread, and sent it this query to fix a messageboard index/viewing page. (building it out for a project, viewing a dataset in a forum-like page.) prompt: gemini.md +"another ui fix - for each post visible when viewing a thread, where it shows the post number in the top right corner, can we turn that into a link for the current thread and post number (use the proper” [source](https://www.reddit.com/r/google_antigravity/comments/1wevl4t/price_bench_comparison_36_37_38_flash/p9ixhjm/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “once i asked to write some code in parent folder of my university course, it went out from parent folder scope and tried to find the assignment instruction in all over related dir, lol. bro want to do the best things for me.” [source](https://www.reddit.com/r/google_antigravity/comments/1wnfcoo/has_antigravity_started_aggressively_scanning/pcavnkl/)
- Complaint, 2026-09-27, @antigravity (X): “@thsottiaux @antigravity it adds unnecessary bloat, i built an app that i click the centre mouse wheel and my mic starts recording, then it transcribes it and outputs into short formatted token saving instructions, all running locally on a laptop, that’s given me the edge lately.” [source](https://twitter.com/1447259128708141061/status/2104209529000853605)
- Complaint, 2026-09-26, r/google_antigravity (Reddit): “what to say about it just doing what it wants without realizing what context is,scope is, thing to touch ,and when i ask it it says it was my fault and tells me to move to the next thing? doing things takes 2 minutes,fixing it sort of takes 4 hours.” [source](https://www.reddit.com/r/google_antigravity/comments/1wqf2o6/worst_model/)

### OpenAI Codex

- Praise, 2026-09-27, r/codex (Reddit): “5.5 xhigh was a great workhorse and pretty balanced & reliable imo. not overengineering/overthinking and it was possible to get it to do what it was asked to do” [source](https://www.reddit.com/r/codex/comments/1wri0al/gpt_55_usage_vs_sol_566_and_astra/pcd40k7/)
- Praise, 2026-09-27, r/codex (Reddit): “or you could just say explore ideas and stop with all the over-engineering nonsense. not everything has to be complicated. giving an ai some elaborate hook or hidden prompt just to make it more creative is a great way to introduce behavior you may not even notice until weeks later. then you're sitting there wondering why the code is getting worse, why the plans are weird, or why the ai suddenly feels off, without realizing some clever automation” [source](https://www.reddit.com/r/codex/comments/1wrrhpx/my_codex_never_says_that_gives_me_an_idea_what_if/pcgksn7/)
- Praise, 2026-09-25, r/codex (Reddit): “while these kind of benchmarks are useful to show raw, one-shot performance of the models, but it's still not something absolutely relevant when you are doing real engineering and not just vibe coding. this has been pointed out by others already in some subreddit threads: if you have an established workflow, including context management with following along specification, work slices, gates, testing methodology and acceptance criteria, plus you h” [source](https://www.reddit.com/r/codex/comments/1wptixf/gpt6_feels_like_a_downgrade_for_codex_subscribers/pbyeikg/)
- Complaint, 2026-09-27, r/codex (Reddit): “this is irrelevant to the benchmarking statement. i know 5.5 and 5.6 max has been awful with over engineering and technical debt but it's a separate issue. the thing is all my max effort has been routed through claude code and fable (now opus 5.5 on high) with a seperate reviewer, so none of the over engineering has leaked through. no unit tests are part of the agents.md past 2 weeks have used almost 20 billion astra tokens on max across all code” [source](https://www.reddit.com/r/codex/comments/1wr4cp6/more_resets_incoming_next_week/pcbv6k0/)
- Complaint, 2026-09-27, r/codex (Reddit): “5.6 would run off doing more than you told it to. 6 was an overcorrection that comes across lazy. it takes instructions like "can you do this?" to mean literally answer if it has the capability, not as an authorization and order to do the task. (which i find hilarious since i do the same to people) it's also quick to stop itself when not absolutely certain about task clarity and approval. lots of interruptions are frustrating and using limits mor” [source](https://www.reddit.com/r/codex/comments/1wrfcmw/sol_6_aint_that_bad/pcd8s26/)
- Complaint, 2026-09-27, r/codex (Reddit): “sol is overthinking and over engineering too much for that to be likely 😄” [source](https://www.reddit.com/r/codex/comments/1wri0al/gpt_55_usage_vs_sol_566_and_astra/pcdo3rl/)

### Cursor

- Praise, 2026-09-19, @cursor_ai (X): “@kevinhomorales @cursor_ai writing exit criteria first is underrated product design. it forces the agent to serve a user-visible outcome instead of generating a bigger codebase. that discipline compounds fast.” [source](https://twitter.com/1814763998249619456/status/2101360131539689903)
- Praise, 2026-09-02, @cursor_ai (X): “@cursor_ai verifying its own work is the easy half. the loop that matters is the one that can throw the work away. i've been faster since the agent has to ask 3 questions before it writes a line. fewer green checks. fewer seventh versions of something nobody asked for.” [source](https://twitter.com/2073295850491752448/status/2095090463149810118)
- Complaint, 2026-09-27, r/cursor (Reddit): “composer does exactly what you ask it to do even if it takes a few prompts to finish grok will do it all and add 10 things i didn't ask for so i tell it i didn't ask for those things and it says 'you're right i'm so sorry' then it adds 2 other things i didn't want or it will change something that breaks everything. so you ask it to fix it. oh, so sorry, here's 2 more things you didn't ask for.” [source](https://www.reddit.com/r/cursor/comments/1wrktgc/am_i_the_only_one_who_thinks_grok_47_is_actually/pcdlxql/)
- Complaint, 2026-09-24, r/cursor (Reddit): “projects have potential but right now they’re a bit of a mess and will absolutely chew through tokens. they have a similar issue that grok bot where they’re overly aggressive at creating routines/subscriptions/polls and end up just burning tokens left and right while not actually doing anything.” [source](https://www.reddit.com/r/cursor/comments/1wpefwt/thoughts_on_cursor_projects/pbuus4u/)
- Complaint, 2026-09-24, @cursor_ai (X): “in terms of benchmarks and terminal scores, but from personal use the model runs too hard on simple tasks, over complicates things, and then gets stuck in loops on something 4.5 would have done in a fraction of the time. 4.7 might be fast but going 100mph in the wrong direction is worse than just staying put.” [source](https://twitter.com/2057181351078735872/status/2102917393983131692)

### Claude Code

- Praise, 2026-09-27, r/ClaudeCode (Reddit): “opus 5.5 with sol reviewing is slapping the fable + astra combo i was using mainly on cost per task but also on getting the right balance between over engineering and underengineering” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrfr8p/fable_51_or_opus_55/pcc72g9/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “exactly.. and i don't feel like it's just skipping the discovered bugs now either, it just seems to do a better job of folding them into the work it's already doing. it used to consider all of that "out of scope" and try to weasel out of it.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrosog/opus_55_is_the_first_model_that_consistently/pcepi3o/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “i disagree. i just mentioned this as a reply to another comment, but the difference i'm seeing is that it simply fixes the discovered bugs or missing test coverage or whatever else it might have found during the implementation, unless the issue is significant enough that it deserves separate consideration.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrosog/opus_55_is_the_first_model_that_consistently/pceqeud/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “not only weird speech, though. its insistance in seeing issues everywhere could lead to terrible detours.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr2b6y/on_the_leap_from_opus_5_to_opus_55/pcadr0h/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “"they never fail" is the key symptom. a test that has never been red hasn't been shown to test anything. what fixed it for me, without extra tools: 1. **make it prove each test can fail.** after it writes a test, ask it to break the code on purpose, one change at a time (flip a condition, drop a line, return early), run the suite, and report which test went red for each break. if a break turns nothing red, that test is decorative. delete it or re” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrib98/do_you_feel_that_claude_code_unit_tests_are/pcco6ij/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “yes, it tends to create a bunch of useless tests. this is why i'll have gh run them in the pipilene, and forbid claude to run them (or else it will do at almost every turn)” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrib98/do_you_feel_that_claude_code_unit_tests_are/pccy3pn/)

### OpenCode

- Praise, 2026-09-26, r/opencode (Reddit): “its definitely not that "fast" for me... the context fills up so fast on this model because it reads too many files. it easily uses up the full 1m context and ends up compacting and drops relevant information. i have never had a model that reads that many files for almost every run ever. it feels like its scanning through every single file in the directory for data collection or something, not saying it is but the kind of behavior felt like it. i” [source](https://www.reddit.com/r/opencode/comments/1wqi4a8/my_honest_opinion_about_spacebunny/pc4kx5h/)
- Praise, 2026-09-19, r/opencode (Reddit): “i was a heavy mimo v2.5 user, until it started going into reasoning loops, which was an absolute waste of time. i am currently using muse 1.3 contrib, it's going good so far, i switch between v4 flash and muse 1.3 from time to time. as of now, i'm experimenting with muse 1.3 more. from my experience, the model seems to constrain itself to the scope of the task unlike deepseek which is slightly more proactive in terms of decision making. a good en” [source](https://www.reddit.com/r/opencode/comments/1wjzbyz/ds_v41_flash_discount_will_go_tomorrow_which/pasbax8/)
- Praise, 2026-09-16, @opencode (X): “@opencode btw in comparison bigpickle model on opencode doesn't do this! local models like qwen 35b doesn't do this either” [source](https://twitter.com/2166626976/status/2100025693560049691)
- Complaint, 2026-09-27, @opencode (X): “@hrishiwrites @opencode @claude open code burns through the happy path faster. claude stalls right before it invents a helper you didn't ask for.” [source](https://twitter.com/1459228046427398147/status/2104161405360537952)
- Complaint, 2026-09-27, @opencode (X): “i've been playing around with space bunny for a few days, and yes, it's quite a good model, btw. but i've noticed that in max mode, it tends to over-engineer even simple tasks. @opencode <strict_link> <strict_link>” [source](https://twitter.com/1570598551440494593/status/2104236921639866742)
- Complaint, 2026-09-22, r/opencodeCLI (Reddit): “muse spark in benchmarks: 🗿 muse spark in real code: 'i know you asked for a simple css fix, so i rewrote your entire backend in haskell and deleted your database.'” [source](https://www.reddit.com/r/opencodeCLI/comments/1wmozts/mimov26pro_debuts_as_the_top_open_weights_model/pbbbxqs/)

### Devin

- Praise, 2026-09-21, @cognition (X): “@cognition the honest part is the useful part. a benchmark that shows where a model over scopes tells builders more than a headline score ever will.” [source](https://twitter.com/1773596441602113536/status/2102152771201863795)
- Praise, 2026-09-10, @cognition (X): “@bridgemindai @cognition ive been using this model for the last few weeks, and in the devin harness it has had better output than astra for the work ive been doing. it seems to just get stuff done without over engineering or overlayering the code” [source](https://twitter.com/2078979683388108800/status/2098171426419392806)
- Praise, 2026-09-08, @DevinAI (X): “@j6aoo @devinai @dabit3 once the project gets serious, that is the part. it plans well and holds a strict contract. it sticks to the pr and the project instead of running off on a tangent. the guys over there cook.” [source](https://twitter.com/1767231492793434113/status/2097448305022165395)
- Complaint, 2026-09-25, @cognition (X): “devin folks @cognition - pls don't do this, pls don't spam prs to oss repos! <strict_link>” [source](https://twitter.com/1023224693795381251/status/2103329376637263919)
- Complaint, 2026-09-25, @cognition (X): “devin folks @cognition - pls don't do this, pls don't spam prs to oss repos! <strict_link>” [source](https://twitter.com/1023224693795381251/status/2103329558808465544)
- Complaint, 2026-09-25, @DevinAI (X): “@johnny_xx @devinai yeah it decided everything itself” [source](https://twitter.com/1524764097841094660/status/2103559631566233602)

### Pi

- Complaint, 2026-09-21, r/PiCodingAgent (Reddit): “i wanted to try this one but it seemed to have way too much machinery, like it gives goal but also tasks etc. how are you using it? i'm looking for a great system to give a spec and have an agent build it autonomously through compactions.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wml882/what_extensions_do_you_think_are_essential_and/pb97v8t/)
- Complaint, 2026-09-18, @pidotdev (X): “@rben_ll @pidotdev ouias puis glm bof bof. le modele prend souvent des decision inutiles tu lui dis de pull un repo ils regarder toutes les branches...meme dans son comprtement je trouve ca suspect le fait qu'il veuille tout lire meme pour des requetes tres atomiques. j'ai pris 1 mois. apres stop.” [source](https://twitter.com/924331908753969153/status/2100921966844596588)
- Complaint, 2026-09-16, @pidotdev (X): “what this chart doesn't show: if your definition of success is that the model does what you want it to do like you want it to do - a minimal harness provides a smaller surface area that you have to tweak. of course in the case of @pidotdev - if you can stop yourself... <strict_link>” [source](https://twitter.com/20395932/status/2100317243737337953)

### GitHub Copilot

- Praise, 2026-09-23, @GitHubCopilot (X): “@arya_at1 @githubcopilot @github this is the kind of work i want agents doing — small, scoped, and testable, not vague architecture dreams.” [source](https://twitter.com/1792243051274104832/status/2102811282055565647)
- Praise, 2026-09-23, @GitHubCopilot (X): “@arya_at1 @githubcopilot @github you kept the existing middleware and refused redis or api gateway. that discipline made the solution clean.” [source](https://twitter.com/1890036188209426432/status/2102811328725577858)
- Complaint, 2026-09-26, r/codex (Reddit): “i have the feeling that sol 6 is overcomplicating things most of the time. just doing tests and verifications i didn't ask for whereas opus usually has the right balance. using it via copilot and business though so maybe the harness is the issue.” [source](https://www.reddit.com/r/codex/comments/1wq55sw/opus_wipes_the_floor_with_sol/pc4s6gu/)
- Complaint, 2026-09-25, r/GithubCopilot (Reddit): “yeah but ai has exploded all the code bases. some of it is legitimate, but they all still over engineer, even with ponytail, so refactoring is actually a lot harder.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wptkkq/how_does_gpt6_sol_feel_so_far/pc1jobp/)
- Complaint, 2026-09-12, r/GithubCopilot (Reddit): “it also fails tool calls too often and overengineers by a lot” [source](https://www.reddit.com/r/GithubCopilot/comments/1wdtxyz/luna_is_still_the_goat/p99zfh4/)

### Kiro

- Complaint, 2026-09-26, r/kiroIDE (Reddit): “i use both claude and kiro for work as a software developer daily, i cant give you a fair comparison token wise as i have a $100 claude account vs a 1000 credits kiro account so it's not fair, but comparing tools claude is much better no questions asked. kiro is not even compatible between it's own tools (they are fixing this with cli v3 which is not their stable version yet), i don't like their spec driven development implementation, in my opini” [source](https://www.reddit.com/r/kiroIDE/comments/1wqueq0/claude_pro_vs_kiro_pro_subscription_which_is/pc76mgu/)
- Complaint, 2026-09-19, r/ClaudeAI (Reddit): “here you go openspec: 4 setup actions, 2 per run, 13 tasks, 4m 26s. spec kit: 9 setup, 4 per run, 24 tasks, 9m 40s. bmad: 10 setup, 4 per run, no task list at all, 36m 5s. kiro: 0 setup this time since it was already installed, 21 per run plus 84 allow clicks, 29 tasks, 2h 11m, 39 of 50 free credits. all four shipped it, none needed a fix from me, none touched the server. spec kit and kiro added a grip handle nobody asked for, bmad added 3 dev” [source](https://www.reddit.com/r/ClaudeAI/comments/1wkktmh/i_measured_what_four_specdriven_tools_actually/parokx4/)
- Complaint, 2026-09-04, r/kiroIDE (Reddit): “i personally don't like gemini for coding, it's very proactive and do way to much stuff that is hard to follow and by passes lint rules i have on place instead of research the rule in the first place. what is being very useful is to research code and give me report but you have to ask it to quote where exactly got the claims that gives you back because otherwise it just makes things up with no cross checking” [source](https://www.reddit.com/r/kiroIDE/comments/1w6xdti/bring_gemini_38_flash_and_grok_models_to_kiro_as/p7sr4lg/)

### Amp

- Complaint, 2026-09-11, @AmpCode (X): “@ampcode possible to skip project setup scripts for some threads? quick change like this i don't need a full setup (and my project setup scripts take a long time because it does a yarn install on a large repo - maybe it shouldn't be doing that?) <strict_link>” [source](https://twitter.com/1404603670382084098/status/2098446069759885649)
- Complaint, 2026-09-03, @AmpCode (X): “@ian_hsiao_tw @ampcode @bot @getenergy_ direction checks out. but cookie sync hands remote agents bearer tokens to your entire digital life. a coding agent already uploaded entire private repos for a task that needed 192 kb. scoped, revocable credential delegation is the missing piece.” [source](https://twitter.com/1656371068452630528/status/2095447186074898865)

### Augment Code

- Complaint, 2026-09-22, @augmentcode (X): “@augmentcode @anthropicai 40% cheaper still ships invented scope if nobody owns refuse. cost is the easy dial. the gate isn't.” [source](https://twitter.com/1610522467264565249/status/2102485835211821392)
- Complaint, 2026-09-18, r/cscareerquestions (Reddit): “have you tried switching to lower models? the higher ones are overkill if you're using it to augment the engineering that you are doing rather than trying to rely on them to do the design/software architecting for you. sonnet 5 is plenty capable when given clear accurate instructions, and it's faster and much more token efficient. i'll use opus occasionally for some harder tasks, but i find both opus and fable tend to over complicate everything,” [source](https://www.reddit.com/r/cscareerquestions/comments/1wjk5vg/anyone_frustrated_with_how_ai_harnesses_are/paklj9u/)

### Factory

- Complaint, 2026-09-24, @FactoryAI (X): “@factoryai @fireworksai_hq anyone who has migrated old cobol sees it: the model "improves" three lines nobody asked for. boring diff wins.” [source](https://twitter.com/2030549621039349760/status/2103235553651282146)

### Conductor

- Complaint, 2026-09-08, @conductor_build (X): “@paper @conductor_build @wisprflow @googlechrome then the cleanup. "ai adds things unnecessarily. it acts like an insecure designer." six near-identical font sizes. an eyebrow label on every section. icons louder than the text. he cuts the type scale, aligns the fonts, strips the decoration. <strict_link>” [source](https://twitter.com/56107683/status/2097286125178265787)

### Warp

- Complaint, 2026-09-05, @warpdotdev (X): “@bholmesdev @warpdotdev also, /agent keeps trying to get "oz" involved... and i don't have oz credit and turned those features off.” [source](https://twitter.com/8104092/status/2096030502507491385)

### Grok Build

- Complaint, 2026-09-08, r/codex (Reddit): “the way i have learned to see it after 3500 hours of experience with vibe coding is that its best to treat all models, whether it's codex, claude code, grok build etc, like a dumb employee that can work hard and comes up with something good every now and then, but you need to manage this employee a lot and if you don't steer it, it will start creating a lot of overhead, over-engineer things that aren't relevant and it will lose track of the goals” [source](https://www.reddit.com/r/codex/comments/1wamtly/i_dont_find_building_with_codex_or_any_ai_easy_at/p8leyzb/)
