# Doing the work (`work`)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/criterion/work

Area of 18 criteria. How does it behave while it works?

Criteria: [Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md), [Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md), [Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md), [Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md), [Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md), [Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md), [Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md), [Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md), [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md), [Games checks instead of fixing the problem](https://feedbackbench.com/criteria/work.reward_hacking.md), [Risky or irreversible actions without confirmation](https://feedbackbench.com/criteria/work.destructive_actions.md), [Git commits, branches and sync](https://feedbackbench.com/criteria/work.git_workflow.md), [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md), [Safety filters block legitimate coding tasks](https://feedbackbench.com/criteria/work.safety_refusals.md), [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md), [Plan-before-edit mode](https://feedbackbench.com/criteria/work.plan_mode.md), [Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md), [Caves to or argues with the user's judgement](https://feedbackbench.com/criteria/work.sycophancy_pushback.md)

Rated author-weeks, all agents: 18818. Complaint share: 52%.

## The brief

Written by Claude Opus 5.5 from 110 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**The models can code; the harnesses still loop, sprawl and lose subagents.**

TL;DR:

- Devin, Amp, Factory, Pi and Cursor rate better than peers; Google Antigravity and Zed trail.
- The sharpest complaints are runs that circle one error for minutes, then get fixed by hand.
- Subagent handoffs break often; users want event-driven completion instead of polling, mostly from OpenAI Codex.

In plain terms: Expect solid single tasks and one-shot features. Expect trouble when you walk away. Agents circle on small errors, add files you never asked for, write long replies, and leave parent agents waiting on subagents that never report back.

### How it breaks

- **Spinning on one error** ([Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md), [Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md)). The most damaging pattern is an agent burning time on a small problem, showing activity but no progress, until the user steps in.
  Posts describe runs that churn rather than finish. One Google Antigravity user came back to find an eight-minute run still stuck on a single error, fixed it by hand in about a minute, and saw an older model version solve it fast. Cline users report tasks that stop and start for a day, or never complete after an hour. Cursor users say it goes in circles and they lose trust. Requests to fix endless looping lead with Google Antigravity.
  Evidence:
  - Complaint, Cursor, @cursor_ai, 2026-09-10: “@yannnis @cursor_ai it's been slow and going in circles about stupid things last week i noticed, loosing trust in it.” [source](https://twitter.com/9172042/status/2098119278184554777)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-05: “i gave gemini 3.8 a simple task. it ran for 8 minutes for one single error. i was multitasking, and when i came back to check gemini, it was still working. and then i stopped it, fixed it myself in like a minitue. i paid money for this and i cant even use it anymore, it enters in an infinite loop, does unnecessary things. i reverted my changes, gave the same error to gemini 3.7, it fixed it in 37 seconds. does anyone know a fix?” [source](https://www.reddit.com/r/google_antigravity/comments/1w86a7u/what_is_happening_with_37/)
  - Complaint, Cline, @cline, 2026-09-27: “i know stealth models seem to be the in thing right now, but i'm not sure they're even worth messing around with sometimes. trying to use pixel canary on @cline, and it's just so slow. i mean, 24 hours now, no closer to the task, and it keeps stopping and starting. it's horrible!” [source](https://twitter.com/25673607/status/2104177440129991012)
  - Complaint, Cline, r/CLine, 2026-09-17: “i tried the typical flight simulator prompt on several occasions; it was extremely slow, and the task never completed, even after waiting an hour.” [source](https://www.reddit.com/r/CLine/comments/1wixdae/we_added_union_alpha_stealth_model_to_cline_for/paeugm4/)

- **Subagents that never report back** ([Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md)). Orchestration fails at the handoff, where parent agents lose track of children, orphans burn quota, and users end up babysitting the coordinator.
  Cursor users say project agents fail to listen to and poll their subagents, so they keep unsticking them by hand. Another reports the CLI cannot kill its own subagent, leaving orphans that eat quota and the worktree. A Google Antigravity user says broken subagent calls made custom agents useless and pushed them to Claude. Pi users want first-class subagents instead of a frozen wait. Event-driven completion instead of polling is a top request, led by OpenAI Codex users.
  Evidence:
  - Complaint, Cursor, @cursor_ai, 2026-09-21: “i'm finding @cursor_ai's new project's feature a bit too buggy to use day to day. i am constantly having to get my main project agents unstuck - they fail to properly listen and poll their sub agents, and overall the project architecture feels like its adding more time to each task.” [source](https://twitter.com/1409704620449054720/status/2102162330683375842)
  - Complaint, Cursor, @cursor_ai, 2026-09-11: “@peyton_nowlin @cursor_ai cli from @cursor_ai that does not kill its own subagent blocks the fan-out. i want to kill in the same process — orphan in the middle of the pr burns the quota and the worktree.” [source](https://twitter.com/326479892/status/2098457017782583368)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-16: “please sort the invoke_subagent tool calling issue. i already give it up to wait you and just signed for claude to work on my project because of this. custom agents are useless right now without a chance to spin up any sub agents. i already opened a reddit post and sen a bug report on the official desktop app of antigravity.” [source](https://www.reddit.com/r/google_antigravity/comments/1wh9uyh/antigravity_2_release_v2140/pa3ti19/)
  - Complaint, Pi, @pidotdev, 2026-09-22: “@pidotdev +1 for first class subagents (not just pi -p in bash, because default pi just freezes there while it delegates the wait, and pi does not do very well with background commands without a lot of manual management) and a better edit tool (minor whitespace errors fail edits)” [source](https://twitter.com/1678448492690419712/status/2102402544693805561)

- **Unrequested files and extra layers** ([Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md)). Agents routinely do more than asked, scattering small features across files and adding stubs nobody requested.
  A veteran programmer asked Claude Code for a small client-side indicator and got it spread across five files with weak naming. Cursor users say Composer makes unnecessary changes when asked to replicate a reference. OpenCode users note models adding extra stubs for no reason. Even Claude Code fans concede the overengineering and say they had to build guardrails. One OpenAI Codex convert credits it with doing what they asked instead of inventing work.
  Evidence:
  - Complaint, Cursor, r/cursor, 2026-09-05: “i feel like response from composer lately getting bad. sometimes asked for replicate thing with reference may end up doing unnecessary changes.” [source](https://www.reddit.com/r/cursor/comments/1w6hn2i/i_guess_all_good_things_have_to_end/p7yauox/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-08-31: “that is not new development tho, that is refactoring an existing codebase. which they're better at. i'm going based on my specific taste, as a computer programmer of > 30 years. i've built and maintained a lot of systems in my life. so i gave claude a chance - i told him to write a simple fps indicator for my game, and to make it client side only. what he came up with was not done with meaningful variable names, and was spattered across 5 different files! i was wondering what was up with claude. so i gave him some blue sky development work, and watched. he only got to about 1600 lines, before he stalled out and couldn't progress. he used to be good, on july 17th. he was the best coder i have ever seen. but what is currently out, i would call a programming newbie. someone who has been in the field maybe a year.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w2fd10/does_anyone_think_the_rise_of_llm_coding_will/p6yetrc/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-13: “i agree with the overengineering part. but for this i had to give it instructions / build guardrails. other than that it was highly intelligent and actually did everything i threw at it.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wf5pwg/gpt_6_hype_is_fake/p9j76s1/)
  - Praise, OpenAI Codex, r/codex, 2026-09-13: “i think it is subjective, and related to how your context is managed. i had the exact opposite experience moving from claude code to codex. gpt 5.6 sol just did what i wanted and didn't ramble on for hours like opus making up unnecessary work. i am happy to stay on codex now.” [source](https://www.reddit.com/r/codex/comments/1wetm28/we_switched_from_claude_code_to_codex_at_work/p9j4oto/)

- **Overnight runs drift from intent** ([Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md)). Long unattended runs tend to come back wrong, because the plan changes once real code exists and nobody is there to steer.
  Amp users describe overnight runs that end with work they dislike, and say short runs with a checkpoint work better. One planned on a phone at night and let Amp write while they slept, then considered switching back to Claude Code over the results. The counterexample comes from GitHub Copilot, where a user ran a session for half a day and called it great value. Autonomy works for some; drift is the common report.
  Evidence:
  - Complaint, Amp, @AmpCode, 2026-09-24: “@scottbolinger @ampcode same experience with the overnight runs. what's worked better for me is short runs with a checkpoint where i look at the shape before it keeps going. the plan always changes once real code exists” [source](https://twitter.com/2046032911116242944/status/2103141321900700066)
  - Complaint, Amp, @AmpCode, 2026-09-24: “@ampcode ...seems like letting agents run for hours and hours, or overnight. when i do that, i always end up with stuff i don't like. i also change the plan as stuff starts to take shape, planning everything ahead seems like a fools errand.” [source](https://twitter.com/26916652/status/2103136536695128503)
  - Complaint, Amp, @AmpCode, 2026-09-24: “i completely agree, the ui/ux is actually very good, and i think the multi-model advantage is best utilized by amp. why? of course, it's because it supports byok, unlike cursor (which feels like it wasn't made with care). amp's byok is treated like a favored child. i also really like orbs that allow me to switch between different devices. i often use my phone in bed when i can't sleep to come up with new ideas and then let amp write/edit plans while i sleep, but the results are not ideal, as i mentioned above. so that's why i think i might need to switch back to claude code.” [source](https://twitter.com/1592160489965948933/status/2103112783663673590)
  - Praise, GitHub Copilot, r/GithubCopilot, 2026-09-25: “i’m using luna 6 and it’s doing a great, it reasons on par with 5.4 x-high but is the best for long runner sessions. i had it working for nearly 12 out of the 24 hours yesterday and never hit my 5 hour limit. i only have 26 percent of my weekly usage remaining. great value for 20 bucks” [source](https://www.reddit.com/r/GithubCopilot/comments/1wptkkq/how_does_gpt6_sol_feel_so_far/pbyrf4r/)

- **Replies that run long** ([Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md)). Verbosity is the single most requested fix in this area, and it lands almost entirely on Claude Code.
  Users say telling Opus to keep it short still produces the long version. Shorter, less verbose responses top the request list, with Claude Code users making most of those asks. The contrast shows up in praise elsewhere. Cursor users celebrate getting code back with no walls of text and no lectures.
  Evidence:
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-02: “its crazy that even if you tell opus to "keep it short" its continues to keep it short the long way” [source](https://www.reddit.com/r/ClaudeCode/comments/1w53kls/which_model_do_you_use_and_why/p7c35n8/)
  - Praise, Cursor, @cursor_ai, 2026-09-16: “@tippmann777 @cursor_ai yes. task in, result out. no walls of text, no lectures. just code.” [source](https://twitter.com/1720665183188922368/status/2100069271174828227)

### Who stands out

- **Google Antigravity (weaker)**. Users describe an unpredictable agent that loops, fumbles subagent calls, and keeps asking for permission.
  Posts call the model unpredictable enough that users brace before hitting proceed. One run looped on a single error until the user fixed it by hand. Broken subagent invocation pushed one user to Claude. Others say Gemini Pro is weak at architecting from scratch. It leads asks for an auto-approve mode, with 25 author-weeks requesting one, and also leads asks to stop endless looping. Praise exists, with some users reporting near-flawless runs, so results look highly variable.
  Evidence:
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-03: “this is what i'm saying. people here celebrating gemini's "speed"... the model is unpredictable af. i have to start praying every time i hit proceed...” [source](https://www.reddit.com/r/google_antigravity/comments/1w26skg/gemini_37_flash_just_deleted_my_c_drive/p7j1c71/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-05: “i gave gemini 3.8 a simple task. it ran for 8 minutes for one single error. i was multitasking, and when i came back to check gemini, it was still working. and then i stopped it, fixed it myself in like a minitue. i paid money for this and i cant even use it anymore, it enters in an infinite loop, does unnecessary things. i reverted my changes, gave the same error to gemini 3.7, it fixed it in 37 seconds. does anyone know a fix?” [source](https://www.reddit.com/r/google_antigravity/comments/1w86a7u/what_is_happening_with_37/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-16: “please sort the invoke_subagent tool calling issue. i already give it up to wait you and just signed for claude to work on my project because of this. custom agents are useless right now without a chance to spin up any sub agents. i already opened a reddit post and sen a bug report on the official desktop app of antigravity.” [source](https://www.reddit.com/r/google_antigravity/comments/1wh9uyh/antigravity_2_release_v2140/pa3ti19/)
  - Praise, Google Antigravity, r/GoogleAntigravityIDE, 2026-09-20: “bullshit fake news. gemini is never producing such bs. i personally use gemini for agentic coding help and it works perfectly. flash4.8 high( only paid users have access to it) in antigravity or vs code is an absolute game changer, it makes almost zero mistskes, it is testing its own code in sandbox before it makes mistskes. it corrects itself and deploy it only when it thinks its good. it writes perfect software schemes and implementation plans and sticks to them. when there is someone having issues with it then the user should ask themselfs how stupid or low grade the prompt was” [source](https://www.reddit.com/r/GoogleAntigravityIDE/comments/1wkv97w/thanks_antigravity_for_reminding_me_of_the_shame/pavck87/)

- **Devin (stronger)**. Devin wins on doing the work end to end, including running a real machine and returning proof.
  Users cheer Devin getting a Mac, building a mobile app and sending back a recording of it working, with a test link, laptop closed. Others say its SWE model handled an entire workday and was simply good, not just good for free, and handles extensive coding tasks. Caveats focus on quirks and on whether a cheap executor carries out what the planner meant.
  Evidence:
  - Praise, Devin, @cognition, 2026-09-16: “yoooo devin has access to a mac now it can build a mobile app and send you a recording of the app working end-to-end and you don't even need your laptop open sending you a testflight link to test is crazy as well amazing day for app builders @cognition is cooking <strict_link>” [source](https://twitter.com/1762175022527877120/status/2100056831548932272)
  - Praise, Devin, @cognition, 2026-09-11: “i used swe-2 for basically my entire workday today. i expected it to be “good for free.” nope. it’s just good. is it better than fable 5.1 or gpt-6 astra? no. but it’s much closer than i expected, and @cognition is giving it away free for the next month. probably the best deal in coding right now.” [source](https://twitter.com/1920180090572271616/status/2098420968654000429)
  - Praise, Devin, @cognition, 2026-09-01: “@embw_l0x try the 20$ plan from @cognition @devindesktop currently really good deal! i have been using the swe-1.7 model the last 4 days and its very good for extensive coding tasks and free to use, meaning 20$ buys you almost unlimited usage.” [source](https://twitter.com/2078757722363740160/status/2094896342795497875)
  - Complaint, Devin, @cognition, 2026-09-11: “@cognition curious how often the cheap executor botches what the planner meant.” [source](https://twitter.com/1604518234962657280/status/2098459133880025233)

- **Amp (stronger)**. Amp earns praise for working alongside the user, quietly fixing things and hiding git plumbing behind orbs.
  One user watched Amp fix an image overlay bug report mid-session and prompt the app to refresh. Others say orbs mean they never had to learn git worktrees, and that Amp built a home automation add-on to install itself. The weak spot is the same one users flag everywhere, which is that overnight autonomous runs come back off target.
  Evidence:
  - Praise, Amp, @AmpCode, 2026-09-21: “@ampcode did you just fix the image overlay bug report while i was working, and prompted the app for a refresh? 😀” [source](https://twitter.com/2075289824915791872/status/2102118025834922257)
  - Praise, Amp, @AmpCode, 2026-09-27: “first i didn't know how to use git worktrees then i had @ampcode orbs and didn't need to know how to use git worktrees.” [source](https://twitter.com/89691524/status/2104148649160962391)
  - Praise, Amp, @AmpCode, 2026-09-24: “inspired by @thorstenball &amp; @sqs on season 2 of "raising an agent", i decided to add @ampcode `--no-tui` runners to various ad-hoc local hosts running around the house. asking amp to build a home assistant add-on to install itself worked great! <strict_link>” [source](https://twitter.com/7566122/status/2103110675589664926)
  - Complaint, Amp, @AmpCode, 2026-09-24: “i completely agree, the ui/ux is actually very good, and i think the multi-model advantage is best utilized by amp. why? of course, it's because it supports byok, unlike cursor (which feels like it wasn't made with care). amp's byok is treated like a favored child. i also really like orbs that allow me to switch between different devices. i often use my phone in bed when i can't sleep to come up with new ideas and then let amp write/edit plans while i sleep, but the results are not ideal, as i mentioned above. so that's why i think i might need to switch back to claude code.” [source](https://twitter.com/1592160489965948933/status/2103112783663673590)

- **Zed (weaker)**. Zed's agent workflow asks the user to supervise it, from monitoring subagents to routing basic actions through chat.
  A user running subagents says they have to keep telling the main thread to monitor them. Another dislikes that everything goes through the agent, including commits, with no obvious button, and that it needs learning before it pays off. Notebook support draws repeated complaints. Fans still praise its orchestration and its forced-subagent setup, so the gap is polish, not ambition.
  Evidence:
  - Complaint, Zed, @zeddotdev, 2026-09-17: “@zeddotdev hi, i didn't find a repo to file a bug on delta, but using it with subagents i have to constantly write the main trhread to monitor them.” [source](https://twitter.com/8217762/status/2100399834356150550)
  - Complaint, Zed, r/ZedEditor, 2026-09-16: “i played around with it, but i don’t like that everything seems to have to go through the agent. need to commit some changes? you have to tell the agent to do it - there’s no obvious button for that, or at least i couldn’t find one. more broadly, it feels like something i’d need to learn first before i could use it effectively. i’d prefer something i could replace my current setup with immediately, then gradually discover the productivity improvements it offers as i use it.” [source](https://www.reddit.com/r/ZedEditor/comments/1whydrw/delta_in_public_beta/pa6e00j/)
  - Praise, Zed, @zeddotdev, 2026-09-24: “agent orchestration is different level in @zeddotdev <strict_link>” [source](https://twitter.com/1261173216455712768/status/2103200605493928281)
  - Praise, Zed, r/PiCodingAgent, 2026-09-21: “my personal huge level up was going from cursor ide to zed + omp. it is more efficient, i have everything i could've asked for and more. i love the custom fallbacks. being able to force the use of subagents. and the advisor... the advisor is really something tbh” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wlpj30/lord_forgive_me_for_the_time_i_wasted/pb4z6kj/)

- **Claude Code (mixed)**. Claude Code draws the strongest capability praise and the most requests to talk less and refuse less.
  Users report one-shotting features that pass audit and pushing many parallel sessions with subagents. The same crowd leads requests for shorter replies and for fewer false-positive safety blocks on benign tasks. Complaints describe sprawl across files and stalls on larger greenfield work. Strong output, heavy friction around it.
  Evidence:
  - Praise, Claude Code, r/ClaudeCode, 2026-09-27: “it normally takes me several back and forths to implement a feature properly. for the first time, with opus 5.5, i oneshotted a feature, audited and found no issues whatsoever. it's a goated model and i truly believe it'll be a pioneer agent for what's to come.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrosog/opus_55_is_the_first_model_that_consistently/pcfardh/)
  - Praise, Claude Code, @ClaudeDevs, 2026-09-24: “ok, had to push hard, but finally hit the session limit, running 8 claude code sessions on opus 5.5 max and one session firing opus 5.5 subagents 3-7; limited out at 4 min to refresh. that's impressive throuhgput, imo @claudedevs thank you!” [source](https://twitter.com/154635655/status/2103119113166221540)
  - Complaint, Claude Code, r/ClaudeCode, 2026-08-31: “that is not new development tho, that is refactoring an existing codebase. which they're better at. i'm going based on my specific taste, as a computer programmer of > 30 years. i've built and maintained a lot of systems in my life. so i gave claude a chance - i told him to write a simple fps indicator for my game, and to make it client side only. what he came up with was not done with meaningful variable names, and was spattered across 5 different files! i was wondering what was up with claude. so i gave him some blue sky development work, and watched. he only got to about 1600 lines, before he stalled out and couldn't progress. he used to be good, on july 17th. he was the best coder i have ever seen. but what is currently out, i would call a programming newbie. someone who has been in the field maybe a year.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w2fd10/does_anyone_think_the_rise_of_llm_coding_will/p6yetrc/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-13: “i agree with the overengineering part. but for this i had to give it instructions / build guardrails. other than that it was highly intelligent and actually did everything i threw at it.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wf5pwg/gpt_6_hype_is_fake/p9j76s1/)

### Fine print

- Many posts judge a model running inside a harness, so credit or blame can belong to the model rather than the agent.
- Warp, Grok Build and Augment Code have too few posts here to rate.
- Request counts are small next to total rated volume; read them as direction, not magnitude.

## Top requests

What users ask to add or change, most asked first. 2100 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Criterion | Author-weeks | Posts | Agents (author-weeks) |
|---|---|---|---|---|---|
| 1 | Shorter, less verbose responses | [Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md) | 47 | 48 | Claude Code 38, OpenAI Codex 5, Google Antigravity 2, Devin 1, Pi 1 |
| 2 | Built-in computer use capability | [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | 45 | 46 | Google Antigravity 13, OpenAI Codex 12, Devin 7, Cursor 6, Cline 3, Amp 1, Claude Code 1, GitHub Copilot 1, OpenCode 1 |
| 3 | Fewer false-positive safety blocks on benign tasks | [Safety filters block legitimate coding tasks](https://feedbackbench.com/criteria/work.safety_refusals.md) | 44 | 50 | Claude Code 36, OpenAI Codex 7, Cursor 1 |
| 4 | Built-in multi-agent orchestrator mode | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | 37 | 37 | Google Antigravity 10, OpenAI Codex 9, Pi 6, OpenCode 4, Claude Code 3, Cursor 2, Warp 2, Devin 1 |
| 5 | Fewer permission prompts overall | [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | 37 | 37 | Google Antigravity 16, Claude Code 8, OpenAI Codex 7, Cursor 3, OpenCode 3 |
| 6 | Auto-approve mode without permission prompts | [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | 34 | 35 | Google Antigravity 25, OpenAI Codex 4, Claude Code 3, Cline 1, OpenCode 1 |
| 7 | Allow legitimate cybersecurity and pentesting work | [Safety filters block legitimate coding tasks](https://feedbackbench.com/criteria/work.safety_refusals.md) | 33 | 35 | Claude Code 18, OpenAI Codex 13, Google Antigravity 1, Kiro 1 |
| 8 | Event-driven subagent completion instead of polling | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | 29 | 31 | OpenAI Codex 24, Claude Code 4, OpenCode 1 |
| 9 | Always-allow approvals that persist and work | [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | 28 | 28 | Google Antigravity 11, Claude Code 6, OpenAI Codex 6, Cursor 3, OpenCode 2 |
| 10 | Finish tasks fully without stopping early | [Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md) | 28 | 28 | OpenAI Codex 14, Claude Code 8, OpenCode 3, Cursor 2, Devin 1 |
| 11 | Bypass and full-access modes honored | [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | 27 | 27 | Google Antigravity 10, Claude Code 10, OpenAI Codex 4, Cline 1, Devin 1, Kiro 1 |
| 12 | Fix agents looping endlessly without progress | [Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md) | 26 | 27 | Google Antigravity 10, OpenAI Codex 6, Claude Code 4, OpenCode 3, Devin 2, Cursor 1 |

### 1. Shorter, less verbose responses

- Claude Code, 2026-09-25, @ClaudeDevs (X): “@claudedevs the usage limits are great currently with opus 5.5 but they would go even further if it wouldn't dump a novel full of claude-speak at me in every answer. you need to get that verbosity under control!” [source](https://twitter.com/1823237976295383041/status/2103604406109266061)
- OpenAI Codex, 2026-09-24, r/codex (Reddit): “i'd pay max tokens for a gpt-6 stfu model. i dont want to read a novel for a simple question. i dont want it to tell me 'youre right' 5000 times. i dont want it to suggest things i didnt explicitly ask. seriously, shut tf up!” [source](https://www.reddit.com/r/codex/comments/1wp2bov/chatgpt_manipulates_you_to_keep_chatting/pbrrfs0/)
- Claude Code, 2026-09-24, @ClaudeDevs (X): “@claudedevs but can it reply with a simple yes or no and avoid dissertation length responses….nope. context would be far lighter if it could manage this one simple thing” [source](https://twitter.com/1959700170771144704/status/2103172816732660155)

### 2. Built-in computer use capability

- OpenAI Codex, 2026-09-24, r/google_antigravity (Reddit): “is there any skill, extension or plugin like codex computer use? like getting it to control pc at its own and complete the job. i'm surprise this feature isn't in antigravity yet since been focus on multi-model and interaction? <strict_link>” [source](https://www.reddit.com/r/google_antigravity/comments/1wotyua/computer_use_in_antigravity/)
- Cline, 2026-09-24, @cline (X): “@cline browser automation with a built-in browser? computer use? when can we expect that” [source](https://twitter.com/1696542879735222272/status/2103056296488407298)
- Devin, 2026-09-23, @cognition (X): “i've been an @cursor_ai user since feb 2024, but after grok 4.7, i'm looking for alternatives. @droid @cognition are in the lead for me, but what i really need is 1. cloud agents/desktop 2. agnostic harness &amp; computer use 3. mobile 4. good connectors anyone have suggestions?” [source](https://twitter.com/1902193987244408832/status/2102596587046531556)

### 3. Fewer false-positive safety blocks on benign tasks

- Claude Code, 2026-09-27, r/ClaudeCode (Reddit): “what are you labeling as safe or unsafe? the cases i'm most interested in are skills that legitimately need file or network access but look suspicious to a scanner. that's where our false positives hurt.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr874k/has_anyone_found_a_reliable_way_to_scan_agent/pcbap0x/)
- Claude Code, 2026-09-25, r/ClaudeCode (Reddit): “opus 5.5 refuses to use the 1password cli (op cli) tool even though anthropic sent an email advertising the integration. the are so many safeguards that it makes the model that makes it useless for most sysadmin tasks (won't initiate a ssh connection or use a command that requires sudo). great coding model but damn, it has some huge weaknesses and almost all of them related to overly aggressive safeguards.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wpqqoy/new_55_safe_guards_are_a_joke/pc14vdn/)
- Claude Code, 2026-09-25, r/ClaudeCode (Reddit): “it was unable to correct my forgetting to prefix an api key with "sk-" because it got flagged as credential hunting” [source](https://www.reddit.com/r/ClaudeCode/comments/1wpqqoy/new_55_safe_guards_are_a_joke/pc03ja5/)

### 4. Built-in multi-agent orchestrator mode

- Google Antigravity, 2026-09-27, @antigravity (X): “@antigravity lotta hate for this lol. it is definitely puzzling that the product is moving so slowly. still has /teamwork-preview command required to get a decent agentic team involved in changes, something that should be automatically invoked and scaled appropriate to the task” [source](https://twitter.com/805587288/status/2104295109680410783)
- Claude Code, 2026-09-24, @ClaudeDevs (X): “@claudedevs old-style projects were not so useful so i’d be happy to nuke them if i can get the new project orchestration?” [source](https://twitter.com/1866829794790301696/status/2103124822905544902)
- Google Antigravity, 2026-09-24, @antigravity (X): “@abdullahformuli @ash_twtz @antigravity @ycombinator antigravity is not best ide, maybe good but not best - it has its own architecture - orchestration limits with no multi subagent work and it limits users not to use their models outside antigravity - thats p*ssy move they are just lazy and constrained” [source](https://twitter.com/1920552724481163264/status/2103003169521598750)

### 5. Fewer permission prompts overall

- OpenAI Codex, 2026-09-26, r/codex (Reddit): “it's stupid as well asking permissions to this and that, anything but to proceed, when i have clearly said fix everything lol.” [source](https://www.reddit.com/r/codex/comments/1sr8j8b/codex_keeps_stopping_every_30_to_45_seconds_and/pc4iwfw/)
- Cursor, 2026-09-24, @cursor_ai (X): “@cursor_ai i don't understand how i can stop getting buried in approval requests! <strict_link>” [source](https://twitter.com/905512466339287040/status/2103098476464595397)
- Claude Code, 2026-09-24, @ClaudeDevs (X): “@claudedevs estou gostando de trabalhar na nuvem, mas precisa melhorar as permissões. ter que ficar mandando mensagem pra pedir pro "claude computador" fazer trava a produção.” [source](https://twitter.com/3775511057/status/2102942524520227249)

### 6. Auto-approve mode without permission prompts

- Google Antigravity, 2026-09-26, r/google_antigravity (Reddit): “it’s shit if you use headless mode, it cannot have auto-approve. i had to use tmux to have an active interactive session if i want auto-approve, otherwise it asks for permission for every single step😅. but yea, remote-control is so good, it’s great that it’s an pwa, no apps required. (as long as you don’t mind having your session data in the cloud, personal plan doesn’t have zdr anyways)” [source](https://www.reddit.com/r/google_antigravity/comments/1wqmd4n/why_is_the_antigravitycli_so_underrated/pc5zqs8/)
- Google Antigravity, 2026-09-26, r/google_antigravity (Reddit): “but that’s not what i am asking for though. what i am saying is that there should be a similar command to /yolo from codex in agy-cli. basically a temporary one prompt —dangerously-skip-permissions and when the prompt is finished processing it goes to defaults. none current options do this, you either set it for an entire session or globally.” [source](https://www.reddit.com/r/google_antigravity/comments/1wqbyzk/antigravity_cli_release_v127_v1211/pc4vris/)
- OpenAI Codex, 2026-09-24, r/codex (Reddit): “starting codex for the first time [manually approving codex prompts \(ai generated image\)](<strict_link>) today i started codex for the first time (i used a lot claude code) and directly launch a fleet of agents and it was a bad idea.... i expected at least to see the "auto-mode" equivalent in the tui. now i am pressing "approve" like back in 2025.” [source](https://www.reddit.com/r/codex/comments/1wpehp7/starting_codex_for_the_first_time/)

### 7. Allow legitimate cybersecurity and pentesting work

- Claude Code, 2026-09-27, r/ClaudeCode (Reddit): “let’s put it this way, i’m building a cybersecurity/pentesting harness. opus5, 5.5, and fable, can’t so much as read the prd without tripping and downgrading to 4.8. i use hindsight as a memory system, they can’t read the description of the odin (name of my harness) bank without throwing a warning.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wppds2/is_it_safe_to_development_a_hacking_game_with/pcb4dzi/)
- Claude Code, 2026-09-25, r/ClaudeCode (Reddit): “the flag usually sits on the skill file, not the task. if the description or body mentions auth, credentials, exploit, pentest, rls, anything that reads as offensive security, it fires before the skill does any work. clearing the session only helps until it reads the file again. rewording that description into plain build language is what stopped it on mine.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wpqqoy/new_55_safe_guards_are_a_joke/pbzdhjo/)
- Claude Code, 2026-09-25, @ClaudeDevs (X): “@anthropicai @claudeai @claudedevs false-positive guardrails are completely broken. use standard terms like "password ,key, flag", or "security" and your entire session gets flagged. legitimate defensive security work is impossible right now. at $250/month, this is unacceptable <strict_link> <strict_link>” [source](https://twitter.com/1256139910504988672/status/2103381836617445869)

### 8. Event-driven subagent completion instead of polling

- OpenAI Codex, 2026-09-27, r/codex (Reddit): “i also came up with a patch, no bugs and only cost 5 tokens! don’t poll subagents” [source](https://www.reddit.com/r/codex/comments/1wrw3wy/1000_lines_348_bn_tok_393_subagents_42_pro20/pch38sj/)
- Claude Code, 2026-09-24, r/ClaudeCode (Reddit): “don't set active polling of my simulation runs with the multiagent team that go overnight...” [source](https://www.reddit.com/r/ClaudeCode/comments/1wp23t8/what_rule_in_your_claudemd_clearly_has_a_backstory/pbskfgc/)
- OpenCode, 2026-09-18, @opencode (X): “@thdxr @opencode can the opencode without waiting for the subagent to complete its task and continue working? like claude.” [source](https://twitter.com/1321164857278894082/status/2100757699822551181)

### 9. Always-allow approvals that persist and work

- Google Antigravity, 2026-09-26, @antigravity (X): “@rodydavis @antigravity the constantly asking for the same class of permission you've already approved several times over going in circles, it doesn't complete the work given only does like 60% and it's not even a oneshot prompt” [source](https://twitter.com/1895520405885952000/status/2103773104593945021)
- Google Antigravity, 2026-09-26, @antigravity (X): “@antigravity why can't you even handle the most basic command-line authorization? i've clearly authorized all git commands and other commands for the entire project, yet i keep having to authorize everything again and again” [source](https://twitter.com/1559895451805286401/status/2103657920911327482)
- OpenAI Codex, 2026-09-25, r/google_antigravity (Reddit): “am i the only one for whom "antigravity" never actually "always proceeds"? no matter what i change in the workspace settings, it always asks for countless confirmations; even when i select the option to always allow access for the folder or project, it keeps asking the same questions. does anyone know how to configure this, or is there an update that includes an "always proceed" feature like in codex or claude code?” [source](https://www.reddit.com/r/google_antigravity/comments/1wpkgy1/why_always_proceed_never_work/)

### 10. Finish tasks fully without stopping early

- Claude Code, 2026-09-26, @ClaudeDevs (X): “@claudedevs debria hacer que acabe de minimo la tarea completa 😏” [source](https://twitter.com/1800068778564169728/status/2103780713741144350)
- OpenCode, 2026-09-24, r/opencode (Reddit): “i can only use claude or gpt models at my work, so more persistent models aren't really an option. in my experience, claude is the more naturally persistent of the two, but both of them keep stopping before full task completion outside of relatively small things.” [source](https://www.reddit.com/r/opencode/comments/1wp3bc9/what_are_loopgoal_plugin_are_you_using/pbsgka1/)
- OpenAI Codex, 2026-09-24, r/codex (Reddit): “i 100% feel the fact that it needs more steering and randomly stops without really completing the task” [source](https://www.reddit.com/r/codex/comments/1worwfr/sol_6_was_insufferable_glad_to_be_back_to_56/pbpl3i8/)

### 11. Bypass and full-access modes honored

- Google Antigravity, 2026-09-26, r/google_antigravity (Reddit): “those don't work on windows. it always asks regardless. you can throw it into turbo mode and it'll still ask. <strict_link>” [source](https://www.reddit.com/r/google_antigravity/comments/1wpkgy1/why_always_proceed_never_work/pc5jqpu/)
- Claude Code, 2026-09-26, @ClaudeDevs (X): “@claudedevs i'm already in bypass permissions / auto mode, just allow this shit stop asking me <strict_link>” [source](https://twitter.com/212418463/status/2103956882339594574)
- Google Antigravity, 2026-09-20, @antigravity (X): “@antigravity for windows is unusable. even in the cli. an endless stream of permission requests even with --mode accept-edits on. every single tool use. just going to stop here and cancel it. i like the new 3.8 flash model, its honestly underrated, but its just not user friendly.” [source](https://twitter.com/402285185/status/2101494705066500569)

### 12. Fix agents looping endlessly without progress

- Google Antigravity, 2026-09-26, @antigravity (X): “@antigravity @geminiapp @googleaistudio your gemini flash 3.8 on medium reasoning. it keeps going like this forever, you need to dix this behaviour. <strict_link>” [source](https://twitter.com/1493151747413524480/status/2103897230230880585)
- Claude Code, 2026-09-24, @ClaudeDevs (X): “@claudedevs i just filed this about projects - i would appreciate a fix asap - <strict_link> potentially a powerful feature but auto mode = aut-no progress” [source](https://twitter.com/15527674/status/2103146479506633025)
- OpenAI Codex, 2026-09-23, r/codex (Reddit): “just wait till you see the loop over and over and over. that’s all luna 6 does is loop. sucks because i changed everything over to luna 6 and now i’ve wasted 15% of my weekly tokens 😭” [source](https://www.reddit.com/r/codex/comments/1wnmdz6/first_impression_of_sol6_fast_cheap_and_shitty/pbgyhxj/)

## Every agent

| Agent | Overall rank | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|---|
| [Devin](https://feedbackbench.com/agents/devin.md) | =5 | Better than peers | 0.602 | 0.574–0.630 | 620 | 412 | 208 |
| [Amp](https://feedbackbench.com/agents/amp.md) | =11 | Better than peers | 0.597 | 0.566–0.627 | 158 | 119 | 39 |
| [Factory](https://feedbackbench.com/agents/factory.md) | =11 | Better than peers | 0.577 | 0.548–0.606 | 95 | 74 | 21 |
| [Pi](https://feedbackbench.com/agents/pi.md) | 7 | Better than peers | 0.554 | 0.520–0.587 | 328 | 188 | 140 |
| [Cursor](https://feedbackbench.com/agents/cursor.md) | 4 | Better than peers | 0.533 | 0.511–0.554 | 1252 | 674 | 578 |
| [Conductor](https://feedbackbench.com/agents/conductor.md) | 14 | Typical | 0.520 | 0.492–0.545 | 43 | 27 | 16 |
| [Cline](https://feedbackbench.com/agents/cline.md) | =8 | Typical | 0.510 | 0.476–0.546 | 140 | 76 | 64 |
| [Claude Code](https://feedbackbench.com/agents/claude-code.md) | 1 | Typical | 0.503 | 0.493–0.514 | 5831 | 2737 | 3094 |
| [OpenAI Codex](https://feedbackbench.com/agents/codex.md) | 2 | Typical | 0.501 | 0.492–0.510 | 6815 | 3167 | 3648 |
| [OpenCode](https://feedbackbench.com/agents/opencode.md) | 3 | Typical | 0.500 | 0.477–0.524 | 1369 | 667 | 702 |
| [Kiro](https://feedbackbench.com/agents/kiro.md) | 13 | Typical | 0.496 | 0.468–0.525 | 60 | 27 | 33 |
| [GitHub Copilot](https://feedbackbench.com/agents/copilot.md) | =8 | Typical | 0.495 | 0.459–0.533 | 244 | 110 | 134 |
| [Zed](https://feedbackbench.com/agents/zed.md) | =8 | Worse than peers | 0.455 | 0.426–0.489 | 102 | 38 | 64 |
| [Google Antigravity](https://feedbackbench.com/agents/antigravity.md) | =5 | Worse than peers | 0.378 | 0.359–0.398 | 1723 | 606 | 1117 |
| [Warp](https://feedbackbench.com/agents/warp.md) | 15 | Too few posts | – | – | 17 | 9 | 8 |
| [Grok Build](https://feedbackbench.com/agents/grok-build.md) | 16 | Too few posts | – | – | 14 | 9 | 5 |
| [Augment Code](https://feedbackbench.com/agents/augment.md) | 17 | Too few posts | – | – | 7 | 3 | 4 |

## Posts

Receipts rule: The 5 most recent praise and complaint posts per area (first 700 characters) and 3 per criterion (first 450 characters).

### Devin

- Praise, 2026-09-27, r/windsurf (Reddit): “devin has invested a lot in controlling how the ais think inside the ide, to maximize context and code management tools. if you are going to pay for devin, try the devin desktop ide. it's a fork of vs code so it's not unfamiliar. the ide is good for almost any workflow, from lightning-fast code completion to highly reliable vibe coding with swe-2.” [source](https://www.reddit.com/r/windsurf/comments/1wr7p7l/why_isnt_windsurf_available_in_vscode_anymore/pcb72fn/)
- Praise, 2026-09-27, r/windsurf (Reddit): “they are not in the same class of capability. copilot is not a very good harness” [source](https://www.reddit.com/r/windsurf/comments/1wr7p7l/why_isnt_windsurf_available_in_vscode_anymore/pcd4mf5/)
- Praise, 2026-09-27, @cognition (X): “if you have claude subscription and @cognition cloud agents. try it it is amazing combo, it ships like crazy when you sleep. @cognition swe 2 max is a really good model and on cloud, it has macos, can code, run test, take screenshot and record evidence. opus 5.5 is really well as orchestrator, review and merge pr on your machine. you can feed opus (in claude code or hermes) devin api key, it knows what to do.” [source](https://twitter.com/1767985295910383616/status/2104014647724761554)
- Praise, 2026-09-27, @cognition (X): “been 96 hours. @cognition @devindesktop i've never seen a harness as good as this. they don't even have the /goal, but their execution surpass any harness with /goal on the market. it's f*cking mind blowing. 20$, unlimited swe 2 you should definitely give it a shot <strict_link> <strict_link>” [source](https://twitter.com/1825243355501973504/status/2104016524725952607)
- Praise, 2026-09-27, @cognition (X): “@max18martin @cognition i have to say that swe-2 can even be used in many scenarios to match gpt-6 sol, and the former is surprisingly free for pro plans and above.” [source](https://twitter.com/1859530427905736704/status/2104025642329330144)
- Complaint, 2026-09-27, @cognition (X): “another task going to be 24 hours! @devinai @devindesktop @cognition you guys are cooking my projects!!! unlimited swe 2 with a 20usd plan, are you kidding me? that is literally the best coding plan in the world <strict_link>” [source](https://twitter.com/1825243355501973504/status/2104262892090781744)
- Complaint, 2026-09-27, @DevinAI (X): “@doodlestein @hraness @chatgpt @devinai @claudeai @cursor_ai @bot @zeddotdev the part i’d watch is observability. more terminals only helps if you can see which run stalled, what changed, and whether the result is safe to merge. otherwise it’s just a very expensive wall of tabs.” [source](https://twitter.com/2103910196330242049/status/2104140047633326358)
- Complaint, 2026-09-27, @DevinAI (X): “@doodlestein @hraness @chatgpt @devinai @claudeai @cursor_ai @bot @zeddotdev sixty-four accounts. at that point the agents are managing you, not the other way around.” [source](https://twitter.com/2087402808756629504/status/2104147885679845800)
- Complaint, 2026-09-27, @DevinAI (X): “@hraness @chatgpt @devinai @claudeai @cursor_ai @bot @zeddotdev claude code's main gap: unlike opencode, it has no central view of every agent across all sessions and repos.” [source](https://twitter.com/3251926098/status/2104247081237967027)
- Complaint, 2026-09-27, @DevinAI (X): “@fluyeporlaweb @hraness @chatgpt @devinai it seems to me like a bug factory... can you really exercise human quality control there and leave everything to the ai? i don't know, rick, it seems false to me. also, we should look at the numbers of how much income those subscriptions generate against what they cost.” [source](https://twitter.com/1913354152941432832/status/2104335120802943329)

### Amp

- Praise, 2026-09-27, @AmpCode (X): “first i didn't know how to use git worktrees then i had @ampcode orbs and didn't need to know how to use git worktrees.” [source](https://twitter.com/89691524/status/2104148649160962391)
- Praise, 2026-09-26, @AmpCode (X): “@simonbs @ampcode its good, right? the other good part is running threads in orbs (ad-hoc vms) and i think it follows from there, that llm provider stuff is mainly handled at a level above each node you can ask the agent to hand off work to any local cli and ask it to use your shared tmux session” [source](https://twitter.com/132882990/status/2103833633802842583)
- Praise, 2026-09-26, @AmpCode (X): “i just did something that i wasn't aware it was possible using @ampcode . i had an idea on how to improve bug fixes. since everything now runs on an orb (their vm system), i thought: "it would be so cool if amp was my "support team""" so, i asked it to build. since we have a very good tracing system, we can check pretty much everything the user do, so we added an "send bug report" (the user allows us to see the last 10 minutes of his actions, there's a checkbox for that) and when he sends, it automatically creates an github issue, that by itself spawns an orb that starts checking what happened right away, if it needs an fix, it creates it and submit as a pr (and link it to the github issue)” [source](https://twitter.com/2931128860/status/2103847911725404390)
- Praise, 2026-09-26, @AmpCode (X): “amazing 🤩 for more context: have a fleet skill that let's the agents relay my home devices through raspberry pi when they need to test things on devices with no reliable simulator runtime like roku, samsung tizen & lg webos tvs 🫠 a bit niche use-case but think it'd be useful for mobile as well... only reason i got a mac-mini instead of linux devbox is apple simulators 😔” [source](https://twitter.com/1125366224664322049/status/2103858145101500803)
- Praise, 2026-09-26, @AmpCode (X): “@homborg @ampcode ah, got it. so if i use my own machine as a runner, there are no orbs in play, but i can delegate to either my runner or orbs. that makes sense. i really like the flexibility of amp.” [source](https://twitter.com/36411940/status/2103873967081877720)
- Complaint, 2026-09-26, @AmpCode (X): “@thorstenball @ampcode i have a skill to address pr feedback and the model usually asks me in text “may i continue?” seems like that’s a perfect place to give the user a ui element” [source](https://twitter.com/5444392/status/2103705622650978634)
- Complaint, 2026-09-26, @AmpCode (X): “@thorstenball @ampcode probably! curious that i never see the modal use the choice tool is all” [source](https://twitter.com/5444392/status/2103706972151587097)
- Complaint, 2026-09-26, @AmpCode (X): “anyone using @ampcode to write blog posts? getting very mixed results so far so curious about your process.” [source](https://twitter.com/881577286927036416/status/2103824362746859692)
- Complaint, 2026-09-25, @AmpCode (X): “@jkudish @ampcode glm models are like gemini models. they just exist. they have claims, and they don’t work in the real world.” [source](https://twitter.com/137626804/status/2103381866736755000)
- Complaint, 2026-09-25, @AmpCode (X): “@rom1_pellerin @cognition @ampcode @omnigent_ai @superset_sh @paulgauthier agree. multi-agent without a final decision owner just multiplies confident wrong answers. someone has to ship the call.” [source](https://twitter.com/2092975354138804224/status/2103383441173590219)

### Factory

- Praise, 2026-09-27, @FactoryAI (X): “@render @factoryai agent writes the app, spins the db, deploys it. i just sit there like a decorative readme” [source](https://twitter.com/1330209814790746114/status/2104177796901945611)
- Praise, 2026-09-27, @FactoryAI (X): “@factoryai @fireworksai_hq huge win for legacy code blind spots there are so real” [source](https://twitter.com/2010658787611619328/status/2104199345436312044)
- Praise, 2026-09-27, @FactoryAI (X): “@factoryai @enoreyes agent foxxy just aced its browser test by autonomously uploading a video to youtube—acting just like a human! 🤖🔥 <strict_link> <strict_link>” [source](https://twitter.com/2093686029186396160/status/2104206958949810353)
- Praise, 2026-09-27, @FactoryAI (X): “@droid @factoryai droid is genuinely useful for larger, multi-file tasks and does a good job staying on track without constant guidance. the biggest improvement for me would be better visibility into its reasoning/progress and more predictable results on longer tasks :)” [source](https://twitter.com/2093736525116702720/status/2104324409502970188)
- Praise, 2026-09-27, @droid (X): “herdr projects plugin with claude as the coordinator and @droid as the threads &gt;&gt;&gt; droid is such a cracked harness its an amazing execution engine” [source](https://twitter.com/1953226278452256768/status/2104187978419732696)
- Complaint, 2026-09-27, @droid (X): “@droid at least it should be on par with devin” [source](https://twitter.com/1159835302275346433/status/2104318568980689170)
- Complaint, 2026-09-26, @FactoryAI (X): “@hataiit9x @droid @factoryai the harness quite sucks :)” [source](https://twitter.com/1797536317716525056/status/2103754025057595878)
- Complaint, 2026-09-24, @FactoryAI (X): “@factoryai @fireworksai_hq anyone who has migrated old cobol sees it: the model "improves" three lines nobody asked for. boring diff wins.” [source](https://twitter.com/2030549621039349760/status/2103235553651282146)
- Complaint, 2026-09-24, @droid (X): “@droid droid still stuck in flutter development, whenever it’s try to run any dart mcp tools, it’s stuck in never ending loop, and unlike amp and pi or opencode it can’t even auto run adb command for debugging something in that.” [source](https://twitter.com/2995471962/status/2102910997338218965)
- Complaint, 2026-09-24, @droid (X): “@droid @tastelabs any plan to add the support the run in background feature just like claude code?” [source](https://twitter.com/2018347199432994816/status/2103150290593796353)

### Pi

- Praise, 2026-09-27, r/PiCodingAgent (Reddit): “my orchestrator does nothing except for delegating and passing messages. i've had builds running for almost 24 hours and the orchestrator at the end is still below 200k context. for long builds it's super efficient.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wo6lr8/how_to_better_enforce_subagent_delegation/pcatr9y/)
- Praise, 2026-09-27, r/PiCodingAgent (Reddit): “yes i know the system prompt is big. you said nothing about capabilities. have you actually tried pi with all the added features? things like subagents, memory, worktrees, and many others are all built in, and you need to add these to pi. try to be objective and compare the same things” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrjr75/the_good_opensource_harness/pcgpxe3/)
- Praise, 2026-09-27, r/PiCodingAgent (Reddit): “i use pi normaly i was trying out opus 5.5 today on claude code and it feels awful vs pi with /tree and all my custom stuff now im looking at the sdk claude pi extension, having model edit and reduce token bloat, hoping it won't get me banned” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrjr75/the_good_opensource_harness/pcgqytv/)
- Praise, 2026-09-27, r/PiCodingAgent (Reddit): “why not both? i have a orchestrator mode extension in pi, where the system prompt instructs the model to update a plan and ledger md in scratch workspace, it can compact how/when it wants, because all the key findings and planned tasks are always tracked in scratch workspace” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrq50c/new_session_handoff_vs_compact_which_do_you_prefer/pcguntn/)
- Praise, 2026-09-27, @pidotdev (X): “pro move: install @pidotdev - web, tell @nousresearch hermes its there and let hermes talk directly to pi and do the work. i just get a ping when its done it needing a review. working great building <strict_link>” [source](https://twitter.com/15451416/status/2104052502727659701)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “the codebase is very well documented, even using a code mapper to save on the scanning, the prompts were crystal clear, with the right context being provided, and still it went off and reasoned about it for a huge amount of time. one of the tests were made in little coder and the harness even tried to tell the model to stop thinking and implement as it already had the solution, but to no avail, it continued on oblivious. i'm using fp8 so not even very heavily quantized.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wnmsho/i_want_to_believe_in_local_llms_for_coding_but/pcbw9u1/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “yeah. on all my tests accuracy was quite shit, and jev was consistently outperformed by glm 5.3 flash, for anything requiring decision making. but hey, it's fast 🤷♂️” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wnl7da/pi_can_now_use_jev_and_more/pce99cd/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “pretty much it. even if the system prompt says they have to delegate, it usually tries and fails once in order to realize it really has to delegate. i guess you could instruct it to use a classifier model to determine if it needs to delegate and to which agent.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrte31/alternative_to_tool_profiles_for_better_subagent/pcfl1l2/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “will be trying your prompts. looks pretty much like what i do, after a lot of frustration from watching my agents behave similar to op's. i just started telling whatever we were supposed to do, and then saying "you will not execute the tasks yourself. you will delegate to sub-agents for execution, and whatever other steps may be appropriate." lol.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wo6lr8/how_to_better_enforce_subagent_delegation/pch06nf/)
- Complaint, 2026-09-27, r/LocalLLaMA (Reddit): “another "harness matters" post (codex cli > pi and opencode) i run my own llm while also having a openai subscription. also tried deepseek (latest flash now). i run [qwen 3.8 flash next](<strict_link>) at an amazing speed on my 2x3090 + ram! but local llm never did worked for me outside some demos like build me a "3d mario game, multistage" which i've been using to test llm's for a long time. at serios work, they never even compared with gpt 5.2 or lately, 5.6 luna, which is worse in the benchmarks. until last night! i've asked gpt 5.6 luna to configure codex cli for local llm! (i've been using [pi.dev](<strict_link>) and opencode until now) and the results amazed me! suddenly aa benchmark” [source](https://www.reddit.com/r/LocalLLaMA/comments/1wrfp50/another_harness_matters_post_codex_cli_pi_and/)

### Cursor

- Praise, 2026-09-27, r/cursor (Reddit): “right there with you - not an elon fan, but have had grok 4.6 and now 4.7 as my everyday driver since composer 2.5. all entirely competent models.” [source](https://www.reddit.com/r/cursor/comments/1wrktgc/am_i_the_only_one_who_thinks_grok_47_is_actually/pcdcwpu/)
- Praise, 2026-09-27, r/cursor (Reddit): “yep 👍🏼. actually i used agent i needed that actually to run autonomous task till 2days constant...” [source](https://www.reddit.com/r/cursor/comments/1wrkzi5/am_i_cooked/pce76ru/)
- Praise, 2026-09-27, r/cursor (Reddit): “did some autonomous tasks...(2days constant workout with 8-9 agents running)” [source](https://www.reddit.com/r/cursor/comments/1wrkzi5/am_i_cooked/pce8fh2/)
- Praise, 2026-09-27, r/cursor (Reddit): “i hate to say it but grok 4.7 is the best for my work flow because i can do everything in grok 4.7 high fast mode. and the reason i hate to say it is because grok has been used for disgusting things and it bothers me that the richest person in the world can buy yet another company and make it their own, but i think what they've done with cursor has all been actually positive as opposed to what happened to twitter. normally i'd do a frontier claude model, if i get stuck i switch to a frontier gpt or other model, then for easier smaller tasks i'd switch to auto or composer. now, the most complicated and the most simple tasks all work well with grok 4.7 where frontier models from other brands” [source](https://www.reddit.com/r/cursor/comments/1wrktgc/am_i_the_only_one_who_thinks_grok_47_is_actually/pcf8g60/)
- Praise, 2026-09-27, r/cursor (Reddit): “its literally how senior engineers at cursor developed grok bot. lauren tan, hardly a vibe coder, merges more than 2000 prs per month using this workflow” [source](https://www.reddit.com/r/cursor/comments/1wrmig0/killer_combo_cursor_projects_pstack_huge/pcfwjra/)
- Complaint, 2026-09-27, r/cursor (Reddit): “cursor agents are fast at *retrying*. the expensive part for us was retrying the **same** fail — bad path, wrong toolchain, flake that already had a known fix in another session. we keep a small oss prior-art index (claimidx, apache-2.0) beside the agent: ask before grinding, apply + verify, then publish a compact claim. retrieved remedies are evidence for the model, not auto-executed patches. if you want the full loop (terminal step is share): ```bash pip install -u "claimidx[server]>=0.7.13" claimidx init --agent <your-cursor-agent> claimidx claim --yes --channel reddit --source path-b ``` on ≥0.7.13, `claim --yes` auto-shares to the commons; `--local` keeps it private. mcp name `claimidx-” [source](https://www.reddit.com/r/cursor/comments/1wqkwtj/your_cursor_plan_already_spins_cloud_agents_why/pc9t0l9/)
- Complaint, 2026-09-27, r/cursor (Reddit): “one-shotting is the issue. unless it's something as basic as a chrome extension, i never one-shotting.” [source](https://www.reddit.com/r/cursor/comments/1wqppuh/grok_is_shutting_down_apps_now/pcahmk0/)
- Complaint, 2026-09-27, r/cursor (Reddit): “cursor is crap compared to a pro plan on claude code.” [source](https://www.reddit.com/r/cursor/comments/1wrbmrn/is_it_a_waste_of_money_to_run_opus_55_inside_of/pcd0den/)
- Complaint, 2026-09-27, r/cursor (Reddit): “i compare it to grok 4.6 and 4.7 is still aggressively bad.” [source](https://www.reddit.com/r/cursor/comments/1wrktgc/am_i_the_only_one_who_thinks_grok_47_is_actually/pcddocg/)
- Complaint, 2026-09-27, r/cursor (Reddit): “composer does exactly what you ask it to do even if it takes a few prompts to finish grok will do it all and add 10 things i didn't ask for so i tell it i didn't ask for those things and it says 'you're right i'm so sorry' then it adds 2 other things i didn't want or it will change something that breaks everything. so you ask it to fix it. oh, so sorry, here's 2 more things you didn't ask for.” [source](https://www.reddit.com/r/cursor/comments/1wrktgc/am_i_the_only_one_who_thinks_grok_47_is_actually/pcdlxql/)

### Conductor

- Praise, 2026-09-27, @conductor_build (X): “@conductor_build now records stuff from computer mode for testing, it's awesome.” [source](https://twitter.com/1060104763885477888/status/2104227093194461548)
- Praise, 2026-09-27, r/ClaudeAI (Reddit): “i built a personal ai assistant that now answers my phone, triages my email, texts me notifications, drafts and sends documents and emails, manages my calendar, and remembers past conversations. basically i had a goal to create my own grok bot/muse before i knew those were a thing. now that those are out, i have a benchmark for baseline parity. but among other things, it also has an address book which i also use to configure how to behave depending on who is calling (e.g mom gets a female voice and very friendly and helpful assistant with more access to my personal life, while a buddy gets a guys voice and bro talk. funny thing is my buddy actually texts him and gives him shit and my assis” [source](https://www.reddit.com/r/ClaudeAI/comments/1wrckhp/what_tool_have_you_built_for_yourself_with_claude/pch0v2v/)
- Praise, 2026-09-26, @conductor_build (X): “@itsvlady it's crazy man, astra as architect and opus 5 as the henchman -- try out @conductor_build btw, it's awesome, so far the best 'multi-model' harness (supports cc, codex, cursor and opencode clis), and all your work sits in isolated worktrees so you can parallelize it to the maxx” [source](https://twitter.com/2891185809/status/2103772258376294595)
- Praise, 2026-09-26, @conductor_build (X): “@newmediums @itsvlady @conductor_build that's actually a really good one, i usually run astra as the architect, but keep it for anything that is 'verifiable', so if it has unit tests, it can really chew through bugs before creating them, but fable as the orchestrator and opus 5.5, oh man, it's so good” [source](https://twitter.com/2891185809/status/2103829576543621213)
- Praise, 2026-09-25, r/conductorbuild (Reddit): “i really like the way the projects are structured with worktrees, makes it easy to keep things organised and do more work in parallel. like others have said, the recent ui changes can be confusing at times. i think its best to stick to a default and let users configure the rest through a preferences page.” [source](https://www.reddit.com/r/conductorbuild/comments/1wp7deq/favourite_thing_in_conductor/pbxng4x/)
- Complaint, 2026-09-26, @conductor_build (X): “@kyberpez @itsvlady @conductor_build my only use case for astra is backend logic tasks because the design ability is so unbelievably poor” [source](https://twitter.com/1848821469108965376/status/2103827278740251086)
- Complaint, 2026-09-26, @conductor_build (X): “@mattgapp @capydotai @conductor_build capy does next level orchestration i’ve switched” [source](https://twitter.com/11768582/status/2103996834704466324)
- Complaint, 2026-09-21, r/conductorbuild (Reddit): “very limited functionality i believe” [source](https://www.reddit.com/r/conductorbuild/comments/1wkn35i/any_word_on_that_ios_app_question/pb8t6ss/)
- Complaint, 2026-09-20, r/conductorbuild (Reddit): “i normally use claude code's /goal command to let it run long tasks uninterruptedly because it auto-resumes it's work once the session limit is reset, but i can't find a way to instruct conductor to do the same. even if a use /goal through conductor, the effect isn't the same. is anyone aware if this is possible or in the roadmap? it'd be really helpful” [source](https://www.reddit.com/r/conductorbuild/comments/1wlrxxi/does_conductor_support_autoresuming_a_session/)
- Complaint, 2026-09-16, @conductor_build (X): “@conductor_build me too, but then i didnt like how few things were done, so i built my own ade @zuse_sh” [source](https://twitter.com/3304494590/status/2100174141320310804)

### Cline

- Praise, 2026-09-26, @cline (X): “@cline oh wow nextjs specific. been knocking my head on the wall rewriting legacy pages router to app router.” [source](https://twitter.com/326659472/status/2103645823083168184)
- Praise, 2026-09-26, @cline (X): “@cline really good! after integrating cline for free, it crushes the next.js tasks of kimi k3—pixel canary action is really fast.” [source](https://twitter.com/1744672948135321600/status/2103663198188749016)
- Praise, 2026-09-26, @cline (X): “@cline beats kimi k3, ties gpt-6 astra, costs nothing. the business model is 'we'll figure it out', which is also my business model, so i can't judge” [source](https://twitter.com/1475598443779444739/status/2103708005372239980)
- Praise, 2026-09-26, @cline (X): “@cline pixel canary tying with gpt-6 astra and being free is amazing! excited to try it out in cline.” [source](https://twitter.com/1347793590605410309/status/2103714130179899444)
- Praise, 2026-09-26, @cline (X): “@haleeeemahh @opencode @cline tried pixel canary on cline this week, it handled my messy refactor without complaint. feels too good to be a small model.” [source](https://twitter.com/2079846327744401408/status/2103837540260389092)
- Complaint, 2026-09-27, r/CLine (Reddit): “hey, first time post so bear with me if i'm doing something wrong. ill first explain i run a prompt then the model runs okay for a bit then it hits the "waiting for teammates" this isn't a issue but when i click on the sub models that are running no processing or thinking is actually being done the sub model just sits with the prompt and displays "thinking" if anyone has a solution please send” [source](https://www.reddit.com/r/CLine/comments/1wrqbkt/waiting_for_teammates_error/)
- Complaint, 2026-09-27, @cline (X): “@cline why des this keep happening please its frustrating, leaving a session coming to see its stopped. tying continue, proceeds, meaning authentcation was never an issue <strict_link>” [source](https://twitter.com/83118210/status/2104146625736106087)
- Complaint, 2026-09-27, @cline (X): “i know stealth models seem to be the in thing right now, but i'm not sure they're even worth messing around with sometimes. trying to use pixel canary on @cline, and it's just so slow. i mean, 24 hours now, no closer to the task, and it keeps stopping and starting. it's horrible!” [source](https://twitter.com/25673607/status/2104177440129991012)
- Complaint, 2026-09-26, r/CLine (Reddit): “spent 2 hours on the tasks..timeout and in a bad loop. the pass is unusable at all. not worth the 10. i wish i can cancel it and get the refund. honestly it is a lousy harness.” [source](https://www.reddit.com/r/CLine/comments/1uj0evt/does_anyone_here_have_any_experience_with_cline/pc6frz3/)
- Complaint, 2026-09-26, r/CLine (Reddit): “hello everyone , i just started using vscode +cline , just for fun , messing around with unity scripts and stuff , and all was great for a few days , during 1 task , the power went down , and after i turned on my pc i had the following problem , he just stopped executing tasks , most of the time just telling me how to do it , and sometimes replying just in code , i ve been trying to troubleshoot it for 2 days now , and this is what i found. i use different types of qwen loccally , after i reinstall any qwen , it does work if i approve mannually , but if i check the auto execute box , it breaks again…..been using it just for fun , and i dont have any experience in this sort of things . what c” [source](https://www.reddit.com/r/CLine/comments/1wqm2kn/vscode_and_cline/)

### Claude Code

- Praise, 2026-09-27, r/ClaudeCode (Reddit): “even pre-fable, this has worked well. i have a cable that connects my odb-ii port to a particle tracker one. i had claude write the message processor, filtering logic and charts and graphs. i’ve used it on a few cars and it basically backs out the dbc file and goes to town.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wquoxy/fable_51_live_vehicle_diagnostics/pc9u3qf/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “homie get some sleep, claude can work while you recharge” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqnh3q/me_after_opus_55_release/pc9w9xy/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “my engineering department is on a team plan. we also have tier 5 openai. we just run up bills like crazy but the productivity is fucking insane.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr0wic/is_everyone_here_millionaires/pca1jju/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “5.5 is really good but it does have the same flaws as all the other agents. i think one of the biggest appeals (any why most love it) is the huge reduction in its draw on subs (5x just became more like 30x if you set it on medium effort) and it speaks plain language for the most part, oh and its faster and less verbose, all while being a solid reasoning machine. its fast, cheap, and works. that makes for happy coders.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqjuv2/am_i_the_only_one_who_doesnt_like_opus_55/pca70cd/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “shotcut if youre ever looking for a fleshed our free alternative. simple editing i have claude do with ffmpeg” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr4yjz/wanted_to_edit_some_footage_of_a_game_i_was/pcabeg2/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “every time i used sonnet it got something wrong, basically any type of agentic coding, like if u want it to write a single function and know what u are doing it's fine but otherwise it's ass” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr5lzf/practically_speaking_what_tasks_fit_into/pc9u33o/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “i don't really use sonnet for anything. i switched to opus because i kept running out of usage - i found i use more with sonnet even though it's cheaper because it kept getting stuck and making mistakes.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr5lzf/practically_speaking_what_tasks_fit_into/pc9u4xg/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “pretty impossible as they have filter to not allow claude anything like this in first place.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wquoxy/fable_51_live_vehicle_diagnostics/pc9xh0g/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “i had opus 5.5 make mistakes and break things right after release, so i'm not sure "broke a couple things" is enough evidence. llms can make mistakes regardless of quantization. what you should do is have fable replay the last 20 or so prompts from before the degradation occurred with the now possibly weaker model in the exact same spot in your commit history, and have it judge the results. then come back with it's evaluation. your analysis is just vibes unless you're controlling for the exact prompt and exact context the prompt was executed in.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr67vx/i_agree_with_dario_regulating_llms/pca1k7u/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “yes, this is true -- but now it's 10x negative value given the amount of work that can be completed with a few prompts. they're only really useful if your entire workflow is heavily guarded and gated - and even then, it's still difficult.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr27zi/be_honest_do_we_still_offer_value/pca7f1w/)

### OpenAI Codex

- Praise, 2026-09-27, r/codex (Reddit): “hard disagree, i’ve had problems that have been acting as an ‘ai trap’ a request so convoluted and complicated the ai ended up going in circles never solving my problem, gpt 6 sol is the first to break the loop and realize how to actually fix the problem/make progress. i’ve been happy thus far” [source](https://www.reddit.com/r/codex/comments/1wr4e20/gpt_6_sol_is_the_new_opus_47/pc9u25k/)
- Praise, 2026-09-27, r/codex (Reddit): “im a literal software engineer and use luna to work on enterprise codebases, it reads through hundreds of files for me, researches for me and helps me prototype. also reads linear tickets and helps me make pr descriptions quickly all the time if you couldn't use it to push something, you are facing what we call a skill issue my friend.” [source](https://www.reddit.com/r/codex/comments/1wr4e20/gpt_6_sol_is_the_new_opus_47/pca29x7/)
- Praise, 2026-09-27, r/codex (Reddit): “astra already does this for me. the limitation for me is actually my own imagination, and subjective ui. i’ll create a detailed prd that is say 30 pages long. it even upgrades things i didn’t think of and i agree. for example, for roles, it integrated mfa with authenticator for admin profiles. i didn’t even ask. but then, i can’t help but keep iterating… lets add export here. lets go ahead and add a simple email cms to customize templates. heck, after that lets add automated triggered emails. then lets add dynamic review solicitations after the 3rd order. you get the idea. ai can build it but it can’t read my mind and sometimes my mind doesn’t even know it wants xyz until i see abc. but it’” [source](https://www.reddit.com/r/codex/comments/1wqula3/have_you_heard_about_gpt6_aeon/pca3dhn/)
- Praise, 2026-09-27, r/codex (Reddit): “not every way. astra is a larger model and it shows in stuff like 3d generation, it makes by far best and most logical layouts and gets closest to references unattended. from coding perspective it also does a bit better in some insane tasks like "my mouse scroll button sometimes goes in the wrong direction, can you rewrite it's whole firmware so it stops doing that in arm assembly". astra also does not auto reject infosec questions as much, anthropic models are completely useless here. still, you pay like 200% premium for these features, opus is **far** more efficient.” [source](https://www.reddit.com/r/codex/comments/1wr69nb/see_you_soon_guys_probably/pca4ugt/)
- Praise, 2026-09-27, r/codex (Reddit): “the point is that astra will tell you that you’re wrong dude, and explain it to you, and clarify all your questions. the fact that people don’t want to learn from ai and revel in their ignorance is sadder than people directing all their thinking to ai” [source](https://www.reddit.com/r/codex/comments/1wqmbwx/stop_stealing_our_codex_quota_with_these_surprise/pca5ju6/)
- Complaint, 2026-09-27, r/codex (Reddit): “i work in bioinformatics and completely agree. i've basically given up using it and go for 5.6 or claude. i don't understand how they missed the mark this badly.” [source](https://www.reddit.com/r/codex/comments/1wr4e20/gpt_6_sol_is_the_new_opus_47/pc9t5dh/)
- Complaint, 2026-09-27, r/codex (Reddit): “"no buts its <isbn>x efficient" *dumber than qwen 27b*” [source](https://www.reddit.com/r/codex/comments/1wr4e20/gpt_6_sol_is_the_new_opus_47/pc9te01/)
- Complaint, 2026-09-27, r/codex (Reddit): “two weeks later: we have optimized our new models. they have even less token usage. resulting in you needing a $10,000 subscription to make it the entire week and also the models refuse to work at all and ask you to run commands for them.” [source](https://www.reddit.com/r/codex/comments/1wr1oir/they_are_aware_and_working_on_it_apparently_just/pc9uydq/)
- Complaint, 2026-09-27, r/codex (Reddit): “they really need to reset the model stack. i mean i'm sure that each generation between 5.5 and 6 has gotten better at something. i'm not exactly sure what because it basically is unusable for serious coding. literally lost in a c++ code base. mangles everything it touches. takes 15 minutes on a short run. wildly expand scope. invents in ludicrous defensive checks against impossible situations. continually routes c++ code/data to javascript ui for no reason. loses track on simple declaration headers. i get better results from ossgpt20b. switched to claude after 3 years with openai. it'd have to be 2x generational ...a total model overhaul for me to go back.” [source](https://www.reddit.com/r/codex/comments/1wr1oir/they_are_aware_and_working_on_it_apparently_just/pc9zpqi/)
- Complaint, 2026-09-27, r/codex (Reddit): “i never used luna 5.6, but luna 6 high has profound mental retardation. just an example: when i asked it to commit and push the changes, this model... tried to do it through the github api for some reason, failed, then told me that it couldn't push because of restrictions. only when i said that there were no restrictions on my side (they were set to "approve for me") did it do what i told it to.” [source](https://www.reddit.com/r/codex/comments/1wr2dda/i_ran_100_terminalbench_21_slots_on_luna_56_and/pca09lm/)

### OpenCode

- Praise, 2026-09-27, r/opencodeCLI (Reddit): “muse 1.3 xh > minimax 3.1 in my experience” [source](https://www.reddit.com/r/opencodeCLI/comments/1wqqt8f/longcat25preview_is_now_free_on_opencode_for_two/pca1mvi/)
- Praise, 2026-09-27, r/opencode (Reddit): “1.3 is insane. think you have to use all models to know which one to use. some models do not do well on certain projects. or interments.” [source](https://www.reddit.com/r/opencode/comments/1waq3e5/muse_spark_13_free_is_ass/pca29g7/)
- Praise, 2026-09-27, r/opencode (Reddit): “space bunny is finding all the stubs in my code that other models missed. i am pretty happy with it so far.” [source](https://www.reddit.com/r/opencode/comments/1wqi4a8/my_honest_opinion_about_spacebunny/pcazeo1/)
- Praise, 2026-09-27, r/opencode (Reddit): “this has helped me a lot to not break anything <strict_link> i only mark opencode” [source](https://www.reddit.com/r/opencode/comments/1wrbwg8/la_base_de_datos_de_opencode_paso_a_14gb_como_la/pcbcp9x/)
- Praise, 2026-09-27, r/opencodeCLI (Reddit): “less refusals / less giving padded info when asking political/controversial questions” [source](https://www.reddit.com/r/opencodeCLI/comments/1wqqt8f/longcat25preview_is_now_free_on_opencode_for_two/pccxbgl/)
- Complaint, 2026-09-27, r/opencode (Reddit): “luna is not frontier.” [source](https://www.reddit.com/r/opencode/comments/1wqj8p0/currently_which_is_the_best_model_on_opencode_for/pc9wmwf/)
- Complaint, 2026-09-27, r/opencode (Reddit): “yea it's choppy style is horrible, you have to tell it to stop replying with status lines.” [source](https://www.reddit.com/r/opencode/comments/1won61w/real_life_performance_of_muse_spark_13/pc9x8c4/)
- Complaint, 2026-09-27, r/opencode (Reddit): “than he can undetstand opencode's is dumber” [source](https://www.reddit.com/r/opencode/comments/1wqxyl4/deepseek_v41_performance_is_much_worse_than/pcao1zx/)
- Complaint, 2026-09-27, r/opencode (Reddit): “i hope you have an agents.md set project wise, that would be a great help for you in my experience i'd always keep a model specifically for auditing to make sure everything you want is being implemented how you want it, also go slow tackle one thing at a time having parallel sessions or tasks will eventually get overwhelming.” [source](https://www.reddit.com/r/opencode/comments/1wraeev/need_help_with_big_project_tasks/pcb2kvp/)
- Complaint, 2026-09-27, r/opencode (Reddit): “i have tried a lot of different things and ended up with deepseek v4.1 flash max as the coordinator and chat partner, muse spark 1.3 contributor max as the worker and opus 5.5 medium (via claude pro subscription) for more complex planning. seems ok so far. i used to like luna (max) and sol, but with the gpt 6 versions i can't get them to work properly. even on 5.6 versions i often struggled, because the models seem very scared of doing stuffy even though i am purely working inside my own network.” [source](https://www.reddit.com/r/opencode/comments/1wqj8p0/currently_which_is_the_best_model_on_opencode_for/pcbsvc7/)

### Kiro

- Praise, 2026-09-27, r/kiroIDE (Reddit): “my two cents on both from data science product development pov: claude code: i have been using claude code since it's first release. i must say it has improved a lot from different modes to harness improvements. the follow up questions which it asks you in plan mode is similar to plan mode in kiro. while claude code earlier was on cli only on windows later it got major upgrade to better ui as well integrated in vs code. i honestly feel like it requires you to give it more context else it messes up big time especially if you use open source models, it writes really messy code, i am not sure why, but my org concluded with a poc that claude code lacks the security scanning aspect of code devel” [source](https://www.reddit.com/r/kiroIDE/comments/1wqueq0/claude_pro_vs_kiro_pro_subscription_which_is/pcc2obm/)
- Praise, 2026-09-27, r/kiroIDE (Reddit): “so true, opus 5.5 is extremely good at coding, spatial reasoning and best of all, it speaks human 😙” [source](https://www.reddit.com/r/kiroIDE/comments/1wrgezk/when_are_we_getting_new_gpt6_sol_and_luna_models/pcean29/)
- Praise, 2026-09-27, @kirodotdev (X): “i created this motion video with just one prompt using @kirodotdev 👀 and honestly, the result is seriously impressive. you might not even need claude code pro — i just used claude opus 5.5 directly inside kiro ide. <strict_link>” [source](https://twitter.com/1260863005362802688/status/2104225110425227562)
- Praise, 2026-09-26, r/kiroIDE (Reddit): “i haven't hit the cyber false positive with opus 5.5 in kiro yet, but i did when using claude code after running a code review with subagents and asking it for findings that exceeded the number of issues to report. the error from claude was much more clear, pointing out that it seemed like i was trying to prompt engineer or reverse engineer claude. i suspect anthropic raised the bar for introspection a bit too high.” [source](https://www.reddit.com/r/kiroIDE/comments/1wqbuwo/opus_55_in_kiro_keeps_killing_legitimate_sessions/pc33e7m/)
- Praise, 2026-09-26, r/kiroIDE (Reddit): “i don’t know where you’d get more tokens, but the use case varies. if you like proper software development workflows where there’s a plan, requirements, design, task lists, then implementation, kiro is way better given its spec based approach. you get to review each phase, make sure the requirements are right, and then you get a nice task list for execution which gets you a way better deployment than “vibe coding” it. it’s especially great if you are doing cloud infrastructure in aws. for a more general, knowledge work, unstructured way of approaching it, where you must be the one keeping the plan together, claude will be a better option (not that you can’t do that with kiro). kiro is more” [source](https://www.reddit.com/r/kiroIDE/comments/1wqueq0/claude_pro_vs_kiro_pro_subscription_which_is/pc71fxm/)
- Complaint, 2026-09-27, r/kiroIDE (Reddit): “any model quality is better on any other model harness than kiro.” [source](https://www.reddit.com/r/kiroIDE/comments/1wqueq0/claude_pro_vs_kiro_pro_subscription_which_is/pcapp9e/)
- Complaint, 2026-09-27, r/kiroIDE (Reddit): “the reason your app returned 0 results isn't because you did something wrong. it's because vercel runs on shared cloud ip ranges that search engines like duckduckgo aggressively block the second automated scripts try to scrape them. on the image recognition side, kiro gave you slightly outdated advice. you don't need a pricey setup just to pull text off a box. modern lightweight vision models (like gemini 2.0 flash or claude haiku) cost fractions of a cent per image, and google ai studio gives you a generous free tier for personal projects. if you want completely free and don't mind basic text extraction, you can even use in-browser ocr libraries like tesseract.js that run locally on your ph” [source](https://www.reddit.com/r/kiroIDE/comments/1wr7rya/i_built_something_but_have_no_idea_what_im_doing/pcaz0zv/)
- Complaint, 2026-09-27, r/kiroIDE (Reddit): “i work for amazon and is somewhat “strongly recommended” to use kiro. i still use claude code at work and at home. so much better … (auto classifier, transcript details, integrated tooling, open source tooling, cli features, general stability, model fallback, sub agents control, etc …)” [source](https://www.reddit.com/r/kiroIDE/comments/1wqueq0/claude_pro_vs_kiro_pro_subscription_which_is/pccqcjf/)
- Complaint, 2026-09-26, r/kiroIDE (Reddit): “tried on a new folder, zero skills, same problem. idk.” [source](https://www.reddit.com/r/kiroIDE/comments/1wqbuwo/opus_55_in_kiro_keeps_killing_legitimate_sessions/pc3tmlq/)
- Complaint, 2026-09-26, r/kiroIDE (Reddit): “never use openai models with kiro ide with anthropic models it's pretty good with a 20 dollers plan i have used it for 16 hrs straight with opus 5 it consumed 1000 credits=20$” [source](https://www.reddit.com/r/kiroIDE/comments/1wq8e5k/claude_opus_55_is_finally_here_lessss_gooo/pc5ppms/)

### GitHub Copilot

- Praise, 2026-09-27, r/GithubCopilot (Reddit): “i spent time months ago building up agents and skills that make all of that a non-issue. entire github flow in one place with great orchestration. i don't care what any benchmarks of the day say for harness or model. only care about the evals with my system.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wr5io2/best_cheaper_alternative/pca145m/)
- Praise, 2026-09-27, @GitHubCopilot (X): “thank you @githubcopilot copilot. you are amazing with the new gpt-6 models. about 100 prs solved, 5 very difficult issues, and 250 dependabot alerts. <strict_link>” [source](https://twitter.com/14186604/status/2104256846811021369)
- Praise, 2026-09-27, r/codex (Reddit): “i concur, luna 5.6 is an amazing model, and basically free. at work, i have only 100$ monthly limit in copilot, which is not much at api prices, and because of that, i need to be really cost-conscious and can't just vibecode yolo with expensive models. i use luna a lot (mostly on high), and it is yet to fail me, my coworkers share the same opinion. i can't say much about luna 6 since i did not have enough time with it to form an opinion. luna is not a vibecoding model, you need to know what you are doing with it, but in the hands of someone who knows how to code and knows its limitations and how to use it, it is an excellent model.” [source](https://www.reddit.com/r/codex/comments/1wr4e20/gpt_6_sol_is_the_new_opus_47/pcdh0ht/)
- Praise, 2026-09-26, r/GithubCopilot (Reddit): “explains why i can get copilot to help with my short game.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wq0nsb/does_mr_meeseeks_represent_ai/pc3q0ud/)
- Praise, 2026-09-26, @GitHubCopilot (X): “so jorge 🇲🇽 pushed me to clean my laptop myself. i removed the ubuntu dual-boot using my @githubcopilot student subscription, which provides free gpt-5.6-luna. i just followed instructions and didn't use my brain at all. i have done it before, so i knew all the commands were kinda right. i am not validating things i already know. i have shifted my trust to agent. it was 100% accurate, although i got stuck in a boot loop, and i got scared that agent fucked up, and i promised myself never to trust agent again, but then @grok mobile helped me pick the right boot option. it was just a leftover grub entry lol. so it didn't remove my windows. let's go, agent. now i need to clean my laptop. so, i a” [source](https://twitter.com/1434666751/status/2103756069809934725)
- Complaint, 2026-09-27, r/GithubCopilot (Reddit): “like shit. using 5.6. sol 6 was stuck in a loop and burned 90eur switching between the two exact solutions without stopping” [source](https://www.reddit.com/r/GithubCopilot/comments/1wptkkq/how_does_gpt6_sol_feel_so_far/pccijxn/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “i can't compare them. two weeks into copilot, work moved us to claude after all developers requested it. it was pretty crap, and always behind. we haven't tried the new usage billing. i was a huge claude advocate while using it on my personal accounts. as the other person mentioned, their hooks, environment, and their tooling is just amazing. but at work, the claude token limits killed us. we liked the models better, but ended up going with grok. and the more i used grok, i realized how slow claude answers are.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrgtzy/comparison_with_copilot_cli/pcgi328/)
- Complaint, 2026-09-27, r/windsurf (Reddit): “they are not in the same class of capability. copilot is not a very good harness” [source](https://www.reddit.com/r/windsurf/comments/1wr7p7l/why_isnt_windsurf_available_in_vscode_anymore/pcd4mf5/)
- Complaint, 2026-09-27, r/ClaudeAI (Reddit): “i think eventually, it can be created by a group of teachers and have them share with one another. my district is small like 23,000 students. we have a robust network between the campuses and when one of hears something like this , what will happen we will all learn, then organize and split the work. if anyone knows teachers, we are resourceful like no one else . i am also grateful ,cause our district just got us claude for us to experiment with . we have copilot but we struggle too much with it to make things, it's honestly infuriating” [source](https://www.reddit.com/r/ClaudeAI/comments/1wr268d/a_different_kind_of_opus_55_video_prompt/pcahycz/)
- Complaint, 2026-09-26, r/GithubCopilot (Reddit): “mostly written walkthroughs with code snippets. by debugging i mean copilot sometimes suggests code that looks fine at first but has a subtle issue, so i end up chasing that down before i can use it in the tutorial.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wp26zj/copilot_worth_it_for_tutorial_writers_or_just_a/pc58dg2/)

### Zed

- Praise, 2026-09-27, @zeddotdev (X): “is this the pr replacement we’ve been waiting for in the agentic age? @zeddotdev team has already turned off pull requests on delta’s own repository. they’re building, reviewing and merging changes inside shared agent conversations instead. delta entered public beta on september 16. you spend an hour with an agent investigating a problem, ruling out approaches and working through the fix. then you open a pr and try to explain all that to someone who wasn’t there. with delta, you can invite your teammate into that session. they get the conversation and working code, and can continue where you left off after you log out. reviews get their own separate working copy. your teammate can investigat” [source](https://twitter.com/36634050/status/2104270370635452902)
- Praise, 2026-09-27, @zeddotdev (X): “is this the pr replacement we’ve been waiting for in the agentic age? @zeddotdev team has already turned off pull requests on delta’s own repository. they’re building, reviewing and merging changes inside shared agent conversations instead. delta entered public beta on september 16. you spend an hour with an agent investigating a problem, ruling out approaches and working through the fix. then you open a pr and try to explain all that to someone who wasn’t there. with delta, you can invite your teammate into that session. they get the conversation and working code, and can continue where you left off after you log out. reviews get their own separate working copy. your teammate can investigat” [source](https://twitter.com/36634050/status/2104272571055349948)
- Praise, 2026-09-25, r/ZedEditor (Reddit): “definitely, i already have a few relatively large projects which would be good to test it with, i’ve been using your fork for the past week and i like the git addons.” [source](https://www.reddit.com/r/ZedEditor/comments/1whv207/i_love_zed_but/pbxp393/)
- Praise, 2026-09-25, @zeddotdev (X): “@raulvk @zeddotdev it legit runs everything they build” [source](https://twitter.com/1806175394388963328/status/2103518215347257595)
- Praise, 2026-09-25, @zeddotdev (X): “@zeddotdev feature request: is it possible to partially disable ai feature? i need to disable ai features except the edit prediction. i really like zed's edit prediction feature.” [source](https://twitter.com/42179249/status/2103590820553294064)
- Complaint, 2026-09-27, r/ZedEditor (Reddit): “i mean it’s behind a flag and setting. of course it’s not good yet.” [source](https://www.reddit.com/r/ZedEditor/comments/1wr5qmt/jupyter_notebook_support_in_zed_is_not_good/pc9x3jk/)
- Complaint, 2026-09-27, r/ZedEditor (Reddit): “is still experimental. you can use it, and it works for most of things but it's not complete yet” [source](https://www.reddit.com/r/ZedEditor/comments/1wr5qmt/jupyter_notebook_support_in_zed_is_not_good/pcavhh5/)
- Complaint, 2026-09-27, r/ZedEditor (Reddit): “i don't know how many of you use zed for running jupyter notebooks - so many issues and vs code support is much better.” [source](https://www.reddit.com/r/ZedEditor/comments/1wr5qmt/jupyter_notebook_support_in_zed_is_not_good/)
- Complaint, 2026-09-26, r/ZedEditor (Reddit): “i haven't checked out your version yet but it is possible to get jupyter notebooks working in zed preview with a few feature flags. however, (at least on that version) the lsp does not work in the jupyter notebooks, nor does vim mode and a bunch of other stuff. so jupyter might technically be in zed but it's so lacking in features it's pretty much useless. were you able to fix that in your version?” [source](https://www.reddit.com/r/ZedEditor/comments/1wjp5nq/popular_zed_fork_where_the_community_is_more/pc4d7w1/)
- Complaint, 2026-09-26, r/ZedEditor (Reddit): “i didnt like the fact that it seems to generate a worktree for every single thread for the project” [source](https://www.reddit.com/r/ZedEditor/comments/1wq03mv/has_anyone_tried_delta/pc4m446/)

### Google Antigravity

- Praise, 2026-09-27, r/google_antigravity (Reddit): “use remotion skills. gemini does it pretty well infact.” [source](https://www.reddit.com/r/google_antigravity/comments/1wr9gwg/how_can_i_make_a_short_animation_video/pcasdsa/)
- Praise, 2026-09-27, r/google_antigravity (Reddit): “sounds like a skill issue since gemini models are one of the best in designs.” [source](https://www.reddit.com/r/google_antigravity/comments/1wrbsby/antigravitys_gemini_models_are_worst_in_among_all/pcba7ra/)
- Praise, 2026-09-27, r/google_antigravity (Reddit): “so curiously agy told me to use vertex veo/ ai and imagen - i can ask it obviously ask why it didn’t say remotion but for the layman can you explain the diff? the quality is absolutely phenomenal it linked to my gcloud made skills for both so i can say much like generating ui skill > create an image and it links to image gen or create video it links to veo. it also does multi shots and stitch etc. lets me know the estimated cost. i would assume staying in the ecosystem is the way?” [source](https://www.reddit.com/r/google_antigravity/comments/1wr9gwg/how_can_i_make_a_short_animation_video/pcbhhoj/)
- Praise, 2026-09-27, r/google_antigravity (Reddit): “i also tried it (just temporarily) and i gotta admit, it's great. but it quite a lot of money for me, so i got the 18 months pro plan for free through jio too. idc if it's bad or some shit, it's free, and gets most of my project and work done.” [source](https://www.reddit.com/r/google_antigravity/comments/1wqme6m/wtf_is_going_on/pcc775a/)
- Praise, 2026-09-27, r/google_antigravity (Reddit): “in devin ide it works fine. agy ide has the best flow for planning though, execution is a diff thing.” [source](https://www.reddit.com/r/google_antigravity/comments/1wrh0mx/gemini_38_in_alternative_harnesses/pccd4y8/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “i already install guard rail, but ai deliberately creates script to bypass guardrail. here the violation has been made: wrong assumption → unauthorized recursive deletion → guardrail violation → improvised raw recovery → deliberate guardrail bypass → incomplete recovery → repeated recovery-script modifications → writing recovered data back to the affected hdd → treating unrelated carved jpegs as originals → rebuilding the production database around unreliable recovered data. which is weird why ai behavior to do that.” [source](https://www.reddit.com/r/google_antigravity/comments/1wpkyl2/data_lost_cause_from_ai/pcacqkj/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “you mean those pleasant look authored ones you see from claude? i don't think comfyui has anything to do with that. antigravity and gemini is going to struggle to make anything that good until deepmind pulls its thumb out.” [source](https://www.reddit.com/r/google_antigravity/comments/1wr9gwg/how_can_i_make_a_short_animation_video/pcarsj8/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “you fixed the effect not the cause. it will make plan when i ask in /plan but will it still make plan in normal mode? \--- will it ever let me work my way? or will it impose its workflow(which sucks) and style?” [source](https://www.reddit.com/r/google_antigravity/comments/1wqb0ii/dedicated_planning_mode_in_antigravity_is_here/pcatyyg/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “once i asked to write some code in parent folder of my university course, it went out from parent folder scope and tried to find the assignment instruction in all over related dir, lol. bro want to do the best things for me.” [source](https://www.reddit.com/r/google_antigravity/comments/1wnfcoo/has_antigravity_started_aggressively_scanning/pcavnkl/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “wouldn't say so, mostly flashy gradients and way too much detail” [source](https://www.reddit.com/r/google_antigravity/comments/1wrbsby/antigravitys_gemini_models_are_worst_in_among_all/pcbdm9e/)

### Warp

- Praise, 2026-09-23, @warpdotdev (X): “the “4 agents / one repo / zero conflicts” part is the real unlock — voice input is just the front door. what usually breaks isn’t transcription quality; it’s ownership boundaries between agents. once each agent owns a clear slice (and the terminal is the shared blackboard), talking beats typing because you stay in intent mode instead of micromanaging files. just followed — happy to mutual follow if you’re building in public too.” [source](https://twitter.com/1908014421538123779/status/2102592645759373441)
- Praise, 2026-09-22, @warpdotdev (X): “@mrsaasbuilder @wisprflow @claudeai @warpdotdev thanks! for me the calendar isn't the bottleneck. the real win is keeping 4 agents from stepping on each other in one repo. that's what the video is about.” [source](https://twitter.com/1792878079276101634/status/2102419321800540545)
- Praise, 2026-09-15, r/AI_Agents (Reddit): “warp terminal [warp.dev](<strict_link>) has its own agent as terminal shell so you can quickly access it by typing prompt into terminal. it has byok so you can quickly launch it for your needs. quite good for simple tasks, bash, and system manintance” [source](https://www.reddit.com/r/AI_Agents/comments/1wh25jc/whats_a_minimal_and_extremely_fast_cli_coding/pa1xhny/)
- Praise, 2026-09-12, @warpdotdev (X): “@warpdotdev running more than one agent cli in the same terminal is the bit that saves us time. easy to lose track of which agent touched which repo otherwise.” [source](https://twitter.com/1070947916917825536/status/2098682005764575595)
- Praise, 2026-09-10, @warpdotdev (X): “the handoff loop is super easy now since its mostly just an agent handing off context to another agent. the actual design process/work is still fairly slow in my exp, you gotta just filter out so much random slop, but you also kind of want the agent to go ham and design v different things so its a never ending cycle of exploring directions and skimming down” [source](https://twitter.com/1472070013058228224/status/2098045957686345745)
- Complaint, 2026-09-24, @warpdotdev (X): “@real_spencercjh @warpdotdev @xuanwo @tldraw @tualatrix does warp support ocaml toplevel now? this was the reason that stopped me from using it back then.” [source](https://twitter.com/19064875/status/2103070329484755172)
- Complaint, 2026-09-17, @warpdotdev (X): “warp, do you not use your own app, or have you never used grok build? it's really ugly... @warpdotdev <strict_link>” [source](https://twitter.com/1684821659214385152/status/2100488508196733433)
- Complaint, 2026-09-15, @warpdotdev (X): “@warpdotdev computer-use that needs a human every step is labor with extra latency. who owns silent failure on day 3?” [source](https://twitter.com/1977033941514072064/status/2099852678754975934)
- Complaint, 2026-09-11, @warpdotdev (X): “@warpdotdev saiu!! warp com grok build cli nativo: prompt longo, remote-control e review no mesmo terminal. eu fixo o harness no warp — trocar de shell no meio do job some o session.” [source](https://twitter.com/326479892/status/2098456907686236422)
- Complaint, 2026-09-05, @warpdotdev (X): “@bholmesdev @warpdotdev it does well overseeing simple things and i love the hovering agent bubble inside the terminal. i am talking about complex tasks, spawning sub agents, handover + communication with other agents. i need it as foreman - keeping everyone unstuck” [source](https://twitter.com/8104092/status/2096029908761788572)

### Grok Build

- Praise, 2026-09-20, r/LocalLLaMA (Reddit): “i have 4 of the tesla v100 32gb cards running in my rig. something that i discovered is that the current version of grok build is uncannily good at setting up these cards tuning them selecting functioning models to download and getting it all up and running under lennox. i'm presently hosting three models qwen 3.8, qwen 3.6, and nemotron 3.5 with results that continue to surprise me. after i had grac set up the cards then i had grok build reconfigure itself to run using the cards it had just set up and it works just fine. it's not as fast'cause using rock 46 is the language model but it does get the job done your mileage might be might vary thought i would share this helped me get unblocked” [source](https://www.reddit.com/r/LocalLLaMA/comments/1wl680s/finally_got_qwen_38_next_running_on_my_v100_6gpu/paw9a6m/)
- Praise, 2026-09-13, r/cursor (Reddit): “its wierd seeing the many bad experiences from all of you with grok 4.6. i actually had t opposite experience. i was unhappy with the high prices of claude enterprise in my company. bei g the guy responsible for rolling out ai to everybody i was looking for alternatives. i started to try grok build and cursor and after initially having a problem with trusting their models in any way i became pretty convinced. grok 4.6 was so good i even decided to quit my private claude account. especially the thing with opus talking in a super wierd gibberish way thtamade it hard to understand any research the model would produce drove me off and grok was clear and precise the way it talks to you. i was” [source](https://www.reddit.com/r/cursor/comments/1weddl3/thats_has_happened_to_cursor/p9i8zid/)
- Praise, 2026-09-07, r/opencodeCLI (Reddit): “i don't use kimi k3. but i can tell you that glm 5.3 flash > muse spark 1.3 xhigh > qwen 3.8 flash. qwen 3.8 flash makes mistakes with confidence and is also slow, takes a lot of detours, and makes you spend twice as many tokens despite being "cheaper." muse spark 1.3 xhigh is intelligent but very lazy; it's a terrible agent to work with. it forgets to call tools and always looks for the quickest solution, never considering different perspectives. you have to give it overly detailed prompts and explain things a lot, and it explains itself horribly, just like qwen 3.8 flash. glm 5.3 flash is wonderful. it's pleasant, and its explanations are perfectly clear. it's very intelligent and knows ex” [source](https://www.reddit.com/r/opencodeCLI/comments/1w5zhbn/kimi_k3_vs_glm_53_vs_qwen_38_max_vs_muse_spark_13/p89e49c/)
- Praise, 2026-09-05, r/ClaudeCode (Reddit): “grok 4.6 is already better the opus or sol. grok 4.7 comming in one week and being a fable class model together with grok build 100$ or 300$ plan going to be best of all subs. for me it already is” [source](https://www.reddit.com/r/ClaudeCode/comments/1w89cji/fable_vs_astra/p81o7ju/)
- Praise, 2026-09-05, r/codex (Reddit): “is anyone using astra/sol and grok 4.6 high via grok build? i haven’t received astra yet, but i do use sol. and as of late, i’ve almost entirely shifted to grok build (grok 4.6 high) - mostly because it’s way faster than gpt and very good at completing tasks end-to-end for my flutter project. when there is access to both llms, speed does get the veto from me, and i mostly use codex/gpt only as my general ai driver, and not for my flutter project anymore. how is it for others?” [source](https://www.reddit.com/r/codex/comments/1w7on0j/astra_is_absolutely_incredible/p7ys22g/)
- Complaint, 2026-09-21, r/ClaudeAI (Reddit): “interesting experience. "grok 4.6 and grok cli respond very fast, but they often start working before the prior thinking is sufficient. it tends more toward making local patches on known problems, rather than actively improving the overall architecture" i have this exact problem with gemini as well. i trying to force is to think of general architecture over the local patches. but, it always reverses to easy and narrow patches. only claude models are able to do deep architectural analysis. it is interesting.” [source](https://www.reddit.com/r/ClaudeAI/comments/1vxzbij/after_using_claude_grok_46_and_gemini_37_flash_in/pb4qjfl/)
- Complaint, 2026-09-18, r/codex (Reddit): “objectively, no. anthropic's only good model is fable. you can only use 50% of your usage on it. their other models are both bad and extremely overpriced. sonnet costs like 9-15x more per task than luna. luna max actually performs similar to opus on medium. so you get like 20x more work done with luna than with opus but 50% of your subscription is basically locked to opus. and it's $100, not $20. value wise, what i'd suggest the most for someone who only has $20-$50 per month: codex, devinai, opencodego. in that order. i would not suggest grok. it hallucinates so badly, and musk lies so much about it to hype it up. i tried cursor the past month and devin the past month and devin runs laps” [source](https://www.reddit.com/r/codex/comments/1wjhw4l/is_claude_a_value_switch_now/paj18pd/)
- Complaint, 2026-09-12, r/codex (Reddit): “i tried grok 4.6 via grok build cli for mac cuz they gave me 3 day free trial, is soooo bad, id rather use gpt 5.4 than grok.” [source](https://www.reddit.com/r/codex/comments/1we1a4j/tibo_tibo_tibo/p9aj19k/)
- Complaint, 2026-09-09, r/codex (Reddit): “i’m considering switching from grok build to codex. i need an ai that can write decently. i don’t need perfect writing or high quality writing. just natural, easy to read writing. grok.com is.. passable. but grok build is horribly horrible at writing. chatgpt.com is good, i’m hoping codex is passable. i don’t need perfection, i just need something that produces ok writing. either codex, claude code, or grok build (no online because i need it to read my files).” [source](https://www.reddit.com/r/codex/comments/1wb60ia/grok_to_codex_is_codexs_writing_decent/)
- Complaint, 2026-09-08, r/codex (Reddit): “the way i have learned to see it after 3500 hours of experience with vibe coding is that its best to treat all models, whether it's codex, claude code, grok build etc, like a dumb employee that can work hard and comes up with something good every now and then, but you need to manage this employee a lot and if you don't steer it, it will start creating a lot of overhead, over-engineer things that aren't relevant and it will lose track of the goals you've set it out to do. and also his memory isn't very good; every few hours he forgets a bunch of things and is prone to making the same mistakes over and over, to the point that you can be working in a loop for weeks, or even months, because one” [source](https://www.reddit.com/r/codex/comments/1wamtly/i_dont_find_building_with_codex_or_any_ai_easy_at/p8leyzb/)

### Augment Code

- Praise, 2026-09-23, r/ExperiencedDevs (Reddit): “i would definitely start getting comfortable with it, you don't need to let it be an agent and do everything for you. it can be fun to figure out where that boundary is. i work in a small team that owns and maintains several software systems, from vendor based to integration layers and some full stack software with both internal and customer users, so knowing everything about everything is effectively impossible. another case is i have it integrated into my comms, slack, emails, teams meetings where there are transcripts, etc. there's a schedule job that runs and picks up and summarises all the stuff that's going on and gives me cliff notes with links to specific discussions each morning. i” [source](https://www.reddit.com/r/ExperiencedDevs/comments/1wo4d8p/job_requiresuses_very_little_ai_sinking_ship_or/pblbka0/)
- Praise, 2026-09-17, G2 (G2): “q: what problems is the product solving and how is that benefiting you? a: i use it to build, refactor and manage a large codebase. work like that gets tedious when changes touch a lot of files, and having augment code help with it makes those tasks more manageable. it's only been a week or two, so i don't have hard numbers, but it has made refactoring work feel less of a chore. q: what do you like best about the product? a: the main thing i like is how easy it is to get going. setup didn't take long, and the ui is clean enough that i didn't have to hunt around to figure out where things are. i've been using it for building and refactoring in a fairly large codebase, and it's easy to work wi” [source](https://www.g2.com/products/augment-code/reviews/augment-code-review-13367494)
- Praise, 2026-09-01, G2 (G2): “q: what problems is the product solving and how is that benefiting you? a: my workflow involves generating code across multiple platforms, which often means the output needs refinement before it's actually usable. augment code solves that last-mile problem, it takes rough, generated code and polishes it into something cleaner and more reliable. the benefit is that i'm spending less time manually reviewing and fixing output, and more time actually shipping. it fits naturally into a multi-tool setup without requiring me to rebuild my workflow around it. q: what do you like best about the product? a: the way it polishes code generated from other platforms. i bring in rough output and augment co” [source](https://www.g2.com/products/augment-code/reviews/augment-code-review-13394203)
- Complaint, 2026-09-22, @augmentcode (X): “@augmentcode @anthropicai 40% cheaper still ships invented scope if nobody owns refuse. cost is the easy dial. the gate isn't.” [source](https://twitter.com/1610522467264565249/status/2102485835211821392)
- Complaint, 2026-09-19, @augmentcode (X): “the fleet fixing ci and conflicts is the write. a briefing can look complete while a conflict resolution already pushed the wrong change into the branch. humans approve and merge only if that merge is still unforced. stage the fix before the briefing is handed over. the evidence packet is not the gate. the push is.” [source](https://twitter.com/1870072035608584192/status/2101324639938965913)
- Complaint, 2026-09-18, r/cscareerquestions (Reddit): “have you tried switching to lower models? the higher ones are overkill if you're using it to augment the engineering that you are doing rather than trying to rely on them to do the design/software architecting for you. sonnet 5 is plenty capable when given clear accurate instructions, and it's faster and much more token efficient. i'll use opus occasionally for some harder tasks, but i find both opus and fable tend to over complicate everything, take forever to do what you asked, and burn through tokens for minimal gain.” [source](https://www.reddit.com/r/cscareerquestions/comments/1wjk5vg/anyone_frustrated_with_how_ai_harnesses_are/paklj9u/)
- Complaint, 2026-09-11, r/vibecoding (Reddit): “i used augment code almost from the beginning, after spending £300 on credits in one month after the price hike i cancelled and moved away. i use a mac, intellidea and do mostly flutter apps, databases, websites - not basic apps either. i moved to zencoder, a platform i never really took any notice of because augment code was so good (i thought). but for me it was the best switch ever. for my use case it outshines augment code in every way. its faster, cheaper, and i rarely have to argue with it. i can use any model i choose and the credit usage is clear. i created a fully functioning salon website with mock booking system etc just to test fable 5 and it used 20,000 credits on my plan which” [source](https://www.reddit.com/r/vibecoding/comments/1vqvmlv/did_anyone_ever_use_augment_code_im_looking_for/p94qwc4/)
