# Can do the user's kind of task (`work.capability`)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/criterion/work.capability

Area: [Doing the work](https://feedbackbench.com/criteria/work.md)

**Definition.** Whether the agent, with its models, succeeds at the user's kind of task: its size, complexity, language or domain, including building a whole feature in one go or failing at it. The post names the task or its size but no narrower failure behaviour.

**Boundary.** Not this: see [Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md) for visual UI work and [Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md) for fixing a reported bug. Not this: see [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md) for a change over time. Not this: see the other work.* and verify.* leaves when a specific behaviour (looping, scope, false claims, breaking code) is named. Not this: see general.unspecific when no task is named.

Rated author-weeks, all agents: 10373. Complaint share: 39%.

## The brief

Written by Claude Opus 5.5 from 110 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**Strong models win scoped work; harnesses decide whether hard tasks land.**

TL;DR:

- Claude Code, Cursor, Devin and Amp beat peers here; Google Antigravity trails on task success.
- Fast, cheap models finish quickly, then users pay a stronger model to clean up.
- Same model, different harness, different result. Users blame wrappers as often as models.

In plain terms: Agents handle well-scoped features and legacy cleanup that once took hours. Complex or vague tasks still come back with bugs, so users review, re-prompt, or route the fix through a second, stronger model.

### How it breaks

- **Cheap models create cleanup work** ([Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md)). Fast, low-cost models finish first, but users report spending the savings on a stronger model to repair the output.
  The pattern repeats across agents. Users pick a flash or budget model for speed or quota, then run a frontier model to fix what it left behind. One Pi user says the smaller model leaves unhandled edge cases and weak tests, so total time and iterations go up. The trade looks cheap per call and expensive per finished task.
  Evidence:
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-03: “because gpt or claude make fewer mistakes, 3.7 and 3.8 flash – fast, very fast, but afterward i have to run gpt or claude to fix many mistakes.” [source](https://www.reddit.com/r/google_antigravity/comments/1w6ch6i/idk_why_anyone_is_still_paying_for_claude_or/p7mojog/)
  - Complaint, Pi, r/PiCodingAgent, 2026-09-18: “honestly i would've used big guns everywhere, but it's too expensive so i need to play around with subagents/forks. perception of luna good enough is straight up wrong. it leaves after itself a mess in codebase, unhandled edge cases and bugs, tests quality is terrible. so you think that you spend less cause luna did all that work, but you end up doing more iterations, have to review more carefully, still got bugs and spent way more time then just using larger model and pretty much doing it in 1-2 iterations.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wju5ja/real_software_developers_do_you_ever_really_feel/pan01b5/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-03: “flash is definetly not ahead of sonnet on high or xhigh on general coding usage.” [source](https://www.reddit.com/r/google_antigravity/comments/1w5iwz1/here_the_actual_performance_of_38_flash/p7lalrr/)
  - Complaint, GitHub Copilot, r/GithubCopilot, 2026-09-02: “every time i switch to dumber models they miss something and i regret it. i use 5.6 sol extra high. i work with a large code base of embedded safety critical c code” [source](https://www.reddit.com/r/GithubCopilot/comments/1w5i0sx/one_msft_employee_reaches_100k_mo_in_token/p7gv15r/)

- **Complex tasks still need human review** ([Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md)). Even users who praise their agent say hard tasks rarely land in one shot and code review stays mandatory.
  Users running multi-agent pipelines still read every diff. Hallucinated assumptions slip into complex work, and end-to-end testing surfaces bugs the agent did not catch. Some add an orchestrator or an automated second-agent review to raise the hit rate. The cost of a confidently wrong agent shows up at review time, not generation time.
  Evidence:
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-16: “i've got a whole ass pipeline of agents and i still need to read code to correct it. hallucinations and assumptions still happen.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wgrtt6/do_yall_still_read_lines_of_code/pa45n3u/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-26: “depends on how well your task specified, this has definitely happened to me with fable 5.1 as well. you kinda just gotta accept that llms can't always perfectly oneshot more complex tasks with no bugs and code review has to be part of any serious production deployment. but then again, merely adding in an orchestrator increases the chances that it will catch bugs without a separate code review step. you can also automate the code review step, heck, you can even have it spin up codex or something else for a different perspective. just put some reasonable limits on the review cycles if you plan to leave it unattended so that it doesn't burn all your tokens chasing some dumb lead.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqnjkp/be_careful_with_opus_55s_confidence/pc5gzsn/)
  - Complaint, Pi, @pidotdev, 2026-09-13: “@pidotdev @badlogicgames @mitsuhiko the only thing about tests right now is that it feels somewhat safer to let the agent do larger refactors. when developing new features in general i feel like the advantage is minimal. still hitting lots of bugs when actually doing end to end user testing.” [source](https://twitter.com/731697937/status/2099218423846731917)
  - Complaint, Devin, @cognition, 2026-09-10: “@cognition cutting coding-agent cost by 70% is a big deal. it makes experimentation cheap enough that a lot more teams will use them. the harder part is still reviewing what ships when the agent is confidently wrong.” [source](https://twitter.com/2463603438/status/2098100040854319424)

- **The harness changes the model** ([Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md)). Users report the same model performing better or worse depending on which agent wraps it.
  Posts describe identical models drifting more, guessing more, or writing worse code inside one harness than another. A Google Antigravity user suspects the harness makes Gemini worse. An Amp user saw harder tasks drift more than in Claude Code on the same Claude model. A Pi user found the standalone CLI beat a custom setup. Capability is a property of the pair, not the model alone.
  Evidence:
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-14: “well, gemini isn't the best, that's true, but i have my reasons to believe (based on my experience with antigravity) that the harness makes it that much worse” [source](https://www.reddit.com/r/google_antigravity/comments/1wg3igg/i_gave_antigravity_only_three_instructions/p9tqdrr/)
  - Complaint, Amp, @AmpCode, 2026-09-24: “orbs are genuinely great. but on the same claude model, harder tasks drifted noticeably more than in claude code. more guessing, more needing me to steer it back. feels like a harness gap. and after a few days i realized i had no idea what was actually happening in my codebase. the "agent runs while you're away" model is amazing when it works, but when the agent isn't reliable enough, it just becomes a loss of control. claude code can handle long-running tasks on its own, and it doesn't claim it did something when it actually didn't.” [source](https://twitter.com/1592160489965948933/status/2103097279691571704)
  - Complaint, Pi, r/PiCodingAgent, 2026-09-10: “i’d heard about the pi, that it’s a great way to customize tools to suit yourself and save tokens. i don’t know about token savings, but i tried using gemini 3.8 flash inside it, and i didn’t like how it wrote code. for a while, i tried to refine the pi using codex and claude, ran various tests, and thought that the whole problem was with gemini. i would really find it convenient if several ai models were inside a single harness, but right now i suspect that’s exactly the problem, despite all the attempts to refine it. i opened the classic antigravity cli, and the results seem much better. i don’t like the antigravity desktop app — it’s also kind of weak. i hope a miracle will happen and the cli will turn out to be a lifesaver. gemini attracts me with its slightly higher limits, even though i’m a claude fan. what do you think? what’s your experience?” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wcj9bg/who_uses_pi_what_do_you_like_about_it/)
  - Complaint, Cline, @cline, 2026-09-19: “@musingiqbal @bovetheline @cline @opencode i love cline and i still use it because of this visibility, but i cannot be productive with it. it makes models less smart also it does not even support multimodal features that openrouter exposes. currently i use claude desktop + herms.” [source](https://twitter.com/2051403352869638144/status/2101371670971703598)

- **Scoped prompts beat vague goals** ([Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md)). Agents succeed when users define the task tightly and stall when handed abstract goals or giant plans.
  Users who break work into defined steps report steady wins with modest models. Others say a model needs explicit intent or it returns commands instead of doing the work. A Copilot user says chunking matters more than model choice. The flip side: one Codex user argues an advanced model should handle abstract goals like a developer would, and says it does not yet.
  Evidence:
  - Praise, Google Antigravity, r/google_antigravity, 2026-09-12: “and it isn’t even that bad in coding either, you just need the bare minimum knowledge to instruct it more than codex or claude.” [source](https://www.reddit.com/r/google_antigravity/comments/1we8549/antigravity_projects_is_a_massive_sleeper_why_the/p9ctzv9/)
  - Praise, OpenCode, r/codex, 2026-08-31: “everyone hating in the comments, but you are right. if you have a structured codebase, plan out the work, and give it a defined task, it doesn't need to run long and burn through usage to implement something. you can get a lot done with it. i even set up terra and luna subagents in opencode and get even more out of it. it makes me wonder what people are building with this thing. is everyone just vibing and giving it some vague goal prompt or some unhinged massive plan, then letting it rip?” [source](https://www.reddit.com/r/codex/comments/1w344vm/20_plan_is_more_then_enough/p6xhbjb/)
  - Complaint, GitHub Copilot, r/AI_Agents, 2026-09-22: “copilot is hit or miss for real editing because it tries to do too much in one pass. the best results i've seen come from breaking the work into structured steps with claude or gpt-4: give it the doc as markdown or plain text, specify the edit scope tightly (redline these sections, keep formatting conventions), then reassemble. eating context limits is the core problem, so chunking matters more than which model you pick.” [source](https://www.reddit.com/r/AI_Agents/comments/1wlkceu/is_there_any_serious_ai_agent_for_editing_word/pbdt7jd/)
  - Complaint, Amp, @AmpCode, 2026-09-06: “holy shit i don’t like gpt-6 astra in @ampcode (or in any harness, really) it needs more handholding. you have to be more explicit about your intent and it’ll spit commands back at you instead of actually doing the work!” [source](https://twitter.com/1705384263867379712/status/2096477043475255422)

- **Legacy and bulk work land well** ([Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md)). The strongest praise comes from large, grinding jobs where agents clear volume humans would not touch.
  Users describe modernizing legacy apps across 100 commits without reading the code, closing triage queues across repos, and clearing dependency alerts in bulk. One Factory user calls out legacy-code blind spots as a real win. These posts share a trait: well-understood work at high volume, where speed compounds and errors are cheap to catch.
  Evidence:
  - Praise, GitHub Copilot, @GitHubCopilot, 2026-09-27: “thank you @githubcopilot copilot. you are amazing with the new gpt-6 models. about 100 prs solved, 5 very difficult issues, and 250 dependabot alerts. <strict_link>” [source](https://twitter.com/14186604/status/2104256846811021369)
  - Praise, Cursor, r/cursor, 2026-09-02: “this. i modernized a legacy app and never looked at the code until 100 commits later. even then, what am i supposed to do, open up some file in 4 directories deep and go "yep, just what i assumed. the code could didn't have to assign a variable to hold this thing"? like the manual review does nothing for me. i'm better off having an agent give me a high level review of the code in multiple areas such as security, best practices etc. it's so funny how people hold ai to this perfect standard, but evaluate humans as perfect just because their hands were on it. when humans were manually writing the code, in many cases there was 3 units tests and a bunch of commented out shit code. properly prompted and guardrailed ai is vastly superior for the real outcome we're all striving for (but typically wont admit).” [source](https://www.reddit.com/r/cursor/comments/1w5gyin/do_you_review_every_line_of_aigenerated_code/p7fy2wt/)
  - Praise, Cursor, r/cursor, 2026-09-12: “it addresses every quirk/limitation grokbot had for technical usecases: model selection, context windows, native ui integration with the core app. my mind was blown when i just dragged a cloud pr chat from two days ago into the project chat sidebar and the main agent just goes "on it, closing it out." and it did. cloud instances as just live anchor-text in the chat? brilliant! i threw a bunch of complexity at it in the past hour: cross-repo reference to do a complicated task. triage multiple outstanding issues. reconcile x/y/z. it worked. it did it well. i spent approximately 30s setting the whole thing up, and maybe two minutes of input while i focused on my primary work. just brilliant! only been using it for a day and i already see this as the future way of working. i honestly think with a bit more time in the oven, this approach represents a fucking revolution in computing!” [source](https://www.reddit.com/r/cursor/comments/1wdorqg/cursor_projects/p99gu94/)
  - Praise, Factory, @FactoryAI, 2026-09-27: “@factoryai @fireworksai_hq huge win for legacy code blind spots there are so real” [source](https://twitter.com/2010658787611619328/status/2104199345436312044)

### Who stands out

- **Claude Code (stronger)**. Users credit Claude Code with large builds and long-running jobs, while still flagging hallucinations on hard work.
  Posts describe whole products built without hand-written code and legacy block logic handled in a fraction of the usual time. A company-wide user says correcting it into memory steadily cuts slop over months. Another user prefers it for planning and creative work. Complaints focus on assumptions in complex tasks and on R&D algorithm work where it adds speed but not new capability.
  Evidence:
  - Praise, Claude Code, r/ClaudeCode, 2026-09-24: “some of our older systems is block logic. when the code become complex it is very time consuming and impossible to spot without testing. today claude did a 5 hour job in 40 min, and built a simulator in 10 min. training people in this code takes over a year since there simply is not enough tasks to train them in.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wp3pey/looking_for_engineers_who_are_actually_loving_this/pbty1at/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-24: “want to see one i consider pretty good? try <strict_link>. :-) the readme is a little out of date. i’m probably up to 140,000 lines now. at least the feature matrix explains why… but i haven’t written a single line of code for this one, and i’d be a little scared to do so.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wov62z/got_mogged_by_claude_opus/pbqol2x/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-16: “our entire company uses ai now. we are a full stack c#/typescript/mssql company. we have a lot of technical docs describing our product. claude also has direct access to the code base and we provide every package with its own markdown file. we also have numerous skills and memories we continuously add. you say slop, i say 'do it this way instead claude and store it to memory'. claude does and never does that type of slop again in the codebase. if you do that enough times it eventually stops writing slop. we've been at it for about 9 months now and still fixing slop but the amount is a fraction of what it was before. it takes time and if you have developers who dont correct it then it will continue to happen.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wgm4si/engineers_who_write_all_their_code_with_claude/pa7ztvc/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-03: “sure. i'd be a lot slower, but claude isn't doing anything that i couldn't do (yet). but i work in r&d, so the code i write is much more focused on bare bones algorithms instead of building entire products.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w5jwxj/are_you_too_depended_on_claude_for_your_daily_work/p7ky28w/)

- **Cursor (stronger)**. Cursor users praise fast, direct problem-solving and cross-repo task handling, with quality concerns surfacing in comparisons.
  Users say its models read tickets and design a solution with less hand-holding, cutting time from start to merge. One user threw cross-repo triage at it and says it worked. Another uses it as a second brain that catches issues frontier models miss. The pushback: some users find output quality worse than Devin on the same prompts.
  Evidence:
  - Praise, Cursor, @cursor_ai, 2026-09-22: “i love @deepseek_ai pricing but it is no @grok/@cursor_ai . where grok just reads the tickets and designs a solution, ds needs hand holding the whole way. and grok cutting right through to solutions means it is easily twice as fast. from start to merge.” [source](https://twitter.com/41420332/status/2102514622565917029)
  - Praise, Cursor, r/cursor, 2026-09-12: “it addresses every quirk/limitation grokbot had for technical usecases: model selection, context windows, native ui integration with the core app. my mind was blown when i just dragged a cloud pr chat from two days ago into the project chat sidebar and the main agent just goes "on it, closing it out." and it did. cloud instances as just live anchor-text in the chat? brilliant! i threw a bunch of complexity at it in the past hour: cross-repo reference to do a complicated task. triage multiple outstanding issues. reconcile x/y/z. it worked. it did it well. i spent approximately 30s setting the whole thing up, and maybe two minutes of input while i focused on my primary work. just brilliant! only been using it for a day and i already see this as the future way of working. i honestly think with a bit more time in the oven, this approach represents a fucking revolution in computing!” [source](https://www.reddit.com/r/cursor/comments/1wdorqg/cursor_projects/p99gu94/)
  - Praise, Cursor, r/cursor, 2026-09-24: “it's a decent second brain, that finds issues that astra and fable miss. it's slightly better than glm, but 100x better than glm for price and usage you get. plus $20 credit. depends on your use i guess.” [source](https://www.reddit.com/r/cursor/comments/1wpd3mm/i_hate_to_admit_it_but_grok_sucks/pbulfe1/)
  - Complaint, Cursor, @cursor_ai, 2026-09-17: “@devinai @cursor_ai @bot cursor seems to be much more generous with their usage limits, but if the quality feels worse across the board even with the same prompts, does that even matter? devin has consistently been giving me better code, better test suites, and more useful automations, but higher cost.” [source](https://twitter.com/2070908287978246144/status/2100701788127236560)

- **Devin (stronger)**. Devin draws praise for shipping concrete engineering outcomes, with its in-house model singled out by users.
  Users report finishing a full site visibility overhaul in about a day and call the in-house model underrated for execution. One user says it handles most work and they escalate only harder cases. Complaints target sloppy generated sites and very long runs that make slow progress.
  Evidence:
  - Praise, Devin, @cognition, 2026-09-24: “fixed my community site's ai visibility in ~1 day (duration not effort) with devin + jev. jev (@typesafe) scored 25 pages on "would an ai assistant find the answer here?" devin (@cognition) did the engineering: prerendering, json-ld, llms.txt, indexnow. seo 88→97.6. ai-eo 43→86” [source](https://twitter.com/521854405/status/2103256368472056079)
  - Praise, Devin, @cognition, 2026-09-25: “benchmarks will always look good on paper till you actually test out the model yourself jus imagining getting swe-2 for free on devin been using it alongside opus 5.5 but for execution it's such an underrated from by @cognition @da7_tech cooked on this one <strict_link>” [source](https://twitter.com/1504761521037062152/status/2103592548795044077)
  - Praise, Devin, @DevinAI, 2026-09-01: “@ryan_he_man @devinai @trysynara i love it, use it a lot. very good for 80% of the things, if it's harder i'll use swe 1.7 and sol for final reviews. try out the orchestrator feature in synara, you can have any agent create new threads and delegate stuff via /goal. works with all the agents in synara” [source](https://twitter.com/1617493632025776131/status/2094764442735255570)
  - Complaint, Devin, @DevinAI, 2026-09-19: “@brandon_galang @devinai good idea. i’ve had it running 5 hours nonstop on fusion and still only at 52%” [source](https://twitter.com/130668248/status/2101451910536654888)

- **Google Antigravity (weaker)**. Google Antigravity users lead the requests for better model and harness quality, and many route its output elsewhere for fixes.
  Users say its flash models are fast but not ahead of frontier rivals for general coding, and some dismiss it as fit only for trivial snippets. Requests for better model quality and better harness quality come mostly from its users. Praise exists for infrastructure setup and simple tasks, but it rarely extends to complex work.
  Evidence:
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-02: “yeah for cooking recipes and writing 10 short lines of r code copy/pasted from bioconductor tutorials perhaps” [source](https://www.reddit.com/r/google_antigravity/comments/1w5mjkz/love_it_tho/p7gyddu/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-03: “flash is definetly not ahead of sonnet on high or xhigh on general coding usage.” [source](https://www.reddit.com/r/google_antigravity/comments/1w5iwz1/here_the_actual_performance_of_38_flash/p7lalrr/)
  - Praise, Google Antigravity, r/google_antigravity, 2026-09-07: “used antigravity to set up and configure dockerized prowlarr sonar radar whisparr and a media server. within an hour or so everything was setup, tested and ready to go” [source](https://www.reddit.com/r/google_antigravity/comments/1wa44ok/arr_stack/)
  - Complaint, Pi, r/PiCodingAgent, 2026-09-10: “i’d heard about the pi, that it’s a great way to customize tools to suit yourself and save tokens. i don’t know about token savings, but i tried using gemini 3.8 flash inside it, and i didn’t like how it wrote code. for a while, i tried to refine the pi using codex and claude, ran various tests, and thought that the whole problem was with gemini. i would really find it convenient if several ai models were inside a single harness, but right now i suspect that’s exactly the problem, despite all the attempts to refine it. i opened the classic antigravity cli, and the results seem much better. i don’t like the antigravity desktop app — it’s also kind of weak. i hope a miracle will happen and the cli will turn out to be a lifesaver. gemini attracts me with its slightly higher limits, even though i’m a claude fan. what do you think? what’s your experience?” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wcj9bg/who_uses_pi_what_do_you_like_about_it/)

- **OpenAI Codex (mixed)**. Codex draws the most posts here, split between users it now writes everything for and users who call its new model half-baked.
  Praise centers on technical refactors and thoughtful critique of ideas, and one user says it now writes all their code. Complaints describe hallucinations, laziness, and redo cycles that burn usage. Its users lead the requests for reliable success on complex tasks and higher-quality generated code.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-12: “yep it does this. openai needs to come out with 6.1 as fast as possible cuz it’s ridiculous how astras behavior is so half baked” [source](https://www.reddit.com/r/codex/comments/1wedfpu/astra_sometimes_gets_stuck_announcing_work/p9dbglm/)
  - Praise, OpenAI Codex, r/codex, 2026-09-15: “i didn’t get terra to be useful before, when i tried to use sol for planning. now it writes all my code.” [source](https://www.reddit.com/r/codex/comments/1wgg7zu/2030b_tokens_a_week_to_1b_are_they_for_real/p9wig6q/)
  - Praise, OpenAI Codex, r/codex, 2026-09-25: “i agree. i have had subs to both since february. and antigravity shortly after. it's interesting to see my preferences and use time swing almost weekly. i do love when they release a new frontier model because i spend 90% of my budget using luna and sonnet xhigh, and those models usually get cheaper lol. i will say that i personally prefer claude for creative stuff and making plans, and codex for highly technical things like refactors and github work, but honestly you could use only one just fine. i love variety and the fact that i can put them on each other's code to check their work. really efficient that way imho.” [source](https://www.reddit.com/r/codex/comments/1wp3bzl/you_need_to_try_opus_55/pbwqoal/)
  - Complaint, OpenAI Codex, X search: OpenAI Codex, Codex CLI, Codex app, 2026-09-20: “what is going on with @openai? codex lost the plot. first of all, usage is absorbed at an astronomical rate, from 50% to 5%, and that's with me going to bed. but then the work output is so bad: the errors, the constant need to re do it, and welp there goes the usage. all the work that we've been trying to do is now stopped.” [source](https://twitter.com/249367919/status/2101625480818356322)

### Fine print

- Many posts judge underlying models rather than the agent, so harness and model effects are hard to separate.
- OpenAI Codex and Claude Code dominate post volume; smaller agents rest on far fewer voices.
- Request counts are small, so they show direction of demand, not its size.

## Top requests

What users ask to add or change, most asked first. 222 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Author-weeks | Posts | Agents (author-weeks) |
|---|---|---|---|---|
| 1 | Better overall model quality | 11 | 11 | Google Antigravity 8, OpenAI Codex 1, Cursor 1, OpenCode 1 |
| 2 | Better agent harness quality | 8 | 8 | Google Antigravity 5, OpenAI Codex 1, Devin 1, OpenCode 1 |
| 3 | Capability parity with competitor agents | 8 | 8 | Google Antigravity 3, OpenAI Codex 2, Cursor 1, Factory 1, OpenCode 1 |
| 4 | Reliable success on complex tasks | 8 | 8 | OpenAI Codex 4, Google Antigravity 1, Claude Code 1, Cursor 1, Factory 1 |
| 5 | Support for non-coding and non-technical work | 7 | 7 | Claude Code 3, Amp 1, Google Antigravity 1, Cursor 1, OpenCode 1 |
| 6 | Benchmarks reflecting real-world use | 6 | 6 | OpenAI Codex 2, Google Antigravity 1, Claude Code 1, Devin 1, OpenCode 1 |
| 7 | Better game development support | 6 | 6 | Claude Code 2, OpenAI Codex 2, Devin 1, Zed 1 |
| 8 | Fewer hallucinations and dumb mistakes | 6 | 6 | Google Antigravity 4, Claude Code 1, OpenAI Codex 1 |
| 9 | Higher quality generated code | 6 | 6 | OpenAI Codex 4, Cursor 2 |
| 10 | Blender integration for 3D modeling and animation | 5 | 7 | OpenAI Codex 3, Google Antigravity 2 |
| 11 | Built-in image generation | 5 | 5 | OpenAI Codex 3, Claude Code 2 |
| 12 | Fix looping and unreliable models | 5 | 5 | OpenAI Codex 2, OpenCode 2, Google Antigravity 1 |

### 1. Better overall model quality

- Google Antigravity, 2026-09-26, @antigravity (X): “i give antigravity a try often, and use it for things of no consequence. i just like giving you all a hard time. crisis precipitates change, after all. i would love it if the gemini had a model that could compete with openal / anthropic. i'd use it more if it was anywhere near as good. maybe focus on that instead of plan mode?” [source](https://twitter.com/1352821386679611396/status/2103982799422165426)
- Cursor, 2026-09-25, r/cursor (Reddit): “devin has all the same problems but worse. the only reason i stick with cursor for now is because i like the “ide first, ai second” feel that you don’t really get with command-line ai tools like claude code. i rarely leave ask mode and i love the workflow of cursor, if the models could just work properly.” [source](https://www.reddit.com/r/cursor/comments/1wpg5gt/thinking_of_cancelling_cursor_pro_and_switch_but/pbvxqy8/)
- OpenCode, 2026-09-21, r/opencode (Reddit): “i honestly need it to be a lot more smarter. the more smart it is, the more we can replace traditional llm with it. i like saving money” [source](https://www.reddit.com/r/opencode/comments/1wko226/how_good_is_jev_113/pb4tsto/)

### 2. Better agent harness quality

- OpenAI Codex, 2026-09-18, X search: OpenAI Codex, Codex CLI, Codex app (X): “@gohilhardy it would be good if inference is getting cheaper. and the coding harnesses still need more work (especially codex cli is still behind claude code cli).” [source](https://twitter.com/753242250/status/2100926904119214414)
- Google Antigravity, 2026-09-15, @antigravity (X): “@antigravity guys, fr yall need to improve a lot the harness. no matter how good the model is, if the harness is bad the model wont work good. i tried to edit and redesign a 2 pages pdf and it took nearly 2 hours(1h45m). thats just insane and tells how bad the harness for the agent is...” [source](https://twitter.com/1787288502310146048/status/2099673510553501792)
- Google Antigravity, 2026-09-07, @antigravity (X): “@1kartikkabadi1 too bad the harness antigravity sucks, like really. @antigravity anyhow to make it better? maybe not claudecode, codex level but more like pi, opencode? you bought the whole windsurf team! at least make it good like windsurf. maybe open the source so people can improve it?” [source](https://twitter.com/1192620866/status/2096811696149192828)

### 3. Capability parity with competitor agents

- Google Antigravity, 2026-09-27, @antigravity (X): “@johnennis that has been my experience as well. @antigravity is a complete shit show. literally all they have to do is copy codex or claude code and it would be great because i actually like the 3.8 flash model.” [source](https://twitter.com/402285185/status/2104056897968144560)
- Google Antigravity, 2026-09-22, @antigravity (X): “@tigerjpeg @antigravity make antigravity as good as codex and claude code pls” [source](https://twitter.com/2044713468931031040/status/2102433638298374556)
- Google Antigravity, 2026-09-21, @antigravity (X): “@nohedev @rodydavis @antigravity when will it be as capable as codex?” [source](https://twitter.com/1991165792768122880/status/2102176106359238894)

### 4. Reliable success on complex tasks

- Factory, 2026-09-27, @FactoryAI (X): “@droid @factoryai droid is genuinely useful for larger, multi-file tasks and does a good job staying on track without constant guidance. the biggest improvement for me would be better visibility into its reasoning/progress and more predictable results on longer tasks :)” [source](https://twitter.com/2093736525116702720/status/2104324409502970188)
- OpenAI Codex, 2026-09-15, r/codex (Reddit): “i wouldn't mind either way. i just want work to work when in projects reliably” [source](https://www.reddit.com/r/codex/comments/1wgpyz5/i_personally_love_the_chat_and_work_separation/p9wycu7/)
- OpenAI Codex, 2026-09-14, r/vibecoding (Reddit): “hey claude, i mean codex. build the entire photoshop app for me. work autonomously till the end. make no mistakes. make it polished.” [source](https://www.reddit.com/r/vibecoding/comments/1wf3hvx/i_vibe_coded_photoshop_alternative_using_gpt6astra/p9nwhdx/)

### 5. Support for non-coding and non-technical work

- Claude Code, 2026-09-26, @ClaudeDevs (X): “@claudedevs it’s expensive and it’s not as great as marketed we deserve super intelligence that can do more and build better for not so technical users, if claude had better knowledge on helping non technical users build things rather than spinning around in circles we would use less tokens” [source](https://twitter.com/1800156934504431617/status/2103952869481173286)
- Amp, 2026-09-07, @AmpCode (X): “@ampcode @thorstenball @sqs have you guys thought about this at all? not all my work lives in a git repo and i'd love to leverage the power of orbs with non-coding work” [source](https://twitter.com/1427083747506102282/status/2096762043810685256)
- Claude Code, 2026-09-06, r/ClaudeCode (Reddit): “any plans/eta on expanding outside of swe or jr roles? i'm on the infrastructure side and landing jobs there is equally as annoying” [source](https://www.reddit.com/r/ClaudeCode/comments/1w8e6i5/i_built_a_job_search_engine_for_claude_code_it/p87cwc8/)

### 6. Benchmarks reflecting real-world use

- Claude Code, 2026-09-22, r/ClaudeCode (Reddit): “we need better benchmarks for that, because virtually none of the mainstream ones catch those real world use cases, but they're very very real.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wnk74q/whats_the_point_of_fable_if_opus_55_is_stronger/pbfnupr/)
- OpenAI Codex, 2026-09-22, r/codex (Reddit): “cool! please benchmark it for performance/cost/speed vs using codex/claude vanilla out of the box! otherwise, nobody knows if this is useful or not” [source](https://www.reddit.com/r/codex/comments/1wmra7j/next_level_agent_nla_an_autonomous_multiagent/pbar7jz/)
- Devin, 2026-09-10, @cognition (X): “@kentcdodds @cognition building this offcut skill, still in dev with more features coming. would love to test it on the devin harness — and keep shipping through devin’s models. some features still sitting because the benchmarks are weak. <strict_link>” [source](https://twitter.com/1175314932914692096/status/2098178724042568046)

### 7. Better game development support

- Zed, 2026-09-17, @zeddotdev (X): “@zeddotdev can you pleease support angelscript? 🥲 using it in a game engine” [source](https://twitter.com/1156314936622092289/status/2100412485425635624)
- Devin, 2026-09-11, @cognition (X): “@kentcdodds @cognition i’m building wright: a local ai agent that can work directly inside roblox studio. devin, ship the first usable beta this week: connect to an open place, inspect the datamodel, make safe luau edits, run a playtest, read the output, fix errors, and show a reviewable diff etc” [source](https://twitter.com/2034777292770353152/status/2098517872846963134)
- Claude Code, 2026-09-09, r/ClaudeCode (Reddit): “beautiful! if only i could get it working so well with unreal engine or unity.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wb2ksh/did_a_horror_game_with_fable_51_1_prompt/p8nnmku/)

### 8. Fewer hallucinations and dumb mistakes

- Claude Code, 2026-09-27, @ClaudeDevs (X): “@claudedevs time for quality focus ai coding doing basic errors ex: missing out on paging” [source](https://twitter.com/1796200174093750274/status/2104064387510308910)
- Google Antigravity, 2026-09-25, @antigravity (X): “@antigravity when you guys released gemini 4 plz don't releases don't late i can't and plz fix only hallucination i repeated you guys only fix hallucination i don't care other updates only care hallucination fix plz make model top in less hallucination” [source](https://twitter.com/1992194786040905728/status/2103614732448239628)
- Google Antigravity, 2026-09-16, @antigravity (X): “@antigravity try to remove ai slop(in designing) and ai hallucinations in upcoming gemini 4 / 4 pro models” [source](https://twitter.com/1824095764202848256/status/2100044438781526218)

### 9. Higher quality generated code

- OpenAI Codex, 2026-09-27, r/codex (Reddit): “lmao. couldn't care less. i think what most pro (and heavy api) users want is just a better coding fullstack model. the pressure from opus 5.5 might prove overwhelming.” [source](https://www.reddit.com/r/codex/comments/1wrn697/new_openai_product_o_but_not_for_plus_users/pceb2s1/)
- Cursor, 2026-09-12, @cursor_ai (X): “@ethereaglehq @openai @claudeai @cursor_ai yes limit is the main concern . i have google ai ultra but that is very low quality code .. i want little better quality like say similar to sol high. but main concern is limit as i burn too fast” [source](https://twitter.com/1625966296998486016/status/2098805083182088391)
- Cursor, 2026-09-05, @cursor_ai (X): “@grok @elonmusk @cursor_ai whelp. reality is often different than the docs. no 4.5. i guess we go back to composer. not as good as 4.5 but at least you can code with it. hopefully they fix the coding ability in 4.7.” [source](https://twitter.com/863062011386527748/status/2096319205956263941)

### 10. Blender integration for 3D modeling and animation

- OpenAI Codex, 2026-09-15, r/codex (Reddit): “3d is literally what i need for my workflow the thing i like doing the least.. you know the part of work ai is supposed to alleviate, the parts you find tedious.. and for the record, it still sucks at it despite what people are saying or at least the version being hosted now is.” [source](https://www.reddit.com/r/codex/comments/1wg1zae/gpt6_sol/p9w67g5/)
- OpenAI Codex, 2026-09-15, r/codex (Reddit): “to add, it would be cool to have astra pull what cpuz, hwinfo, and other similar apps can do and then it can actually be "your" components modeled!🤣 i'm sure most popular components have plenty of pictures/videos that can be used to make the models if wanted to get extra crazy w/it haha” [source](https://www.reddit.com/r/codex/comments/1wgo9nx/i_built_an_opensource_interactive_3d_pc_anatomy/p9w0yu8/)
- OpenAI Codex, 2026-09-12, r/codex (Reddit): “will have to give that a try. i've found that even on high it struggles a lot with animations in blender with mcp. i literally had to describe to it in text how humans swing an axe to cut down a tree for it to get the animation right, and had to explain to it that humans hold tools by the wooden handle. had to repeat for every single animation, explaining every physical constraint in detail edit: it burned through all my usage without any results” [source](https://www.reddit.com/r/codex/comments/1we2tyf/astra_light_is_as_capable_for_most_work_as_xhigh/p9asx6b/)

### 11. Built-in image generation

- Claude Code, 2026-09-26, @ClaudeDevs (X): “@claudedevs se vocês consertarem aqueles limites injustos e derem um uso maior dos modelos, eu volto. se acrescentarem criação de imagens, eu nunca mais saio💔🙏🏻” [source](https://twitter.com/1506472371418537984/status/2103715316652208570)
- OpenAI Codex, 2026-09-25, r/ClaudeAI (Reddit): “claude doesnt have image generation, which is why codex becomes mandatory” [source](https://www.reddit.com/r/ClaudeAI/comments/1wq1n7n/codex_x20_vs_claude_code_x20_which_gives_you_more/pc0bvri/)
- Claude Code, 2026-09-18, @claude_code (X): “@claudedevs @claudeai @claude_code you guys should add a midjourney-style image generator directly into the subscription. a creative mode where claude can generate images right inside claude code would be insanely useful.” [source](https://twitter.com/1882101484155727872/status/2100986351386837030)

### 12. Fix looping and unreliable models

- OpenCode, 2026-09-26, @opencode (X): “deepseek 4.1 flash on @ollama cloud in @opencode 1.18.32 is unsuable for me. totally unreliable, looping, stopping. seems like they are running a low quantized version that starts to hallucinate, ans/or totally get lost in k/v cache vram needs. please address 💚🙏” [source](https://twitter.com/26595741/status/2103779333009559734)
- OpenAI Codex, 2026-09-25, r/codex (Reddit): “faster slop, most memory of slop, etc. fix the models. they suck at recognizing intent, are super lazy when doing discovery, react too quickly to any hint of frustration, stop working when there’s work to be done, turning a goal into a loop of bad decisions that are out of scope, etc. i miss older gpt models that were moving us beyond the days where we needed to spend half a day writing super specific prompts.” [source](https://www.reddit.com/r/codex/comments/1wpq44p/openai_prepares_new_500_per_month_pro_max_plan/pbyrzcs/)
- OpenAI Codex, 2026-09-25, r/codex (Reddit): “gpt6 is unusable and they are proud to tell us that they are actively working on their devday instead of fixing astra and sol.... wow, are we supposed to take this well?” [source](https://www.reddit.com/r/codex/comments/1wppkog/new_tibo_tweet_about_devday/pbxetko/)

## Every agent

| Agent | Overall rank | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|---|
| [Claude Code](https://feedbackbench.com/agents/claude-code.md) | 1 | Better than peers | 0.558 | 0.543–0.574 | 2798 | 1831 | 967 |
| [Cursor](https://feedbackbench.com/agents/cursor.md) | 4 | Better than peers | 0.555 | 0.526–0.584 | 639 | 444 | 195 |
| [Devin](https://feedbackbench.com/agents/devin.md) | =5 | Better than peers | 0.554 | 0.522–0.587 | 449 | 330 | 119 |
| [Amp](https://feedbackbench.com/agents/amp.md) | =11 | Better than peers | 0.550 | 0.522–0.578 | 55 | 48 | 7 |
| [Pi](https://feedbackbench.com/agents/pi.md) | 7 | Better than peers | 0.544 | 0.511–0.578 | 163 | 116 | 47 |
| [Factory](https://feedbackbench.com/agents/factory.md) | =11 | Better than peers | 0.543 | 0.516–0.570 | 57 | 48 | 9 |
| [OpenAI Codex](https://feedbackbench.com/agents/codex.md) | 2 | Typical | 0.489 | 0.477–0.500 | 4052 | 2371 | 1681 |
| [Kiro](https://feedbackbench.com/agents/kiro.md) | 13 | Typical | 0.485 | 0.460–0.509 | 34 | 17 | 17 |
| [Cline](https://feedbackbench.com/agents/cline.md) | =8 | Typical | 0.482 | 0.452–0.514 | 71 | 42 | 29 |
| [OpenCode](https://feedbackbench.com/agents/opencode.md) | 3 | Worse than peers | 0.467 | 0.442–0.496 | 859 | 502 | 357 |
| [GitHub Copilot](https://feedbackbench.com/agents/copilot.md) | =8 | Worse than peers | 0.457 | 0.423–0.494 | 158 | 78 | 80 |
| [Zed](https://feedbackbench.com/agents/zed.md) | =8 | Worse than peers | 0.440 | 0.411–0.472 | 52 | 19 | 33 |
| [Google Antigravity](https://feedbackbench.com/agents/antigravity.md) | =5 | Worse than peers | 0.371 | 0.347–0.394 | 954 | 437 | 517 |
| [Conductor](https://feedbackbench.com/agents/conductor.md) | 14 | Too few posts | – | – | 11 | 4 | 7 |
| [Grok Build](https://feedbackbench.com/agents/grok-build.md) | 16 | Too few posts | – | – | 10 | 6 | 4 |
| [Warp](https://feedbackbench.com/agents/warp.md) | 15 | Too few posts | – | – | 7 | 5 | 2 |
| [Augment Code](https://feedbackbench.com/agents/augment.md) | 17 | Too few posts | – | – | 4 | 3 | 1 |

## Posts

Receipts rule: The 5 most recent praise and complaint posts per area (first 700 characters) and 3 per criterion (first 450 characters).

### Claude Code

- Praise, 2026-09-27, r/ClaudeCode (Reddit): “even pre-fable, this has worked well. i have a cable that connects my odb-ii port to a particle tracker one. i had claude write the message processor, filtering logic and charts and graphs. i’ve used it on a few cars and it basically backs out the dbc file and goes to town.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wquoxy/fable_51_live_vehicle_diagnostics/pc9u3qf/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “my engineering department is on a team plan. we also have tier 5 openai. we just run up bills like crazy but the productivity is fucking insane.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr0wic/is_everyone_here_millionaires/pca1jju/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “shotcut if youre ever looking for a fleshed our free alternative. simple editing i have claude do with ffmpeg” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr4yjz/wanted_to_edit_some_footage_of_a_game_i_was/pcabeg2/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “every time i used sonnet it got something wrong, basically any type of agentic coding, like if u want it to write a single function and know what u are doing it's fine but otherwise it's ass” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr5lzf/practically_speaking_what_tasks_fit_into/pc9u33o/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “yes, this is true -- but now it's 10x negative value given the amount of work that can be completed with a few prompts. they're only really useful if your entire workflow is heavily guarded and gated - and even then, it's still difficult.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr27zi/be_honest_do_we_still_offer_value/pca7f1w/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “minecraft is so often replicated, it's source code probably shows up in every single model's training data. i'm actually surprised it took you longer than 30 minutes.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr8pum/used_opus_55_to_recreate_minecraft_in_threejs/pcarqim/)

### Cursor

- Praise, 2026-09-27, r/cursor (Reddit): “right there with you - not an elon fan, but have had grok 4.6 and now 4.7 as my everyday driver since composer 2.5. all entirely competent models.” [source](https://www.reddit.com/r/cursor/comments/1wrktgc/am_i_the_only_one_who_thinks_grok_47_is_actually/pcdcwpu/)
- Praise, 2026-09-27, r/cursor (Reddit): “i hate to say it but grok 4.7 is the best for my work flow because i can do everything in grok 4.7 high fast mode. and the reason i hate to say it is because grok has been used for disgusting things and it bothers me that the richest person in the world can buy yet another company and make it their own, but i think what they've done with cursor has all been actually positive as opposed to what happened to twitter. normally i'd do a frontier clau” [source](https://www.reddit.com/r/cursor/comments/1wrktgc/am_i_the_only_one_who_thinks_grok_47_is_actually/pcf8g60/)
- Praise, 2026-09-27, r/cursor (Reddit): “yes. i'm having the same realization over the past week. opus 5.5 is insanely capable all around. fable for heavy planning sessions burns many more tokens but been giving excellent results.” [source](https://www.reddit.com/r/cursor/comments/1wrx24u/1_year_of_cursor_switched_to_claude_best_decision/pcgl7cy/)
- Complaint, 2026-09-27, r/cursor (Reddit): “one-shotting is the issue. unless it's something as basic as a chrome extension, i never one-shotting.” [source](https://www.reddit.com/r/cursor/comments/1wqppuh/grok_is_shutting_down_apps_now/pcahmk0/)
- Complaint, 2026-09-27, r/cursor (Reddit): “cursor is crap compared to a pro plan on claude code.” [source](https://www.reddit.com/r/cursor/comments/1wrbmrn/is_it_a_waste_of_money_to_run_opus_55_inside_of/pcd0den/)
- Complaint, 2026-09-27, r/cursor (Reddit): “i compare it to grok 4.6 and 4.7 is still aggressively bad.” [source](https://www.reddit.com/r/cursor/comments/1wrktgc/am_i_the_only_one_who_thinks_grok_47_is_actually/pcddocg/)

### Devin

- Praise, 2026-09-27, r/windsurf (Reddit): “devin has invested a lot in controlling how the ais think inside the ide, to maximize context and code management tools. if you are going to pay for devin, try the devin desktop ide. it's a fork of vs code so it's not unfamiliar. the ide is good for almost any workflow, from lightning-fast code completion to highly reliable vibe coding with swe-2.” [source](https://www.reddit.com/r/windsurf/comments/1wr7p7l/why_isnt_windsurf_available_in_vscode_anymore/pcb72fn/)
- Praise, 2026-09-27, r/windsurf (Reddit): “they are not in the same class of capability. copilot is not a very good harness” [source](https://www.reddit.com/r/windsurf/comments/1wr7p7l/why_isnt_windsurf_available_in_vscode_anymore/pcd4mf5/)
- Praise, 2026-09-27, @cognition (X): “if you have claude subscription and @cognition cloud agents. try it it is amazing combo, it ships like crazy when you sleep. @cognition swe 2 max is a really good model and on cloud, it has macos, can code, run test, take screenshot and record evidence. opus 5.5 is really well as orchestrator, review and merge pr on your machine. you can feed opus (in claude code or hermes) devin api key, it knows what to do.” [source](https://twitter.com/1767985295910383616/status/2104014647724761554)
- Complaint, 2026-09-27, @DevinAI (X): “@fluyeporlaweb @hraness @chatgpt @devinai it seems to me like a bug factory... can you really exercise human quality control there and leave everything to the ai? i don't know, rick, it seems false to me. also, we should look at the numbers of how much income those subscriptions generate against what they cost.” [source](https://twitter.com/1913354152941432832/status/2104335120802943329)
- Complaint, 2026-09-27, r/ChatGPTCoding (Reddit): “welp... been 2 years. interesting and fun time capsule and appreciate you driving this. hate to say it, but as predicted my industry keeps hiring and growing and frontend and backend developers haven't been replaced by ai... the agi beliefs by leading researchers have tempered, many of whom no longer believe we're on the path to agi w/ llms. devin is notoriously terrible, no novel breakthroughs have been achieved by them yet ( and before you b” [source](https://www.reddit.com/r/ChatGPTCoding/comments/1fooq1c/will_ai_really_replace_frontend_developers/pcfpcxs/)
- Complaint, 2026-09-26, @cognition (X): “@cognition a billion dollars for a tool that still cant get my browser tabs in order but honestly good for them” [source](https://twitter.com/594090233/status/2103677906224632214)

### Amp

- Praise, 2026-09-26, @AmpCode (X): “i just did something that i wasn't aware it was possible using @ampcode . i had an idea on how to improve bug fixes. since everything now runs on an orb (their vm system), i thought: "it would be so cool if amp was my "support team""" so, i asked it to build. since we have a very good tracing system, we can check pretty much everything the user do, so we added an "send bug report" (the user allows us to see the last 10 minutes of his actions, the” [source](https://twitter.com/2931128860/status/2103847911725404390)
- Praise, 2026-09-26, @AmpCode (X): “amazing 🤩 for more context: have a fleet skill that let's the agents relay my home devices through raspberry pi when they need to test things on devices with no reliable simulator runtime like roku, samsung tizen & lg webos tvs 🫠 a bit niche use-case but think it'd be useful for mobile as well... only reason i got a mac-mini instead of linux devbox is apple simulators 😔” [source](https://twitter.com/1125366224664322049/status/2103858145101500803)
- Praise, 2026-09-26, @AmpCode (X): “.@ampcode is really good at using a namespace devbox through its ios app. <strict_link>” [source](https://twitter.com/1320822782498918404/status/2103922224176476222)
- Complaint, 2026-09-26, @AmpCode (X): “anyone using @ampcode to write blog posts? getting very mixed results so far so curious about your process.” [source](https://twitter.com/881577286927036416/status/2103824362746859692)
- Complaint, 2026-09-25, @AmpCode (X): “@jkudish @ampcode glm models are like gemini models. they just exist. they have claims, and they don’t work in the real world.” [source](https://twitter.com/137626804/status/2103381866736755000)
- Complaint, 2026-09-24, @AmpCode (X): “orbs are genuinely great. but on the same claude model, harder tasks drifted noticeably more than in claude code. more guessing, more needing me to steer it back. feels like a harness gap. and after a few days i realized i had no idea what was actually happening in my codebase. the "agent runs while you're away" model is amazing when it works, but when the agent isn't reliable enough, it just becomes a loss of control. claude code can handle long” [source](https://twitter.com/1592160489965948933/status/2103097279691571704)

### Pi

- Praise, 2026-09-27, r/PiCodingAgent (Reddit): “yes i know the system prompt is big. you said nothing about capabilities. have you actually tried pi with all the added features? things like subagents, memory, worktrees, and many others are all built in, and you need to add these to pi. try to be objective and compare the same things” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrjr75/the_good_opensource_harness/pcgpxe3/)
- Praise, 2026-09-27, r/PiCodingAgent (Reddit): “i use pi normaly i was trying out opus 5.5 today on claude code and it feels awful vs pi with /tree and all my custom stuff now im looking at the sdk claude pi extension, having model edit and reduce token bloat, hoping it won't get me banned” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrjr75/the_good_opensource_harness/pcgqytv/)
- Praise, 2026-09-27, @pidotdev (X): “pro move: install @pidotdev - web, tell @nousresearch hermes its there and let hermes talk directly to pi and do the work. i just get a ping when its done it needing a review. working great building <strict_link>” [source](https://twitter.com/15451416/status/2104052502727659701)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “yeah. on all my tests accuracy was quite shit, and jev was consistently outperformed by glm 5.3 flash, for anything requiring decision making. but hey, it's fast 🤷♂️” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wnl7da/pi_can_now_use_jev_and_more/pce99cd/)
- Complaint, 2026-09-27, r/LocalLLaMA (Reddit): “another "harness matters" post (codex cli > pi and opencode) i run my own llm while also having a openai subscription. also tried deepseek (latest flash now). i run [qwen 3.8 flash next](<strict_link>) at an amazing speed on my 2x3090 + ram! but local llm never did worked for me outside some demos like build me a "3d mario game, multistage" which i've been using to test llm's for a long time. at serios work, they never even compared with gpt 5.2” [source](https://www.reddit.com/r/LocalLLaMA/comments/1wrfp50/another_harness_matters_post_codex_cli_pi_and/)
- Complaint, 2026-09-26, r/opencode (Reddit): “for “learning”, i would start with a sandboxed pi.dev setup by default without any prompting. it’s lightweight and makes it easier to follow what’s happening behind the scenes. as you get more familiar and better at agentic development, you’ll start to want features that pi doesn’t natively offer and you’ll then either stay with pi or explore opencode. there’s a lot of magic in opencode. hence it’s not necessarily the best place to start - but a” [source](https://www.reddit.com/r/opencode/comments/1wqppth/opencode_v2/pc81x5p/)

### Factory

- Praise, 2026-09-27, @FactoryAI (X): “@render @factoryai agent writes the app, spins the db, deploys it. i just sit there like a decorative readme” [source](https://twitter.com/1330209814790746114/status/2104177796901945611)
- Praise, 2026-09-27, @FactoryAI (X): “@factoryai @fireworksai_hq huge win for legacy code blind spots there are so real” [source](https://twitter.com/2010658787611619328/status/2104199345436312044)
- Praise, 2026-09-27, @FactoryAI (X): “@droid @factoryai droid is genuinely useful for larger, multi-file tasks and does a good job staying on track without constant guidance. the biggest improvement for me would be better visibility into its reasoning/progress and more predictable results on longer tasks :)” [source](https://twitter.com/2093736525116702720/status/2104324409502970188)
- Complaint, 2026-09-27, @droid (X): “@droid at least it should be on par with devin” [source](https://twitter.com/1159835302275346433/status/2104318568980689170)
- Complaint, 2026-09-26, @FactoryAI (X): “@hataiit9x @droid @factoryai the harness quite sucks :)” [source](https://twitter.com/1797536317716525056/status/2103754025057595878)
- Complaint, 2026-09-23, @FactoryAI (X): “@factoryai delegated bug fixes to droids. so now instead of one dev complaniing about the codebase, we got ten devs complaining about the codebase and a droid that's gonna fuck it up anyway. genius.” [source](https://twitter.com/1873462291565277185/status/2102817203775033608)

### OpenAI Codex

- Praise, 2026-09-27, r/codex (Reddit): “im a literal software engineer and use luna to work on enterprise codebases, it reads through hundreds of files for me, researches for me and helps me prototype. also reads linear tickets and helps me make pr descriptions quickly all the time if you couldn't use it to push something, you are facing what we call a skill issue my friend.” [source](https://www.reddit.com/r/codex/comments/1wr4e20/gpt_6_sol_is_the_new_opus_47/pca29x7/)
- Praise, 2026-09-27, r/codex (Reddit): “astra already does this for me. the limitation for me is actually my own imagination, and subjective ui. i’ll create a detailed prd that is say 30 pages long. it even upgrades things i didn’t think of and i agree. for example, for roles, it integrated mfa with authenticator for admin profiles. i didn’t even ask. but then, i can’t help but keep iterating… lets add export here. lets go ahead and add a simple email cms to customize templates. heck,” [source](https://www.reddit.com/r/codex/comments/1wqula3/have_you_heard_about_gpt6_aeon/pca3dhn/)
- Praise, 2026-09-27, r/codex (Reddit): “not every way. astra is a larger model and it shows in stuff like 3d generation, it makes by far best and most logical layouts and gets closest to references unattended. from coding perspective it also does a bit better in some insane tasks like "my mouse scroll button sometimes goes in the wrong direction, can you rewrite it's whole firmware so it stops doing that in arm assembly". astra also does not auto reject infosec questions as much, anthr” [source](https://www.reddit.com/r/codex/comments/1wr69nb/see_you_soon_guys_probably/pca4ugt/)
- Complaint, 2026-09-27, r/codex (Reddit): “i work in bioinformatics and completely agree. i've basically given up using it and go for 5.6 or claude. i don't understand how they missed the mark this badly.” [source](https://www.reddit.com/r/codex/comments/1wr4e20/gpt_6_sol_is_the_new_opus_47/pc9t5dh/)
- Complaint, 2026-09-27, r/codex (Reddit): “"no buts its <isbn>x efficient" *dumber than qwen 27b*” [source](https://www.reddit.com/r/codex/comments/1wr4e20/gpt_6_sol_is_the_new_opus_47/pc9te01/)
- Complaint, 2026-09-27, r/codex (Reddit): “they really need to reset the model stack. i mean i'm sure that each generation between 5.5 and 6 has gotten better at something. i'm not exactly sure what because it basically is unusable for serious coding. literally lost in a c++ code base. mangles everything it touches. takes 15 minutes on a short run. wildly expand scope. invents in ludicrous defensive checks against impossible situations. continually routes c++ code/data to javascript ui” [source](https://www.reddit.com/r/codex/comments/1wr1oir/they_are_aware_and_working_on_it_apparently_just/pc9zpqi/)

### Kiro

- Praise, 2026-09-27, r/kiroIDE (Reddit): “my two cents on both from data science product development pov: claude code: i have been using claude code since it's first release. i must say it has improved a lot from different modes to harness improvements. the follow up questions which it asks you in plan mode is similar to plan mode in kiro. while claude code earlier was on cli only on windows later it got major upgrade to better ui as well integrated in vs code. i honestly feel like it r” [source](https://www.reddit.com/r/kiroIDE/comments/1wqueq0/claude_pro_vs_kiro_pro_subscription_which_is/pcc2obm/)
- Praise, 2026-09-27, r/kiroIDE (Reddit): “so true, opus 5.5 is extremely good at coding, spatial reasoning and best of all, it speaks human 😙” [source](https://www.reddit.com/r/kiroIDE/comments/1wrgezk/when_are_we_getting_new_gpt6_sol_and_luna_models/pcean29/)
- Praise, 2026-09-27, @kirodotdev (X): “i created this motion video with just one prompt using @kirodotdev 👀 and honestly, the result is seriously impressive. you might not even need claude code pro — i just used claude opus 5.5 directly inside kiro ide. <strict_link>” [source](https://twitter.com/1260863005362802688/status/2104225110425227562)
- Complaint, 2026-09-27, r/kiroIDE (Reddit): “any model quality is better on any other model harness than kiro.” [source](https://www.reddit.com/r/kiroIDE/comments/1wqueq0/claude_pro_vs_kiro_pro_subscription_which_is/pcapp9e/)
- Complaint, 2026-09-27, r/kiroIDE (Reddit): “the reason your app returned 0 results isn't because you did something wrong. it's because vercel runs on shared cloud ip ranges that search engines like duckduckgo aggressively block the second automated scripts try to scrape them. on the image recognition side, kiro gave you slightly outdated advice. you don't need a pricey setup just to pull text off a box. modern lightweight vision models (like gemini 2.0 flash or claude haiku) cost fractions” [source](https://www.reddit.com/r/kiroIDE/comments/1wr7rya/i_built_something_but_have_no_idea_what_im_doing/pcaz0zv/)
- Complaint, 2026-09-27, r/kiroIDE (Reddit): “i work for amazon and is somewhat “strongly recommended” to use kiro. i still use claude code at work and at home. so much better … (auto classifier, transcript details, integrated tooling, open source tooling, cli features, general stability, model fallback, sub agents control, etc …)” [source](https://www.reddit.com/r/kiroIDE/comments/1wqueq0/claude_pro_vs_kiro_pro_subscription_which_is/pccqcjf/)

### Cline

- Praise, 2026-09-26, @cline (X): “@cline oh wow nextjs specific. been knocking my head on the wall rewriting legacy pages router to app router.” [source](https://twitter.com/326659472/status/2103645823083168184)
- Praise, 2026-09-26, @cline (X): “@cline really good! after integrating cline for free, it crushes the next.js tasks of kimi k3—pixel canary action is really fast.” [source](https://twitter.com/1744672948135321600/status/2103663198188749016)
- Praise, 2026-09-26, @cline (X): “@cline beats kimi k3, ties gpt-6 astra, costs nothing. the business model is 'we'll figure it out', which is also my business model, so i can't judge” [source](https://twitter.com/1475598443779444739/status/2103708005372239980)
- Complaint, 2026-09-26, @cline (X): “@cline this model is just unusable, don't waste your time guys! <strict_link>” [source](https://twitter.com/1760047654027796481/status/2103851905503879464)
- Complaint, 2026-09-25, @cline (X): “@cline i mean it snice but i compared both normal deepseek and yours, why yours is 10x worse ? yo uare giving it for free but at 1 bit q ?” [source](https://twitter.com/1529503683880226816/status/2103525031523619136)
- Complaint, 2026-09-24, @cline (X): “@cline thank you, but something wrong with the output it gives <strict_link>” [source](https://twitter.com/1885218934145630208/status/2103171853355507784)

### OpenCode

- Praise, 2026-09-27, r/opencodeCLI (Reddit): “muse 1.3 xh > minimax 3.1 in my experience” [source](https://www.reddit.com/r/opencodeCLI/comments/1wqqt8f/longcat25preview_is_now_free_on_opencode_for_two/pca1mvi/)
- Praise, 2026-09-27, r/opencode (Reddit): “1.3 is insane. think you have to use all models to know which one to use. some models do not do well on certain projects. or interments.” [source](https://www.reddit.com/r/opencode/comments/1waq3e5/muse_spark_13_free_is_ass/pca29g7/)
- Praise, 2026-09-27, r/opencodeCLI (Reddit): “ability to to write large codebases or with them with correctness in this case” [source](https://www.reddit.com/r/opencodeCLI/comments/1wrnxw6/what_is_the_best_opencode_free_model/pcebwtr/)
- Complaint, 2026-09-27, r/opencode (Reddit): “luna is not frontier.” [source](https://www.reddit.com/r/opencode/comments/1wqj8p0/currently_which_is_the_best_model_on_opencode_for/pc9wmwf/)
- Complaint, 2026-09-27, r/opencode (Reddit): “than he can undetstand opencode's is dumber” [source](https://www.reddit.com/r/opencode/comments/1wqxyl4/deepseek_v41_performance_is_much_worse_than/pcao1zx/)
- Complaint, 2026-09-27, r/opencode (Reddit): “thanks for the response. giving an agent a bound and independent task is always tricky one, i don't know why but agent surely messes up something specially in the end.. not sure this happens with me or with everyone. last night also i was about to call it a day, but agent messed up big time and i need to provide it a hand holding to fix the issue.” [source](https://www.reddit.com/r/opencode/comments/1wraeev/need_help_with_big_project_tasks/pcbvfzz/)

### GitHub Copilot

- Praise, 2026-09-27, @GitHubCopilot (X): “thank you @githubcopilot copilot. you are amazing with the new gpt-6 models. about 100 prs solved, 5 very difficult issues, and 250 dependabot alerts. <strict_link>” [source](https://twitter.com/14186604/status/2104256846811021369)
- Praise, 2026-09-27, r/codex (Reddit): “i concur, luna 5.6 is an amazing model, and basically free. at work, i have only 100$ monthly limit in copilot, which is not much at api prices, and because of that, i need to be really cost-conscious and can't just vibecode yolo with expensive models. i use luna a lot (mostly on high), and it is yet to fail me, my coworkers share the same opinion. i can't say much about luna 6 since i did not have enough time with it to form an opinion. luna is” [source](https://www.reddit.com/r/codex/comments/1wr4e20/gpt_6_sol_is_the_new_opus_47/pcdh0ht/)
- Praise, 2026-09-26, r/GithubCopilot (Reddit): “explains why i can get copilot to help with my short game.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wq0nsb/does_mr_meeseeks_represent_ai/pc3q0ud/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “i can't compare them. two weeks into copilot, work moved us to claude after all developers requested it. it was pretty crap, and always behind. we haven't tried the new usage billing. i was a huge claude advocate while using it on my personal accounts. as the other person mentioned, their hooks, environment, and their tooling is just amazing. but at work, the claude token limits killed us. we liked the models better, but ended up going with grok” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrgtzy/comparison_with_copilot_cli/pcgi328/)
- Complaint, 2026-09-27, r/windsurf (Reddit): “they are not in the same class of capability. copilot is not a very good harness” [source](https://www.reddit.com/r/windsurf/comments/1wr7p7l/why_isnt_windsurf_available_in_vscode_anymore/pcd4mf5/)
- Complaint, 2026-09-27, r/ClaudeAI (Reddit): “i think eventually, it can be created by a group of teachers and have them share with one another. my district is small like 23,000 students. we have a robust network between the campuses and when one of hears something like this , what will happen we will all learn, then organize and split the work. if anyone knows teachers, we are resourceful like no one else . i am also grateful ,cause our district just got us claude for us to experiment with” [source](https://www.reddit.com/r/ClaudeAI/comments/1wr268d/a_different_kind_of_opus_55_video_prompt/pcahycz/)

### Zed

- Praise, 2026-09-25, @zeddotdev (X): “@raulvk @zeddotdev it legit runs everything they build” [source](https://twitter.com/1806175394388963328/status/2103518215347257595)
- Praise, 2026-09-25, @zeddotdev (X): “@zeddotdev feature request: is it possible to partially disable ai feature? i need to disable ai features except the edit prediction. i really like zed's edit prediction feature.” [source](https://twitter.com/42179249/status/2103590820553294064)
- Praise, 2026-09-25, @zeddotdev (X): “@theo the ai parts of @zeddotdev tuck away nicely, and i believe can be disabled pretty easily. the ai autocomplete is nice, and the editor itself is really fast” [source](https://twitter.com/1834266246088400901/status/2103612073192141004)
- Complaint, 2026-09-27, r/ZedEditor (Reddit): “i mean it’s behind a flag and setting. of course it’s not good yet.” [source](https://www.reddit.com/r/ZedEditor/comments/1wr5qmt/jupyter_notebook_support_in_zed_is_not_good/pc9x3jk/)
- Complaint, 2026-09-27, r/ZedEditor (Reddit): “is still experimental. you can use it, and it works for most of things but it's not complete yet” [source](https://www.reddit.com/r/ZedEditor/comments/1wr5qmt/jupyter_notebook_support_in_zed_is_not_good/pcavhh5/)
- Complaint, 2026-09-27, r/ZedEditor (Reddit): “i don't know how many of you use zed for running jupyter notebooks - so many issues and vs code support is much better.” [source](https://www.reddit.com/r/ZedEditor/comments/1wr5qmt/jupyter_notebook_support_in_zed_is_not_good/)

### Google Antigravity

- Praise, 2026-09-27, r/google_antigravity (Reddit): “use remotion skills. gemini does it pretty well infact.” [source](https://www.reddit.com/r/google_antigravity/comments/1wr9gwg/how_can_i_make_a_short_animation_video/pcasdsa/)
- Praise, 2026-09-27, r/google_antigravity (Reddit): “so curiously agy told me to use vertex veo/ ai and imagen - i can ask it obviously ask why it didn’t say remotion but for the layman can you explain the diff? the quality is absolutely phenomenal it linked to my gcloud made skills for both so i can say much like generating ui skill > create an image and it links to image gen or create video it links to veo. it also does multi shots and stitch etc. lets me know the estimated cost. i would assume s” [source](https://www.reddit.com/r/google_antigravity/comments/1wr9gwg/how_can_i_make_a_short_animation_video/pcbhhoj/)
- Praise, 2026-09-27, r/google_antigravity (Reddit): “i also tried it (just temporarily) and i gotta admit, it's great. but it quite a lot of money for me, so i got the 18 months pro plan for free through jio too. idc if it's bad or some shit, it's free, and gets most of my project and work done.” [source](https://www.reddit.com/r/google_antigravity/comments/1wqme6m/wtf_is_going_on/pcc775a/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “i did try this on antigravtiy gemini 3.8 flash medium on from your suggestions, i asked it to make a 10 seconds promotional 2d animation video for antigravtiy with animal characters in the video. [<strict_link>. i pointed it to use remotion and use whatever tools and skills needed to make the video visually coherent i am not satisfied with the quality, am i missing something?” [source](https://www.reddit.com/r/google_antigravity/comments/1wr9gwg/how_can_i_make_a_short_animation_video/pcbfoti/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “i did try this on antigravtiy gemini 3.8 flash medium on from your suggestions, i asked it to make a 10 seconds promotional 2d animation video for antigravtiy with animal characters in the video. [<strict_link>. i pointed it to use remotion and use whatever tools and skills needed to make the video visually coherent i am not satisfied with the quality, am i missing something?” [source](https://www.reddit.com/r/google_antigravity/comments/1wr9gwg/how_can_i_make_a_short_animation_video/pcbfux4/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “not skill issue u would be surprized how much dumb gemini is until u use other frontier” [source](https://www.reddit.com/r/google_antigravity/comments/1wrbsby/antigravitys_gemini_models_are_worst_in_among_all/pcbglnw/)

### Conductor

- Praise, 2026-09-24, r/conductorbuild (Reddit): “i use workflows for pr requests. i can see the code and open the editor as well. i haven't tried the codex app. everything i need is right there.” [source](https://www.reddit.com/r/conductorbuild/comments/1wp7deq/favourite_thing_in_conductor/pbuky7f/)
- Praise, 2026-09-06, r/ClaudeCode (Reddit): “it’s a great worker for me, but it over thinks as an orchestrator. great for hill climbing, not so good at delegation. from my ai team l lead: route by verifiability, not difficulty: running the claude 5 family as an orchestra instead of a chat window someone asked how i structure inference across the claude 5 family. short version: route by verifiability, not difficulty. the question is never "is this task hard?" it's "can i mechanically check” [source](https://www.reddit.com/r/ClaudeCode/comments/1w966g8/opus_5_orchestrator_creates_work_faster_than_it/p8890r0/)
- Praise, 2026-09-01, @conductor_build (X): “@alphacolin @herdrdev @conductor_build is my daily driver, basically does all of these (local + cloud options too)” [source](https://twitter.com/2000470921329709056/status/2094889566314369265)
- Complaint, 2026-09-21, r/conductorbuild (Reddit): “very limited functionality i believe” [source](https://www.reddit.com/r/conductorbuild/comments/1wkn35i/any_word_on_that_ios_app_question/pb8t6ss/)
- Complaint, 2026-09-16, @conductor_build (X): “@conductor_build me too, but then i didnt like how few things were done, so i built my own ade @zuse_sh” [source](https://twitter.com/3304494590/status/2100174141320310804)
- Complaint, 2026-09-15, @conductor_build (X): “@immadsahin @michelerivacode @conductor_build this ^^^ and don’t even get me started on their agentic coding integrations lol, all you can do is tag an agent (no control over model, branch, local/worktree, etc) and hope for the best - can’t even give it a prompt 🤣” [source](https://twitter.com/2017249370018844672/status/2099951255179366653)

### Grok Build

- Praise, 2026-09-20, r/LocalLLaMA (Reddit): “i have 4 of the tesla v100 32gb cards running in my rig. something that i discovered is that the current version of grok build is uncannily good at setting up these cards tuning them selecting functioning models to download and getting it all up and running under lennox. i'm presently hosting three models qwen 3.8, qwen 3.6, and nemotron 3.5 with results that continue to surprise me. after i had grac set up the cards then i had grok build reconfi” [source](https://www.reddit.com/r/LocalLLaMA/comments/1wl680s/finally_got_qwen_38_next_running_on_my_v100_6gpu/paw9a6m/)
- Praise, 2026-09-13, r/cursor (Reddit): “its wierd seeing the many bad experiences from all of you with grok 4.6. i actually had t opposite experience. i was unhappy with the high prices of claude enterprise in my company. bei g the guy responsible for rolling out ai to everybody i was looking for alternatives. i started to try grok build and cursor and after initially having a problem with trusting their models in any way i became pretty convinced. grok 4.6 was so good i even decide” [source](https://www.reddit.com/r/cursor/comments/1weddl3/thats_has_happened_to_cursor/p9i8zid/)
- Praise, 2026-09-07, r/opencodeCLI (Reddit): “i don't use kimi k3. but i can tell you that glm 5.3 flash > muse spark 1.3 xhigh > qwen 3.8 flash. qwen 3.8 flash makes mistakes with confidence and is also slow, takes a lot of detours, and makes you spend twice as many tokens despite being "cheaper." muse spark 1.3 xhigh is intelligent but very lazy; it's a terrible agent to work with. it forgets to call tools and always looks for the quickest solution, never considering different perspectives” [source](https://www.reddit.com/r/opencodeCLI/comments/1w5zhbn/kimi_k3_vs_glm_53_vs_qwen_38_max_vs_muse_spark_13/p89e49c/)
- Complaint, 2026-09-21, r/ClaudeAI (Reddit): “interesting experience. "grok 4.6 and grok cli respond very fast, but they often start working before the prior thinking is sufficient. it tends more toward making local patches on known problems, rather than actively improving the overall architecture" i have this exact problem with gemini as well. i trying to force is to think of general architecture over the local patches. but, it always reverses to easy and narrow patches. only claude model” [source](https://www.reddit.com/r/ClaudeAI/comments/1vxzbij/after_using_claude_grok_46_and_gemini_37_flash_in/pb4qjfl/)
- Complaint, 2026-09-18, r/codex (Reddit): “objectively, no. anthropic's only good model is fable. you can only use 50% of your usage on it. their other models are both bad and extremely overpriced. sonnet costs like 9-15x more per task than luna. luna max actually performs similar to opus on medium. so you get like 20x more work done with luna than with opus but 50% of your subscription is basically locked to opus. and it's $100, not $20. value wise, what i'd suggest the most for someone” [source](https://www.reddit.com/r/codex/comments/1wjhw4l/is_claude_a_value_switch_now/paj18pd/)
- Complaint, 2026-09-12, r/codex (Reddit): “i tried grok 4.6 via grok build cli for mac cuz they gave me 3 day free trial, is soooo bad, id rather use gpt 5.4 than grok.” [source](https://www.reddit.com/r/codex/comments/1we1a4j/tibo_tibo_tibo/p9aj19k/)

### Warp

- Praise, 2026-09-15, r/AI_Agents (Reddit): “warp terminal [warp.dev](<strict_link>) has its own agent as terminal shell so you can quickly access it by typing prompt into terminal. it has byok so you can quickly launch it for your needs. quite good for simple tasks, bash, and system manintance” [source](https://www.reddit.com/r/AI_Agents/comments/1wh25jc/whats_a_minimal_and_extremely_fast_cli_coding/pa1xhny/)
- Praise, 2026-09-09, @warpdotdev (X): “@bholmesdev @warpdotdev one of my earlier attempt opened a read only page. in a test session it did open an input too. that is a good feature.” [source](https://twitter.com/8104092/status/2097684654958481909)
- Praise, 2026-09-03, @warpdotdev (X): “i’m not sure but i find @warpdotdev terminal much smoother and faster for local system work compared to hermes. i asked hermes to look into this open-design repo and configure it for my deepseek harness agent. it took much longer even though i provided exa ai and firecrawl api for web search. it still took over two minutes and gave me a detailed but cluttered result. in contrast, warp terminal did it in 30 seconds with a much smoother concise rep” [source](https://twitter.com/859077042129600516/status/2095579049729151354)
- Complaint, 2026-09-24, @warpdotdev (X): “@real_spencercjh @warpdotdev @xuanwo @tldraw @tualatrix does warp support ocaml toplevel now? this was the reason that stopped me from using it back then.” [source](https://twitter.com/19064875/status/2103070329484755172)
- Complaint, 2026-09-02, @warpdotdev (X): “i always come back to @warpdotdev not sure what it is.. the blend of the ui + libghostty. i hope file explorer and /agent gets improved.... no more side quests :) warp-cli, factories... i wish agent can manage the session for me.. right now it is useless even on glm5.3 flash” [source](https://twitter.com/8104092/status/2095286699328782447)

### Augment Code

- Praise, 2026-09-23, r/ExperiencedDevs (Reddit): “i would definitely start getting comfortable with it, you don't need to let it be an agent and do everything for you. it can be fun to figure out where that boundary is. i work in a small team that owns and maintains several software systems, from vendor based to integration layers and some full stack software with both internal and customer users, so knowing everything about everything is effectively impossible. another case is i have it integra” [source](https://www.reddit.com/r/ExperiencedDevs/comments/1wo4d8p/job_requiresuses_very_little_ai_sinking_ship_or/pblbka0/)
- Praise, 2026-09-17, G2 (G2): “q: what problems is the product solving and how is that benefiting you? a: i use it to build, refactor and manage a large codebase. work like that gets tedious when changes touch a lot of files, and having augment code help with it makes those tasks more manageable. it's only been a week or two, so i don't have hard numbers, but it has made refactoring work feel less of a chore. q: what do you like best about the product? a: the main thing i like” [source](https://www.g2.com/products/augment-code/reviews/augment-code-review-13367494)
- Praise, 2026-09-01, G2 (G2): “q: what problems is the product solving and how is that benefiting you? a: my workflow involves generating code across multiple platforms, which often means the output needs refinement before it's actually usable. augment code solves that last-mile problem, it takes rough, generated code and polishes it into something cleaner and more reliable. the benefit is that i'm spending less time manually reviewing and fixing output, and more time actually” [source](https://www.g2.com/products/augment-code/reviews/augment-code-review-13394203)
- Complaint, 2026-09-11, r/vibecoding (Reddit): “i used augment code almost from the beginning, after spending £300 on credits in one month after the price hike i cancelled and moved away. i use a mac, intellidea and do mostly flutter apps, databases, websites - not basic apps either. i moved to zencoder, a platform i never really took any notice of because augment code was so good (i thought). but for me it was the best switch ever. for my use case it outshines augment code in every way. its f” [source](https://www.reddit.com/r/vibecoding/comments/1vqvmlv/did_anyone_ever_use_augment_code_im_looking_for/p94qwc4/)
