# Quality got worse or better over time (`models.quality_drift`)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/criterion/models.quality_drift

Area: [Choosing models](https://feedbackbench.com/criteria/models.md)

**Definition.** The post compares the agent or one of its models with an earlier time and says it got worse or better: 'nerfed', 'dumber since last week', 'better than at launch'. An explicit comparison with the past is required.

**Boundary.** Not this: see general.unspecific for a verdict with no comparison over time. Not this: see [Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md) for what it can do now. Not this: see [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) for changed quotas or prices.

Rated author-weeks, all agents: 5153. Complaint share: 78%.

## The brief

Written by Claude Opus 5.5 from 83 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**Users judge every release against last month, and Codex loses most.**

TL;DR:

- OpenAI Codex draws the bulk of drift complaints, centered on newer versions underperforming the ones they replaced.
- Claude Code rates better than peers, yet its users file the most requests to stop nerfing.
- Google Antigravity and Devin earn praise for upgrades that land. Cursor users say Grok 4.7 regressed.

In plain terms: In practice, users pin an older version that worked, watch for silent mid-week changes, and re-test after every release. When a new model ships worse than the last one, many switch agents rather than wait.

### How it breaks

- **New version ships worse than old** ([Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md)). The sharpest complaint is a version bump that reasons worse than its predecessor, pushing users to roll back by hand.
  Posts describe a newer model doing tasks halfway and needing several prompts where the old one finished in one. The pattern crosses agents. Codex and Copilot users flag the same GPT release, Cursor users say Grok 4.7 fell below 4.6, and Antigravity users say 3.8 turned slower and loopier while 3.7 stayed fine. The fix users reach for is manual pinning, which only works while the old version stays available.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-25: “luna 6 is much worse than luna 5.6. i really noticed that version 6 has poor reasoning skills because it does things halfway. to complete a task, it has to send several prompts. i went back to using 5.6 and hope they don't change it so i can continue using it.” [source](https://www.reddit.com/r/codex/comments/1wptlvk/gpt6_luna_in_codex_does_zero_reasoning_high/pby9ao6/)
  - Complaint, GitHub Copilot, r/GithubCopilot, 2026-09-26: “gpt 6 luna is worst than 5.6 for coding, becarefull to adjust your model when necesarry <strict_link>” [source](https://www.reddit.com/r/GithubCopilot/comments/1wq39rv/did_they_10x_the_cost_of_luna_56_since_the/pc4rmf2/)
  - Complaint, Cursor, r/cursor, 2026-09-24: “at this point i really don't know what's going to keep people coming to cursor at all. just some weeks ago i was fine with grok 4.6, but now openai and anthropic (especially anthropic) are advancing so much that i am starting to regret getting another month of cursor. grok4.7 is such a failure that it was able to be worse than the previous model, which is absolutely terrible because i also can't 'pick' my model in grokbot, which is one of the reasons i wanted to stay with cursor. either way, now i just hope that 4.8 will be a better model so that i can have fun with grokbot again, even though i highly doubt it” [source](https://www.reddit.com/r/cursor/comments/1wp6j44/is_this_supposed_to_be_good_news/pbtfr06/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-21: “3.8 now is not only 3 times more expensive than what it used to be, but also 5 times slower and much more stupid. on each thinking level. what i noticed is that it enter meaningless 'analysis' loops of the same range of the same file all the time. at the same time 3.7 is pefectly fine - super fast and token-efficient.” [source](https://www.reddit.com/r/google_antigravity/comments/1wmm9nj/what_is_going_on_it_drains_the_quotea_like_crazy/pb84xta/)

- **Same model, dumber this week** ([Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md)). Users report the same named model getting worse between sessions with no announcement, and read it as a silent nerf.
  These posts carry no version change. Users say the model they used yesterday now behaves like an older generation, or stays degraded even after a fix shipped. Antigravity users compare notes across regions and conclude something broke recently. Without a changelog, users treat their own day-to-day sessions as the only benchmark, and community threads amplify each report.
  Evidence:
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-17: “opus today is definitely a secret nerfed checkpoint. we got opus 3 at the wheel” [source](https://www.reddit.com/r/ClaudeCode/comments/1wimzqe/seen_a_few_posts_comments_on_anthropic_doing_ab/pac3lkk/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-05: “sol is horrible and im so furious they dosturbed that. even after the release its still nerfed. no ones gonna use ill drain your qouta in 10 mins astra pver 5.6 sol . most people will still use sol. give me back my sol” [source](https://www.reddit.com/r/codex/comments/1w7qwt3/sol_v_astra/p80eesy/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-23: “yes, it's been happening to me for about a week, but only 3.8. 3.7 seems to stay stable. i also see similar reports almost every day lately. so you are not alone and you are right assuming that something broke recently.” [source](https://www.reddit.com/r/google_antigravity/comments/1woea7p/30_minutes_for_a_single_prompt_20_40_from_the_5/pbmm8kt/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-15: “it is incredibly throttled compared to what it was before. belived it to be regional but seems to be happening in usa and brazil.” [source](https://www.reddit.com/r/google_antigravity/comments/1wguex8/just_me_or_is_38_flash_still_slow_af/pa017dr/)

- **Monitoring the model instead of working** ([Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md)). Undisclosed changes turn users into part-time evaluators, and some pick agents by which vendor they believe does not degrade.
  Users describe constantly watching response quality and lining up fallbacks. Others say changes happen on the sly while prices climb. One defender argues downgrades are not hidden, just tied to effort levels and pricing. Disclosure requests here mostly target OpenAI Codex. The other side of the same distrust shows up when users wonder whether positive posters are even on the same model.
  Evidence:
  - Praise, Cursor, r/codex, 2026-09-21: “i've been dealing with the same frustration. it's causing me a lot of stress and a complete loss of productivity because i have to look for alternatives while constantly monitoring codex and the quality of its responses. my quota reset yesterday, and it's hard to just ignore it, but for now i'm using cursor. as far as i know, they don't degrade their models(imagine grok degrading lol).” [source](https://www.reddit.com/r/codex/comments/1wmdc58/i_just_want_codex_to_feel_reliable_again/pb6vl7i/)
  - Complaint, OpenCode, r/opencode, 2026-09-16: “prices keep going up, performance is dropping, and subscriptions offer less for the same price... it’s unsustainable; they know it, and they’re making changes on the sly. it’s a shame. the only reason i’d pay for ai is for a cheap subscription that allows for heavy use of chinese models... i’d be willing to pay $20 a month, provided it meant practically unlimited monthly access to cheap chinese models... but no—instead of moving in that direction, we’re heading the opposite way: less usage and higher costs... it’s a shame.” [source](https://www.reddit.com/r/opencode/comments/1whtq1k/glm53_nerfed/pa4ys2v/)
  - Praise, OpenAI Codex, r/codex, 2026-08-31: “point taken, they could be more transparent. but i think you're also demanding more than the technology can provide. it's next to impossible to calculate how hard a task is before starting it, that's why they had to track back on their automated model switcher introduced with gpt-5. and i really don't think they are hiding downgrades. newer models have higher potential effort levels, and cost more per token. if you want to have an experience like half a year ago, terra low/medium should be smarter than the models back then, and allow for a lot of usage. plus, if you want concrete billing per task, i'm sure other services offer that for certain tasks. it'd just have to be less flexible than chatgpt/codex to make it possible to offer more predictable pricing.” [source](https://www.reddit.com/r/codex/comments/1w3a75y/weve_all_been_sludged/p70wupv/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-23: “my experience exactly. whenever anyone posts something positive, i wonder if they’re using a different model than i am.” [source](https://www.reddit.com/r/google_antigravity/comments/1woay30/yo_google_fix_your_product_its_not_funny_anymore/pbmxe41/)

- **Upgrades that actually land** ([Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md)). Praise appears when a new version clearly beats the one it replaced, often after earlier releases disappointed.
  Antigravity users say Gemini 3.8 beats the 3.5 and 3.6 releases that let them down, and one says 3.8 flash got far smarter overnight. Claude Code users report fewer mistakes and smoother iteration on old work after a new model arrived. Codex users who stay positive call releases slightly better at a lower price. The good news is real but framed as relief, not trust.
  Evidence:
  - Praise, Google Antigravity, r/google_antigravity, 2026-09-07: “personnellement je fais des plans détaillés avec opus, gemini 3.8 est étonnamment bon comparé à 3.5 et 3.6 qui m'avaient vraiment déçus” [source](https://www.reddit.com/r/google_antigravity/comments/1w96dbj/antigravity_with_gemini_38_flash_highly_unreliable/p8d3cyt/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-27: “i just had it check things that i set up a couple of months ago and improved and iterated on it. things are working \*so\* much smoother and more efficient. this is the first model i'm seeing that doesn't make me want to pull my hair out after it completes a task. it actually, for the *most* part, doesn't fuck up and makes fewer mistakes than before. also, remember to audit your codebase every now and then, especially when a newer model comes out!” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrgldx/remember_to_have_opus_55_review_your_old_workflow/)
  - Praise, Google Antigravity, r/google_antigravity, 2026-09-27: “i've been heavily using 3.6 flash in android studio, 3.8 flash in antigravity and qwen-3.5-122b-a10b locally hosted all weekend on some projects, and i swear 3.8 flash in antigravity got about 4x smarter overnight. the comparison to the other two models which i have been using side-by-side is astonishing. i am almost certain they are testing gemini 4 pro stealthily--it is incredibly capable. feels like the opus 5 family models i use at work.” [source](https://www.reddit.com/r/google_antigravity/comments/1wrjwli/gemini_4_in_antigravity/pcg70i4/)
  - Praise, OpenAI Codex, r/codex, 2026-09-23: “i don't need a hype. the release is good. little bit better models i already use with a cheaper price. looks good for me.” [source](https://www.reddit.com/r/codex/comments/1wo3ixp/how_do_you_think_openai_will_respond_next_week/pbk7yqz/)

### Who stands out

- **OpenAI Codex (weaker)**. Codex carries the heaviest drift complaint load, with users saying new releases underdeliver and pushing them toward rivals.
  Posts split the lineup by tier and version, preferring older releases over newer ones and calling some current models unusable or quota-hungry. Users ask to restore earlier quality, roll back versions and disclose downgrades more than any other agent's users. Praise exists, from users who see steady gains, but it is outnumbered.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-25: “i agree, after the quality of opus 5.5 and how much generous limit it has. and openai gpt 6 sol being not usable. and astra being good. burns my limit too fast. i am switching to claude, once my openai sub expires next week” [source](https://www.reddit.com/r/codex/comments/1wp5x1a/gpt6_sol_is_not_good/pbxofzx/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-22: “luna i hope so. i’ve never understood or realised why terra exists. it does neither of the things luna or sol does well. planning in sol high and executing in luna xhigh, for me at least, bears so many cost benefits and better completion.” [source](https://www.reddit.com/r/codex/comments/1wl80mt/gpt_6_luna/pbbni9g/)
  - Praise, OpenAI Codex, r/codex, 2026-09-09: “codex. mostly writing, to fix itself and other models. now it's finally writing good prompts for other agents, and it stopped ignoring the obvious next step. i can just say "okay." and it'll keep going and take the lead. it still fails 20% of the time, but it's getting better. when they cook a version of astra with rl as good as 5.6 sol it's going to be glorious, seriously. this one is definitely gpt-5.5-like.” [source](https://www.reddit.com/r/codex/comments/1wbaote/how_are_you_guys_using_gpt_astra_right_now/p8okrtq/)
  - Praise, OpenAI Codex, r/codex, 2026-09-07: “fair enough, i just read reddit and watch it be tested lol, not gonna use my usage i pay for just to test it when i'm probably won't use it since last models were just fine for my workflow.” [source](https://www.reddit.com/r/codex/comments/1w99ef8/i_dont_really_understand_the_point_of_astra_on/p8btzvn/)

- **Claude Code (mixed)**. Claude Code rates better than peers on drift, yet its users lead requests to stop nerfing models.
  Fans say the current Opus improved sharply over earlier versions and makes fewer mistakes. Critics say the model swaps into a weaker checkpoint on some days, and tie it to shifting usage. The volume of stop-nerfing requests, 59 author-weeks, shows the fear stays alive even when the experience is good.
  Evidence:
  - Praise, Claude Code, @ClaudeDevs, 2026-09-24: “@vcg_run @claudedevs it’s peak and improved extremely from opus 5” [source](https://twitter.com/1734444968259829760/status/2103080649628217387)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-15: “interesting. i tried the prompt on my enterprise work plan and it failed, but on my personal max 20x plan the prompt worked with opus but failed on sonnet. i can't check on fable because i've maxed my fable limit haha. i also have noted an improvement in opus over the last week as i've had to route more and more work to it due to fable use vanishing extremely quickly on 5.1. but it's all perception so it might not be true, and it could just as easily be due to fable having a greater influence on the overall health of the project.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wg85gb/opus_52_stealth_routing/p9vdl7l/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-08-31: “i give up. i used claude code for months. got lots done. now though, the plans are shit. usage is endlessly being fucked with on a week to week basis. opus 5 is a fucking loser. fable is alright for the 14 minutes of usage you get before it stops out, or its overly sensitive detection downgrades it back to the loser fuck. the 20x plan with or without boosted limits is a blatant lie. this was my first session of the week. maxed out in 2.5 hours and 20% of the weekly usage. do the math. boosted 50% weekly translates to 30% usage of the 20x under normal usage, or 60% of the normal 5x plan in a sub 3-hour session?” [source](https://www.reddit.com/r/ClaudeCode/comments/1w3lw7f/just_cancelled_my_claude_code_bullshit_20x_plan/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-24: “i dont care about a little nerf in intelligence as long as they don't make it speak claudish ever again 😭” [source](https://www.reddit.com/r/ClaudeCode/comments/1woba1t/opus_55/pboyqc1/)

- **Google Antigravity (stronger)**. Antigravity users mostly report Gemini versions getting better, though the same 3.8 release draws both praise and regression reports.
  Users credit 3.8 with beating earlier Gemini releases and call flash dramatically smarter. Others say 3.8 became slower, costlier and loopier while 3.7 held steady, and some sidestep Gemini by running Claude through an extension. The net reads positive but unstable.
  Evidence:
  - Praise, Google Antigravity, @antigravity, 2026-09-15: “@rodydavis @ibocodes @antigravity i just use ag ide exclusively. we still have the issue of the model confidently stating that work was completed but, in fact, didn't complete it. how would you fix that? 3.8 high is certainly better than before, overall staying with 3.1 pro high.” [source](https://twitter.com/15162579/status/2099886373507645661)
  - Praise, Google Antigravity, r/google_antigravity, 2026-09-27: “i've been heavily using 3.6 flash in android studio, 3.8 flash in antigravity and qwen-3.5-122b-a10b locally hosted all weekend on some projects, and i swear 3.8 flash in antigravity got about 4x smarter overnight. the comparison to the other two models which i have been using side-by-side is astonishing. i am almost certain they are testing gemini 4 pro stealthily--it is incredibly capable. feels like the opus 5 family models i use at work.” [source](https://www.reddit.com/r/google_antigravity/comments/1wrjwli/gemini_4_in_antigravity/pcg70i4/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-23: “my experience exactly. whenever anyone posts something positive, i wonder if they’re using a different model than i am.” [source](https://www.reddit.com/r/google_antigravity/comments/1woay30/yo_google_fix_your_product_its_not_funny_anymore/pbmxe41/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-17: “working fine for me, probably because im using the claude extension? i cant stand the gemini models.” [source](https://www.reddit.com/r/google_antigravity/comments/1wiptxc/is_it_slow_again_or_am_i_just_being_paranoid/pafjo00/)

- **Cursor (weaker)**. Cursor users blame a model regression and limited model choice for slipping quality.
  Posts say Grok 4.7 performed worse than 4.6 and that users cannot pick their model in some flows, leaving them stuck with the downgrade. Some heavy users say they see no dumbing down at all, only slight slowness. The complaints focus on the shipped model, not the editor.
  Evidence:
  - Complaint, Cursor, @cursor_ai, 2026-09-22: “@cursor_ai dont make me switch to claude? bro make 4.7 better” [source](https://twitter.com/1581256784/status/2102499093088202885)
  - Complaint, Cursor, r/cursor, 2026-09-24: “at this point i really don't know what's going to keep people coming to cursor at all. just some weeks ago i was fine with grok 4.6, but now openai and anthropic (especially anthropic) are advancing so much that i am starting to regret getting another month of cursor. grok4.7 is such a failure that it was able to be worse than the previous model, which is absolutely terrible because i also can't 'pick' my model in grokbot, which is one of the reasons i wanted to stay with cursor. either way, now i just hope that 4.8 will be a better model so that i can have fun with grokbot again, even though i highly doubt it” [source](https://www.reddit.com/r/cursor/comments/1wp6j44/is_this_supposed_to_be_good_news/pbtfr06/)
  - Praise, Cursor, r/cursor, 2026-09-20: “i'm using cursor 8-10 hours a day and i didn't notice that cursor models got dumber. slower? yes maybe a bit, but not unreasonably. i'm doing the same exact tasks every day and they seem to be quite complex, including security analysis, etc. don't see anything wrong recently.” [source](https://www.reddit.com/r/cursor/comments/1wk16qb/so_what_happened_to_cursor_in_the_past_few_weeks/paxoenz/)
  - Praise, Cursor, r/cursor, 2026-09-23: “in my experience 4.7 is just 4.6 but faster, and i never hit my quotas so it's good for me” [source](https://www.reddit.com/r/cursor/comments/1woad6b/grok_47_performance_in_cursor/pbmtian/)

- **Devin (stronger)**. Devin users praise SWE-2 as a real step up at lower cost, tempered by suspicion of benchmark tuning.
  Users call the new model stronger and cheaper, and say it beats a rival release. Skeptics recall an earlier version as benchmaxxed and ask for a reproducible eval protocol. The split is close, so the praise reads as early.
  Evidence:
  - Praise, Devin, @cognition, 2026-09-03: “@cognition lower costs with stronger performance makes this a compelling upgrade” [source](https://twitter.com/1519550915132669952/status/2095522442923983153)
  - Praise, Devin, @cognition, 2026-09-24: “people are sleeping on swe models (by cognition) honestly. it is literally far better than grok 4.7 in every possible way. and the harness (devin) is honestly amazing. @cognition you guys are cooking!” [source](https://twitter.com/1888764529766502400/status/2103104155195961481)
  - Complaint, Devin, @cognition, 2026-09-13: “@djlougen @cognition swe 2 or just fusion? i didn't try swe-2 yet but their 1.7 was hilariously benchmaxxed, so i am still biased” [source](https://twitter.com/2085263271645396992/status/2099213056790180089)
  - Complaint, Devin, @cognition, 2026-09-10: “@cognition congrats on the launch. one request: publish the eval protocol alongside the scores. environment setup, contamination controls, and a path for independent reruns. scores 'on par' without a public reproducible harness read as marketing, not measurement.” [source](https://twitter.com/2023185508428304384/status/2098102888023068954)

### Fine print

- Drift claims are user perceptions without controlled tests, and some authors admit their impression may be wrong.
- The same model names appear across several agents, so a regression reported in one harness may reflect the underlying model.

## Top requests

What users ask to add or change, most asked first. 308 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Author-weeks | Posts | Agents (author-weeks) |
|---|---|---|---|---|
| 1 | Stop nerfing or degrading models over time | 88 | 92 | Claude Code 59, OpenAI Codex 22, Cursor 3, OpenCode 3, Devin 1 |
| 2 | Fix currently degraded model quality | 41 | 43 | Claude Code 24, OpenAI Codex 12, Cursor 2, OpenCode 2, Google Antigravity 1 |
| 3 | Smarter, higher-quality models | 37 | 37 | OpenAI Codex 15, Claude Code 10, Google Antigravity 8, Cursor 3, OpenCode 1 |
| 4 | Restore earlier model quality level | 30 | 30 | OpenAI Codex 16, Claude Code 8, Cursor 3, Google Antigravity 1, Devin 1, OpenCode 1 |
| 5 | Consistent, stable model quality | 21 | 21 | OpenAI Codex 14, Claude Code 3, Cursor 3, Google Antigravity 1 |
| 6 | Roll back to earlier model version | 16 | 16 | OpenAI Codex 10, Claude Code 4, Google Antigravity 2 |
| 7 | Disclose model changes and quality downgrades | 15 | 15 | OpenAI Codex 13, Google Antigravity 1, Claude Code 1 |
| 8 | Cheaper, more efficient models without quality loss | 12 | 12 | OpenAI Codex 7, Claude Code 3, Google Antigravity 1, OpenCode 1 |
| 9 | Prioritize model quality and testing over new releases | 8 | 8 | OpenAI Codex 4, Claude Code 2, Google Antigravity 1, Cursor 1 |
| 10 | Release improved models sooner | 8 | 8 | OpenAI Codex 6, Claude Code 2 |
| 11 | Public tracking of model quality over time | 7 | 7 | OpenAI Codex 5, Claude Code 1, OpenCode 1 |
| 12 | Serve full unquantized models | 6 | 9 | OpenAI Codex 5, Claude Code 1 |

### 1. Stop nerfing or degrading models over time

- Claude Code, 2026-09-27, r/ClaudeCode (Reddit): “please, dario, don't 'optimize' opus 5.5! seriously, why cannot we just have nice things??” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrpq38/any_idea_what_anthropic_figured_out/pcfk19s/)
- Claude Code, 2026-09-27, @ClaudeDevs (X): “opus 5.5 has been so fucking great, fast, much less verbose, objective, very efficient with long shot tasks and loops, as orchestrator and less token burning. @claudeai @claudedevs give us a huge huge favor: don't dare to nerf it.” [source](https://twitter.com/2057822650752200704/status/2104022567157731336)
- Claude Code, 2026-09-26, r/ClaudeCode (Reddit): “i pray for this to not be subsized to death and won't be nerfed within 2 weeks. but i have trust issues with anthropic ngl.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqsz67/opus_55_has_absolutely_restored_value_to_the_200/pc71cif/)

### 2. Fix currently degraded model quality

- OpenAI Codex, 2026-09-27, r/codex (Reddit): “they need to release a model which fucking fixes this gpt-6.5, alongside usage limit fixes, even for the low-paying customers, especially the $100 package. otherwise people are just screwed over” [source](https://www.reddit.com/r/codex/comments/1wrftcs/gpt6_sol_is_massive_downgrade/pccqpm7/)
- OpenAI Codex, 2026-09-26, r/codex (Reddit): “if they dont fix gpt6 sol being worse than 5.6 luna immedietly im canceling and not going back.” [source](https://www.reddit.com/r/codex/comments/1wpysxs/openai_teaser_in_x/pc35k7x/)
- OpenAI Codex, 2026-09-25, r/codex (Reddit): “it will be pathetic if they don't fix 6 sol first. at this rate even sonnet 5.5 might end up being better than 5.6 sol” [source](https://www.reddit.com/r/codex/comments/1wpysxs/openai_teaser_in_x/pbzmxsv/)

### 3. Smarter, higher-quality models

- Claude Code, 2026-09-27, r/ClaudeCode (Reddit): “wtf u doing blizzard codex underpowered as fuck for 3 patches now and still buffing claude code fix ur shit or i reroll” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrg9xq/is_it_me_or_did_opus_55_get_much_better/pcd8725/)
- OpenAI Codex, 2026-09-25, r/codex (Reddit): “fix everything tbh: the models, the intelligence, but also the goddamn user limits. if people cannot afford your shit, who the fuck are you selling it to?” [source](https://www.reddit.com/r/codex/comments/1wpq44p/openai_prepares_new_500_per_month_pro_max_plan/pbzzsec/)
- OpenAI Codex, 2026-09-25, r/codex (Reddit): “both models are trash, they should’ve put the effort into making astra 6.1” [source](https://www.reddit.com/r/codex/comments/1wpu2b5/this_didnt_age_too_well/pbyi76d/)

### 4. Restore earlier model quality level

- OpenAI Codex, 2026-09-25, r/codex (Reddit): “fix the braindead ai first. they gave us an amazing model just slightly behind fable, maybe similarly capable, and then silently nerfed it.” [source](https://www.reddit.com/r/codex/comments/1wpq44p/openai_prepares_new_500_per_month_pro_max_plan/pbzyulk/)
- OpenAI Codex, 2026-09-22, r/codex (Reddit): “just bring back the capability astra had when i was using it with blender the first few days it became available. because it's now donkey balls.” [source](https://www.reddit.com/r/codex/comments/1wnh5j8/gpt_6_sol_and_luna/pbez7ct/)
- OpenAI Codex, 2026-09-22, r/codex (Reddit): “i think we've probably been using astra minor this past week, and maybe hopefully astra will go back to pre-nerf although it's still going to be a token monster” [source](https://www.reddit.com/r/codex/comments/1wndst2/gpt6_sol_luna_and_astra_minor_reportedly_just/pbeahab/)

### 5. Consistent, stable model quality

- Claude Code, 2026-09-25, r/codex (Reddit): “i feel like i'm switching from one subscription to the next and eventually either the model gets decapitated or limits increase. will there ever be a frontier model that stays the same? i am getting sick of porting my projects between programs (codex and claude code mainly).” [source](https://www.reddit.com/r/codex/comments/1wpfoxh/the_downhill_begins/pbx665h/)
- OpenAI Codex, 2026-09-24, r/codex (Reddit): “that’s the quiet part said out loud! ;) i have both 5x subscriptions and constantly alternate between the two. it’s maddening when a model starts chewing through tokens or becomes plain dumb. like some consistency would be nice, i don’t trust the same exact model version to perform the same way tomorrow because they both mess with them and hope we don’t notice.” [source](https://www.reddit.com/r/codex/comments/1wp3bzl/you_need_to_try_opus_55/pbti6y4/)
- Claude Code, 2026-09-22, r/ClaudeCode (Reddit): “it keeps happening. i don't think i have gone more than 3 or 4 days tops in the past 3 months where there were not issues. i need claude to be optimal otherwise my $200 a month is really spent in vain and i keep having to re-do the work.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wmvw5r/any_idea/pba8an1/)

### 6. Roll back to earlier model version

- OpenAI Codex, 2026-09-25, r/codex (Reddit): “did you not like opus 5.5? what are the current issues? i happily dropped claude for 5.6 sol when it came out, that was cinema. bring me back pre nerf astra and 5.6” [source](https://www.reddit.com/r/codex/comments/1wpq44p/openai_prepares_new_500_per_month_pro_max_plan/pc00kg0/)
- OpenAI Codex, 2026-09-24, r/codex (Reddit): “literally, just give me a stable 5.6 sol for code generation and astra high for deep analysis and forget about everything else.” [source](https://www.reddit.com/r/codex/comments/1wp09a6/6_sol_means_regression_to_6000bc/pbrp6jo/)
- OpenAI Codex, 2026-09-23, r/codex (Reddit): “just give us back the full quant astra 6 instead of this nerfed model. it was better than opus 5.5 already but we only got it for like 3 days before they nerfed it lmao.” [source](https://www.reddit.com/r/codex/comments/1wo0mmf/copex/pbjh3rr/)

### 7. Disclose model changes and quality downgrades

- OpenAI Codex, 2026-09-26, r/codex (Reddit): “brooo now it makes sense why i felt codex was doing way worse than usual yesterday. they should be at least transparent about it as its very sketchy to bill same for sub par model. if this continues i would consider some alternatives. claude code any better?” [source](https://www.reddit.com/r/codex/comments/1wqjxyy/i_thought_the_model_nerf_posts_were_bullshit/pc5qkvn/)
- OpenAI Codex, 2026-09-25, r/codex (Reddit): “i think every ai provider should be required by law to disclose exactly what model is being served, and it should be illegal to change the endpoint behavior under the same model/api version. if the performance regresses because it's actually a smaller model or quantized differently, they should have to call it out with data like parameter count and quantization.” [source](https://www.reddit.com/r/codex/comments/1wpvp0i/absolutely_0_doubt_in_my_mind_astra_has_been/pc12mhl/)
- OpenAI Codex, 2026-09-23, r/codex (Reddit): “that’s my biggest issue right now. they clearly changed something and it’s not the same astra - at the very least be transparent about it. there should be some regulation around this.” [source](https://www.reddit.com/r/codex/comments/1wo0mmf/copex/pbmab64/)

### 8. Cheaper, more efficient models without quality loss

- OpenAI Codex, 2026-09-27, r/codex (Reddit): “\+1 unless we get astra 6.1 with opus 5.5 quota-use, i'm out (until openai catches up again)” [source](https://www.reddit.com/r/codex/comments/1wrri5h/been_running_astra_high_100month_and_opus55_high/pcf1s9e/)
- Claude Code, 2026-09-26, r/ClaudeCode (Reddit): “it doesn't feel that much smarter for the increased token spend imo. the new, more concise writing style is nice, but if they could keep it while "nerfing" the cost and the quality to around opus 5.0 i wouldn't complain” [source](https://www.reddit.com/r/ClaudeCode/comments/1wp4ywp/please_dont_nerf_opus_55/pc5oaey/)
- OpenAI Codex, 2026-09-25, r/codex (Reddit): “i dont understand why they dont use cheaper and more efficient architecture like moe, ngram, new attention mechanisms like gated delta, instead of these huge dense models, at this point they wont loose much if they do it for something like luna and its now for a while proven that these tricks actually works for something like glm5, they can do this” [source](https://www.reddit.com/r/codex/comments/1wpspww/gpt6_astra_seems_unusable_due_to_token_burn_gpt6/pby2bsf/)

### 9. Prioritize model quality and testing over new releases

- Google Antigravity, 2026-09-26, r/google_antigravity (Reddit): “instead of releasing updates can you focus on fixing gemini 3.8?” [source](https://www.reddit.com/r/google_antigravity/comments/1wqbyzk/antigravity_cli_release_v127_v1211/pc4s50x/)
- Cursor, 2026-09-26, @cursor_ai (X): “i wasn't able to hit 30% of my @cursor_ai usage with the supergrok heavy + cursor ultra bundle on @grok 4.6. but just five days after 4.7 was released, i've already reached 45%. the massive increase in cost doesn't justify the slight improvement in intelligence. not to mention, it now takes way longer to finish tasks. @spacexai needs to focus more on the model instead of products that rely on it at the end!” [source](https://twitter.com/3245961241/status/2103902869321597219)
- Claude Code, 2026-09-26, @ClaudeDevs (X): “@ziwenxu_ @claudeai @claudedevs @bcherny please don’t plan the minor version update games rather don’t release only all version after 4.6 seemed like optimisation or rl until opus 5.5 so those version only benefit provider cost performance and hype creation” [source](https://twitter.com/2013981009927401472/status/2103863701690618348)

### 10. Release improved models sooner

- OpenAI Codex, 2026-09-27, r/codex (Reddit): “yeah, they really need to bring out the big guns for tuesday.” [source](https://www.reddit.com/r/codex/comments/1wruupn/opus_55_blows_astra_out_of_the_water_for/pcfy5yt/)
- OpenAI Codex, 2026-09-25, r/codex (Reddit): “they need ti release sol 6.5 as soon as possible honestly this doesn't even feel like sol 6 they should have named this sol 5.7 or something very disappointed with this release” [source](https://www.reddit.com/r/codex/comments/1wqat13/openai_is_becoming_incompetent/pc2msks/)
- OpenAI Codex, 2026-09-25, r/codex (Reddit): “i hope openai will drop something soon to compete, time gaps between new models getting shorter and shorter so hopefully will be released soon” [source](https://www.reddit.com/r/codex/comments/1wpxxwa/i_did_it_you_got_me/pbzd8yy/)

### 11. Public tracking of model quality over time

- OpenAI Codex, 2026-09-27, r/codex (Reddit): “we need multiple dated deepswe benchmarks over time to confirm model degradation.” [source](https://www.reddit.com/r/codex/comments/1wrfqeh/wdyt_theyre_going_to_do_with_us/pccgsfe/)
- OpenAI Codex, 2026-09-27, r/codex (Reddit): “wish they had a website where u could see the output of models over the same prompt across a period of time” [source](https://www.reddit.com/r/codex/comments/1wqjxyy/i_thought_the_model_nerf_posts_were_bullshit/pcbbfhu/)
- Claude Code, 2026-09-22, r/ClaudeCode (Reddit): “need some kind of intelligence test suite/benchmark to measure today's iq for any llm to compare with previous days” [source](https://www.reddit.com/r/ClaudeCode/comments/1wn5rpt/thats_wild_how_stupid_opus_and_fable_became_for/pbc8jcj/)

### 12. Serve full unquantized models

- OpenAI Codex, 2026-09-25, r/codex (Reddit): “yeah we basically have to wait 6 months for the quantized models that we actually get served to be as good as the first few days of a frontier model release. it’s highly frustrating” [source](https://www.reddit.com/r/codex/comments/1wpvp0i/absolutely_0_doubt_in_my_mind_astra_has_been/pc0hcq6/)
- Claude Code, 2026-09-25, r/ClaudeCode (Reddit): “please don't quantize the model and maintain performance t\_t” [source](https://www.reddit.com/r/ClaudeCode/comments/1wpnyw2/how_is_opus_55_even_real_incredible_efficiency/pbx4eba/)
- OpenAI Codex, 2026-09-23, r/codex (Reddit): “the us government literally told them to do so because it was too good, to prevent distillation and bad actors. but shouldn't we be able to verify we're not chinese citizens or taliban to get the full model reliably? i had daybreak access, verified as a us citizen, and i was still getting nerfed astra.” [source](https://www.reddit.com/r/codex/comments/1woekw4/astra_prompts_are_getting_silently_rerouted_to/pbmvqq0/)

## Every agent

| Agent | Overall rank | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|---|
| [Google Antigravity](https://feedbackbench.com/agents/antigravity.md) | =5 | Better than peers | 0.589 | 0.557–0.623 | 336 | 115 | 221 |
| [Claude Code](https://feedbackbench.com/agents/claude-code.md) | 1 | Better than peers | 0.563 | 0.544–0.583 | 1616 | 442 | 1174 |
| [Devin](https://feedbackbench.com/agents/devin.md) | =5 | Better than peers | 0.551 | 0.520–0.583 | 66 | 32 | 34 |
| [OpenCode](https://feedbackbench.com/agents/opencode.md) | 3 | Typical | 0.531 | 0.489–0.575 | 213 | 60 | 153 |
| [Cursor](https://feedbackbench.com/agents/cursor.md) | 4 | Worse than peers | 0.443 | 0.404–0.483 | 314 | 54 | 260 |
| [OpenAI Codex](https://feedbackbench.com/agents/codex.md) | 2 | Worse than peers | 0.422 | 0.403–0.440 | 2537 | 398 | 2139 |
| [GitHub Copilot](https://feedbackbench.com/agents/copilot.md) | =8 | Too few posts | – | – | 26 | 5 | 21 |
| [Cline](https://feedbackbench.com/agents/cline.md) | =8 | Too few posts | – | – | 19 | 7 | 12 |
| [Factory](https://feedbackbench.com/agents/factory.md) | =11 | Too few posts | – | – | 9 | 4 | 5 |
| [Pi](https://feedbackbench.com/agents/pi.md) | 7 | Too few posts | – | – | 6 | 0 | 6 |
| [Amp](https://feedbackbench.com/agents/amp.md) | =11 | Too few posts | – | – | 6 | 4 | 2 |
| [Warp](https://feedbackbench.com/agents/warp.md) | 15 | Too few posts | – | – | 2 | 0 | 2 |
| [Kiro](https://feedbackbench.com/agents/kiro.md) | 13 | Too few posts | – | – | 1 | 0 | 1 |
| [Grok Build](https://feedbackbench.com/agents/grok-build.md) | 16 | Too few posts | – | – | 1 | 1 | 0 |
| [Augment Code](https://feedbackbench.com/agents/augment.md) | 17 | Too few posts | – | – | 1 | 0 | 1 |
| [Zed](https://feedbackbench.com/agents/zed.md) | =8 | Too few posts | – | – | 0 | 0 | 0 |
| [Conductor](https://feedbackbench.com/agents/conductor.md) | 14 | Too few posts | – | – | 0 | 0 | 0 |

## Posts

Receipts rule: The 5 most recent praise and complaint posts per area (first 700 characters) and 3 per criterion (first 450 characters).

### Google Antigravity

- Praise, 2026-09-27, r/google_antigravity (Reddit): “skill issue. using 3.1 pro when 3.8 flash is definitely better is just dumb. and use skills there are user made skills for this kinda stuff and making a vpn is not that easy too.” [source](https://www.reddit.com/r/google_antigravity/comments/1wrbsby/antigravitys_gemini_models_are_worst_in_among_all/pcbag2m/)
- Praise, 2026-09-27, r/google_antigravity (Reddit): “yeah exactly i think so too. at night it edited a video in davinci much better than it usually does 🤔” [source](https://www.reddit.com/r/google_antigravity/comments/1wrjwli/gemini_4_in_antigravity/pcd3zyv/)
- Praise, 2026-09-27, r/google_antigravity (Reddit): “now we have gemini 3.8, 3.7 flash and all. i created my app (rust + react) which is like very big in the times of gemini 2.5 pro and 3.0 pro. i dont understand why ppl can't utlize much smarter model that we have now.” [source](https://www.reddit.com/r/google_antigravity/comments/1wrbsby/antigravitys_gemini_models_are_worst_in_among_all/pcdme2e/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “wtf u talking about 3.8flash isn't better than 3.1 pro and his right antigravity start to really fucking suck compared to the others.” [source](https://www.reddit.com/r/google_antigravity/comments/1wrbsby/antigravitys_gemini_models_are_worst_in_among_all/pcbh0g2/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “thanks for sharing! my issue might be different. but maybe also similar? it’s not burning limits, not using /boost. even though “working…” appears for hours, hardly any tokens are used (99% quota remains). i can cancel after it’s clearly stuck and ask it if it finished, it usually admits it didn’t finish then spends tokens figuring out where it left off, sometimes makes more progress, then stalls again. it wasn’t always like this, feels incredibl” [source](https://www.reddit.com/r/google_antigravity/comments/1wrle6u/working_forever_until_cancelled_but_only_a_couple/pcew96b/)
- Complaint, 2026-09-27, @antigravity (X): “@antigravity you just destroying a good harness day by day your windsuf fork was much better than at current. if you can't do anything better just fork opencode/deepseek/zcode or let your subscribers use those instead” [source](https://twitter.com/151309638/status/2104017857126531200)

### Claude Code

- Praise, 2026-09-27, r/ClaudeCode (Reddit): “the efficiency gains are insane. it's opus 5.5 medium is more efficient than sonnet 5 high. it's honesty become my daily use because the quality is so much better lol.” [source](https://www.reddit.com/r/ClaudeCode/comments/1woe3lo/is_opus_55_really_better_than_fable_in_your/pca63pb/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “it is truly vastly different from opus 5. the cost and response speed are also completely different. that helps me stay better focused.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqzt2x/opus_55_experience_of_an_engineer_at_big_tech/pcb4klf/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “i had a web app floating around 700mg - now sub 100. 5.5 is base level ‘working’ now.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wquop8/i_dont_think_anyone_has_ever_seen_this_before/pcbfafp/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “same feel with yesterday as well. still very good, but i am noticing slips and slightly worse tool calls then before. so it seems to less capable in identifying knowletge gaps also” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqy8rp/am_i_to_understand_the_nerf_has_begun_or_theres/pc9zd6p/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “yea. noticed yesterday too. still very good, but missing gaps, taking more turns on execution and worse tool handling. but yea, stiil fine so far...but we will see” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr67vx/i_agree_with_dario_regulating_llms/pca1vld/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “50% of what? they don't give the actual numbers. they fiddle with the actual token use. the overall value in a year is steadily down down down.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqjb89/i_ran_out_of_codex_on_chatgpt_pro_with_4_days/pca7gi1/)

### Devin

- Praise, 2026-09-27, @cognition (X): “wtf is that pareto??? did swe 2 just fucking break it??? @cognition @devindesktop i can't wait for swe 3!!! your team is crazy!!! <strict_link> <strict_link>” [source](https://twitter.com/1825243355501973504/status/2104017046799372411)
- Praise, 2026-09-27, @cognition (X): “@bnistordev @learnmore_smart @cognition @devindesktop opus 5.5 + swe-2 cheaper and even better” [source](https://twitter.com/17719163/status/2104128141895643174)
- Praise, 2026-09-25, @DevinAI (X): “@notjazii @devinai damn. at this rate opus 6 will be agi” [source](https://twitter.com/1530248240821592064/status/2103555407553933556)
- Complaint, 2026-09-27, @DevinAI (X): “@markfenner @devinai i suspect they have used a quantized version causing the models iq to drop” [source](https://twitter.com/1916897001922506752/status/2104092718213583286)
- Complaint, 2026-09-25, @DevinAI (X): “@notjazii @devinai yeah fr they shouldn't nerf it” [source](https://twitter.com/2012475539324559360/status/2103554876320137216)
- Complaint, 2026-09-25, @DevinAI (X): “@notjazii @devinai don't worry, they won't nerf it down until the release of opus 5.6” [source](https://twitter.com/1388715421864402947/status/2103563089996337471)

### OpenCode

- Praise, 2026-09-27, r/opencode (Reddit): “this model is fast af, and probably better than muse 1.3” [source](https://www.reddit.com/r/opencode/comments/1wqqtgx/longcat25preview_is_now_free_on_opencode_for_two/pcah74r/)
- Praise, 2026-09-27, r/opencode (Reddit): “i honestly don't know where people get the idea that opencode's model is quantized. opencode mostly uses proxy rather than host the model by themself and if you think about it, it might be actually cheaper for them. also, the idea that deepseek from opencode is slower, cannot give same quality of work is not true for me, i've used deepseek v4.1 flash provided from opencode and deepseek official api in deepseek harness and they give same speed (t” [source](https://www.reddit.com/r/opencode/comments/1wqy8pq/how_true_is_it_that_opencode_go_models_are/pcapfjm/)
- Praise, 2026-09-27, r/opencodeCLI (Reddit): “3.1 feels better than 3.0 :)” [source](https://www.reddit.com/r/opencodeCLI/comments/1wrddp1/minimax_is_actively_testing_m31flashpreview/pcchvqd/)
- Complaint, 2026-09-27, r/opencode (Reddit): “so both of us agree that, deepseek v4.1 flash more dumber than api right?” [source](https://www.reddit.com/r/opencode/comments/1wqxyl4/deepseek_v41_performance_is_much_worse_than/pcap2tv/)
- Complaint, 2026-09-27, r/opencodeCLI (Reddit): “opencode go is their $10 subscription plan for use inside the opencode cli. your $10 of payment get you what you would get for $60 at full api pricing, so 6x factor. i would be afraid of quantized models running in stupid mode with it. see other comments asking the same thing. by comparison, i think a chatgpt sub gets you roughly 20x multiplier (your $20 subscription lets you spend $400 of api value), but not sure how that changes in the past wee” [source](https://www.reddit.com/r/opencodeCLI/comments/1wpzekf/operation_cheepseek_phase_2/pcarwqn/)
- Complaint, 2026-09-27, r/opencodeCLI (Reddit): “it was on day 1. thought in caveman and spoke normally. now it's a completely different model that thinks normally and speaks in claudish.” [source](https://www.reddit.com/r/opencodeCLI/comments/1wrddp1/minimax_is_actively_testing_m31flashpreview/pccdbz7/)

### Cursor

- Praise, 2026-09-27, r/cursor (Reddit): “4.6 seemed better in my experience. it was faster and seemed to take less shortcuts.” [source](https://www.reddit.com/r/cursor/comments/1wrktgc/am_i_the_only_one_who_thinks_grok_47_is_actually/pcfhkvq/)
- Praise, 2026-09-27, @cursor_ai (X): “@stillnesshum @cursor_ai opus 5.5 is incredibly good. 🩵✨ it is creative and a workhorse orchestrator.” [source](https://twitter.com/1982798442272616448/status/2104213006209216652)
- Praise, 2026-09-26, r/cursor (Reddit): “opus 5.5 has been so good for me. fast and way less verbose than 5. it’s my new favorite and i love how it spins out subagents.” [source](https://www.reddit.com/r/cursor/comments/1wq3m71/im_out/pc4847j/)
- Complaint, 2026-09-27, r/cursor (Reddit): “if it came out two years ago, it would be amazing. competing with modern models, it’s garbage.” [source](https://www.reddit.com/r/cursor/comments/1wrktgc/am_i_the_only_one_who_thinks_grok_47_is_actually/pcdvs3p/)
- Complaint, 2026-09-27, r/cursor (Reddit): “if you want the cursor models usage to last, stick with composer and grok. don't use fast mode and you'll last the whole month on a $60 plan. they changed the auto to pick api models recently. maybe people jumped ship to other ides and they needed to up the api usage to keep the partnership going since cursor was acquired. who knows, either way it was the death of auto. i myself am looking to alternatives. i love the speed and ux of cursor but ma” [source](https://www.reddit.com/r/cursor/comments/1wrddz6/why_is_the_other_models_usage_is_being_used_so/pceihvh/)
- Complaint, 2026-09-27, r/cursor (Reddit): “i don’t think 4.7 is bad at all, but for me 4.6 was noticeably faster and better overall. 4.7 can absolutely handle complex tasks, but the reasoning time feels way heavier. with 4.6, it was so fast i could barely scratch myself off the chair before it was already done. that’s probably the biggest regression for me. i’d rather have 4.6’s speed and consistency back than wait longer for 4.7 to maybe give me a slightly better answer.” [source](https://www.reddit.com/r/cursor/comments/1wrktgc/am_i_the_only_one_who_thinks_grok_47_is_actually/pcenf29/)

### OpenAI Codex

- Praise, 2026-09-27, r/codex (Reddit): “no point in using 5.6 terra anymore, 6 sol is more or less a drop in represent.” [source](https://www.reddit.com/r/codex/comments/1wr4e20/gpt_6_sol_is_the_new_opus_47/pca4xum/)
- Praise, 2026-09-27, r/codex (Reddit): “agree. it's too good. they can't let it last unless it really is just that efficient it could be the first model they aren't forced to nerf.” [source](https://www.reddit.com/r/codex/comments/1wp3bzl/you_need_to_try_opus_55/pcabhky/)
- Praise, 2026-09-27, r/codex (Reddit): “i hope they wont nerf astra, this model is so damn good, we need the same astra but cheaper :x let me dream guys!” [source](https://www.reddit.com/r/codex/comments/1wr1oir/they_are_aware_and_working_on_it_apparently_just/pcauu3q/)
- Complaint, 2026-09-27, r/codex (Reddit): “it's more like a sonnet with thinking off” [source](https://www.reddit.com/r/codex/comments/1wr4e20/gpt_6_sol_is_the_new_opus_47/pc9vdq0/)
- Complaint, 2026-09-27, r/codex (Reddit): “last week-2 weeks have been not good for gpt. very good for claude, compounding effects.” [source](https://www.reddit.com/r/codex/comments/1wqtd1g/openai_gpu_are_really_cooling_down_theo_just/pc9w5q0/)
- Complaint, 2026-09-27, r/codex (Reddit): “they really need to reset the model stack. i mean i'm sure that each generation between 5.5 and 6 has gotten better at something. i'm not exactly sure what because it basically is unusable for serious coding. literally lost in a c++ code base. mangles everything it touches. takes 15 minutes on a short run. wildly expand scope. invents in ludicrous defensive checks against impossible situations. continually routes c++ code/data to javascript ui” [source](https://www.reddit.com/r/codex/comments/1wr1oir/they_are_aware_and_working_on_it_apparently_just/pc9zpqi/)

### GitHub Copilot

- Praise, 2026-09-24, r/GithubCopilot (Reddit): “the real question is why you're still using opus 4.7. especially if usage is a concern, now would be a good time to switch to opus 5.5 as it is better, cheaper, and more efficient with token use.” [source](https://www.reddit.com/r/GithubCopilot/comments/1woeusl/github_copilot_opus_47_usage_and_cost_details/pbowiz6/)
- Praise, 2026-09-24, r/codex (Reddit): “if you have a 12 months+ agreement and you can show they're nerfing usage in an non-subjective way, get a refund or threaten them. if you're on a 1 month plan, shouldn't have any attachment to an ai provider. there's so many now. let the month expire and move to the next one. these products are not sticky and soon enough there will be very little value in the model infrastructure sans the ecosystem around it. with that said, both openai and anthr” [source](https://www.reddit.com/r/codex/comments/1wpfoxh/the_downhill_begins/pbv3vft/)
- Praise, 2026-09-19, r/GithubCopilot (Reddit): “models are more intelligent now. all we need is to know when to use which model and yes, the prompt is everything.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wkdn9s/copilot_is_allowing_27_in_overage_despite_my/paqi9pu/)
- Complaint, 2026-09-26, r/GithubCopilot (Reddit): “claude code and codex are where i’ve landed for most real work. copilot feels less essential than it did a year ago. the bigger improvement for larger codebases wasn’t switching models though. it was giving the agent a way to look up our own systems instead of trying to infer everything from the repo. we use port.io for that over mcp; backstage can play a similar role. model quality is getting pretty close. the bigger difference now is how much o” [source](https://www.reddit.com/r/GithubCopilot/comments/1u95cce/which_ai_coding_assistant_are_developers_actually/pc3pg02/)
- Complaint, 2026-09-26, r/GithubCopilot (Reddit): “5.6 max is better by far at tool use, 6 max thinks way too much on how to perform a skill differently than the instructions state, struggles mightily and eventually fails. it seems about the same with writing code, but i need my agents to be able to read figma designs and test uis with playwright, so poor tool use is a show stopper, even if 6 is half the price.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wptkkq/how_does_gpt6_sol_feel_so_far/pc3x681/)
- Complaint, 2026-09-26, r/GithubCopilot (Reddit): “gpt 6 luna is worst than 5.6 for coding, becarefull to adjust your model when necesarry <strict_link>” [source](https://www.reddit.com/r/GithubCopilot/comments/1wq39rv/did_they_10x_the_cost_of_luna_56_since_the/pc4rmf2/)

### Cline

- Praise, 2026-09-26, @cline (X): “@cline 41 on the intelligence index at that speed and price is an insane combo” [source](https://twitter.com/1889631970667405317/status/2103747538025325008)
- Praise, 2026-09-24, @cline (X): “cline has improved a lot in a mean time, the cache hit rate is absolutely insane now great work, guys @cline <strict_link>” [source](https://twitter.com/2017132628361822208/status/2103104236540375166)
- Praise, 2026-09-23, @cline (X): “@cline opus cheaper and still beating fable?? sol luna half price sticking for you tho” [source](https://twitter.com/1434380824678309889/status/2102778896332582977)
- Complaint, 2026-09-27, @cline (X): “@cline totally worthless model!” [source](https://twitter.com/82187574/status/2104250853783662640)
- Complaint, 2026-09-26, @cline (X): “@cline this feels more like a deepseek/glm model than a gemini model... 👀 <strict_link>” [source](https://twitter.com/1803325156955480064/status/2103666923435331858)
- Complaint, 2026-09-25, @cline (X): “@dynamicwebpaige @cline trash model” [source](https://twitter.com/944978898927964161/status/2103382184887194021)

### Factory

- Praise, 2026-09-24, @FactoryAI (X): “@factoryai desktop app is constantly improving at an insane rate, it’s a treat to watch 🥹 <strict_link>” [source](https://twitter.com/1721143727043887104/status/2102998403496268015)
- Praise, 2026-09-21, @droid (X): “@droid it is a massive step up for the model, especially with those self-verification capabilities. we actually went deeper on this here: <strict_link>” [source](https://twitter.com/1213502906332110848/status/2102172305392635915)
- Praise, 2026-09-18, @FactoryAI (X): “i'd use it to build out an eval suite for the mcp layer of my ios app. have been a droid user for over a year but i mainly use it as my backup and code-review agent made this open source skill for making the droid feedback loop even tighter and more automated: <strict_link> with the max plan i'd basically use the hell out of it and really test the models. in return i'd write up a blog post since i find factory's research articles like: <strict_li” [source](https://twitter.com/377179355/status/2100786695604146485)
- Complaint, 2026-09-23, @FactoryAI (X): “@sqs @harrystuck77 @factoryai @droid @badlogicgames @ampcode @cursor_ai @opencode @anthropicai @claudeai @cognition @xai @grok @build amp is definitely better, i used to love droid, but it went downhill majorly when they started focusing on enterprise customers” [source](https://twitter.com/1475837160590880773/status/2102574979594535291)
- Complaint, 2026-09-14, @droid (X): “it was good but its really fallen off, they are insanely lazy too cant even keep their release notes current for weeks at a time for a supposedly autonomous software factory. like hello why isnt this autonomously updated by an agent? i only use it occasionally now with my glm api key but i used to be a $200/month subscriber” [source](https://twitter.com/1445863287804022785/status/2099392033324417363)
- Complaint, 2026-09-08, @FactoryAI (X): “@droid @factoryai i'm sad, as i've been using it daily for almost 6 months. since grok build was giving me the results i needed, i dropped it. am i missing something. i'd love to have droid again, giving me more than just a model cli.” [source](https://twitter.com/1738636938616115200/status/2097389013287968781)

### Pi

- Complaint, 2026-09-26, @pidotdev (X): “@shantanugoel @pidotdev its not a good model, just use luna6” [source](https://twitter.com/1448626313619705856/status/2103773838320550066)
- Complaint, 2026-09-20, r/PiCodingAgent (Reddit): “yeah i've been live-patching features i wanted into cc for some time now before finally planned to roll my own on pi to get easier access to more models (been hating sonnet lately). then i found omp and realised it has almost everything i wanted and more (idle timeout compaction is best quota saver ever). being able to turn on/off the tools/extensions makes it perfectly flexible for me. while i'm still building some new extensions, the fact i ca” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wlpj30/lord_forgive_me_for_the_time_i_wasted/pb1poql/)
- Complaint, 2026-09-18, @pidotdev (X): “yeah it's basically a persistent deno notebook kernel/shell (close to these rlm things, but i wanted to avoid python). generally still runs the codemode v8 isolate, but it allows for declaring constants etc. adds a notebook management tool and requires a couple of extra calls for new session for the agent to get its bearings arount its current notebook state, maybe that's where the increased token cost came from, but upsides in long, multi contex” [source](https://twitter.com/183580487/status/2100907414253892010)

### Amp

- Praise, 2026-09-24, @AmpCode (X): “@leuler2718 @ampcode @synthetic_new i haven't noticed any degraded performance with gpt models in amp, no. i think that was a characteristic of older codex models.” [source](https://twitter.com/1637683046395592705/status/2103058329194856588)
- Praise, 2026-09-23, @AmpCode (X): “sry to hijack this tweet 😓 but... - cleanest/most seamless cloud agent impl. skills/mcps/integrations/oidcs syncing seamlessly to both local and orbs is amazing) (i've been using a ton of cursor cloud agents too but it was rough around the edges and i hate using it). this is for sure amp's best feature imo. - being able to use a terminal or open something in desktop or spawn a local web app in portals for testing is really good (there are instanc” [source](https://twitter.com/129354616/status/2102767352886514024)
- Praise, 2026-09-20, @AmpCode (X): “@sqs @ampcode @vercel getting better and better!” [source](https://twitter.com/2881611/status/2101641625352949889)
- Complaint, 2026-09-14, @AmpCode (X): “@sqs @andreaslbigger @ampcode next step:train amp's own model or just buy openai, so that we can have a model that works stable.” [source](https://twitter.com/99872071/status/2099492722528911364)
- Complaint, 2026-09-05, @AmpCode (X): “@ampcode, a ui/ux (macos app) task wuas given both to gpt 5.6 sol - medium and gpt 6 astra and they came up with almost similar output, not much difference is quality of output - both consumed similar amount of tokens (13m,14m) and similar # of requests (167,170) - i was expecting astra to do better” [source](https://twitter.com/38651218/status/2096353560682443178)

### Warp

- Complaint, 2026-09-25, @warpdotdev (X): “@mitchellh really cool feature. i loved it in @warpdotdev , but then it become dumber on this aspect” [source](https://twitter.com/1049728717/status/2103597816169762855)
- Complaint, 2026-09-09, @warpdotdev (X): “i miss when @warpdotdev was incredible was excited about the opensourcing and the concept of 0z and everything but its diabolically bad, slow and laggy when it used to be truely blazingly fast might have to fork and rip out all the bs or just drop it” [source](https://twitter.com/2970558232/status/2097834207552618664)

### Kiro

- Complaint, 2026-09-25, r/codex (Reddit): “that's all a very good point. so i was using claude all last year, coming from kiro which never seemed to evolve. after some bugs that completely depleted my credits several times, i decided to move to codex after gpt 5 was released and it was a breath of fresh air. the last week and a half or so i find my normal usage depletes significantly faster. this post was more a light hearted take but holy cow so many people got so offended? when moving t” [source](https://www.reddit.com/r/codex/comments/1wq7zh6/so_now_that_were_jumping_ship_to_claude_whats_the/pc2m88l/)

### Grok Build

- Praise, 2026-09-24, r/opencodeCLI (Reddit): “<strict_link> from my experience, muse spark 1.3 at xhigh in opencode give me wrong answers all the time. it might flare better with muse code as the model is trained and refined around the harness, the same way grok inside opencode feels dumber compared to when it's inside grok build.” [source](https://www.reddit.com/r/opencodeCLI/comments/1wopbtv/what_muse_spark_14_contributor_is_already_here/pbswa6q/)

### Augment Code

- Complaint, 2026-09-12, @augmentcode (X): “@simplygandan @augmentcode also, sonnet used to feel good enough. now, even opus feels dumb! maybe, augment code was that good or we are spoilt by fable and astra” [source](https://twitter.com/2959524282/status/2098864871924478317)
