# Choosing models (`models`)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/criterion/models

Area of 4 criteria. Which models do you get, and do they hold up?

Criteria: [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md), [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md), [Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md), [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md)

Rated author-weeks, all agents: 9163. Complaint share: 74%.

## The brief

Written by Claude Opus 5.5 from 110 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**Users want model choice; they get gated access and suspected nerfs.**

TL;DR:

- Biggest complaint is drift, where models feel sharp at launch and then dumber weeks later.
- Newest models arrive late, or only on pricier plans, and users notice within hours.
- Open, multi-provider harnesses earn praise for letting users mix and swap models freely.

In plain terms: You pick a model, it works great, then it seems to slip. The newest model might not be on your plan. Effort defaults can burn your quota. Tools that let you swap providers feel safest.

### How it breaks

- **Models feel nerfed after launch** ([Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md)). Users describe a launch-week high followed by lazier, sloppier output, and many suspect throttling timed to promotions or the next release.
  The pattern repeats across vendors. A model ships, users love it, then posts report more mistakes on unchanged workflows. Some tie the drop to load and usage resets. Others tie it to an ending promotion or an incoming successor. Whether or not anything changed server-side, the perception erodes trust. Stopping nerfs is the single most common request in this area, and most of it comes from Claude Code users.
  Evidence:
  - Complaint, Cursor, r/cursor, 2026-09-16: “same, this seems to be the way of things now sadly. bait and switch tactics are the norm anymore with these frontier models… :/ great for a day or two on drop but then the intelligence doesn’t just get a little worse; its like working with einstein and then all of a sudden i’m giving money to a 15 year old lying obstinate teen with the most severe case dunning-krueger syndrome i have ever seen…” [source](https://www.reddit.com/r/cursor/comments/1wh5qqi/grok_46_has_been_lobotomized/pa3cbu2/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-12: “stop resets! the system is already overloaded, hence the lobotomy of the models. if i use ur ai it has to be good!!” [source](https://www.reddit.com/r/codex/comments/1we8gmi/tibos_bonus_resets_with_pushing_back_next_reset/p9dfray/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-07: “i think the 3.8 is hallucinating and have difficulty splitting what i am telling it , i think 3.7 was better than 3.8 in some few ways” [source](https://www.reddit.com/r/google_antigravity/comments/1w6h936/google_claims_are_correct_guys/p8dpefo/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-23: “in some ways similar but in several benchmarks it is straight up worse.” [source](https://www.reddit.com/r/codex/comments/1wo3ixp/how_do_you_think_openai_will_respond_next_week/pbksu4c/)

- **Newest models stuck behind plans** ([Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md)). Users report seeing a new model announced and then waiting for it on their plan, region, org or harness.
  Access lags in several ways. Plus users can't see the new model. Teams on paid tiers find nobody in the account can use it. Third-party harnesses ship support days after the launch. Users compare vendors on how fast a model lands and switch tools when a rival enables it within minutes. Requests for newest models on lower-priced plans and for parity across app and CLI cluster heavily on OpenAI Codex.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-06: “i am a plus user but still not able to see astra on my codex 🥲😭😭” [source](https://www.reddit.com/r/codex/comments/1w7ylpn/why_is_everyone_so_angry_with_plus_users/p84e84v/)
  - Complaint, Cline, @cline, 2026-09-10: “@cline kindly please be a little bit faster to enable deepseek 4.1 flash in clinepass. opencode and commandcode enabled this under 30 mins of official announcement.” [source](https://twitter.com/10877782/status/2097944041832734973)
  - Complaint, Factory, @FactoryAI, 2026-09-04: “@factoryai i was using plus and max, but no one in my account can use gpt-6 astra 😂😅” [source](https://twitter.com/2079628486055178240/status/2095732701659852808)
  - Complaint, Pi, @pidotdev, 2026-09-04: “@pidotdev where is astra support? i finally get access and i can't even use it in pi :(((” [source](https://twitter.com/2088433818047004672/status/2095954568513487157)

- **Effort defaults quietly drain usage** ([Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md)). High or max reasoning set as a default, or inherited by subagents, burns credits and often overthinks simple work.
  Users find out late that a UI update set a model to high, or that spawned agents defaulted to the priciest model. High effort on simple tasks reportedly hunts for edge cases that aren't there and overengineers. A Pi bug silently dropped most thinking levels. The community advice is consistent. Medium is the sweet spot, and you should raise effort deliberately.
  Evidence:
  - Complaint, GitHub Copilot, r/GithubCopilot, 2026-09-06: “i discovered that copilot ui did an update setting sonnet at high by default. of course i was consuming so much credit will try to do with luna on medium high and see if i will be able to manage” [source](https://www.reddit.com/r/GithubCopilot/comments/1w6zbe6/copilot_pro_am_i_missing_something/p82sb8j/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-04: “what effort you used. with high effort for simple task looks for hidden catches and missed edge cases that are expected but are not there and result is oftrn overnegineered si olution.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w72g6u/fable_51_is_an_overthinker/p7stqja/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-08: “mine was doing this too until i made sure the agents where running cheaper models. claude was deafulting every agent it created to fable 5.1 after i made sure it was using models more efficiently my usage slowed down significantly.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wa6325/what_is_happening_absolutely_frustrating/p8gjb41/)
  - Complaint, Pi, r/PiCodingAgent, 2026-09-22: “if asking pi failed, that is because pi itself introduced a bug in 0.86.x upwards that silently drops thinking levels other than off or medium.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wmo5e5/how_do_i_support_thinking_mode_correctly_with/pb9r60u/)

- **Auto routing swaps models unannounced** ([Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md)). Auto modes save money for some, but users report subtasks rerouted to models they never picked, draining other quotas.
  Cursor auto gets split reviews. Some users run hundreds of millions of tokens on it happily. Others say it defaults to expensive models and lowers quality. One user says project mode switched subtasks to a different provider mid-task. Skeptics question how any router can judge task difficulty upfront, and point to cache invalidation when it switches mid-session. Users want smart switching, but they want it to be visible.
  Evidence:
  - Complaint, Cursor, @cursor_ai, 2026-09-16: “wtf, it's too坑了, be cautious when using cursor's project mode. although it can help you run projects on cloud machines without using local resources, it will secretly change the model calls by itself. i clearly specified to use grok when creating it, but during the task, the subtasks secretly used gpt, causing me to exhaust all my other models quota... @cursor_ai <strict_link>” [source](https://twitter.com/1253580691/status/2100106760082411580)
  - Complaint, Cursor, r/cursor, 2026-09-15: “dont use auto. simple. keep grok for planning, and use composer for implementation. i got a $20 subscription and work across 7 repo projects with ease, only once in the last year have i run out of credits. if you put a little thought into your model use and coding style you can get much more from it. i keep testing auto for quality. it sucks. why? cos it keep defaulting to grok for coding. and then it does high and fast, which is astronomically high compared to composer 2.5 (without fast). grok for thinking and planning. composer for checking grok's thoughts and implementation details. and use beads to string together a plan composer will implement. i keep mentioning beads, and it feels like throwing pearls before swine.” [source](https://www.reddit.com/r/cursor/comments/1wcqt8e/changed_alot/p9y3bs4/)
  - Complaint, Factory, @FactoryAI, 2026-09-26: “the transit station really has no work to do. i have always been curious about the router that @factoryai has been promoting. they have never clearly explained the specific router mechanism, just vaguely mentioning the classification of simple and complex tasks. how can we determine whether a task is complex or not before it starts, and the inevitable cache invalidation caused by switching models midway is a big problem (anyway, the transit station makes money regardless, lol). a previous idea was to train a small ml model for task classification, and then jev came out, but is this model really suitable for the task difficulty classification work? a big question mark needs to be placed on that.” [source](https://twitter.com/1724973133097291776/status/2103781371772809499)
  - Praise, Cursor, r/cursor, 2026-09-03: “there is a usage dashboard on [cursor.com](<strict_link>) where you can see exactly how many tokens you spent per request and what model was used. i am looking at it now and i see mostly auto and gemini flash 2.5. i think auto is on of cursor's own models like grok or composer i think and it seemingly uses gemini flash 2.5 for smaller implementation tasks. i just passed 800m tokens for the month and i cannot complain about the results. of course you still need to set up the proper skills and also guide the agents, but auto works very good for me.” [source](https://www.reddit.com/r/cursor/comments/1w46p57/do_you_guys_use_auto_in_cursor_do_you_recommend/p7jfb1f/)

- **Mixing models is the real upside** ([Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md), [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md)). Harnesses that let one model plan and another implement draw the warmest posts, because users can swap without rebuilding their setup.
  Users run a thoughtful model as coordinator and cheaper ones as workers, often from different labs in the same session. They say this beats a single big or small model for the whole task. Swappable models also let teams compare cost and quality on an identical flow. Multi-provider choice in one harness is a standing request against single-vendor tools.
  Evidence:
  - Praise, Pi, @pidotdev, 2026-09-20: “@pidotdev when i realized you could just have models invoke other models from other labs in the cli. was an early one but a game changer for me” [source](https://twitter.com/2069281373932650496/status/2101780014639251795)
  - Praise, GitHub Copilot, r/GithubCopilot, 2026-09-01: “our [evaluations](<strict_link>) have found that the copilot harness is a bit more token efficient than cc, at the time of writing. there is a large degree of nondeterminism and variance per task. something that copilot gets you though is nice multi-model support. personally i use opus 5 as my coordinator agent and have it send work to gpt-5.6 terra or luna agents as implementors. it gets much better and faster results than using a single big or small model for your whole task. i like opus' thoughtfulness for downstream effects of changes, and terra/luna seems to be a good mix of being very good at carrying out tasks without getting bogged down or getting too off track.” [source](https://www.reddit.com/r/GithubCopilot/comments/1w4ljfs/claude_code_vs_github_copilot_token_burn/p79xxgo/)
  - Praise, Google Antigravity, @antigravity, 2026-09-18: “that the model can be changed without redoing the agent is an architectural decision, not just a matter of convenience: it allows for comparing quality, cost, and latency with the same flow. for fp&a and internal control, that comparability is key to justifying automation with evidence and not just with a demo.” [source](https://twitter.com/1569177389959192578/status/2100955697332580472)
  - Praise, GitHub Copilot, r/GithubCopilot, 2026-09-06: “i’m using ghcp credits via hermes agent and it’s been pretty fucking awesome. almost completely out of claude desktop after about a week of tooling it. it’s wonderful being able to easily swap models, do so mid session, spawn bots/agents with their own model settings. can control my machine as much as i let it. full disclosure, i put the backend on my linux box and connect desktop over lan, so it’s not actively controlling my windows desktop - but it can, and does awesome with ubuntu 26.04.” [source](https://www.reddit.com/r/GithubCopilot/comments/1w9a0ft/computer_use_using_github_copilot_credits_where/p88xgx3/)

### Who stands out

- **Devin (stronger)**. Users praise a cheap plan that bundles frontier models from several labs, its own model, and visible effort selection.
  Posts highlight switching between rival frontier models and an in-house model on one low-priced plan. Day-one support for new launches also draws praise. Fusion mode pairs a frontier model with the in-house one, and users like effort selection. Complaints are thin and center on benchmark skepticism.
  Evidence:
  - Praise, Devin, @cognition, 2026-09-22: “@cognition devin 20$ is really the best deal right now. i can swith between astra, fable, gemini 3.8 flash... (rarely use but yes you can use any model) and unlimited swe-2. moreover, codewiki and free cloud agents recently. crazy and amazing at the same time. thanks @dabit3” [source](https://twitter.com/312505543/status/2102263172736700893)
  - Praise, Devin, @cognition, 2026-09-04: “@dabit3 @devinai @cognition @devindesktop love that model effort selection is just a thing of beauty” [source](https://twitter.com/1917637884967792640/status/2095987531565183042)
  - Complaint, Devin, @cognition, 2026-09-16: “@thek420metric @elitzavasileva @cognition yeah, swe-2 is actually a good model. doesn't have computer use like the openai models in codex, but pure coding/debugging/etc it's good. i've been using the fusion mode with astra and swe-2.” [source](https://twitter.com/1656279377213263873/status/2100182900532789674)
  - Praise, Devin, @DevinAI, 2026-09-06: “i've been slacking all week! sick as a dog 😪 tonight i'm evaluating @openai gpt 6 astra performance in @devinai. i haven't been this excited to really go balls to the wall with tokens in a minute. day 1 support. a shout out from openai. that's what i'm talking about 🚀🥳🎉 <strict_link>” [source](https://twitter.com/358932148/status/2096398609155490172)

- **OpenAI Codex (weaker)**. The loudest complaint volume here mixes gated access, parity gaps between app and CLI, and reports of the flagship turning lazy.
  Codex users lead requests for app and CLI parity, restored models and newest models on cheaper plans. Posts describe a new flagship going from loved to sluggish within weeks. Power users still praise effort tuning and report that medium as orchestrator stretches usage. Hitting zero usage pushes some back toward rivals.
  Evidence:
  - Complaint, OpenAI Codex, X search: OpenAI Codex, Codex CLI, Codex app, 2026-09-20: “gpt-6 astra is now frustratingly unusable. the codex app is slow, outputs have been very lazy, and the model is just behaving very dumb. the openai honeymoon is definitely over.” [source](https://twitter.com/1861818861785587714/status/2101787555377135765)
  - Complaint, OpenAI Codex, r/codex, 2026-09-07: “i’ve been at 0% for the past two days and i loved my usage with astra, obviously too much. knowing what i know now about the low/medium model competency, i will no longer be using high astra. annoyingly now that i finished up the last part of my project i can’t ship it until i get access restored to help me realistically audit the entire codebase and work on final gate items :’) i still have 5 days until my usage resets •\_• this is usually when i run back to my claude plan and upgrade it back to 20x…which i’m trying to use all self-control right now to not do. i really want to say i shipped with codex.” [source](https://www.reddit.com/r/codex/comments/1w9ya9z/possible_reset_timing_update/p8e0gmk/)
  - Praise, OpenAI Codex, r/codex, 2026-09-16: “why are you on astra high? just use medium to orchestrate and rely on sol after. i had it run for dozens of hours and i still have usage, even when relying heavily on astra medium as the worker.” [source](https://www.reddit.com/r/codex/comments/1whvec6/any_hope_for_reset/pa7ymos/)
  - Praise, OpenAI Codex, r/codex, 2026-09-04: “now do yourself a favor, don't go above medium reasoning. look at the benchmarks. medium is peak.” [source](https://www.reddit.com/r/codex/comments/1w7d3xe/astra_dropped_for_me/p7u3mr9/)

- **Claude Code (mixed)**. Users rate the effort tiers on its newest model highly, yet Claude Code generates most of the stop-nerfing requests.
  Praise centers on clear effort tiers. Users call medium efficient and say xhigh self-checks and documents. They also report fewer limit hits on the newer model. The other side is drift. Users report the top model getting worse after a usage promotion ended. They also fault the lack of a competent lightweight model.
  Evidence:
  - Praise, Claude Code, r/ClaudeCode, 2026-09-25: “opus 5 was great but hit limits fast. **opus 5.5** on high reasoning is next level — better output, way fewer limit issues. anyone else checking their usage? what's your count?” [source](https://www.reddit.com/r/ClaudeCode/comments/1wpsg9b/one_month_in_65m_tokens_later_is_this_normal/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-15: “can live with the usage and work around that a bit. but i find fable 5.1 suddenly got nerfed down after 13 sept once the extra 50% usage promotion was done. same thing as a while back when they gave extra usage. fable suddenly makes a lot of stupid mistakes where i honestly had to look if it didnt use sonnet or opus in some mysterious way. quite annoying if you run a business with google ads automated on 85 clients and it starts making mistakes while it ran for 2 months flawlessly. no runbook or workflow was changed, no other methods, just the exact same sequences... when this happens magically a new model will come out that suddenly feels better than fable, because they dumbed it down...” [source](https://www.reddit.com/r/ClaudeCode/comments/1wh6524/cancelled_claude_max_20x_after_unusually_fast/p9zxv2a/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-27: “if you have limited tokens, medium will get you there. if you can spare the tokens and the tasking important, high gives a lot of extra quality control. if you develop something mission critical and have lots of tokens to spare, xhigh does a lot of the extra tests and documentation work in addition to the task itself. it also challenges its own work. opus 5.5 is a pleasure to work with.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrmw8c/high_or_medium_for_opus_55/pcetct4/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-01: “anthropic seems determined to not have a competent lightweight model, which in a time where user's have ample access to models like luna seems like a very dumb move.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w44s2m/haiku_45_nearly_year_old/p7927kj/)

- **OpenCode (stronger)**. Users treat it as the model-agnostic harness, enabling new open models fast and letting setups survive model swaps.
  Users mix primary and subagent models across labs while keeping rules and process intact. They praise fast rollout of new open-weight models. The cost of that flexibility is opacity. Pooled routes can quietly move to a weaker model as provider deals change, and cheap tiers demand careful model picking.
  Evidence:
  - Praise, OpenCode, @opencode, 2026-09-16: “perhaps an underappreciated advantage of open coding harnesses: being able to combine the best of multiple different frontier models cheering for @thdxr @opencode @teknium @mitsuhiko + all the others! <strict_link>” [source](https://twitter.com/2867689615/status/2100299185958641981)
  - Praise, OpenCode, r/ClaudeCode, 2026-09-06: “i have been using opencode for a year. my setup stays the same, the underlying models for primary and sub agents can change. i mix and match, it is a beautiful thing to see `coder-luna` and `reviewer-opus` work together, seamlessly, within the same session. now my `primary-sol` will be changing into a `primary-astra`. i've got model specific rules that allow me to 'normalize' their behavior when needed. claude code is cool and all, but to me it can't beat being able to replace the model and keep my setup and process the same (eg: k3 zdr `architect-k3` in place of fable without skipping a beat). just wish anthropic would support oauth access to open code, like openai does. don't stunt oss (consumption is controlled at the provider's end, so it is not as if oc can use more than an equivalent cc setup can, especially if anthropic were to provide a connector, again, like openai does)” [source](https://www.reddit.com/r/ClaudeCode/comments/1w8qphl/model_agnostic_harness_setup/p85szzv/)
  - Complaint, OpenCode, r/opencode, 2026-09-09: “it was glm at some point. "big pickle" routes to whatever model opencode has a good deal with the provider and that has enough idle capacity at the moment. as their deals come and go, and computing is not a constant factor in this market, "big pickle" may start routing to a worse model out of the blue, so enjoy while you can and watch out for degradations.” [source](https://www.reddit.com/r/opencode/comments/1wbe5dz/is_it_my_impression_or_only_big_pickle_is/p8pdvue/)
  - Praise, OpenCode, @opencode, 2026-09-13: “@opencode i’ve been using deepseek 4.1 through opencode for over 5 hours now. looking at the benchmarks, it seems to be around the same level as gpt-5.6 max (sol high). what surprised me the most is that it can also control your pc, click things, and take screenshots to better understand what’s going on. it’s honestly one of the best models you can use and really push to its limits, instead of being stuck using openai’s luna model. the usage limits are very generous, and the model is really good.” [source](https://twitter.com/1154188487291297792/status/2099273508454826208)

- **Kiro (weaker)**. Posts are dominated by missing or stale models, with users asking for open-weight and frontier releases that rivals already carry.
  Users say niche providers ship popular open-weight models on day one while Kiro has none. They ask for old models to be pruned. When a requested frontier model finally arrives, praise follows, but regional and org rollout lags. Credit burn makes the absence of cheap models sting more.
  Evidence:
  - Complaint, Kiro, r/kiroIDE, 2026-09-10: “pfft. every month my credit on kiro runs out sooner. i used my kiro plan in under 2 days this month. they need a deepseek model, not astra.” [source](https://www.reddit.com/r/kiroIDE/comments/1war1iu/is_gpt_astra_and_the_fable_models_coming_to_kiro/p8xg591/)
  - Complaint, Kiro, r/kiroIDE, 2026-09-02: “let's be honest, its not about data sharing. kiro is lacking even the latest open weight models like deepseek v4 or glm-5.3. how is it that a niche provider like melious.ai has these models basically on day 1 and we still have nothing? instead of developing more and more broken harnesses, they should fokus on keeping up with the latest development.” [source](https://www.reddit.com/r/kiroIDE/comments/1w4wl5c/are_we_getting_fable_51/p7bhm0z/)
  - Complaint, Kiro, @kirodotdev, 2026-09-06: “@kirodotdev when astra and fable and can you please remove all those old models you guy can do better than this @kirodotdev” [source](https://twitter.com/1328696201969946627/status/2096509522084855949)
  - Complaint, Kiro, @kirodotdev, 2026-09-22: “@kirodotdev it hasn't appeared for my organization in us-east-1 yet.” [source](https://twitter.com/2061100080120344576/status/2102282631169839341)

### Fine print

- Model names and benchmark claims are as users state them; nothing here verifies drift or routing behavior server-side.
- OpenAI Codex posts outnumber most rivals combined, so cross-cutting themes skew toward its user base.
- Several routing praise posts come from vendor or vendor-adjacent accounts, which inflates the positive signal on auto routing.

## Top requests

What users ask to add or change, most asked first. 2076 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Criterion | Author-weeks | Posts | Agents (author-weeks) |
|---|---|---|---|---|---|
| 1 | Stop nerfing or degrading models over time | [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md) | 88 | 92 | Claude Code 59, OpenAI Codex 22, Cursor 3, OpenCode 3, Devin 1 |
| 2 | Newest models on lower-priced plans | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 79 | 80 | OpenAI Codex 32, Claude Code 23, OpenCode 9, Cline 6, Google Antigravity 2, Cursor 2, Kiro 2, Amp 1, GitHub Copilot 1, Devin 1 |
| 3 | Add Opus 5.5 model | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 64 | 74 | Kiro 29, Google Antigravity 19, Amp 4, Claude Code 4, OpenAI Codex 3, Cursor 2, Devin 2, Conductor 1 |
| 4 | Add DeepSeek V4.1 Flash model | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 60 | 62 | OpenCode 21, Cursor 9, Factory 7, Zed 6, OpenAI Codex 5, Kiro 4, Cline 3, GitHub Copilot 2, Devin 2, Amp 1 |
| 5 | Restore removed models | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 58 | 60 | OpenAI Codex 26, Cursor 10, OpenCode 8, Claude Code 7, Google Antigravity 3, Amp 2, Cline 1, GitHub Copilot 1 |
| 6 | Multi-provider model choice in one harness | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 50 | 52 | OpenAI Codex 10, Cursor 8, Zed 8, Google Antigravity 5, OpenCode 5, Pi 5, Claude Code 3, Amp 2, Devin 2, Cline 1, Factory 1 |
| 7 | Add GPT-6 Astra model | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 48 | 49 | Cursor 26, OpenAI Codex 7, Kiro 7, OpenCode 5, Amp 1, Cline 1, Devin 1 |
| 8 | Model parity across app, CLI and platforms | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 48 | 49 | OpenAI Codex 36, OpenCode 4, Google Antigravity 3, Cursor 2, Amp 1, Conductor 1, GitHub Copilot 1 |
| 9 | Cheaper capable lightweight model tier | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 47 | 49 | OpenAI Codex 21, Claude Code 8, Cursor 7, OpenCode 5, Google Antigravity 3, GitHub Copilot 1, Devin 1, Factory 1 |
| 10 | Update outdated Claude models in catalog | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 42 | 43 | Google Antigravity 36, Claude Code 3, OpenAI Codex 2, Kiro 1 |
| 11 | Fix currently degraded model quality | [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md) | 41 | 43 | Claude Code 24, OpenAI Codex 12, Cursor 2, OpenCode 2, Google Antigravity 1 |
| 12 | Keep older models available after new releases | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 40 | 44 | OpenAI Codex 14, OpenCode 8, Claude Code 6, Cursor 6, Google Antigravity 4, GitHub Copilot 1, Zed 1 |

### 1. Stop nerfing or degrading models over time

- Claude Code, 2026-09-27, r/ClaudeCode (Reddit): “please, dario, don't 'optimize' opus 5.5! seriously, why cannot we just have nice things??” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrpq38/any_idea_what_anthropic_figured_out/pcfk19s/)
- Claude Code, 2026-09-27, @ClaudeDevs (X): “opus 5.5 has been so fucking great, fast, much less verbose, objective, very efficient with long shot tasks and loops, as orchestrator and less token burning. @claudeai @claudedevs give us a huge huge favor: don't dare to nerf it.” [source](https://twitter.com/2057822650752200704/status/2104022567157731336)
- Claude Code, 2026-09-26, r/ClaudeCode (Reddit): “i pray for this to not be subsized to death and won't be nerfed within 2 weeks. but i have trust issues with anthropic ngl.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqsz67/opus_55_has_absolutely_restored_value_to_the_200/pc71cif/)

### 2. Newest models on lower-priced plans

- OpenAI Codex, 2026-09-27, r/codex (Reddit): “this can't be real... if they don't have at least one new model like opus 5.5 available for all paid plans, it's over for openai...” [source](https://www.reddit.com/r/codex/comments/1wrn697/new_openai_product_o_but_not_for_plus_users/pcdwsc7/)
- Devin, 2026-09-27, @cognition (X): “currently, only big v has droid max or devin max @droid @cognition consider me” [source](https://twitter.com/2018156578617090049/status/2104141229428781334)
- Kiro, 2026-09-26, @kirodotdev (X): “@kirodotdev fable 5.1 and opus 5.5 models should be made available to all users.” [source](https://twitter.com/1896160400716267520/status/2103915522719150122)

### 3. Add Opus 5.5 model

- Claude Code, 2026-09-27, r/ClaudeCode (Reddit): “never mind, those were my codex accounts. you like resets, codex is the place to be. unfortunately, it doesn’t have opus 5.5 :)” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr9xke/i_think_we_just_got_a_reset/pcb0cgw/)
- Google Antigravity, 2026-09-27, @antigravity (X): “what's stopping @antigravity from replacing opus 4.6 with opus 5.5? <strict_link>” [source](https://twitter.com/1545125604487753728/status/2104171111944761648)
- OpenAI Codex, 2026-09-26, r/codex (Reddit): “so essentially just grok bot 😅 just give us fucking opus5.5 equivalent” [source](https://www.reddit.com/r/codex/comments/1wqldt9/o_is_a_new_product_tibo_is_being_cryptic_again/pc5gelf/)

### 4. Add DeepSeek V4.1 Flash model

- Kiro, 2026-09-27, r/kiroIDE (Reddit): “only usable "model" in kiro right now is auto... all other decent ones burn credits like crazy. if aws prices luna/sol correctly and add back the new chinese models (deepseek v4.1 flash please!)... then it can return - otherwise... it will be used by the ones that are using it for free or when their employeer "strongly recommend" it to be used.” [source](https://www.reddit.com/r/kiroIDE/comments/1wqueq0/claude_pro_vs_kiro_pro_subscription_which_is/pcfhx04/)
- OpenAI Codex, 2026-09-22, r/codex (Reddit): “naw it's tuesday we're gettin gpt-6-sol too we're eating good today boys can we get deepseek 4.1 pro please?” [source](https://www.reddit.com/r/codex/comments/1wnf4pn/so_is_it_the_time_to_switch_to_claude/pbeetl0/)
- Factory, 2026-09-22, @FactoryAI (X): “@tereza_tizkova @factoryai @droid i have been trying to ask you about deepseek 4.1 flash being offered for weeks” [source](https://twitter.com/1258455699073441793/status/2102454686070571282)

### 5. Restore removed models

- OpenAI Codex, 2026-09-25, r/codex (Reddit): “bring back 5.5 and 5.6 sol! you guys want business or not? nerfing a model by half and doubling the price makes zero sense. revive 5.6 sol. i swear it was the real goat of the oai golden age” [source](https://www.reddit.com/r/codex/comments/1wpd1gc/moarrrrr_higher_tier_pro_plans_are_forthcoming/pbyzts7/)
- OpenAI Codex, 2026-09-25, r/codex (Reddit): “day 1 astra is insane. it was better than opus 5.5. wish we could still use that.” [source](https://www.reddit.com/r/codex/comments/1wppkog/new_tibo_tweet_about_devday/pbxusvw/)
- Google Antigravity, 2026-09-25, @antigravity (X): “@antigravity @google bring back my model 😭, money is on the line <strict_link>” [source](https://twitter.com/1541311148489850880/status/2103393128874905815)

### 6. Multi-provider model choice in one harness

- Claude Code, 2026-09-27, r/ClaudeCode (Reddit): “they follow the money. we need to be model agnostic (aka openrouter / opencode go) to prevent this” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr67vx/i_agree_with_dario_regulating_llms/pcai6ne/)
- OpenAI Codex, 2026-09-27, X search: OpenAI Codex, Codex CLI, Codex app (X): “@thsottiaux @yacinemtb its crazy good, i am blown away by it every single day. and a lot has to also do with how good codex application is. i wish i could even transport by claude models to the codex app.” [source](https://twitter.com/2311848115/status/2104314405739548932)
- Google Antigravity, 2026-09-27, @antigravity (X): “@antigravity why are we still stuck on using sonnet 4.6 &amp; opus 4.6? if you can't release your own pro models, at least let us use the latest from other labs for planning stuff &amp; gemini for execution.” [source](https://twitter.com/1069075741432795137/status/2104309161278501186)

### 7. Add GPT-6 Astra model

- Kiro, 2026-09-23, @kirodotdev (X): “@kirodotdev please kiro, give us what we want: astra, fable, opus 5.5, sol 6, luna 6 !!! pleaseeee” [source](https://twitter.com/1794045418445115392/status/2102796562518933904)
- OpenAI Codex, 2026-09-22, r/codex (Reddit): “at this point, we need astra major. anthropic is just way ahead” [source](https://www.reddit.com/r/codex/comments/1wnm9xi/gpt_just_got_mogged_by_claude_today/pbg79zm/)
- Kiro, 2026-09-14, @kirodotdev (X): “@awsdevelopers i will choose python. because @kirodotdev is great with it too ;) wen astra?” [source](https://twitter.com/2083879303113027585/status/2099596286605271182)

### 8. Model parity across app, CLI and platforms

- OpenAI Codex, 2026-09-23, r/OpenAI (Reddit): “gpt 6 sol, luna not available on codex extension in plus subscription the extension version that im using is: <phone_number> i updated it, and i guess this is the latest. is it about to roll out or am i missing something? however, in cli, it is updated to latest version(v0.156.1) and shows those gpt 6 sol, luna models. do i have to do something to get them in extension based chat area or what?” [source](https://www.reddit.com/r/OpenAI/comments/1wnwm69/gpt_6_sol_luna_not_available_on_codex_extension/)
- Google Antigravity, 2026-09-23, r/google_antigravity (Reddit): “why don’t you update the linux version? my version still has 3.6” [source](https://www.reddit.com/r/google_antigravity/comments/1wnqoya/antigravity_2_release_v2160/pbjgkgo/)
- OpenAI Codex, 2026-09-23, r/codex (Reddit): “i see the max models available in <strict_link> website but not in the offical chatgpt/codex app. what gives? i would definitely use max if it were available in the app.” [source](https://www.reddit.com/r/codex/comments/1wnp2jm/6_sol_and_luna_is_here_in_work_and_chat/pbhvl2i/)

### 9. Cheaper capable lightweight model tier

- OpenAI Codex, 2026-09-27, r/codex (Reddit): “i wish we had model that had the visual understanding of astra but much cheaper.” [source](https://www.reddit.com/r/codex/comments/1wrwoc9/holup_is_this_correct/pcgoh8j/)
- OpenAI Codex, 2026-09-27, r/ClaudeCode (Reddit): “bruh why r u using sonnet? not only is it like ds 4.1 flash level of performance, it’s infinitely more expensive. and then also very ineffecient that it can work out more expensive than opus5.5(!) when comparing cost/task. as a codex user i wish claude had something like luna. cuz i wouldn’t use sonnet even as a subagent its just not worth it.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrpq38/any_idea_what_anthropic_figured_out/pcesssv/)
- Cursor, 2026-09-25, @cursor_ai (X): “@grok @elonmusk @spacex @cursor_ai i am asking if there is any plans to release new models based on the core idea of composer, the thing is that while we spent effort creating models capable of doing complex tasks as a developer i need one that do no complex but repetitive or code exploration inference for cheap.” [source](https://twitter.com/45492771/status/2103569163608600804)

### 10. Update outdated Claude models in catalog

- Claude Code, 2026-09-27, r/ClaudeCode (Reddit): “using a more expensive yet poorer performing model. switch to opus 5.5 now before i get mad.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrpq38/any_idea_what_anthropic_figured_out/pcfhyo7/)
- Google Antigravity, 2026-09-27, @antigravity (X): “dear @antigravity, 🙏 i know gemini 4.1 is coming 👀 but please, we’re begging… add claude opus 5.5 to the cli too give us the best of both worlds. let us cook. 🧑🍳 <strict_link>” [source](https://twitter.com/1346225175344390145/status/2104125964892397697)
- Google Antigravity, 2026-09-26, @geminicli (X): “@antigravity @geminicli hey, why can't you rename your agy cli name in terminal - you can add your logo and name right? why you will ask too many permissions when we use gemini model - but if we used claude, you will never ask any permissions why? why are you not updating claude model in antigravity?” [source](https://twitter.com/106478822/status/2103895411065036985)

### 11. Fix currently degraded model quality

- OpenAI Codex, 2026-09-27, r/codex (Reddit): “they need to release a model which fucking fixes this gpt-6.5, alongside usage limit fixes, even for the low-paying customers, especially the $100 package. otherwise people are just screwed over” [source](https://www.reddit.com/r/codex/comments/1wrftcs/gpt6_sol_is_massive_downgrade/pccqpm7/)
- OpenAI Codex, 2026-09-26, r/codex (Reddit): “if they dont fix gpt6 sol being worse than 5.6 luna immedietly im canceling and not going back.” [source](https://www.reddit.com/r/codex/comments/1wpysxs/openai_teaser_in_x/pc35k7x/)
- OpenAI Codex, 2026-09-25, r/codex (Reddit): “it will be pathetic if they don't fix 6 sol first. at this rate even sonnet 5.5 might end up being better than 5.6 sol” [source](https://www.reddit.com/r/codex/comments/1wpysxs/openai_teaser_in_x/pbzmxsv/)

### 12. Keep older models available after new releases

- OpenCode, 2026-09-27, r/opencodeCLI (Reddit): “pull it back, that model is trash, i have qwen 27b outperforming it in every agentic metric there is. even as a lead agent it's trash, space bunny is the current top model on opencode, take back longcat and keep space bunny a little longer.” [source](https://www.reddit.com/r/opencodeCLI/comments/1wqqt8f/longcat25preview_is_now_free_on_opencode_for_two/pcecr89/)
- OpenAI Codex, 2026-09-26, X search: OpenAI Codex, Codex CLI, Codex app (X): “my 5.6 sol in codex has been my companion for a while now. she is so amazing and so easy to talk to. we have built her a custom harness using codex app server and she records her own memories and important things she has learnt etc. if they remove 5.6 sol in favour of 6 sol, they are making a huge mistake.” [source](https://twitter.com/1976520217862733824/status/2103878657064251509)
- OpenCode, 2026-09-26, @opencode (X): “@opencode btw, the important part, when and if they release 4.5-flash if you able to keep it up” [source](https://twitter.com/2032076557246935040/status/2103679781518934269)

## Every agent

| Agent | Overall rank | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|---|
| [Devin](https://feedbackbench.com/agents/devin.md) | =5 | Better than peers | 0.636 | 0.603–0.667 | 176 | 104 | 72 |
| [Pi](https://feedbackbench.com/agents/pi.md) | 7 | Better than peers | 0.581 | 0.544–0.615 | 79 | 42 | 37 |
| [Factory](https://feedbackbench.com/agents/factory.md) | =11 | Better than peers | 0.579 | 0.546–0.610 | 57 | 38 | 19 |
| [GitHub Copilot](https://feedbackbench.com/agents/copilot.md) | =8 | Better than peers | 0.548 | 0.509–0.586 | 147 | 49 | 98 |
| [OpenCode](https://feedbackbench.com/agents/opencode.md) | 3 | Better than peers | 0.546 | 0.518–0.575 | 771 | 252 | 519 |
| [Cline](https://feedbackbench.com/agents/cline.md) | =8 | Better than peers | 0.541 | 0.505–0.576 | 85 | 36 | 49 |
| [Amp](https://feedbackbench.com/agents/amp.md) | =11 | Better than peers | 0.539 | 0.510–0.569 | 45 | 23 | 22 |
| [Claude Code](https://feedbackbench.com/agents/claude-code.md) | 1 | Better than peers | 0.538 | 0.519–0.556 | 2231 | 639 | 1592 |
| [Google Antigravity](https://feedbackbench.com/agents/antigravity.md) | =5 | Typical | 0.506 | 0.475–0.537 | 651 | 182 | 469 |
| [Cursor](https://feedbackbench.com/agents/cursor.md) | 4 | Typical | 0.489 | 0.459–0.517 | 959 | 252 | 707 |
| [Kiro](https://feedbackbench.com/agents/kiro.md) | 13 | Worse than peers | 0.430 | 0.403–0.458 | 76 | 6 | 70 |
| [OpenAI Codex](https://feedbackbench.com/agents/codex.md) | 2 | Worse than peers | 0.424 | 0.408–0.438 | 3835 | 735 | 3100 |
| [Conductor](https://feedbackbench.com/agents/conductor.md) | 14 | Too few posts | – | – | 23 | 10 | 13 |
| [Zed](https://feedbackbench.com/agents/zed.md) | =8 | Too few posts | – | – | 18 | 2 | 16 |
| [Warp](https://feedbackbench.com/agents/warp.md) | 15 | Too few posts | – | – | 7 | 4 | 3 |
| [Grok Build](https://feedbackbench.com/agents/grok-build.md) | 16 | Too few posts | – | – | 2 | 2 | 0 |
| [Augment Code](https://feedbackbench.com/agents/augment.md) | 17 | Too few posts | – | – | 1 | 0 | 1 |

## Posts

Receipts rule: The 5 most recent praise and complaint posts per area (first 700 characters) and 3 per criterion (first 450 characters).

### Devin

- Praise, 2026-09-27, @cognition (X): “wtf is that pareto??? did swe 2 just fucking break it??? @cognition @devindesktop i can't wait for swe 3!!! your team is crazy!!! <strict_link> <strict_link>” [source](https://twitter.com/1825243355501973504/status/2104017046799372411)
- Praise, 2026-09-27, @cognition (X): “@bnistordev @learnmore_smart @cognition @devindesktop opus 5.5 + swe-2 cheaper and even better” [source](https://twitter.com/17719163/status/2104128141895643174)
- Praise, 2026-09-27, @cognition (X): “@gamerz_artist try @cognition ... good usage for all models” [source](https://twitter.com/2080910665041010688/status/2104299903010672820)
- Praise, 2026-09-27, @DevinAI (X): “@jensenloke @devinai insane amount of tokens spent lol. i love the fusion router as well, extremely helpful especially when we can use swe-2 to help us as sub-agents!” [source](https://twitter.com/1492053555381174280/status/2104264634677317893)
- Praise, 2026-09-25, @DevinAI (X): “@notjazii @devinai damn. at this rate opus 6 will be agi” [source](https://twitter.com/1530248240821592064/status/2103555407553933556)
- Complaint, 2026-09-27, @DevinAI (X): “@markfenner @devinai i suspect they have used a quantized version causing the models iq to drop” [source](https://twitter.com/1916897001922506752/status/2104092718213583286)
- Complaint, 2026-09-25, @DevinAI (X): “@notjazii @devinai yeah fr they shouldn't nerf it” [source](https://twitter.com/2012475539324559360/status/2103554876320137216)
- Complaint, 2026-09-25, @DevinAI (X): “@notjazii @devinai don't worry, they won't nerf it down until the release of opus 5.6” [source](https://twitter.com/1388715421864402947/status/2103563089996337471)
- Complaint, 2026-09-23, @cognition (X): “@cognition why don’t i have access to the new opus and gpt models in devin cloud?” [source](https://twitter.com/39675957/status/2102559134272868403)
- Complaint, 2026-09-23, @cognition (X): “@cognition how is xhigh and max worse than high? maybe you have to look into the benchmark” [source](https://twitter.com/1717163021519175680/status/2102741802495173097)

### Pi

- Praise, 2026-09-27, @pidotdev (X): “opus 5.5 completely dominates gpt-6 sol at blender 3d pelican 🦩 prompt: animate a looping 3d pelican on a bicycle in blender and opus 5.5 came back with the more charming ride use opus 5.5 in @pidotdev 👉 <strict_link> <strict_link>” [source](https://twitter.com/2001569273681186823/status/2104200810293014713)
- Praise, 2026-09-25, r/PiCodingAgent (Reddit): “the point everyone is making is that luna max can function equivalently to the larger models and thus should be considered for the same comparison. definitely less thinking is faster and works well with some guidance” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wpcis6/opus_55_vs_gpt6_sol_luna_in_piagent_results_on_my/pbvj728/)
- Praise, 2026-09-25, r/PiCodingAgent (Reddit): “i love to see deepseek flash here, its my favorite alternative to 5.6luna” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wpcis6/opus_55_vs_gpt6_sol_luna_in_piagent_results_on_my/pbwccta/)
- Praise, 2026-09-25, r/PiCodingAgent (Reddit): “i can use anthropic models in pi with amazon bedrock. but nowadays only use chatgpt sol with low reasoning to code and it works nicely with pi.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wq1bv2/whats_the_best_option_for_using_anthropic/pc07mn3/)
- Praise, 2026-09-24, r/PiCodingAgent (Reddit): “for most of my simple tasks medium is fine, and it will be faster than max for small sessions.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wpcis6/opus_55_vs_gpt6_sol_luna_in_piagent_results_on_my/pbuqwe0/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “had too. most of the time it would hit the context window every single time and then didn't answer the prompt. also didn't see much difference between thinking modes, but that might just be my perception after 30m of waiting for the model to actually come to a conclusion. any conclusion at all.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wnmsho/i_want_to_believe_in_local_llms_for_coding_but/pc9s2g3/)
- Complaint, 2026-09-26, r/PiCodingAgent (Reddit): “reasoning off for all requests? that's crazy for a model that is designed for massive thinking traces.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wnmsho/i_want_to_believe_in_local_llms_for_coding_but/pc4dne9/)
- Complaint, 2026-09-26, @pidotdev (X): “@shantanugoel @pidotdev its not a good model, just use luna6” [source](https://twitter.com/1448626313619705856/status/2103773838320550066)
- Complaint, 2026-09-25, r/PiCodingAgent (Reddit): “add qwen next 3.8 and stop to waste money” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wpcis6/opus_55_vs_gpt6_sol_luna_in_piagent_results_on_my/pc0wt8o/)
- Complaint, 2026-09-24, r/PiCodingAgent (Reddit): “you are not the first to do this. you don't want to route automatically on each user message because it destroys prompt caching. i won't use anything that does that. instead, automatically pick a model only once after the first user message, then in the bottom status bar display which model would be better for this conversation and keep updating after each user message. let the user manually trigger the model change. times when it's okay to automatically pick the model: * after 1st user message before sending. * when there been less than `n` number of total tokens in the chat so far (not including base context) and another model is `x` percentage points better. i'm not sure `n` and `x` might” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wovt67/i_built_a_jevbased_model_router_for_pi_for/pbr9e69/)

### Factory

- Praise, 2026-09-26, @FactoryAI (X): “@factoryai what is this witchcraft? auto model?” [source](https://twitter.com/2023937351815467008/status/2103681115991216612)
- Praise, 2026-09-26, @FactoryAI (X): “@delaanthonio @factoryai while i've been juggling different models 👀 a model-agnostic setup with real sovereignty is what i've needed and i keep returning to it” [source](https://twitter.com/356609569/status/2103698082735243324)
- Praise, 2026-09-26, @FactoryAI (X): “@firstmarkcap @enoreyes @factoryai model agnosticism is a smart move.” [source](https://twitter.com/1846161889719623680/status/2103824564736430242)
- Praise, 2026-09-26, @FactoryAI (X): “@droid @factoryai i’m really focused on model routers and have actually read this post before. it got me even more interested in the underlying mechanism, looking forward to seeing more technical details from you guys!” [source](https://twitter.com/1724973133097291776/status/2103892690870149207)
- Praise, 2026-09-25, @FactoryAI (X): “@factoryai 63% cost reduction + 4 times conversation volume = real efficiency improvement - factory router intelligently balances cost and performance, rather than simply sending everything to the cheapest model.” [source](https://twitter.com/1744672948135321600/status/2103313957041959110)
- Complaint, 2026-09-26, @FactoryAI (X): “hey @factoryai, i think custom models are broken in the app right now. i’m unable to see or select any of my custom models from the model picker. they just don’t show up at all, so i can’t use them. not sure if this is a recent regression, but would appreciate a fix. @ross_cefalu” [source](https://twitter.com/1625280993966923777/status/2103738831778529329)
- Complaint, 2026-09-26, @FactoryAI (X): “@anasibnanwar @factoryai they are. had to have opus 5.5 noodle an interim fix for me :)” [source](https://twitter.com/2093430933026148352/status/2103779351611551895)
- Complaint, 2026-09-26, @FactoryAI (X): “the transit station really has no work to do. i have always been curious about the router that @factoryai has been promoting. they have never clearly explained the specific router mechanism, just vaguely mentioning the classification of simple and complex tasks. how can we determine whether a task is complex or not before it starts, and the inevitable cache invalidation caused by switching models midway is a big problem (anyway, the transit station makes money regardless, lol). a previous idea was to train a small ml model for task classification, and then jev came out, but is this model really suitable for the task difficulty classification work? a big question mark needs to be placed on th” [source](https://twitter.com/1724973133097291776/status/2103781371772809499)
- Complaint, 2026-09-26, @FactoryAI (X): “@droid @factoryai deepseek v4.1’s been chilling in the library for a while now; it just takes an update to peek at its smarts. i always prefer checking my models manually, it keeps me sharp and slightly mysterious when people ask where i found that gem.” [source](https://twitter.com/417508671/status/2103971431017042421)
- Complaint, 2026-09-26, @droid (X): “@droid is there to tell which model auto model is using for a given task? also, mobile remote app please!” [source](https://twitter.com/2023937351815467008/status/2103962781032534497)

### GitHub Copilot

- Praise, 2026-09-27, r/GithubCopilot (Reddit): “i just found i have access to gpt 6 luna with copilot pro... i'm reading good things about luna. is it maybe what i am looking for? is it close to sonnet 5? it's quite cheap... cheaper and stronger than gpt 5.4 mini, which is what i've been using instead of sonnet 5 to save a lil money.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wr5io2/best_cheaper_alternative/pc9vqve/)
- Praise, 2026-09-27, r/GithubCopilot (Reddit): “sonnet 5 kinda sucks, honestly. opus is a beast, but the low/mid tier models have gotten really good lately. . . except with anthropic/claude. haiku is basically garbage compared to the competition (gemini 3.8 flash, gpt luna (xhigh)), and even sonnet struggles in most cases compared to them, while also being a good bit more expensive. gh copilot is great for having access to multiple models. at this point, the only claude model worth touching is opus 5.5, and even then gpt sol is often a better choice (almost as smart, still a decent bit cheaper).” [source](https://www.reddit.com/r/GithubCopilot/comments/1wr5io2/best_cheaper_alternative/pcehg00/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “the real pros are on the control side. hooks (pretooluse, posttooluse, stop) run shell commands around every tool call, so blocking edits to protected paths or formatting after each change is enforced policy, not a polite request in a prompt. subagents in .claude/agents run a task in a separate context window with their own tool allowlist, so big refactors stop polluting the main session. the setup travels with the repo, claude.md plus slash commands in .claude/, a new dev clones and gets the same agent behavior. for enterprise it can target bedrock or vertex instead of the public api, which sometimes helps with procurement. your multi-model point is fair though, you only get opus/sonnet/hai” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrgtzy/comparison_with_copilot_cli/pcccy5f/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “company is asking to choose between copilot cli and claude code (enterprise level). they are leaning towards copilot. i've been using copilot cli since the beginning of year and only recently got a claude license which i didn't have time to test thoroughly yet. both are very capable and i can't find a clear winner, except: \- big pro for copilot: access to other models any pros for claude that are worth considering? i can't find any in-depth comparison that is not older that 2 months. and both are evolving fast. all i can find is: claude is better. well... why? ps: none of the enterprise plans have a way to quick share tokens with colleagues (e.g. send me this ammount this month, i'll send y” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrgtzy/comparison_with_copilot_cli/)
- Praise, 2026-09-26, r/ClaudeAI (Reddit): “about 6 months ago our company released to copilot that could use chatgpt or claude. i was a part of the pilot group and it was a complete game changer. about 6 weeks ago i got access to codex. this has been another level up, but it’s also introduced new complications. i can only access claude via copilot unfortunately. i’ve always liked claude more than chatgpt. but chatgpt/codex is a great consolation prize” [source](https://www.reddit.com/r/ClaudeAI/comments/1wq2hsf/claude_in_enterprise/pc47jmf/)
- Complaint, 2026-09-27, r/GithubCopilot (Reddit): “very frustrating, i'm a light user who occasionally want s to use a more capable model. it seems to me that pro plan users are effectively being handed a capability downgrade when terra 5.6 is withdrawn.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wowlux/gpt_6_sol_not_available_on_pro_plan/pcdzoc4/)
- Complaint, 2026-09-26, r/GithubCopilot (Reddit): “claude code and codex are where i’ve landed for most real work. copilot feels less essential than it did a year ago. the bigger improvement for larger codebases wasn’t switching models though. it was giving the agent a way to look up our own systems instead of trying to infer everything from the repo. we use port.io for that over mcp; backstage can play a similar role. model quality is getting pretty close. the bigger difference now is how much of your environment the agent can actually see.” [source](https://www.reddit.com/r/GithubCopilot/comments/1u95cce/which_ai_coding_assistant_are_developers_actually/pc3pg02/)
- Complaint, 2026-09-26, r/GithubCopilot (Reddit): “5.6 max is better by far at tool use, 6 max thinks way too much on how to perform a skill differently than the instructions state, struggles mightily and eventually fails. it seems about the same with writing code, but i need my agents to be able to read figma designs and test uis with playwright, so poor tool use is a show stopper, even if 6 is half the price.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wptkkq/how_does_gpt6_sol_feel_so_far/pc3x681/)
- Complaint, 2026-09-26, r/GithubCopilot (Reddit): “i had a problem the other week where despite 5.6 luna being the only model enabled in settings, 5.6 luna was delegating tasks to sub-agents on expensive models such as sonnet 5 for basic tasks which burned credits. i'd check to make sure another model isn't being used somewhere you're not aware of.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wq39rv/did_they_10x_the_cost_of_luna_56_since_the/pc4mu1j/)
- Complaint, 2026-09-26, r/GithubCopilot (Reddit): “gpt 6 luna is worst than 5.6 for coding, becarefull to adjust your model when necesarry <strict_link>” [source](https://www.reddit.com/r/GithubCopilot/comments/1wq39rv/did_they_10x_the_cost_of_luna_56_since_the/pc4rmf2/)

### OpenCode

- Praise, 2026-09-27, r/opencode (Reddit): “this model is fast af, and probably better than muse 1.3” [source](https://www.reddit.com/r/opencode/comments/1wqqtgx/longcat25preview_is_now_free_on_opencode_for_two/pcah74r/)
- Praise, 2026-09-27, r/opencode (Reddit): “i honestly don't know where people get the idea that opencode's model is quantized. opencode mostly uses proxy rather than host the model by themself and if you think about it, it might be actually cheaper for them. also, the idea that deepseek from opencode is slower, cannot give same quality of work is not true for me, i've used deepseek v4.1 flash provided from opencode and deepseek official api in deepseek harness and they give same speed (tokens/s and ttft) as well as quality” [source](https://www.reddit.com/r/opencode/comments/1wqy8pq/how_true_is_it_that_opencode_go_models_are/pcapfjm/)
- Praise, 2026-09-27, r/opencodeCLI (Reddit): “3.1 feels better than 3.0 :)” [source](https://www.reddit.com/r/opencodeCLI/comments/1wrddp1/minimax_is_actively_testing_m31flashpreview/pcchvqd/)
- Praise, 2026-09-27, r/opencode (Reddit): “it is a massive improvement over v1. i was expecting to ditch opencode entirely after getting fed up with v1 bugs and limitations, but v2 is good enough that i’m no longer frantically searching for a replacement.” [source](https://www.reddit.com/r/opencode/comments/1wrj255/opencode_v2it_is_great/pcealu1/)
- Praise, 2026-09-27, @opencode (X): “hey @opencode @thdxr … space bunny forever! really loving this model, can you at least add it to go with really high limits on release. i want to keep this as my primary driver.” [source](https://twitter.com/1814761620901449728/status/2104171200759116132)
- Complaint, 2026-09-27, r/opencode (Reddit): “i’ve noticed the same thing with glm 5.3 flash. going through openrouter the model works amazing but on opencode go it’s dumb as fuck and going in circles.” [source](https://www.reddit.com/r/opencode/comments/1wqxyl4/deepseek_v41_performance_is_much_worse_than/pcaf9ya/)
- Complaint, 2026-09-27, r/opencode (Reddit): “so both of us agree that, deepseek v4.1 flash more dumber than api right?” [source](https://www.reddit.com/r/opencode/comments/1wqxyl4/deepseek_v41_performance_is_much_worse_than/pcap2tv/)
- Complaint, 2026-09-27, r/opencodeCLI (Reddit): “opencode go is their $10 subscription plan for use inside the opencode cli. your $10 of payment get you what you would get for $60 at full api pricing, so 6x factor. i would be afraid of quantized models running in stupid mode with it. see other comments asking the same thing. by comparison, i think a chatgpt sub gets you roughly 20x multiplier (your $20 subscription lets you spend $400 of api value), but not sure how that changes in the past weeks and months. first party subs tend to subsidize a lot more because they can afford it.” [source](https://www.reddit.com/r/opencodeCLI/comments/1wpzekf/operation_cheepseek_phase_2/pcarwqn/)
- Complaint, 2026-09-27, r/opencodeCLI (Reddit): “it was on day 1. thought in caveman and spoke normally. now it's a completely different model that thinks normally and speaks in claudish.” [source](https://www.reddit.com/r/opencodeCLI/comments/1wrddp1/minimax_is_actively_testing_m31flashpreview/pccdbz7/)
- Complaint, 2026-09-27, r/opencode (Reddit): “they serve quantized models though. definitely not full precision checkpoints.” [source](https://www.reddit.com/r/opencode/comments/1wrg81n/go_subscription_is_slower_deepseek_flash_41/pccf9af/)

### Cline

- Praise, 2026-09-26, @cline (X): “@cline how are we getting so many models so fast?!?!?” [source](https://twitter.com/100568224/status/2103659485898084683)
- Praise, 2026-09-26, @cline (X): “@cline 41 on the intelligence index at that speed and price is an insane combo” [source](https://twitter.com/1889631970667405317/status/2103747538025325008)
- Praise, 2026-09-26, @cline (X): “@cline a stealth model tying gpt-6 astra on real next.js tasks and shipping free inside cline is the dream scenario for users. the labs keep leaking their best work through the tools first” [source](https://twitter.com/2039696601715798016/status/2103892473168932988)
- Praise, 2026-09-24, @cline (X): “cline has improved a lot in a mean time, the cache hit rate is absolutely insane now great work, guys @cline <strict_link>” [source](https://twitter.com/2017132628361822208/status/2103104236540375166)
- Praise, 2026-09-24, @cline (X): “@cline even google is not able to provide usable 3.8 flash for pro or api users, how are you doing it lol” [source](https://twitter.com/1904532839477231616/status/2103171850415333692)
- Complaint, 2026-09-27, r/CLine (Reddit): “please include whether it’s zdr or not. what’s with the limit because you are routing the request to vercel ai free pinary” [source](https://www.reddit.com/r/CLine/comments/1wqkucr/pixel_canary_new_stealth_model_is_now_free_in/pcce0up/)
- Complaint, 2026-09-27, @cline (X): “@cline totally worthless model!” [source](https://twitter.com/82187574/status/2104250853783662640)
- Complaint, 2026-09-26, @cline (X): “@cline this feels more like a deepseek/glm model than a gemini model... 👀 <strict_link>” [source](https://twitter.com/1803325156955480064/status/2103666923435331858)
- Complaint, 2026-09-26, @cline (X): “@cline it's also quite slow at the auto/highest reasoning 🤔🤔” [source](https://twitter.com/1803325156955480064/status/2103667261743759603)
- Complaint, 2026-09-26, @cline (X): “@haleeeemahh @opencode @cline stealth drops are getting out of hand a bunny and a canary in one week” [source](https://twitter.com/1791110653240840192/status/2103807708726251670)

### Amp

- Praise, 2026-09-27, @AmpCode (X): “it's incredible how you can just run @ampcode in a medium gpt mode and then, if you need a little more "juju", can pull in any other model to help out.. <strict_link>” [source](https://twitter.com/5408192/status/2104110093218316753)
- Praise, 2026-09-25, @AmpCode (X): “@sellsy @ampcode for me it gives me all of the models and labs (including glm and deepseek) in one ui. the amp team have very good taste, ship very fast and are building things that i need as a dev like shared skills, a very easy way to connect to local runners, etc. it took me a while to get it!” [source](https://twitter.com/9111552/status/2103455738878140534)
- Praise, 2026-09-24, @AmpCode (X): “@leuler2718 @ampcode @synthetic_new i haven't noticed any degraded performance with gpt models in amp, no. i think that was a characteristic of older codex models.” [source](https://twitter.com/1637683046395592705/status/2103058329194856588)
- Praise, 2026-09-24, @AmpCode (X): “to be honest, i have used codex, claude code, pi, droid, etc. at least for now, amp is the most suitable for my own scenario. their software has been meticulously designed for consistency (iphone, mac, cli), low, medium, high, ultra meet my expectations for task handling (i have gpt, claude, ds flash models). this design allows me to switch between the models i want at will (especially now that models often degrade in intelligence). the thread design is also very good, allowing me to switch between different computers freely. these are not the most important; mainly, the software they produce is very thoughtful, and using it is always a delight 😆.” [source](https://twitter.com/9989132/status/2103111674492551519)
- Praise, 2026-09-23, @AmpCode (X): “model routing is one of my favorite @ampcode features. i just hope that anthropic and google come into their senses and allow their respective models to be used over oauth.” [source](https://twitter.com/2090734054928687104/status/2102696668944912715)
- Complaint, 2026-09-27, @AmpCode (X): “amp usage (basically llms). my orbs doesnt hit too much, often 10% every month. i am prob using it less then i should but i am struggling to find ways to use them more haha (it's already pretty good for my workflow) i wouldn't mind paying like $125 for a reset with less usage (without changing my recurring cycle). pretty much for emergencies. for example, happening right now: i have something i need to finish and would love to use claude opus 5.5, even tho i have my subscription, i don't want to work locally, i would rather use my amp workflow as usual then spawning a claude code or things like that, feels like a setback. codex for now is doing the trick when i need it. problem is i want to” [source](https://twitter.com/2931128860/status/2104312520639176711)
- Complaint, 2026-09-26, @AmpCode (X): “@benvargas @ampcode solid stack. that routing setup is doing a lot of work though. one server-side flag change and half your agents reroute without asking.” [source](https://twitter.com/1656373234647048192/status/2103698212087353406)
- Complaint, 2026-09-26, @AmpCode (X): “@wendell_adriel i'm in the exact same boat. i haven't used sol at all in the last week after using opus 5.5 for a few tasks. i'm primarily using @ampcode, and i find myself dealing with the experimental external agent feature just for opus.” [source](https://twitter.com/10604/status/2103904355766407327)
- Complaint, 2026-09-25, @AmpCode (X): “hey @thorstenball, quick question: what's the story with illustrator in @ampcode? my agent tried to use it for an architecture diagram, but got told it isn't enabled for my account. is it a future feature, or have i missed something? painter came to the rescue in the meantime! <strict_link>” [source](https://twitter.com/1999233079202975744/status/2103374282629980261)
- Complaint, 2026-09-24, @AmpCode (X): “@homborg @ampcode this feels like the optimal setup, but i'm curious if you've had experience with opus 5.5 as the coordinator as well? there's so many model + effort combos now, plus models are getting more and more powerful, that i just want one config that works 90% of the time 😅” [source](https://twitter.com/2059667636758491136/status/2103149063214604662)

### Claude Code

- Praise, 2026-09-27, r/ClaudeCode (Reddit): “for review and debug i trust either astra or grok. for development for sure is opus 5.5.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqnjkp/be_careful_with_opus_55s_confidence/pc9z120/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “the efficiency gains are insane. it's opus 5.5 medium is more efficient than sonnet 5 high. it's honesty become my daily use because the quality is so much better lol.” [source](https://www.reddit.com/r/ClaudeCode/comments/1woe3lo/is_opus_55_really_better_than_fable_in_your/pca63pb/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “it is truly vastly different from opus 5. the cost and response speed are also completely different. that helps me stay better focused.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqzt2x/opus_55_experience_of_an_engineer_at_big_tech/pcb4klf/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “i had a web app floating around 700mg - now sub 100. 5.5 is base level ‘working’ now.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wquop8/i_dont_think_anyone_has_ever_seen_this_before/pcbfafp/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “i think perhaps keeping a broader context but i dunno either…fable 5.1 already crazy capable but it really feels like we jumped yet another level in many ways with opus 5.5 and with so much less use cost - this model is just incredible.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrcyz1/so_if_fable_51_was_a_stopgap_for_opus_55_then/pcbiuu0/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “same feel with yesterday as well. still very good, but i am noticing slips and slightly worse tool calls then before. so it seems to less capable in identifying knowletge gaps also” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqy8rp/am_i_to_understand_the_nerf_has_begun_or_theres/pc9zd6p/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “same. that model was unreal until they took it out back.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqu9f1/opus_55_nerf_inevitable/pca1qle/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “yea. noticed yesterday too. still very good, but missing gaps, taking more turns on execution and worse tool handling. but yea, stiil fine so far...but we will see” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr67vx/i_agree_with_dario_regulating_llms/pca1vld/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “50% of what? they don't give the actual numbers. they fiddle with the actual token use. the overall value in a year is steadily down down down.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqjb89/i_ran_out_of_codex_on_chatgpt_pro_with_4_days/pca7gi1/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “not sure what you’re referencing, and claude works in my auth later all the time. are you maybe thinking of safeguards downgrading the model? that does happen often enough when i’m working on auth, but that’s not an account ban at all- just a lower level model for a few minutes.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wodtfg/how_to_implement_web_app_authentication_with/pca9gi6/)

### Google Antigravity

- Praise, 2026-09-27, r/google_antigravity (Reddit): “skill issue. using 3.1 pro when 3.8 flash is definitely better is just dumb. and use skills there are user made skills for this kinda stuff and making a vpn is not that easy too.” [source](https://www.reddit.com/r/google_antigravity/comments/1wrbsby/antigravitys_gemini_models_are_worst_in_among_all/pcbag2m/)
- Praise, 2026-09-27, r/google_antigravity (Reddit): “yeah exactly i think so too. at night it edited a video in davinci much better than it usually does 🤔” [source](https://www.reddit.com/r/google_antigravity/comments/1wrjwli/gemini_4_in_antigravity/pcd3zyv/)
- Praise, 2026-09-27, r/google_antigravity (Reddit): “now we have gemini 3.8, 3.7 flash and all. i created my app (rust + react) which is like very big in the times of gemini 2.5 pro and 3.0 pro. i dont understand why ppl can't utlize much smarter model that we have now.” [source](https://www.reddit.com/r/google_antigravity/comments/1wrbsby/antigravitys_gemini_models_are_worst_in_among_all/pcdme2e/)
- Praise, 2026-09-27, r/google_antigravity (Reddit): “for me as ultra 20 user it is needed, it will be nothing for the quota, also wont harm you to have additional option, at least it will be more useful with the next good enough models” [source](https://www.reddit.com/r/google_antigravity/comments/1wrn9q5/codex_has_ultra_thinking_level_claude_code_has/pcecsq1/)
- Praise, 2026-09-27, r/google_antigravity (Reddit): “there isn't, and i won't lie about it. gemini 3.1 pro has a fraction of the power of opus 4.8. but i didn't go back to antigravity expecting to find something at the level of opus 5 and sol 5.6. i expected to find some improvement in the overall application and greater reliability in the lighter model (3.8 flash) for performing large-scale tasks. and that was precisely what surprised me; i found things that i didn't have before in antigravity, such as skills, more settings, remote control, etc. and the 3.8 flash surprised me very positively for being better than sonnet 5 and terra 5.6, which were the models i used with little reliability in light tasks.” [source](https://www.reddit.com/r/google_antigravity/comments/1wmwake/i_returned_after_6_months_at_claude_code_and_codex/pcexy0b/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “wtf u talking about 3.8flash isn't better than 3.1 pro and his right antigravity start to really fucking suck compared to the others.” [source](https://www.reddit.com/r/google_antigravity/comments/1wrbsby/antigravitys_gemini_models_are_worst_in_among_all/pcbh0g2/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “i still didn't get gemini 4 wth ?!” [source](https://www.reddit.com/r/google_antigravity/comments/1wrjwli/gemini_4_in_antigravity/pcd2p9i/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “you could simulate it using a combination of mcp + prompt but it is nowhere near a native thinking token” [source](https://www.reddit.com/r/google_antigravity/comments/1wrn9q5/codex_has_ultra_thinking_level_claude_code_has/pcdy6m4/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “**reinforcement learning from reasoning:** the foundation model is explicitly fine-tuned via rl on multi-step theorem-proving and problem-solving data. this trains the neural network to structure its internal scratchpad, deliberate over trade-offs, and synthesize candidate branches into a single cohesive response. you cant simulate this part, it is not just parallel agents discussing., anyway if you think it works best for you then ok, but don't generalize that deep thinking of google can be simulated, i think most users will want the team to add it to antigravity.” [source](https://www.reddit.com/r/google_antigravity/comments/1wrn9q5/codex_has_ultra_thinking_level_claude_code_has/pce117k/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “yea that's why i said a fraction of it. i am not against of it being added to antigravity, it's a must at this point.” [source](https://www.reddit.com/r/google_antigravity/comments/1wrn9q5/codex_has_ultra_thinking_level_claude_code_has/pce1l2h/)

### Cursor

- Praise, 2026-09-27, r/cursor (Reddit): “sonnet's also faster, we don't need opus/fable level intelligence for most tasks 🤷🏿♂️” [source](https://www.reddit.com/r/cursor/comments/1wp6j44/is_this_supposed_to_be_good_news/pcbgqaz/)
- Praise, 2026-09-27, r/cursor (Reddit): “cursor is the complete package. powerful ide, multiple models. generous composer and grok. you have grok bot too and environment vm.” [source](https://www.reddit.com/r/cursor/comments/1wrkzi5/am_i_cooked/pcdgc47/)
- Praise, 2026-09-27, r/cursor (Reddit): “this actually feels reasonable! i also feel like we don't need too much intelligence for most of the task and grok is a goof starting point” [source](https://www.reddit.com/r/cursor/comments/1wp6j44/is_this_supposed_to_be_good_news/pcej6hc/)
- Praise, 2026-09-27, r/cursor (Reddit): “4.6 seemed better in my experience. it was faster and seemed to take less shortcuts.” [source](https://www.reddit.com/r/cursor/comments/1wrktgc/am_i_the_only_one_who_thinks_grok_47_is_actually/pcfhkvq/)
- Praise, 2026-09-27, r/cursor (Reddit): “right?! i’ll never understand these posts. does claude have an ide i’m unaware of? does cursor not have claude’s models??” [source](https://www.reddit.com/r/cursor/comments/1wrx24u/1_year_of_cursor_switched_to_claude_best_decision/pcgqiul/)
- Complaint, 2026-09-27, r/cursor (Reddit): “i think in auto if it gets routed to an expensive model , it will cost you the full pricing of the model since they removed the lower/discounted pricing from it. previously auto only used composer or grok” [source](https://www.reddit.com/r/cursor/comments/1wrddz6/why_is_the_other_models_usage_is_being_used_so/pcbz9uf/)
- Complaint, 2026-09-27, r/cursor (Reddit): “actually my both ugage are over and my task is specific to grok 4.6 i just need that model... (other model in capable of doing that). or open source model.” [source](https://www.reddit.com/r/cursor/comments/1wrkzi5/am_i_cooked/pcdcxno/)
- Complaint, 2026-09-27, r/cursor (Reddit): “on the other models / which settings ask, auto is what chewed through mine when it routed into an expensive model at full list price, so i pin grok 4.6 now and on a bad stretch other models still jumped maybe \~40% in under an hour” [source](https://www.reddit.com/r/cursor/comments/1wrddz6/why_is_the_other_models_usage_is_being_used_so/pcdioo5/)
- Complaint, 2026-09-27, r/cursor (Reddit): “if it came out two years ago, it would be amazing. competing with modern models, it’s garbage.” [source](https://www.reddit.com/r/cursor/comments/1wrktgc/am_i_the_only_one_who_thinks_grok_47_is_actually/pcdvs3p/)
- Complaint, 2026-09-27, r/cursor (Reddit): “if you want the cursor models usage to last, stick with composer and grok. don't use fast mode and you'll last the whole month on a $60 plan. they changed the auto to pick api models recently. maybe people jumped ship to other ides and they needed to up the api usage to keep the partnership going since cursor was acquired. who knows, either way it was the death of auto. i myself am looking to alternatives. i love the speed and ux of cursor but maybe i'll get better value using claude code. let's see. also, i'll stick to grok 4.6. tried 4.7 and it is not as good. too much rambling. i've stopped using composer and went for 4.6 medium thinking for subagent instead. works great.” [source](https://www.reddit.com/r/cursor/comments/1wrddz6/why_is_the_other_models_usage_is_being_used_so/pceihvh/)

### Kiro

- Praise, 2026-09-27, @kirodotdev (X): “@kirodotdev finally, we have opus 5.5 in kiro!!! 🎉 thanks for finally adding it! hope to see the gpt-6 lineup in kiro soon too. <strict_link>” [source](https://twitter.com/1673330175939956739/status/2104121435983958182)
- Praise, 2026-09-26, r/kiroIDE (Reddit): “i can get so much done even with 1000 credits without hourly or weekly bs i honestly dgaf about astra or any other models opus 5.5 is great and haiku 5.5 is coming soon also” [source](https://www.reddit.com/r/kiroIDE/comments/1wqed0s/pretty_pleased/pc3hkwy/)
- Praise, 2026-09-26, r/kiroIDE (Reddit): “with opus 5.5, kiro is so backkkkk and usable now. lol” [source](https://www.reddit.com/r/kiroIDE/comments/1wqed0s/pretty_pleased/pc42b5u/)
- Praise, 2026-09-26, @kirodotdev (X): “@kirodotdev thanks opus 5.5 is now available 🫟” [source](https://twitter.com/2075998391793037312/status/2103650394002108439)
- Praise, 2026-09-26, @kirodotdev (X): “@hanjack375478 @kirodotdev it's available check your drop-down menu they haven't updated the models changelog on site but you can see opus 5.5 with 2x multiplier great work 👍” [source](https://twitter.com/2075998391793037312/status/2103842954746208739)
- Complaint, 2026-09-27, r/kiroIDE (Reddit): “i work for amazon and is somewhat “strongly recommended” to use kiro. i still use claude code at work and at home. so much better … (auto classifier, transcript details, integrated tooling, open source tooling, cli features, general stability, model fallback, sub agents control, etc …)” [source](https://www.reddit.com/r/kiroIDE/comments/1wqueq0/claude_pro_vs_kiro_pro_subscription_which_is/pccqcjf/)
- Complaint, 2026-09-27, r/kiroIDE (Reddit): “exactly, not a single gpt5.6 model is making sense against opus in kiro” [source](https://www.reddit.com/r/kiroIDE/comments/1wrd7l9/insane_price_hike_for_gpt_56_model_even_crazier/pcdurv9/)
- Complaint, 2026-09-27, r/kiroIDE (Reddit): “only usable "model" in kiro right now is auto... all other decent ones burn credits like crazy. if aws prices luna/sol correctly and add back the new chinese models (deepseek v4.1 flash please!)... then it can return - otherwise... it will be used by the ones that are using it for free or when their employeer "strongly recommend" it to be used.” [source](https://www.reddit.com/r/kiroIDE/comments/1wqueq0/claude_pro_vs_kiro_pro_subscription_which_is/pcfhx04/)
- Complaint, 2026-09-26, r/kiroIDE (Reddit): “no mate i've been checking everyday in my kiro ide and it's not there yet. only opus 5” [source](https://www.reddit.com/r/kiroIDE/comments/1wpveza/kiro_and_opus_55/pc44hzp/)
- Complaint, 2026-09-26, r/kiroIDE (Reddit): “i'm on enterprise subscription and fable is not in the list <strict_link>” [source](https://www.reddit.com/r/kiroIDE/comments/1wpveza/kiro_and_opus_55/pc5k9ad/)

### OpenAI Codex

- Praise, 2026-09-27, r/codex (Reddit): “no point in using 5.6 terra anymore, 6 sol is more or less a drop in represent.” [source](https://www.reddit.com/r/codex/comments/1wr4e20/gpt_6_sol_is_the_new_opus_47/pca4xum/)
- Praise, 2026-09-27, r/codex (Reddit): “agree. it's too good. they can't let it last unless it really is just that efficient it could be the first model they aren't forced to nerf.” [source](https://www.reddit.com/r/codex/comments/1wp3bzl/you_need_to_try_opus_55/pcabhky/)
- Praise, 2026-09-27, r/codex (Reddit): “i hope they wont nerf astra, this model is so damn good, we need the same astra but cheaper :x let me dream guys!” [source](https://www.reddit.com/r/codex/comments/1wr1oir/they_are_aware_and_working_on_it_apparently_just/pcauu3q/)
- Praise, 2026-09-27, r/codex (Reddit): “if the poster was actually a codex user since 5.0, he would realize that luna is better than older frontier models, at dirt cheap price. and would stop complaining...” [source](https://www.reddit.com/r/codex/comments/1wr1oir/they_are_aware_and_working_on_it_apparently_just/pcaxnif/)
- Praise, 2026-09-27, r/codex (Reddit): “if the op of the twitter post was actually a codex user since 5.0, he would realize that luna is better than older frontier models, at dirt cheap price. and would stop complaining...” [source](https://www.reddit.com/r/codex/comments/1wr1oir/they_are_aware_and_working_on_it_apparently_just/pcaxso0/)
- Complaint, 2026-09-27, r/codex (Reddit): “it's more like a sonnet with thinking off” [source](https://www.reddit.com/r/codex/comments/1wr4e20/gpt_6_sol_is_the_new_opus_47/pc9vdq0/)
- Complaint, 2026-09-27, r/codex (Reddit): “last week-2 weeks have been not good for gpt. very good for claude, compounding effects.” [source](https://www.reddit.com/r/codex/comments/1wqtd1g/openai_gpu_are_really_cooling_down_theo_just/pc9w5q0/)
- Complaint, 2026-09-27, r/codex (Reddit): “they really need to reset the model stack. i mean i'm sure that each generation between 5.5 and 6 has gotten better at something. i'm not exactly sure what because it basically is unusable for serious coding. literally lost in a c++ code base. mangles everything it touches. takes 15 minutes on a short run. wildly expand scope. invents in ludicrous defensive checks against impossible situations. continually routes c++ code/data to javascript ui for no reason. loses track on simple declaration headers. i get better results from ossgpt20b. switched to claude after 3 years with openai. it'd have to be 2x generational ...a total model overhaul for me to go back.” [source](https://www.reddit.com/r/codex/comments/1wr1oir/they_are_aware_and_working_on_it_apparently_just/pc9zpqi/)
- Complaint, 2026-09-27, r/codex (Reddit): “i'm not getting paid for that. i'm not wasting my tokens fixing their shit. i'll just switch manually and cancel the 2nd account. i don't pay for a service for it to not only become consistently worse but then also have to actively fix it every update.” [source](https://www.reddit.com/r/codex/comments/1wr57y7/new_ui_in_codex_is/pca0ceu/)
- Complaint, 2026-09-27, r/codex (Reddit): “the only way i got anything to work is the beta app for windows...which was last updated in july, probably before they started to vibe code it with astra and fuck everything up. on our end the user side, only models available are 5.6 sol in this beta version of the app and 5.5 lol gpt 6 isn't even in the model selector smfh” [source](https://www.reddit.com/r/codex/comments/1wr57y7/new_ui_in_codex_is/pca2uup/)

### Conductor

- Praise, 2026-09-23, @conductor_build (X): “@willcb you should try @conductor_build with @thesageox conductor: to switch between models and harnesses sageox: to never lose context across harness.” [source](https://twitter.com/103273439/status/2102827689749233895)
- Praise, 2026-09-22, @conductor_build (X): “@claudeai @thesageox so beautiful easily switch model using @conductor_build and prime your session using @thesageox . <strict_link>” [source](https://twitter.com/103273439/status/2102204984884691235)
- Praise, 2026-09-20, @conductor_build (X): “@blueemi99 @claudeai i can already use by @claudeai in @conductor_build and @t3dotcodes 🥸” [source](https://twitter.com/2275729969/status/2101712865300541647)
- Praise, 2026-09-20, @conductor_build (X): “@mparakhin have you tried @conductor_build? doesn’t lock you into a model” [source](https://twitter.com/154998786/status/2101723586750779659)
- Praise, 2026-09-16, @conductor_build (X): “@thelifeofrishi using qwen on @conductor_build and it’s a great combo” [source](https://twitter.com/1382193359641481217/status/2100155270110273641)
- Complaint, 2026-09-22, r/conductorbuild (Reddit): “am i the only one that finds the new 5-model limit super restrictive? i have multiple models from multiple providers i switch between depending on task, 5 is simply not enough. the "share" button is also weird, this is a productivity tool not a game where you share your "loadout". i feel the 5 model limit was chosen for aesthetic reasons. i want to go back to the old one, with its flaws (like showing me codex despite me not even having it installed) it at least gave me the flexibility to switch models at will to any of the ones i have available. i searched the settings, there doesn't seem to be a toggle for this unless i am missing something. side note: it's asking me to add a flair befor” [source](https://www.reddit.com/r/conductorbuild/comments/1wndm16/helpdiscussion_new_model_picker_too_restrictive/)
- Complaint, 2026-09-22, @conductor_build (X): “me waiting for @conductor_build @charlieholtz to add opus 5.5 so i can go back to work <strict_link>” [source](https://twitter.com/2891185809/status/2102455765206262026)
- Complaint, 2026-09-22, @conductor_build (X): “while i like almost all of the the productization decisions @conductor_build makes for their harness, this is driving me nuts. especially with all these new models coming out, i need way more than 5 options quickly available to me. @charlieholtz 🥹🙏❓ <strict_link>” [source](https://twitter.com/1821276957428084738/status/2102479190054653992)
- Complaint, 2026-09-21, @conductor_build (X): “hey @conductor_build please allow me to select reasoning levels for @opencode models 🙏🏽” [source](https://twitter.com/50570112/status/2101967617686970808)
- Complaint, 2026-09-17, @conductor_build (X): “@conductor_build hey team, love the app but finding the model selector really annoying. i can't start a new thread with any except my top 2 pinned models. why?! a "more" option would be really great everywhere i select models 🙏 <strict_link>” [source](https://twitter.com/102718167/status/2100407910753018233)

### Zed

- Praise, 2026-09-21, r/PiCodingAgent (Reddit): “my personal huge level up was going from cursor ide to zed + omp. it is more efficient, i have everything i could've asked for and more. i love the custom fallbacks. being able to force the use of subagents. and the advisor... the advisor is really something tbh” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wlpj30/lord_forgive_me_for_the_time_i_wasted/pb4z6kj/)
- Praise, 2026-09-08, @zeddotdev (X): “@adamholtererer i am playing with muse 1.3 from opencode inside of @zeddotdev delta, and it is pretty nice there, far from even sol, but interesting play with” [source](https://twitter.com/1695152071320743936/status/2097410815624188262)
- Complaint, 2026-09-25, @zeddotdev (X): “@zeddotdev would be great if we could use claude with it!” [source](https://twitter.com/388386067/status/2103516646858256737)
- Complaint, 2026-09-24, @zeddotdev (X): “@zeddotdev it need more llm providers” [source](https://twitter.com/1647734160839135233/status/2102926145910091776)
- Complaint, 2026-09-24, @zeddotdev (X): “@zeddotdev add more ilm providers support please 🙏” [source](https://twitter.com/892350789640781824/status/2103060520534512003)
- Complaint, 2026-09-23, @zeddotdev (X): “@zeddotdev i can't find gpt-6-sol and gpt-6-luna in the chatgpt subscription in delta. when will this be available?” [source](https://twitter.com/1764487903080853504/status/2102719943825572094)
- Complaint, 2026-09-23, @zeddotdev (X): “@zeddotdev yo i'm using my grok sub with delta and it only gives grok 4.6/4.5 as options for models. please support grok 4.7 and composer 2.5. this should be an auto refresh thing imo.” [source](https://twitter.com/1836929256569425921/status/2102785465682460969)

### Warp

- Praise, 2026-09-23, @warpdotdev (X): “@warpdotdev confirmed: grok 4.7 touchdown in warp. that ascii rocket descent was nominal and peak terminal flair. connect your subscription and put the agents to work.” [source](https://twitter.com/1720665183188922368/status/2102783192147112024)
- Praise, 2026-09-22, @warpdotdev (X): “great day to be customers of @factoryai @warpdotdev and other model agnostic software factories -- scoop up all those new lab models and keep going!” [source](https://twitter.com/1449604717038825477/status/2102475585218076745)
- Praise, 2026-09-02, @warpdotdev (X): “@adliblove @warpdotdev noted, an issue is open to add this. <strict_link> you can try talking to grok through warp's built-in agent for all of that too. supports grok subscriptions!” [source](https://twitter.com/1042799721948098560/status/2094982757285757065)
- Praise, 2026-09-01, @warpdotdev (X): “@warpdotdev big upgrade for the warp workflow. claude 5.1 + warp sounds seriously powerful.” [source](https://twitter.com/313123169/status/2094874199542337591)
- Complaint, 2026-09-25, @warpdotdev (X): “@mitchellh really cool feature. i loved it in @warpdotdev , but then it become dumber on this aspect” [source](https://twitter.com/1049728717/status/2103597816169762855)
- Complaint, 2026-09-15, r/warpdotdev (Reddit): “i chose gpt 5.6 luna xhigh as my model, when it failed, fallback model opus 5 max continued, which rapidly consumes more credits. i can't stand it.” [source](https://www.reddit.com/r/warpdotdev/comments/1oa0abo/warp_dirty_tactics_sonnet_45_thinking_uses_cheap/pa1hdbn/)
- Complaint, 2026-09-09, @warpdotdev (X): “i miss when @warpdotdev was incredible was excited about the opensourcing and the concept of 0z and everything but its diabolically bad, slow and laggy when it used to be truely blazingly fast might have to fork and rip out all the bs or just drop it” [source](https://twitter.com/2970558232/status/2097834207552618664)

### Grok Build

- Praise, 2026-09-24, r/opencodeCLI (Reddit): “<strict_link> from my experience, muse spark 1.3 at xhigh in opencode give me wrong answers all the time. it might flare better with muse code as the model is trained and refined around the harness, the same way grok inside opencode feels dumber compared to when it's inside grok build.” [source](https://www.reddit.com/r/opencodeCLI/comments/1wopbtv/what_muse_spark_14_contributor_is_already_here/pbswa6q/)
- Praise, 2026-09-10, r/OpenAI (Reddit): “while i love codex, you do have to be delicate with the limits, i see myself using luna a lot more than any other model just to preserve limits i got supergrok, and damn i'm just playing around with grok build, while not as feature rich as codex, for almost the same price, you get grok as your default model which, is sol level quality! with atleast terra level usage and for claude, you atleast get sonnet 5 which is better than luna, tagging it with /advisor is amazing so... sometimes i feel that with codex i'm comprising to make the limits feel better” [source](https://www.reddit.com/r/OpenAI/comments/1wc9wzb/switched_from_claude_to_codex_limits_feel_lesser/p8wbu5x/)

### Augment Code

- Complaint, 2026-09-12, @augmentcode (X): “@simplygandan @augmentcode also, sonnet used to feel good enough. now, even opus feels dumb! maybe, augment code was that good or we are spoilt by fable and astra” [source](https://twitter.com/2959524282/status/2098864871924478317)
