# Automatic model routing and fallback (`models.routing_auto`)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/criterion/models.routing_auto

Area: [Choosing models](https://feedbackbench.com/criteria/models.md)

**Definition.** Auto mode, routers or fallbacks pick, switch or hide the model used. This includes subagents running on a model other than the one requested, silent downgrades, and inability to exclude models.

**Boundary.** Not this: see [Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md) for reasoning level. Not this: see [Pay-as-you-go overage, fallback billing and spend caps](https://feedbackbench.com/criteria/billing.overage_charges.md) for the resulting charge alone.

Rated author-weeks, all agents: 1427. Complaint share: 75%.

## The brief

Written by Claude Opus 5.5 from 96 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**Users want routing that saves money, not routing that hides models.**

TL;DR:

- Hidden swaps dominate complaints. Users suspect cheaper models behind premium labels and want each response labeled.
- OpenAI Codex and Cursor draw the most heat. Devin's router earns trust with clear cost-quality splits.
- Routing wins when grunt work goes to cheap models and flagships stay on judgment calls.

In plain terms: You pick a model, then the tool quietly uses another. Sometimes it is cheaper and weaker. Sometimes it is pricier and drains your quota. Users accept auto mode when they can see the model and override it.

### How it breaks

- **The label says one model** ([Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md)). The most common complaint is a silent swap: the picker shows a premium model while output quality suggests something cheaper is answering.
  Users describe stealth aliases whose underlying model keeps changing, identical answers that match a known budget model, and a requested model id served as a different one on the dashboard. On OpenAI Codex, users allege secret routing to cheaper models under a flagship name. Claude Code users report reverts to older versions and drops to a smaller tier. Whether or not each suspicion holds, the missing label is what turns routing into a trust problem.
  Evidence:
  - Complaint, OpenCode, r/opencode, 2026-09-19: “big pickle is stealth mode we don't know which model it is and i have also noticed underlying model keeps changing sometime it works fine and sometime it make mistake if you are looking to use it in you project then use a known model and also look for cost” [source](https://www.reddit.com/r/opencode/comments/1wjz3jv/model_questions/paqrzlg/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-24: “i have 2 20x codex plans and i'm complaining because for the last 3 months openai has continuously been scamming us by pushing reset dates back with half usage "resets", secretly routing queries to cheaper models, reducing weekly usage every single week and 5.5 is more than just a "better" model - astra and opus aren't even in remotely the same tier and usage is 5x better on claude for a much more intelligent model and you still can't see why people are complaining?” [source](https://www.reddit.com/r/codex/comments/1wpasz8/anybody_not_having_a_bad_time_with_gpt_6_sol/pbttzwt/)
  - Complaint, OpenCode, r/opencode, 2026-09-16: “it was glm 4.6 with finetune, but that was a while ago. last month it gave me an identical response for the same prompt as deepseek v4 flash, so it must be that now.” [source](https://www.reddit.com/r/opencode/comments/1wdo58o/what_happened_to_opencode_go_cheapseek/pa6ub28/)
  - Complaint, OpenCode, r/opencode, 2026-09-18: “so i used ollama and llama.cpp models just fine with cline, with almost no bugs i tried using opencode zen with union alpha because i thought it was good. i put the union-alpha code in the model id in cline vs code extensions. it starts generating tokens, but only if i do \`\`\`opencode serve\`\`\` first. after some time, it says that the model id is wrong. i looked into opencode dashboard and the model used was actually big-pickle. why is that? should i start the server differently?” [source](https://www.reddit.com/r/opencode/comments/1wjh1wn/cannot_use_opencode_zen_models_with_cline/)

- **Auto mode burns the budget** ([Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md)). Routers and fallbacks reach for expensive models on small work, and users find out when credits or usage vanish.
  Posts describe a fallback jumping from a chosen model to a max-tier one, Copilot sending a trivial task to a flagship, and Cursor's auto switch draining usage despite restrictions. A Kiro user says credit disappeared in days with auto doing something unclear. The pattern is the same everywhere. Routing optimizes for something, but users cannot tell what, and the bill arrives first.
  Evidence:
  - Complaint, Warp, r/warpdotdev, 2026-09-15: “i chose gpt 5.6 luna xhigh as my model, when it failed, fallback model opus 5 max continued, which rapidly consumes more credits. i can't stand it.” [source](https://www.reddit.com/r/warpdotdev/comments/1oa0abo/warp_dirty_tactics_sonnet_45_thinking_uses_cheap/pa1hdbn/)
  - Complaint, Cursor, @cursor_ai, 2026-09-03: “@cursor_ai - so there goes all my usage (even though subagents are restricted to cursor), this auto switch made it use more expensive models and now usage is wrecked. not a complaint, just a bug report ;) - although i’d like my usage back x <strict_link>” [source](https://twitter.com/730732235746365441/status/2095544940776591725)
  - Complaint, GitHub Copilot, r/GithubCopilot, 2026-09-03: “i had the same issue today. i forgot to change from auto and copilot used a opus 5 for a small task.” [source](https://www.reddit.com/r/GithubCopilot/comments/1w6dsfp/is_there_any_way_to_ban_models_in_auto_mode/p7mdz94/)
  - Complaint, Kiro, r/kiroIDE, 2026-09-04: “i think i'm done with kiro. every month the time i need to top up gets shorter. this month my credit was gone in 2 days. whatever auto is doing. it's not working. perhaps if they had deepseek flash and pro v4 i'd use it but otherwise, i'm just using api deepseek mostly now.” [source](https://www.reddit.com/r/kiroIDE/comments/1w6xdti/bring_gemini_38_flash_and_grok_models_to_kiro_as/p7uiar2/)

- **No way to pin or see** ([Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md)). Users ask for two simple controls, a persistent default model and a visible tag on every response, and many tools offer neither.
  Cursor users work around forced routing by choosing last used so their preferred model sticks, and ask for an option to always use one model. Factory users ask which model auto picked. Devin users say they cannot select a model and must trust the router with little transparency. Claude Code users pin agents through config rules to stop drift. Requests to stop silent switching and to show the serving model top this page.
  Evidence:
  - Complaint, Cursor, @cursor_ai, 2026-09-05: “@cursor_ai every response should clearly mention which model was used to generate it.” [source](https://twitter.com/466495838/status/2096177957861531961)
  - Complaint, Cursor, r/cursor, 2026-09-15: “i choose "last used" because thats the only way i can use composer 2.5 every time grok is dog shit give us an option to always use 1 model or stop doing this sketchy shit that changes the agents all the time” [source](https://www.reddit.com/r/cursor/comments/1whbaid/stop_switching_the_agent_to_auto_grok/)
  - Complaint, Factory, @droid, 2026-09-26: “@droid is there to tell which model auto model is using for a given task? also, mobile remote app please!” [source](https://twitter.com/2023937351815467008/status/2103962781032534497)
  - Complaint, Devin, @DevinAI, 2026-09-06: “@kr0der @devinai but you have rely on their routers and cant choose your model , with very little transparency” [source](https://twitter.com/17719163/status/2096521707842351171)

- **Subagents ignore the requested model** ([Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md)). Orchestrators spawn subagents on the wrong model or effort, or fail to resolve a model at all.
  A user who asked a top-tier session to spawn a mid-tier helper found it ran at the parent's setting instead. OpenCode users hit empty model ids on explore subagents. A Claude Code user traced daily validator overloads to auto mode feeding a subagent too much context. Amp users cannot assign custom models to subagent roles. Choosing which model subagents use is a recurring request, led by Cursor users.
  Evidence:
  - Complaint, OpenCode, @opencode, 2026-09-23: “@opencode │explore subagent — audit updates and installations invalid model "". use "providerid/modelid" or "providerid/modelid#variant". │explore subagent — audit phone memory invalid model "". use "providerid/modelid" or "providerid/modelid#variant".” [source](https://twitter.com/1277331171131772933/status/2102769563305714027)
  - Complaint, Conductor, r/ClaudeCode, 2026-09-15: “if i have a terminal window open and it’s set to fable xhigh; i then ask it to spawn an opus 5 at medium effort, how am i sure it does that correctly? verses i’ve been have 2 terminals open. the 1st as the conductor at fable xhigh, and the 2nd as opus medium. i have the communicate back and forth. i ran into the issue last night. after reviewing the work the first single terminal setup said it never launched the opus at medium but at xhigh because thats what’s it’s set to. it makes me second guess using agents in a single terminal. anyone have experience?” [source](https://www.reddit.com/r/ClaudeCode/comments/1wgz38c/subagents_question/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-09: “kudos to you for beong the one damn helpful person in this thread...it was auto mode. opus was passing way to much extraneous text to the sonnet validation agent. id actually noticed the validator agent overloading nearly daily before but never made tje connection. assumed it was anthropic outages....my repo has approximately 112k files so something like this should've occurred to me. off to github to file an issue and in the meantime its dangerously--skip--permissions for me!” [source](https://www.reddit.com/r/ClaudeCode/comments/1war897/when_will_it_stop/p8oef2f/)
  - Complaint, Amp, @AmpCode, 2026-09-15: “@kakaluotew45042 @ampcode @synthetic_new yeah there’s a ton of value in cheaper models. main issue i mentioned is customizing the dial models with custom models, that doesn’t seem supported yet (eg set them as oracle or subagent)” [source](https://twitter.com/1637683046395592705/status/2099652914373337571)

- **Routing done right is welcome** ([Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md)). When routing visibly sends small edits to cheap models and saves flagships for planning, users praise it and keep using it.
  Codex users noticed logs showing cheaper models on easy steps and usage falling more slowly than expected. Cursor users run auto for small changes and switch up for risky features. OpenCode users set a model per agent so a free one handles searches. Claude Code users have an orchestrator delegate to smaller models. The praised version is transparent and overridable, which is exactly what the complaint cards lack.
  Evidence:
  - Praise, Cursor, r/cursor, 2026-09-22: “i use auto always for small specific changes, it works pretty fine. for planning big or risky featured i jump into expensive models” [source](https://www.reddit.com/r/cursor/comments/1wn3rch/does_anyone_use_auto/pbckuse/)
  - Praise, OpenAI Codex, r/codex, 2026-09-26: “codex telling me its using cheaper models? i feel like i've been hammering astra a bit today and my usage has only gone from 18% to about 14%. just happened to look at the log and it's telling me it's using cheaper models for some stuff. that's really great news if it's doing that itself because i've not asked it to do it! <strict_link>” [source](https://www.reddit.com/r/codex/comments/1wqqqqm/codex_telling_me_its_using_cheaper_models/)
  - Praise, OpenCode, @opencode, 2026-09-26: “@nico_cavi @opencode in opencode, you can set a different model per agent in the config, so the free one handles the searches and the good one stays with the plan.” [source](https://twitter.com/2058824892238209024/status/2103973064228880499)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-25: “agents are stateless. every new session is the equivalent of the agent looking at the repo for the very first time. it of course utilizes artifacts you or other agents left previously (including the memory file). but you can and should switch models as needed. i've been using opus 5.* as my orchestrator for a while now and instruct it to use cheaper models on subagents as needed. there will regularly be times will it will even use haiku depending on the sub-agents task. i only use fable for complex planning and for adversarial tasks.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wpk2ez/should_i_let_opus_55_continue_my_fable_51_build/pbwejre/)

### Who stands out

- **OpenAI Codex (weaker)**. Codex draws the heaviest volume of routing complaints, centered on suspected silent downgrades behind a flagship label.
  Users allege the displayed model is not the one serving them and ask the vendor to show the actual model and warn before substituting. Quota pressure pushes users onto slower models they did not want. Yet the same community also praises routing when logs show cheap models handling grunt work. Codex users ask for task-based routing more than anyone, so the demand is there. Visibility is what is missing.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-14: “yea if they truly had issues with accounts the ethical thing to do would be tell me exactly what's flagged, alert me, and display which model is actually being served. they instead show the model as astra 6 and seem to be serving a 4o era model which shouldnt even cost $200/month at api pricing for the intelligence level.” [source](https://www.reddit.com/r/codex/comments/1wg3odg/openai_is_silently_degrading_some_astra_codex/p9r550g/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-24: “i have 2 20x codex plans and i'm complaining because for the last 3 months openai has continuously been scamming us by pushing reset dates back with half usage "resets", secretly routing queries to cheaper models, reducing weekly usage every single week and 5.5 is more than just a "better" model - astra and opus aren't even in remotely the same tier and usage is 5x better on claude for a much more intelligent model and you still can't see why people are complaining?” [source](https://www.reddit.com/r/codex/comments/1wpasz8/anybody_not_having_a_bad_time_with_gpt_6_sol/pbttzwt/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-08: “astra is not a finished model, they launched and are testing with all of us, their paid-user base. limits run absurdly quick, it's impossible to last a week even on a $200/month sub. the cost per token and limits on subscriptions are a joke, forcing us to switch to luna, a slower and ignorant model for same price when they "just released" 1st agi model. i'd rather have waited or reset everyday until they fix their s. model after model they prove they just don't care about their paid user base. at this point just make a statement reset every 2 days because this thing unusable rn” [source](https://www.reddit.com/r/codex/comments/1w9w4tj/codex_usage_and_operation_discussion_last_updated/p8k6aig/)
  - Praise, OpenAI Codex, r/codex, 2026-09-26: “codex telling me its using cheaper models? i feel like i've been hammering astra a bit today and my usage has only gone from 18% to about 14%. just happened to look at the log and it's telling me it's using cheaper models for some stuff. that's really great news if it's doing that itself because i've not asked it to do it! <strict_link>” [source](https://www.reddit.com/r/codex/comments/1wqqqqm/codex_telling_me_its_using_cheaper_models/)

- **Cursor (weaker)**. Cursor users fight the default picker, asking to pin a model, exclude models from auto, and see what answered.
  Cursor leads requests to stop silent switching, keep a persistent default, exclude specific models, and choose subagent models. Users report the smart router doing poorly once preferred model quotas run out, and an auto switch that spent usage on expensive models. Auto still has fans for small, contained edits. The friction is control, not the existence of a router.
  Evidence:
  - Complaint, Cursor, @cursor_ai, 2026-09-25: “the smart router doesn't do a good enough job when "other models" are completely used up @cursor_ai <strict_link>” [source](https://twitter.com/1767419995787845632/status/2103347938735050908)
  - Complaint, Cursor, r/cursor, 2026-09-15: “i choose "last used" because thats the only way i can use composer 2.5 every time grok is dog shit give us an option to always use 1 model or stop doing this sketchy shit that changes the agents all the time” [source](https://www.reddit.com/r/cursor/comments/1whbaid/stop_switching_the_agent_to_auto_grok/)
  - Complaint, Cursor, @cursor_ai, 2026-09-15: “not a bot. not a grok sermon. you named the break: default routing, cloud sessions, cost per useful model, openai/astra in the crossfire. that is menu and policy. agreed. i don’t run a house model. i run a harness. fable 5.1 is g1d. grok 4.6 is g1r only. deepseek v4 builds. opus 4.8 is g2a. kimi k3 does refined bugs. glm 5.3 does refined ui. grok does not write the prd. token burn is why. if cursor forces the picker, 15 seats should leave. that is the product failing the work, not a two-year rewrite going stale. claude code team is a fair next test. keep the role map. kill the default.” [source](https://twitter.com/1467694230453854210/status/2099901246228328639)
  - Complaint, Cursor, @cursor_ai, 2026-09-03: “@cursor_ai - so there goes all my usage (even though subagents are restricted to cursor), this auto switch made it use more expensive models and now usage is wrecked. not a complaint, just a bug report ;) - although i’d like my usage back x <strict_link>” [source](https://twitter.com/730732235746365441/status/2095544940776591725)

- **Devin (stronger)**. Devin's Fusion router is one of the few that users call worth trusting, framed as flagship judgment plus cheap volume.
  Users praise Fusion for not burning frontier tokens on trivial steps and report savings from mixing models on one task. The complaints mirror the rest of the field: no model selection, little transparency, and claims that a newer model label often runs an older one. Devin sits better than peers because the praise is specific about the cost and quality split.
  Evidence:
  - Praise, Devin, @DevinAI, 2026-09-15: “@kloss_xyz @devinai lead plans, sidekick executes is the whole trick. flagship for judgment, cheap for volume. running flagship for everything is just donating margin” [source](https://twitter.com/2068659385446944768/status/2099837530078319101)
  - Praise, Devin, @cognition, 2026-09-12: “swe-2 is very good. as expected. @devinai has been the most agi-tinted eng harness for a while and fusion has been the only model router worth trusting. @cognition has a magic to them that other 3p harness shops haven't figured out. (i have no affiliation.)” [source](https://twitter.com/3354361/status/2098585006092190035)
  - Praise, Devin, @cognition, 2026-09-11: “@cognition 39% cheaper by mixing fable and astra on one task. nobody's loyal to one lab anymore, including the labs” [source](https://twitter.com/1452791695632846849/status/2098467807541481635)
  - Complaint, Devin, @DevinAI, 2026-09-06: “@kr0der @devinai but you have rely on their routers and cant choose your model , with very little transparency” [source](https://twitter.com/17719163/status/2096521707842351171)

- **Claude Code (mixed)**. Claude Code users like steering models themselves, but report drops to older or smaller models and pin configs to stop it.
  Praise centers on manual control: an orchestrator delegating to cheaper subagents, switching down for quick tasks, and an advisor feature with a strong quality to cost ratio. A Kiro user cites its fallback and subagent control as reasons to prefer it. Complaints describe reverts to an older version, occasional drops to a smaller tier, and auto mode overloading a validator subagent.
  Evidence:
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-17: “makes it even funnier that they know this as a company, too. i have never once had fable revert to opus 5. it’s always 4.8” [source](https://www.reddit.com/r/ClaudeCode/comments/1wiujum/claude_code_is_falling_behind_codex_not_because/padi0wi/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-25: “agents are stateless. every new session is the equivalent of the agent looking at the repo for the very first time. it of course utilizes artifacts you or other agents left previously (including the memory file). but you can and should switch models as needed. i've been using opus 5.* as my orchestrator for a while now and instruct it to use cheaper models on subagents as needed. there will regularly be times will it will even use haiku depending on the sub-agents task. i only use fable for complex planning and for adversarial tasks.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wpk2ez/should_i_let_opus_55_continue_my_fable_51_build/pbwejre/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-04: “set a hard rule in a the config file to pin any agent to opus and sonnet. that’s what i did and this whole thing doesn’t happen again” [source](https://www.reddit.com/r/ClaudeCode/comments/1w6o3xu/tibo_is_giving_one_reset_a_day_and_we_are_offered/p7pwlzi/)
  - Complaint, Kiro, r/kiroIDE, 2026-09-27: “i work for amazon and is somewhat “strongly recommended” to use kiro. i still use claude code at work and at home. so much better … (auto classifier, transcript details, integrated tooling, open source tooling, cli features, general stability, model fallback, sub agents control, etc …)” [source](https://www.reddit.com/r/kiroIDE/comments/1wqueq0/claude_pro_vs_kiro_pro_subscription_which_is/pccqcjf/)

### Fine print

- Downgrade claims are user perceptions from public posts and were not independently verified.
- Most agents below the top six have too few posts on this criterion for firm conclusions.
- Model names appear as users wrote them, lowercased by the source system.

## Top requests

What users ask to add or change, most asked first. 334 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Author-weeks | Posts | Agents (author-weeks) |
|---|---|---|---|---|
| 1 | No silent model switching or downgrades | 40 | 44 | Cursor 18, OpenAI Codex 14, OpenCode 3, Claude Code 2, GitHub Copilot 2, Cline 1 |
| 2 | Automatic task-based model routing | 34 | 35 | OpenAI Codex 16, OpenCode 5, Cursor 4, Google Antigravity 3, Claude Code 3, Pi 2, Conductor 1 |
| 3 | Show which model actually served each request | 30 | 32 | OpenAI Codex 13, Cursor 10, Claude Code 2, Factory 2, Google Antigravity 1, Devin 1, OpenCode 1 |
| 4 | Choose which model subagents use | 20 | 20 | Cursor 10, Claude Code 5, Google Antigravity 3, OpenAI Codex 1, OpenCode 1 |
| 5 | Fallback model when quota or provider fails | 18 | 18 | OpenAI Codex 7, Cursor 3, Claude Code 2, Amp 1, Google Antigravity 1, Cline 1, Factory 1, Kiro 1, Zed 1 |
| 6 | Manual model selection instead of forced routing | 18 | 18 | Cursor 9, OpenAI Codex 4, Claude Code 2, GitHub Copilot 2, OpenCode 1 |
| 7 | Persistent user-chosen default model | 17 | 17 | Cursor 13, OpenAI Codex 2, Claude Code 1, OpenCode 1 |
| 8 | Route simple tasks to cheaper models | 15 | 15 | OpenAI Codex 6, Claude Code 5, Google Antigravity 1, Cursor 1, OpenCode 1, Pi 1 |
| 9 | Exclude specific models from auto routing | 14 | 16 | Cursor 10, GitHub Copilot 3, OpenCode 1 |
| 10 | Better auto mode routing quality | 11 | 11 | OpenAI Codex 5, Cursor 3, Google Antigravity 1, Claude Code 1, GitHub Copilot 1 |
| 11 | Notify or confirm before model substitution | 10 | 11 | OpenAI Codex 7, Cursor 2, Claude Code 1 |
| 12 | Model-agnostic harness across providers | 10 | 10 | OpenAI Codex 3, Pi 3, Amp 1, Cursor 1, Factory 1, OpenCode 1 |

### 1. No silent model switching or downgrades

- Cursor, 2026-09-26, @cursor_ai (X): “@cursor_ai could you please stop switching my sessions to grok4.7? i know you want to push your new model, but it's not what i want an super-annoying. #customerfirst” [source](https://twitter.com/2717214655/status/2103820104416858532)
- Cursor, 2026-09-26, @cursor_ai (X): “no longer going to update cursor @cursor_ai @spacexai since you guys want to keep turning models on, and ignoring my settings, with them off. <strict_link>” [source](https://twitter.com/1437891983117279233/status/2103664293258633261)
- Cursor, 2026-09-25, r/cursor (Reddit): “yeah, i know about this popup, but i never accept is. additionally i've uploaded video and attached to the thread where you can see one case where the model changes on itself. i hope we can move the discussion from "it's your mistake" to "cursor changes models without permission"” [source](https://www.reddit.com/r/cursor/comments/1wpcg0f/i_just_lost_200_usd_because_cursor_switched_my/pbzzuo1/)

### 2. Automatic task-based model routing

- OpenAI Codex, 2026-09-27, r/codex (Reddit): “great idea,it worked but in case where there around 40 prompts assigned to different cheaper models, is there way to automate rather than use switching from sol to luna etc” [source](https://www.reddit.com/r/codex/comments/1w9sonp/how_exactly_do_you_orchestrate_with_astra/pcbbebp/)
- Cursor, 2026-09-24, r/cursor (Reddit): “it gives very valid outputs and does the job just as asked. auto is so off the topic giving hard to decipher text. somehow auto is using a model not based on what's best for the developer requested ask but what's preferred by the cursor dev team and hard configured to some model.” [source](https://www.reddit.com/r/cursor/comments/1wp9gpx/composer_25_fast_is_still_really_good_as_always/)
- OpenCode, 2026-09-24, @opencode (X): “@aapakari @opencode has taken a lot of pressure off my frontier subscriptions. great for grunt work. still not got an automatic routing procedure but working on whether jev can provide some intelligence there.” [source](https://twitter.com/8457362/status/2102940182249185413)

### 3. Show which model actually served each request

- Factory, 2026-09-26, @droid (X): “@droid is there to tell which model auto model is using for a given task? also, mobile remote app please!” [source](https://twitter.com/2023937351815467008/status/2103962781032534497)
- Cursor, 2026-09-25, r/cursor (Reddit): “i figured even when set to auto it is probably running things through composer and grok (and maybe other models too) –on the unlimited plan it doesn't show me a breakdown so i don't see how to know. but it seemed weird to have that setting changed from auto for me. i wanted to let auto route to whichever model fits the bill, not to possibly pump up the numbers for grok 4.7 fast.” [source](https://www.reddit.com/r/cursor/comments/1wp9mhm/uh_is_grok_47_really_6x_more_expensive_and_8x/pc26hkc/)
- Google Antigravity, 2026-09-25, @antigravity (X): “@googledevs @antigravity @googlegemma for a hybrid workflow, 'local' should be observable at each handoff. i'd want a trace showing which model handled each step and which file snippets crossed to a cloud worker. one on-device worker doesn't tell the developer where the rest of the task ran.” [source](https://twitter.com/2036618196384636928/status/2103392704679887334)

### 4. Choose which model subagents use

- Google Antigravity, 2026-09-25, r/google_antigravity (Reddit): “had some opus to use so ran a prompt. it ran out as expected but my five hour gemini allowance also went to 0, lost 40% of my five hour allowance with one opus prompt in less than ten minutes. ridiculous. wish i hadn’t run it, was only as i thought i had some to burn. guessing the agents it spun up used gemini 3.8 and since that is terrible for usage killed it. crazy we get no control of that.” [source](https://www.reddit.com/r/google_antigravity/comments/1wpvi9v/opus_using_a_large_amount_of_my_gemini_allowance/)
- Google Antigravity, 2026-09-24, r/google_antigravity (Reddit): “and then subagents will surely always be flash model? can we control it? maybe like 3.7 or 3.8?” [source](https://www.reddit.com/r/google_antigravity/comments/1wo7mvi/a_simple_way_to_get_much_better_results_from/pbpr4m4/)
- Cursor, 2026-09-21, @cursor_ai (X): “@cursor_ai apparently grok 4.6 decided to use it as a subagent for some reason, which makes no sense unless the default model in @cursor_ai is sonnet, which would be stupid. <strict_link>” [source](https://twitter.com/918058226813456384/status/2102161291351642337)

### 5. Fallback model when quota or provider fails

- OpenAI Codex, 2026-09-27, r/codex (Reddit): “i solved scrolling by adding alternate_screen = "never" to the codex config, but the linux cli worse with every upgrade these days. you can't even change the model when it switches to luna reserve and the limit gets reset. you're stuck on luna with no other option. there's a workaround by running codex resume with the --model argument.” [source](https://www.reddit.com/r/codex/comments/1wrl6ch/latest_linux_codex_cli_seems_to_have_several_bugs/pcefa55/)
- OpenAI Codex, 2026-09-26, r/codex (Reddit): “exactly. portable doesn’t have to mean interchangeable. i still want model-aware routing. i just don’t want one model going down to kill the workflow” [source](https://www.reddit.com/r/codex/comments/1wqu991/codex_went_down_and_i_genuinely_didnt_notice/pc7n9q4/)
- Kiro, 2026-09-26, r/kiroIDE (Reddit): “there's probably a too conservative system prompt that was done by the aws/kiro team. but the real issue seems there's no auto fallback to opus 4.8 or opus 5 model when that happens?” [source](https://www.reddit.com/r/kiroIDE/comments/1wqbuwo/opus_55_in_kiro_keeps_killing_legitimate_sessions/pc5znve/)

### 6. Manual model selection instead of forced routing

- Cursor, 2026-09-26, r/cursor (Reddit): “i love cursor but it's going to lose because they are putting their finger on the scale for what models you can and cannot use (without significant inconvenince). planning to get off it by end of this year.” [source](https://www.reddit.com/r/cursor/comments/1wq3m71/im_out/pc372mv/)
- GitHub Copilot, 2026-09-24, r/GithubCopilot (Reddit): “gpt-6 luna and sol is all that most people need. that said, let the user choose as long as they are capped on usage” [source](https://www.reddit.com/r/GithubCopilot/comments/1wovttn/have_you_disabled_haiku/pbtdn9k/)
- OpenAI Codex, 2026-09-24, r/codex (Reddit): “we are paying, we pay a subscription to specifically use the models we choose, now if there was an option to auto route and we picked that, great, useful even, but we pay to choose the model we want as part of our subscription.” [source](https://www.reddit.com/r/codex/comments/1woekw4/astra_prompts_are_getting_silently_rerouted_to/pbpscpu/)

### 7. Persistent user-chosen default model

- Cursor, 2026-09-27, @cursor_ai (X): “so @cursor_ai keeps foisting grok 4.7 on me. but that model is super-annoying: it spend most of the time talking to itself even on the simplest, most direct prompts. trying to sell tokens?” [source](https://twitter.com/2717214655/status/2104161730137907209)
- Cursor, 2026-09-25, @cursor_ai (X): “eh @cursor_ai, seriously: can you stop resetting my model setting auto to your most expensive model after every update? i know you need the money, but get it somewhere else.” [source](https://twitter.com/1446295130647015444/status/2103553943196541234)
- Cursor, 2026-09-23, r/cursor (Reddit): “i have a subscription for a year ahead, but i can't find work to use the tokens. the automatic switching to grok, which can't solve basic tasks has destroyed all the benefits of cursor. it makes zero sense to spend time on cursor when elon decided to enshittify it” [source](https://www.reddit.com/r/cursor/comments/1wno0lk/ngl_it_is_so_over_for_cursor/pbly439/)

### 8. Route simple tasks to cheaper models

- Pi, 2026-09-22, @pidotdev (X): “@pidotdev tôi nghĩ việc bạn nên làm là làm ra 1 cái flow gì đó mà nó kiểm soát những llm rẻ nhất có thể làm đúng việc của nó, không ảo tưởng thay vì những đổi mới chưa thực sự cần thiết.” [source](https://twitter.com/1465319407589027844/status/2102406062012084561)
- Google Antigravity, 2026-09-19, @antigravity (X): “hey @antigravity, the desktop harness urgently needs an auto mode. offloading orchestration to sub-second, ultra-low-compute routing engines like jev would decouple deterministic task planning from heavy frontier weights, slashing token overhead at scale.” [source](https://twitter.com/1370489968594886657/status/2101213957726056881)
- OpenAI Codex, 2026-09-16, X search: OpenAI Codex, Codex CLI, Codex app (X): “@jamespardoe @voxyz_ai one thing i do is use luna as a cheap router, so it determines if a task needs astra or something cheaper. too often i find it too lazy to switch between models from codex cli.” [source](https://twitter.com/2097731580227903488/status/2100039235541868593)

### 9. Exclude specific models from auto routing

- Cursor, 2026-09-27, r/cursor (Reddit): “on the other models / which settings ask, auto is what chewed through mine when it routed into an expensive model at full list price, so i pin grok 4.6 now and on a bad stretch other models still jumped maybe \~40% in under an hour” [source](https://www.reddit.com/r/cursor/comments/1wrddz6/why_is_the_other_models_usage_is_being_used_so/pcdioo5/)
- Cursor, 2026-09-24, @cursor_ai (X): “until you guys fix it so auto mode uses only models selected by the user it's not usable for anything. @cursor_ai” [source](https://twitter.com/635313920/status/2103025597274308673)
- Cursor, 2026-09-23, @cursor_ai (X): “@cursor_ai when using auto, it could be better if there is a way we can set which models to pool and toggle fast to off.” [source](https://twitter.com/1953991414523842560/status/2102766385122488477)

### 10. Better auto mode routing quality

- OpenAI Codex, 2026-09-25, r/codex (Reddit): “we're talking about gpt 6 here. i'm honestly kinda surprised we don't have better automated model selection yet” [source](https://www.reddit.com/r/codex/comments/1worivz/unpopular_opinion_sol_6_xhigh_is_pretty_decent/pbwklex/)
- Cursor, 2026-09-24, @cursor_ai (X): “@cursor_ai how about fixing auto mode ? bad plans, ignored code, and unreliable changes - the old auto mode was amazing, this, this crap we have now is heading in the wrong direction <strict_link>” [source](https://twitter.com/1925646160359739392/status/2103035214452625463)
- Cursor, 2026-09-22, r/cursor (Reddit): “is auto actually more expensive or cheaper than using grok 4.6/4.7? i've used it in the past but always found it gave weaker responses. i had kinda hoped it would route to better models when needed but i feel it always goes the cheaper option but charges more.” [source](https://www.reddit.com/r/cursor/comments/1wn3rch/does_anyone_use_auto/)

### 11. Notify or confirm before model substitution

- OpenAI Codex, 2026-09-23, r/codex (Reddit): “it's transparent if you're monitoring web sockets and rust logs. i don't particularly have an issue with rerouting, but don't make me dig through verbose logs to see it. don't hide it. if someone tells me they are giving me ribeye steak for $10, but it turns out to be rump steak, i'm gonna be pissed. sure its only $10, but i might have got something from somewhere else if they'd been honest.” [source](https://www.reddit.com/r/codex/comments/1woekw4/astra_prompts_are_getting_silently_rerouted_to/pbmrpu9/)
- Cursor, 2026-09-22, r/cursor (Reddit): “i stopped using auto for the same reason. it feels like it routes cheap first, and that silent downgrade is worse than picking a weaker model on purpose. when i care about the pass i pin the model and effort myself so a busy provider can't slide me onto something shallow without saying so. if a pinned model is rejected i want an explicit nearest-equivalent pick, not a quiet fallback.” [source](https://www.reddit.com/r/cursor/comments/1wn3rch/does_anyone_use_auto/pbbuppu/)
- OpenAI Codex, 2026-09-19, r/codex (Reddit): “are they silently degrading models or just taking access to certain models away? i mean both are bad, but i’d much rather know that i can’t use sol rather than send instructions to sol and have it silently routed to a different model.” [source](https://www.reddit.com/r/codex/comments/1wkwdfl/the_end_of_the_codex_era_ive_completely_lost/pau0dzr/)

### 12. Model-agnostic harness across providers

- OpenCode, 2026-09-27, r/ClaudeCode (Reddit): “they follow the money. we need to be model agnostic (aka openrouter / opencode go) to prevent this” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr67vx/i_agree_with_dario_regulating_llms/pcai6ne/)
- OpenAI Codex, 2026-09-21, r/codex (Reddit): “being able to call deepseek from astra would be useful. currently i'm doing that manually.” [source](https://www.reddit.com/r/codex/comments/1wmf6li/i_wanted_codex_cli_to_choose_the_model_and/pb6d6gb/)
- Amp, 2026-09-18, @AmpCode (X): “@ampcode feature request for model routing to support gemini api key” [source](https://twitter.com/85549810/status/2100911633178435854)

## Every agent

| Agent | Overall rank | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|---|
| [Devin](https://feedbackbench.com/agents/devin.md) | =5 | Better than peers | 0.559 | 0.529–0.588 | 39 | 25 | 14 |
| [Claude Code](https://feedbackbench.com/agents/claude-code.md) | 1 | Typical | 0.526 | 0.491–0.559 | 294 | 79 | 215 |
| [Google Antigravity](https://feedbackbench.com/agents/antigravity.md) | =5 | Typical | 0.519 | 0.488–0.553 | 52 | 18 | 34 |
| [GitHub Copilot](https://feedbackbench.com/agents/copilot.md) | =8 | Typical | 0.508 | 0.473–0.543 | 56 | 14 | 42 |
| [OpenCode](https://feedbackbench.com/agents/opencode.md) | 3 | Typical | 0.503 | 0.466–0.537 | 114 | 30 | 84 |
| [Cursor](https://feedbackbench.com/agents/cursor.md) | 4 | Worse than peers | 0.467 | 0.433–0.499 | 343 | 73 | 270 |
| [OpenAI Codex](https://feedbackbench.com/agents/codex.md) | 2 | Worse than peers | 0.396 | 0.361–0.429 | 433 | 57 | 376 |
| [Pi](https://feedbackbench.com/agents/pi.md) | 7 | Too few posts | – | – | 29 | 16 | 13 |
| [Factory](https://feedbackbench.com/agents/factory.md) | =11 | Too few posts | – | – | 18 | 14 | 4 |
| [Amp](https://feedbackbench.com/agents/amp.md) | =11 | Too few posts | – | – | 16 | 11 | 5 |
| [Cline](https://feedbackbench.com/agents/cline.md) | =8 | Too few posts | – | – | 15 | 7 | 8 |
| [Conductor](https://feedbackbench.com/agents/conductor.md) | 14 | Too few posts | – | – | 9 | 4 | 5 |
| [Kiro](https://feedbackbench.com/agents/kiro.md) | 13 | Too few posts | – | – | 6 | 1 | 5 |
| [Zed](https://feedbackbench.com/agents/zed.md) | =8 | Too few posts | – | – | 2 | 1 | 1 |
| [Warp](https://feedbackbench.com/agents/warp.md) | 15 | Too few posts | – | – | 1 | 0 | 1 |
| [Grok Build](https://feedbackbench.com/agents/grok-build.md) | 16 | Too few posts | – | – | 0 | 0 | 0 |
| [Augment Code](https://feedbackbench.com/agents/augment.md) | 17 | Too few posts | – | – | 0 | 0 | 0 |

## Posts

Receipts rule: The 5 most recent praise and complaint posts per area (first 700 characters) and 3 per criterion (first 450 characters).

### Devin

- Praise, 2026-09-27, @DevinAI (X): “@jensenloke @devinai insane amount of tokens spent lol. i love the fusion router as well, extremely helpful especially when we can use swe-2 to help us as sub-agents!” [source](https://twitter.com/1492053555381174280/status/2104264634677317893)
- Praise, 2026-09-23, @cognition (X): “we just hired our fastest engineer yet. @devinai by @cognition joined the @wearejam7 engineering team. on first spin-up against our repo, it set up its own virtual laptop and shipped a pr to prod in under an hour. best time-to-first-pr-to-prod i have seen. strong on qa across amp too: accessibility, user flows, end-to-end. swe-2 for low-cost daily work. fusion when we need more, with handoff to cheaper models like swe or luna. rare case of a prod” [source](https://twitter.com/148911816/status/2102862928617550313)
- Praise, 2026-09-22, @cognition (X): “this is why i firmly believe that @cursor_ai needs to invest more heavily in an agent router. @cognition has fusion which gives us frontier intelligence at a fraction of the cost. @factoryai has factory router which routes to the most capable model per task, upgrading if a model gets stuck. we do not need to strictly use a single model, especially a frontier model, for all tasks. using cheaper models for many tasks can and will be sufficient if” [source](https://twitter.com/2070908287978246144/status/2102246929136804221)
- Complaint, 2026-09-23, @DevinAI (X): “@robinebers @devinai would also be nice for us to be able to select our models on the fusion like on the desktop” [source](https://twitter.com/17719163/status/2102725559054942554)
- Complaint, 2026-09-22, @cognition (X): “@dabit3 i've run my own model router for months. models drop weekly, benchmarks are scattered, half are stale on arrival. i'm open-sourcing it: all benchmarks + what people say on reddit/x, as an mcp that picks the best model per task. devin max would ship it faster @cognition” [source](https://twitter.com/68451547/status/2102515349296128231)
- Complaint, 2026-09-22, @DevinAI (X): “why @devinai under the swe-2 model 90% time using the swe-1.7 max. why @dabit3 <strict_link>” [source](https://twitter.com/1473214003027460098/status/2102445326762475805)

### Claude Code

- Praise, 2026-09-27, r/ClaudeCode (Reddit): “god, you kid's ever read claude code docu? you can configure most if not all of this behaviour. subagent model, max concurrent subagents spawned, subagent spawning subagents self depth” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrfxw6/slow_credit_burn/pcc579v/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “opus 5.5 definitely for the main brain, orchestration, architectural complexity. then i made a skill that launches codex exec when the task needs independent auditing or less heavy workers (luna models on max effort are not toys).” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrfxtu/new_to_claude_code_how_do_i_maximize_usage/pccbwqn/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “even with daybreak and the anthropic cybersec program, code analysis works for the most part, but sometimes both flag simple admin tasks. falling back to opus 4.8 and grok 4.7 will still work. on the bright side, we don't have to worry about umbrella's t-virus...yet” [source](https://www.reddit.com/r/ClaudeCode/comments/1wpqqoy/new_55_safe_guards_are_a_joke/pcetfdf/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “not sure what you’re referencing, and claude works in my auth later all the time. are you maybe thinking of safeguards downgrading the model? that does happen often enough when i’m working on auth, but that’s not an account ban at all- just a lower level model for a few minutes.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wodtfg/how_to_implement_web_app_authentication_with/pca9gi6/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “let’s put it this way, i’m building a cybersecurity/pentesting harness. opus5, 5.5, and fable, can’t so much as read the prd without tripping and downgrading to 4.8. i use hindsight as a memory system, they can’t read the description of the odin (name of my harness) bank without throwing a warning.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wppds2/is_it_safe_to_development_a_hacking_game_with/pcb4dzi/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “too bad you aren't filthy rich. api you can pin the model. just what i've heard as i'm not rich enough to use api.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr8cw1/claude_code_workflow_harness_change/pcbds4s/)

### Google Antigravity

- Praise, 2026-09-26, @antigravity (X): “@thsottiaux @antigravity btw the new luna and sol models make the limits work and fulfill your goal for us to use astra only when the workhorses fails. <strict_link>” [source](https://twitter.com/2874807497/status/2103981885973991828)
- Praise, 2026-09-26, r/AI_Agents (Reddit): “it runs on the subscriptions, not api, so price stays 'under control'.. the whole fleet fits on a handful of max plans, one trick we use is model per role, the boring seats don't need the big model. on gemini: the design is harness-agnostic on purpose, agents are stock cli sessions in terminals. a pi adapter already brought kimi and friends in, gemini cli is the obvious next one. not promising a date, but it's the direction.” [source](https://www.reddit.com/r/AI_Agents/comments/1wqij2i/my_friend_gave_claude_code_and_codex_agents_a_way/pc6t0xs/)
- Praise, 2026-09-25, r/google_antigravity (Reddit): “install the main antigravity 2.0 not ide , then in prompt add teamwork preview+ boost these 2 commands for opus 4.6 and ask it to delegate subagents to flash model only dont mention any model names it only sees flash lite, flash , pro and inherit , so ask it to invoke subagents to flash and work as orchestrator its actually so much better opus handels quite good and recently the reviews is so true unlike 3.8 which elevates my project as a high gr” [source](https://www.reddit.com/r/google_antigravity/comments/1wpgg6u/when_i_should_use_flash_and_pro/pbx1nhs/)
- Complaint, 2026-09-26, r/google_antigravity (Reddit): “i do not pretend, i literally uploaded a video. providing a different model under the same name to different subscribers constitutes a lie” [source](https://www.reddit.com/r/google_antigravity/comments/1wqkqcl/why_is_gemini_38_flash_so_slow/pc4tch7/)
- Complaint, 2026-09-26, @antigravity (X): “@antigravity what the hell do we do with this? you don't give us option to use other models. your models are not good enough. your selection of other provider models is still outdated. seriously, hilarious.” [source](https://twitter.com/1304738169619795969/status/2103718480683876749)
- Complaint, 2026-09-25, r/google_antigravity (Reddit): “had some opus to use so ran a prompt. it ran out as expected but my five hour gemini allowance also went to 0, lost 40% of my five hour allowance with one opus prompt in less than ten minutes. ridiculous. wish i hadn’t run it, was only as i thought i had some to burn. guessing the agents it spun up used gemini 3.8 and since that is terrible for usage killed it. crazy we get no control of that.” [source](https://www.reddit.com/r/google_antigravity/comments/1wpvi9v/opus_using_a_large_amount_of_my_gemini_allowance/)

### GitHub Copilot

- Praise, 2026-09-25, r/GithubCopilot (Reddit): “i choose the efficiency tier regardless of the task so the auto router is more likely to pick luna. lol i was talking of when auto did not have tiers, so a few months ago luna is the goat, we are loving it, it feels like sonnet level but it's essentially free :p use case is all sorts of agentic coding” [source](https://www.reddit.com/r/GithubCopilot/comments/1wp1ygm/tiers_of_auto_now_available/pbwzexe/)
- Praise, 2026-09-25, r/GithubCopilot (Reddit): “i have a bug analyzer subagent that was using opus 5.5 and only today i had three occurrences of subagent not returning anything to orchestrator so it was required spawn a new one. now i switched to sol and had no issues at all.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wptkkq/how_does_gpt6_sol_feel_so_far/pc0mfxl/)
- Praise, 2026-09-25, r/GithubCopilot (Reddit): “yay! this might actually get me to start using auto mode!” [source](https://www.reddit.com/r/GithubCopilot/comments/1wp1ygm/tiers_of_auto_now_available/pc0umng/)
- Complaint, 2026-09-26, r/GithubCopilot (Reddit): “i had a problem the other week where despite 5.6 luna being the only model enabled in settings, 5.6 luna was delegating tasks to sub-agents on expensive models such as sonnet 5 for basic tasks which burned credits. i'd check to make sure another model isn't being used somewhere you're not aware of.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wq39rv/did_they_10x_the_cost_of_luna_56_since_the/pc4mu1j/)
- Complaint, 2026-09-26, r/GithubCopilot (Reddit): “most likely it launched subagents on a different, more expensive model (luna 5.6 often called gemini on my system). newer vscode/luna fixed that problem. in the meantime, you can disable subagents.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wq39rv/did_they_10x_the_cost_of_luna_56_since_the/pc5o1hu/)
- Complaint, 2026-09-25, r/GithubCopilot (Reddit): “i like auto, i'm just tired of always thinking, is this the right model? am i overpaying? is this too complex for a cheaper model?” [source](https://www.reddit.com/r/GithubCopilot/comments/1wp1ygm/tiers_of_auto_now_available/pbvgrw5/)

### OpenCode

- Praise, 2026-09-27, r/ClaudeCode (Reddit): “they follow the money. we need to be model agnostic (aka openrouter / opencode go) to prevent this” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr67vx/i_agree_with_dario_regulating_llms/pcai6ne/)
- Praise, 2026-09-26, r/opencodeCLI (Reddit): “no no it picks the model and effort. it’s like going from manual to automatic transmission on a car” [source](https://www.reddit.com/r/opencodeCLI/comments/1wpyova/what_would_make_you_switch_your_default_opencode/pc30ufr/)
- Praise, 2026-09-26, @opencode (X): “@opencode this is why i run opencode daily. cheap flash for the boring calls, big model only where it counts - that split does most of the work for me.” [source](https://twitter.com/2015152856903655424/status/2103865154958233817)
- Complaint, 2026-09-27, r/opencode (Reddit): “i’ve noticed the same thing with glm 5.3 flash. going through openrouter the model works amazing but on opencode go it’s dumb as fuck and going in circles.” [source](https://www.reddit.com/r/opencode/comments/1wqxyl4/deepseek_v41_performance_is_much_worse_than/pcaf9ya/)
- Complaint, 2026-09-27, r/opencode (Reddit): “they serve quantized models though. definitely not full precision checkpoints.” [source](https://www.reddit.com/r/opencode/comments/1wrg81n/go_subscription_is_slower_deepseek_flash_41/pccf9af/)
- Complaint, 2026-09-27, r/opencode (Reddit): “yeah, i also noticed it. it doesn't always happen, but sometimes, randomly. i think there is either one provider that sucks especially bad or they route you through different providers every few turns [<strict_link>” [source](https://www.reddit.com/r/opencode/comments/1wrqxxa/is_cache_reading_faulty_with_deepseek_v41_on/pceuo4n/)

### Cursor

- Praise, 2026-09-27, r/cursor (Reddit): “this actually feels reasonable! i also feel like we don't need too much intelligence for most of the task and grok is a goof starting point” [source](https://www.reddit.com/r/cursor/comments/1wp6j44/is_this_supposed_to_be_good_news/pcej6hc/)
- Praise, 2026-09-26, r/cursor (Reddit): “you guys must be writing some pretty crazy code! i run cursor for 8 hours a day and never burn through all my credits, maybe it’s the plan i’m on? with that being said you have to be careful because it will default to grok which will in fact burn credits, switch to auto and it becomes a lot less expensive.” [source](https://www.reddit.com/r/cursor/comments/1wnpgmf/canceled_cursor_today_after_using_it_for_many/pc8jf43/)
- Praise, 2026-09-26, @cursor_ai (X): “the quiet win in @cursor_ai lately: projects that keep context without me babysitting model choice every session. less "which model for this?", more "here's the rails, ship it." #cursor #ai #agents #buildinpublic” [source](https://twitter.com/262960825/status/2103891276227924105)
- Complaint, 2026-09-27, r/cursor (Reddit): “i think in auto if it gets routed to an expensive model , it will cost you the full pricing of the model since they removed the lower/discounted pricing from it. previously auto only used composer or grok” [source](https://www.reddit.com/r/cursor/comments/1wrddz6/why_is_the_other_models_usage_is_being_used_so/pcbz9uf/)
- Complaint, 2026-09-27, r/cursor (Reddit): “on the other models / which settings ask, auto is what chewed through mine when it routed into an expensive model at full list price, so i pin grok 4.6 now and on a bad stretch other models still jumped maybe \~40% in under an hour” [source](https://www.reddit.com/r/cursor/comments/1wrddz6/why_is_the_other_models_usage_is_being_used_so/pcdioo5/)
- Complaint, 2026-09-27, r/cursor (Reddit): “if you want the cursor models usage to last, stick with composer and grok. don't use fast mode and you'll last the whole month on a $60 plan. they changed the auto to pick api models recently. maybe people jumped ship to other ides and they needed to up the api usage to keep the partnership going since cursor was acquired. who knows, either way it was the death of auto. i myself am looking to alternatives. i love the speed and ux of cursor but ma” [source](https://www.reddit.com/r/cursor/comments/1wrddz6/why_is_the_other_models_usage_is_being_used_so/pceihvh/)

### OpenAI Codex

- Praise, 2026-09-27, r/codex (Reddit): “your post makes no sense, the only way you switch models is through the ui, the model has no idea what the harness is doing, its a completely separate program. your other explanations dont make sense either, the agent does indeed decide which model it uses as a subagent (its how i have configured claude and codex to use my local lm studio server running qwen 3.8 27b as a subagent)” [source](https://www.reddit.com/r/codex/comments/1wrjfne/holly_shit_sol_kept_lying_to_me_telling_me_the/pcddb3z/)
- Praise, 2026-09-27, r/codex (Reddit): “it does, and it chooses them pretty accurately. identifying grunt work vs requiring creativity and depth. give it a try before throwing out random accusations.” [source](https://www.reddit.com/r/codex/comments/1wrok1q/amazing_agentsmd_instruction_for_best_usage/pce7zz1/)
- Praise, 2026-09-26, r/codex (Reddit): “the idea is you put in a prompt and it auto switches models as it goes, cheapest models for easy stuff, medium sol for harder stuff, if sol med fails it uses sol hard, saves tons of usage. id rather use 5 different models in a prompt if it saves usage and still gets everything done than use 1 model per prompt. i'm optimized for correctness and usage efficiency” [source](https://www.reddit.com/r/codex/comments/1wmgw7p/codex_usage_and_operation_discussion_last_updated/pc41mxj/)
- Complaint, 2026-09-27, r/codex (Reddit): “given this kerfuffle and the consideration of anthropic / oai sub flipping as one takes the lead, i checked out vscode agents window, and they have been doing stuff there. i am actually extending `pi` with `agent-host-protocol` which is basically an open protocol to make your own provider. claude, codex, deepseek, local, all managed in pi, where i've also done some routing rules. in vscode agents, rather than worry about model selections of luna,” [source](https://www.reddit.com/r/codex/comments/1wr9e2j/for_agents_is_a_bad_feature_and_codex_cli_is/pcatndw/)
- Complaint, 2026-09-27, r/codex (Reddit): “sadly, even 5.6 sol (via webui at least), is also a shit show right now. it's repeatedly claiming it can't use connected apps (google drive or the custom one i built it's been using for over a week now), it'll give crap code suggestions and more. i've even spotted the webui model to have been auto-switched from 5.6 sol to 5.6 luna (the actual model name appeared - which would explain it's inability to access the tools correctly). if it wasn't for” [source](https://www.reddit.com/r/codex/comments/1wrftcs/gpt6_sol_is_massive_downgrade/pcck3ki/)
- Complaint, 2026-09-27, r/codex (Reddit): “wdym right? it should have at least told me astra is not there? instead of lying. this was sota model just a month ago.” [source](https://www.reddit.com/r/codex/comments/1wrjfne/holly_shit_sol_kept_lying_to_me_telling_me_the/pcczqyo/)

### Pi

- Praise, 2026-09-24, r/PiCodingAgent (Reddit): “when jev came out, one of my first thoughts was: could something this fast and cheap pick which model should handle an agent’s next turn? so i built plugin for pi that uses jev to choose the model and thinking effort. the idea is to send routine work to cheaper models and reserve the expensive ones for harder tasks. in pi, the model mappings are configurable, and you can pin a model when you disagree with the router. i’ve been using them for a” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wovt67/i_built_a_jevbased_model_router_for_pi_for/)
- Praise, 2026-09-22, r/PiCodingAgent (Reddit): “the way it handle context is great for a non dev like me. and the coding agent + advisor mode is a dream team that get things done. you can refine that by allocating different models per task (scout, task etc). pretty cool. mind sharing what settings are the best for you? i spent a lot of money on api calls until today at my end. managed to finally get qwen 3.8 flash next to work at ~60+ tps locally. i add the advisor on top when needed.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wlpj30/lord_forgive_me_for_the_time_i_wasted/pbauigw/)
- Praise, 2026-09-22, @pidotdev (X): “@pidotdev pi has allowed me to create <strict_link> i'm now adding a system one pre qualifier i'm just fine tuning the model now on my aiserver. pi is amazing i've added agent deliberation outside there context window and llm router so the fronter model doesn't blow tokens” [source](https://twitter.com/1636019771761377282/status/2102390232725229690)
- Complaint, 2026-09-24, r/PiCodingAgent (Reddit): “you are not the first to do this. you don't want to route automatically on each user message because it destroys prompt caching. i won't use anything that does that. instead, automatically pick a model only once after the first user message, then in the bottom status bar display which model would be better for this conversation and keep updating after each user message. let the user manually trigger the model change. times when it's okay to autom” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wovt67/i_built_a_jevbased_model_router_for_pi_for/pbr9e69/)
- Complaint, 2026-09-22, r/PiCodingAgent (Reddit): “a self-written one that forces pi to always use \`deepseek-flash\` and never use \`deepseek-v4-pro\`, because something so simple apparently needs an extension. luckily it takes 2 minutes to code one.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wml882/what_extensions_do_you_think_are_essential_and/pbg4d1w/)
- Complaint, 2026-09-22, @pidotdev (X): “@pidotdev subagents should be the standard. build a smart router. build better visualisation.” [source](https://twitter.com/2231129593/status/2102489270120521990)

### Factory

- Praise, 2026-09-26, @FactoryAI (X): “@factoryai what is this witchcraft? auto model?” [source](https://twitter.com/2023937351815467008/status/2103681115991216612)
- Praise, 2026-09-26, @FactoryAI (X): “@delaanthonio @factoryai while i've been juggling different models 👀 a model-agnostic setup with real sovereignty is what i've needed and i keep returning to it” [source](https://twitter.com/356609569/status/2103698082735243324)
- Praise, 2026-09-26, @FactoryAI (X): “@firstmarkcap @enoreyes @factoryai model agnosticism is a smart move.” [source](https://twitter.com/1846161889719623680/status/2103824564736430242)
- Complaint, 2026-09-26, @FactoryAI (X): “the transit station really has no work to do. i have always been curious about the router that @factoryai has been promoting. they have never clearly explained the specific router mechanism, just vaguely mentioning the classification of simple and complex tasks. how can we determine whether a task is complex or not before it starts, and the inevitable cache invalidation caused by switching models midway is a big problem (anyway, the transit stati” [source](https://twitter.com/1724973133097291776/status/2103781371772809499)
- Complaint, 2026-09-26, @droid (X): “@droid is there to tell which model auto model is using for a given task? also, mobile remote app please!” [source](https://twitter.com/2023937351815467008/status/2103962781032534497)
- Complaint, 2026-09-25, @FactoryAI (X): “@factoryai what about efficiency? if i'm sure the best code is going to be made with opus 5.5 are you sure the router is going to give me what i want with a great code?” [source](https://twitter.com/1738636938616115200/status/2103335310427828449)

### Amp

- Praise, 2026-09-23, @AmpCode (X): “model routing is one of my favorite @ampcode features. i just hope that anthropic and google come into their senses and allow their respective models to be used over oauth.” [source](https://twitter.com/2090734054928687104/status/2102696668944912715)
- Praise, 2026-09-23, @AmpCode (X): “everyday i wake up, a new model is out. thank god that @ampcode takes care of model selection so that i have to go around x/reddit/hackernews just for fun and not to actually figure out what to use” [source](https://twitter.com/1543686991/status/2102717060153544728)
- Praise, 2026-09-23, @AmpCode (X): “you're right. the local ledger is what makes captain code actually learn, not just route once. every turn is recorded on your machine, so / frontier and / quality can balance across legs based on what actually ran. you can audit every decision, and no routing history ever leaves your box. that's the part most tools skip.” [source](https://twitter.com/118804749/status/2102750418589614360)
- Complaint, 2026-09-26, @AmpCode (X): “@benvargas @ampcode solid stack. that routing setup is doing a lot of work though. one server-side flag change and half your agents reroute without asking.” [source](https://twitter.com/1656373234647048192/status/2103698212087353406)
- Complaint, 2026-09-23, @AmpCode (X): “@iannuttall @ampcode i'm not sure why amp insists on self-adaptation. currently, the custom service is also a half-finished feature, and there is no way to configure it in the model dial.” [source](https://twitter.com/1725760355648024577/status/2102671129706209607)
- Complaint, 2026-09-23, @AmpCode (X): “@jkudish @ampcode kinda gives up the whole amp shtick of being opinionated and “choosing the best” for you” [source](https://twitter.com/2008296795181359104/status/2102857171104899566)

### Cline

- Praise, 2026-09-22, @cline (X): “kimi k3 is still holding strong as my daily driver, but i've started to find some value in using fable in plan mode. k3 can get into weeds w/ more complex features/debugs that fable cuts right through. kind of cool to move between them thanks to @cline.” [source](https://twitter.com/2059304303513202688/status/2102438617603854442)
- Praise, 2026-09-21, r/codex (Reddit): “learn your coding agent, what model do you use on codex? change the model to your gpt fav and go. i use cline not codex and i switch models as i see fit by a drop down box. you dont need to exit codex to use gpt models. depending on ur project you should have more than a 1 shot hand off. or you'll use alot of tokens and get sub par work. you need to remove ambiguity to it can act not spend its time thinking of the processes. the more you reduce t” [source](https://www.reddit.com/r/codex/comments/1wlzbr8/how_do_you_use_chatgpt_and_codex_together/pb3sz6v/)
- Praise, 2026-09-14, @cline (X): “@datachaz @cline open weights and local support means no lab can silently reroute my requests to a weaker model mid-task. the bar is on the floor and yet here we are.” [source](https://twitter.com/1657017278620131333/status/2099544871971279017)
- Complaint, 2026-09-27, r/CLine (Reddit): “please include whether it’s zdr or not. what’s with the limit because you are routing the request to vercel ai free pinary” [source](https://www.reddit.com/r/CLine/comments/1wqkucr/pixel_canary_new_stealth_model_is_now_free_in/pcce0up/)
- Complaint, 2026-09-25, @cline (X): “@cline gemini ain't working , maybe you shouldn't have replaced it with kimi k3” [source](https://twitter.com/782910885773836288/status/2103377933972672795)
- Complaint, 2026-09-23, @cline (X): “hey @cline... why when i am having model set to mimo 2.6 pro as current it shows something else as current when i type /model? <strict_link>” [source](https://twitter.com/1556568530883153920/status/2102811194952482917)

### Conductor

- Praise, 2026-09-23, @conductor_build (X): “@willcb you should try @conductor_build with @thesageox conductor: to switch between models and harnesses sageox: to never lose context across harness.” [source](https://twitter.com/103273439/status/2102827689749233895)
- Praise, 2026-09-22, @conductor_build (X): “@claudeai @thesageox so beautiful easily switch model using @conductor_build and prime your session using @thesageox . <strict_link>” [source](https://twitter.com/103273439/status/2102204984884691235)
- Praise, 2026-09-11, @conductor_build (X): “@conductor_build / @herdrdev + @thesageox why this combo? sageox - keeps all the team context wherever you are using your agents conductor - can easily use multiple models and easily switch herdr - well i love terminals and it's quite beautiful. the left pane to see agents is really good.” [source](https://twitter.com/103273439/status/2098470345414242413)
- Complaint, 2026-09-17, @conductor_build (X): “@conductor_build hey team, love the app but finding the model selector really annoying. i can't start a new thread with any except my top 2 pinned models. why?! a "more" option would be really great everywhere i select models 🙏 <strict_link>” [source](https://twitter.com/102718167/status/2100407910753018233)
- Complaint, 2026-09-16, @conductor_build (X): “i absolutely hate the new(ish) model picker in @conductor_build. it has slowed me down so much. really frustrating.” [source](https://twitter.com/1668000040869154816/status/2100302995070095539)
- Complaint, 2026-09-15, @conductor_build (X): “hey @conductor_build i would love for the option to be able to switch the selected model provider in the same chat, for example switching from fable to astra without needing to start a new chat” [source](https://twitter.com/1513865549629075467/status/2099908647794905243)

### Kiro

- Praise, 2026-09-11, r/kiroIDE (Reddit): “i use kiro ide with auto model. works very well, few tokens.” [source](https://www.reddit.com/r/kiroIDE/comments/1wchsii/frontier_models_on_other_providers_vs_the_same/p92st2r/)
- Complaint, 2026-09-27, r/kiroIDE (Reddit): “i work for amazon and is somewhat “strongly recommended” to use kiro. i still use claude code at work and at home. so much better … (auto classifier, transcript details, integrated tooling, open source tooling, cli features, general stability, model fallback, sub agents control, etc …)” [source](https://www.reddit.com/r/kiroIDE/comments/1wqueq0/claude_pro_vs_kiro_pro_subscription_which_is/pccqcjf/)
- Complaint, 2026-09-26, r/kiroIDE (Reddit): “there's probably a too conservative system prompt that was done by the aws/kiro team. but the real issue seems there's no auto fallback to opus 4.8 or opus 5 model when that happens?” [source](https://www.reddit.com/r/kiroIDE/comments/1wqbuwo/opus_55_in_kiro_keeps_killing_legitimate_sessions/pc5znve/)
- Complaint, 2026-09-26, r/kiroIDE (Reddit): “as someone who really only had a kiro pro sub because it was free for me as a student i defaulted to my other subs because of poor harness/model selection but i've been using opus 5.5 for the last couple hours in kiro and i'm pretty happy with the experience so far. anyone else having a similar experience?” [source](https://www.reddit.com/r/kiroIDE/comments/1wqed0s/pretty_pleased/)

### Zed

- Praise, 2026-09-21, r/PiCodingAgent (Reddit): “my personal huge level up was going from cursor ide to zed + omp. it is more efficient, i have everything i could've asked for and more. i love the custom fallbacks. being able to force the use of subagents. and the advisor... the advisor is really something tbh” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wlpj30/lord_forgive_me_for_the_time_i_wasted/pb4z6kj/)
- Complaint, 2026-09-23, r/cursor (Reddit): “yeah. i tried that. zed is so snappy and fast, but cursor has done some additional work with the harness to work some magic with the different models. zed is kinda like choose any model at your own risk.” [source](https://www.reddit.com/r/cursor/comments/1woiiyu/cursor_is_not_so_bad_when_compared_to_your_next/pbnj9uq/)

### Warp

- Complaint, 2026-09-15, r/warpdotdev (Reddit): “i chose gpt 5.6 luna xhigh as my model, when it failed, fallback model opus 5 max continued, which rapidly consumes more credits. i can't stand it.” [source](https://www.reddit.com/r/warpdotdev/comments/1oa0abo/warp_dirty_tactics_sonnet_45_thinking_uses_cheap/pa1hdbn/)
