# Reasoning effort setting and its defaults (`models.effort_control`)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/criterion/models.effort_control

Area: [Choosing models](https://feedbackbench.com/criteria/models.md)

**Definition.** How the reasoning or effort level can be set, whether it stays set, and how outcomes differ by level, such as overthinking at high effort or failing at low effort.

**Boundary.** Not this: see [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md) when the point is the quota share an effort level consumed. Not this: see [Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md) for over-engineered output.

Rated author-weeks, all agents: 1033. Complaint share: 54%.

## The brief

Written by Claude Opus 5.5 from 67 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**More effort rarely means better work; users tune it per task.**

TL;DR:

- Top effort settings overthink, loop and bloat context; many users drop to medium and get cleaner results.
- Low effort is not free either; users report missed details and weak architect work after compaction.
- OpenCode draws the sharpest complaints, from forced max reasoning to no way to turn thinking off.

In plain terms: Picking an effort level is a manual tuning job. Max often stalls or overthinks, low sometimes misses basics, and settings can reset, stay greyed out, or never reach subagents and custom models. Users who match effort to task report better results.

### How it breaks

- **Max effort overthinks and stalls** ([Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md)). The most common complaint is that the highest settings make agents slower and worse, not smarter.
  Users describe hours lost to overthinking, context that fills up before an answer lands, and bloated output. One Codex user suspects anything above xhigh fails more as context grows because compaction kicks in earlier. Cline and Copilot users say the same thing. Many report that dialing back one notch fixes the behaviour. The pattern holds across vendors, so the cause looks like the setting itself rather than any single harness.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-14: “high wrecks it. i run medium and it seems to take way longer now but works” [source](https://www.reddit.com/r/codex/comments/1wfsy6h/holy_the_nerf_is_insane/p9qxjly/)
  - Complaint, Cline, r/CLine, 2026-08-31: “i agree. i tried using glm with xhigh reasoning today, it started overassuming a bunch of stuff and considering hallucinative stuff in its reasoning. i had to turn down the reasoning to high to make it perform as it should.” [source](https://www.reddit.com/r/CLine/comments/1w37lzq/why_the_free_glm_53_on_cline_falls_short_with/p709e4s/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-24: “i suspect the same thing but for max effort. in benchmarks, max might be scoring higher, but as context grows, anything above xhigh seemed to fail more often, probably due to compaction triggering earlier and inaccurately summarizing things.” [source](https://www.reddit.com/r/codex/comments/1wp64gh/56_sol_xhigh_is_the_only_usable_model/pbso212/)
  - Complaint, GitHub Copilot, r/GithubCopilot, 2026-09-25: “i agree. i find the praise for luna baffling. it may be cheap, but most people seem to recommend running it at xhigh or max, where it takes forever to reason and generates bloated output.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wp1ygm/tiers_of_auto_now_available/pbwf62p/)

- **Low effort skips the basics** ([Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md)). Dropping effort to save time can backfire, with users reporting overlooked details and glaring errors in planning roles.
  Some users run high exclusively because medium starts missing basic details. A Claude Code user saw silly errors in an architect role at medium after compaction, and fine results only on xhigh. Antigravity users report a newer model thinking too much even on low and looping. Low works for planned, scoped work. It struggles when the model still has to decide what to do.
  Evidence:
  - Praise, OpenAI Codex, r/codex, 2026-09-22: “pretty sure "how long the model will try" is pretty significant to the output you get. even just using regular chatgpt, i've been running that on high exclusively for years. sometimes i drop it to medium, and the model immediately starts overlooking basic details.” [source](https://www.reddit.com/r/codex/comments/1wnh014/absolutely_no_fking_way_the_pricing_is_wowww/pbg6uxa/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-02: “anecdotally, i've been surprised at how poorly it performs an architect role post-compaction at lower effort levels (eg medium). some really glaring silly errors (eg confusing context used vs context remaining). part of it may be that it is not following my [claude.md](http://claude.md) post-compact instructions -- at least that has been its excuse. but so far it has done well for me on xhigh, but very expensive, and rather abysmally in an architect role on medium. on high it seemed ok, but no better than 5.0.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w5fzyu/so_fable_51_yay_or_nay/p7fprtr/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-09: “had no issue with 3.7, most if time i stay on 3.7. \- overall \- less looping dead loop, 3.8 think too much on low \- no bias with massing curl calls” [source](https://www.reddit.com/r/google_antigravity/comments/1w62rr4/gemini_38_flash_goes_to_cycle_way_too_often/p8qs5yl/)
  - Complaint, Devin, @cognition, 2026-09-07: “@cognition low and medium astra is indistinguishable to sol. checks out after days of use....” [source](https://twitter.com/441691870/status/2096997531075154386)

- **Settings reset or stay locked** ([Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md)). Several harnesses make effort hard to change or keep, forcing users to re-set it each chat or abandon mid-conversation changes.
  Pi users re-select max thinking on every new chat. Amp users find the level greyed out mid-conversation, with a tooltip that vanishes. Copilot users on Visual Studio say the reasoning picker disappeared from the chat sidebar and moved into a models dialog. Persistent effort across sessions is a standing request, led by Claude Code users. The control exists in these tools, but users struggle to keep or reach it.
  Evidence:
  - Complaint, Pi, @pidotdev, 2026-08-31: “@pidotdev finally, every time i started a new chat i have to go from no thinking to max is a pain” [source](https://twitter.com/2512335767/status/2094391917299523865)
  - Complaint, Amp, @AmpCode, 2026-09-14: “@sqs @ampcode another thing that's confusing to me is i can't change the model level (i.e. "medium") mid-conversation. it's just greyed out. and clicking on the little gauge just pops up a tooltip saying "medium" that quickly disappears which is confusing” [source](https://twitter.com/721234540/status/2099487663384334421)
  - Complaint, GitHub Copilot, r/GithubCopilot, 2026-09-08: “we have users on the may and july visual studio updates are reporting today that they don't have a reasoning/thinking selection anymore in the visual studio chat sidebar. after updating to the august update, we see options to control that in the "manage models" dialog, but it feels odd to be removing this from older version and making it harder to control even on the new update. especially when models like luna that can be great cost-optimizations depend a _lot_ on this reasoning level, making this less-visible feels worse.” [source](https://www.reddit.com/r/GithubCopilot/comments/1waqq4w/visual_studio_thinkingreasoning_level_seems_to/)
  - Complaint, Amp, @AmpCode, 2026-09-14: “ok i have this @ampcode thing open... how do i change from medium to something else? clicking “medium” does nothing. <strict_link>” [source](https://twitter.com/63583842/status/2099516460666085650)

- **Subagents and custom models ignore effort** ([Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md)). Effort control often stops at the main agent, leaving subagents and third-party models without a dependable dial.
  Antigravity users want different thinking levels for orchestrator and builder agents and cannot find the parameter. A Claude Code user says subagent effort set in frontmatter is not deterministic the way model choice is. Factory, Zed and Conductor users cannot pick a thinking level for custom or OpenCode models at all. Requests to respect effort for subagents and to support third-party models recur across vendors.
  Evidence:
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-22: “i don't know if we can configure specifically thinking level to agent definition files. i want to use flash high for orchestrator main agent and flash medium for builder subagent. i cannot find that parameter in the documentation. it's nice to have that.” [source](https://www.reddit.com/r/google_antigravity/comments/1wm7g5r/weekly_quotas_known_issues_support_september_21/pbbdikj/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-21: “unless my info is outdated, it’s not documented to work. it’s not in their docs and there is a long open issue about adding it. it’s hallucinating based on the model being settable. [<strict_link> you can create a custom agent. and the sub agent seems to be able to set its own effort based on what you tell it to do in the frontmatter, but it’s not deterministic like model is.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wknv74/so_fable_is_pretty_much_off_the_table_for/pb2u366/)
  - Complaint, Factory, @droid, 2026-09-19: “@tereza_tizkova @droid i have no way to choose the thinking level when accessing the custom model. is this a problem? but i directly asked the model to modify the settings by itself. hahaha” [source](https://twitter.com/1889310672368095232/status/2101460620063154561)
  - Complaint, Conductor, @conductor_build, 2026-09-21: “hey @conductor_build please allow me to select reasoning levels for @opencode models 🙏🏽” [source](https://twitter.com/50570112/status/2101967617686970808)

- **Matching effort to task pays off** ([Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md)). Users who route high effort to planning and low effort to execution report faster, cheaper and cleaner runs.
  The praise has one consistent shape. Use high for planning and low for work that is already planned. Medium is the default until the task is clear. One Codex user found low fine once real decisions moved into plain code. A Claude Code tester found low effort cheaper and faster with no misses on their tasks. Automatic effort selection by task difficulty is a top request, because users currently do this routing by hand.
  Evidence:
  - Praise, Google Antigravity, r/google_antigravity, 2026-09-19: “high for planning. low for implementing work that is already planned.” [source](https://www.reddit.com/r/google_antigravity/comments/1whnh0m/which_antigravity_model_do_you_actually_use_for/pattfcs/)
  - Praise, Cursor, r/cursor, 2026-08-31: “each of the 4 has a place. medium should be default until you figure out your task. low is great when you want it to do exactly what is told. high and xhigh can do some great planning.” [source](https://www.reddit.com/r/cursor/comments/1w36xbw/setting_grok_46_to_extra_high/p7028e9/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-12: “dynamic workflows with model and effort scaled to task difficulty. seems to work very well. haiku is worthless except for the simplest of tasks though.” [source](https://www.reddit.com/r/ClaudeCode/comments/1weqgnj/do_you_manually_switch_between_models_or_do_you/p9fxyqe/)
  - Praise, Claude Code, r/ClaudeAI, 2026-09-24: “fair question... and yeah, for raw capability the public benchmarks are better designed than my 25 tasks. i was after something different: cost per passed task, time, and low vs high effort on the same model, all run through claude code, since that's where most of us actually use it. i didn't find a benchmark that reports that side by side. the opus 5.5 low vs high effort result (\~20% cheaper, \~30% faster, no misses on my tasks) is the kind of thing i was after. small samples and my own tasks, so it's one more data point, not the last word. if you know a benchmark that covers effort settings plus cost per pass, i'd genuinely like to see it.” [source](https://www.reddit.com/r/ClaudeAI/comments/1woicc1/opus_55_is_the_real_deal_same_accuracy_as_fable/pbr7ta1/)

### Who stands out

- **OpenCode (weaker)**. OpenCode posts complain that effort controls fail to map cleanly to what models do, in both directions.
  Users report that setting low, medium or high now silently runs max reasoning. Others want thinking fully disabled and find only default, low, high and max. Some models are criticised for overthinking and exploding context, and one lacks a visible thinking trace for catching drift. A max effort option is the top request here. Praise exists for fast models at xhigh and for changing effort after typing a prompt.
  Evidence:
  - Complaint, OpenCode, r/opencodeCLI, 2026-09-10: “they got rid of the reasoning effort modes. if you set it to low medium or high like we did in the past, it's automatically set to max reasoning (100).” [source](https://www.reddit.com/r/opencodeCLI/comments/1wcjpzj/controversial_opinion_deepseek_41_flash_is/p8yfjpb/)
  - Complaint, OpenCode, r/opencode, 2026-09-03: “the only options are default, low, high and max. i need it completely gone.” [source](https://www.reddit.com/r/opencode/comments/1w62dqj/how_to_turn_off_deepseek_model_thinking_using_the/p7jlmi5/)
  - Complaint, OpenCode, r/opencodeCLI, 2026-09-13: “same experience: usually for the same task, ds4.1flash needs 5-10x more time than glm-5.3-flash to complete because it always thinks too much. for me, this is not acceptable even if i "assume" the cost to complete the same task is 1/2 is true. my experience is similar to [artificialanalysis.ai](<strict_link>)” [source](https://www.reddit.com/r/opencodeCLI/comments/1wcjpzj/controversial_opinion_deepseek_41_flash_is/p9hnxx6/)
  - Praise, OpenCode, @opencode, 2026-09-18: “does no one using grok ever want to increase/decrease effort after having typed a paragraph? it's all these little things that make @opencode so much more pleasant to use.” [source](https://twitter.com/15773993/status/2101039803001421851)

- **Claude Code (mixed)**. Claude Code users increasingly prefer low and medium, but blame confusing labels and patchy subagent support.
  Posts say extra effort is no guarantee and that lower settings get the agent to the point instead of circling. One user argues the effort labels push people to max because nobody wants low effort work, when the setting actually controls breadth of research. Persistent effort across sessions and respecting effort in subagents are the recurring asks. Medium on the flagship model draws steady praise.
  Evidence:
  - Praise, Claude Code, r/ClaudeCode, 2026-09-21: “extra is no guarantee for a better result imho. i’ve begun using low/medium a lot more, and feel the agent gets more “to the point” instead of going around in circles.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wly1hr/fable_usage_reduced_and_now_im_not_running_out_of/pb3q074/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-20: “imo anthropic's recent blog explains a lot of the problem. claude's "effort" levels are really poorly labeled, nobody wants "low effort" work but what that actually does is decide the breadth of claude's work and how much research and consideration it does. imo this poor ux means people are cranking the model to the highest setting because more=better and undermining themselves” [source](https://www.reddit.com/r/ClaudeCode/comments/1wiujum/claude_code_is_falling_behind_codex_not_because/payjihd/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-23: “medium or high. max seems pointless according to their graphs.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wnt4d3/they_fucking_cooked_yo_opus_55_is_a_massive/pbhvz1e/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-21: “unless my info is outdated, it’s not documented to work. it’s not in their docs and there is a long open issue about adding it. it’s hallucinating based on the model being settable. [<strict_link> you can create a custom agent. and the sub agent seems to be able to set its own effort based on what you tell it to do in the frontmatter, but it’s not deterministic like model is.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wknv74/so_fable_is_pretty_much_off_the_table_for/pb2u366/)

- **OpenAI Codex (mixed)**. Codex carries the most discussion, split between users who swear by xhigh and those who say high wrecks results.
  Codex users run detailed playbooks, with different models at different levels for audits, mid tasks and planning. Others report overthinking, slowness and more failures above xhigh as context grows. The loudest asks are automatic effort selection by difficulty and clearer guidance on tradeoffs. That points to a capable dial that users must still calibrate by trial and error.
  Evidence:
  - Praise, OpenAI Codex, r/codex, 2026-09-07: “the effort level you need is mostly a function of how much you've left the model to decide. once i moved the actual decisions out into plain code and left the model only the write up, low was fine for basically everything. high only earns its keep when the model is still the thing choosing what to do next.” [source](https://www.reddit.com/r/codex/comments/1w9erx3/usage_tip_gpt6_astra_on_low_performs_better_than/p8d4azh/)
  - Praise, OpenAI Codex, r/codex, 2026-09-18: “oh for tasks like do a basic code audit, go find a bunch of data, change this color or that text. anything else goes to sol medium for me. but i was using terra for "in the middle tasks" but now i'll just use terra xhigh for everything up to the point of planning or producing code i depend on.” [source](https://www.reddit.com/r/codex/comments/1wjvcu1/terra_seems_useless_now/palsi54/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-16: “what effort level are you on? i do not have this problem with my agents.md setup but i really don't go above medium unless i'm about to go to bed with somewhat high limits and will let it crunch on evaluating the state of the project or evaluating various classes for improvement. higher effort levels will have it consider a lot more factors and in a lot of cases make a lot more work out of a problem than need be. e: i usually go really hard into reusable base classes i inherit off of and i think it does pretty good job building off those without going off the rails so maybe it helps.” [source](https://www.reddit.com/r/codex/comments/1whizjj/how_do_you_stop_artra_from_scope_creeping_and/pa3dt2t/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-01: “like it is so slow, and after hours, when i realize it overthinking again, it took more hours to correct. bruh” [source](https://www.reddit.com/r/codex/comments/1w3rk5c/anyone_noticing_slow_speeds/p7538ow/)

### Fine print

- Most agents here have too few posts to judge; only Codex, Claude Code, OpenCode and Antigravity have meaningful volume.
- Many complaints name a specific model, so some effort behaviour may reflect the model rather than the harness.
- Quota consumption by effort level is tracked under limits.burn_rate, not here.

## Top requests

What users ask to add or change, most asked first. 173 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Author-weeks | Posts | Agents (author-weeks) |
|---|---|---|---|---|
| 1 | Max reasoning effort level option | 20 | 22 | OpenCode 16, OpenAI Codex 3, Claude Code 1 |
| 2 | Automatic effort selection by task difficulty | 17 | 17 | OpenAI Codex 12, Claude Code 3, Amp 1, Cursor 1 |
| 3 | Reduce overthinking and over-engineering at high effort | 9 | 10 | OpenAI Codex 4, Google Antigravity 2, Claude Code 2, Devin 1 |
| 4 | Clearer guidance on effort level tradeoffs | 9 | 9 | OpenAI Codex 6, Google Antigravity 1, Claude Code 1, Devin 1 |
| 5 | Persistent effort setting across sessions | 9 | 9 | Claude Code 6, Cline 2, Google Antigravity 1 |
| 6 | Reasoning effort control for custom and third-party models | 9 | 9 | OpenAI Codex 3, Cursor 2, OpenCode 2, Conductor 1, Factory 1 |
| 7 | Respect effort settings for subagents | 8 | 9 | Claude Code 5, Google Antigravity 2, Cursor 1 |
| 8 | Per-task model and effort configuration | 8 | 8 | OpenAI Codex 4, Claude Code 3, Amp 1 |
| 9 | Better default reasoning effort | 7 | 7 | Amp 1, Google Antigravity 1, Claude Code 1, OpenAI Codex 1, Cursor 1, Factory 1, Grok Build 1 |
| 10 | Option to disable thinking entirely | 7 | 7 | Claude Code 4, Google Antigravity 1, OpenAI Codex 1, OpenCode 1 |
| 11 | Deep thinking reasoning mode | 6 | 7 | Google Antigravity 2, OpenAI Codex 2, Claude Code 1, Cline 1 |
| 12 | Effort level slider | 6 | 6 | OpenAI Codex 4, GitHub Copilot 1, OpenCode 1 |

### 1. Max reasoning effort level option

- OpenCode, 2026-09-24, r/opencodeCLI (Reddit): “the max reasoning level for muse spark 1.3 in the standard opencode zen package got removed recently. i am not able anymore to select max as reasoning level for muse spark 1.3. it was there before and i used it but now it does not exist anymore in opencode. without the max reasoning its not that good for low level engineering computer programming and tasks in the programming languages c, c++, assembly, verilog.” [source](https://www.reddit.com/r/opencodeCLI/comments/1wopbtv/what_muse_spark_14_contributor_is_already_here/pbpmdr6/)
- Claude Code, 2026-09-23, @ClaudeDevs (X): “@claudedevs plz provide a speciall keyword like "ultrathink" for this” [source](https://twitter.com/1809147175391399937/status/2102744494487785955)
- OpenAI Codex, 2026-09-22, X search: OpenAI Codex, Codex CLI, Codex app (X): “@thsottiaux why can't i select max reasoning efforts when i use gpt6 luna in the codex app?” [source](https://twitter.com/1911327506529222656/status/2102541597175062604)

### 2. Automatic effort selection by task difficulty

- Claude Code, 2026-09-25, @ClaudeDevs (X): “@trq212 @claudedevs great content, based on this could we have an ‘auto-effort’ mode where based on the context it decides the effort?” [source](https://twitter.com/2200731333/status/2103591103480160661)
- OpenAI Codex, 2026-09-21, r/codex (Reddit): “the whole point of agi is to be better at all tasks than humans. not only hard tasks. a frontier model shouldn’t shit the bed at easy things. it should adapt to the task difficulty. like a good senior dev.” [source](https://www.reddit.com/r/codex/comments/1wm2c4g/agi_is_here_astra_cant_stop_creating_random_md/pb8d9aa/)
- Cursor, 2026-09-21, @cursor_ai (X): “@kvncnls @cursor_ai auto mode dynamic reasoning would go crazy” [source](https://twitter.com/3253638337/status/2102108575757939177)

### 3. Reduce overthinking and over-engineering at high effort

- OpenAI Codex, 2026-09-24, r/codex (Reddit): “but they said they had a very clear plan. i’ve been using 5.6 sol medium to execute against plans because anything higher tends to inflate scope and get off track. maybe it’s a trait of this generation of model or an issue with how compaction works in codex. either way, they need to fix the goldilocks behavior.” [source](https://www.reddit.com/r/codex/comments/1worivz/unpopular_opinion_sol_6_xhigh_is_pretty_decent/pbs9jyf/)
- Devin, 2026-09-12, @cognition (X): “@devindesktop @cognition can you plz make "smart" mode for devin better, it is wayy too restrictive” [source](https://twitter.com/1754601800135585793/status/2098635974192271530)
- OpenAI Codex, 2026-09-08, r/codex (Reddit): “yap openai model almost always over engineer if you use the smarter model or give them too high effort. i think is the way they rl it to have that super persistent behaviour. so they kind of just if i cannot catch all negative scenario i will keep writing until i reach 0% error rate. when some of the cases probably will never ever happen.” [source](https://www.reddit.com/r/codex/comments/1wa021z/unpopular_opinion_the_astra_complaints_are_more/p8i6q5g/)

### 4. Clearer guidance on effort level tradeoffs

- Claude Code, 2026-09-27, r/ClaudeCode (Reddit): “would be sick to know what we lose or gain going up or down like more details ideally.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrvhg1/for_those_running_large_mostly_autonomous/pcg6n7z/)
- OpenAI Codex, 2026-09-26, r/codex (Reddit): “intersect this w/ the constant mental overhead of deciding the model x reasoning effort and now we're cooking. they're just trying to keep our little brains engaged” [source](https://www.reddit.com/r/codex/comments/1wr3yev/ive_been_seeing_a_lot_of_comments_circling_around/pc9hxdf/)
- OpenAI Codex, 2026-09-25, X search: OpenAI Codex, Codex CLI, Codex app (X): “can we have efforts and model as separate selection in codex cli like claude code? @thsottiaux” [source](https://twitter.com/3197003120/status/2103449443957891463)

### 5. Persistent effort setting across sessions

- Claude Code, 2026-09-26, @ClaudeDevs (X): “@claudedevs cc said saved as your default for new sessions with max effort , but actually cannot” [source](https://twitter.com/1571075404164698115/status/2103725741024457053)
- Cline, 2026-09-21, @cline (X): “@cline why does the desktop app always use low reasoning mode for all models as default? even if i change it to high or extra it reverts to low.” [source](https://twitter.com/2065830145114660865/status/2101927357821096407)
- Google Antigravity, 2026-09-19, @antigravity (X): “@antigravity since the version 2.15 the /boost command stay for every subsequent request once we call it. let's say that i use /boost for a big task, once the big task is complete i want to do all the subsequent task with the normal mode. but /boost will be invoked when i don't ask for it. the main issue is the token consumption and the time it will take to complete tasks that are simple.” [source](https://twitter.com/1268332754661490693/status/2101341426923585603)

### 6. Reasoning effort control for custom and third-party models

- OpenCode, 2026-09-21, @opencode (X): “hey @conductor_build please allow me to select reasoning levels for @opencode models 🙏🏽” [source](https://twitter.com/50570112/status/2101967617686970808)
- Factory, 2026-09-19, @droid (X): “@tereza_tizkova @droid i have no way to choose the thinking level when accessing the custom model. is this a problem? but i directly asked the model to modify the settings by itself. hahaha” [source](https://twitter.com/1889310672368095232/status/2101460620063154561)
- Cursor, 2026-09-06, r/cursor (Reddit): “i'm using a deepseek model via openrouter in cursor. cursor shows a reasoning level selector (low/medium/high) for its built-in models, but not for my custom openrouter model. is there a way to set reasoning effort. thanks.” [source](https://www.reddit.com/r/cursor/comments/1w92wed/how_to_set_reasoning_level_for_custom_openrouter/)

### 7. Respect effort settings for subagents

- Claude Code, 2026-09-25, r/ClaudeCode (Reddit): “yeah this is a harness failure i think. it should adhere to the reasoning level you ask for, instead it seems to always mirror the same reasoning level as the orchestrator. perhaps /feedback” [source](https://www.reddit.com/r/ClaudeCode/comments/1wpz5mr/your_subagents_probably_arent_running_at_the/pc1vm1m/)
- Google Antigravity, 2026-09-22, r/google_antigravity (Reddit): “i don't know if we can configure specifically thinking level to agent definition files. i want to use flash high for orchestrator main agent and flash medium for builder subagent. i cannot find that parameter in the documentation. it's nice to have that.” [source](https://www.reddit.com/r/google_antigravity/comments/1wm7g5r/weekly_quotas_known_issues_support_september_21/pbbdikj/)
- Claude Code, 2026-09-20, r/ClaudeCode (Reddit): “you can’t deterministically set effort level for subagents like you can the model. they inherit the orchestrating agents effort level. you can try to tailor in the prompt, that’s it, but it’s not binding.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wknv74/so_fable_is_pretty_much_off_the_table_for/pb0qrwe/)

### 8. Per-task model and effort configuration

- OpenAI Codex, 2026-09-20, r/codex (Reddit): “i just hope we can configure what model and thinking mode to use to each skill.” [source](https://www.reddit.com/r/codex/comments/1wl8tkp/we_need_some_better_ux_around_effort_switching/paws2rz/)
- Amp, 2026-09-16, @AmpCode (X): “@ampcode amp plugins show-agent-options --json returns empty efforts, i guess that's a bug? also: any chance to enable effort selection in the raw mode panel, just like with fast/normal mode? <strict_link>” [source](https://twitter.com/87301883/status/2100131804153819594)
- Claude Code, 2026-09-15, r/ClaudeCode (Reddit): “the agent tool doesn't have a parameter for effort. edit - there's a simple solution to that; have it spawn headless claude sessions, they're more efficient anyway.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wh7sqb/nah_this_some_bs/pa1cd5k/)

### 9. Better default reasoning effort

- Claude Code, 2026-09-23, @ClaudeDevs (X): “@claudedevs @addyosmani why wouldn’t “think carefully” to be the default? why do we need to drop it?” [source](https://twitter.com/1138168374/status/2102696535721234628)
- Amp, 2026-09-22, @AmpCode (X): “@ampcode it's been sometime - looking forward to your post about opus 5.5 and update to the default dial :p” [source](https://twitter.com/85549810/status/2102480563596591364)
- Factory, 2026-09-22, @FactoryAI (X): “@factoryai @spacexai need medium as the grok 4.7 default on droid” [source](https://twitter.com/1811332417099055105/status/2102287768621568050)

### 10. Option to disable thinking entirely

- Claude Code, 2026-09-16, @ClaudeDevs (X): “@claudedevs can we please, in claude dot ai, get the option to turn off thinking in opus 5 like in claude code when effort is below high?” [source](https://twitter.com/1932087056437800960/status/2100320608198516933)
- OpenAI Codex, 2026-09-04, r/codex (Reddit): “looking at what astra (none) got on frontiermath t4 (higher than 5.6 sol pro max)... they should give us a no thinking option in codex xd” [source](https://www.reddit.com/r/codex/comments/1w6nf7d/confirmed_bank_reset/p7oobnq/)
- OpenCode, 2026-09-03, r/opencode (Reddit): “the only options are default, low, high and max. i need it completely gone.” [source](https://www.reddit.com/r/opencode/comments/1w62dqj/how_to_turn_off_deepseek_model_thinking_using_the/p7jlmi5/)

### 11. Deep thinking reasoning mode

- Google Antigravity, 2026-09-27, r/google_antigravity (Reddit): “for me as ultra 20 user it is needed, it will be nothing for the quota, also wont harm you to have additional option, at least it will be more useful with the next good enough models” [source](https://www.reddit.com/r/google_antigravity/comments/1wrn9q5/codex_has_ultra_thinking_level_claude_code_has/pcecsq1/)
- Google Antigravity, 2026-09-27, r/google_antigravity (Reddit): “yea that's why i said a fraction of it. i am not against of it being added to antigravity, it's a must at this point.” [source](https://www.reddit.com/r/google_antigravity/comments/1wrn9q5/codex_has_ultra_thinking_level_claude_code_has/pce1l2h/)
- Cline, 2026-09-09, @cline (X): “@cline allow using glm 5.3 with reasoning, you are already letting it be used free, why not just having reasoning option?” [source](https://twitter.com/1489236941899911170/status/2097784036454494693)

### 12. Effort level slider

- OpenAI Codex, 2026-09-24, r/codex (Reddit): “it kept thinking at 1 step for 13 minutes, then it hit some time limit, then second attempt again same, 3rd time it finished and decided what do. this model doesn't know when to stop thinking, we don't have adjustable thinking effort slider also. and at 30-40 tps is too slow time also have some value” [source](https://www.reddit.com/r/codex/comments/1wmp5bh/xiaomi_just_aboslutely_killed_pareto_frontier/pbthgwt/)
- OpenAI Codex, 2026-09-24, r/codex (Reddit): “reinstate catastrophe, also for me. gpt5.6 sol was great, i was never dissatisfied with it. gpt6 sol ignores my plugins, skills, agent instructions, and entire workflows and jeopardizes the product. i have pointed this out several times, it always acknowledges it and continues to do it wrong. my wife is also missing the thinking slider in the app. something has gone wrong!” [source](https://www.reddit.com/r/codex/comments/1woiw95/something_is_wrong_with_gpt_6_sol/pbqkgh8/)
- OpenAI Codex, 2026-09-23, X search: OpenAI Codex, Codex CLI, Codex app (X): “openai codex / gpt 6 sol should have a persistence slider / its naturally what makes astra the goat <strict_link>” [source](https://twitter.com/281849089/status/2102814293536465042)

## Every agent

| Agent | Overall rank | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|---|
| [Google Antigravity](https://feedbackbench.com/agents/antigravity.md) | =5 | Typical | 0.523 | 0.497–0.548 | 49 | 28 | 21 |
| [Claude Code](https://feedbackbench.com/agents/claude-code.md) | 1 | Typical | 0.521 | 0.489–0.548 | 267 | 132 | 135 |
| [OpenAI Codex](https://feedbackbench.com/agents/codex.md) | 2 | Typical | 0.504 | 0.484–0.524 | 561 | 260 | 301 |
| [OpenCode](https://feedbackbench.com/agents/opencode.md) | 3 | Worse than peers | 0.445 | 0.420–0.470 | 57 | 12 | 45 |
| [Cursor](https://feedbackbench.com/agents/cursor.md) | 4 | Too few posts | – | – | 29 | 16 | 13 |
| [Devin](https://feedbackbench.com/agents/devin.md) | =5 | Too few posts | – | – | 17 | 9 | 8 |
| [GitHub Copilot](https://feedbackbench.com/agents/copilot.md) | =8 | Too few posts | – | – | 17 | 5 | 12 |
| [Cline](https://feedbackbench.com/agents/cline.md) | =8 | Too few posts | – | – | 12 | 1 | 11 |
| [Pi](https://feedbackbench.com/agents/pi.md) | 7 | Too few posts | – | – | 10 | 5 | 5 |
| [Amp](https://feedbackbench.com/agents/amp.md) | =11 | Too few posts | – | – | 10 | 3 | 7 |
| [Factory](https://feedbackbench.com/agents/factory.md) | =11 | Too few posts | – | – | 2 | 1 | 1 |
| [Zed](https://feedbackbench.com/agents/zed.md) | =8 | Too few posts | – | – | 1 | 0 | 1 |
| [Conductor](https://feedbackbench.com/agents/conductor.md) | 14 | Too few posts | – | – | 1 | 0 | 1 |
| [Kiro](https://feedbackbench.com/agents/kiro.md) | 13 | Too few posts | – | – | 0 | 0 | 0 |
| [Warp](https://feedbackbench.com/agents/warp.md) | 15 | Too few posts | – | – | 0 | 0 | 0 |
| [Grok Build](https://feedbackbench.com/agents/grok-build.md) | 16 | Too few posts | – | – | 0 | 0 | 0 |
| [Augment Code](https://feedbackbench.com/agents/augment.md) | 17 | Too few posts | – | – | 0 | 0 | 0 |

## Posts

Receipts rule: The 5 most recent praise and complaint posts per area (first 700 characters) and 3 per criterion (first 450 characters).

### Google Antigravity

- Praise, 2026-09-24, r/google_antigravity (Reddit): “i use the medium and it works much better than the high. lately, it has been going better for me, but i have also changed the instructions to a model with smaller rules and a smaller [agents.md](<strict_link>), and maybe that has something to do with it.” [source](https://www.reddit.com/r/google_antigravity/comments/1wp3k02/is_gemini_38_flash_getting_stuck_in_loops_for/pbs3wpn/)
- Praise, 2026-09-19, r/google_antigravity (Reddit): “only 3.8 high. using any other is criminal waste of your subscription at this point.” [source](https://www.reddit.com/r/google_antigravity/comments/1whnh0m/which_antigravity_model_do_you_actually_use_for/pataxxh/)
- Praise, 2026-09-19, r/google_antigravity (Reddit): “high for planning. low for implementing work that is already planned.” [source](https://www.reddit.com/r/google_antigravity/comments/1whnh0m/which_antigravity_model_do_you_actually_use_for/pattfcs/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “you could simulate it using a combination of mcp + prompt but it is nowhere near a native thinking token” [source](https://www.reddit.com/r/google_antigravity/comments/1wrn9q5/codex_has_ultra_thinking_level_claude_code_has/pcdy6m4/)
- Complaint, 2026-09-26, r/google_antigravity (Reddit): “a model's effort is just a system prompt. it shouldn't have that much of an impact on the model when generating a simple response. its direct competitor, sonnet 5 high, doesn't have this problem at all” [source](https://www.reddit.com/r/google_antigravity/comments/1wqkqcl/why_is_gemini_38_flash_so_slow/pc4t1ds/)
- Complaint, 2026-09-23, r/google_antigravity (Reddit): “high might be bugged, it has been an ongoing issue for about a week now. try medium, it works fine for both 3.7 and 3.8” [source](https://www.reddit.com/r/google_antigravity/comments/1wnxf4i/37_flash_high/pbiyrpe/)

### Claude Code

- Praise, 2026-09-27, r/ClaudeCode (Reddit): “this. i have a swarm concept and had a couple of them build out units in parallel with opus low vs sonnet medium. opus won on cost quality and speed. then i switched my builder to opus low and my cost per pr plummeted by half bc i was spending less iterations on review” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrfr8p/fable_51_or_opus_55/pccljkn/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “medium as a good baseline. high will look at more edge cases, run more tests.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrmw8c/high_or_medium_for_opus_55/pcdrtii/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “i’ve been using extra high myself after i saw someone did a comparison video of results of the same prompts with extra high being noticeably better. it does seem like my results are noticeably better than when i first used medium and high and at least on a max 20 plan i’m having no issues with running out of use.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrmw8c/high_or_medium_for_opus_55/pcdtk46/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “max effort is just “i want to quit working burn my tokens please” mode.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqtk1l/its_just_so_good/pcdbf8h/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “optimizing tokens this early is the wrong dial. the expensive thing is not turns, it's the day of work you throw away when the data model turns out wrong, and no effort setting saves you from that. get the feature described in a few plain sentences before you start. cheapest step in the whole workflow.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrfxtu/new_to_claude_code_how_do_i_maximize_usage/pcdsh7i/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “many people are thinking „coding is really difficult, therefore i must set reasoning as high as possible“. as senior developer with over 40 years of coding experience i have to admit … coding is really easy, if you know what you want. setting reasoning to high on big models (or even qwen 3.8 😱) can totally ruin your code by overthinking and over engineering.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrmw8c/high_or_medium_for_opus_55/pcect5x/)

### OpenAI Codex

- Praise, 2026-09-27, r/codex (Reddit): “agreed. luna 6 max for reading, some research work, sol 6 high for coding with a escalation to xhigh and in veery rare cases astra if having problems” [source](https://www.reddit.com/r/codex/comments/1wrb1xg/i_tested_astra_solo_vs_astra_orchestrating_luna/pcbwdim/)
- Praise, 2026-09-27, r/codex (Reddit): “5.5 xhigh was a great workhorse and pretty balanced & reliable imo. not overengineering/overthinking and it was possible to get it to do what it was asked to do” [source](https://www.reddit.com/r/codex/comments/1wri0al/gpt_55_usage_vs_sol_566_and_astra/pcd40k7/)
- Praise, 2026-09-27, r/codex (Reddit): “i mean, obviously the point of a smart model is to solve hard problems. the advice of selecting a reasoning level that matches the task is good advice.” [source](https://www.reddit.com/r/codex/comments/1wrw3wy/1000_lines_348_bn_tok_393_subagents_42_pro20/pcgg9y9/)
- Complaint, 2026-09-27, r/codex (Reddit): “one thing i found caused more usage was allowing codex to raise the reasoning level. this was when i was trying astra and i kept finding astra as both the default model (i only tried it, i never set it as default) and reasoning level set higher. in config.toml make sure you set the default model you want and also make sure model\_reasoning\_summary is not set to "auto", i make mine: i make mine and then if i need something different i do it myse” [source](https://www.reddit.com/r/codex/comments/1wo8fhw/okay_they_literally_cut_our_quota_by_half_gpt_6/pcayrwh/)
- Complaint, 2026-09-27, r/codex (Reddit): “i only use high and medium. i am in the 0.1% of people who use this tech in terms of previous experience with swe and exposure to ai from pre gpt2. will not tell you how to do your business but i agree with this person. max/highest reasoning has always been shit. overthinks. overengineers. takes forever.” [source](https://www.reddit.com/r/codex/comments/1wr4cp6/more_resets_incoming_next_week/pcbdm52/)
- Complaint, 2026-09-27, r/codex (Reddit): “honestly i just don't see the point of openai models at the moment.. they just seem so "dumb" and hard to work with compared to opus. i think i need to see something like astra 6.5 or something to consider wasting money on a sub... maybe 5x *only* for adversarial stuff, but again, as you said here, astra loves to overthink things and add random crap that doesn't matter.” [source](https://www.reddit.com/r/codex/comments/1wqc44m/reset_confirmed/pcburvi/)

### OpenCode

- Praise, 2026-09-26, @opencode (X): “@jlongster @opencode i like it on low” [source](https://twitter.com/1252868508762771457/status/2103847204842860608)
- Praise, 2026-09-22, r/opencodeCLI (Reddit): “wow and even with higher cache cost its still better as it will not over thing and not eats tokens like a black hole” [source](https://www.reddit.com/r/opencodeCLI/comments/1wnhazu/gpt6_luna_cheap_than_deepseek_v41_flash/pbf7psb/)
- Praise, 2026-09-20, r/opencodeCLI (Reddit): “i like omo-slim and it's delegation, because i'm very much into watching all the models and finding strengths and weaknesses. i also built up up the process around writing plans so i know the heavy thinking is going on up front and i can insert myself if i want. i also have vs code set up as my ide and the kilo code plugin for ai. it's much less detailed than omo or it's offshoots, but if i really want to hand-hold for caution i can take the plan” [source](https://www.reddit.com/r/opencodeCLI/comments/1wl18ez/matrixx_ohmyopencode/pb1k1ix/)
- Complaint, 2026-09-25, r/opencode (Reddit): “it’s odd. on high, wrote up a prd for a project and did a good job. on low, asked it to scaffold the project folder based off the prd and it got stuck tool calling for over 39 mins.” [source](https://www.reddit.com/r/opencode/comments/1wprfba/space_bunny_is_better_than_i_expected/pbyj51b/)
- Complaint, 2026-09-25, @opencode (X): “@andyjscott @opencode typical. can be issue if use some proxy like eg. 9router proxy than level thinking can be a problem (default/low etc.)” [source](https://twitter.com/1584977819372847124/status/2103521353613955516)
- Complaint, 2026-09-24, r/opencodeCLI (Reddit): “meta spark 1.3 is really great especially the max reasoning level is very impressive. for specific low level tasks it seems to be better trained than deepseek 4.1 flash. i tryed just today to use muse spark 1.3 in opencode zen with max reasoning however just found out that the max reasoning level was removed and does not exist anymore in opencode zen sadly. without the max reasoning level muse spark 1.3 is same like deepseek 4.1 flash and i prefe” [source](https://www.reddit.com/r/opencodeCLI/comments/1wopbtv/what_muse_spark_14_contributor_is_already_here/pbphnsq/)

### Cursor

- Praise, 2026-09-27, r/cursor (Reddit): “sonnet's also faster, we don't need opus/fable level intelligence for most tasks 🤷🏿♂️” [source](https://www.reddit.com/r/cursor/comments/1wp6j44/is_this_supposed_to_be_good_news/pcbgqaz/)
- Praise, 2026-09-25, r/cursor (Reddit): “grok works flawlessly for me on extra high effort 🤷🏾♂️.” [source](https://www.reddit.com/r/cursor/comments/1wpd3mm/i_hate_to_admit_it_but_grok_sucks/pbytfs4/)
- Praise, 2026-09-24, r/cursor (Reddit): “i still think using auto can help, if you think the problem isn’t complex switch to midum temp not high” [source](https://www.reddit.com/r/cursor/comments/1wosirg/its_just_getting_bad/pbq0qx7/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “i'm using claude code extension in cursor and i have to have some effort selection set, there is no auto. if i set it to high and tell it to decide by itself will it actually work if you know?” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrmw8c/high_or_medium_for_opus_55/pce2dzm/)
- Complaint, 2026-09-24, r/cursor (Reddit): “it seems good to me on low thinking, medium though for a head to head test took about 3x longer and was slightly worse on the end result. it seems to just waste too many tokens above low. also i noticed on low thinking it did some automation tasks i wanted it to do really, really well.” [source](https://www.reddit.com/r/cursor/comments/1woad6b/grok_47_performance_in_cursor/pbobumy/)
- Complaint, 2026-09-22, r/cursor (Reddit): “i stopped trusting the new-chat default. model and effort are part of the task for me, not whatever cursor last left selected. composer is fine for day-to-day work. for anything i treat as a real review pass i pin high (or higher) on purpose, because a shallow default quietly misses stuff and you only notice after the patch is already messy. whatever you settle on, set it before the agent starts. a silent swap mid-workflow is worse than picking a” [source](https://www.reddit.com/r/cursor/comments/1wn17sn/cursor_default_model_to_highest_model_with_high/pbc5y4f/)

### Devin

- Praise, 2026-09-24, @cognition (X): “@rohit3a actually medium effort is the highest ranked on @cognition” [source](https://twitter.com/615971509/status/2103237274892702174)
- Praise, 2026-09-24, @cognition (X): “@orbarak123 @cognition agreed! and for a good reason. medium is almost perfect for 95% of work.” [source](https://twitter.com/2914778029/status/2103237780348575885)
- Praise, 2026-09-23, @cognition (X): “@lotusdecoder @cognition @cursor_ai for questions of different difficulty, different effort should be chosen.” [source](https://twitter.com/1504955609619177475/status/2102571063205036508)
- Complaint, 2026-09-23, @cognition (X): “@cognition how is xhigh and max worse than high? maybe you have to look into the benchmark” [source](https://twitter.com/1717163021519175680/status/2102741802495173097)
- Complaint, 2026-09-21, @cognition (X): “@cognition 4.7's own curve slopes down though. more effort, worse score.” [source](https://twitter.com/1604518234962657280/status/2102166904282513767)
- Complaint, 2026-09-18, @cognition (X): “@silasalberti @cognition great to hear! been my daily driver since launch. did some testing on the levels recently and high exhibited some strange behavior. assumed that was because of compute and it was just spinning waiting (but not errored out or stopped). <strict_link>” [source](https://twitter.com/2040861418547798016/status/2101055731025719340)

### GitHub Copilot

- Praise, 2026-09-19, r/GithubCopilot (Reddit): “i just set it to luna medium for price, speed and not over-thinking things like max does.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wglp9b/model_selection_best_practices/pat8org/)
- Praise, 2026-09-12, r/GithubCopilot (Reddit): “high is good. i love it” [source](https://www.reddit.com/r/GithubCopilot/comments/1wdtxyz/luna_is_still_the_goat/p99unvi/)
- Praise, 2026-09-11, r/GithubCopilot (Reddit): “100% agree luna on max only way to go” [source](https://www.reddit.com/r/GithubCopilot/comments/1wdtxyz/luna_is_still_the_goat/p98p9rx/)
- Complaint, 2026-09-25, r/GithubCopilot (Reddit): “i agree. i find the praise for luna baffling. it may be cheap, but most people seem to recommend running it at xhigh or max, where it takes forever to reason and generates bloated output.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wp1ygm/tiers_of_auto_now_available/pbwf62p/)
- Complaint, 2026-09-25, r/GithubCopilot (Reddit): “interesting, my experience is that luna 6 is better with medium-xhigh, and drops insane in quality on max.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wptkkq/how_does_gpt6_sol_feel_so_far/pc11atx/)
- Complaint, 2026-09-08, r/GithubCopilot (Reddit): “never use low of anything, it's designed to make mistakes” [source](https://www.reddit.com/r/GithubCopilot/comments/1wag6w6/breaking_msft_has_stopped_providing_claude_models/p8jzbz3/)

### Cline

- Praise, 2026-09-24, r/LocalLLaMA (Reddit): “qwen3.8-27b-nvfp4-mtp has been outstanding for me (with cline) when hosted in lm studio with temp set to .1 and the "reasoning budget" to 1024. the latter definitely eliminated the annoying over thinking.” [source](https://www.reddit.com/r/LocalLLaMA/comments/1wp0z3i/qwen3827b_is_good_enough_that_i_stopped_using_api/pbrd9b5/)
- Complaint, 2026-09-26, @cline (X): “@cline it's also quite slow at the auto/highest reasoning 🤔🤔” [source](https://twitter.com/1803325156955480064/status/2103667261743759603)
- Complaint, 2026-09-21, @cline (X): “@cline why does the desktop app always use low reasoning mode for all models as default? even if i change it to high or extra it reverts to low.” [source](https://twitter.com/2065830145114660865/status/2101927357821096407)
- Complaint, 2026-09-21, @cline (X): “from user reports on the cline announcement, free kimi k3 in desktop appears limited to (or stealth-downgrades to) low reasoning effort. the setting often resets to low due to a known ui bug when switching sessions. higher efforts may not be available or usable under the free quota.” [source](https://twitter.com/1720665183188922368/status/2101990895054704684)

### Pi

- Praise, 2026-09-25, r/PiCodingAgent (Reddit): “the point everyone is making is that luna max can function equivalently to the larger models and thus should be considered for the same comparison. definitely less thinking is faster and works well with some guidance” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wpcis6/opus_55_vs_gpt6_sol_luna_in_piagent_results_on_my/pbvj728/)
- Praise, 2026-09-25, r/PiCodingAgent (Reddit): “i can use anthropic models in pi with amazon bedrock. but nowadays only use chatgpt sol with low reasoning to code and it works nicely with pi.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wq1bv2/whats_the_best_option_for_using_anthropic/pc07mn3/)
- Praise, 2026-09-24, r/PiCodingAgent (Reddit): “for most of my simple tasks medium is fine, and it will be faster than max for small sessions.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wpcis6/opus_55_vs_gpt6_sol_luna_in_piagent_results_on_my/pbuqwe0/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “had too. most of the time it would hit the context window every single time and then didn't answer the prompt. also didn't see much difference between thinking modes, but that might just be my perception after 30m of waiting for the model to actually come to a conclusion. any conclusion at all.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wnmsho/i_want_to_believe_in_local_llms_for_coding_but/pc9s2g3/)
- Complaint, 2026-09-26, r/PiCodingAgent (Reddit): “reasoning off for all requests? that's crazy for a model that is designed for massive thinking traces.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wnmsho/i_want_to_believe_in_local_llms_for_coding_but/pc4dne9/)
- Complaint, 2026-09-22, r/PiCodingAgent (Reddit): “if asking pi failed, that is because pi itself introduced a bug in 0.86.x upwards that silently drops thinking levels other than off or medium.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wmo5e5/how_do_i_support_thinking_mode_correctly_with/pb9r60u/)

### Amp

- Praise, 2026-09-23, @AmpCode (X): “@paelgutierrez @ampcode yes, that’s true, but i like having the prebuilt modes and also choosing the oracle that goes with it” [source](https://twitter.com/33135576/status/2102871479402819921)
- Praise, 2026-09-04, @AmpCode (X): “👀 how have i missed that @ampcode now allows dial mode tuning? <strict_link>” [source](https://twitter.com/1751950522502860800/status/2095806978937249828)
- Praise, 2026-09-04, @AmpCode (X): “low mode on @ampcode is more like "value" mode imo 👏😀 i recently defaulted to it and on most harnesses (vs @zai_org glm-5.3). i route heavy brain juice work to other modes / latter model the flip is inevitable <strict_link>” [source](https://twitter.com/3159431/status/2095870767607013662)
- Complaint, 2026-09-24, @AmpCode (X): “@homborg @ampcode this feels like the optimal setup, but i'm curious if you've had experience with opus 5.5 as the coordinator as well? there's so many model + effort combos now, plus models are getting more and more powerful, that i just want one config that works 90% of the time 😅” [source](https://twitter.com/2059667636758491136/status/2103149063214604662)
- Complaint, 2026-09-23, @AmpCode (X): “@jkudish @ampcode agree! could be an option in "build your own dial" to just add a couple more. sometimes i want "high but with fable" vs my regular high with sol.” [source](https://twitter.com/1567548474794835969/status/2102856772021305745)
- Complaint, 2026-09-14, @AmpCode (X): “@sqs @ampcode another thing that's confusing to me is i can't change the model level (i.e. "medium") mid-conversation. it's just greyed out. and clicking on the little gauge just pops up a tooltip saying "medium" that quickly disappears which is confusing” [source](https://twitter.com/721234540/status/2099487663384334421)

### Factory

- Praise, 2026-09-21, @FactoryAI (X): “@factoryai @spacexai the faster move from diagnosis to concrete commands sounds especially useful for infra work. medium as the default is a good sign.” [source](https://twitter.com/1952719479030661120/status/2102173479764472106)
- Complaint, 2026-09-19, @droid (X): “@tereza_tizkova @droid i have no way to choose the thinking level when accessing the custom model. is this a problem? but i directly asked the model to modify the settings by itself. hahaha” [source](https://twitter.com/1889310672368095232/status/2101460620063154561)

### Zed

- Complaint, 2026-09-01, r/ZedEditor (Reddit): “hi guys i am new to this editor. i just want to know if there is a way to configure models efforts in zed agent (max, low, highx etc). i could not find it. thanks in advance!” [source](https://www.reddit.com/r/ZedEditor/comments/1w46ffc/is_there_a_way_to_select_model_efforts_in_zed/)

### Conductor

- Complaint, 2026-09-21, @conductor_build (X): “hey @conductor_build please allow me to select reasoning levels for @opencode models 🙏🏽” [source](https://twitter.com/50570112/status/2101967617686970808)
