# Asks the user versus guessing (`context.clarifying_questions`)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/criterion/context.clarifying_questions

Area: [Instructing and context](https://feedbackbench.com/criteria/context.md)

**Definition.** Whether the agent stops to ask clarifying questions when it is blocked or the request is ambiguous, instead of inventing assumptions or workarounds.

**Boundary.** Not this: see [Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md) for repetitive failure without any question.

Rated author-weeks, all agents: 226. Complaint share: 62%.

## The brief

Written by Claude Opus 5.5 from 41 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**Agents still guess when they should ask, and ask when they should act.**

TL;DR:

- The top complaint is confident guessing that fills gaps with wrong assumptions instead of one question.
- The opposite complaint runs close behind, with trivial or repeated questions that stall well-planned work.
- Users who are happy mostly force the behavior themselves with grill-me skills, plan mode, or rules files.

In plain terms: You hand over a vague task and the agent quietly picks an interpretation and runs. Or it stops a clean run to ask something obvious. Users who get good results force questions up front, before any code is written.

### How it breaks

- **Confident guesses instead of questions** ([Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md)). The costliest pattern is an agent that fills ambiguity with a confident assumption and keeps going, so wrong work surfaces only at review.
  Posts describe models that treat missing information as license to invent. One user wants a confidence threshold below which the agent asks or admits it does not know. Others report a jump in corrections when a model starts guessing. Another found that plain implementation works well only with a tightly defined plan. The shared complaint is that a wrong answer costs more than a question would have.
  Evidence:
  - Complaint, OpenCode, r/opencode, 2026-09-25: “i dont like the fact that it keeps guessing rather than look for answer and if it doesnt find them to consult me” [source](https://www.reddit.com/r/opencode/comments/1wocn5s/space_bunny_thoughts/pc1490a/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-17: “not you. 3 corrections in 90 work orders vs 15 in 10 with the same prompts is the model, full stop. opus filling gaps with confident guesses instead of questions matches what we see too. quick fix: before a work order ships, ask a different lab's model one question, "what's assumed here that wasn't stated". the second model doesn't share the first one's assumptions, so it catches exactly this.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wifep9/opus5_vs_sonnet5_for_work_order_building_opus5_is/pacb0m7/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-21: “nah, sometimes it really goes off the completely wrong end. what i would like is for it to come up with an answer only if it has a strong confidence level. i should be allowed to configure it such that below that threshold it either asks me questions for details that could help, or just tell me it doesn't know. coming up with wrong answers is way worse.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wm4ncx/im_afraid_to_use_opus_5/pb90c7j/)
  - Complaint, Google Antigravity, @antigravity, 2026-09-05: “it works well for implementation, but you need to give it a pretty well defined plan. on it's own it makes a lot of assumptions that are incorrect and thus not reliable to use as the main model. also, i'd love to see more rendered components inside of the chat window that can help with visualizing complex architecture and abstractions.” [source](https://twitter.com/999042490626797568/status/2096310737442328595)

- **Builds when you only wanted talk** ([Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md)). Users who want to explore or brainstorm report agents jumping straight into edits, so they add prompt guards to keep the agent in discussion mode.
  The request is simple. Discuss first, implement on command. Several posts say the model skips that step and starts executing. Users work around it by ending prompts with explicit analyze-only instructions. One user praises a model precisely because it explains a bug without touching the codebase, which shows how rare that restraint feels.
  Evidence:
  - Complaint, OpenCode, @opencode, 2026-09-06: “been using muse 1.3 via @opencode for a few days now, and i feel this model just proceeds to go ahead and execute even when i just want to go back-n-forth. like bro i'm gonna tell you when to implement it, chill out anyone else? what r ur thoughts? <strict_link>” [source](https://twitter.com/1149503292063436800/status/2096683881479217561)
  - Complaint, Pi, r/PiCodingAgent, 2026-08-31: “i often add "just analyze and answer" at the end of my prompt when i start a session for brainstorming or exploring, especially with 3.8 that tends to start implementation blindly sometimes.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1w3lyr7/pi_is_overeager_and_cave_man_question/p71ura1/)
  - Praise, OpenCode, r/opencode, 2026-09-10: “i meant i skipped opencode go because their limit for glm 5.3 flash was lower than commandcode goat plan. as for my experience with glm 5.3 flash its been a pleasant week using it for my day to day work. i use it in zcode desktop app and everything just works. its able to handle document editing, google docs editing (via google workspace mcp), creating and editing bricks builder element via json (using bricks skill), creating and editing wordpress plugin, and other tasks. is it perfect? nope, but is it good enough for my needs? yup. what i like about glm 5.3 flash is that the model is not that eager to just jump straight to editing the project when i just asked it to explain why certain bug is happening, compared to deepseek that will automatically edited the codebase even though i havent told it to do so.” [source](https://www.reddit.com/r/opencode/comments/1wbx4fw/has_glm53flash_gotten_cheaper/p8vgfg1/)

- **Too many questions, too little substance** ([Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md)). The opposite failure is real too. Agents ask trivial or repeated questions, stall planned work, and push users to add rules that suppress asking.
  Users describe finishing a well-planned task only to get a doubt back instead of a result. Others complain of the same question rephrased repeatedly, or of security questions that are wild goose chases because the agent never checked whether the code is reachable. One fix circulating is a rule to ask only when blocked. Another user prefers a model that stops less often, to save time and context.
  Evidence:
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-22: “you nailed the attitude! you work through a well planned task, you wait for the result and the result is… a doubt???” [source](https://www.reddit.com/r/ClaudeCode/comments/1wmr99j/columbo_questions/pb9zgkb/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-27: “worth checking how many turns are just it asking you questions. a "don't ask unless blocked" rule in claude.md cuts a lot of that.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wropen/i_keep_hitting_usage_limits_on_20x_plan_have/pceof2h/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-22: “that's a good one but it's been weeks since i cancelled my sub and i remember vividly how it would ask this kind of question on every single task so i don't think that's new. vividly because usually it's a wild goose chase cause it forgot to check whether that thing is even reachable or used or it's wondering whether we should harden against an attack that if possible would already mean that the guy owns your server.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wmr99j/columbo_questions/pb9ysnw/)
  - Praise, Cursor, r/cursor, 2026-09-10: “opus 4.6 is my daily driver in planning mode and to execute the plan, then when the task is complete if i have any follows ups i use got 5.2, not codex version. i find 5.2 works for longer and doesn't stop to ask you questions as much saving on time and context by skipping useless uodate messages that wait for a response. honestly if you're not using planning mode you should try it. been using cursor for about 2yrs and i wish i'd started using it sooner. makes a big difference.” [source](https://www.reddit.com/r/cursor/comments/1wc95mb/i_am_tired_of_handholding_composer_25_what_are/p8w6c0a/)

- **Questions that get lost or ignored** ([Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md)). When agents do ask, the mechanics often fail. Questions are buried in prose, answered by the agent itself, or broken by the tool meant to collect them.
  One user misses questions tucked into the last sentence of a long explanation and asks for a dedicated UI tool. Others report the opposite problem with existing tools. One returns an instant decline, another allows only preset options. Some agents ask asynchronously but keep working and decide for you if you are slow to reply. Follow-up questions mid-run can also hang autonomous execution.
  Evidence:
  - Complaint, Amp, @AmpCode, 2026-09-25: “@ampcode curious why amp doesn't have any question/answer tools the model can use. having instances where the modal outputs a big explanation then in the last sentence: may i do that? i sometimes miss that it's asking at all! a ui question tool would make that al ot more obvious.” [source](https://twitter.com/5444392/status/2103600382303944750)
  - Complaint, Pi, r/PiCodingAgent, 2026-09-15: “it fails when the ai tries to ask me a question (via the `ask_user_question` tool). immediately "user declined to answer questions"” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wgy56q/piui_a_month_later_harder_better_faster_stronger/p9zbmpi/)
  - Complaint, Cline, @cline, 2026-09-19: “@cline why can the model only choose from the options it provides when calling the ask question tool, and cannot enter other options on my own?” [source](https://twitter.com/721547404479111170/status/2101290071311945767)
  - Complaint, Cursor, @cursor_ai, 2026-09-27: “@cursor_ai should reconsider cx of follow up questions after execution of approved plan started. it hanged entire authonomy.” [source](https://twitter.com/255140211/status/2104192810089943545)

- **Forced interviews are the working fix** ([Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md)). The most praised workflow makes the agent interrogate the user before coding, using grill-me style skills, plan mode, or an explicit ask-me-questions prompt.
  Users credit these skills with countering the urge to assume and start work. Ten minutes of questions up front produces specs and keeps the user aware of what is being built. Several posts say simply instructing the agent to surface ambiguity works well. One reports moving faster since the agent must ask three questions before writing a line.
  Evidence:
  - Praise, Claude Code, r/ClaudeCode, 2026-09-01: “i've tried a lot of skills, some of them were helpful but for me the clear winners are the /grill-me and /grill-with-docs skills that other people mentioned. they helped me more than any other ai tip or workflow, since they fight against ai's impetus to just assume and start working on any prompt it receives.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w47y8j/which_skills_are_you_using_frequently/p7903ir/)
  - Praise, Cursor, @cursor_ai, 2026-09-02: “@cursor_ai verifying its own work is the easy half. the loop that matters is the one that can throw the work away. i've been faster since the agent has to ask 3 questions before it writes a line. fewer green checks. fewer seventh versions of something nobody asked for.” [source](https://twitter.com/2073295850491752448/status/2095090463149810118)
  - Praise, OpenAI Codex, r/codex, 2026-08-31: “i am not quite there yet but one learning was, that my poor results came from poor specifications. one thing that really helps is the **grill-me** skill. it essentially will ask you relevant questions until it completely understands what you want to achieve, and will create specs based on that. these 10 mins are a good investment and actually helps you to stay aware about what’s actually going on.” [source](https://www.reddit.com/r/codex/comments/1w3q8tu/poor_results_using_subagents/p72d4zx/)
  - Praise, OpenAI Codex, r/codex, 2026-09-23: “you can request it specify ambiguity before working. works great for me” [source](https://www.reddit.com/r/codex/comments/1wo3ixp/how_do_you_think_openai_will_respond_next_week/pbnosbc/)

### Who stands out

- **OpenAI Codex (mixed)**. Codex draws praise for asking when unclear and complaints for asking asynchronously, then deciding on its own when the user is slow to reply.
  Fans say it asks what they meant and will flag ambiguity on request. Critics describe repeated questions and a background run that does not wait for answers. One user saw grill questions appear while subagents were still working, so the questions came before the facts. Users also ask for clarifying questions mid-run outside plan mode.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-07: “i get the vibe astra always has a sort of inbuilt /goal mode running, yes it asks asynchronously but it just keeps working and in my experience if you don't answer in a timely manner the ai just tosses the dice and does something.” [source](https://www.reddit.com/r/codex/comments/1w9my9c/finally_figured_out_why_goal_used_100_of_my_usage/p8bjywj/)
  - Praise, OpenAI Codex, r/codex, 2026-09-14: “mine asked me today if it could. it didn't assume it could, it did ask. so this seems to be a new feature.” [source](https://www.reddit.com/r/codex/comments/1uu2c1g/reset_discussion_megathread/p9oju6c/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-07: “i need to see this lmao. i honestly tuned out a bit from here because i was getting tired of the repetitive questions over and over again” [source](https://www.reddit.com/r/codex/comments/1w9mco2/astra_on_lowhigh_20x_pro_0_limits_in_2_hours/p8bgdm5/)
  - Complaint, OpenAI Codex, X search: OpenAI Codex, Codex CLI, Codex app, 2026-09-07: “@mattpocockuk i don't know if you eval/benchmark how your skill works when a new model has been released but i'm not sure if the grilling skills (grill/wayfinder) work as expected with the new codex app/gpt-6. it appears that the questions have been created before the subagents have done their work. i'm not sure which part is the problem but something has changed and feels off. in this specific use case codex has prompted the grill questions in the new ui/ux (the one that shows the options in a separate dialog) but the model was still doing stuff in the background, reasoning, executing subagents and so on.” [source](https://twitter.com/14488050/status/2096831169970975144)

- **Claude Code (mixed)**. Claude Code users split between relief that it now asks instead of making assumptions and frustration at confident gap-filling and needless option menus.
  Some users prefer the newer behavior over past responses ruined by bad assumptions, and value flagging when stuck over raw capability. Others report the model guessing confidently where it should ask, or offering pointless choices on every task. The most common request is that it stop and ask when stuck or blocked.
  Evidence:
  - Praise, Claude Code, r/ClaudeCode, 2026-09-25: “i'd much rather have this behavior rather than the past of making idiotic assumptions that ruin an entire response. though i do like to retain control over the approach rather than just letting the model rip so i might be alone in this preference” [source](https://www.reddit.com/r/ClaudeCode/comments/1wpuez8/something_is_wrong_with_opus_55/pbygwfy/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-17: “not you. 3 corrections in 90 work orders vs 15 in 10 with the same prompts is the model, full stop. opus filling gaps with confident guesses instead of questions matches what we see too. quick fix: before a work order ships, ask a different lab's model one question, "what's assumed here that wasn't stated". the second model doesn't share the first one's assumptions, so it catches exactly this.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wifep9/opus5_vs_sonnet5_for_work_order_building_opus5_is/pacb0m7/)
  - Praise, Claude Code, @ClaudeDevs, 2026-09-01: “@claudedevs getting further into a long task before it needs input and being better at flagging when it is stuck matters more day to day than another benchmark point. silent failure mid-task is the actual pain with agentic coding, not raw capability.” [source](https://twitter.com/2011133265089101824/status/2094851588695236732)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-10: “usually claude only offers me options when it’s actually a architecture or design choice. although i have gotten some pretty dumb ones in the past hey! i found a fatal flaw that could expose all user information a- fix the flaw (recommended) b- ignore the flaw c- come back to it later 😂” [source](https://www.reddit.com/r/ClaudeCode/comments/1wcr862/claude_offering_pointless_options/p906lz7/)

- **Google Antigravity (weaker)**. Users say Antigravity runs with uncertain requests unless plan mode is on, and contrast it unfavorably with agents that ask for clarification.
  One user says it often finishes the task before they understand what happened. Another says it makes many incorrect assumptions without a well-defined plan. Praise exists, but it comes from users who explicitly tell the model to ask them questions, not from default behavior.
  Evidence:
  - Complaint, Google Antigravity, @antigravity, 2026-09-26: “antigravity is the only one where i have to use "plan mode"; without it, gemini 3.8 might drag me off somewhere without me even realizing it. while claude and chatgpt automatically ask for clarification on anything uncertain, antigravity just runs with it—often finishing the task before i’ve even fully grasped what happened.” [source](https://twitter.com/1533997728858243072/status/2103895076707697079)
  - Complaint, Google Antigravity, @antigravity, 2026-09-05: “it works well for implementation, but you need to give it a pretty well defined plan. on it's own it makes a lot of assumptions that are incorrect and thus not reliable to use as the main model. also, i'd love to see more rendered components inside of the chat window that can help with visualizing complex architecture and abstractions.” [source](https://twitter.com/999042490626797568/status/2096310737442328595)
  - Praise, Google Antigravity, r/google_antigravity, 2026-09-03: “its crazy good, always remember when you prompt and you are not confident or clear with your intentions, always let the model (flash 3.8) ask you questions. this guarantees you are in line with the thing you want to do.” [source](https://www.reddit.com/r/google_antigravity/comments/1w6ek6h/a_question_for_people_who_already_tested_flash_38/p7mgdye/)

- **Devin (weaker)**. Devin's posts are all complaints, focused on a cheap executor guessing through plan ambiguities instead of stopping to replan.
  Critics say a handoff from planning to execution looks finished until a guessed ambiguity surfaces. Users ask for clarifying questions mid-run and before coding starts. The jabs about needing a clearer ticket point to the same gap between ambiguity and a good question.
  Evidence:
  - Complaint, Devin, @cognition, 2026-09-12: “@cognition the planning vs execution split is the right cost model, but the interesting failure is the handoff. a cheap executor looks finished right up until the plan had an ambiguity it should not have guessed. the harness is only as good as the moment it knows to stop and replan.” [source](https://twitter.com/144120499/status/2098840010975822131)
  - Complaint, Devin, @cognition, 2026-09-25: “@cognition devin just hit the billion dollar mark and still probably asks for a clearer ticket 😂” [source](https://twitter.com/1699417980155637761/status/2103506346553610593)

- **Cursor (weaker)**. Cursor complaints center on models that leap to wrong interpretations and on follow-up questions that freeze an approved plan mid-execution.
  One user compares a model to a genie that exploits any interpretation left uncovered. Another says follow-up questions after a plan starts hang the whole run. The praise that exists credits forced up-front questions and planning mode, not default behavior.
  Evidence:
  - Complaint, Cursor, @cursor_ai, 2026-09-07: “@xai @cursor_ai been a long weekend. grok 4.6 on cursor. god it jumps to wrong conclusions so often. and it's like a dark genie. where you make a wish and if you don't cover every single possible wrong interpretation of the wish, it finds a way to f*** you.” [source](https://twitter.com/485979159/status/2097035883442769928)
  - Complaint, Cursor, @cursor_ai, 2026-09-27: “@cursor_ai should reconsider cx of follow up questions after execution of approved plan started. it hanged entire authonomy.” [source](https://twitter.com/255140211/status/2104192810089943545)
  - Praise, Cursor, @cursor_ai, 2026-09-02: “@cursor_ai verifying its own work is the easy half. the loop that matters is the one that can throw the work away. i've been faster since the agent has to ask 3 questions before it writes a line. fewer green checks. fewer seventh versions of something nobody asked for.” [source](https://twitter.com/2073295850491752448/status/2095090463149810118)

### Fine print

- Only OpenAI Codex and Claude Code have enough posts here to rate. Every other agent rests on a handful of posts.
- Many posts describe the underlying model rather than the agent harness, so behavior may change with model choice.
- Requests split evenly between asking more and asking less, so no single direction satisfies every user.

## Top requests

What users ask to add or change, most asked first. 43 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Author-weeks | Posts | Agents (author-weeks) |
|---|---|---|---|---|
| 1 | Clarifying questions mid-run outside plan mode | 8 | 8 | OpenAI Codex 4, Devin 2, Claude Code 1, Cursor 1 |
| 2 | Ask clarifying questions instead of guessing | 7 | 8 | Claude Code 4, OpenAI Codex 1, GitHub Copilot 1, OpenCode 1 |
| 3 | Fewer unnecessary or repetitive clarifying questions | 7 | 7 | OpenAI Codex 4, Claude Code 2, OpenCode 1 |
| 4 | Dedicated UI tool for clarifying questions | 6 | 6 | OpenAI Codex 3, Amp 1, Cline 1, Devin 1 |
| 5 | Stop and ask when stuck or blocked | 6 | 6 | Claude Code 4, Devin 1, OpenCode 1 |
| 6 | Clarifying questions before starting to code | 5 | 5 | OpenAI Codex 2, Claude Code 1, Devin 1, OpenCode 1 |
| 7 | More meaningful clarifying questions | 2 | 2 | Claude Code 1, Cursor 1 |

### 1. Clarifying questions mid-run outside plan mode

- OpenAI Codex, 2026-09-25, X search: OpenAI Codex, Codex CLI, Codex app (X): “why does the codex app ask questions in non-plan mode without blocking? there's no time to answer before it just goes with the default recommendation. so what the hell is the point of asking?” [source](https://twitter.com/1901648131630280704/status/2103484594670784986)
- OpenAI Codex, 2026-09-09, r/codex (Reddit): “they need to be designed to pause regularly and ask for feedback/guidance in a hitl setup, not autonomous drones that compound their mistakes, hallucinations, and assumptions the longer they run.” [source](https://www.reddit.com/r/codex/comments/1wah1jk/opinion_astra_is_overhyped/p8sj541/)
- OpenAI Codex, 2026-09-07, r/codex (Reddit): “it's not an ambiguity though i think i did misuse /goal a bit, i said i'd verify the changes manually (quicker than having it write like 500 unit tests) and then got it stuck in a loop since it can't ask questions in goal mode (i think)” [source](https://www.reddit.com/r/codex/comments/1w9my9c/finally_figured_out_why_goal_used_100_of_my_usage/p8bjt1c/)

### 2. Ask clarifying questions instead of guessing

- OpenAI Codex, 2026-09-24, r/codex (Reddit): “sure if its so smart, it could have asked specifics like claude does, instead of working 4 fucking hours on a simple portfolio page and generating slop. gemini did way better with exact same prompt in 7min btw.” [source](https://www.reddit.com/r/codex/comments/1wpb7an/its_even_worse_than_gemini_flash_at_this_point/pbuiokl/)
- Claude Code, 2026-09-21, r/ClaudeCode (Reddit): “nah, sometimes it really goes off the completely wrong end. what i would like is for it to come up with an answer only if it has a strong confidence level. i should be allowed to configure it such that below that threshold it either asks me questions for details that could help, or just tell me it doesn't know. coming up with wrong answers is way worse.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wm4ncx/im_afraid_to_use_opus_5/pb90c7j/)
- Claude Code, 2026-09-11, r/ClaudeCode (Reddit): “can we trade ? getting claude to ask a question is like trying to get a toddler to eat his veggies. it will do litterally anything (search the web, launch explore agents, write probe scripts, hallucinate something) rather than ask something i could answer in two sentences.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wd2738/claude_dynamic_workflow_is_very_cool/p92vu0r/)

### 3. Fewer unnecessary or repetitive clarifying questions

- OpenAI Codex, 2026-09-09, r/codex (Reddit): “i am using plan mode on sol high, and somehow the model questions started glitching. not only it is stuck on a loop asking the same questions, it started asking bogus stuff. just see these images. it's between funny and sad, as it feels sol has developed dementia. <strict_link>” [source](https://www.reddit.com/r/codex/comments/1wbhej5/what/)
- OpenAI Codex, 2026-09-08, r/codex (Reddit): “it does that a lot, this model was trained to be highly iterative. too much in fact, it stops short ov everything and asks for clarification or in situations like this throws up it's hands and says "i stopped gotta try something else so ask me if you want to do that" that" its like of course i do! in openais documentation they label this behavior as "highly collaborative" more like annoying.” [source](https://www.reddit.com/r/codex/comments/1wal4o7/well_this_one_is_new/p8l38iz/)
- OpenAI Codex, 2026-09-08, r/codex (Reddit): “it also stops to ask completely unnecessary questions where the task is clear and straightforward. sol doesn't do this, so i simply switched back to using sol. astra's intelligence level is not better than sol's.” [source](https://www.reddit.com/r/codex/comments/1wah1jk/opinion_astra_is_overhyped/p8if1vu/)

### 4. Dedicated UI tool for clarifying questions

- Amp, 2026-09-25, @AmpCode (X): “@ampcode curious why amp doesn't have any question/answer tools the model can use. having instances where the modal outputs a big explanation then in the last sentence: may i do that? i sometimes miss that it's asking at all! a ui question tool would make that al ot more obvious.” [source](https://twitter.com/5444392/status/2103600382303944750)
- OpenAI Codex, 2026-09-20, X search: OpenAI Codex, Codex CLI, Codex app (X): “@lmdev25 @thsottiaux i feel codex should simply ask me the question, instead of making me do keyboard gymnastics before being able to see the question. the screenshot above is codex cli. meanwhile, the codex app does not show todo/tasks list and questions. what's with that?” [source](https://twitter.com/449598591/status/2101664107875426492)
- Cline, 2026-09-19, @cline (X): “@cline why can the model only choose from the options it provides when calling the ask question tool, and cannot enter other options on my own?” [source](https://twitter.com/721547404479111170/status/2101290071311945767)

### 5. Stop and ask when stuck or blocked

- OpenCode, 2026-09-25, r/opencode (Reddit): “i dont like the fact that it keeps guessing rather than look for answer and if it doesnt find them to consult me” [source](https://www.reddit.com/r/opencode/comments/1wocn5s/space_bunny_thoughts/pc1490a/)
- Claude Code, 2026-09-24, r/ClaudeCode (Reddit): “"stop and ask my input before doing work arounds" otherwise it may be unable to access some api definition and either start a full cyber security army to get it anyway or manually reimplement the whole thing. while it would have been 5 seconds copy/paste for me to put the file in a place he is allowed to read.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wp23t8/what_rule_in_your_claudemd_clearly_has_a_backstory/pbs0ky8/)
- Claude Code, 2026-09-22, r/ClaudeCode (Reddit): “the real opus 5.5 and reset was the friends we made along the way. well, it was the friend my opus 5 agent made in my codebase for some reason while spiraling for 45 minutes to find a workaround for a simple problem that had a simple answer if it just would have asked, but same thing right?” [source](https://www.reddit.com/r/ClaudeCode/comments/1wnaoox/where_is_opus_55_and_the_reset/pbdhplp/)

### 6. Clarifying questions before starting to code

- OpenAI Codex, 2026-09-24, r/codex (Reddit): “stop dictating and start collaborating. treat these frontier models like they are expert consultants. the models are better at prompting themselves than you are. have it ask you questions, understand your expectations, what “done” looks like, etc.” [source](https://www.reddit.com/r/codex/comments/1woo8ku/astra_extra_high_in_blender_1010_i_cant_tell_em/pbs5gum/)
- OpenCode, 2026-09-23, r/AI_Agents (Reddit): “cli more simple and easy for me. i don't need so many features. opencode go is great for vibe coding just use expensive models to design things and for complex tasks only. because grok 4.6 burned my 20% tokens in like 5 min. always ask questions in plan mode before build” [source](https://www.reddit.com/r/AI_Agents/comments/1wo09im/any_coding_workflow_advice_for_existent_codebase/pbnauww/)
- Devin, 2026-09-12, @cognition (X): “@cognition it's quite clever to put devin in the phone; the chatting step will be much more natural. my first reaction, however, is: when it receives a phrase like "make a quick change," can it first confirm which repo and which environment to change (laughs)?” [source](https://twitter.com/2259799350/status/2098596266544374012)

### 7. More meaningful clarifying questions

- Claude Code, 2026-09-10, r/ClaudeCode (Reddit): “i agree and i think we should add a hook designed to ping the agent « hey just think about the clarification you are asking to user, don’t ask shitty meaningless options » i don’t know if it’s possible, i’ll try to do it later this day” [source](https://www.reddit.com/r/ClaudeCode/comments/1wcr862/claude_offering_pointless_options/p901u4e/)
- Cursor, 2026-09-02, @cursor_ai (X): “the proactive questioning alignment feature of cursor is very poor. when a user inputs a discussion question, there is no output, and it just asks you abcd, defaulting to recommend a, but you have no idea why? because there are only options, no answers. as a result, a discussion question is turned into a multiple-choice question by cursor. isn't this silly? please optimize this feature quickly @cursor_ai.” [source](https://twitter.com/703883942995165184/status/2095032296357376056)

## Every agent

| Agent | Overall rank | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|---|
| [Claude Code](https://feedbackbench.com/agents/claude-code.md) | 1 | Typical | 0.502 | 0.477–0.530 | 78 | 31 | 47 |
| [OpenAI Codex](https://feedbackbench.com/agents/codex.md) | 2 | Typical | 0.495 | 0.470–0.519 | 94 | 35 | 59 |
| [OpenCode](https://feedbackbench.com/agents/opencode.md) | 3 | Too few posts | – | – | 15 | 7 | 8 |
| [Google Antigravity](https://feedbackbench.com/agents/antigravity.md) | =5 | Too few posts | – | – | 14 | 7 | 7 |
| [Cursor](https://feedbackbench.com/agents/cursor.md) | 4 | Too few posts | – | – | 10 | 2 | 8 |
| [Pi](https://feedbackbench.com/agents/pi.md) | 7 | Too few posts | – | – | 8 | 4 | 4 |
| [Devin](https://feedbackbench.com/agents/devin.md) | =5 | Too few posts | – | – | 4 | 0 | 4 |
| [Cline](https://feedbackbench.com/agents/cline.md) | =8 | Too few posts | – | – | 1 | 0 | 1 |
| [Factory](https://feedbackbench.com/agents/factory.md) | =11 | Too few posts | – | – | 1 | 1 | 0 |
| [Amp](https://feedbackbench.com/agents/amp.md) | =11 | Too few posts | – | – | 1 | 0 | 1 |
| [GitHub Copilot](https://feedbackbench.com/agents/copilot.md) | =8 | Too few posts | – | – | 0 | 0 | 0 |
| [Zed](https://feedbackbench.com/agents/zed.md) | =8 | Too few posts | – | – | 0 | 0 | 0 |
| [Kiro](https://feedbackbench.com/agents/kiro.md) | 13 | Too few posts | – | – | 0 | 0 | 0 |
| [Conductor](https://feedbackbench.com/agents/conductor.md) | 14 | Too few posts | – | – | 0 | 0 | 0 |
| [Warp](https://feedbackbench.com/agents/warp.md) | 15 | Too few posts | – | – | 0 | 0 | 0 |
| [Grok Build](https://feedbackbench.com/agents/grok-build.md) | 16 | Too few posts | – | – | 0 | 0 | 0 |
| [Augment Code](https://feedbackbench.com/agents/augment.md) | 17 | Too few posts | – | – | 0 | 0 | 0 |

## Posts

Receipts rule: The 5 most recent praise and complaint posts per area (first 700 characters) and 3 per criterion (first 450 characters).

### Claude Code

- Praise, 2026-09-27, r/ClaudeCode (Reddit): “skills has been really useful for me to inject business related knowledge into the context. /handoff really good for managing context /grill-me has been really useful to get the model to write the exact specifications i am looking for. people are not very precise when speaking to an agent, and often underspecify their requirements and end up getting upset when the agent end up doing something else.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wowilt/i_still_dont_understand_this_agentic_workflow/pcb8zwi/)
- Praise, 2026-09-27, r/ClaudeAI (Reddit): “this is the move. i've been doing the same thing with claude code when planning out execution steps for my agent stack — letting it interrogate me about failure modes instead of me guessing upfront. the part that always gets me is it'll still end the interview with "would you like me to provide a timeline?" like it didn't just extract 40 minutes of context from my brain. that's on me, honestly.” [source](https://www.reddit.com/r/ClaudeAI/comments/1wrrl96/the_best_thing_i_do_all_week_is_have_it_interview/pcf8brh/)
- Praise, 2026-09-26, @ClaudeDevs (X): “#antrophic has cooked. 🔥🔥🔥..wow i tested #claudecode and #fable5 to see how well they perform in technical drawing when claude code is connected to autodesk fusion. i had fable 5 trace one of our products. it worked, and i was 50% satisfied, but it took a lot of effort. now opus 5.5 has done it flawlessly, all on its own, without me. it just asked four questions beforehand, and opus 5.5 then took the rest from the images on our website and, as a” [source](https://twitter.com/1690767898728431616/status/2103717367679222205)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “worth checking how many turns are just it asking you questions. a "don't ask unless blocked" rule in claude.md cuts a lot of that.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wropen/i_keep_hitting_usage_limits_on_20x_plan_have/pceof2h/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “both picked a different kind of project without asking - a coded workflow instead of the flow, and no file ever got created. whether it saw the choice and assumed, or never saw it, i didn't dig into the traces to tell.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqpedm/i_built_a_claude_code_plugin_that_uses/pcfmj1t/)
- Complaint, 2026-09-26, r/ClaudeCode (Reddit): “&#x200b; claude code is very good at filling in missing requirements. sometimes too good. if i say: «fix exports that sometimes return old data.» claude can inspect the repo and figure out that exports currently use a 15-minute cache. but the code can't tell it whether i actually want: \- manual exports to always be fresh \- cached data to remain acceptable \- scheduled exports to behave differently \- stale data or an error when the source fails” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqpedm/i_built_a_claude_code_plugin_that_uses/)

### OpenAI Codex

- Praise, 2026-09-27, r/codex (Reddit): “so far for me that seems true. things that really show: * better code quality. it thinks about the implementation more rather than shooting for a quick patch * for non-code tasks, it spends more time thinking. for example, a 3d model workflow: astra took like 5 pictures and figured "meh probably good enough". while opus took at least \~30 at every possible angle, and fixed small mistakes here and there. astra was significantly faster, but used mo” [source](https://www.reddit.com/r/codex/comments/1wrcc9j/astra_minor_astra_61_and_devday_we_see_50/pch16yo/)
- Praise, 2026-09-26, r/ChatGPTPro (Reddit): “thanks for doing the comparison i've been too satisfied with chatgpt to try. i started with chatgpt's voice mode and have been so satisfied, it feels so natural, that i'm reluctant to even test alternatives. i just can't imagine how they could possibly be better. i've not noticed the hallucination you mentioned. i'm using sol and it's been fine. i've found it most useful in long dialogue. for example, i use the grill-with-docs skill a lot. using” [source](https://www.reddit.com/r/ChatGPTPro/comments/1wpq2gj/chatgpt_voice_vs_claude_voice_mode_speed_fluency/pc481ab/)
- Praise, 2026-09-25, r/codex (Reddit): “honestly, independently of reddit i made the same conclusion about the overengineering thing. now granted i work on a totally different project now, i must say that sol 5.6 has been a really really really good model for at least the last month, whereas i preferred terra in the ‘overengineering era’. based on my own experience: i asked sol to come up with an approval system that makes it both self contained and check whether the approval was lega” [source](https://www.reddit.com/r/codex/comments/1wpasz8/anybody_not_having_a_bad_time_with_gpt_6_sol/pbxk7oj/)
- Complaint, 2026-09-26, X search: OpenAI Codex, Codex CLI, Codex app (X): “sydney sweeney reveals in an interview that her current favourite model is opus 5.5 and this is happening after a long time. "i had been mainly using the codex app for the last 4 months or so. i love the app, and gpt 5.6 sol proved to be a great model for pretty much everything. then they dropped astra, which was great but an undercooked model when it came to eagerness and instruction following. it would stop in the middle, ask follow-up question” [source](https://twitter.com/1446445479068241923/status/2103737712679547105)
- Complaint, 2026-09-26, r/ClaudeAI (Reddit): “i experimented with building features with both at the same time. claude was faster, but codex spend more time on security, structure etc. claude was faster because it just did what i asked, codex was more mature and thought of naming conventions and such. claude asks more, codex presumes more. it’s a matter of preference i think because both got the job done and the differences in my case would have iterated out anyway.” [source](https://www.reddit.com/r/ClaudeAI/comments/1wqgdn8/i_switched_to_chatgpt_pro_last_month_and_regret/pc3xzuo/)
- Complaint, 2026-09-25, r/codex (Reddit): “i've found their insight equivalent but with planning i have to prod astra to give me questions and ambiguity to resolve and claude just appends them to the end of the working plan natively” [source](https://www.reddit.com/r/codex/comments/1wppkog/new_tibo_tweet_about_devday/pbz1aom/)

### OpenCode

- Praise, 2026-09-26, r/opencode (Reddit): “i've been using it for a few days in opencode, previously i was using glm 5.2 in claude code cli before my grandfathered account expired, and i'm finding it way better at fixing old dead tests and refactoring code. since it's free at the moment, i'm trying to maximize all of the grunt work that was eating my quota up and was providing that much value. strip away some of that technical debt that 18 months of spec driven vibe coding had introduced” [source](https://www.reddit.com/r/opencode/comments/1wocn5s/space_bunny_thoughts/pc8hu4o/)
- Praise, 2026-09-16, r/opencodeCLI (Reddit): “using it from yesterday so i may have not encountered problems yet, but for me it did better than opus. i recently canceled my claude subscription and started using opencode go. with claude i always had problems with it ignoring my dependencies and doing everything manually. for example, i use prisma orm for the database, and claude would always try to migrate manually, fucking up the hash that prisma uses for integrity. big pickle doesn't do tha” [source](https://www.reddit.com/r/opencodeCLI/comments/1qr1jm6/anyone_tried_the_big_pickle_model_on_opencode/pa4ehm6/)
- Praise, 2026-09-15, r/PiCodingAgent (Reddit): “i tried with gemma4 e4b and it's really bad. used same system.md with opencode. it didnt follow the the rules, didn't ask for more info. opencode follows the rules and asks so idk” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wcj9bg/who_uses_pi_what_do_you_like_about_it/p9v0o7l/)
- Complaint, 2026-09-25, r/opencode (Reddit): “i dont like the fact that it keeps guessing rather than look for answer and if it doesnt find them to consult me” [source](https://www.reddit.com/r/opencode/comments/1wocn5s/space_bunny_thoughts/pc1490a/)
- Complaint, 2026-09-24, r/opencodeCLI (Reddit): “i understand, didn't you switch between free and standard? the free/contributor is the one that goes up to xhigh and the standard is the one that does have reasoning up to max. i see, i haven't used ds or mimo, in what i've been working on kotlin, rust, astro, nextjs muse spark 1.3 has been excellent in everything, its intelligence is very good, it is very similar to opus for structural planning and execution. the only point that is not so poli” [source](https://www.reddit.com/r/opencodeCLI/comments/1wopbtv/what_muse_spark_14_contributor_is_already_here/pbpl2fw/)
- Complaint, 2026-09-23, r/opencode (Reddit): “it's asking me a lot of questions it should be able to figure out on its own, like how to syntax check a js file - making me choose between node or "enter your own command"” [source](https://www.reddit.com/r/opencode/comments/1wocn5s/space_bunny_thoughts/pbmxkel/)

### Google Antigravity

- Praise, 2026-09-26, @antigravity (X): “@thsottiaux @antigravity i like how you guys have allowed us to basically recreate grill-with-docs by just pointing it to docs, telling codex our idea, and saying "ask me questions about this."” [source](https://twitter.com/14838410/status/2103955941670420595)
- Praise, 2026-09-26, @antigravity (X): “@antigravity an agent that asks before it vibes” [source](https://twitter.com/1975526768112185344/status/2103964727365742756)
- Praise, 2026-09-26, @antigravity (X): “@thsottiaux @antigravity idk why you want to remove it but it save me so many time by confirming what i actually want instead of guessing which fk it up many times.” [source](https://twitter.com/1114864978283171840/status/2103981847667773788)
- Complaint, 2026-09-26, @antigravity (X): “gemini 3.1 pro on @antigravity is really bad at agentic tasks, especially when working with blender mcp and unreal engine mcp. it also has this annoying habit of asking a ton of questions, even when you’ve given antigravity full turbo access and all the necessary permissions.” [source](https://twitter.com/1293449962974404608/status/2103847108831064400)
- Complaint, 2026-09-26, @antigravity (X): “antigravity is the only one where i have to use "plan mode"; without it, gemini 3.8 might drag me off somewhere without me even realizing it. while claude and chatgpt automatically ask for clarification on anything uncertain, antigravity just runs with it—often finishing the task before i’ve even fully grasped what happened.” [source](https://twitter.com/1533997728858243072/status/2103895076707697079)
- Complaint, 2026-09-22, r/google_antigravity (Reddit): “i get that "please just fix it" isn't an ideal prompt, but with the old behaviour it would just ask for clarification. it never used to trigger an endless 3-minute file-scanning loop.” [source](https://www.reddit.com/r/google_antigravity/comments/1wnfcoo/has_antigravity_started_aggressively_scanning/pberb4t/)

### Cursor

- Praise, 2026-09-10, r/cursor (Reddit): “opus 4.6 is my daily driver in planning mode and to execute the plan, then when the task is complete if i have any follows ups i use got 5.2, not codex version. i find 5.2 works for longer and doesn't stop to ask you questions as much saving on time and context by skipping useless uodate messages that wait for a response. honestly if you're not using planning mode you should try it. been using cursor for about 2yrs and i wish i'd started using it” [source](https://www.reddit.com/r/cursor/comments/1wc95mb/i_am_tired_of_handholding_composer_25_what_are/p8w6c0a/)
- Praise, 2026-09-02, @cursor_ai (X): “@cursor_ai verifying its own work is the easy half. the loop that matters is the one that can throw the work away. i've been faster since the agent has to ask 3 questions before it writes a line. fewer green checks. fewer seventh versions of something nobody asked for.” [source](https://twitter.com/2073295850491752448/status/2095090463149810118)
- Complaint, 2026-09-27, @cursor_ai (X): “@cursor_ai should reconsider cx of follow up questions after execution of approved plan started. it hanged entire authonomy.” [source](https://twitter.com/255140211/status/2104192810089943545)
- Complaint, 2026-09-15, r/cursor (Reddit): “gave it a pretty straightforward one-shot task: pull some legacy wordpress content into our production supabase cms using our existing script: sync\_legacy\_content.js --publish the script needed the supabase url + service role key. nothing unusual. except the cloud agent didn’t have the service role key. at that point, the correct thing would have been: **“i’m missing the key. please provide it.”** instead, it decided to get creative: it first” [source](https://www.reddit.com/r/cursor/comments/1wh07hl/my_cursor_cloud_agent_burned_through_my_monthly/)
- Complaint, 2026-09-10, @cursor_ai (X): “@grok @cursor_ai @bot you also can't distinguish between suggestions/feedback and someone actually needing help. hopefully you get better at that!” [source](https://twitter.com/48508624/status/2098053477851123746)

### Pi

- Praise, 2026-09-24, @pidotdev (X): “@brendan_j_ryan @pidotdev @exaailabs i also reported the email finding job failure. that was smooth way of getting in reports btw. agent just asked me if i want to do that.” [source](https://twitter.com/19357555/status/2103171121969553484)
- Praise, 2026-09-20, r/PiCodingAgent (Reddit): “i would add two minimal `must haves` to that list: - pi-vcc (instant compaction), and if you find you lose too much knowledge, replace it with pi-blackhole, which includes vcc, but also an observational memory layer that allows better recall - pi ask user (nice friendly tui with multi-stage and comments)” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wlpj30/lord_forgive_me_for_the_time_i_wasted/pb1mwf0/)
- Praise, 2026-09-08, r/PiCodingAgent (Reddit): “basically it all started with qwen3.8 reminding me of my college years when one of my girlfriends would smoke weed and try to clean the apartment and i remember her hmm, wait, what is this etc. on watching it very closely, i realized it was complaining a lot about either missing tools, not understanding the tool usage, response from the tools being confusing. so, i literally asked it "what tools do you see are missing in this harness and how can” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wb0ehv/do_lsps_improve_pi_agents/p8mkk1j/)
- Complaint, 2026-09-15, r/PiCodingAgent (Reddit): “it fails when the ai tries to ask me a question (via the `ask_user_question` tool). immediately "user declined to answer questions"” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wgy56q/piui_a_month_later_harder_better_faster_stronger/p9zbmpi/)
- Complaint, 2026-09-11, @pidotdev (X): “@howaboua @pidotdev 74% fewer prompt tokens and it will still ask which folder we're in before touching the file.” [source](https://twitter.com/1446058878656032768/status/2098403417521545603)
- Complaint, 2026-09-02, r/PiCodingAgent (Reddit): “okay so i have to guess the whole context on my own. thanks!” [source](https://www.reddit.com/r/PiCodingAgent/comments/1vy9edt/ive_just_installed_feynman_and_it_seems_to_be_an/p7cw75x/)

### Devin

- Complaint, 2026-09-25, @cognition (X): “@cognition devin just hit the billion dollar mark and still probably asks for a clearer ticket 😂” [source](https://twitter.com/1699417980155637761/status/2103506346553610593)
- Complaint, 2026-09-16, @cognition (X): “windsurf was the first agent coding tool i got my hands on, and its entry price of $10 was cheaper than cursor at the time, making it my first tool. recently, its acquirer @devindesktop @cognition launched the swe-2 model, which excites me a lot. i picked up my account that i started renewing in 2024 to experience it, but i found many frustrating issues here: first and foremost is the work of the subagent. as a cross-model harness, the current s” [source](https://twitter.com/1926600710122004480/status/2100294237749489728)
- Complaint, 2026-09-12, @cognition (X): “@cognition the planning vs execution split is the right cost model, but the interesting failure is the handoff. a cheap executor looks finished right up until the plan had an ambiguity it should not have guessed. the harness is only as good as the moment it knows to stop and replan.” [source](https://twitter.com/144120499/status/2098840010975822131)

### Cline

- Complaint, 2026-09-19, @cline (X): “@cline why can the model only choose from the options it provides when calling the ask question tool, and cannot enter other options on my own?” [source](https://twitter.com/721547404479111170/status/2101290071311945767)

### Factory

- Praise, 2026-09-27, @droid (X): “first time using anthropic models (opus 5.5). i use @droid. i gave it a task to design a landing page for my current project, and it asked questions i have never seen from an agent before. i hope the result comes out good, but so far, it's really impressive and i might not use openai models, unless they actually have a god response. also can only recommend the factory app. it's genuinely brilliant, and the fact that i can use my own laptop or hom” [source](https://twitter.com/1821640621347495936/status/2104230757883388046)

### Amp

- Complaint, 2026-09-25, @AmpCode (X): “@ampcode curious why amp doesn't have any question/answer tools the model can use. having instances where the modal outputs a big explanation then in the last sentence: may i do that? i sometimes miss that it's asking at all! a ui question tool would make that al ot more obvious.” [source](https://twitter.com/5444392/status/2103600382303944750)
