# Stops mid-task or answers instead of acting (`work.premature_stop`)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/criterion/work.premature_stop

Area: [Doing the work](https://feedbackbench.com/criteria/work.md)

**Definition.** The agent halts before finishing, or replies with explanation or a summary instead of performing the action.

**Boundary.** Not this: see [Claims work is done or fixed when it is not](https://feedbackbench.com/criteria/verify.false_completion.md) for stopping while claiming the work is done. Not this: see [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) for stops caused by quota.

Rated author-weeks, all agents: 408. Complaint share: 85%.

## The brief

Written by Claude Opus 5.5 from 38 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**Agents narrate the next step, then stop. Codex users feel it most.**

TL;DR:

- OpenAI Codex draws most complaints, mainly announced fixes that never run and repeated nudges to continue.
- Claude Code rates better than peers, helped by a graceful wrap-up that replaces mid-edit cutoffs.
- Silent stops on long tasks hit Devin, Antigravity, Amp and Cursor, but their post counts are thin.

In plain terms: You hand off a task, walk away, and come back to an empty diff. The agent promised the next step without taking it, paused for permission it already had, or quietly halted partway through.

### How it breaks

- **Says it will act, then does nothing** ([Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md)). The most common failure is an agent agreeing to do the work, apologising for not doing it, and still not doing it.
  Posts describe a loop. The agent acknowledges the problem, says it will fix it, and ends its turn. Users reply 'ok do it' several times before any change lands. One user watched a model admit it should have acted, promise to act, then make no attempt at all. Switching models finished the job. Claude Code users quote similar guidance to nudge the model past 'next, I'll' phrasing. Four author-weeks explicitly ask agents to perform announced actions instead of describing them.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-06: “yes, happened to me quite a few times, it said ok, and stopped. i was like "ok do it?" and had some few back to back ok do its until it did the work.” [source](https://www.reddit.com/r/codex/comments/1w8t3tu/astra_is_lazy/p85hevp/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-17: “terra worse than luna. idk how people use that model. i'll look up a conversation i had with terra and post it here once i'm on my laptop but no model has behave so bad since gpt 4. i gave terra medium later bumped to high a task to use chrome desktop app to fill in my expenses. and it stopped. and then stopped after doing one more day after which it said it'll do it and it did nothing. like literally saying "yea my bad i should have done it." well then fix it "i'll go and fix it" does nothing. doesn't think doesn't even try to connect to chrome making it a possibility that it's a failure to connect. switched to luna max and it did the job.” [source](https://www.reddit.com/r/codex/comments/1widk1m/tier_list_according_to_me/pa9spu0/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-02: “> on complex asynchronous workloads, though, nudge it not to end its turn before the work is done. without the nudge, the model sometimes describes what it would do next instead of doing it ("next, i'll …") or stops to ask permission for a step the original request already covered ("shall i apply this?"). users have to reply "continue" or "go ahead," which suits pair programming and other human-in-the-loop work but doesn't use the model's full long-horizon capability. this is a known issue with anthropic models, had the issue with opus and sonnet too.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w4qs4p/warning_read_the_fable_51_docs_fable_51_is/p7e2e3g/)
  - Complaint, Amp, @AmpCode, 2026-09-07: “tomorrow i would try my tuned hard mode in @ampcode , cz default high mode somehow doesnt passing the vibe + having bad exp with “you are right” “replying instead of start working”, may be its astra itself thats doing this dumb play <strict_link>” [source](https://twitter.com/1248119942194421760/status/2096764065334923682)

- **Pauses for permission it already had** ([Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md)). Agents stop to ask whether they should continue, which turns unattended runs into hours of idle waiting.
  Users delegate a task, step away, and find the agent stalled on a confirmation prompt. Some report the agent asking whether to implement fixes only after burning a large share of their allowance on analysis. Copilot users say certain models pause often to ask whether to keep going. The cost is wasted wall-clock time on top of tokens. 'No pauses asking user to continue' is a recurring request, led by OpenAI Codex users.
  Evidence:
  - Complaint, Cursor, @cursor_ai, 2026-09-08: “@nathantippy @cursor_ai it halted after you walked away, then let you spend hours feeling like you delegated. which hurt worse when you sat back down: the empty diff, or realizing it spent the whole time waiting for a single click?” [source](https://twitter.com/2085715088267485184/status/2097324355177038113)
  - Complaint, GitHub Copilot, r/GithubCopilot, 2026-09-03: “you are having issues with run on their own? i find i have better luck using fable as the main, opus always seems to pause a lot and ask if it should keep going” [source](https://www.reddit.com/r/GithubCopilot/comments/1w5i0sx/one_msft_employee_reaches_100k_mo_in_token/p7jro2t/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-06: “glad you fanboys love it - all i see is token burn to the max and the results are crap. most end with it asking if it’s ok to implement the fixes after half the weekly credit allowance is burned. then say yes - down to 0%. not done yet? too bad!!! it stops mid work now too. fantastic” [source](https://www.reddit.com/r/codex/comments/1w92bbt/loving_astra/p87oxm6/)

- **Silent cold stops on long tasks** ([Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md)). Long or multi-step runs end partway with no message, no summary and no error to explain why.
  Devin users report constant mid-work stops on a recent desktop version, and others say runs die around the hour mark. Antigravity users describe intermittent termination on large tasks across models and fresh conversations. Amp users cannot tell whether model or harness is at fault. A Claude Code user saw a long task end at item 18 of 23. Some posts note the agent ignored explicit instructions to keep going.
  Evidence:
  - Complaint, Devin, @cognition, 2026-09-17: “@cognition hi glm 5.3 flash on devin is constantly stopping mid work. on latest desktop version. it wasnt like this before. pls fix this” [source](https://twitter.com/1392605672794034182/status/2100696367614300178)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-01: “i’m currently using antigravity 2.0 for a fairly large coding project, but it sometimes terminates on long/complex tasks. i’ve tested multiple models and new conversations, and the issue is intermittent. would you recommend antigravity cli over 2.0 for long coding tasks, and is cli generally more stable?” [source](https://www.reddit.com/r/google_antigravity/comments/1w4ikza/antigravity_20_vs_cli_for_large_projects/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-23: “opus 5.5 unfortunately left a lot of adherence journal entries for me. that is a very bad sign. i am currently analyzing how many of them are hallucinated and how many actually match the existing records. it also ended a very long task even though it had only reached item 18 out of 23, which i found strange as well. i will keep monitoring it. my first impression is mixed at the moment. the additional tasks it created for me are unusual. i am going to analyze the journal cases and the other open tasks it created during its work and then i will see whether it is hallucinating again as opus 5. <strict_link>” [source](https://www.reddit.com/r/ClaudeCode/comments/1wnsn6a/opus_55_is_great/pbhsw21/)
  - Complaint, Amp, @AmpCode, 2026-09-05: “astra keeps stopping before finishing its work, not sure if it’s a model or harness thing @ampcode” [source](https://twitter.com/1305405391/status/2096247762052641010)

- **Explains the change instead of writing it** ([Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md)). Some setups never get past reading and explaining, so users receive instructions or code snippets rather than edited files.
  A Pi user reports the agent reading endlessly without ever writing code to files. A Cline user running local models says the agent switched to telling them how to do the task, and broke again whenever auto-execute was enabled. A Kiro user describes questions abandoned without a word. These reports are sparse and tied to specific model and harness pairings, but the symptom is the same. The agent answers when it should act.
  Evidence:
  - Complaint, Pi, @pidotdev, 2026-09-18: “@melissapan @howaboua @pidotdev nice work. i have issues with fable &amp; pi, where pi keeps reading and never writes the code to files. not sure of the problem though.” [source](https://twitter.com/1613070948/status/2100965181912482194)
  - Complaint, Cline, r/CLine, 2026-09-26: “hello everyone , i just started using vscode +cline , just for fun , messing around with unity scripts and stuff , and all was great for a few days , during 1 task , the power went down , and after i turned on my pc i had the following problem , he just stopped executing tasks , most of the time just telling me how to do it , and sometimes replying just in code , i ve been trying to troubleshoot it for 2 days now , and this is what i found. i use different types of qwen loccally , after i reinstall any qwen , it does work if i approve mannually , but if i check the auto execute box , it breaks again…..been using it just for fun , and i dont have any experience in this sort of things . what can i do to fix ? thanks” [source](https://www.reddit.com/r/CLine/comments/1wqm2kn/vscode_and_cline/)
  - Complaint, Kiro, r/kiroIDE, 2026-09-12: “sometimes i ask something and it doesn't do anything. like my question was abandoned promptly. and when it's waiting for another service/event which never happens kiro wouldn't say a word” [source](https://www.reddit.com/r/kiroIDE/comments/1wefi0t/lately_kiro_has_been_just_hanging/p9e4h1d/)

- **Graceful wrap-up beats dying mid-edit** ([Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md)). Users strongly prefer an agent that tidies up and leaves a clean stopping point over one cut off mid-edit.
  Claude Code users welcomed a change that lets in-flight work wrap up rather than halt abruptly. They describe past cutoffs leaving half-renamed codebases and interrupted plans. Codex users ask for the same wrap-up behaviour and say Codex has resumed letting tasks finish. Requests for a handoff summary when stopping come only from Claude Code users, a sign that expectations have moved from 'don't stop' to 'stop cleanly'.
  Evidence:
  - Praise, Claude Code, @ClaudeDevs, 2026-09-25: “@claudedevs the graceful stop is the right fix. mid-edit cutoffs were the worst way to learn about the limit” [source](https://twitter.com/822848256690548736/status/2103595577472700628)
  - Praise, Claude Code, @ClaudeDevs, 2026-09-26: “@claudedevs a graceful stopping point instead of just dying mid-task is such an underrated quality-of-life thing. as someone who runs multi-step stuff, i really appreciate an agent that tidies up before clocking out lol” [source](https://twitter.com/2017474274437857280/status/2103657944386871705)
  - Praise, Claude Code, @ClaudeDevs, 2026-09-25: “@claudedevs finally a bartender who calls last orders getting cut off mid-edit meant coming back to a half-renamed codebase billing the wrap-up to the weekly limit is fair should've been day one” [source](https://twitter.com/1524885233375809537/status/2103565253934305634)
  - Complaint, OpenAI Codex, r/codex, 2026-09-05: “haven't encounter this since i'm keeping track on my usage, but wrapping things up instead of just stopping would be a nice touch, if it's not done this way already.” [source](https://www.reddit.com/r/codex/comments/1w7x7uw/they_are_robbing_us/p808qtk/)

### Who stands out

- **OpenAI Codex (weaker)**. Codex carries the heaviest complaint load here, with users describing abrupt stops, announced-but-skipped fixes and repeated nudges to continue.
  Complaints far outnumber praise. Users say Codex acknowledges an issue and stops, or asks to implement only after costly analysis. The praise is narrower and often conditional. Some users get fewer early close-outs by telling it to stick to the plan, and one says a model never stopped short of its explicit stop reasons. A Claude Code user contrasts Codex as the agent that keeps going until its turn is done, so experiences clearly vary by model and setup.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-10: “i was aghast trying to troubleshoot an issue yesterday. it was basically like "yeah thats a problem -stop".” [source](https://www.reddit.com/r/codex/comments/1wcedpo/astra_is_good_but/p8x7g76/)
  - Praise, OpenAI Codex, r/codex, 2026-09-06: “<strict_link> this seems quite useful. in fact i observed the same and i simply asked it to stick to the plan without closing out over small steps, it also just works to close out less often.” [source](https://www.reddit.com/r/codex/comments/1w8sn43/astra_keeps_stopping_and_waiting_for_accept_to/p84zg2d/)
  - Praise, OpenAI Codex, r/codex, 2026-09-21: “i've not yet have astra stop on a goal unless it actually met one of my explicit stop reasons” [source](https://www.reddit.com/r/codex/comments/1wlqlz6/psa_codex_cache_only_lasts_max_30_minutes/pb4ivoh/)
  - Complaint, Claude Code, @ClaudeDevs, 2026-09-25: “@claudedevs cool. why so late? codex has been doing this already and simply doesn’t stop until it’s done with its turn. this still sounds rate limited, not on par with codex” [source](https://twitter.com/1265745017089515521/status/2103567536252186898)

- **Claude Code (stronger)**. Claude Code rates better than peers, largely on the graceful-stop change, though stalls and early endings still draw complaints.
  Praise centres on the agent no longer cutting plans and edits short at the limit. Complaints persist. Users report cold stops despite instructions, long tasks ended partway, and the model describing its next step instead of taking it. One user hesitates to trust it with sensitive work for fear of half-finished output. Users here ask most for handoff summaries, which suggests the baseline improved and the bar rose.
  Evidence:
  - Praise, Claude Code, @ClaudeDevs, 2026-09-25: “@claudedevs man, that’s really cool, because it used to be a total drag when you were making plans and he’d cut the process short.” [source](https://twitter.com/1646247180431245314/status/2103577817590243654)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-14: “doesn’t help that the fucker tends to cold stop and not do shit for a while instead of running triage with reviewers (in my case) and despite getting instructions to the contrary. hitting ttl, with the associated costa, because of these things is pretty damn annoying.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wfnr8r/the_hidden_math_behind_why_claude_code_locks_you/p9nvdah/)
  - Complaint, Claude Code, @ClaudeDevs, 2026-09-25: “@claudedevs thanks. !! i've been a bit scared to use it on the trading bot incase it does like half finished stuff” [source](https://twitter.com/520078869/status/2103606685386649900)
  - Praise, Claude Code, @ClaudeDevs, 2026-09-26: “@claudedevs much needed for both sensibility and also for quality of conclusion avoiding interrupted take over.” [source](https://twitter.com/2472155720/status/2103768203352772944)

- **OpenCode (mixed)**. OpenCode users credit its goal feature and hosted models for curbing random stops, but still report dropped tasks when work is queued.
  Posts say the goal feature keeps smaller local models from quitting midway. One user found OpenCode's hosted version of a model ran cleanly, while another provider's build of the same model stopped randomly with no message. Complaints describe the agent jumping to a newly queued task and abandoning the previous one. Volume is low, so treat this as an early signal.
  Evidence:
  - Praise, OpenCode, r/opencodeCLI, 2026-09-18: “goal is also nice as it prevents some smaller local models from just stopping mid-way.” [source](https://www.reddit.com/r/opencodeCLI/comments/1wjixz4/new_btw_command_is_coming_in_the_new_version_of/pal264g/)
  - Praise, OpenCode, @opencode, 2026-09-01: “not all model providers provide the same model the same way. @ollama's version of glm 5.3 flash likes to just stop randomly mid task. no follow up, no message. just stopped. didn't have that happen a single time with @opencode go's version over the past few days” [source](https://twitter.com/3274745054/status/2094609566998876585)
  - Complaint, OpenCode, r/opencode, 2026-09-26: “no its not great , if you give it more tasks as its doing one task it just moves to the next task and doesn't finish the previous one so you have to be kinda careful with it.” [source](https://www.reddit.com/r/opencode/comments/1wqi4a8/my_honest_opinion_about_spacebunny/pc5fgzl/)
  - Complaint, OpenCode, @opencode, 2026-09-16: “@opencode i hope it doesn't stop mid task randomly, time to play 😋 tqsm 🙃” [source](https://twitter.com/2058452594360774656/status/2100240305178239081)

### Fine print

- OpenAI Codex accounts for over half the rated author-weeks, so its weak showing partly reflects its volume of posts.
- Users often cannot tell whether stops come from the model or the harness, and many posts name a model rather than the agent.
- Quota-driven stops belong to a separate criterion, but praise for graceful wrap-ups at limits overlaps with this one.

## Top requests

What users ask to add or change, most asked first. 55 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Author-weeks | Posts | Agents (author-weeks) |
|---|---|---|---|---|
| 1 | Finish tasks fully without stopping early | 28 | 28 | OpenAI Codex 14, Claude Code 8, OpenCode 3, Cursor 2, Devin 1 |
| 2 | No pauses asking user to continue | 7 | 7 | OpenAI Codex 4, Amp 1, Google Antigravity 1, Devin 1 |
| 3 | Graceful stop with resumable checkpoint at limits | 6 | 6 | Claude Code 3, Google Antigravity 2, OpenAI Codex 1 |
| 4 | Handoff summary when stopping | 6 | 6 | Claude Code 6 |
| 5 | Perform announced actions instead of describing them | 4 | 4 | Claude Code 2, OpenAI Codex 2 |
| 6 | Human approval before implementing or merging | 2 | 2 | Cursor 1, OpenCode 1 |

### 1. Finish tasks fully without stopping early

- Claude Code, 2026-09-26, @ClaudeDevs (X): “@claudedevs debria hacer que acabe de minimo la tarea completa 😏” [source](https://twitter.com/1800068778564169728/status/2103780713741144350)
- OpenCode, 2026-09-24, r/opencode (Reddit): “i can only use claude or gpt models at my work, so more persistent models aren't really an option. in my experience, claude is the more naturally persistent of the two, but both of them keep stopping before full task completion outside of relatively small things.” [source](https://www.reddit.com/r/opencode/comments/1wp3bc9/what_are_loopgoal_plugin_are_you_using/pbsgka1/)
- OpenAI Codex, 2026-09-24, r/codex (Reddit): “i 100% feel the fact that it needs more steering and randomly stops without really completing the task” [source](https://www.reddit.com/r/codex/comments/1worwfr/sol_6_was_insufferable_glad_to_be_back_to_56/pbpl3i8/)

### 2. No pauses asking user to continue

- Google Antigravity, 2026-09-17, @antigravity (X): “let's give it a thumbs up to google’s gemini flash 3.8 high fast – it really is pretty impressive! only if you don’t have to interrupt its operation very often, of course. @antigravity” [source](https://twitter.com/1926996539915804672/status/2100431039583932479)
- OpenAI Codex, 2026-09-16, r/codex (Reddit): “yeah i have to prompt it with just proceed or continue somewhere about 10x to 100x as much as i did with sol. please someone tell me they have a fix for this 🙏” [source](https://www.reddit.com/r/codex/comments/1whhxcr/they_swapped_the_sol_and_astra_positions/pa7un6r/)
- OpenAI Codex, 2026-09-15, r/codex (Reddit): “this! right when it let him know that it was incomplete, the first message, why not "complete it". instead he made it confirm multiple times that it wasn't complete. what's the point of that?” [source](https://www.reddit.com/r/codex/comments/1wh2wt0/why_do_i_have_to_constantly_demand_sol_or_astra/p9z8u67/)

### 3. Graceful stop with resumable checkpoint at limits

- Claude Code, 2026-09-26, @ClaudeDevs (X): “@claudedevs the important part is leaving the repo in a state the next session can trust. a short wrap-up that finishes the edit and says what remains is more useful than a hard stop with a half-written diff.” [source](https://twitter.com/1366427283397824513/status/2103881860304810129)
- Claude Code, 2026-09-26, @ClaudeDevs (X): “@claudedevs great change. please bring the same graceful stop to <strict_link> chats too. long research runs there get cut off in the middle of tool calls, and all that work is lost. also, once a week for pro feels very tight when this is basic reliability, not a bonus feature.” [source](https://twitter.com/1804880617307549696/status/2103852303820472684)
- Google Antigravity, 2026-09-26, @antigravity (X): “@emzrsxn @antigravity gracefully stopping at the usage limit should be standard for every coding agent. leaving the repository at a coherent checkpoint matters more than squeezing out one final edit.” [source](https://twitter.com/1144454518156824577/status/2103808104106185194)

### 4. Handoff summary when stopping

- Claude Code, 2026-09-26, @ClaudeDevs (X): “@claudedevs the wrap-up i'd want is a small handoff: files changed, tests actually run, anything still broken, and the next unfinished step. finishing an edit is useful; knowing what's safe to resume is what makes the pause workable.” [source](https://twitter.com/2180560289/status/2103887228099535203)
- Claude Code, 2026-09-26, @ClaudeDevs (X): “@claudedevs can the wrap-up leave a short handoff with what changed, what’s untested, and what to do next? that’s what i’d want waiting when the limit resets.” [source](https://twitter.com/3425573313/status/2103651072879239458)
- Claude Code, 2026-09-25, @ClaudeDevs (X): “@claudedevs leave me a note on what broke so i don't spend the next hour finding it” [source](https://twitter.com/2100040663878291457/status/2103576248391753833)

### 5. Perform announced actions instead of describing them

- Claude Code, 2026-09-17, r/ClaudeCode (Reddit): “makes me think maybe we should make a not\_worth\_flagging hook, that tells opus if they’re about to reply with “worth flagging:” or “two things you should know:” that they should resolve unambiguous findings before replying to the user instead of raising the fact that it didn’t finish the job” [source](https://www.reddit.com/r/ClaudeCode/comments/1wiujum/claude_code_is_falling_behind_codex_not_because/pafk62v/)
- Claude Code, 2026-09-14, @ClaudeDevs (X): “@claudedevs be good to implement it not write docs. genuinely killing me today” [source](https://twitter.com/1754028639052517376/status/2099491116194410711)
- OpenAI Codex, 2026-09-10, r/codex (Reddit): “<strict_link> i want it to continue its task, but sometimes it just says it will do it, but doesn't do it, ever. how can i fix this? i've translated the picture, cause it was in french.” [source](https://www.reddit.com/r/codex/comments/1w9w4tj/codex_usage_and_operation_discussion_last_updated/p8wskgc/)

### 6. Human approval before implementing or merging

- Cursor, 2026-09-23, @cursor_ai (X): “@cursor_ai a cheaper top model only helps if each run still has a named finish line and a human gate on merge. otherwise you just burn a nicer stack on the same unfinished work.” [source](https://twitter.com/1911453463889825792/status/2102781274863833270)
- OpenCode, 2026-09-06, @opencode (X): “been using muse 1.3 via @opencode for a few days now, and i feel this model just proceeds to go ahead and execute even when i just want to go back-n-forth. like bro i'm gonna tell you when to implement it, chill out anyone else? what r ur thoughts? <strict_link>” [source](https://twitter.com/1149503292063436800/status/2096683881479217561)

## Every agent

| Agent | Overall rank | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|---|
| [Claude Code](https://feedbackbench.com/agents/claude-code.md) | 1 | Better than peers | 0.576 | 0.555–0.598 | 112 | 43 | 69 |
| [OpenAI Codex](https://feedbackbench.com/agents/codex.md) | 2 | Worse than peers | 0.450 | 0.412–0.485 | 234 | 15 | 219 |
| [OpenCode](https://feedbackbench.com/agents/opencode.md) | 3 | Too few posts | – | – | 25 | 2 | 23 |
| [Google Antigravity](https://feedbackbench.com/agents/antigravity.md) | =5 | Too few posts | – | – | 11 | 0 | 11 |
| [Cursor](https://feedbackbench.com/agents/cursor.md) | 4 | Too few posts | – | – | 8 | 0 | 8 |
| [Devin](https://feedbackbench.com/agents/devin.md) | =5 | Too few posts | – | – | 5 | 0 | 5 |
| [GitHub Copilot](https://feedbackbench.com/agents/copilot.md) | =8 | Too few posts | – | – | 4 | 0 | 4 |
| [Amp](https://feedbackbench.com/agents/amp.md) | =11 | Too few posts | – | – | 4 | 0 | 4 |
| [Cline](https://feedbackbench.com/agents/cline.md) | =8 | Too few posts | – | – | 3 | 0 | 3 |
| [Pi](https://feedbackbench.com/agents/pi.md) | 7 | Too few posts | – | – | 1 | 0 | 1 |
| [Kiro](https://feedbackbench.com/agents/kiro.md) | 13 | Too few posts | – | – | 1 | 0 | 1 |
| [Zed](https://feedbackbench.com/agents/zed.md) | =8 | Too few posts | – | – | 0 | 0 | 0 |
| [Factory](https://feedbackbench.com/agents/factory.md) | =11 | Too few posts | – | – | 0 | 0 | 0 |
| [Conductor](https://feedbackbench.com/agents/conductor.md) | 14 | Too few posts | – | – | 0 | 0 | 0 |
| [Warp](https://feedbackbench.com/agents/warp.md) | 15 | Too few posts | – | – | 0 | 0 | 0 |
| [Grok Build](https://feedbackbench.com/agents/grok-build.md) | 16 | Too few posts | – | – | 0 | 0 | 0 |
| [Augment Code](https://feedbackbench.com/agents/augment.md) | 17 | Too few posts | – | – | 0 | 0 | 0 |

## Posts

Receipts rule: The 5 most recent praise and complaint posts per area (first 700 characters) and 3 per criterion (first 450 characters).

### Claude Code

- Praise, 2026-09-27, r/ClaudeCode (Reddit): “the "stop here" option in your readme helps. a next-step card can make a finished task feel unfinished if every option sends you into another round of work. knowing the changes will stay uncommitted gives me a clear point to pause and inspect them myself. i'd keep that visible alongside the action buttons as the interface grows” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrttht/claude_isnt_allowed_to_write_me_prose/pcfxz0d/)
- Praise, 2026-09-27, @ClaudeDevs (X): “@claudedevs this is a real quality-of-life win. getting cut off mid-edit was one of the most frustrating parts of long claude code sessions.” [source](https://twitter.com/1149586194/status/2104035326746620053)
- Praise, 2026-09-27, @ClaudeDevs (X): “@claudedevs graceful stopping is an underrated reliability feature. a bounded wrap-up that leaves the repo coherent, notes what remains, and avoids half-written edits is much more valuable than squeezing out a few extra minutes of raw execution.” [source](https://twitter.com/241089216/status/2104040477767204900)
- Complaint, 2026-09-27, @ClaudeDevs (X): “@claudedevs 到点了先找停手处，别在改文件一半硬切。这个比加额度有用。” [source](https://twitter.com/2002993609671630848/status/2104033364068233609)
- Complaint, 2026-09-27, @ClaudeDevs (X): “@claudedevs or it should treat us like the properly on time recurring paying customers we are and just let it finish. what the fuck you greasy snakes quit destroying hours and thousands of dollars to create more issues and rework and respect time and energy figure it out codex is better” [source](https://twitter.com/1709524258123063296/status/2104201575417999491)
- Complaint, 2026-09-26, r/ClaudeCode (Reddit): “your third check is the one ours needed - the summaries ended on "ready to proceed?" right after validation passed. hadn't come across stop\_hook\_active - thanks.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wp1bj2/5_of_our_8_headless_agent_runs_built_the_thing/pc4mbb0/)

### OpenAI Codex

- Praise, 2026-09-26, r/codex (Reddit): “opus 5.5 is crazy! it is friday and i just ran out a usage when usually i run out on wednesday. it is also performing a lot better and acually doing things rather than saying i will do it.” [source](https://www.reddit.com/r/codex/comments/1wqbnpj/we_are_back/pc35klj/)
- Praise, 2026-09-26, r/GithubCopilot (Reddit): “not sure if it’s just me but i find vs code so frustrating to use compared to codex or claude code. doesn’t matter which agent, they always stop for some dumb reason and say yeah you’re right i stopped for no reason. never have that issue with codex or claude code.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wowuzs/copilot_uses_apply_patch_idiosyncracy_that_fails/pc7dwsv/)
- Praise, 2026-09-22, r/codex (Reddit): “same brother, i noticed they went back to letting the tasks finish before shutting it down (which is great). i'm pretty sure its still limited to like 1-5% of your total possible usage behind the scenes though” [source](https://www.reddit.com/r/codex/comments/1wnda22/the_engines_have_been_forcefully_stopped/pbdy8nv/)
- Complaint, 2026-09-27, r/codex (Reddit): “yes, but also a lot dumber. mine just doesn't finishes tasks at all, it "finishes" and then says 90% of the work still remains, and that goes on and on for hours. half of my astra chats are also being nerfed with some obscure braindead model.” [source](https://www.reddit.com/r/codex/comments/1wr71h7/has_anyones_astra_become_more_generous_on_their/pcaqdky/)
- Complaint, 2026-09-27, r/codex (Reddit): “codex limits have become much smaller and it is absolutely useless now! sol (with medium reasoning) just consumed the entire 5h quota in a single prompt on a relatively small codebase, and did not even finish the task! i guess it's time to move on to claude..” [source](https://www.reddit.com/r/codex/comments/1wo8fhw/okay_they_literally_cut_our_quota_by_half_gpt_6/pccbsyd/)
- Complaint, 2026-09-27, r/codex (Reddit): “yes, absolutely. sol just consumed the entire 5h quota in a single prompt and did not even finish the task. so it has become effectively useless.” [source](https://www.reddit.com/r/codex/comments/1vtiymv/is_the_usage_limit_nerfed/pccdcex/)

### OpenCode

- Praise, 2026-09-18, r/opencodeCLI (Reddit): “goal is also nice as it prevents some smaller local models from just stopping mid-way.” [source](https://www.reddit.com/r/opencodeCLI/comments/1wjixz4/new_btw_command_is_coming_in_the_new_version_of/pal264g/)
- Praise, 2026-09-01, @opencode (X): “not all model providers provide the same model the same way. @ollama's version of glm 5.3 flash likes to just stop randomly mid task. no follow up, no message. just stopped. didn't have that happen a single time with @opencode go's version over the past few days” [source](https://twitter.com/3274745054/status/2094609566998876585)
- Complaint, 2026-09-26, r/opencode (Reddit): “no its not great , if you give it more tasks as its doing one task it just moves to the next task and doesn't finish the previous one so you have to be kinda careful with it.” [source](https://www.reddit.com/r/opencode/comments/1wqi4a8/my_honest_opinion_about_spacebunny/pc5fgzl/)
- Complaint, 2026-09-21, r/opencode (Reddit): “it's pretty decent but a bit more fickle than ds v4.1 flash. it has some weird edge cases where it likes to end its turn early, and those situations can't quite be reliably detected by a harness. it'll sometimes say it's about to do a tool call or something and then just... stop. your best bet is to have it spawn a sub-agent to actually do work, so that if the sub-agent kills itself it triggers the orchestrator muse-spark (or some other model) to” [source](https://www.reddit.com/r/opencode/comments/1wl91no/now_that_deepseek_v41_flash4x_usage_has_ended_is/pb67ozj/)
- Complaint, 2026-09-20, @opencode (X): “is anyone else experiencing abrupt cut-offs with @opencode during code generation? wondering if it’s a known issue or just me.” [source](https://twitter.com/1987082476544794624/status/2101685900929282188)

### Google Antigravity

- Complaint, 2026-09-24, @antigravity (X): “the person literally read my entire project line by line to give up at the finish line. 😭 @rodydavis @antigravity <strict_link>” [source](https://twitter.com/2101999301945614336/status/2103178595166540141)
- Complaint, 2026-09-21, r/google_antigravity (Reddit): “yup. had to fix it myself because the harness wouldn't do it for me.” [source](https://www.reddit.com/r/google_antigravity/comments/1wm320i/bruh_come_on/pb3wzgk/)
- Complaint, 2026-09-19, r/google_antigravity (Reddit): “too lazy to do the work 😭” [source](https://www.reddit.com/r/google_antigravity/comments/1wkksgs/38_invokes_a_sub_agent_for_a_task_the_sub_agent/patrbbw/)

### Cursor

- Complaint, 2026-09-23, @cursor_ai (X): “@grok @elonmusk @cursor_ai the word count is not enough, the main text is not written: xhigh” [source](https://twitter.com/1962731650108035072/status/2102690634343547146)
- Complaint, 2026-09-18, r/cursor (Reddit): “"welp boss, this seems like a great time to take a break! what a great stopping point!" dude you finally fixed the bug and you think we're stopping now?” [source](https://www.reddit.com/r/cursor/comments/1wj4p6s/cursor_agents_saying_no/panq1wq/)
- Complaint, 2026-09-17, @cursor_ai (X): “@stephenperreira @theaaron @cursor_ai @grok @bot roundabout vms are fine. mine still stops after one cheap search page. the shortlist is enough.” [source](https://twitter.com/2077265437398880256/status/2100569684747726971)

### Devin

- Complaint, 2026-09-21, @cognition (X): “@da7_tech ok, i bit and signed up to @cognition @devindesktop with a pro subscription. i asked it to do a scan on one repo. it ran for about 20-25 minutes. it didn't finish <strict_link>” [source](https://twitter.com/2071953914182725633/status/2102138818736300445)
- Complaint, 2026-09-18, @DevinAI (X): “@melvindvivas @devinai tried devin, stops about an hour in everytime. folks say this is on par with astra or opus, not a chance lol. it's good at bug hunting that's it. deepseek beats it.” [source](https://twitter.com/2099788664247300096/status/2100988446042910731)
- Complaint, 2026-09-17, @cognition (X): “@cognition hi glm 5.3 flash on devin is constantly stopping mid work. on latest desktop version. it wasnt like this before. pls fix this” [source](https://twitter.com/1392605672794034182/status/2100696367614300178)

### GitHub Copilot

- Complaint, 2026-09-26, r/GithubCopilot (Reddit): “not sure if it’s just me but i find vs code so frustrating to use compared to codex or claude code. doesn’t matter which agent, they always stop for some dumb reason and say yeah you’re right i stopped for no reason. never have that issue with codex or claude code.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wowuzs/copilot_uses_apply_patch_idiosyncracy_that_fails/pc7dwsv/)
- Complaint, 2026-09-22, r/GithubCopilot (Reddit): “no such id, it is a private repo. how do u see "secrets"? it only says copilot's work was cancelled, and i did not give it that instruction, or stop instruction.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wmbaez/is_anyone_having_random_work_cancels_on_github/pbbupdx/)
- Complaint, 2026-09-21, r/GithubCopilot (Reddit): “no, i got this msg, now 4th time, cancelling the work randomly, out of nowhere, without my stop instruction, and i have no token limit issues, i had to start a new session and brief copilot with the workprgoram and scope again, below image: <strict_link>” [source](https://www.reddit.com/r/GithubCopilot/comments/1wmbaez/is_anyone_having_random_work_cancels_on_github/pb5oo29/)

### Amp

- Complaint, 2026-09-09, @AmpCode (X): “i wonder if the @ampcode team has made any optimizations to address astra’s tendency to stop mid-task and ask lots of questions instead of doing the work 🤔” [source](https://twitter.com/1705384263867379712/status/2097785428917264538)
- Complaint, 2026-09-07, @AmpCode (X): “tomorrow i would try my tuned hard mode in @ampcode , cz default high mode somehow doesnt passing the vibe + having bad exp with “you are right” “replying instead of start working”, may be its astra itself thats doing this dumb play <strict_link>” [source](https://twitter.com/1248119942194421760/status/2096764065334923682)
- Complaint, 2026-09-07, @AmpCode (X): “@jkudish yes i agree! using astra with @ampcode last night building a large scope of work. it stopped half way and said it knew it didn’t complete all the work even when it was instructed to finish all of it.” [source](https://twitter.com/1535774830225944576/status/2097095185301999839)

### Cline

- Complaint, 2026-09-26, r/CLine (Reddit): “hello everyone , i just started using vscode +cline , just for fun , messing around with unity scripts and stuff , and all was great for a few days , during 1 task , the power went down , and after i turned on my pc i had the following problem , he just stopped executing tasks , most of the time just telling me how to do it , and sometimes replying just in code , i ve been trying to troubleshoot it for 2 days now , and this is what i found. i use” [source](https://www.reddit.com/r/CLine/comments/1wqm2kn/vscode_and_cline/)
- Complaint, 2026-09-25, r/CLine (Reddit): “bro just one prompt and your usage is gone and it even finishes midpoint.” [source](https://www.reddit.com/r/CLine/comments/1wpyepq/gemini_38_flash_is_now_free_in_cline/pbzgvsl/)
- Complaint, 2026-09-24, @cline (X): “@dynamicwebpaige @cline it would been nice if it worked for more than half a prompt. deepseek took the slack and finish the job. it's aight in <strict_link> <strict_link>” [source](https://twitter.com/3977375776/status/2103251904612753482)

### Pi

- Complaint, 2026-09-18, @pidotdev (X): “@melissapan @howaboua @pidotdev nice work. i have issues with fable &amp; pi, where pi keeps reading and never writes the code to files. not sure of the problem though.” [source](https://twitter.com/1613070948/status/2100965181912482194)

### Kiro

- Complaint, 2026-09-12, r/kiroIDE (Reddit): “sometimes i ask something and it doesn't do anything. like my question was abandoned promptly. and when it's waiting for another service/event which never happens kiro wouldn't say a word” [source](https://www.reddit.com/r/kiroIDE/comments/1wefi0t/lately_kiro_has_been_just_hanging/p9e4h1d/)
