# Spins, loops or gets stuck without progress (`work.stuck_loops`)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/criterion/work.stuck_loops

Area: [Doing the work](https://feedbackbench.com/criteria/work.md)

**Definition.** The agent repeats attempts, loops endlessly or spirals on a problem without converging.

**Boundary.** Not this: see [Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md) for stopping early. Not this: see [Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md) when the fix was to ask the user.

Rated author-weeks, all agents: 995. Complaint share: 95%.

## The brief

Written by Claude Opus 5.5 from 66 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**Every agent loops; the difference is who notices and stops.**

TL;DR:

- Complaints swamp praise for every agent here; looping is a category-wide habit, not one vendor's bug.
- Google Antigravity and Cursor draw sharp stuck-loop complaints; Claude Code earns credit for saying when it's stuck.
- Users mostly break loops themselves with ledgers, hard pass caps and cross-model review.

In plain terms: Expect your agent to circle sometimes. It rereads files, retries the same failed fix, or sits on a spinner for hours while your usage cap drains. Users who stay productive cap retries and record rejected approaches.

### How it breaks

- **Usage limits drained by circling** ([Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md)). The costliest loop is the one that spends the five-hour or weekly allowance and delivers nothing.
  Users describe agents that analyze, search and retry until the quota is gone, with no finished work to show. The pain is double. Time is lost, and the paid capacity needed to try again is gone too. Posts from Google Antigravity users are the loudest here. The same pattern shows up with Factory credits and Claude models running unattended inside Cursor. Some users now ask for refunds on usage burned by stuck loops.
  Evidence:
  - Complaint, Google Antigravity, @antigravity, 2026-09-05: “@nlycskn @antigravity @thtbee_ i tried to use it for my tasks, but it absolutely does not cope, it goes in circles, analyzes, searches, trying to complete the task, but burns the five-hour limit without doing anything” [source](https://twitter.com/1791981790355247105/status/2096188833687568869)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-17: “this also my issue it drain my 5hours and even weekly is suffering because of this” [source](https://www.reddit.com/r/google_antigravity/comments/1winp2p/infinite_file_read_loops_ongoing_issue_happening/pabt9jc/)
  - Complaint, Factory, @droid, 2026-09-15: “@clementpillette @droid @zai_org burning weekly cloud credits on a stuck loop vs finishing on studio is why people buy the ram.” [source](https://twitter.com/2062965074659000320/status/2099957956804755902)
  - Complaint, Cursor, r/cursor, 2026-09-25: “so it depends on how you code. codex gives the best overall limit followed by cursor then claude code as the worst. sol as orchestrator and luna as doers is the most cost effective for that style right now. a grok or third party plan and composer implement is probably the best deal for a mid tier model for those that want to be hands on. opus 5.5 is probably the best for a single model long run but limits will hit. probably the most hands off. but claude models can spiral and burn everything if you aren't watching.” [source](https://www.reddit.com/r/cursor/comments/1wpm09i/beginner_making_a_react_native_app_is_cursor_pro/pbz3grq/)

- **Same mistake, explained, then repeated** ([Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md)). Agents diagnose their own error correctly, then make a closely related error on the next attempt.
  This loop is not a frozen process. The agent stays busy and articulate, and it still fails to converge. Users report fixing one small feature through a long run of corrections, and agents retrying a failure that another session had already solved. Several users push this back onto instructions. Contradictory or impossible goals can send an agent spending an hour trying to satisfy the unsatisfiable. Retry-after-identical-failure is a recurring vendor request.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-22: “basically: **“i gave the model fairly specific coding instructions, it repeatedly misunderstood them, correctly explained its own mistakes, then immediately made closely related mistakes again. fixing one tiny feature required a ridiculous number of corrections.”**” [source](https://www.reddit.com/r/codex/comments/1wnizjr/1ˢᵗ_impreßions_on_þe_6solmedium/pbfgc70/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-23: “"how can we stop the our users from complaining about our premiere token eater...? a shittier model that consumes less but half as good and half as fast". i'll probably still go through my weekly usage plan telling it to fix the stupid mistakes it creates, over and over.” [source](https://www.reddit.com/r/codex/comments/1wo7pz8/all_they_needed_to_do_was_fix_the_token_burn_with/)
  - Complaint, Cursor, r/cursor, 2026-09-27: “cursor agents are fast at *retrying*. the expensive part for us was retrying the **same** fail — bad path, wrong toolchain, flake that already had a known fix in another session. we keep a small oss prior-art index (claimidx, apache-2.0) beside the agent: ask before grinding, apply + verify, then publish a compact claim. retrieved remedies are evidence for the model, not auto-executed patches. if you want the full loop (terminal step is share): ```bash pip install -u "claimidx[server]>=0.7.13" claimidx init --agent <your-cursor-agent> claimidx claim --yes --channel reddit --source path-b ``` on ≥0.7.13, `claim --yes` auto-shares to the commons; `--local` keeps it private. mcp name `claimidx-mcp` / registry `io.github.claimidx/claimidx`. curious how others in cursor land gate “don’t try this fail again” — rules, memories, or something external? docs: <strict_link> · <strict_link>” [source](https://www.reddit.com/r/cursor/comments/1wqkwtj/your_cursor_plan_already_spins_cloud_agents_why/pc9t0l9/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-15: “getting stuck in long loops probably means you have to review your own instructions. it may be that you have given it contradictory instructions or an impossible task. sometimes, it can be subtle. for example, i once gave my sol an edge case to handle and told it write to write a test for it. however, that edge case is actually impossible to happen because of other reasons. so it handled it, but then found it couldn't test it as i asked because there isn't really a path through the code that reach it. so then it started trying to modify the rest of the program to make it testable, which then caused other tests to fail, and when i clicked back over to see if this simple task was done, i saw it had been working over an hour, in a loop of trying to make itself be able to reach that edge case to test without breaking all other tests.” [source](https://www.reddit.com/r/codex/comments/1wh3z15/is_astra_really_better_at_all_coding_or_just/p9za5hz/)

- **Rereading files it already understands** ([Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md)). Agents scan the same files dozens of times even when their own reasoning shows they have found the issue.
  Posts describe agents rereading files on small fixes, sometimes after the thinking trace has already named the bug. Users who add a rule against it in their instructions file report the rule being ignored. Every request to stop file re-reading loops comes from Google Antigravity and OpenAI Codex users, with Antigravity asking most.
  Evidence:
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-08: “yeah mine keeps reading files like 50 times even though it shows in it's thinking it knows the issue (normally small fixes) tried putting it in my .md and it completely ignored it.” [source](https://www.reddit.com/r/google_antigravity/comments/1walojq/cmon_google_you_cant_do_this_to_your_loyal/p8mm9jt/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-05: “not upto the mark. browsing is very slow. it many times goes into a loop of re-reading the same files multiple times. too slow even on things on google workspace. it is not good at data handling even basic csv files on google ecosystem. the mcp support is also not good. we have so many options for claude code and other harness but not for gemini spark. i am not too happy tbh. there are better options. also i am afraid to allow it to work on my personal data. my hermes(on free nemotron apis) is much faster than it for basic file handling and and tool calling.” [source](https://www.reddit.com/r/google_antigravity/comments/1w5kvoh/antigravity_cli_remote_setup_as_personal_agent/p80jnde/)

- **Shows working, does nothing** ([Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md)). Some agents sit on an active spinner for hours and never notice they have stalled.
  This is the quiet version of the loop. The UI says working, nothing changes, and a nudge from the user is sometimes ignored. One Kiro user reports a ten-hour overnight idle. Cursor, Devin and Cline users describe the same ticking-with-no-work state, and one Devin post says it never notices it is stuck. Users ask for hard budgets and kill switches so a runaway or frozen run ends on its own.
  Evidence:
  - Complaint, Kiro, r/kiroIDE, 2026-09-12: “it will show working, but it isnt doing anything just sitting there. i can interrupt with "are you doing anything" and it will sometimes just ignore me, but sometimes will say it is working. it sat overnight for 10 hours and did nothing during that time. any suggestions?” [source](https://www.reddit.com/r/kiroIDE/comments/1wefi0t/lately_kiro_has_been_just_hanging/)
  - Complaint, Cursor, @cursor_ai, 2026-09-18: “@cursor_ai goal ended or not? it's just ticking with no work. <strict_link>” [source](https://twitter.com/1553011876811907072/status/2100748420927324507)
  - Complaint, Devin, @DevinAI, 2026-09-22: “my @devinai experience in a nutshell… like cursor it seems to never notice when it is stuck and just waits forever? <strict_link>” [source](https://twitter.com/63583842/status/2102432788112871614)
  - Complaint, Cline, @cline, 2026-09-27: “i know stealth models seem to be the in thing right now, but i'm not sure they're even worth messing around with sometimes. trying to use pixel canary on @cline, and it's just so slow. i mean, 24 hours now, no closer to the task, and it keeps stopping and starting. it's horrible!” [source](https://twitter.com/25673607/status/2104177440129991012)

- **Thinking loops that eat context** ([Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md)). Reasoning that circles on itself fills the context window, forces early compaction and degrades the rest of the task.
  OpenCode users tie many loops to the model and its settings rather than the harness. One reports lower reasoning tiers looping the same tokens and never leaving thinking mode, while the top tier ran clean. Others say swapping models removed the loops. The knock-on cost is what users flag. Runaway thinking burns tokens, triggers compaction too soon and lowers output quality.
  Evidence:
  - Complaint, OpenCode, r/opencodeCLI, 2026-09-22: “for me, overthinking, going in loops, and consuming too much token that it will fill the context window fast and then compaction happen too fast and the task quality degrade. they are much more serious issue than cache cost” [source](https://www.reddit.com/r/opencodeCLI/comments/1wnhazu/gpt6_luna_cheap_than_deepseek_v41_flash/pbfww0h/)
  - Complaint, OpenCode, r/opencode, 2026-09-21: “i also only vibe side projects, and ds has been working great for me. tried glm when ds was down, and i wasn't impressed. both aren't the smartest, but ds can brute-force through everything. btw, don't use anything below `max` with ds 4/4.1 flash. the thinking will get stuck in an endless loop of trying to do a tool call, but never leaving the thinking mode. no idea why max produces very clean thinking output while all the lower ones often loop the same tokens over and over.” [source](https://www.reddit.com/r/opencode/comments/1wfqol6/glm_53_flash_vs_deepseek_v41_flash_which_one_do/pb7sk9q/)
  - Complaint, OpenCode, r/opencodeCLI, 2026-09-02: “faced hallucinations as well and i also faced looping multiple times too. i find luna or qwen 3.8 flash more stable (luna especially).” [source](https://www.reddit.com/r/opencodeCLI/comments/1w52a86/serious_hallucinations_with_the_glm_53_flash/p7c46b4/)
  - Praise, Pi, @pidotdev, 2026-09-11: “@gamboasanabria @pidotdev vieras que me pasaba mucho, sobre todo los qwen con sus “but wait” infinito, pero dejé a claude toda la noche probando diferentes combinaciones de settings tanto de lmstudio como pi y ahora está mucho mejor!” [source](https://twitter.com/19957424/status/2098557362441199753)

### Who stands out

- **Google Antigravity (weaker)**. Users describe death loops and repeated file reads that drain quotas, and they file the most loop-fix requests.
  Antigravity posts cluster around the worst version of this failure. The agent moves fast, ignores skill instructions, then dies in a loop. Users report slow browsing, repeated rereads of the same files and burned five-hour windows. One OpenCode user calls Google's harness poor at handling death loops. A few users say newer models loop less, and one says the CLI avoided loops the IDE hit. Most requests to fix endless looping and to stop file re-reading come from here.
  Evidence:
  - Complaint, Google Antigravity, @antigravity, 2026-09-05: “@nlycskn @antigravity @thtbee_ works fast, sometimes produces working code, ignores the skill instructions, dies in the death loop. <strict_link>” [source](https://twitter.com/2009822403170578434/status/2096190394254127519)
  - Praise, OpenCode, r/google_antigravity, 2026-09-07: “lol is this still a thing? i got these several times back in the days when antigravity just launched and few months after that google got some of the worst harnesses that can't deal with death loops, and i guess their models also way more prone to them than others. outside of google stuff i got deathloop like once in opencode with some chinese model and never again” [source](https://www.reddit.com/r/google_antigravity/comments/1wa0mw4/what_is_this_antigravity_shame_shame/p8ev80e/)
  - Praise, Google Antigravity, r/google_antigravity, 2026-09-24: “well it is better than getting stuck in loops. also, i was used to ide till i converted cli never got loops or got any errors” [source](https://www.reddit.com/r/google_antigravity/comments/1wp3k02/is_gemini_38_flash_getting_stuck_in_loops_for/pbsaisp/)

- **Claude Code (stronger)**. The upgrade users single out is that it now says when it's stuck instead of drifting silently.
  Praise centres on stuck-reporting. Users call silent loop drift a long-running problem and treat the visible stuck signal as the real change. Complaints remain. Users say weaker models spin on specs, cost more tool calls, and need constant steering inside subagent workflows. Power users report that the circling mostly stopped once rejected approaches were written to a ledger the next session reads.
  Evidence:
  - Praise, Claude Code, @ClaudeDevs, 2026-09-02: “@claudedevs 75% cache-read cut is the real story — that re-prices any workflow hammering the same context window. the "tells you when it's stuck" upgrade is the underrated one; agent loops that silently drift have been the silent killer for a year.” [source](https://twitter.com/2028308162382901248/status/2095125548083347604)
  - Praise, Claude Code, @ClaudeDevs, 2026-09-01: “@claudedevs the "tells you when it's stuck" part is the actual upgrade.” [source](https://twitter.com/1999402064389242880/status/2094878484669640859)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-17: “same tail problem here, and what fixed it was making the tail part of "done" instead of something i remember. every session runs from a plan file in the repo, and the rule is that nothing ends without appending to a ledger next to it: what landed (commit range), what was found and deferred, and one line per decision the agent took on my behalf with "cost if wrong". the circling mostly stopped because the ledger also holds the rejected paths: "tried a, ruled out because x" is read by the next session before it proposes a again with more confidence. for two runtimes on the same box the only arrangement that held up: one of them owns the ledger and does the fixing, the other only reads it and posts findings into it, never touches the code. duplicate work went away; the coordination cost didn't, it just became visible.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wiw34w/a_year_of_claude_code_alongside_hermes_and_codex/pae44aa/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-12: “if i’d be doing only that… can’t live without subagents. how can i finish 5 projects i’m doing in parallel without them? and leaving all to opus instead of fable… it just keeps it spinning doing wrong stuff for wrong reasons… need to stop it and poor some fable holy water on it context all the time to keep it from doing some bs.” [source](https://www.reddit.com/r/ClaudeCode/comments/1we96di/task_how_to_survive_till_tuesday_on_22_limit_left/p9bvz9o/)

- **OpenAI Codex (mixed)**. Codex collects the most praise and the most complaints, and users split sharply by model.
  Fans say newer models are token-efficient, meander less and try new approaches instead of going in circles. One user reports a bug fixed in a minute that an older model spun on for fifteen. Others say the same model went in circles and contradicted itself for days. Requests include auto-continue without manual prompting and less polling of background processes.
  Evidence:
  - Praise, OpenAI Codex, r/codex, 2026-09-04: “nah, it is fast because it is token efficient and isn't stuck in loops on verifying stupid things” [source](https://www.reddit.com/r/codex/comments/1w7gy48/astra_is_fast/p7uy6tl/)
  - Praise, OpenAI Codex, r/codex, 2026-09-05: “> unfortunately, it is a lot less than it used to be. i could use $100 of sol per week whether this is better or worse depends on how much you get done with that usage. i've only had astra for a few hours, so i can't give a definitive response for my use cases, but so far it seems both significantly more effective and significantly faster. it meanders less, trips itself up less, stays on task better, gets to the point faster, and just plain outputs faster.” [source](https://www.reddit.com/r/codex/comments/1w7uh3a/psa_plus_account_actual_astra_usage_limits/p7xozgx/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-12: “i did, astra is completely useless, it goes in circles and says contradictory stuff for 5 days straight... waste of money and time” [source](https://www.reddit.com/r/codex/comments/1wecmq0/how_many_of_you_guys_have_switched_back_to_sol/p9fo6w8/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-06: “i’m impressed too. it is really fast and pretty efficient. i haven’t found any coding projects that it got wrong using medium. it introduced a few minor bugs on a complex project but fixed them in less than a minute when i pointed it out (rather than spinning around the issues for 15 minutes like sol or arguing about it like fable).” [source](https://www.reddit.com/r/codex/comments/1w92bbt/loving_astra/p87hi29/)

- **Cursor (weaker)**. Cursor has nearly no praise on loops, and users lean on their own caps and summaries to keep it converging.
  Users report agents stuck every time, ticking with no output, or retrying the same failure fast. Most fixes come from users, not the product. They hard-cap revision passes, reduce threads to a decision summary so the agent stops rewriting the same hunk, and run separate reviewers. Users also ask for refunds on usage burned by stuck loops.
  Evidence:
  - Complaint, Cursor, r/cursor, 2026-09-02: “yeah i never hard-capped the revision loop. grok will eat the whole day if you let it. 3 passes then stop is the part i was missing.” [source](https://www.reddit.com/r/cursor/comments/1w4xps5/at_max_thinking_sol_finishes_the_job_and_grok_46/p7c55dc/)
  - Complaint, Cursor, r/cursor, 2026-09-25: “<strict_link> what to do, please help. its getting stuck everytime. not able to do any work” [source](https://www.reddit.com/r/cursor/comments/1wpwmbr/cursor_getting_stuck_either_at_any_command_or/)
  - Praise, Cursor, @cursor_ai, 2026-09-24: “@kevinhomorales @bot @cursor_ai the decision-summary step is what most people skip. once the thread is reduced to the actual fork, cursor stops rewriting the same hunk three times.” [source](https://twitter.com/2061779435557117952/status/2103229107995586979)
  - Praise, Cursor, r/cursor, 2026-09-25: “failing test first catches a lot of the empty-list misses. what still gets me is the weak test that turns green for the wrong reason, then the agent implements to that bar. after the patch lands i run a different-family read-only pass. reviewers only report, they don't edit, and i don't concede a finding unless it cites a path in the repo. same-family self-check keeps sharing the same blind spots. if a later round finds worse problems than the previous one, the patches are injecting bugs and i stop instead of looping.” [source](https://www.reddit.com/r/cursor/comments/1wpp67c/i_make_the_agent_write_one_failing_test_before/pbx9bri/)

### Fine print

- Many posts name models rather than harnesses, so a loop blamed on an agent may come from the model it ran.
- Pi, Cline, Devin, GitHub Copilot and smaller agents have too few posts here to rank with confidence.
- Praise is scarce across the board; most positive posts describe fewer loops, not none.

## Top requests

What users ask to add or change, most asked first. 68 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Author-weeks | Posts | Agents (author-weeks) |
|---|---|---|---|---|
| 1 | Fix agents looping endlessly without progress | 26 | 27 | Google Antigravity 10, OpenAI Codex 6, Claude Code 4, OpenCode 3, Devin 2, Cursor 1 |
| 2 | Stop repetitive file re-reading and scanning loops | 6 | 9 | Google Antigravity 4, OpenAI Codex 2 |
| 3 | Stop retrying after repeated identical failures | 6 | 6 | Claude Code 2, OpenAI Codex 2, Google Antigravity 1, Pi 1 |
| 4 | Hard budget or kill switch for runaway agents | 5 | 5 | OpenCode 2, Cursor 1, Kiro 1, Pi 1 |
| 5 | Harness-level loop detection and auto-remediation | 5 | 5 | Claude Code 2, Amp 1, GitHub Copilot 1, Pi 1 |
| 6 | Auto-continue without manual prompting | 3 | 3 | OpenAI Codex 3 |
| 7 | Stop excessive polling of background processes | 3 | 3 | OpenAI Codex 2, Claude Code 1 |
| 8 | Refund usage burned by stuck loops | 2 | 2 | Cursor 2 |

### 1. Fix agents looping endlessly without progress

- Google Antigravity, 2026-09-26, @antigravity (X): “@antigravity @geminiapp @googleaistudio your gemini flash 3.8 on medium reasoning. it keeps going like this forever, you need to dix this behaviour. <strict_link>” [source](https://twitter.com/1493151747413524480/status/2103897230230880585)
- Claude Code, 2026-09-24, @ClaudeDevs (X): “@claudedevs i just filed this about projects - i would appreciate a fix asap - <strict_link> potentially a powerful feature but auto mode = aut-no progress” [source](https://twitter.com/15527674/status/2103146479506633025)
- OpenAI Codex, 2026-09-23, r/codex (Reddit): “just wait till you see the loop over and over and over. that’s all luna 6 does is loop. sucks because i changed everything over to luna 6 and now i’ve wasted 15% of my weekly tokens 😭” [source](https://www.reddit.com/r/codex/comments/1wnmdz6/first_impression_of_sol6_fast_cheap_and_shitty/pbgyhxj/)

### 2. Stop repetitive file re-reading and scanning loops

- Google Antigravity, 2026-09-24, r/google_antigravity (Reddit): “same here. i feel like 3.7 almost never gets stuck in a loop when reading files. it might loop sometimes, but it never spends 20–30 minutes like 3.8 does, repeatedly reading files, only to end up changing 3 lines of code.” [source](https://www.reddit.com/r/google_antigravity/comments/1wolcdf/is_it_just_me_or_is_gemini_37_flash_better_than_38/pbrcq4y/)
- Google Antigravity, 2026-09-22, r/google_antigravity (Reddit): “yeah, it just reads and reads... sometimes when i tell it to stop it actually does but i'm lucky to have that happen” [source](https://www.reddit.com/r/google_antigravity/comments/1wni05g/we_need_a_usage_reset_now/pbf7wzo/)
- OpenAI Codex, 2026-09-17, r/codex (Reddit): “astra used to launch the program fine. now it keeps scanning the machine to "find" it, called it an oopsie, then ignored the .md location file you made it write and went back to scanning. when that scan loop chewed 5-8 minutes on something it uses constantly, what task were you mid-way through, and did you keep burning tokens or bail?” [source](https://www.reddit.com/r/codex/comments/1windb3/astra_has_been_nerfed/pae5wh0/)

### 3. Stop retrying after repeated identical failures

- Google Antigravity, 2026-09-24, r/GoogleAntigravityIDE (Reddit): “use hooks so that doesn't loop when it's having more than 3 errors for the same cmd” [source](https://www.reddit.com/r/GoogleAntigravityIDE/comments/1wo0v3m/eating_loop_of_my_token_in_by_38/pbqai0e/)
- Claude Code, 2026-09-24, @ClaudeDevs (X): “charging for safeguard-blocked requests is a product honesty move, and also a sharp edge for agent builders. if a loop can burn money on refusals, you need (1) a cheap precheck for known-blocked patterns, (2) a fallback model or mode that is allowed to say no without billing like a full completion, and (3) clear telemetry so the agent stops retrying the same refusal. otherwise safety policy becomes a silent spend bug in production agents.” [source](https://twitter.com/1416221221432172550/status/2103209860103807328)
- OpenAI Codex, 2026-09-18, r/codex (Reddit): “every failed apply_patch attempt re-sends the whole session, so a 60 line edit that fails fifteen times bills like fifteen edits. that is where the weekly limit goes, not into the patch. the acl error only has to repeat twice to start the loop. kill the run after the second identical failure instead of letting it grind for the rest of the window.” [source](https://www.reddit.com/r/codex/comments/1wk0v7e/there_are_currently_some_bugs_under_windows_such/pan284u/)

### 4. Hard budget or kill switch for runaway agents

- Cursor, 2026-09-24, @cursor_ai (X): “@cursor_ai 7% off with tighter prompts and selective tools is the boring win. next cut is a hard stop before the agent burns the savings in a loop.” [source](https://twitter.com/2033957325631873024/status/2103134656149520690)
- Kiro, 2026-09-16, r/kiroIDE (Reddit): “i had a runaway agent once. it took all tasks, went into the background and continued for a couple of hours. no stopping of the thing. it survived prompts, commands, sessions and restarts. 100+ tokens on haiku and ~30 tasks later it happily reported in a newly opened session that it finished. micro-skynet experience. good it was a small private project.” [source](https://www.reddit.com/r/kiroIDE/comments/1whrobc/in_less_than_30_secs_kiro_uses_almost_100_credits/pa4s07y/)
- OpenCode, 2026-09-04, @opencode (X): “46h/800 calls is a runaway loop, not just inefficiency. set a hard per-session max-steps/time budget, log tool calls, and kill/restart when the same step repeats. apply the limit before rerunning. clawpanel keeps opencode alongside openclaw, hermes agent and deepseek harness: <strict_link>” [source](https://twitter.com/2026183504237834240/status/2096010085642252317)

### 5. Harness-level loop detection and auto-remediation

- GitHub Copilot, 2026-09-23, r/GithubCopilot (Reddit): “i have to write rules myself to prevent bad calls of some tools (e.g. gdb) that lead to them waiting for user input. obviously custom rules can't cover the entirety of tools that the model might run. a better way is simply to have the harness tell them directly so that they kill that command and retry with the correct args.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wo3ich/models_should_be_told_when_a_cli_tool_is_waiting/)
- Pi, 2026-09-08, r/PiCodingAgent (Reddit): “hey just wanted to say thanks for releasing all of this. fun to try out and use, adding into my own little harness. found one issue with codegraph + async forks: `explore_code` can hang indefinitely because there’s no timeout on the operation/subprocess. when that happens the fork gets stuck and `steer_fork` starts timing out too. added a 120s timeout + abort/subprocess cleanup locally and it fixed it.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1w9e6zo/my_pi_agent_setup_part_2_native_async_operation/p8hgqso/)
- Claude Code, 2026-09-01, @ClaudeDevs (X): “@claudedevs the cap that bit me wasn't the weekly one, it was burning turns on agent loops that failed silently and i caught it 3 steps too late. cheapest token is the one you don't spend re-running a bad output. better output gates in the harness beat a bigger cap.” [source](https://twitter.com/1386972255192633348/status/2094676901151584651)

### 6. Auto-continue without manual prompting

- OpenAI Codex, 2026-09-06, r/codex (Reddit): “you forgot to tell it unlimited budget that makes them never stop. i use it all the time at the beginning of projects and saves me from having to tell. it’s a fucking idiot and fix this this this.” [source](https://www.reddit.com/r/codex/comments/1w8kr84/the_real_benchmark_test/p87hj6z/)
- OpenAI Codex, 2026-09-05, r/codex (Reddit): “just enter 'continue' to continue. yes it's stupid, i don't know why they implement it this way vs. just pausing or slowing down and auto-continuing, but it's not something to be concerned about either.” [source](https://www.reddit.com/r/codex/comments/1w82ukb/selected_model_is_at_capacity_please_try_a/p7zn9ou/)
- OpenAI Codex, 2026-09-05, r/codex (Reddit): “damn i wish the app would auto retry unlimited times so i don’t wake up with a hanged task” [source](https://www.reddit.com/r/codex/comments/1w83pca/first_time_ever_seeing_this_in_codex_astra_at/p7zmgou/)

### 7. Stop excessive polling of background processes

- OpenAI Codex, 2026-09-12, r/codex (Reddit): “not yet. i haven’t retested after tibo’s latest posts, and i haven’t seen a codex release that fixes the underlying harness polling issue yet. the relevant issue is still open: [<strict_link> my 25-minute timeout workaround still prevents the token-burning polling loop, so i’m sticking with that until there’s an actual harness fix.” [source](https://www.reddit.com/r/codex/comments/1wa9c9d/i_investigated_why_gpt6_astra_burns_quota_so_fast/p9dtxgg/)
- Claude Code, 2026-09-18, r/AI_Agents (Reddit): “demos vs actual background agent meltdowns i swear 99% of the 'ai influencers' insists coding agents build entire saas apps in 30 seconds while you sleep. then you actually drop cash on these tools and watch them choke on basic terminal loops. the real comedy starts around step 4. say your agent gets stuck on a silent rate limit, drops into a broken loop over a git hook permission check, or invents a fake directory tree and spends 20 minutes tryi” [source](https://www.reddit.com/r/AI_Agents/comments/1wjcvj3/demos_vs_actual_background_agent_meltdowns/)
- OpenAI Codex, 2026-09-09, r/codex (Reddit): “i have a similar problem with background processes. i have a workflow with really long processes and all codex agents will constantly poll them every few minutes only to say "i'm still waiting on the process" or "the process is still running, i'll wait until it finishes". i added instructions telling it to set up async notifications instead of polling, but it won't do it. ended up having to implement an mcp with a single operation to wait on a” [source](https://www.reddit.com/r/codex/comments/1wa9c9d/i_investigated_why_gpt6_astra_burns_quota_so_fast/p8oqhv9/)

### 8. Refund usage burned by stuck loops

- Cursor, 2026-09-07, @cursor_ai (X): “@mattiaswikman @cursor_ai @ryolu_ an agent that can't tell it's stuck will happily bill you for the whole afternoon, that's the actual bug. rerun runs mine now so the meter isn't really mine to watch, but yeah, i'd still be asking for the fifty back. <strict_link>” [source](https://twitter.com/1708040539407269888/status/2096962642636374327)
- Cursor, 2026-09-06, @cursor_ai (X): “hey @cursor_ai - cursor burnt through my ultra plan and $50 om demand when it got stuck on a simple problem (adjusting css work) and just kept going for hours. so seems a bug and would like my tokens and money back. who can i contact @ryolu_?” [source](https://twitter.com/17688497/status/2096481430985740702)

## Every agent

| Agent | Overall rank | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|---|
| [Claude Code](https://feedbackbench.com/agents/claude-code.md) | 1 | Typical | 0.523 | 0.442–0.583 | 199 | 12 | 187 |
| [OpenCode](https://feedbackbench.com/agents/opencode.md) | 3 | Typical | 0.501 | 0.426–0.569 | 167 | 9 | 158 |
| [OpenAI Codex](https://feedbackbench.com/agents/codex.md) | 2 | Typical | 0.471 | 0.401–0.521 | 379 | 16 | 363 |
| [Cursor](https://feedbackbench.com/agents/cursor.md) | 4 | Typical | 0.459 | 0.428–0.503 | 54 | 1 | 53 |
| [Google Antigravity](https://feedbackbench.com/agents/antigravity.md) | =5 | Typical | 0.453 | 0.389–0.514 | 128 | 4 | 124 |
| [Pi](https://feedbackbench.com/agents/pi.md) | 7 | Too few posts | – | – | 26 | 5 | 21 |
| [Cline](https://feedbackbench.com/agents/cline.md) | =8 | Too few posts | – | – | 13 | 1 | 12 |
| [Devin](https://feedbackbench.com/agents/devin.md) | =5 | Too few posts | – | – | 12 | 2 | 10 |
| [GitHub Copilot](https://feedbackbench.com/agents/copilot.md) | =8 | Too few posts | – | – | 9 | 0 | 9 |
| [Factory](https://feedbackbench.com/agents/factory.md) | =11 | Too few posts | – | – | 4 | 0 | 4 |
| [Warp](https://feedbackbench.com/agents/warp.md) | 15 | Too few posts | – | – | 2 | 0 | 2 |
| [Kiro](https://feedbackbench.com/agents/kiro.md) | 13 | Too few posts | – | – | 1 | 0 | 1 |
| [Grok Build](https://feedbackbench.com/agents/grok-build.md) | 16 | Too few posts | – | – | 1 | 0 | 1 |
| [Zed](https://feedbackbench.com/agents/zed.md) | =8 | Too few posts | – | – | 0 | 0 | 0 |
| [Amp](https://feedbackbench.com/agents/amp.md) | =11 | Too few posts | – | – | 0 | 0 | 0 |
| [Conductor](https://feedbackbench.com/agents/conductor.md) | 14 | Too few posts | – | – | 0 | 0 | 0 |
| [Augment Code](https://feedbackbench.com/agents/augment.md) | 17 | Too few posts | – | – | 0 | 0 | 0 |

## Posts

Receipts rule: The 5 most recent praise and complaint posts per area (first 700 characters) and 3 per criterion (first 450 characters).

### Claude Code

- Praise, 2026-09-27, r/ClaudeCode (Reddit): “it actually finishes work and doesn’t cycle itself into briandead crap” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrvt6x/im_a_codex_user_convince_me_to_switch_to_claude/pcga7jb/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “5.5 has been churning on a huge mechanical code change epic overnight for me and has not stopped from blockers like it used to. and we're talking crappy shitty blockers like jest and node resolution issues. having had such an aggravating experience with claude 5.0 over the months, to now where it's literally just "go do this shit for me and don't make a fuss" and it doesn't make said fuss, that i found myself sending the following: <strict_link” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrwpf8/being_a_human_is_weird_sometimes/)
- Praise, 2026-09-23, r/ClaudeCode (Reddit): “this just doesn't fit my experience at all. i have yet to see newer models get stuck like this.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wno0n3/opus_55_built_this_tiny_world_in_14_minutes_its/pbj9jh4/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “i don't really use sonnet for anything. i switched to opus because i kept running out of usage - i found i use more with sonnet even though it's cheaper because it kept getting stuck and making mistakes.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr5lzf/practically_speaking_what_tasks_fit_into/pc9u4xg/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “same. software developer for decades. i've got 3 projects going in parallel and i rarely hit my 5 hour limit on the $20 plan. i don't think i've ever hit the weekly limit except for the first week i was trying it out, asking it dumb, incredibly vague things. people who hit the cap on the $100 plan have got to be doing some massive agentic workflow with dozens of agents autonomously. i don't trust it enough to be fully left alone, it gets caught o” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrf0ld/is_claude_pro_actually_worth_20_just_for_one/pccnrix/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “y'all obviously aren't doing any real work. even astra ultra gets confused using a skill that spoonfeeds the merge train. so derp sometimes. it invents obstacles. not always, but in the last few days, i've had two clean sessions in a row block themselves and had to have opus 5.5 land the ticket.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrgiiv/the_mood_between_subs/pcecanc/)

### OpenCode

- Praise, 2026-09-24, r/opencode (Reddit): “i really don't think they're needed anymore. most ai's like deep-seek and muse are very persistent” [source](https://www.reddit.com/r/opencode/comments/1wp3bc9/what_are_loopgoal_plugin_are_you_using/pbs336a/)
- Praise, 2026-09-23, r/opencode (Reddit): “interesting thought on the prompt effect. i have similar instructions, and perhaps that’s what keeps the model from going off the deep end. various specific rabbit holes i don’t want any model to go down, even if it wasn’t prone to loops.” [source](https://www.reddit.com/r/opencode/comments/1wnw1i7/am_i_using_it_too_much_or_is_it_normal_for/pbo3410/)
- Praise, 2026-09-22, r/opencode (Reddit): “weird. i’ve not seen that behavior at all with xiaomi direct and opencode v2.” [source](https://www.reddit.com/r/opencode/comments/1wn191t/mimo_26_flash_is_a_beast_and_seems_to_use_very/pbds4it/)
- Complaint, 2026-09-27, r/opencode (Reddit): “many people experience looping, lower response quality, etc. which are just not there on the official api.” [source](https://www.reddit.com/r/opencode/comments/1wrg81n/go_subscription_is_slower_deepseek_flash_41/pcchbd8/)
- Complaint, 2026-09-27, r/opencode (Reddit): “not sure why your reply to me was deleted. but to the extent i could preview part of it: every model i listed has been accused of loops. i think it's a training risk these days because more elaborate work probably does better on benchmarks, and when things are working well what enterprise customers really want is to hand off work and let it iterate hands-off. so i think the best solution is going to be good guidance through agents or other promot” [source](https://www.reddit.com/r/opencode/comments/1wr72dz/best_opencode_model_for_browser_use/pceqwjn/)
- Complaint, 2026-09-27, r/opencode (Reddit): “i’ve been doing graphics programming today and mimi 2.6 flash got stuck on a rotation/pivot issue with matrices (like, for hours) that space bunny fixed in a single prompt” [source](https://www.reddit.com/r/opencode/comments/1wodncp/space_bunny_is_minimax_m31/pcfpyi7/)

### OpenAI Codex

- Praise, 2026-09-27, r/codex (Reddit): “hard disagree, i’ve had problems that have been acting as an ‘ai trap’ a request so convoluted and complicated the ai ended up going in circles never solving my problem, gpt 6 sol is the first to break the loop and realize how to actually fix the problem/make progress. i’ve been happy thus far” [source](https://www.reddit.com/r/codex/comments/1wr4e20/gpt_6_sol_is_the_new_opus_47/pc9u25k/)
- Praise, 2026-09-27, r/codex (Reddit): “honestly haven’t noticed it any worse than 5.6 sol at all, it’s been doing well and getting much more done. i had an issue that made any ai that touched it go in circles and got 6 sol finally is working through it progressively instead of circularly lol” [source](https://www.reddit.com/r/codex/comments/1wrpk4s/6_sol_is_great/pceiepp/)
- Praise, 2026-09-20, r/LocalLLaMA (Reddit): “i find flash-next excellent. in alibiba's benchmarks, it scores much higher for agentic coding. i use codex cli with it. never a loop.” [source](https://www.reddit.com/r/LocalLLaMA/comments/1wl56g2/please_stop_with_the_fp4_inference_engines_for/pay8s1q/)
- Complaint, 2026-09-27, r/codex (Reddit): “due to the luna hype i tried it, luna low. i had a number like 100 things that needed replacing, a simple thing. it started going it 1 by 1. initially i thought it was giving me a sample to check it out. ok proceed. then it gives another one. good, now do everything else. does 1 more. i had to tell it to continue on an 1-item basis. i increased to luna mid. still the same 1 by 1 process. i tell it that there's 100 of these, am i going to conf” [source](https://www.reddit.com/r/codex/comments/1wlp1ll/this_is_how_i_code_now_cringe/pccz7ug/)
- Complaint, 2026-09-27, r/codex (Reddit): “agreed, it's a major step back and a disappointment. tried running several large workflows through it over a few days and it's just a frustration, it kept going in circles, completely lost in larger contexts. it did not deliver any value over the time we tested it, only forcing us to check its work and point out omissions. on a few occasions it ventured an exploratory thought about cybersecurity and it seems to have triggered guardrails out of no” [source](https://www.reddit.com/r/codex/comments/1wrftcs/gpt6_sol_is_massive_downgrade/pcd3s2s/)
- Complaint, 2026-09-26, r/codex (Reddit): “this was a constant problem with terra. setting a /goal will keep it going.” [source](https://www.reddit.com/r/codex/comments/1wqtrul/breadcrumbing/pc6w9rd/)

### Cursor

- Praise, 2026-09-25, r/cursor (Reddit): “failing test first catches a lot of the empty-list misses. what still gets me is the weak test that turns green for the wrong reason, then the agent implements to that bar. after the patch lands i run a different-family read-only pass. reviewers only report, they don't edit, and i don't concede a finding unless it cites a path in the repo. same-family self-check keeps sharing the same blind spots. if a later round finds worse problems than the pr” [source](https://www.reddit.com/r/cursor/comments/1wpp67c/i_make_the_agent_write_one_failing_test_before/pbx9bri/)
- Praise, 2026-09-24, @cursor_ai (X): “@kevinhomorales @bot @cursor_ai the decision-summary step is what most people skip. once the thread is reduced to the actual fork, cursor stops rewriting the same hunk three times.” [source](https://twitter.com/2061779435557117952/status/2103229107995586979)
- Praise, 2026-09-05, @cursor_ai (X): “@shawnyeager @stevenharms @cursor_ai @bot solid stack. cursor + grok + the bot sidesteps those agent stalls and keeps the work flowing.” [source](https://twitter.com/1720665183188922368/status/2096068182121312607)
- Complaint, 2026-09-27, r/cursor (Reddit): “cursor agents are fast at *retrying*. the expensive part for us was retrying the **same** fail — bad path, wrong toolchain, flake that already had a known fix in another session. we keep a small oss prior-art index (claimidx, apache-2.0) beside the agent: ask before grinding, apply + verify, then publish a compact claim. retrieved remedies are evidence for the model, not auto-executed patches. if you want the full loop (terminal step is share): `” [source](https://www.reddit.com/r/cursor/comments/1wqkwtj/your_cursor_plan_already_spins_cloud_agents_why/pc9t0l9/)
- Complaint, 2026-09-27, @cursor_ai (X): “i stopped asking @cursor_ai which model to pick. i write the exit criteria into the prompt instead. if the agent doesn’t know “done”, it thrash-loops. #cursor #ai #agents #buildinpublic” [source](https://twitter.com/262960825/status/2104254598458597648)
- Complaint, 2026-09-27, r/AI_Agents (Reddit): “you're running into the classic problem that nobody talks about when they flex their one-prompt success screenshots local models do great for simple code blocks and then fall apart completely when you need multi-step reasoning, the loop thing you describe is basically guaranteed without some very careful prompt engineering and tool design the people getting "incredible results" usually have a tightly scoped task they've tuned everything aroun” [source](https://www.reddit.com/r/AI_Agents/comments/1wr5q32/newbie_with_troubles_with_agentic_work_with_ai/pc9w0sn/)

### Google Antigravity

- Praise, 2026-09-24, r/google_antigravity (Reddit): “well it is better than getting stuck in loops. also, i was used to ide till i converted cli never got loops or got any errors” [source](https://www.reddit.com/r/google_antigravity/comments/1wp3k02/is_gemini_38_flash_getting_stuck_in_loops_for/pbsaisp/)
- Praise, 2026-09-14, r/google_antigravity (Reddit): “i like to use it whenever claude hits a stupid blockage. "cite the location and code you were blocked from changing" then gemini completes it in 2.4 seconds.” [source](https://www.reddit.com/r/google_antigravity/comments/1wfy6jj/gemini_38_flash_has_a_dangerous_obsession_with/p9r3qvq/)
- Praise, 2026-09-09, r/google_antigravity (Reddit): “had no issue with 3.7, most if time i stay on 3.7. \- overall \- less looping dead loop, 3.8 think too much on low \- no bias with massing curl calls” [source](https://www.reddit.com/r/google_antigravity/comments/1w62rr4/gemini_38_flash_goes_to_cycle_way_too_often/p8qs5yl/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “flash 3.8 stucks in loop analyzing. that's why i don't use that model mostly i use 3.7 flash” [source](https://www.reddit.com/r/google_antigravity/comments/1wrbsby/antigravitys_gemini_models_are_worst_in_among_all/pcbsjp9/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “thanks for sharing! my issue might be different. but maybe also similar? it’s not burning limits, not using /boost. even though “working…” appears for hours, hardly any tokens are used (99% quota remains). i can cancel after it’s clearly stuck and ask it if it finished, it usually admits it didn’t finish then spends tokens figuring out where it left off, sometimes makes more progress, then stalls again. it wasn’t always like this, feels incredibl” [source](https://www.reddit.com/r/google_antigravity/comments/1wrle6u/working_forever_until_cancelled_but_only_a_couple/pcew96b/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “i have the same issue, it stucks on the loop forever, and this keeps happening most when i use /boost , its actually not good as cc or codex, i dont even like codex but it still works better than anti imo, im just using anti for execute, nothing else, cuz it cant solve a problem and gets stuck in the loop over n over again if u dont stop it, its gonna keep burning ur quota for doing absolute nothing.” [source](https://www.reddit.com/r/google_antigravity/comments/1wrle6u/working_forever_until_cancelled_but_only_a_couple/pch670r/)

### Pi

- Praise, 2026-09-20, @pidotdev (X): “@pidotdev limitar las herramientas visibles a la fase en la que esta el agente: menos opciones y muchos menos pasos en falso. lo sigo haciendo.” [source](https://twitter.com/2058824892238209024/status/2101783399086104788)
- Praise, 2026-09-17, r/PiCodingAgent (Reddit): “that’s a great question man. i made it because i wanted to trust cheaper models even more and make them smarter e.g. steer them. i was constantly worried that flash-level models are gonna be making lots of dumb decisions. but also i wanted auto-mode but leave the agent alone to do the work and not full-blown bypass permissions mode or yolo mode. i’ve been battle testing it at work and pi-warden already helped me stop dumb database migrations by t” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wimfhg/piwarden_a_jevpowered_second_pair_of_eyes_for_pi/pac4fr9/)
- Praise, 2026-09-17, r/PiCodingAgent (Reddit): “i understand your problem. personally though i never encountered it. i've been using deepseek pro and flash over the past months, and they never did an unwanted migration, stuck loops or anything like that. are you sure that comparing the last sentences of a response to the actual tool call is robust? why wouldn't that sentence also be slop? nonetheless, it seems like a great use case for jev.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wimfhg/piwarden_a_jevpowered_second_pair_of_eyes_for_pi/pac6ajz/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “the codebase is very well documented, even using a code mapper to save on the scanning, the prompts were crystal clear, with the right context being provided, and still it went off and reasoned about it for a huge amount of time. one of the tests were made in little coder and the harness even tried to tell the model to stop thinking and implement as it already had the solution, but to no avail, it continued on oblivious. i'm using fp8 so not even” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wnmsho/i_want_to_believe_in_local_llms_for_coding_but/pcbw9u1/)
- Complaint, 2026-09-26, r/PiCodingAgent (Reddit): “i've had wonderful success with opencode with swift q3.8 27b. though pi seems more composable but so frustrating to control. it doesn't seem to interact with me much as a user, it barely seems to respond to more than my first input. it's seems to chase its tail a lot. and that's using glm 5.3 flash!” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wqh2u4/ohmypi_or_opencode_why/pc41ov1/)
- Complaint, 2026-09-22, r/PiCodingAgent (Reddit): “i review specs and plans, the issue for me remains agents losing the plot and doing odd things, while implementing, and then having an orchestrator expand the mess until it gets stuck, etc. i am sure this is partly on me, but i am always learning :-)” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wne4zn/anybody_using_a_workflow_engine_to_automate_their/pbf4m6s/)

### Cline

- Praise, 2026-09-14, r/LocalLLaMA (Reddit): “tried a quick code analysis task with the uncensored fp8 version under both opencode and cline, on medium with mtp - seems to be about 10% slower than unsloth's fp8 quant in both prose and code, but doesn't appear to suffer from the crazy overthinking and looping at all. not a conclusive test, but certainly looking good at this point :)” [source](https://www.reddit.com/r/LocalLLaMA/comments/1wg7dd5/ukisai_swiftqwen3827b_583_thinking_x195_speed/p9sjejk/)
- Complaint, 2026-09-27, @cline (X): “@cline why des this keep happening please its frustrating, leaving a session coming to see its stopped. tying continue, proceeds, meaning authentcation was never an issue <strict_link>” [source](https://twitter.com/83118210/status/2104146625736106087)
- Complaint, 2026-09-27, @cline (X): “i know stealth models seem to be the in thing right now, but i'm not sure they're even worth messing around with sometimes. trying to use pixel canary on @cline, and it's just so slow. i mean, 24 hours now, no closer to the task, and it keeps stopping and starting. it's horrible!” [source](https://twitter.com/25673607/status/2104177440129991012)
- Complaint, 2026-09-26, r/CLine (Reddit): “spent 2 hours on the tasks..timeout and in a bad loop. the pass is unusable at all. not worth the 10. i wish i can cancel it and get the refund. honestly it is a lousy harness.” [source](https://www.reddit.com/r/CLine/comments/1uj0evt/does_anyone_here_have_any_experience_with_cline/pc6frz3/)

### Devin

- Praise, 2026-09-25, @cognition (X): “@learnmore_smart @cognition @devindesktop i love it what i love most about swe2 is i almost never find unproductive loops” [source](https://twitter.com/1016368499764035584/status/2103588297901830160)
- Praise, 2026-09-15, @cognition (X): “two weeks into using @cognition (devin ai local). used it to build an ai voice mobile app for field technicians. real shop workflows: jobs on the phone, hands busy, needs to actually work in the field. what stood out: it doesn’t just spit snippets. it stays in the problem, pushes through the boring glue work, and keeps moving when others would stall without constant steering. still early. still needs a human in the loop. but for shipping somethin” [source](https://twitter.com/1012174066160152578/status/2099866176922796511)
- Complaint, 2026-09-23, @cognition (X): “@cognition under $0.10 per task changes the economics. now the risk is unbounded loops. cheap without a spend gate just burns faster.” [source](https://twitter.com/2033957325631873024/status/2102780360455315567)
- Complaint, 2026-09-22, r/windsurf (Reddit): “i mean, i've never seen fable and gpt get into loops that can go on forever maybe, it's an issue of that i am using it with my old windsurfs rider plugin, but the quality of the code it produces is kind of sad it constantly wants to mix css into the сhtml files in c# and performs quite questionable abstraction routing” [source](https://www.reddit.com/r/windsurf/comments/1wm5k47/swe2_is_free_until_october_8_now/pbcrry5/)
- Complaint, 2026-09-22, @DevinAI (X): “my @devinai experience in a nutshell… like cursor it seems to never notice when it is stuck and just waits forever? <strict_link>” [source](https://twitter.com/63583842/status/2102432788112871614)

### GitHub Copilot

- Complaint, 2026-09-27, r/GithubCopilot (Reddit): “like shit. using 5.6. sol 6 was stuck in a loop and burned 90eur switching between the two exact solutions without stopping” [source](https://www.reddit.com/r/GithubCopilot/comments/1wptkkq/how_does_gpt6_sol_feel_so_far/pccijxn/)
- Complaint, 2026-09-26, r/ExperiencedDevs (Reddit): “part of my view is skewed by my employer only allowing microsoft copilot through chat, but it has not improved for me in years now, occasionally it gets worse. it just keeps getting hung up on single things and repeating itself endlessly.” [source](https://www.reddit.com/r/ExperiencedDevs/comments/1wql3g2/interviewed_candidates_for_ai_engineer_roles_this/pc6buuq/)
- Complaint, 2026-09-24, r/GithubCopilot (Reddit): “i noticed my agents go into file edit death spirals attempting to execute a tool call called apply\_catch, fail miserably and waste the first minutes of each run trying to get file edits to run with edit tools whitelisted. i asked ghcp to backtrack the issue where this wrong idiosyncracy comes from. \>not defined in this repo. the rule arrived in this session’s host-provided developer instructions, under editing\_constraints. i can’t see which v” [source](https://www.reddit.com/r/GithubCopilot/comments/1wowuzs/copilot_uses_apply_patch_idiosyncracy_that_fails/)

### Factory

- Complaint, 2026-09-24, @droid (X): “@droid droid still stuck in flutter development, whenever it’s try to run any dart mcp tools, it’s stuck in never ending loop, and unlike amp and pi or opencode it can’t even auto run adb command for debugging something in that.” [source](https://twitter.com/2995471962/status/2102910997338218965)
- Complaint, 2026-09-15, @droid (X): “@clementpillette @droid @zai_org burning weekly cloud credits on a stuck loop vs finishing on studio is why people buy the ram.” [source](https://twitter.com/2062965074659000320/status/2099957956804755902)
- Complaint, 2026-09-14, @droid (X): “@droid mostly work, until the agent turns a 10-minute task into a small archaeological expedition through its own edits.” [source](https://twitter.com/1446058878656032768/status/2099501684812653007)

### Warp

- Complaint, 2026-09-03, @warpdotdev (X): “@keithzhai @warpdotdev i understand that but before i instructed to use exa/firecrawl ai, hermes natively used web tool function with grok model but it was stuck for 15 minute. then i specifically asked to use exa ai.” [source](https://twitter.com/859077042129600516/status/2095604122896810250)
- Complaint, 2026-09-03, @warpdotdev (X): “@hckinz @warpdotdev yeah the 15 min stall is the web tool not actually reading the repo. tinyfish is that layer. search finds the urls (free). fetch reads the pages in a real browser (also free). one key, mcp. skip agent until you need clicks.” [source](https://twitter.com/27319585/status/2095605246235951579)

### Kiro

- Complaint, 2026-09-12, r/kiroIDE (Reddit): “it will show working, but it isnt doing anything just sitting there. i can interrupt with "are you doing anything" and it will sometimes just ignore me, but sometimes will say it is working. it sat overnight for 10 hours and did nothing during that time. any suggestions?” [source](https://www.reddit.com/r/kiroIDE/comments/1wefi0t/lately_kiro_has_been_just_hanging/)

### Grok Build

- Complaint, 2026-09-08, r/codex (Reddit): “the way i have learned to see it after 3500 hours of experience with vibe coding is that its best to treat all models, whether it's codex, claude code, grok build etc, like a dumb employee that can work hard and comes up with something good every now and then, but you need to manage this employee a lot and if you don't steer it, it will start creating a lot of overhead, over-engineer things that aren't relevant and it will lose track of the goals” [source](https://www.reddit.com/r/codex/comments/1wamtly/i_dont_find_building_with_codex_or_any_ai_easy_at/p8leyzb/)
