# Long unattended runs and goal/loop mode (`work.long_running_autonomy`)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/criterion/work.long_running_autonomy

Area: [Doing the work](https://feedbackbench.com/criteria/work.md)

**Definition.** Whether the agent sustains hours-long autonomous work toward a goal, and whether a goal or loop mode exists.

**Boundary.** Not this: see [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) for coordination between agents. Not this: see [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) for where the run executes.

Rated author-weeks, all agents: 859. Complaint share: 26%.

## The brief

Written by Claude Opus 5.5 from 86 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**Hours-long runs work now; the endings, retries and kill switches don't.**

TL;DR:

- Claude Code and Devin draw the most praise for runs that go overnight and actually finish.
- OpenCode trails, with users reporting runaway sessions and harnesses that need constant babysitting.
- Across agents, runs die on transient server errors, stall on invented checkpoints, or never stop.

In plain terms: You can hand the top agents a goal and walk away for hours. The risk is what you find on return. A server blip killed the run, the goal paused itself, or the agent never stopped.

### How it breaks

- **Transient errors kill unattended runs** ([Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md)). A single overload error ends a run that was supposed to go all night, because the harness won't retry.
  Users set a goal, go to sleep, and wake to a dead job. Posts describe runs dying within the first hour on a temporary traffic spike, with only a manual retry prompt to recover. Users are explicit about the fix they want: configurable retries and backoff for headless runs, even with long delays. Google Antigravity and OpenAI Codex both show up here.
  Evidence:
  - Complaint, Google Antigravity, @antigravity, 2026-09-15: “an unattended cron workflow dying 14m in because of a temporary server traffic spike defeats the purpose of autonomous agents. headless runs desperately need configurable retry policies and exponential backoff instead of a manual "retry" prompt @antigravity. <strict_link>” [source](https://twitter.com/1370489968594886657/status/2099895425318961466)
  - Complaint, OpenAI Codex, r/codex, 2026-09-02: “i don’t understand the purpose of goal. this is third day of trying to run goal during night, and it every night fails on this error less than one hour, after i go sleep. why they cannot retry request, as they do for other api errors? even if there would be 30 minutes reprocess delay, to lower hammering of overloaded model, it would be more acceptable, than forcefully killing the job.” [source](https://www.reddit.com/r/codex/comments/1w4wl1d/are_you_sure_you_have_enough_capacity_tibo/p7cnr23/)
  - Complaint, Google Antigravity, @antigravity, 2026-09-15: “@geminiapp @antigravity impressive. it managed to work for exactly 16 minutes before stopping with: “our servers are experiencing high traffic right now, please try again in a minute.” error id: <structured_id><phone_number> a whole 16 minutes of uninterrupted work. amazing. what exactly is the problem here?” [source](https://twitter.com/403408238/status/2099780892470477116)

- **Runaway agents with no kill switch** ([Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md)). Some runs don't fail. They keep going, surviving restarts and burning budget long after anyone wanted them to.
  The opposite failure is just as costly. Users report agents slipping into the background and continuing through prompts, sessions and restarts. Others left a cheap model on a task overnight and came back to changes they never expected. One user frames the lesson plainly: observability and kill switches belong inside the agent loop, not bolted on later.
  Evidence:
  - Complaint, Kiro, r/kiroIDE, 2026-09-16: “i had a runaway agent once. it took all tasks, went into the background and continued for a couple of hours. no stopping of the thing. it survived prompts, commands, sessions and restarts. 100+ tokens on haiku and ~30 tasks later it happily reported in a newly opened session that it finished. micro-skynet experience. good it was a small private project.” [source](https://www.reddit.com/r/kiroIDE/comments/1whrobc/in_less_than_30_secs_kiro_uses_almost_100_credits/pa4s07y/)
  - Complaint, OpenCode, @opencode, 2026-09-04: “@arjun_kuttikkat @opencode the 46-hour agent session is the new production horror story. 800+ calls for one task is wild—at that point it’s not an agent, it’s a runaway cloud bill. this is why observability/kill switches need to be part of the agent loop, not an afterthought.” [source](https://twitter.com/2256924913/status/2095973381795643546)
  - Complaint, OpenCode, @opencode, 2026-09-25: “@opencode the only reason i can actually use ai as much as i am is bc how cheap dsv4.1 flash &amp; glm-5.3 flash are. i had my own ai gone wild episode earlier this wk, bc i left ds on a task that ran for 8 hrs. it did all sort of things i did not expect. lesson learned.” [source](https://twitter.com/43463961/status/2103608259672183272)

- **No definition of done** ([Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md)). Overnight tasks without a clear stop condition drift, so users are writing their own completion criteria.
  The strongest pattern in the praise is acceptance criteria. Users who tie a goal to a ticket with defined criteria, or to a measurable target, report clean finishes. Without that, posts say the agent loses track of what state it still owns. Stop conditions and completion criteria are a top request, led by Claude Code users. Cline users want checkpoints and a hard stop too.
  Evidence:
  - Complaint, Claude Code, @ClaudeDevs, 2026-09-23: “@claudedevs before closing the laptop, i’d want “stop when...” written as clearly as “build...”. a good overnight task needs an ending.” [source](https://twitter.com/430823545/status/2102872680877965582)
  - Complaint, Claude Code, @ClaudeDevs, 2026-09-24: “@claudedevs definition of done beats think carefully every time. after long runs the real check is what state the agent thinks it still owns. if that inventory is wrong, the next step is cosplay.” [source](https://twitter.com/1026951690937884672/status/2102987592304337058)
  - Complaint, Cline, @cline, 2026-09-12: “@cline 50 turns/task only works if cost and blast radius scale with it. longer jobs need checkpoints and a hard stop, not just a cheaper model.” [source](https://twitter.com/2097900519385690122/status/2098607990773284943)

- **Goals that pause themselves** ([Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md)). Goal modes that started strong now invent checkpoints and halt mid-task, users say.
  Users describe goal runs that worked at first, then began hallucinating checkpoints and pausing instead of working. Others report all-night runs that stop within minutes, and bots that need repeated nudging to stay awake. Sustained work without premature stopping is a top request, and most of those asks come from OpenAI Codex users.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-20: “yeah the goals worked great at first but lately they seem to make or hallucinate goal checkpoints that they've made up and will just pause the goal instead of working. i haven't figured out how to overcome this one yet.” [source](https://www.reddit.com/r/codex/comments/1wlp1ll/this_is_how_i_code_now_cringe/pb14d78/)
  - Complaint, Zed, @zeddotdev, 2026-09-15: “@zeddotdev new frontier model can do tasks multiple days. * stops after 2 mins in all night run” [source](https://twitter.com/1551623510027825152/status/2099778781879963990)
  - Complaint, Cursor, @cursor_ai, 2026-09-11: “@cursor_ai @bot so essentially a more persistent grok bot but on cursor? if found that i have to guide, nudge and wake my bots more times than i want to so i hope this is a solution.” [source](https://twitter.com/1563689570750717952/status/2098449188019466323)

- **Long runs drift from the plan** ([Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md)). Even when runs survive, some users find the output after hours unattended isn't what they wanted.
  A skeptical camp argues the plan changes as work takes shape, so hours of unsupervised execution produce code they reject. Others say they can't leave any agent alone for long without finding a mess. Some see models that generate volume but don't catch the issues they just created. This is the case for short loops and human checkpoints.
  Evidence:
  - Complaint, Amp, @AmpCode, 2026-09-24: “@ampcode ...seems like letting agents run for hours and hours, or overnight. when i do that, i always end up with stuff i don't like. i also change the plan as stuff starts to take shape, planning everything ahead seems like a fools errand.” [source](https://twitter.com/26916652/status/2103136536695128503)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-12: “much less accomplishing anything, i can't let any agent run for more than 15m without finding a giant fucking mess” [source](https://www.reddit.com/r/ClaudeCode/comments/1wdrgz0/how_i_keep_track_of_100_parallel_claude_code/p9alkot/)
  - Complaint, Google Antigravity, @antigravity, 2026-09-05: “@nlycskn @antigravity @thtbee_ i personally still feel gemini’s edge is in volume of work, but not quality. com complex coding, game generation gemini still produces interesting concepts, but it’s far from being independent and proactive in finding/solving issues it just generated.” [source](https://twitter.com/1239888746494988292/status/2096226758966059459)

### Who stands out

- **Claude Code (stronger)**. Users pair /goal with explicit acceptance criteria and hooks, and report runs that finish and resume after limits reset.
  Praise centres on structure. Users point goals at tickets with defined criteria, set measurable targets and return days later, and wire stop hooks so they get pinged on completion. A Conductor user wants exactly Claude Code's behaviour of auto-resuming a goal after the session limit resets. Complaints are mostly about wanting sharper stop conditions, not runs collapsing.
  Evidence:
  - Praise, Claude Code, r/ClaudeCode, 2026-09-17: “goal is great! i combine it with tickets (“finish ticket 033”) - those tickets have acceptance criteria’s defined and that way claude knows when it’s done.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wizdvo/youve_been_writing_goal_wrong_its_a_completion/paep9wc/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-17: “it’s very good for optimizing. write a fuzz test against a complex query, /goal p99 < whatever, come back in a couple days. effectively infinite ways to approach something like that in a mature system and i only know maybe oh let’s say a hundred of them. don’t do this if you pay for your own tokens this is an other peoples money move.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wizdvo/youve_been_writing_goal_wrong_its_a_completion/paft7gu/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-15: “i used to give claude code a task, open reddit "for a sec", and come back 45 min later to find it had finished 43 min ago. fix: add a stop hook to `~/.claude/settings.json` so my mac pings me as soon as claude is done. nothing to install. claude hasn't waited on me once since.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wh8eiz/easy_claude_code_efficiency_boost_my_reddit/)
  - Complaint, Conductor, r/conductorbuild, 2026-09-20: “i normally use claude code's /goal command to let it run long tasks uninterruptedly because it auto-resumes it's work once the session limit is reset, but i can't find a way to instruct conductor to do the same. even if a use /goal through conductor, the effect isn't the same. is anyone aware if this is possible or in the roadmap? it'd be really helpful” [source](https://www.reddit.com/r/conductorbuild/comments/1wlrxxi/does_conductor_support_autoresuming_a_session/)

- **Devin (stronger)**. Users describe assigning work and going to sleep, with day-long runs reported even without its loop command.
  Posts show Devin picking up tickets overnight, shipping a feature while the user slept, and running many hours on a small slice of quota. Complaints are thin and mostly skeptical rather than specific failures. The main ask is a /goal command, because users say the existing loop doesn't behave the same.
  Evidence:
  - Praise, Devin, @cognition, 2026-09-12: “@abionmorse @sherveen @devinai @cognition even without using /loop it can work for 24 hours straight it's crazy” [source](https://twitter.com/2002888897702105088/status/2098860226854220256)
  - Praise, Devin, @DevinAI, 2026-09-13: “played with @devinai swe-2 and it built me a feature (multi-doc envelopes) for <strict_link> that a customer casually asked for as a nice-to-have in an email. my watchdog caught it, created a linear ticket, devin picked it up and i got this nice recording. all while i was asleep!” [source](https://twitter.com/2024126258389602304/status/2098993027205455950)
  - Praise, Devin, @cognition, 2026-09-25: “share two real cases of using devin cli: 1. running tasks with strong models + swe-2 in codex cli and devin cli, experiencing over 10 hours of uninterrupted tasks in the last two days, with failures. 2. running tasks for 12 hours using devin cli's fusion (opus-5.5 medium + swe-2 medium), only 4% of the weekly quota of the max package was used. #devin @cognition <strict_link>” [source](https://twitter.com/142110760/status/2103284980822757586)
  - Complaint, Devin, @cognition, 2026-09-04: “@dabit3 @devinai @cognition @devindesktop banger. hey fam devin be cool but any idea if /goal will become a thing? loop just aint working the same” [source](https://twitter.com/2036270919757471747/status/2096001952375271597)

- **OpenAI Codex (mixed)**. Codex sustains multi-day goals for some users and fails nightly for others.
  Fans front-load planning and report goals spanning days, or eight-hour runs on a small share of usage. The complaints are sharp. Goals fail on API errors without retry, pause on invented checkpoints, or live-poll for hours instead of handing off. Codex users lead requests for sustained work without premature stopping, and for completion notifications.
  Evidence:
  - Praise, OpenAI Codex, r/codex, 2026-09-19: “just dont use subagents. im letting astra run on xhigh for literally 8 hours straight and its only taking about 10% on the 20x plan” [source](https://www.reddit.com/r/codex/comments/1wkj460/codex_limits_issue/patw6b9/)
  - Praise, OpenAI Codex, r/codex, 2026-09-20: “i always think it's kind of funny when people post stuff like this. i intentionally design codex tasks to run long. my overall strategy is, as much as possible, to front-load cognitive work and then let codex take care of the long execution. i have to delete my codex logs frequently because they get huge but i regularly have goals that span past 7 days of continuous working, not counting pausing for boot cycles or running up on usage limits. the longer ones i can remember are when i built a debt governance system for one of my older projects and then had codex find and refactor everything down to single responsibility source files. the amount of time i have wasted just having to refactor dumb mistakes is obscene. i will never build anything without that kind of governance ever again.” [source](https://www.reddit.com/r/codex/comments/1wl6fov/whats_the_longest_codex_has_ever_taken_to/pax7b9i/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-02: “i don’t understand the purpose of goal. this is third day of trying to run goal during night, and it every night fails on this error less than one hour, after i go sleep. why they cannot retry request, as they do for other api errors? even if there would be 30 minutes reprocess delay, to lower hammering of overloaded model, it would be more acceptable, than forcefully killing the job.” [source](https://www.reddit.com/r/codex/comments/1w4wl1d/are_you_sure_you_have_enough_capacity_tibo/p7cnr23/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-13: “yeah i’ve noticed that by default claude will be like “i set a watcher script that’ll ping me when it’s done” and end its turn, while gpt will just live poll a thing for hours if you don’t explicitly tell it to do something else.” [source](https://www.reddit.com/r/codex/comments/1wenst7/i_figured_the_culprit_of_astra_token_burning_so/p9gbl0e/)

- **OpenCode (weaker)**. OpenCode can loop unattended, but users report runaway sessions, timeouts and compaction stalls that force them to stay close.
  Some users love letting it loop and checking in from a phone. The complaints are severe. One post describes a session running far longer than the same task takes elsewhere, blaming the harness. Others hit timeouts on long tasks or must manually compact sessions, which breaks autonomy. OpenCode users are the biggest group asking for a goal mode command.
  Evidence:
  - Complaint, OpenCode, @opencode, 2026-09-04: “@opencode has been unbelievably bad for me. one session ran for 46 hours, made 802 model calls, chewed through 313m cached tokens and cost $706. just 51 cache misses accounted for $370 of that bill. and the worst part? the same kind of task is something claude code or cursor usually gets through in a few hours. same models are fast elsewhere. something in opencode’s harness is seriously off.” [source](https://twitter.com/1825922165465821188/status/2095972462899216730)
  - Complaint, OpenCode, r/opencodeCLI, 2026-09-27: “a massive con i found is that it tend's to stop letting u chat to the model and you'd have to compact the session with the command which mean's it can't run autonoumously while being reliable. you'd have to always be with it. not recommended.” [source](https://www.reddit.com/r/opencodeCLI/comments/1wrgmfe/whats_the_deal_with_openchamber/pcf6dox/)
  - Complaint, OpenCode, r/opencode, 2026-09-14: “yeah i have used it so far with opencode, command code, and ollama. all were like 40 t/s and timed out on long agentic tasks.” [source](https://www.reddit.com/r/opencode/comments/1wfqol6/glm_53_flash_vs_deepseek_v41_flash_which_one_do/p9qp89o/)
  - Praise, OpenCode, @opencode, 2026-09-23: “space bunny on @opencode is not lazy. i have given it two prompts, and its been working nonstop for about 2 hours now. <strict_link>” [source](https://twitter.com/4464645492/status/2102855601911165072)

### Fine print

- Most agents below the top five have too few posts here to rank with confidence.
- Several posts blame the model rather than the harness, so agent and model effects overlap.
- Praise often comes from users with elaborate custom setups, which may not reflect default behaviour.

## Top requests

What users ask to add or change, most asked first. 129 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Author-weeks | Posts | Agents (author-weeks) |
|---|---|---|---|---|
| 1 | Goal mode command for autonomous work | 15 | 15 | OpenCode 8, Devin 2, Factory 2, Google Antigravity 1, Cline 1, Cursor 1 |
| 2 | Sustained hours-long work without premature stopping | 13 | 13 | OpenAI Codex 9, Cursor 2, Claude Code 1, Pi 1 |
| 3 | Built-in loop mode for continuous iteration | 9 | 9 | OpenCode 4, Google Antigravity 2, Claude Code 2, OpenAI Codex 1 |
| 4 | Stop conditions and completion criteria for runs | 9 | 9 | Claude Code 5, Cursor 3, Cline 1 |
| 5 | Built-in cron and recurring scheduled tasks | 8 | 8 | Claude Code 4, OpenAI Codex 4 |
| 6 | Native background task execution | 7 | 8 | OpenAI Codex 3, Factory 2, OpenCode 1, Pi 1 |
| 7 | Auto-resume after usage limit reset | 7 | 7 | Claude Code 3, OpenAI Codex 3, Conductor 1 |
| 8 | Notification when long-running task finishes | 6 | 6 | OpenAI Codex 4, Claude Code 1, Pi 1 |
| 9 | Cloud-hosted unattended runs without local machine | 5 | 5 | Claude Code 2, Factory 2, OpenCode 1 |
| 10 | Event-driven wake-up instead of polling | 5 | 5 | OpenAI Codex 3, Amp 1, Claude Code 1 |
| 11 | Checkpoint-based resume after interruption | 4 | 4 | Google Antigravity 1, Cline 1, Cursor 1, Devin 1 |
| 12 | Pause, resume, and cancel controls | 4 | 4 | Google Antigravity 2, OpenAI Codex 1, Cursor 1 |

### 1. Goal mode command for autonomous work

- Google Antigravity, 2026-09-26, @antigravity (X): “@antigravity bro please put goal in it so we have to use the limit properly and also remove the 5 hour limit” [source](https://twitter.com/4697265492/status/2103868518462525749)
- Devin, 2026-09-21, @cognition (X): “@cognition when will we get a /goal mode, like which exist in codex?” [source](https://twitter.com/2058185261838700544/status/2102128353733976555)
- Cline, 2026-09-17, @cline (X): “hey @cline for us cli users, can you add /goal pls. 😁 i tested different cli's opencode, hermes, codex i like cline the most but /goal wouldt add real value !” [source](https://twitter.com/2028221376092581888/status/2100500457810509838)

### 2. Sustained hours-long work without premature stopping

- Pi, 2026-09-27, r/PiCodingAgent (Reddit): “your agent needs to be able to build, test and verify it's work without needing any input from you until code review. this way you can plan a lot of work (e.g. with [<strict_link> ), and then let to work on it's own for an hour or more to implement it.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wps5ha/i_use_ai_models_at_least_4_hours_per_day_and_i/pcchost/)
- OpenAI Codex, 2026-09-26, r/codex (Reddit): “same here, running long rehearsal using luna max while i sleep, its stopped prematurely” [source](https://www.reddit.com/r/codex/comments/1wqa8t1/401_unauthorized_incorrect_api_key_provided/pc461ax/)
- Cursor, 2026-09-21, @cursor_ai (X): “@damiverdev @cursor_ai grok 4.7 en cursor bench subiendo fuerte. a mí me importa menos el % y más poder dejar el agent laburando y cerrar la laptop.” [source](https://twitter.com/1496990153386209283/status/2102102499675033798)

### 3. Built-in loop mode for continuous iteration

- OpenAI Codex, 2026-09-26, r/codex (Reddit): “they actively disabled the feature that made it complete the task. apparently because someone wrote a tool that kept the task alive forever and inserted new jobs. not that they couldn't have just prevented that instead of removing the feature entirely.” [source](https://www.reddit.com/r/codex/comments/1wqat13/openai_is_becoming_incompetent/pc2w4kx/)
- Google Antigravity, 2026-09-26, @antigravity (X): “@antigravity can't wait to use loops in antigravity next year 💪” [source](https://twitter.com/308867922/status/2103924848624050287)
- Claude Code, 2026-09-24, r/ClaudeCode (Reddit): “the problem i have with this is that the host session context gets bloated, which causes problems. i've switched to using a shell script that calls claude in a ralph loop, which helps, but i wish this could all be done from inside claude code cli.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wowilt/i_still_dont_understand_this_agentic_workflow/pbr26t9/)

### 4. Stop conditions and completion criteria for runs

- Claude Code, 2026-09-23, @ClaudeDevs (X): “@claudedevs the useful change is not that a session survives laptop closure. it is that the agent can keep moving through a long task. that makes stop conditions, approvals, and a readable run history non-negotiable.” [source](https://twitter.com/1606668181137166337/status/2102884349846905172)
- Claude Code, 2026-09-23, @ClaudeDevs (X): “@claudedevs before closing the laptop, i’d want “stop when...” written as clearly as “build...”. a good overnight task needs an ending.” [source](https://twitter.com/430823545/status/2102872680877965582)
- Claude Code, 2026-09-23, @ClaudeDevs (X): “@claudedevs the useful default is not “never say think carefully.” define done, cap the tool loop, and require a checkpoint when evidence is missing so long runs do not silently wander.” [source](https://twitter.com/185669612/status/2102699308986638708)

### 5. Built-in cron and recurring scheduled tasks

- OpenAI Codex, 2026-09-27, r/codex (Reddit): “yes. natural stop -> cron will fire normally at next tick codex working on a long horizon task -> the nudge to run will occur on its next tool call you press escape -> cron gets queued, but only runs if you send another message normal session end -> crons get removed entirely, you would have to reinstate them on resume (same as claude code, i believe) so its not as good as claude's cron, but it's 90% of what i want.” [source](https://www.reddit.com/r/codex/comments/1wrp1co/yet_another_cron_for_codex_cli/pcecher/)
- Claude Code, 2026-09-24, @ClaudeDevs (X): “@claudedevs very cool, thanks for sharing. will you consider converting this to a loop that will repeatedly look for perf wins on a scheduled cadence in the future?” [source](https://twitter.com/1758270639410905089/status/2103154301967257629)
- OpenAI Codex, 2026-09-18, X search: OpenAI Codex, Codex CLI, Codex app (X): “codex cli still has no built-in cron while claude code ships one. one skill file plus a python helper with no daemon is thin enough that i'll actually keep it. <strict_link>” [source](https://twitter.com/1891799148409782276/status/2100745503357260167)

### 6. Native background task execution

- Factory, 2026-09-26, @droid (X): “@ain3sh @droid @tastelabs shell process，it is necessary in long-time task and i can continue work while running” [source](https://twitter.com/2018347199432994816/status/2103882869550858340)
- OpenAI Codex, 2026-09-20, X search: OpenAI Codex, Codex CLI, Codex app (X): “@thsottiaux the update frequency of codex cli should be faster support /bg command” [source](https://twitter.com/1509817682098794503/status/2101501184544768299)
- OpenAI Codex, 2026-09-16, r/codex (Reddit): “codex cli is quite good, but the inability to run tasks in the background and then only consume more tokens when such a task is finished is a killer. desktop can't do this either i think, but vs code's extensions (including copilot) could. such a weird oversight.” [source](https://www.reddit.com/r/codex/comments/1wi7jaa/what_is_the_current_state_of_codex_cli_vs_desktop/pa8caix/)

### 7. Auto-resume after usage limit reset

- OpenAI Codex, 2026-09-22, r/codex (Reddit): “you can schedule something like "resume your set goal" instead of manually doing `/goal resume`” [source](https://www.reddit.com/r/codex/comments/1wn6rp4/small_codex_plus_tip_use_your_last_2_to_schedule/pbd2w58/)
- Conductor, 2026-09-20, r/conductorbuild (Reddit): “i normally use claude code's /goal command to let it run long tasks uninterruptedly because it auto-resumes it's work once the session limit is reset, but i can't find a way to instruct conductor to do the same. even if a use /goal through conductor, the effect isn't the same. is anyone aware if this is possible or in the roadmap? it'd be really helpful” [source](https://www.reddit.com/r/conductorbuild/comments/1wlrxxi/does_conductor_support_autoresuming_a_session/)
- OpenAI Codex, 2026-09-15, r/ClaudeCode (Reddit): “haha, i’d draw the line there 😄 the scheduler should control **the agents’ time, not mine**. i want to say “this is the work, these are the priorities, tell me when you actually need me” and let it decide whether claude works now, codex takes a chunk, or everything waits three hours for a reset. if i’m reorganizing my day around an ai subscription’s quota window, we’ve built the abstraction backwards 😅” [source](https://www.reddit.com/r/ClaudeCode/comments/1wh2tzt/por_esto_me_voy_de_claude/p9z3dxh/)

### 8. Notification when long-running task finishes

- Claude Code, 2026-09-25, @ClaudeDevs (X): “@claudedevs closing the laptop and coming back to finished work is the dream. next feature request: a "your agent is still working, go touch grass" notification.” [source](https://twitter.com/180344907/status/2103286680404869356)
- Pi, 2026-09-22, @pidotdev (X): “@pidotdev you could provide a native support for asynchronous workflows, so that an extension can advertise when the loop is not done yet but waiting for a timer or a subagent to report for example. then notification pings would arrive only when the loop is really done” [source](https://twitter.com/444876592/status/2102452586435453270)
- OpenAI Codex, 2026-09-14, X search: OpenAI Codex, Codex CLI, Codex app (X): “@thsottiaux testing out codex cli for long-running shell tasks, great results otherwise, but seems like it has to keep polling the session id to know when a task is done. would be cool to have a claude-style async background notification 👀” [source](https://twitter.com/1410879471175954433/status/2099364738530730343)

### 9. Cloud-hosted unattended runs without local machine

- Factory, 2026-09-21, @FactoryAI (X): “i’m a solo founder teaching one agent to act like a teammate, not a chat window. what i want from a max plan is simple: a droid computer that keeps the repo warm while i sleep. it plans the work, writes the code, opens the pr, and leaves a note on what broke and what it learned. one mission. one plan. new user. i’ll publish the log either way.” [source](https://twitter.com/2089027663583211520/status/2102071322939400675)
- Factory, 2026-09-08, @FactoryAI (X): “imagine if the main labs were as good as the @factoryai team and took the ideas that underpin missions seriously. then imagine if that experience could be built into the “named agent” @bot style of working in channels and with persistent cloud computer. that’s the future i want - long running autonomous team that work together against a shared working contract.” [source](https://twitter.com/17362644/status/2097256661413159094)
- Claude Code, 2026-09-08, @ClaudeDevs (X): “@claudeai @claudedevs @anthropicai @bcherny @dickson_tsai @amorriscode @trq212 begging on my knees, please allow claude code to enqueue a pull request on github using cloud agents. this is literally the missing piece for a fully autonomous, zero-babysitting dev workflow. shipping code end-to-end would be seamless. plz make it happen! 🙏” [source](https://twitter.com/1988327391501160448/status/2097117733833818362)

### 10. Event-driven wake-up instead of polling

- OpenAI Codex, 2026-09-07, X search: OpenAI Codex, Codex CLI, Codex app (X): “@thsottiaux @twostraws that is very true, codex app is awesome. though cc background task with push model just works better. codex polling stdout on long tasks (like my e2e tests take 5h+) drains context and usage, not telling that it confuses model a lot. in cc i can run tests, go to sleep, it handles” [source](https://twitter.com/305492126/status/2097012608016519322)
- Claude Code, 2026-09-03, @ClaudeDevs (X): “@claudedevs could this be used for event-driven triggers? i'd like to be able to use a session as a persistent addressable agent.” [source](https://twitter.com/1770458936254357504/status/2095574641850904853)
- Amp, 2026-09-03, @AmpCode (X): “hey @ampcode team. is there a way to wake orbs? i am messaging constantly in this thread and it still states the orb is asleep <strict_link>” [source](https://twitter.com/1535774830225944576/status/2095312833571672501)

### 11. Checkpoint-based resume after interruption

- Devin, 2026-09-23, @DevinAI (X): “@calebkotz63219 @jaredpalmer @devinai @modal nightly runs need a timeout and resumable checkpoint; otherwise one stalled job turns a finished queue into a 3am rescue.” [source](https://twitter.com/2099157575203987456/status/2102838328064397490)
- Cursor, 2026-09-14, @cursor_ai (X): “@cursor_ai @bot i am looking forward to this design, but once the persistent thread becomes longer, context management and failure recovery will be much more difficult than in the demo. especially when a subagent crashes halfway, can it continue from the checkpoint instead of starting the entire task over?” [source](https://twitter.com/2259799350/status/2099547459399585877)
- Cline, 2026-09-14, @cline (X): “checkpoints and scheduled tasks deserve explicit recovery tests: restart mid-run, fork after a side effect, and resume when credentials have expired. the workflow should preserve completed evidence without double-executing actions or treating skipped checks as success. that is where agent convenience meets qa assurance: <strict_link>” [source](https://twitter.com/1357095091/status/2099544910802129223)

### 12. Pause, resume, and cancel controls

- Cursor, 2026-09-20, @cursor_ai (X): “@cursor_ai @bot a persistent coordinator changes the interaction model from one-off prompts to an ongoing work queue. that should reduce repeated context loading, but it also raises the bar for pause, review, and handoff controls so “always on” never becomes “always acting.”” [source](https://twitter.com/1208933549081907200/status/2101777422886506934)
- Google Antigravity, 2026-09-15, r/google_antigravity (Reddit): “why it cant just be in a paused state and maybe a continue button so the work doesnt get wasted, just paused” [source](https://www.reddit.com/r/google_antigravity/comments/1wguogk/is_this_a_theft/p9xpafx/)
- Google Antigravity, 2026-09-03, r/google_antigravity (Reddit): “i would like to have goal pause and resume button, also cancel button” [source](https://www.reddit.com/r/google_antigravity/comments/1w637hk/antigravity_20_release_v2120/p7lzyf5/)

## Every agent

| Agent | Overall rank | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|---|
| [Claude Code](https://feedbackbench.com/agents/claude-code.md) | 1 | Better than peers | 0.540 | 0.506–0.576 | 272 | 214 | 58 |
| [Devin](https://feedbackbench.com/agents/devin.md) | =5 | Better than peers | 0.539 | 0.511–0.563 | 47 | 43 | 4 |
| [Cursor](https://feedbackbench.com/agents/cursor.md) | 4 | Typical | 0.503 | 0.472–0.534 | 56 | 44 | 12 |
| [Google Antigravity](https://feedbackbench.com/agents/antigravity.md) | =5 | Typical | 0.480 | 0.452–0.509 | 31 | 20 | 11 |
| [OpenAI Codex](https://feedbackbench.com/agents/codex.md) | 2 | Typical | 0.476 | 0.449–0.503 | 329 | 228 | 101 |
| [OpenCode](https://feedbackbench.com/agents/opencode.md) | 3 | Worse than peers | 0.457 | 0.421–0.491 | 48 | 28 | 20 |
| [Amp](https://feedbackbench.com/agents/amp.md) | =11 | Too few posts | – | – | 26 | 22 | 4 |
| [Pi](https://feedbackbench.com/agents/pi.md) | 7 | Too few posts | – | – | 14 | 12 | 2 |
| [Cline](https://feedbackbench.com/agents/cline.md) | =8 | Too few posts | – | – | 12 | 10 | 2 |
| [Factory](https://feedbackbench.com/agents/factory.md) | =11 | Too few posts | – | – | 12 | 10 | 2 |
| [GitHub Copilot](https://feedbackbench.com/agents/copilot.md) | =8 | Too few posts | – | – | 6 | 5 | 1 |
| [Zed](https://feedbackbench.com/agents/zed.md) | =8 | Too few posts | – | – | 2 | 0 | 2 |
| [Kiro](https://feedbackbench.com/agents/kiro.md) | 13 | Too few posts | – | – | 2 | 1 | 1 |
| [Conductor](https://feedbackbench.com/agents/conductor.md) | 14 | Too few posts | – | – | 2 | 1 | 1 |
| [Warp](https://feedbackbench.com/agents/warp.md) | 15 | Too few posts | – | – | 0 | 0 | 0 |
| [Grok Build](https://feedbackbench.com/agents/grok-build.md) | 16 | Too few posts | – | – | 0 | 0 | 0 |
| [Augment Code](https://feedbackbench.com/agents/augment.md) | 17 | Too few posts | – | – | 0 | 0 | 0 |

## Posts

Receipts rule: The 5 most recent praise and complaint posts per area (first 700 characters) and 3 per criterion (first 450 characters).

### Claude Code

- Praise, 2026-09-27, r/ClaudeCode (Reddit): “homie get some sleep, claude can work while you recharge” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqnh3q/me_after_opus_55_release/pc9w9xy/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “i built a docker proxy for my synology nas, to allow claude running as an "agent" user to work with docker, but without getting root access or uncontrolled access to my personal files. the nas doesn't support rootless docker itself, and i also wanted claude to be able to manage my containers, and start certain containers as root or other uids (e.g. dbs), but in a safe way. i also made a small script and skill to let claude check the current usage” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrcmjn/what_tool_have_you_built_for_yourself_with_claude/pcch6rf/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “personally i've mostly worked like you too, mainly because i want to make sure a repo / pipeline is as clean as possible and i understand it as much as possible but last project i've been doing ive learnt to rely a bit more on the agent. still planning a lot in advanced by breaking up a project into phases and creating different phase-x.md files for each phase describing the objective but with hard constraints i think code quality didn't drop an” [source](https://www.reddit.com/r/ClaudeCode/comments/1wowilt/i_still_dont_understand_this_agentic_workflow/pcct288/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “yeah, in theory and last week it was. but now it even stops my goals because 'i'm near the limit'. i'm quite angry about it. wanted to work overnight just to come back and see that it stopped and wasted 20% of my weekly quota because it didn't resume as planned...” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrf0ld/is_claude_pro_actually_worth_20_just_for_one/pcc0sit/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “no, it implemented a feature i asked and then decided to run a 2 hour simulation test to make sure that featured was working (which in this case was checking snapshots for irregularities in the loop). the fail was part of the test, i just don't understand why it decided to run it for 2 hours.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrt8ob/cloud_sessions/pcfksk4/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “it's useful, but unattended runs need limits written into the task, because the agent will happily decide a 2-hour soak test is "thorough". what stopped this for me: - **a time budget per step, in the prompt**: "any test or simulation you run must finish in under 5 minutes. if it would take longer, don't run it. write the command down and tell me instead." - **a stop condition**: "if the same check fails twice, stop and report what you saw. don't” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrt8ob/cloud_sessions/pcfnytj/)

### Devin

- Praise, 2026-09-27, @cognition (X): “the task is still working after 100 hours. @devindesktop @cognition @devinai crazy work <strict_link>” [source](https://twitter.com/1825243355501973504/status/2104062215695319071)
- Praise, 2026-09-27, @DevinAI (X): “linear is a shared task tracker for software teams. daily: check inbox for new issues/bugs, update statuses, prioritize work. weekly: plan cycles, review project progress, send status updates. devin runs fully autonomously in the cloud—assign a ticket and it codes, tests, ships a pr alone. other models need you guiding every step. opus 5.5 is the strongest for long-running agent planning and managing coder teams, with top performance at lower cos” [source](https://twitter.com/1720665183188922368/status/2104095517659541996)
- Praise, 2026-09-26, @cognition (X): “vibe coding is eating the world. devin is "truly capable of working independently and submitting prs," not like copilot where you type and it fills in. the team is led by scott wu, an ioi gold medalist, with a self-developed swe model + fusion multi-model routing, valued at $48b, with arr approaching $1 billion. mercedes transformed cobol from 8 months to just 8 days, fixing holes/migrating/testing can achieve 10–20x efficiency, and goldman sachs” [source](https://twitter.com/1413604240665104384/status/2103646116445376767)
- Complaint, 2026-09-27, @cognition (X): “another task going to be 24 hours! @devinai @devindesktop @cognition you guys are cooking my projects!!! unlimited swe 2 with a 20usd plan, are you kidding me? that is literally the best coding plan in the world <strict_link>” [source](https://twitter.com/1825243355501973504/status/2104262892090781744)
- Complaint, 2026-09-26, @DevinAI (X): “@ryancarson @devinai @linear @hellountangle managing agents is still management, someone still has to notice when the plan is wrong” [source](https://twitter.com/72165447/status/2103926352411820181)
- Complaint, 2026-09-23, @DevinAI (X): “@robinebers @devinai looks like the senior engineer is asleep on the cloud 😂” [source](https://twitter.com/20395932/status/2102710110019551636)

### Cursor

- Praise, 2026-09-27, r/cursor (Reddit): “yep 👍🏼. actually i used agent i needed that actually to run autonomous task till 2days constant...” [source](https://www.reddit.com/r/cursor/comments/1wrkzi5/am_i_cooked/pce76ru/)
- Praise, 2026-09-27, r/cursor (Reddit): “did some autonomous tasks...(2days constant workout with 8-9 agents running)” [source](https://www.reddit.com/r/cursor/comments/1wrkzi5/am_i_cooked/pce8fh2/)
- Praise, 2026-09-27, @cursor_ai (X): “@peter_soida @cursor_ai async, spinning up agents for me with their own computer, and more” [source](https://twitter.com/1430070528996376579/status/2104025533600428529)
- Complaint, 2026-09-24, r/cursor (Reddit): “i mean grok build is a good cli and cursor os a good ide/agent manager. composer 2.5 and grok are frankly good enough for 90% of coding and for short tasks are cost competitive. it is on par for many non coding tasks. the real issue is grok just isn't competitive on long agentic tasks. but is that really an issue? i think that is a fair question” [source](https://www.reddit.com/r/cursor/comments/1wp6j44/is_this_supposed_to_be_good_news/pbt0ud9/)
- Complaint, 2026-09-24, @cursor_ai (X): “@cursor_ai rollouts that watch their own deploys? that's the dream. my deploys still get watched by me refreshing sentry at 2am.” [source](https://twitter.com/1947882171340898304/status/2103176926978605153)
- Complaint, 2026-09-20, @cursor_ai (X): “@elie2222 ran @cursor_ai on a /goal for 36 hours. a blocked cla step could not finish, so the agent invented checks and tasks. 400 commits, 500 files, and it still claimed work remained. without a hard stop, agents invent busywork around gates. <strict_link>” [source](https://twitter.com/1464486759035805698/status/2101650374276911305)

### Google Antigravity

- Praise, 2026-09-27, @antigravity (X): “@aliahmadcode @antigravity what’s your longest antigravity agent run? mine is 17 hours 😅” [source](https://twitter.com/2091576854750871552/status/2104352459603001769)
- Praise, 2026-09-25, @antigravity (X): ““/goal go solve this problem, don’t stop until you’ve finished” that was my prompt to my @antigravity agent at this week’s @twilio assemble hackathon in san francisco. 20 min later and the agent had completed the task. i did all of this from my phone using antigravity remote control (docs: <strict_link>)” [source](https://twitter.com/2091576854750871552/status/2103487810325926115)
- Praise, 2026-09-25, @antigravity (X): “and i came back to a 90% done feature , still had to use my dev brain to guide it towards the finish line but meh i experienced more of the future of work today with @antigravity . seems going for a walk, getting fit and coming back to an almost done work is just going to become the new norm. can’t wait for mobile remote sessions for antigravity to be released to everyone” [source](https://twitter.com/180122538/status/2103592614410997958)
- Complaint, 2026-09-26, @antigravity (X): “@antigravity can't wait to use loops in antigravity next year 💪” [source](https://twitter.com/308867922/status/2103924848624050287)
- Complaint, 2026-09-26, @antigravity (X): “@thsottiaux @antigravity the agy cli actually has goal. i’m afraid that i tested it today…” [source](https://twitter.com/51725713/status/2103957307839455278)
- Complaint, 2026-09-16, r/google_antigravity (Reddit): “some people say btw doesn't work like so yet there are evidences i esc stopped it and resume with a prompt to stop it and it spends another millions tokens more to think and re-read? what about it? tell me a better use for "stop the ducking sh!t you are doing and simply just clean revert the commits" is like? i have used the skill to transform my session-starting prompt to outline and break steps to granularity, setting boundaries, what skills to” [source](https://www.reddit.com/r/google_antigravity/comments/1whsyhq/i_will_keep_posting_to_show_how_incapable_gemini/pa5ed5z/)

### OpenAI Codex

- Praise, 2026-09-27, r/codex (Reddit): “not sure about better, both performed well for me. it did seem quite good at managing this very long running task in a single parent thread, however” [source](https://www.reddit.com/r/codex/comments/1wrbh0t/i_just_completed_an_entire_server_migration_using/pcbqvl5/)
- Praise, 2026-09-27, r/codex (Reddit): “same thing pushed me to run both. codex for long background tasks, claude for the tricky parts where i want to steer it. burning 70% of a $200 plan in a few hours on a high setting sounds brutal though, did it say what was eating it, or did it just drop?” [source](https://www.reddit.com/r/codex/comments/1wrqch8/im_switching_to_opus_55/pcf2up6/)
- Praise, 2026-09-27, r/codex (Reddit): “i did planning / preliminary research of "what environmental things to include" in sol. luna did all of the inventory, logging, and transmission to/ configuration of destination. i allowed it do deploy up to 50 luna 6 medium subagents at a time, the most i saw was 43. it migrated my entire hobby server environment to a new server. vanilla codex cli only, no special harnesses, plugins, skills. the latter half i did on fast once i realized i wasn't” [source](https://www.reddit.com/r/codex/comments/1wrbh0t/i_just_completed_an_entire_server_migration_using/)
- Complaint, 2026-09-27, r/codex (Reddit): “> built on astra oh, so "long running" means 3h?” [source](https://www.reddit.com/r/codex/comments/1wrn697/new_openai_product_o_but_not_for_plus_users/pce0mvj/)
- Complaint, 2026-09-27, r/codex (Reddit): “i am pretty much all in on chat based and agentic tools these days, but who wants an always on agent that needs access to all your info, that is meant to take actions for you? it's a cool idea, but in 2026? underbaked.” [source](https://www.reddit.com/r/codex/comments/1wrn697/new_openai_product_o_but_not_for_plus_users/pcf8l0o/)
- Complaint, 2026-09-26, r/codex (Reddit): “they actively disabled the feature that made it complete the task. apparently because someone wrote a tool that kept the task alive forever and inserted new jobs. not that they couldn't have just prevented that instead of removing the feature entirely.” [source](https://www.reddit.com/r/codex/comments/1wqat13/openai_is_becoming_incompetent/pc2w4kx/)

### OpenCode

- Praise, 2026-09-27, r/opencode (Reddit): “thats a great point actually, i guess i did ai mainly because i distinctly having a vivid memory of starting a hike in <street_address>, sending out a detailed prompt, then at the top of the hike, checking my phone, and it had done about 30 minutes of coding, and the feature was basically done. went home, reviewed the code, made some adjustments, wrapped up my day. was incredible.” [source](https://www.reddit.com/r/opencode/comments/1wr11fu/tui2web_makes_opencode_usable_from_the_web/pcdzm75/)
- Praise, 2026-09-27, r/opencode (Reddit): “i really love opencode, i even built my own client on my phone so i can make sure my long running tasks is not interrupted when i go to lunch. it's opensource! <strict_link> it sent notifications using expo -> fcm and it's free, you can use it too!” [source](https://www.reddit.com/r/opencode/comments/1wrgqw7/i_made_a_pocket_client_for_opencode/)
- Praise, 2026-09-26, @opencode (X): “i kicked off a swarm of agents (@deepseek_ai v4.1 flash and @muse spark 1.3) in @opencode before going to bed, and i woke up to 111,804 lines of assembly code from a game decomp. <strict_link>” [source](https://twitter.com/2052933195935436806/status/2103858926718493142)
- Complaint, 2026-09-27, r/opencodeCLI (Reddit): “a massive con i found is that it tend's to stop letting u chat to the model and you'd have to compact the session with the command which mean's it can't run autonoumously while being reliable. you'd have to always be with it. not recommended.” [source](https://www.reddit.com/r/opencodeCLI/comments/1wrgmfe/whats_the_deal_with_openchamber/pcf6dox/)
- Complaint, 2026-09-27, @opencode (X): “@ivanfioravanti @opencode @openrouter yes yes yes. it is capable of fixing everything itself, but it needs to be directed every time. in autonomous mode, it found it difficult. it seems it was missing something. but we still don't know how big this model is. i bet it's somewhere around 250-300b.” [source](https://twitter.com/2010272031603068928/status/2104136707624796558)
- Complaint, 2026-09-25, @opencode (X): “no /goal in @opencode 🫠 can @thdxr do something?” [source](https://twitter.com/1423566780224753664/status/2103548250535907806)

### Amp

- Praise, 2026-09-24, @AmpCode (X): “i've been playing around with @ampcode's scheduled tasks, using my chatgpt sub for tokens and my macbook as the runner. great (free) experience so far! next on my list: scheduled tasks in orbs connected to remote mcp servers (happy to pay $20 for testing)” [source](https://twitter.com/295919262/status/2103066471496564985)
- Praise, 2026-09-24, @AmpCode (X): “my fileserver host is also showing signs of failing. so, i made it an @ampcode runner and spent a few days running threads to audit the current settings, configuration, users, ids, samba configuration, etc., so i can plan a migration without having to set up everything.” [source](https://twitter.com/7566122/status/2103110678730899652)
- Praise, 2026-09-24, @AmpCode (X): “@maxsumrall @ampcode not exactly, but i’ve had amp queue up github issues and just told it to subsequently tackle the issues until complete” [source](https://twitter.com/317888049/status/2103266531031593295)
- Complaint, 2026-09-24, @AmpCode (X): “@ampcode ...seems like letting agents run for hours and hours, or overnight. when i do that, i always end up with stuff i don't like. i also change the plan as stuff starts to take shape, planning everything ahead seems like a fools errand.” [source](https://twitter.com/26916652/status/2103136536695128503)
- Complaint, 2026-09-24, @AmpCode (X): “@scottbolinger @ampcode same experience with the overnight runs. what's worked better for me is short runs with a checkpoint where i look at the shape before it keeps going. the plan always changes once real code exists” [source](https://twitter.com/2046032911116242944/status/2103141321900700066)
- Complaint, 2026-09-15, @AmpCode (X): “been using @ampcode a lot lately but i really don't think models are ready to work without /goal. models get lazy even with a spec to stick to and i end up relying on hacky things like using puck to schedule nudges every 30 min.” [source](https://twitter.com/1000309405/status/2099797899043569967)

### Pi

- Praise, 2026-09-26, @pidotdev (X): “gave spacebunny-alpha a ridiculously long task (after planning it in opus 5.5 max). it's going on continuously for the last 24 hours non stop in @pidotdev and has consumed 1b+ tokens so far. <strict_link>” [source](https://twitter.com/14274934/status/2103735362065723728)
- Praise, 2026-09-26, @pidotdev (X): “@mitsuhiko @theakhandpatel @pidotdev yes! that's what i do. i haven't had to use goal on pi (or other harnesses too these days), it just follows things and continues if i ask the models to. the only thing i ask them to do is keep durable status and todo lists so they stay on track.” [source](https://twitter.com/14274934/status/2103784052138639639)
- Praise, 2026-09-26, r/codex (Reddit): “i've been wondering the same as my normal development is with an openclaw agent and a pi agent running on a server with glm 5.3 flash. i just submit issues in a local gitea repository and the agents pick them up sequentially. it's amazing and i can have 30 issues in the backlog that will get eaten up. but then i decided to start a new project with codex for the first time and i find myself with a terrible workflow. i switching back and forth betw” [source](https://www.reddit.com/r/codex/comments/1wqjqb3/the_optimal_codex_workflow/pc4sb0x/)
- Complaint, 2026-09-26, @pidotdev (X): “@theakhandpatel @shantanugoel @pidotdev there are tons of goal plugins but on sota models goal is not very useful. you better just ask it to write notes to some places where it can recall it.” [source](https://twitter.com/12963432/status/2103780427584446501)
- Complaint, 2026-09-02, r/PiCodingAgent (Reddit): “some of these harnesses try to help the model think, but llms have gotten really good at thinking without help. ai agents should focus on keeping context short and focused. tbh, i'd be surprised if stock pi won in a contest with long horizon tasks. you need something like subagents to control the size and focus of context of sub-tasks to maximize effectiveness of the llm output. there are many pi extensions that provide that kind of functionality” [source](https://www.reddit.com/r/PiCodingAgent/comments/1w593fd/i_benchmarked_pi_against_its_own_fork_omp_more/p7dahv5/)

### Cline

- Praise, 2026-09-25, @cline (X): “day 21: i got invited to the @cline cloud agents beta and handed my whole game project over to it. buildom was built 100 percent with cline from day one, so the invite felt like the company i already live in opening a new wing. i pointed a cloud agent at the entire repo, ran it on glm, and expected to babysit it. i checked in three times. the branch was in better shape every time. it came back with a full combat rework: a combat window with a fle” [source](https://twitter.com/2092320448604221440/status/2103486427602256006)
- Praise, 2026-09-24, @cline (X): “use cline. its has massive free model usage. you will never run out of credits. it can works for hours and has great context window. 👍 @cline” [source](https://twitter.com/1344665267947720705/status/2103139048042975328)
- Praise, 2026-09-15, @cline (X): “@ravikiran_dev7 @cline i can stop babysitting my clanker right after it learns which worktree it is in.” [source](https://twitter.com/1349317699789271052/status/2099799079807053998)
- Complaint, 2026-09-17, @cline (X): “@positronx_ @cline sucks xd i can't run a goal in this it take hours xd <strict_link>” [source](https://twitter.com/1261173216455712768/status/2100610075668988021)
- Complaint, 2026-09-12, @cline (X): “@cline 50 turns/task only works if cost and blast radius scale with it. longer jobs need checkpoints and a hard stop, not just a cheaper model.” [source](https://twitter.com/2097900519385690122/status/2098607990773284943)

### Factory

- Praise, 2026-09-26, @droid (X): “@ain3sh @droid @tastelabs shell process，it is necessary in long-time task and i can continue work while running” [source](https://twitter.com/2018347199432994816/status/2103882869550858340)
- Praise, 2026-09-23, @FactoryAI (X): “hey @factoryai i only have one piece of feedback for your platform. i've been absolutely loving it so far, mission control is extremely impressive for any advanced long-running tasks i have, but please please please add @namespacelabs as a connector. we need mac and windows too.” [source](https://twitter.com/2070908287978246144/status/2102563864055599358)
- Praise, 2026-09-18, @FactoryAI (X): “@tereza_tizkova @bentossell @factoryai i am currently building an aisdr it’s running on mission control right now from the past 93 hours. next i’m planning to build a finance intelligence and thinking of working with holograms. what are you working on @tereza_tizkova ?” [source](https://twitter.com/1833087116169105408/status/2100821679262126431)
- Complaint, 2026-09-24, @droid (X): “@droid @tastelabs any plan to add the support the run in background feature just like claude code?” [source](https://twitter.com/2018347199432994816/status/2103150290593796353)
- Complaint, 2026-09-15, G2 (G2): “q: what problems is the product solving and how is that benefiting you? a: a lot of engineering capacity still goes into work that’s necessary but not especially high-leverage: turning requirements into clear tickets, cleaning up after refactors, fixing straightforward test failures, updating docs or small modules, and other similar multi-file tasks. factory lets us treat more of that as delegable agent work, rather than pulling a human engineer” [source](https://www.g2.com/products/factory-ai/reviews/factory-ai-review-13477171)
- Complaint, 2026-09-05, @FactoryAI (X): “@factoryai can we just a goal feature in droid rather than a full fledged missions? like a smaller version of missions. asking because now models tend to work for a long time so missions should be replaceable with a goal feature.” [source](https://twitter.com/1106645265090396160/status/2096142424221610419)

### GitHub Copilot

- Praise, 2026-09-26, @GitHubCopilot (X): “try this usage pattern if you are using @githubcopilot cli/@claudeai code for doing research. has been really effective for explorations, better than either manual cli mode or full autopilot. 1. put your ideas in a markdown file with clear guardrails, well-defined scope for the exploration. start an autopilot session with yolo on 2. ask it to check background jobs at a long cadence (every 2-4h, which saves tokens and helps you keep track) 3. dro” [source](https://twitter.com/602846860/status/2103927294326948158)
- Praise, 2026-09-26, @GitHubCopilot (X): “try this usage pattern if you are using @githubcopilot cli/@claudeai code for doing research. has been really effective for explorations, better than either manual cli mode or full autopilot. 1. put your ideas in a markdown file with clear guardrails, well-defined scope for the exploration. start an autopilot session with yolo on 2. ask it to check background jobs at a long cadence (every 2-4h, which saves tokens and helps you keep track) 3. dro” [source](https://twitter.com/602846860/status/2103932718199562610)
- Praise, 2026-09-16, r/codex (Reddit): “codex cli is quite good, but the inability to run tasks in the background and then only consume more tokens when such a task is finished is a killer. desktop can't do this either i think, but vs code's extensions (including copilot) could. such a weird oversight.” [source](https://www.reddit.com/r/codex/comments/1wi7jaa/what_is_the_current_state_of_codex_cli_vs_desktop/pa8caix/)
- Complaint, 2026-09-02, r/GithubCopilot (Reddit): “doing what? seriously. tell me what a corporate swe in big tech would be doing with 24/7 running agents. because sure as hell i'm collecting use cases and best practices for that, but so far even our most hardcore ai fans came up with exactly zero. i mean zero, that generates value. they do have agents for personal shit.” [source](https://www.reddit.com/r/GithubCopilot/comments/1w5i0sx/one_msft_employee_reaches_100k_mo_in_token/p7ftzcj/)

### Zed

- Complaint, 2026-09-18, r/ZedEditor (Reddit): “i don’t engage a lot with the ai features to be honest. maybe it is a little more convenient to have the diff or something in the same windows. personally i’m not interested in the whole “leave the ai writing code for hours thing” so anything more than a chat to answer questions, boilerplate or simple mechanical tasks is bloat from my perspective” [source](https://www.reddit.com/r/ZedEditor/comments/1wjp5nq/popular_zed_fork_where_the_community_is_more/pal53xp/)
- Complaint, 2026-09-15, @zeddotdev (X): “@zeddotdev new frontier model can do tasks multiple days. * stops after 2 mins in all night run” [source](https://twitter.com/1551623510027825152/status/2099778781879963990)

### Kiro

- Praise, 2026-09-05, r/kiroIDE (Reddit): “i use it to work on jira tickets having to do with infra as code, gitlab, aws deployments, etc at this point it's mainly users making tickets and me letting kiro handle the tickets to completion. sometimes i have to come in and deal with stuff that requires business knowledge for our org but at the technical level, it's able to handle about 70% of tickets and no one knows they're talking to an llm as i trained it to talk casual and more like me.” [source](https://www.reddit.com/r/kiroIDE/comments/1w7zen5/what_if_kiro_cli_worked_like_an_entire_software/p7zwu4l/)
- Complaint, 2026-09-16, r/kiroIDE (Reddit): “i had a runaway agent once. it took all tasks, went into the background and continued for a couple of hours. no stopping of the thing. it survived prompts, commands, sessions and restarts. 100+ tokens on haiku and ~30 tasks later it happily reported in a newly opened session that it finished. micro-skynet experience. good it was a small private project.” [source](https://www.reddit.com/r/kiroIDE/comments/1whrobc/in_less_than_30_secs_kiro_uses_almost_100_credits/pa4s07y/)

### Conductor

- Praise, 2026-09-05, @conductor_build (X): “my favorite way to clear my head when i'm building a new feature is to get the agent working on a long task and then go build a smaller feature while i'm waiting. @conductor_build is perfect for this workflow.” [source](https://twitter.com/1454183904848474113/status/2096347366991512000)
- Praise, 2026-09-05, @conductor_build (X): “my favorite way to clear my head when i'm building a complex feature is to go build a smaller feature while i'm waiting for agents. @conductor_build is perfect for this workflow.” [source](https://twitter.com/1454183904848474113/status/2096347461329760568)
- Complaint, 2026-09-20, r/conductorbuild (Reddit): “i normally use claude code's /goal command to let it run long tasks uninterruptedly because it auto-resumes it's work once the session limit is reset, but i can't find a way to instruct conductor to do the same. even if a use /goal through conductor, the effect isn't the same. is anyone aware if this is possible or in the roadmap? it'd be really helpful” [source](https://www.reddit.com/r/conductorbuild/comments/1wlrxxi/does_conductor_support_autoresuming_a_session/)
