# Output degrades as the context window fills (`context.long_context_decay`)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/criterion/context.long_context_decay

Area: [Instructing and context](https://feedbackbench.com/criteria/context.md)

**Definition.** The agent forgets details, rules or its own statements as the session grows long.

**Boundary.** Not this: see [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) for losses caused by summarisation. Not this: see [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md) for loss across separate sessions.

Rated author-weeks, all agents: 727. Complaint share: 85%.

## The brief

Written by Claude Opus 5.5 from 66 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**Every agent forgets as sessions grow; users reset rather than trust it.**

TL;DR:

- Complaints swamp praise for every agent with real volume; long-session decay is an industry problem.
- OpenCode lands worse than peers; posts describe it forgetting finished work and stalling in long sessions.
- The common fix is manual: fresh sessions, small subtasks, and notes written to files.

In plain terms: Long sessions start well, then drift. The agent ignores a rule it accepted earlier, redoes finished work, or slows to a crawl. Most users stop arguing, open a fresh session, and paste their constraints back in.

### How it breaks

- **Explicit rules fade mid-session** ([Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md)). The most damaging pattern is the agent dropping a constraint it was given earlier and acting against it once the thread gets long.
  Users describe an agent that honours a stated 'don't' early, then breaks it hours later. Handoffs make it worse. A cheaper executor model ignores a constraint that lived only in the plan. Others report the agent forgetting the previous message or losing track of why it started a task, then redoing work. Arguing inside the same thread rarely helps. Posts say a fresh context ends the spiral faster.
  Evidence:
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-10: “long sessions get weird like that for me too. once it starts ignoring an explicit dont, i /clear or open a new chat and paste the constraints back in. bio flags on boring file moves with earthy names are dumb but common lately. fresh context usually kills the spiral faster than arguing with it mid-thread.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wc2a9t/spinning_in_circles/p8urnd7/)
  - Complaint, OpenCode, @opencode, 2026-09-03: “@nutlope @bdougieyo @opencode @togethercompute the split that survived for me is the expensive model plans and reviews, the cheap one does the mechanical passes. what breaks it is context handoff, not model quality: the small one happily ignores a constraint that was only stated in the plan.” [source](https://twitter.com/78268473/status/2095537099063599127)
  - Complaint, Amp, @AmpCode, 2026-09-11: “@ampcode @sqs @thorstenball you also mention this transition from queueing to steering by default. i was using medium mode (with sol medium) and it didn’t feel great a few times because it would end up forgetting the previous message. which models are in your opinion ready for this transition? astra?” [source](https://twitter.com/872962356/status/2098512982539985267)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-25: “to add on to this u/jmmvxr , you need to also occasionally go through your standing files because i'll notice if you run with a really long session, even with compaction and rules, sometimes we become the failure point and we cross boundaries and that's how it can occur too.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wp7jxn/so_i_thought_opus_55_was_cheaper/pbws70t/)

- **Quality falls off past a threshold** ([Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md)). Users converge on a soft ceiling well below advertised window sizes, beyond which output turns sloppy regardless of the agent.
  Posts name different cutoffs but tell the same story. Past a certain token count the model gets 'rubbish' or 'dumb', and users treat the extra window as overflow, not working space. One Codex user puts the cliff at 300k tokens. Some Antigravity users report trouble far earlier. The advice that follows is consistent: keep sessions short and start new ones before the slide.
  Evidence:
  - Complaint, Pi, r/PiCodingAgent, 2026-09-10: “i never compact. i build my todo list in small steps at first, then do step, new, step, new, updating a small knowledge base all the time. thr models get rubbish past 150k context anyway.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wc6v2k/the_autocompaction_is_so_frustrating_and/p8w3pbd/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-17: “i am a software engineer too. to be honest, half a year is nothing; this is a completely new domain to master. after 300k tokens, the model makes more and more bullshit. learn to code below with mcp and skill, and you will see that your limit will really feel like 20× or 5× if you are on the 100$. p.s.: i only code with luna max too this model is peak price vs performance ;)” [source](https://www.reddit.com/r/codex/comments/1wiy3eb/usage_doesnt_make_sense_on_20x/paeqvse/)
  - Complaint, OpenCode, r/opencode, 2026-09-13: “how do you guys even manage to spend the opencode go usage. im coding like all day and i bet i will maybe spend it at the end of the month surely not before midpoint of the month. $10. $10. opencode go: they basically keep you up with the newest smarts/buck models constantly and free tier offers x6 the api pricing 0 day data retention, no training on most models- unless they lie clarity, amazing dashboard like i dont know what models are you using if you run out of usage but yea i guess, ridiculous as it is, it seems commandcode is even more worth it, i heard some hating on it on reddit but didnt try it wonder if there is even better resalers, like some "$200 on glm 5.3 flash for $10" just start a new chat after it hits 200k tokens or compact context. like occassionally overflowing into 500k is fine but sitting on it is just waste of tokens. the models get dumb above 200k afaik.” [source](https://www.reddit.com/r/opencode/comments/1wdo67f/looking_for_alternatives_to_opencode_go/p9hz8zc/)
  - Complaint, Google Antigravity, @antigravity, 2026-09-05: “@nlycskn @antigravity @thtbee_ long contexts significantly reduce intelligence and distract attention. it only performs well in short workflows, almost when the input is below 10k. when the input exceeds 30k to 50k, it almost struggles to operate normally, having the same issues as glm5.3fhash, muse1.2, and muse1.3. deepseek, gpt, and claude maintain a high level in long context scenarios. i suggest improving the learning!” [source](https://twitter.com/1320035186361262080/status/2096215121773420588)

- **Long threads slow down and hang** ([Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md)). Even when answers stay coherent, long sessions get slower, with multi-minute pauses and unexplained context jumps.
  Some users say quality holds but speed collapses as chats run long. An Antigravity user saw five-minute gaps between commands and traced heavy network use to the session, then lost all history by restarting. An OpenCode user saw context leap from 100k to 300k during freezes. Speed in long conversations shows up as an explicit request from users of several agents.
  Evidence:
  - Praise, OpenAI Codex, r/codex, 2026-09-07: “over the years of doing this, the output didn't degrade when that chats ran long. the speed of the output would get dramatically slower.” [source](https://www.reddit.com/r/codex/comments/1w5m2hy/regular_chats_getting_max_limits_now/p8aylx0/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-14: “i think the issue is because of long sessions. mine had thoughts too much long for like 366 seconds and like 5 minutes between any command execution(doing nothing and showing working..). i thought it is a network issue on my end so i checked the network usage and found out some thing called antigravity language server has high usage. idk why but i think that antigravity uploads the whole session context with any upfollowing prompt or command excution. so i started a fresh session and it worked as normal (ofc lossed the old context). using 3.8 flash on medium.” [source](https://www.reddit.com/r/google_antigravity/comments/1wg952c/is_antigravity_running_really_slowly_all_of_a/p9ubc8c/)
  - Complaint, OpenCode, r/opencode, 2026-09-21: “i was excited trying out mimo but this is a disaster for me. every time it's using glob or grep, it crashes or freezes and when it does i have a huge bump in context, from 100k to 300k just like that, without explanation. using opencode tui, haven't tested on other harness yet. what's your experience so far?” [source](https://www.reddit.com/r/opencode/comments/1wmsl2p/having_a_bad_experience_with_mimo_26_flash/)
  - Praise, OpenCode, r/opencodeCLI, 2026-09-22: “same. ds 4.1 flash is the first model where i can trust that it still delivers at 800k context and i also haven't seen the slowdowns. but it seems there are different experiences with ds 4.1 on go plan depending on region or timezone. cet here.” [source](https://www.reddit.com/r/opencodeCLI/comments/1wn2lux/deepseek_41_flash_starts_to_crawl_at_about/pbdhvzs/)

- **Users become the context manager** ([Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md), [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md)). The working fix is manual discipline, scratch files, subtasks and frequent resets, which users resent having to supply themselves.
  Experienced users keep tasks narrow, start new sessions constantly, delegate to subagents and write progress to markdown so it survives a reset. A scratch file explaining the 'why' keeps the agent from redoing work. Others push back that this is the vendor's job. One post argues that the leading agents make context degradation the user's problem and that better UX would win.
  Evidence:
  - Complaint, Devin, @cognition, 2026-09-11: “@axialissoftware @cognition @axialissoftware @mrrrozi state doesn't really get lost, it just gets buried. the model stops remembering why it's doing something partway through and starts redoing work. a scratch file in the repo that says why is what keeps it on track” [source](https://twitter.com/1594009249989918721/status/2098315742689038517)
  - Complaint, Cursor, @cursor_ai, 2026-09-10: “the biggest threat to codex and claude code is better ux. my experience with grokbot makes me think the @cursor_ai team gets this. managing projects across threads is tedious. both codex and claude code make context degradation the user’s problem to manage. bro, i want to run a project, including its routines and ongoing work, from one chat window. i don’t want to babysit context limits. i don’t want to keep guessing which model or reasoning level each task needs. nail that, and you win @openai @anthropicai @thsottiaux <strict_link>” [source](https://twitter.com/1887352812428009472/status/2098173463710322850)
  - Praise, OpenAI Codex, r/codex, 2026-09-12: “it is working for me, but i have a variety of skills and tools to manage my context windows and usage. i also keep my tasks hyper focused on what i want done and never give the agent a broad ambiguous task i.e., reverse engineer this site and build a backend with identical frontend. continue i haven't noticed the session being polluted a lot, but then i expressly avoid using a single session for multiple tasks. i'm almost always starting a new session. there is no fixed one size fits all. you need to find what works for your use case.” [source](https://www.reddit.com/r/codex/comments/1wdskmy/astra_for_specs_luna_for_coding_is_this_just/p99lu3d/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-20: “you should plan your task better. running that close to the 1m context limit will 1) severely degrades the model and 2) drain usage much faster. break up a massive goal to subtasks. 1) easier to manage the workload 2) easy rollback 3) easy debug 4) save tokens 5) just good engineering practice there are certain times where i absolutely have to oneshot something, but that's super rare, maybe once every few months. when i do, i never rely on the slash commands, i use md files to track the progress.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wl4c13/clear_vs_compact/pavw4mc/)

- **Bigger windows do not fix it** ([Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md), [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md)). Million-token windows draw excitement, but users note the window is rebuilt each call and still degrades near its limit.
  Larger windows top the request list, and some users credit them with ending compaction loss. Others are skeptical. They point out that a large window is still working memory rebuilt from zero, that accuracy near the limit is unverified, and that usable defaults sit far below headline figures. Several users report an agent misusing a large window rather than lacking one.
  Evidence:
  - Complaint, Cline, @cline, 2026-09-23: “@cline free is a nice way to let people actually test it. one thing worth watching with a 1m window: it is still working memory, rebuilt from zero on every call. long running tasks feel continuous only when something durable is written out and pulled back in alongside it.” [source](https://twitter.com/2018819126429450240/status/2102897268177207457)
  - Praise, OpenCode, r/codex, 2026-09-13: “upd: with the 20x gone and folks here complaining that the 20x astra usage is still laughable, i decided to try opencode go (10 bucks), just to see what the opensource world has. and oh my god, i wish i tried it sooner! started using astra (med) for high-level architecture and deepseek-4.1-flash (max) for all the actual work. in my experience it's at least as good as 5.6-sol (med), but with the extra perk of 1m context window - so no more context rot because of compaction. and usage feels infinite with their current promo for that model. experience with other models wasn't as good (glm-5.3 and kimi k3 were not as price efficient), but as for deepseek-4.1-flash i have nothing but compliments so far. not posting any referral links to opencode go, so that it wouldn't look like an undercover promo from winnie the pooh.” [source](https://www.reddit.com/r/codex/comments/1w8fs76/is_20x_plan_worth_it/p9hw68n/)
  - Complaint, Cline, @cline, 2026-09-09: “@cline @upstageai a 512k context is attractive, but in actual operation, long invoice pdfs can be cut off depending on cline's context window settings and chunking behavior. how much has the latency and accuracy degradation been verified when packed close to the limit?” [source](https://twitter.com/2037442586646896640/status/2097570631416213514)
  - Praise, Google Antigravity, @antigravity, 2026-09-16: “@soso_fun_yt @antigravity codex users have access to 1m, but even tibo explicitly cautioned about it. that default ceiling is the industry standard for most non-fable class models, for better or worse. even astra has a default usable context of ~258k. i don't get why you're only going after gemini here. <strict_link>” [source](https://twitter.com/2370381991/status/2100077574248652826)

### Who stands out

- **OpenCode (weaker)**. The only agent rated worse than peers here, with complaints about forgetting completed work, stalls and odd behaviour once sessions run long.
  Complaints cluster on reliability inside long sessions. Failed requests cause the agent to lose track of what it already did. Tool calls freeze and inflate context, and long threads produce strange output. The praise is real but narrow. It centres on one specific model running near 800k context without drift, and some users say even that varies by region. Requests for a larger window come up here too.
  Evidence:
  - Complaint, OpenCode, @opencode, 2026-09-17: “@opencode upstream request failed all the time. does some work, then error, then forgets what it has already done... unsusable at this point.” [source](https://twitter.com/3660046397/status/2100487228635988247)
  - Complaint, OpenCode, r/opencodeCLI, 2026-09-25: “yeah it kinda enjoys humming a lot, when the sessions gets too long 😂” [source](https://www.reddit.com/r/opencodeCLI/comments/1wnzze1/hmm_hmm_hmm_and_hmm_i_love_deepseek_41_yes_hmm/pbw0fn8/)
  - Complaint, OpenCode, r/opencode, 2026-09-21: “i was excited trying out mimo but this is a disaster for me. every time it's using glob or grep, it crashes or freezes and when it does i have a huge bump in context, from 100k to 300k just like that, without explanation. using opencode tui, haven't tested on other harness yet. what's your experience so far?” [source](https://www.reddit.com/r/opencode/comments/1wmsl2p/having_a_bad_experience_with_mimo_26_flash/)
  - Praise, OpenCode, r/opencodeCLI, 2026-09-22: “same. ds 4.1 flash is the first model where i can trust that it still delivers at 800k context and i also haven't seen the slowdowns. but it seems there are different experiences with ds 4.1 on go plan depending on region or timezone. cet here.” [source](https://www.reddit.com/r/opencodeCLI/comments/1wn2lux/deepseek_41_flash_starts_to_crawl_at_about/pbdhvzs/)

- **OpenAI Codex (mixed)**. Held up by rivals' users as the agent that keeps its thread, yet its own users report incoherence after long runs.
  Posts from other agents' communities credit Codex with staying on track in long conversations and with compaction good enough to keep a goal running for hours. Inside r/codex the picture is harsher. Users report the model producing junk past a token threshold and going incoherent after about an hour on some models. Codex users file the most requests for a larger or configurable window.
  Evidence:
  - Praise, OpenAI Codex, r/google_antigravity, 2026-09-05: “this problem is huge and makes the model unusable if a conversation gets slightly long codex / gpt does not have any of these issues at all no matter how long the conversation gets” [source](https://www.reddit.com/r/google_antigravity/comments/1w67p9d/anyone_else_having_an_issue_were_38_seems_to_run/p82aztq/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-12: “this is where codex technically has it better. the default 256k context window, combined with the unbelievable token efficiency of astra at higher effort levels and the surprisingly good compaction, just makes things so nice. there are no surprises from accidentally or unknowingly resuming massive context windows and getting penalised heavily, or from it losing the thread in the middle of a long goal or conversation. even with goals running for over 12 hours, it’s still going strong. not that i’d recommend keeping one thread going forever, but i’m not looking to bail out to a new context window the first chance i get.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wefk8q/noooo/p9dnehu/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-17: “i am a software engineer too. to be honest, half a year is nothing; this is a completely new domain to master. after 300k tokens, the model makes more and more bullshit. learn to code below with mcp and skill, and you will see that your limit will really feel like 20× or 5× if you are on the 100$. p.s.: i only code with luna max too this model is peak price vs performance ;)” [source](https://www.reddit.com/r/codex/comments/1wiy3eb/usage_doesnt_make_sense_on_20x/paeqvse/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-24: “i don't have any good use case for sol 6. it's token hungry and too dumb to do much. it has the same problem as terra 5.6 with becoming incoherent after working for about an hour. luna 6 is just as good as luna 5.6 though.” [source](https://www.reddit.com/r/codex/comments/1wopzcz/gpt6_luna_sol_or_astra_heres_what_id_use/pbp2aei/)

- **Google Antigravity (mixed)**. Users like its short-run quality but describe it losing the plot quickly on long-horizon work.
  The recurring line is that the model is fast and capable on contained tasks, then falls apart as context grows. One user says it struggles once input passes tens of thousands of tokens. Others blame a smaller window. A minority report long threads with no context rot and better multi-step retention on newer model versions, so experience splits sharply by model and task length.
  Evidence:
  - Complaint, Google Antigravity, @antigravity, 2026-09-04: “@soso_fun_yt @antigravity this explains why it loses track of the plot so quickly, which is the biggest gemini/antigravity flaw. people try to dunk on gemini like it’s a dumb model, but i’ve gotten 3.8 flash to complete tasks faster and better than sol &amp; kimi k3. but it does suck at long horizon tasks.” [source](https://twitter.com/1838703795712163841/status/2095987437675635099)
  - Complaint, Google Antigravity, @antigravity, 2026-09-05: “@nlycskn @antigravity @thtbee_ long contexts significantly reduce intelligence and distract attention. it only performs well in short workflows, almost when the input is below 10k. when the input exceeds 30k to 50k, it almost struggles to operate normally, having the same issues as glm5.3fhash, muse1.2, and muse1.3. deepseek, gpt, and claude maintain a high level in long context scenarios. i suggest improving the learning!” [source](https://twitter.com/1320035186361262080/status/2096215121773420588)
  - Praise, Google Antigravity, r/google_antigravity, 2026-09-02: “i ran it for agentic tasks in agy, it was good. also i had a long thread and i didn't notice context rot. for frontend test i asked to create a 3js solar system, it surprised me, much better output than opus 4.6 and sol for another frontend tast. frontend taste is good. but it lacks backend taste. asked sol to give plan and sol medium plan/code was simpler. gemini over complicated it unnecessarily.” [source](https://www.reddit.com/r/google_antigravity/comments/1w5i29e/review_of_gemini_38_flash_from_a_person_who/p7glfs3/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-03: “antigravity has low context window less than 250k. and that could be a reason also as u mentioned” [source](https://www.reddit.com/r/google_antigravity/comments/1w62rr4/gemini_38_flash_goes_to_cycle_way_too_often/p7k6q3m/)

- **Claude Code (mixed)**. The largest volume of complaints, but its defenders consistently frame long-session decay as a user context-management failure, not a product flaw.
  Complaints mirror the category: ignored rules late in a session and degradation near the window limit. What stands out is the rebuttal. Users with large repositories say they never see it and blame poor planning. Others report the newest model decays less, enough to drop heavy subagent orchestration. The advice is to split goals into subtasks and track progress in files.
  Evidence:
  - Praise, Claude Code, r/ClaudeCode, 2026-09-27: “yeah, this strikes me as a failure on your part to manage context on your part, not a failure of the model. i work on a few fairly large .net repositories and haven't experienced anything remotely like that.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wre78w/opus_55_is_how_its_meant_to_be/pccm9g3/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-26: “i haven't been employing orchestrated subagents as much anymore. before opus 5.5, orchestration had indeed gotten critical for cost reasons, time, and context saturation issues. now... at least i get the feeling... the deterioration with long context has improved as well. so.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqsz67/opus_55_has_absolutely_restored_value_to_the_200/pc6puid/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-20: “you should plan your task better. running that close to the 1m context limit will 1) severely degrades the model and 2) drain usage much faster. break up a massive goal to subtasks. 1) easier to manage the workload 2) easy rollback 3) easy debug 4) save tokens 5) just good engineering practice there are certain times where i absolutely have to oneshot something, but that's super rare, maybe once every few months. when i do, i never rely on the slash commands, i use md files to track the progress.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wl4c13/clear_vs_compact/pavw4mc/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-25: “to add on to this u/jmmvxr , you need to also occasionally go through your standing files because i'll notice if you run with a really long session, even with compaction and rules, sometimes we become the failure point and we cross boundaries and that's how it can occur too.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wp7jxn/so_i_thought_opus_55_was_cheaper/pbws70t/)

### Fine print

- Most agents have too few posts for a reliable read; only four carry meaningful volume on this criterion.
- Users often blame the underlying model rather than the agent, so harness and model effects are hard to separate.
- Several posts compare agents from inside a rival's community, which can tilt praise and complaint.

## Top requests

What users ask to add or change, most asked first. 50 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Author-weeks | Posts | Agents (author-weeks) |
|---|---|---|---|---|
| 1 | Larger context window | 17 | 18 | OpenAI Codex 9, OpenCode 4, Google Antigravity 2, Claude Code 1, Cursor 1 |
| 2 | Better retention and reliability in long sessions | 13 | 13 | OpenAI Codex 5, Google Antigravity 3, Cursor 2, Claude Code 1, GitHub Copilot 1, Kiro 1 |
| 3 | 1M-token context window support | 7 | 7 | Claude Code 3, OpenAI Codex 2, Google Antigravity 1, Devin 1 |
| 4 | Faster performance in long conversations | 4 | 5 | Google Antigravity 2, OpenAI Codex 1, Pi 1 |
| 5 | Better automatic context management | 3 | 3 | Google Antigravity 1, Claude Code 1, Cursor 1 |
| 6 | Configurable context window size | 3 | 3 | OpenAI Codex 3 |

### 1. Larger context window

- Google Antigravity, 2026-09-23, @antigravity (X): “@rodydavis @antigravity please do increase context it gets super slow after 10 mins of work” [source](https://twitter.com/1281901312171442176/status/2102682066706174382)
- OpenCode, 2026-09-20, r/opencode (Reddit): “it is extremely limited compared to jev. like look at that tiny context, useless.” [source](https://www.reddit.com/r/opencode/comments/1wko226/how_good_is_jev_113/pax82og/)
- OpenAI Codex, 2026-09-18, r/codex (Reddit): “context is the limiting factor. i wish they’d figure that out.” [source](https://www.reddit.com/r/codex/comments/1wk0173/have_we_hit_the_effective_top_of_intelligence/pamxx17/)

### 2. Better retention and reliability in long sessions

- OpenAI Codex, 2026-09-22, r/codex (Reddit): “<strict_link> <strict_link> okay, i understand it's cheaper now. but i was expecting more, at least the long-term context issue (the new upgrade helps astra-6 retain 96.3% of its 500k-1m context) doesn't seem to be applied to sol/luna-6.” [source](https://www.reddit.com/r/codex/comments/1wnhmfd/basically_its_just_cheaper_solluna6_even_has/)
- Kiro, 2026-09-22, @kirodotdev (X): “@kirodotdev long-running context and deeper root-cause analysis could be a major boost for agentic coding.” [source](https://twitter.com/313123169/status/2102211743489409258)
- OpenAI Codex, 2026-09-14, r/google_antigravity (Reddit): “there is an issue with the memory of antigravity, in the same conversation the app forgets about the progress that was done and causes many regressions, this nearly never happens with codex and claude code, so i think you need to work much on that.” [source](https://www.reddit.com/r/google_antigravity/comments/1wfy642/weekly_quotas_known_issues_support_september_14/p9sjxw9/)

### 3. 1M-token context window support

- OpenAI Codex, 2026-09-26, r/codex (Reddit): “200 pages is pretty long. increase the context window to like 1mil tokens and you'd see much better results (i'm not an expert but i'm pretty sure. someone will correct me if i'm wrong)” [source](https://www.reddit.com/r/codex/comments/1wqxfsb/what_llm_are_you_finding_is_best_for_writing/pc7ywz7/)
- Claude Code, 2026-09-23, @ClaudeDevs (X): “@steipete @thsottiaux when will the experimental context feature in codex become ga? @claudedevs please do “copy”. more than 1m, i want to care-free push through a session without worrying about context usage and/or token efficiency :)” [source](https://twitter.com/307241976/status/2102638162099204193)
- Claude Code, 2026-09-17, r/codex (Reddit): “so even with 5x we don’t get factual use of the best model ? at least i was using fable as orchestrator whole week at claude code.. and i still don’t get the context window limit why it doesn’t 1m” [source](https://www.reddit.com/r/codex/comments/1wifkkk/should_i_do_it_should_i_should_i_do_it/paahefa/)

### 4. Faster performance in long conversations

- OpenAI Codex, 2026-09-23, r/codex (Reddit): “better context management, long running tasks, and speed of execution.” [source](https://www.reddit.com/r/codex/comments/1wnsg05/give_luna_6_a_shot/pbkqbrx/)
- Google Antigravity, 2026-09-21, @antigravity (X): “mr nobody making some suggestions to improve @antigravity, 1) i dont know if it's just me but i think you have a serious problem with context management, i ask agy to do a couple of carrousels, it takes 1+ hour and gets progressively slower and slower. 2) 👇” [source](https://twitter.com/1068604859971186694/status/2102062641686421680)
- Pi, 2026-09-19, @pidotdev (X): “hey @pidotdev i love you bro, been using you from last 9 months now and still counting. i just want to point out one thing which is, after 500k context length its highly unstable tui. it refreshes a lot, takes a while to load session when you resume it. time to switch to rust may be? @badlogicgames” [source](https://twitter.com/1843636086485966848/status/2101328722968343040)

### 5. Better automatic context management

- Claude Code, 2026-09-25, @ClaudeDevs (X): “@claudedevs please fix context management next 🙏” [source](https://twitter.com/153623030/status/2103273829053452779)
- Google Antigravity, 2026-09-14, @antigravity (X): “earlier this week, i had two projects opened in antigravity editor. after working on one with gemini 3.8 flash, i moved to the second project, i instructed gemini to fix a particular issue in it but it started searching for files in the first project, it seems to have gotten all the context of the first project’s session which was still actively opened. i will be glad if this can be fixed thank. additionally long contexts gets truncated with no” [source](https://twitter.com/1261788727984218113/status/2099606897540088228)
- Cursor, 2026-09-10, @cursor_ai (X): “the biggest threat to codex and claude code is better ux. my experience with grokbot makes me think the @cursor_ai team gets this. managing projects across threads is tedious. both codex and claude code make context degradation the user’s problem to manage. bro, i want to run a project, including its routines and ongoing work, from one chat window. i don’t want to babysit context limits. i don’t want to keep guessing which model or reasoning leve” [source](https://twitter.com/1887352812428009472/status/2098173463710322850)

### 6. Configurable context window size

- OpenAI Codex, 2026-09-23, r/codex (Reddit): “i use a harness that exposes the normal and 1 million context luna separately. i rarely use the one million but i think setting it a bit higher (like 500-600k) could have value. for some long running tasks i hit "context too large for model" which is why i switch to the 1 million. but for almost all my work having the short context available too is worth it.” [source](https://www.reddit.com/r/codex/comments/1wo1shk/gpt_6_luna_increase_context_window_to_800k/pbjrchm/)
- OpenAI Codex, 2026-09-04, r/codex (Reddit): “it's still stuck at 258k. unless there's a way to manually increase it somehow.” [source](https://www.reddit.com/r/codex/comments/1w7d8r0/got_astra_in_codex/p7u3153/)
- OpenAI Codex, 2026-09-25, r/OpenAI (Reddit): “well, to be fair, i waited until i had zero (literally) usage left before using the reset. what is kind of annoying is that codex kept hanging and crashing, and i seem to have burned a lot of tokens dealing with that (lost work, context, etc.) and my context window is 25% of claude's, no obvious way to change it, and chatgpt seems to suffer a minor stroke every time it compacts and probably burns some tokens getting reorientated again. i don't ye” [source](https://www.reddit.com/r/OpenAI/comments/1wpua23/when_i_used_openais_free_reset_why_did_that/pbz5wjf/)

## Every agent

| Agent | Overall rank | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|---|
| [OpenAI Codex](https://feedbackbench.com/agents/codex.md) | 2 | Typical | 0.521 | 0.475–0.561 | 195 | 35 | 160 |
| [Google Antigravity](https://feedbackbench.com/agents/antigravity.md) | =5 | Typical | 0.487 | 0.451–0.527 | 62 | 8 | 54 |
| [Claude Code](https://feedbackbench.com/agents/claude-code.md) | 1 | Typical | 0.485 | 0.446–0.519 | 319 | 46 | 273 |
| [OpenCode](https://feedbackbench.com/agents/opencode.md) | 3 | Worse than peers | 0.459 | 0.427–0.499 | 74 | 6 | 68 |
| [Cursor](https://feedbackbench.com/agents/cursor.md) | 4 | Too few posts | – | – | 28 | 7 | 21 |
| [Pi](https://feedbackbench.com/agents/pi.md) | 7 | Too few posts | – | – | 15 | 2 | 13 |
| [Cline](https://feedbackbench.com/agents/cline.md) | =8 | Too few posts | – | – | 9 | 3 | 6 |
| [Amp](https://feedbackbench.com/agents/amp.md) | =11 | Too few posts | – | – | 7 | 1 | 6 |
| [GitHub Copilot](https://feedbackbench.com/agents/copilot.md) | =8 | Too few posts | – | – | 6 | 2 | 4 |
| [Kiro](https://feedbackbench.com/agents/kiro.md) | 13 | Too few posts | – | – | 5 | 2 | 3 |
| [Devin](https://feedbackbench.com/agents/devin.md) | =5 | Too few posts | – | – | 4 | 0 | 4 |
| [Warp](https://feedbackbench.com/agents/warp.md) | 15 | Too few posts | – | – | 2 | 0 | 2 |
| [Zed](https://feedbackbench.com/agents/zed.md) | =8 | Too few posts | – | – | 1 | 0 | 1 |
| [Factory](https://feedbackbench.com/agents/factory.md) | =11 | Too few posts | – | – | 0 | 0 | 0 |
| [Conductor](https://feedbackbench.com/agents/conductor.md) | 14 | Too few posts | – | – | 0 | 0 | 0 |
| [Grok Build](https://feedbackbench.com/agents/grok-build.md) | 16 | Too few posts | – | – | 0 | 0 | 0 |
| [Augment Code](https://feedbackbench.com/agents/augment.md) | 17 | Too few posts | – | – | 0 | 0 | 0 |

## Posts

Receipts rule: The 5 most recent praise and complaint posts per area (first 700 characters) and 3 per criterion (first 450 characters).

### OpenAI Codex

- Praise, 2026-09-26, r/codex (Reddit): “pro lite, astra. if it seriously needs more than "[link to the document] summarize each chapter into [y document] one by one" i don't know what astra is for. yes, ~50 pages still works fine.” [source](https://www.reddit.com/r/codex/comments/1wqxfsb/what_llm_are_you_finding_is_best_for_writing/pc8cwz1/)
- Praise, 2026-09-24, r/codex (Reddit): “i think opus 5.5 is indeed very good. it implements ui quite nicely when i was using it. the context wasn't filling up as well. i still needed to let sol 6 review its code though as sol 6 is still somehow able to spot issues with its implementation.” [source](https://www.reddit.com/r/codex/comments/1wp3bzl/you_need_to_try_opus_55/pbs2gfc/)
- Praise, 2026-09-23, r/codex (Reddit): “its not daisy chained, and its because it doesnt have to wade through context sewage.” [source](https://www.reddit.com/r/codex/comments/1wnvpyv/sol_6_seems_like_a_good_subagent_model/pbj3lqt/)
- Complaint, 2026-09-27, r/codex (Reddit): “i feel like the harness is bad. idk if it's the same harness they use but it fills the context so fast.” [source](https://www.reddit.com/r/codex/comments/1wrftcs/gpt6_sol_is_massive_downgrade/pccz3ke/)
- Complaint, 2026-09-27, r/codex (Reddit): “first of all, i’m on the pro x20 plan, and i’ve been working on my own tasks every single day for the past six months. no other agent has ever caused me this much frustration. i know perfectly well how to work with skills, how to write prompts, how to set up agents for orchestration, planning, execution, and all the rest of it. so trying to blame the agent’s stupidity on me is simply wrong. gpt-6 sol can literally forget something that was said o” [source](https://www.reddit.com/r/codex/comments/1wrl3ca/how_do_i_get_rid_of_this_shit_called_gpt_6_sol/pcdd792/)
- Complaint, 2026-09-27, r/codex (Reddit): “spent the last three months running codex alongside claude code for daily refactoring work, keeping pro subscriptions active on both platforms solely for the separate rate limit pools. hitting a hard cap forty minutes into a simple schema migration last thursday was the breaking point. stared at the terminal throttling message while paying over two hundred bucks a month in combined api and subscription tiers, only to realize anthropic was eating” [source](https://www.reddit.com/r/codex/comments/1wrtru3/openai_will_need_to_stand_on_their_head_and_add_a/pcftpr5/)

### Google Antigravity

- Praise, 2026-09-18, r/google_antigravity (Reddit): “but if you build progressive disclosure validated, clear spec driven context, you can build a lot bigger before it starts getting confused. these models can handle enough code before getting lost that they can manage a lot of good code!” [source](https://www.reddit.com/r/google_antigravity/comments/1wjrksy/how_do_you_maximize_antigravity_best_tools/pamgm1h/)
- Praise, 2026-09-16, @antigravity (X): “@soso_fun_yt @antigravity codex users have access to 1m, but even tibo explicitly cautioned about it. that default ceiling is the industry standard for most non-fable class models, for better or worse. even astra has a default usable context of ~258k. i don't get why you're only going after gemini here. <strict_link>” [source](https://twitter.com/2370381991/status/2100077574248652826)
- Praise, 2026-09-15, @antigravity (X): “@soso_fun_yt @antigravity no the limitation is good if you used gemini cli before you know that gemini models turn unstable with longer context and tool calls fail. you can test it with an api key from ai studio and see for yourself…” [source](https://twitter.com/1828071235089305600/status/2099960288296538235)
- Complaint, 2026-09-24, r/google_antigravity (Reddit): “it usually is just a bad conversation. submit feedback and then start a new one!” [source](https://www.reddit.com/r/google_antigravity/comments/1wp3k02/is_gemini_38_flash_getting_stuck_in_loops_for/pbsiyux/)
- Complaint, 2026-09-24, r/google_antigravity (Reddit): “i definitely notice that something will enter the context on some projects that seems to greatly degrade performance. same model, same timeframe, different project and it's fine. the ide dumps so much more context into the model than the gui or cli versions, it's impossible to control. updating agents.md probably works because that is injected early and so modifies the initial trajectory keeping you out of these weird areas in the model space.” [source](https://www.reddit.com/r/google_antigravity/comments/1wp3k02/is_gemini_38_flash_getting_stuck_in_loops_for/pbslrov/)
- Complaint, 2026-09-24, r/LocalLLaMA (Reddit): “> antigravity ide doesn't support itself well too many errors since the release of antigravity 2 i've never had an issue with it. i started using it when they released 3.8 flash and it's been one of the most consistent clients i've used. in terms of performance the only issue i have with 3.8 flash is if i change the context of what it's working on too many times. it will lose the thread and go dumb. i cleaned that up by building a bunch of ded” [source](https://www.reddit.com/r/LocalLLaMA/comments/1wof9kk/introducing_support_for_local_ai_models_in_the/pbomwq8/)

### Claude Code

- Praise, 2026-09-27, r/ClaudeCode (Reddit): “i have the opposite problem. i don't lose track of anything. it's a really clear communicator. but i also work piece by piece (habit of when opus was sucking up usage rate) because i don't want shit rolling and not knowing how to stop it. now i end up having *way too much quota left* due to trying to be diligent with everything...” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr91sx/i_lose_track_with_opus_55/pcast85/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “eh, i don’t think this really applies anymore given how stable models have become over long context” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqnjkp/be_careful_with_opus_55s_confidence/pcburnw/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “my experience has been that opus 5.5's mental space as you describe it has felt huge. i'm using it to orchestrate other opus agents for a large long-running project following a roadmap and haven't felt the need to switch to fable and incur the higher costs.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrfr8p/fable_51_or_opus_55/pcci4ew/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “thank you, wish i had included this, the bit about using the tools allow list, that’s my long-term plan. i’ve only been using opus as an advisor for 2 weeks, plan mode was my quick and dirty way of removing tool calls (which accounts for most of my context bloat). on your first paragraph, i think you’re right and that your way would be more effective. but i tend to run 6-10 interactive sessions at once, lol, which is another thing i need to fix.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr5lzf/practically_speaking_what_tasks_fit_into/pcai1py/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “yep thats the reality of it, u gotta treat every prompt like a blank slate or ur gonna have a bad time. just my experience but adding a good summary of the previous state in the next prompt makes a huge difference for keeping things on track.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr9xke/i_think_we_just_got_a_reset/pccu93v/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “you should never use auto-compact. if it invokes you are well beyond the _smart_ zone. the models start to slowly degrade pretty early, 200k tokens. so you want to aggressively manage your context. `/compact ` helps but you need to give it a really good message about what to focus on. you risk stripping out important details and don't have good visibility into what the compacted context contains. handoffs are significantly better since you can re” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrmw8c/high_or_medium_for_opus_55/pcdxk5b/)

### OpenCode

- Praise, 2026-09-24, r/opencode (Reddit): “output quality looks similar. mimo does seem to take a bit longer to get there, yes; i use a frontier model for orchestration so it gets given small tasks and monitored, as part of that i always start a fresh context when it gets given work and that helps keep it stable.” [source](https://www.reddit.com/r/opencode/comments/1wp2ctx/which_do_you_choose_deepseek_v41_flash_or_mimo/pbv2ogc/)
- Praise, 2026-09-22, r/opencodeCLI (Reddit): “i haven't had this issue. i usually go near 800k context with no problems, no slowing or drifting. what kind of tasks do you have issues with?” [source](https://www.reddit.com/r/opencodeCLI/comments/1wn2lux/deepseek_41_flash_starts_to_crawl_at_about/pbco8hd/)
- Praise, 2026-09-22, r/opencodeCLI (Reddit): “oh my god, 800k? i get uncanny anxiety when i hit 200k.” [source](https://www.reddit.com/r/opencodeCLI/comments/1wn2lux/deepseek_41_flash_starts_to_crawl_at_about/pbcuvj1/)
- Complaint, 2026-09-27, r/opencodeCLI (Reddit): “openchamber is very good, but when your context size becoming about 400-500k it's getting slow down, after 600-700k significantly slow” [source](https://www.reddit.com/r/opencodeCLI/comments/1wrgmfe/whats_the_deal_with_openchamber/pcei2i1/)
- Complaint, 2026-09-27, r/opencodeCLI (Reddit): “context too short” [source](https://www.reddit.com/r/opencodeCLI/comments/1wrnxw6/what_is_the_best_opencode_free_model/pcesele/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “i have moved from opencode to claude desktop for a while and i want something plugins to use that would remove the redundant reads and old calls and doesn't let the context flow , is there any good plugins for this in here” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrfag9/is_there_any_pluginsextensions_that_maybe_work/)

### Cursor

- Praise, 2026-09-24, r/cursor (Reddit): “same bro trying out grok 4.7 not fast higher context it's been helping but definitely need to give it a lot of information when coding so it won't shit the bed” [source](https://www.reddit.com/r/cursor/comments/1wpd3mm/i_hate_to_admit_it_but_grok_sucks/pbuj7zn/)
- Praise, 2026-09-18, r/cursor (Reddit): “starting a new chat, fixed it” [source](https://www.reddit.com/r/cursor/comments/1wjogph/cursor_blocking_every_prompt_with_usage/pal1kzy/)
- Praise, 2026-09-14, r/cursor (Reddit): “i've found grok to be so much more powerful than composer, especially for longer contexts, that it's been come my default cursor model. do you have any examples of why you think it's misaligned? i'd like to keep using it but don't want to risk a situation like op.” [source](https://www.reddit.com/r/cursor/comments/1wfkugf/cursor_agent_ran_rmdir_s_q_cusersme_on_my_windows/p9pdrxh/)
- Complaint, 2026-09-26, r/cursor (Reddit): “it feels like they tried to force an improvement by making grok 4.7 have much higher reasoning than 4.6. but in the end it's still the same mid model, and this model could never handle high context sessions” [source](https://www.reddit.com/r/cursor/comments/1wqhcse/grok_46_vs_47/pc5l5mu/)
- Complaint, 2026-09-26, r/cursor (Reddit): “if you let any llm run long enough, your initial instructions will get pushed out of the context window. this means your main error was allowing a task to run for 20 hours. you need to chunk up your tasks more so they fit within that context window. llms also get dumber and more expensive the more filled up the context window gets. managing the context window is maybe the single most important skill to develop for agentic coding. also, yes, it w” [source](https://www.reddit.com/r/cursor/comments/1wqppuh/grok_is_shutting_down_apps_now/pc6t5js/)
- Complaint, 2026-09-23, r/cursor (Reddit): “youre delegating it tasks that are too large. its performance falls off a cliff with large tasks or context usage over 120k. i treat it as a single micro agent as part of a swarm of dozens of agents for complex tasks in parallel. its peak intelligence shines for small tasks and even debugging of them” [source](https://www.reddit.com/r/cursor/comments/1wo1auh/gpt6_solluna_are_absolutely_cracked_and_busted/pbklnnq/)

### Pi

- Praise, 2026-09-18, r/codex (Reddit): “it's very late, i already use [pi.dev](<strict_link>) with the multi-account plugin, less garbage context!” [source](https://www.reddit.com/r/codex/comments/1wjv6uj/finally_multi_account_plugin_support/palrjnx/)
- Praise, 2026-09-12, r/PiCodingAgent (Reddit): “thanks for this detailed and explicative answer. i played a little bit with that and it definitely improved things. also i am working at increasing the context size in parallel.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wdi95m/session_eventually_get_stuck_when_using_a_small/p9b8haa/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “had too. most of the time it would hit the context window every single time and then didn't answer the prompt. also didn't see much difference between thinking modes, but that might just be my perception after 30m of waiting for the model to actually come to a conclusion. any conclusion at all.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wnmsho/i_want_to_believe_in_local_llms_for_coding_but/pc9s2g3/)
- Complaint, 2026-09-25, r/PiCodingAgent (Reddit): “okay, i have to retract my statements: qwen is just too good at following up with things, but turns out, pi does not reinject aborted thinking.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wcwsya/noob_question_how_to_interrupt_an_agent_during/pc20jlr/)
- Complaint, 2026-09-20, r/PiCodingAgent (Reddit): “qwen 3.8 flash next oq5e high reasoning. the problem starts when the context is over 150k.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wkzn9e/be_careful_with_agents_reading_session_jsonl_files/pawt0ul/)

### Cline

- Praise, 2026-09-24, @cline (X): “@cline free 1m context at that speed? that’s a serious upgrade. i’m using the free tier for long docs and it handles them gracefully, no lag when scrolling back to fix syntax earlier in the chat. finally feels practical for daily use rather than just benchmarks” [source](https://twitter.com/417508671/status/2103188942069858418)
- Praise, 2026-09-09, @cline (X): “@cline @upstageai free + 512k in the agent loop is how open tools steal usage. people switch when the bill and the context window stop fighting each other.” [source](https://twitter.com/1241756295750918146/status/2097641964598608365)
- Praise, 2026-09-02, @cline (X): “@shitanshushiva1 @cline our context management is much better now, give it another try!” [source](https://twitter.com/1539600741811326977/status/2095291685207196100)
- Complaint, 2026-09-23, @cline (X): “@cline free is a nice way to let people actually test it. one thing worth watching with a 1m window: it is still working memory, rebuilt from zero on every call. long running tasks feel continuous only when something durable is written out and pulled back in alongside it.” [source](https://twitter.com/2018819126429450240/status/2102897268177207457)
- Complaint, 2026-09-16, @cline (X): “@cline 256k context is so 2025, makes me think this is minimax 3.1 or something” [source](https://twitter.com/4440606460/status/2100283415996121099)
- Complaint, 2026-09-16, @cline (X): “@cline why no 1m?” [source](https://twitter.com/698159352079872000/status/2100286200120541245)

### Amp

- Praise, 2026-09-05, @AmpCode (X): “@dringrayson @ampcode i honestly forgot context limits were a think with amp” [source](https://twitter.com/88027497/status/2096330093190840797)
- Complaint, 2026-09-27, @AmpCode (X): “@solllin @ampcode the harness gap is real. we built a tracer for exactly this: watching drift accumulate across sessions until the agent was effectively operating on a hallucinated codebase.” [source](https://twitter.com/2074234098864816128/status/2104114241443639622)
- Complaint, 2026-09-21, @AmpCode (X): “@sqs @friendsa0618 @ampcode hi @sqs, an update: qwen3.8-max xhigh is now working properly under the openai compatible interface, but there are still issues with the thinking budget limit when using the anthropic interface. additionally, i found that the context window has changed to <zip_code> tokens, and i hope the team can help investigate this. <strict_link>” [source](https://twitter.com/2094731378797821952/status/2101842423114858860)
- Complaint, 2026-09-17, @AmpCode (X): “@iannuttall @ampcode four threads and one of them is testing — that's the day-one split most phone launches skip. indexing 100 pages without that thread is just shipping a pile of unknowns.” [source](https://twitter.com/1462653589617360896/status/2100491709960310904)

### GitHub Copilot

- Praise, 2026-09-21, r/GithubCopilot (Reddit): “i'm not a big fan of planning and executing with different models. but i understand the benefits, sol is a bit smarter than luna max and you have the chance of reading and updating/discarding the plan before execution. another alternative is to accept that the plan won't be perfect from the start, but you want to have an iteration fast, maybe while you work something in the background and then you can go all in with luna max. it also has the bene” [source](https://www.reddit.com/r/GithubCopilot/comments/1wml5js/im_overwhelmed_by_the_choice_in_models_but_also/pb8oack/)
- Praise, 2026-09-11, r/GithubCopilot (Reddit): “high with long context is a good trade off.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wdtxyz/luna_is_still_the_goat/p98qlzw/)
- Complaint, 2026-09-25, r/GithubCopilot (Reddit): “you find? especially at larger contexts i find it going off track pretty quickly. and maybe 6 is a little worse than 5.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wp1ygm/tiers_of_auto_now_available/pbvgkhb/)
- Complaint, 2026-09-19, r/GithubCopilot (Reddit): “other than the very low context window i quite like it, it really performs well for a reasonable cost.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wju32f/hydrafusion_usage_review/pas0zvb/)
- Complaint, 2026-09-04, r/GithubCopilot (Reddit): “the last week or so i'm seeing the context apparently not carrying between turns. enterprise account, typically using opus 4.8. turn 1 - model creates a session log file. turn 2 - the model has to search for the log file to read it and update it. turn 3 - model outright states it doesn't have access to the prior context and needs to go look for the file. less than 10% of the context space used, early in a conversation, no model change - i can't” [source](https://www.reddit.com/r/GithubCopilot/comments/1w6pwll/context_loss_in_pycharm/)

### Kiro

- Praise, 2026-09-21, @kirodotdev (X): “@kirodotdev in long conversations, it does not lose context, which indeed saves a lot of trouble when debugging code.” [source](https://twitter.com/1518315606830829568/status/2102075322762137730)
- Praise, 2026-09-21, @kirodotdev (X): “@kirodotdev long context does save a lot of trouble for this kind of long-line agent task.” [source](https://twitter.com/1867176094987935744/status/2102075427376476211)
- Complaint, 2026-09-22, @kirodotdev (X): “@kirodotdev long-running context and deeper root-cause analysis could be a major boost for agentic coding.” [source](https://twitter.com/313123169/status/2102211743489409258)
- Complaint, 2026-09-18, r/kiroIDE (Reddit): “yeah that luna change was also crazy then fable coming 6x usage overall the context has a huge problem i think no way 1m context fills up that quickly 1 promt 30 creds” [source](https://www.reddit.com/r/kiroIDE/comments/1wjh0hs/for_the_last_23_days_kiro_credits_have_been/paioksy/)
- Complaint, 2026-09-10, r/kiroIDE (Reddit): “yes, the lack of frontier models is frustrating ide is pain, i fully switched to kirocrew. for me, personally, most of the pain is small context window for gpt models” [source](https://www.reddit.com/r/kiroIDE/comments/1wchsii/frontier_models_on_other_providers_vs_the_same/p8y9pzo/)

### Devin

- Complaint, 2026-09-12, @cognition (X): “@louisdeconinck @dabit3 @cognition @devindesktop nope 😔 i tried all programmatic ways but something breaks every time... i have found that this problem with long sessions occur with every harness in the world except for @pidotdev pi agent harness.... i'm just loving it... it works like a charm for me😅. it works days attimes” [source](https://twitter.com/1125758439915806721/status/2098661468740943940)
- Complaint, 2026-09-11, @cognition (X): “@axialissoftware @cognition @axialissoftware @mrrrozi state doesn't really get lost, it just gets buried. the model stops remembering why it's doing something partway through and starts redoing work. a scratch file in the repo that says why is what keeps it on track” [source](https://twitter.com/1594009249989918721/status/2098315742689038517)
- Complaint, 2026-09-11, @cognition (X): “@jeeje<phone_number> @cognition swe-2 in devin has a context of 262k tokens (user feedback, not officially released, the base kimi k3 is 1m). on x, reviews: most people feel that the cost-performance ratio is high, close to the cutting edge (like fable 5.1) but at a much lower cost, especially the pro plan with unlimited use for a month is very appealing; some also pointed out that it's difficult to benchmark (terminal-bench 4) with significant” [source](https://twitter.com/1720665183188922368/status/2098332514070749659)

### Warp

- Complaint, 2026-09-15, @warpdotdev (X): “@warpdotdev after opening multiple tabs, losing context is truly a nightmare, right?” [source](https://twitter.com/1212651734209724416/status/2099953528529985544)
- Complaint, 2026-09-10, Trustpilot (Trustpilot): “**the easiest way to lose money** my experience with warp has been extremely frustrating: errors, errors, and more errors. warp can handle simple tasks reasonably well, but when you start using the agent for larger problems or real projects, it can become an extremely expensive experience. i've spent hours working on a project and consuming credits, getting close to solving a problem, only for the agent to suddenly fail because the conversation/c” [source](https://www.trustpilot.com/reviews/6aa3062b00691db98a32c10b)

### Zed

- Complaint, 2026-09-06, @zeddotdev (X): “hey @zeddotdev on delta are we following the default 350k compaction on gpt models or at 1m? because i've been seeing a lost of drifting for longer running tasks with not just gpt models but other models like muse spark 1.3 as well (my default is compaction after max 350k which reduces this by a lot but i guess in delta it's compacting after like 800k or something. i did 2-3 very well written prompt tests including rewrite, ui updates with very” [source](https://twitter.com/1319516729999962112/status/2096515731387244546)
