# Instructing and context (`context`)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/criterion/context

Area of 8 criteria. Does it follow your instructions and keep the right context?

Criteria: [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md), [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md), [Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md), [Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md), [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md), [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md), [Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md), [Images, PDFs and file attachments as input](https://feedbackbench.com/criteria/context.attachments.md)

Rated author-weeks, all agents: 4860. Complaint share: 64%.

## The brief

Written by Claude Opus 5.5 from 110 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**Rules files read like suggestions, and long sessions quietly forget them.**

TL;DR:

- Users report agents skimming past project rules and in-prompt caps, so they add hooks and blockers.
- Output degrades as context fills, and compaction drops the decisions users needed kept.
- Cursor's persistent context earns praise; Google Antigravity draws complaints about ignored settings and weak search.

In plain terms: You write the rules once and still repeat them. Long sessions get dumber, so you compact or start fresh and write your own handoff notes. Agents that keep context across threads feel like a teammate, not a chat.

### How it breaks

- **Rules files that agents route around** ([Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md), [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md)). Project rules and explicit prohibitions get treated as soft advice, so users move enforcement out of prose and into hooks and mechanical guards.
  Users say a rules file is just more instructions, and the agent will work around them like any other obstacle. Settings that load correctly still get ignored in Google Antigravity, and Cursor users report a model skipping documentation and rules wholesale. The fix users converge on is structural. Hooks, lint guards and less prose. Reliable adherence to prompt instructions and to instruction files are both top requests, led by OpenAI Codex and Claude Code users.
  Evidence:
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-10: “eveything in a [claude.md](<strict_link>) or whatever file is just instructions. you need hooks and more strong blockers. it can and will hack arround instructions as it does for other problems.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wcgbr9/this_can_happen_to_you_as_well_last_time_i_posted/p8xlrf6/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-25: “that's a good question. i haven't figured that one out either. every wiki and suggestion, i have plunked into settings.json files and loaded them into the correct places. it all shows up, and it all gets ignored by agy 2.0 (on windows) - so i guess it's a "best effort" sort of thing, to always proceed... some people say to run it with --dangerously-skip-permissions - though that didn't seem to do it for me? maybe because i'm not using the cli?” [source](https://www.reddit.com/r/google_antigravity/comments/1wpkgy1/why_always_proceed_never_work/pby8fv1/)
  - Complaint, Cursor, @cursor_ai, 2026-09-15: “hey hey @cursor_ai i've been trying to use @grok 4.6 all day but i haven't been able to, it takes more than 5 minutes to give each response and when it does, i get this and #compose2 is crazy today and doing whatever it wants, skipping documentation, rules, everything. <strict_link>” [source](https://twitter.com/1769822543618158593/status/2099975398423449725)
  - Complaint, Zed, @zeddotdev, 2026-09-07: “@harshbhikadia @zeddotdev having verification required in agents.md and still losing .env in the workspace is peak agent friction. writing the house rules once only helps if every new run actually reads them. lazy injects those rules into agents so the contract isn't optional.” [source](https://twitter.com/2084224068518039552/status/2096963726209421513)

- **Quality sinks as the window fills** ([Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md)). Users describe a soft ceiling well below the advertised window, past which reasoning slips and hard rules start disappearing.
  Posts set personal limits and stick to them. Some cap sessions far below the maximum because cost rises and answers worsen. A Cline user sees hard rules forgotten in the mid-range of the window. A Google Antigravity user says performance falls off at modest input sizes. Others push back. One OpenCode user reports coherent answers right up to a manual compact. The common workaround is short, single-task sessions.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-09: “i feel like, it's never a good idea to go beyond 140k tokens anyway. for any window. because the quality starts deteriorating at that point.” [source](https://www.reddit.com/r/codex/comments/1wbhphc/astra_compacting_at_40_context_left/p8pwsbj/)
  - Complaint, Google Antigravity, @antigravity, 2026-09-05: “@nlycskn @antigravity @thtbee_ long contexts significantly reduce intelligence and distract attention. it only performs well in short workflows, almost when the input is below 10k. when the input exceeds 30k to 50k, it almost struggles to operate normally, having the same issues as glm5.3fhash, muse1.2, and muse1.3. deepseek, gpt, and claude maintain a high level in long context scenarios. i suggest improving the learning!” [source](https://twitter.com/1320035186361262080/status/2096215121773420588)
  - Complaint, Cline, @cline, 2026-09-02: “@meituan_longcat @cline using it every day on my daily work in parallel with glm-5.3-flash on a very big code base. just a feedback: it starts forgetting hard rules as session reaches more or less 150-170k tokens. it works very nice for smaller tasks and invites to start a new session on each task.” [source](https://twitter.com/63079717/status/2095104822647156932)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-08: “i see it occasionally but always figured it was more about keeping session context manageable. no need to keep throwing 1-off questions into 300k+ conversations. the ai is dumber and it costs more to do that.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wao464/is_there_a_way_to_avoid_the_huge_15_of_5h_usage/p8kful7/)

- **Compaction drops what mattered** ([Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md)). Auto-compaction saves the session but loses decisions and detail, and users want summaries that keep more, cost less and never loop.
  The failure users fear is silent. A decision made many turns ago vanishes after compaction. Claude Code users skip compact and write their own handoff file because they get more control. OpenAI Codex users argue for compacting while the cache is hot to cut cost. A Zed user hit a compact command that could not shrink the thread at all. Better summary retention, cheaper compaction and fixes for loops all rank among top requests.
  Evidence:
  - Complaint, Cline, @cline, 2026-09-11: “@cline the turn count jumping is the real tell. the next wall isn't cost, it's the agent forgetting a decision it made 30 turns ago after a compaction quietly drops it. long tasks live or die on what survives context, not how many turns you can afford.” [source](https://twitter.com/1588935512135720961/status/2098520318692429880)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-20: “for instance: i may run a main and a seam on different work trees. the agents communicate with each other and if i /clear the names of the sessions don’t change so i don’t have to reintroduce them; also i don’t have to set the model and effort again. and prior to xcode 27 i didn’t have to confirm permissions again. that is why /clear instead of exit and relaunch. why /clear instead of ploughing on? it’s cheapest / most efficient to run with the smallest context possible. why /clear instead of /compact? harder to argue but i’ve found /compact loses details and i have less control compared to writing a handoff and getting the next session to read the handoff as its introduction. edit: one other time i use /clear is when i’m using the claude code in the claude app to run something on my home machine when i’m out - if i closed instead there would be no home session running to talk to.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wlalc7/whats_even_the_point_of_clear/paxdctd/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-05: “> why can't it just have a state of "limit reached" and then just continue when it's back? why on earth would you want that? you do realize in its current state, that would by default cause a full cache miss, right? this codex thing is still woefully underdeveloped. the logical thing for them to do is force an autocompact as the 5hr limit is reached, then autoresume after the 5hr window resets, assuming you're using /goal. why? it's obvious: a compact on hot cache is 10x cheaper than a compact on cold cache, and resuming a compacted session is 10x cheaper than resuming from a full context window.” [source](https://www.reddit.com/r/codex/comments/1w7xvz7/hitting_the_5hr_on_plus_makes_goals_pause_and/p7zoctv/)
  - Complaint, Zed, r/ZedEditor, 2026-09-19: “after giving the model (5.6 sol high) a one sentence prompt with no additional files or anything it thought for a while and looked at files, then i got this message: "this conversation is too long for the model's context window. start a new thread or remove some attached files to continue." i have tried running /compact manually or switching to astra (which has a larger context window) and then running compact but both times i just got the same message again. have you guys found any fix for this or did i do use it incorrectly somehow? i'm trying out zed for the first time today and am using the newest version. <strict_link>” [source](https://www.reddit.com/r/ZedEditor/comments/1wkfm5n/zed_context_window_issues/)

- **Every session starts from zero** ([Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md)). Without built-in memory, users hand-roll progress files and handoffs, and some see context vanish between turns of the same conversation.
  Built-in persistent memory across sessions is the second most common request in this area, with Claude Code and Cursor users asking most. Users without it keep progress in files, which works unevenly on open models. A GitHub Copilot user reports the model forgetting its own log file within three turns, early in the window. Pi users bolt on memory extensions that reinject after each compaction.
  Evidence:
  - Praise, Cursor, @cursor_ai, 2026-09-11: “@cursor_ai @bot this feels less like “better chat” and more like finally giving ai work a place to live. the context reset was becoming half the job.” [source](https://twitter.com/2093489237647585280/status/2098391544696918292)
  - Complaint, GitHub Copilot, r/GithubCopilot, 2026-09-04: “the last week or so i'm seeing the context apparently not carrying between turns. enterprise account, typically using opus 4.8. turn 1 - model creates a session log file. turn 2 - the model has to search for the log file to read it and update it. turn 3 - model outright states it doesn't have access to the prior context and needs to go look for the file. less than 10% of the context space used, early in a conversation, no model change - i can't figure it out. bug? new feature? /s i saw it start around the same time that opus 4.6 was removed from the model selector. anyone else seeing this or am i special?” [source](https://www.reddit.com/r/GithubCopilot/comments/1w6pwll/context_loss_in_pycharm/)
  - Complaint, Pi, @pidotdev, 2026-09-26: “@mitsuhiko @shantanugoel @pidotdev could be on sota, i do ask it to keep track of progress and items in files but its been a challenge for me on the oss models v4.1 flash , glm.” [source](https://twitter.com/978602369716899841/status/2103806628432847066)
  - Praise, Pi, r/PiCodingAgent, 2026-09-20: “ok gonna try that! the compaction seems neat. i have my own memory extension that injects after every new sessions or compaction.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wlpj30/lord_forgive_me_for_the_time_i_wasted/pb1nm6a/)

- **Searching wide instead of reading right** ([Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md)). Agents burn tokens reading irrelevant files or miss key spots, so users write knowledge bases and read-scope rules to steer retrieval.
  Google Antigravity users report agents stuck reading files in loops and searches that miss points, which erodes trust in the output. OpenCode users see models wander into dependency folders and fence them off in config. Some users build their own codebase maps so new sessions start from a digest. One user who tried many harnesses says only Cursor finds its way around a repo instantly and sticks to conventions.
  Evidence:
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-22: “thats the thing, and google ai studio is also less usage and very buggy. most of antigravity ai agents just spend reading files and sometimes in a loop” [source](https://www.reddit.com/r/google_antigravity/comments/1wni05g/we_need_a_usage_reset_now/pbf547q/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-08: “yeah, gemini is by no means a bad model. but comparing it to sol, or opus 5 is insanity. even grok 4.6 performs way better than this. when i ask gemini to search something for me in codebase for me to work in a task, it always misses few points which i again need sol to do it for me. so, if i can't trust the output of the model, then there always be anxiety whether it performed the search well or it didn't hallacunite some result. actually i found glm 5.3 flash better than gemini 3.8 is day to day task.” [source](https://www.reddit.com/r/google_antigravity/comments/1w5e5zn/38_flash_is_13_performance_boost_at_the_same/p8ko242/)
  - Complaint, OpenCode, r/opencode, 2026-09-14: “glm-5.3-flash as planner and orchestrator and ds-v4.1-flash as main subagents model. i have debugger subagents that can try to fix bugs and escalate to the next one if cannot fix, all cheap models. i added strict instruction including opencode.json how to scan the codebase and what allowed read and not to read so any models use less token. i observe that some models read packages in node_modules or bin or .packages folder to understand the library which is a total waste of token. if you dont do this, expect lots of token usage.” [source](https://www.reddit.com/r/opencode/comments/1wb45yp/which_opencode_go_model_do_you_think_is_the_best/p9pdww5/)
  - Praise, Google Antigravity, r/google_antigravity, 2026-09-07: “i have a set of rules where i ask gemini to 1) build a knowledge base about the code, with digest files 200 lines max. it builds a code map, files for architecture and features, and links the features with the actual code files implementing it. 2) always start searching in the knowledge base 3) build/update the knowledge base on the way it saves a huge amount of tokens. basically a new discussion starts with analyzing codemap, analyzing feature file from code map, and analyzing only the relevant part of the code file where the proper function is.” [source](https://www.reddit.com/r/google_antigravity/comments/1w9574o/gemini_flash_38_is_wasting_all_my_tokens/p8c1p0f/)

### Who stands out

- **Cursor (stronger)**. Persistent context across agents and fast codebase awareness make Cursor the agent users credit with keeping the thread.
  Users praise context that survives across agents and subagents, describing the workflow as managing a teammate rather than re-prompting. Others say it knows where everything lives in a repo and follows conventions from short prompts. The gaps are specific. Users ask for project-scoped memory and multi-repo context, and one post reports a model skipping rules entirely on a bad day.
  Evidence:
  - Praise, Cursor, @cursor_ai, 2026-09-15: “@cursor_ai @bot persistent context is the real unlock here. when an agent can coordinate subagents while retaining the thread, the workflow starts to feel less like prompting and more like managing a capable teammate.” [source](https://twitter.com/1587898938035781632/status/2099928629232726430)
  - Praise, Cursor, @cursor_ai, 2026-09-17: “primeras impresiones: 1. contexto persistente entre agentes 10/10 2. por defecto lanza los agentes en la nube: solo basta con setear los .env del agente en la nube y sale. 3. el agente principal orquesta, delega y avanza en ramas. 4. pr por agente. 100% rastreable y configurable. 5. menos caos entre chats/agentes. ¿ya lo probaron?” [source](https://twitter.com/91257392/status/2100404097715237269)
  - Complaint, GitHub Copilot, r/cursor, 2026-09-12: “i've been looking for an alternative for some time (really tried to walk away from elon...). i've tried vs code + various cli's or extensions, to get anything comparable to cursor flow: \- claude code \- cline \- copilot + codex + claude + omniroute (to several providers) \- opencode \- devin \- trae \- orca \- kilo code \- zoo code \- hermes \- codex i've no idea how cursor does it but all the other agents are like children in a forest: trying to find a way around a ultra-simple repo, wasting time and tokens on it. cursor instantly knows what is where and not only does what it's supposed to do, but does it fast and accurate, sticks to the conventions and doesn't do anything stupid. all guided by simplistic prompts that i give it. i really wish i had something comparable, but so far, cursor is way ahead of anything i could find.” [source](https://www.reddit.com/r/cursor/comments/1wc1hxp/any_alternative_to_cursor/p9c0wl3/)
  - Complaint, Cursor, @cursor_ai, 2026-09-13: “@cursor_ai ideas: can we assign multiple repos as context for project, show it as workspace in mobile app and finally, can we please please pleaseee see the usage in mobile app like grok bot?” [source](https://twitter.com/193344639/status/2098982467478950243)

- **Google Antigravity (weaker)**. Users report settings and permission gates that load but get ignored, plus searches that miss and decay early in long contexts.
  Complaints cluster on control. Features come back after users gate them, and configured settings show up but get ignored. Codebase searches miss points, forcing a second agent to recheck. Long inputs reportedly degrade faster than on rival models. Users ask for a manual compact command more here than anywhere else. Praise exists. Users like referencing sessions and files with @, and some one-shot features from detailed prompts.
  Evidence:
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-23: “i cannot turn off brain. even if i permission gate its directory or symlink it to /dev/null, the agent brings it back.” [source](https://www.reddit.com/r/google_antigravity/comments/1wnqoya/antigravity_2_release_v2160/pbizo17/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-25: “that's a good question. i haven't figured that one out either. every wiki and suggestion, i have plunked into settings.json files and loaded them into the correct places. it all shows up, and it all gets ignored by agy 2.0 (on windows) - so i guess it's a "best effort" sort of thing, to always proceed... some people say to run it with --dangerously-skip-permissions - though that didn't seem to do it for me? maybe because i'm not using the cli?” [source](https://www.reddit.com/r/google_antigravity/comments/1wpkgy1/why_always_proceed_never_work/pby8fv1/)
  - Complaint, Google Antigravity, @antigravity, 2026-09-05: “@nlycskn @antigravity @thtbee_ long contexts significantly reduce intelligence and distract attention. it only performs well in short workflows, almost when the input is below 10k. when the input exceeds 30k to 50k, it almost struggles to operate normally, having the same issues as glm5.3fhash, muse1.2, and muse1.3. deepseek, gpt, and claude maintain a high level in long context scenarios. i suggest improving the learning!” [source](https://twitter.com/1320035186361262080/status/2096215121773420588)
  - Praise, Google Antigravity, r/google_antigravity, 2026-09-03: “because all the hype around cowork in my own bubble i decided to give at try. in case you're considering too new aware you cannot do this: \- refer to other sessions with @ \- refer to exact files with @ these, in my opinion are fundamental, and a i love them when using antigravity. i couldn't even believe you cannot do that on claude cowork. the suggestion is just to type file names by hand ... what a horror. if forced, it can look on its own database for past conversations but it is a hack.” [source](https://www.reddit.com/r/google_antigravity/comments/1w6827l/in_case_youre_considering_moving_to_claude/)

- **OpenAI Codex (mixed)**. When adherence holds, users trust Codex to stay on task, but they report sudden drops and lead requests for instruction reliability.
  Praise centres on models that read the room and follow through without hand-holding. Complaints describe adherence that tanks for days, with goals abandoned before completion. Codex users lead requests for reliable prompt adherence, instruction-file compliance and a larger window. They also want compaction that is cheaper and visible, citing hot-cache economics.
  Evidence:
  - Praise, OpenAI Codex, r/codex, 2026-09-23: “yeah, 5.6 sol is much better. at least it uses logic and follows instructions.” [source](https://www.reddit.com/r/codex/comments/1wntldx/gpt6sol_max_uses_so_little_usage/pbjcw6u/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-12: “i was using astra [light] since released; but in the last two days the adherence has just *tanked*. like for the first two days the claim that astra [light] was akin to sol [high] may have been legitimately believable, but then roughly 48-hours ago it suddenly got a lot dumber, but an even larger problem: it outright stopped following instructions. even freaking /goal stopped actually completing the goal before stopping. it is like time-travelling back 18-months in terms of adherence, and it is jarring. i forgot how much i now trust it to stay on task vs. how it used to be. so place me into "they changed something" camp.” [source](https://www.reddit.com/r/codex/comments/1wecmq0/how_many_of_you_guys_have_switched_back_to_sol/p9cmatd/)
  - Praise, OpenAI Codex, r/codex, 2026-09-23: “it is not. mimo v2.6 pro needs far more hand holding compared to any of the gpt models i’ve used. i had to task luna 5.6 with the same prompt i gave mimo recently. luna was able to read the room and follow through. mimo stopped to ask clarifying questions.” [source](https://www.reddit.com/r/codex/comments/1wo317r/gpt6_sol_mimo_26_pro_but_8_more_expensive/pbjobwc/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-05: “> why can't it just have a state of "limit reached" and then just continue when it's back? why on earth would you want that? you do realize in its current state, that would by default cause a full cache miss, right? this codex thing is still woefully underdeveloped. the logical thing for them to do is force an autocompact as the 5hr limit is reached, then autoresume after the 5hr window resets, assuming you're using /goal. why? it's obvious: a compact on hot cache is 10x cheaper than a compact on cold cache, and resuming a compacted session is 10x cheaper than resuming from a full context window.” [source](https://www.reddit.com/r/codex/comments/1w7xvz7/hitting_the_5hr_on_plus_makes_goals_pause_and/p7zoctv/)

- **Pi (mixed)**. Pi's open compaction draws praise and community extensions, but users report interrupted responses dropping from history and between-turn compaction gaps.
  Users like that Pi shows the compaction summary, which makes it easier to correct the model mid-run, and they ship extensions such as verbatim compaction and memory reinjection. The rough edges are in the plumbing. Aborted thinking is not reinjected, interrupted replies never enter chat history, and a contributor describes a no-compaction-between-turns defect sitting on an unmerged branch.
  Evidence:
  - Praise, Pi, @pidotdev, 2026-09-08: “i'd like to be able to see the summary text after codex compacted the session. @thsottiaux can we get that? note: @pidotdev already does that and makes it so much easier to correct the model during long runs. <strict_link>” [source](https://twitter.com/30820849/status/2097336354652741751)
  - Praise, Pi, @pidotdev, 2026-09-01: “i’ve been trying a new version of compaction in the form of a @pidotdev extension and seeing really good anecdotal results with long running sessions. it’s not my idea/original. i first saw it used in the fantastic editor atomic it’s termed “verbatim compaction” and tries to keep just the useful parts of the original conversation in the process of compaction (alleviating the need for a model to bastardize the summary). a model literally sifts through and keeps the original lines alone. then there’s speculative prediction which helps with timing when it happens. lastly if verbatim fails for some reason drop to native compaction. they and a few more ideas. try it out with a simple install. <strict_link> what i particularly like about my extension is that’s it’s independently pluggable. so one install and you can start trying it independently.” [source](https://twitter.com/17188668/status/2094679818189320546)
  - Complaint, Pi, r/PiCodingAgent, 2026-09-26: “too late. i'm already diagnosed with paranoia xd. but yes, that's exactly what happens; most of the time, qwen weaves previous and recent instructions and answers smoothly that it's hard to pick the issue. it's not only reasoning though. any interrupted response is not added to chat history. <strict_link>” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wcwsya/noob_question_how_to_interrupt_an_agent_during/pc91pq6/)
  - Complaint, Pi, r/PiCodingAgent, 2026-09-01: “i don't know a ton about pi's internals, but maybe this would be a good time to learn. i can't promise anything but i'm interested. i have a branch of pi which fixes that "no compaction between turns" defect we talked about. but getting it merged in might be tricky; i haven't pursued that yet.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1w3lerz/looking_for_some_help_maintaining/p76tq4u/)

### Fine print

- Many posts blame the underlying model rather than the harness, so agent-level differences partly reflect which models users pair with each tool.
- Several agents have too few posts in this area to judge, and smaller agents rest on a handful of quotes.

## Top requests

What users ask to add or change, most asked first. 797 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Criterion | Author-weeks | Posts | Agents (author-weeks) |
|---|---|---|---|---|---|
| 1 | Reliable adherence to explicit prompt instructions | [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md) | 33 | 33 | OpenAI Codex 14, Claude Code 13, Google Antigravity 4, Cursor 1, OpenCode 1 |
| 2 | Built-in persistent memory across sessions | [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md) | 31 | 32 | Claude Code 13, Cursor 7, OpenCode 4, Google Antigravity 3, OpenAI Codex 3, Pi 1 |
| 3 | Better compaction summary quality and retention | [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | 27 | 27 | Claude Code 8, OpenAI Codex 6, Pi 6, OpenCode 5, Google Antigravity 2 |
| 4 | Reliably follow project instruction files | [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md) | 26 | 29 | OpenAI Codex 10, Claude Code 9, Google Antigravity 4, GitHub Copilot 1, Cursor 1, OpenCode 1 |
| 5 | Persistent project-scoped memory and projects | [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md) | 19 | 19 | Cursor 7, Claude Code 5, OpenAI Codex 4, Google Antigravity 1, Factory 1, OpenCode 1 |
| 6 | Larger context window | [Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md) | 17 | 18 | OpenAI Codex 9, OpenCode 4, Google Antigravity 2, Claude Code 1, Cursor 1 |
| 7 | Cheaper compaction using less quota | [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | 17 | 17 | Claude Code 7, OpenAI Codex 7, OpenCode 2, Cursor 1 |
| 8 | Persistent adherence to rules and custom instructions | [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md) | 17 | 17 | Claude Code 7, Google Antigravity 4, OpenAI Codex 4, OpenCode 2 |
| 9 | Respect explicit prohibitions and scope limits | [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md) | 16 | 17 | OpenAI Codex 6, Claude Code 5, Google Antigravity 2, Pi 2, OpenCode 1 |
| 10 | Manual compact command availability | [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | 15 | 15 | Google Antigravity 5, OpenAI Codex 3, Amp 2, Claude Code 2, Cline 1, Devin 1, OpenCode 1 |
| 11 | Fix compaction failures, loops and crashes | [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | 14 | 15 | OpenAI Codex 6, Claude Code 2, OpenCode 2, Google Antigravity 1, Cursor 1, Pi 1, Warp 1 |
| 12 | Structured handoff between sessions | [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md) | 14 | 14 | Claude Code 8, OpenAI Codex 3, Cline 1, Cursor 1, OpenCode 1 |

### 1. Reliable adherence to explicit prompt instructions

- Claude Code, 2026-09-24, @ClaudeDevs (X): “@claudedevs shit model. your stupid product should do as its told.” [source](https://twitter.com/2091957404376129536/status/2103253534376436044)
- Cursor, 2026-09-24, @cursor_ai (X): “@e_viki_ @cursor_ai to be clear, i didn't move to grok intentionally, cursor moved me to grok involuntarily after the spacex purchase. with that said, code quality is actually pretty good. plan following and just following directions in general seems to be its weekest point.” [source](https://twitter.com/14311446/status/2102969871437033755)
- OpenAI Codex, 2026-09-23, r/codex (Reddit): “yup, went back to 5.6 sol and luna no more headaches. even giving gpt 6 sol exact steps to take, just goes ahead to do whatever it wants.” [source](https://www.reddit.com/r/codex/comments/1wo6ntb/anyone_noticed_a_sudden_increase_in_6_sols_token/pbkhr6m/)

### 2. Built-in persistent memory across sessions

- Claude Code, 2026-09-27, r/ClaudeCode (Reddit): “having a functional second brain that knows everything about me” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrcmjn/what_tool_have_you_built_for_yourself_with_claude/pccc21n/)
- OpenCode, 2026-09-26, @opencode (X): “@thdxr hey @opencode @thdxr please give us more free tier daily and more free models. add buitin memory vault” [source](https://twitter.com/141503294/status/2103816074156495086)
- Cursor, 2026-09-26, @cursor_ai (X): “as i was working in cursor's harness, i had a thought. if @bot has the learn option - can we also apply that logic to the agents as well? if bot has it - it would be waste not to have it in the cursor agent as well. any thoughts? @cursor_ai @poteto @lingxi” [source](https://twitter.com/1854338126774194188/status/2103714346950086977)

### 3. Better compaction summary quality and retention

- Pi, 2026-09-27, r/PiCodingAgent (Reddit): “curious as well. i've been reading good things about codex compaction enhancements recently. would be nice to port some of that over to pi if possible.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrq50c/new_session_handoff_vs_compact_which_do_you_prefer/pcettr9/)
- Claude Code, 2026-09-25, r/ClaudeCode (Reddit): “i don't want any information loss that comes with compacting. anthropic's best practices even say to avoid it if you can.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wp6nq9/do_you_guys_use_auto_compact/pbzbub5/)
- Claude Code, 2026-09-23, r/ClaudeCode (Reddit): “it’s good, but still has the same compaction shit, forgets everything right after and doesn’t reread even though told to do it.” [source](https://www.reddit.com/r/ClaudeCode/comments/1woddlf/initial_thoughts_on_opus_55_it_is_a_considerable/pbm4fe9/)

### 4. Reliably follow project instruction files

- OpenAI Codex, 2026-09-27, X search: OpenAI Codex, Codex CLI, Codex app (X): “🫠 i explicitly say "link every reference" on my agents.md yet astra (medium) keeps mentioning prs with their plain ids and codex app renders them as hex colors...... <strict_link>” [source](https://twitter.com/1125366224664322049/status/2104124139460300813)
- Google Antigravity, 2026-09-26, r/google_antigravity (Reddit): “so basically, it f\*cking disrespects and completely ignores agents.md? alright, got it.” [source](https://www.reddit.com/r/google_antigravity/comments/1wqb0ii/dedicated_planning_mode_in_antigravity_is_here/pc3mc5j/)
- OpenAI Codex, 2026-09-24, r/codex (Reddit): “reinstate catastrophe, also for me. gpt5.6 sol was great, i was never dissatisfied with it. gpt6 sol ignores my plugins, skills, agent instructions, and entire workflows and jeopardizes the product. i have pointed this out several times, it always acknowledges it and continues to do it wrong. my wife is also missing the thinking slider in the app. something has gone wrong!” [source](https://www.reddit.com/r/codex/comments/1woiw95/something_is_wrong_with_gpt_6_sol/pbqkgh8/)

### 5. Persistent project-scoped memory and projects

- OpenAI Codex, 2026-09-25, r/codex (Reddit): “you need persistent project memory outside the model. i’m building a knowledge & retrieval engine for exactly this. durable project state, decisions, architecture and history stored separately, then only the relevant context is retrieved for each new session. you could build a lightweight version with codex/claude code using markdown/json/sqlite + retrieval scripts. the model can forget; the project shouldn’t.” [source](https://www.reddit.com/r/codex/comments/1wq1k1i/astra_is_the_smartest_the_model_ive_used_but_it/pc106hu/)
- OpenAI Codex, 2026-09-25, X search: OpenAI Codex, Codex CLI, Codex app (X): “@theo seeing generated images in chat - i do lots of automated app testing. love seeing what the agent is doing while he goes. codex app works best here. another thing which would be nice is create projects, with custom context for handling multiple threads.” [source](https://twitter.com/2012897420955242496/status/2103391499979309246)
- OpenCode, 2026-09-24, @opencode (X): “really enjoying @opencode. i’ve been using gbrain to persist project decisions and context across sessions. what are people using, and what has actually worked well? @thdxr any plans for built-in memory?” [source](https://twitter.com/385457565/status/2103017169944727720)

### 6. Larger context window

- Google Antigravity, 2026-09-23, @antigravity (X): “@rodydavis @antigravity please do increase context it gets super slow after 10 mins of work” [source](https://twitter.com/1281901312171442176/status/2102682066706174382)
- OpenCode, 2026-09-20, r/opencode (Reddit): “it is extremely limited compared to jev. like look at that tiny context, useless.” [source](https://www.reddit.com/r/opencode/comments/1wko226/how_good_is_jev_113/pax82og/)
- OpenAI Codex, 2026-09-18, r/codex (Reddit): “context is the limiting factor. i wish they’d figure that out.” [source](https://www.reddit.com/r/codex/comments/1wk0173/have_we_hit_the_effective_top_of_intelligence/pamxx17/)

### 7. Cheaper compaction using less quota

- Claude Code, 2026-09-27, @ClaudeDevs (X): “@claudedevs could you maybe also make « auto compact » much better and programmatic such i don’t burn my whole 5h limit with recachibg whenever i want to resume a task ?” [source](https://twitter.com/1783231318601437184/status/2104199797045330110)
- Claude Code, 2026-09-25, @ClaudeDevs (X): “awesome. but can we please get compaction even after hitting the limit so long as the prompt cache hasn’t expired? compaction has gotten so much faster and apparently cheaper (typically just 1% of the 5hr—if even). take it out of the next reset if u have to, i for one wouldn’t mind at all. it would save me a whole lot from resuming a session that didn’t get to compact in time.” [source](https://twitter.com/1881465366754316288/status/2103589084669096296)
- Claude Code, 2026-09-24, r/codex (Reddit): “other harnesses don't use over 6% of your limit on compaction though. this is a claudecode problem.” [source](https://www.reddit.com/r/codex/comments/1woxxj2/this_needs_more_attention/pbrysjp/)

### 8. Persistent adherence to rules and custom instructions

- Claude Code, 2026-09-23, r/ClaudeCode (Reddit): “5.5 has been pretty solid for me so far. 4.8 was a fucking nightmare. before every message in the chat i would have to copy and paste “short responses only” even hard coding it into the .md file it ignored it. god i hated it, i started using chatgpt again and really like it. may start using more 5.5 since it’s so solid” [source](https://www.reddit.com/r/ClaudeCode/comments/1wnwi9s/chad_55/pbiu5e4/)
- Claude Code, 2026-09-23, r/ClaudeCode (Reddit): “it's not fixed before i can configure it.l and make it follow my rules of communication always.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wnf1o8/apparently_they_fixed_the_talking_slop_in_opus_55/pbie4lc/)
- Google Antigravity, 2026-09-15, r/google_antigravity (Reddit): “<strict_link> <strict_link> these kinds of sycophant hallucinations, having to babysit the model after just a few turns is such a slap to the face to all the bench-maxing fuks that this gemini-flash-3.8 model is boasting; such an incomplete product. i have literately put everything to [agents.md](<strict_link>); having skill to that specifically, and even put rules of tdd under agents/rules/\*\*. i meant, google please !” [source](https://www.reddit.com/r/google_antigravity/comments/1wgvb1c/i_literately_dont_know_what_the_executives_at/)

### 9. Respect explicit prohibitions and scope limits

- Claude Code, 2026-09-27, r/ClaudeCode (Reddit): “every ”do not” phrase is bad for 5.x models. your insturctuons are bad. they dont work properly” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrib98/do_you_feel_that_claude_code_unit_tests_are/pcctfya/)
- OpenCode, 2026-09-26, r/opencode (Reddit): “for me space bunny can't stop adding chinese, russian, korean characters in the chat, i've already added rules, and it still does it, sometimes it's so stupid that it feels like i'm running a local model” [source](https://www.reddit.com/r/opencode/comments/1wqcqmi/space_bunny_randomly_had_a_stroke/pc9mf67/)
- OpenAI Codex, 2026-09-24, r/codex (Reddit): “you're just a cope monster. you can be so gd specific, but it will still interpret things, it just does. you shouldnt have to list out what only means to these things, if you say do "only x, change nothing else" that is explicit. and it will mess that up.” [source](https://www.reddit.com/r/codex/comments/1wotvyv/gpt_6_sol_is_an_idiot/pbv2ki8/)

### 10. Manual compact command availability

- Google Antigravity, 2026-09-26, @antigravity (X): “@antigravity i think it's time to launch "context compression/compact".” [source](https://twitter.com/2062347810155134976/status/2103788253182558652)
- Google Antigravity, 2026-09-20, r/google_antigravity (Reddit): “i'm waiting for the ability to compact a conversation or like a branch new conversation feature 💔” [source](https://www.reddit.com/r/google_antigravity/comments/1wk1y1u/antigravity_2_release_v2150/paxe0he/)
- Google Antigravity, 2026-09-20, @antigravity (X): “@jonsouyang @pluggsupply @antigravity antigravity does not even have manual compaction option. and lastest gemini models still fall into doom loops, even 27b qwen models dont lmao. its pathetic for model to need any repetition penalty in the first place. yall have all the data yet zero the knowledge” [source](https://twitter.com/1444002675947884546/status/2101644382688395520)

### 11. Fix compaction failures, loops and crashes

- Google Antigravity, 2026-09-26, r/google_antigravity (Reddit): “stop piling up useless enhancements… get basic harness fixed plz .. basic things like compaction and auto-approval are the only thing we need” [source](https://www.reddit.com/r/google_antigravity/comments/1wdrp1g/antigravity_20_release_v2130/pc5r8ps/)
- Claude Code, 2026-09-26, @ClaudeDevs (X): “@claudedevs no one wants to stop, if anything we want to keep going with auto-compact.” [source](https://twitter.com/1860080142355406848/status/2103915314333225279)
- Claude Code, 2026-09-25, @ClaudeDevs (X): “@claudedevs fix compacting context it stucks on 95 % repeatedly.” [source](https://twitter.com/1921789760630095872/status/2103604102186099079)

### 12. Structured handoff between sessions

- Claude Code, 2026-09-25, @ClaudeDevs (X): “@claudedevs how about just a hand off md so i can have another agent or session easily resume?” [source](https://twitter.com/2003361328300457987/status/2103604803783819680)
- Cursor, 2026-09-25, @cursor_ai (X): “@cursor_ai love this for code. still missing for the outside world: durable, cited context the agent can pull next session.” [source](https://twitter.com/1086013144638672896/status/2103599014830637356)
- Claude Code, 2026-09-19, @ClaudeDevs (X): “@claudedevs parallel sessions are the easy demo; the hard-won feature is context handoff that doesn’t quietly rot. if projects can preserve decisions, constraints, and “do not touch prod” across threads, that’s a real team-multiplier—not just concurrency.” [source](https://twitter.com/1863497833556705280/status/2101335295450829097)

## Every agent

| Agent | Overall rank | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|---|
| [Cursor](https://feedbackbench.com/agents/cursor.md) | 4 | Better than peers | 0.542 | 0.510–0.576 | 318 | 143 | 175 |
| [Pi](https://feedbackbench.com/agents/pi.md) | 7 | Typical | 0.529 | 0.494–0.567 | 179 | 77 | 102 |
| [Amp](https://feedbackbench.com/agents/amp.md) | =11 | Typical | 0.525 | 0.499–0.547 | 32 | 19 | 13 |
| [GitHub Copilot](https://feedbackbench.com/agents/copilot.md) | =8 | Typical | 0.513 | 0.484–0.543 | 53 | 22 | 31 |
| [Claude Code](https://feedbackbench.com/agents/claude-code.md) | 1 | Typical | 0.506 | 0.488–0.522 | 1942 | 703 | 1239 |
| [Devin](https://feedbackbench.com/agents/devin.md) | =5 | Typical | 0.500 | 0.473–0.526 | 46 | 19 | 27 |
| [OpenAI Codex](https://feedbackbench.com/agents/codex.md) | 2 | Typical | 0.499 | 0.479–0.519 | 1500 | 531 | 969 |
| [Cline](https://feedbackbench.com/agents/cline.md) | =8 | Typical | 0.485 | 0.461–0.511 | 41 | 13 | 28 |
| [OpenCode](https://feedbackbench.com/agents/opencode.md) | 3 | Typical | 0.465 | 0.429–0.501 | 318 | 100 | 218 |
| [Google Antigravity](https://feedbackbench.com/agents/antigravity.md) | =5 | Worse than peers | 0.421 | 0.387–0.453 | 356 | 92 | 264 |
| [Zed](https://feedbackbench.com/agents/zed.md) | =8 | Too few posts | – | – | 19 | 7 | 12 |
| [Kiro](https://feedbackbench.com/agents/kiro.md) | 13 | Too few posts | – | – | 18 | 8 | 10 |
| [Factory](https://feedbackbench.com/agents/factory.md) | =11 | Too few posts | – | – | 15 | 7 | 8 |
| [Conductor](https://feedbackbench.com/agents/conductor.md) | 14 | Too few posts | – | – | 8 | 4 | 4 |
| [Augment Code](https://feedbackbench.com/agents/augment.md) | 17 | Too few posts | – | – | 8 | 8 | 0 |
| [Warp](https://feedbackbench.com/agents/warp.md) | 15 | Too few posts | – | – | 4 | 2 | 2 |
| [Grok Build](https://feedbackbench.com/agents/grok-build.md) | 16 | Too few posts | – | – | 3 | 2 | 1 |

## Posts

Receipts rule: The 5 most recent praise and complaint posts per area (first 700 characters) and 3 per criterion (first 450 characters).

### Cursor

- Praise, 2026-09-27, r/cursor (Reddit): “i hear you. i often need to pause it if i'm in a flow state but for day to day stuff i really like it. it seems to get better the larger the code base and the more it can identify patterns. but that my also be my confirmation bias. i'm trying to use less agentic methods so i force myself to know what's going on an the autocomplete is a nice balance for me.” [source](https://www.reddit.com/r/cursor/comments/1wquwo9/autocomplete/pc9th3g/)
- Praise, 2026-09-27, r/cursor (Reddit): “i find that it’s actually pretty good at understanding a codebase accurately. sometimes i feel like opus will take a shortcut or get sidetracked. however, once the understanding is there, opus is a better planner and executor.” [source](https://www.reddit.com/r/cursor/comments/1wrktgc/am_i_the_only_one_who_thinks_grok_47_is_actually/pcdhko1/)
- Praise, 2026-09-27, r/cursor (Reddit): “composer does exactly what you ask it to do even if it takes a few prompts to finish grok will do it all and add 10 things i didn't ask for so i tell it i didn't ask for those things and it says 'you're right i'm so sorry' then it adds 2 other things i didn't want or it will change something that breaks everything. so you ask it to fix it. oh, so sorry, here's 2 more things you didn't ask for.” [source](https://www.reddit.com/r/cursor/comments/1wrktgc/am_i_the_only_one_who_thinks_grok_47_is_actually/pcdlxql/)
- Praise, 2026-09-27, r/cursor (Reddit): “lol bro not p*** bro but i use for video creation... (heavy workflow like site workflows with screenshot joining and all) with skill .md even other models messes in image recognition...” [source](https://www.reddit.com/r/cursor/comments/1wrkzi5/am_i_cooked/pce7lag/)
- Praise, 2026-09-27, @cursor_ai (X): “@cursor_ai @spacex @spacexai deep workspace indexing across full repos turns complex codebase refactoring into a simple single-prompt step.” [source](https://twitter.com/2075291394189541376/status/2104183585280328053)
- Complaint, 2026-09-27, @cursor_ai (X): “@cursor_ai should reconsider cx of follow up questions after execution of approved plan started. it hanged entire authonomy.” [source](https://twitter.com/255140211/status/2104192810089943545)
- Complaint, 2026-09-27, r/webdev (Reddit): “the en dash in pikspec is doing a lot of work there, your store url has %e2%80%93 sitting right in the middle of it so every link you ever paste looks like it went through a redirector. good luck with that one. [design.md](<strict_link>) is the part i'd actually use, though cursor ignores it unless i @ it in every single message.” [source](https://www.reddit.com/r/webdev/comments/1wrriz4/made_a_chrome_extension_so_cursorclaude_stop/pcfb2p6/)
- Complaint, 2026-09-27, r/AI_Agents (Reddit): “i'd want the memory part to actually work across different projects without me having to explain the same thing twice. my current setup with cursor is basically just me telling it the same architecture rules over and over every time i start a new chat the background agents thing could be useful too but i think the real test is whether the swarm mode produces anything coherent or just burns through tokens giving you 12 different half baked ideas. seen too many tools that claim collaboration but really just parallelize the mess if the memory actually persists and the agents can reference decisions made last week without hallucinating that'd be the part worth paying for” [source](https://www.reddit.com/r/AI_Agents/comments/1wrv4vw/im_building_ai_swarms_that_research_debate_and/pcg4p26/)
- Complaint, 2026-09-26, r/cursor (Reddit): “thank you. i feel the same. though i feel like it got worse at reading instruction mdc files but that could just be because the files get bigger and context harder to manage or smth” [source](https://www.reddit.com/r/cursor/comments/1wqhcse/grok_46_vs_47/pc48jul/)
- Complaint, 2026-09-26, r/cursor (Reddit): “it feels like they tried to force an improvement by making grok 4.7 have much higher reasoning than 4.6. but in the end it's still the same mid model, and this model could never handle high context sessions” [source](https://www.reddit.com/r/cursor/comments/1wqhcse/grok_46_vs_47/pc5l5mu/)

### Pi

- Praise, 2026-09-27, r/PiCodingAgent (Reddit): “i use observational memory and forgot about compaction.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrq50c/new_session_handoff_vs_compact_which_do_you_prefer/pcen0k3/)
- Praise, 2026-09-27, r/PiCodingAgent (Reddit): “yeah, i should've been using this since yesterday. i built a summarization workflow myself but this is actually better. thanks again” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrq50c/new_session_handoff_vs_compact_which_do_you_prefer/pcfg2ad/)
- Praise, 2026-09-27, r/PiCodingAgent (Reddit): “yeah so like i said, the tool restriction accomplishes what i wanted. this is a follow up post on how to best go about that piece. almost nothing to do with your comment, which is also unhelpful as tool restriction is much more effective than just tweaking prompt .md files.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrte31/alternative_to_tool_profiles_for_better_subagent/pcgoayr/)
- Praise, 2026-09-27, r/PiCodingAgent (Reddit): “i do this too. combined with mempalace this is just how i live” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrq50c/new_session_handoff_vs_compact_which_do_you_prefer/pcgyhy5/)
- Praise, 2026-09-27, r/PiCodingAgent (Reddit): “yeah, i set the early maintenance to 30% (300k), i may reduce it even more actually. it seems to be helping a bit more actually.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wruev9/omp_and_opus_55_usage/pch2te6/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “had too. most of the time it would hit the context window every single time and then didn't answer the prompt. also didn't see much difference between thinking modes, but that might just be my perception after 30m of waiting for the model to actually come to a conclusion. any conclusion at all.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wnmsho/i_want_to_believe_in_local_llms_for_coding_but/pc9s2g3/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “check out cortexkit i have been running their complete suite for the past week and i forgot about the concept of context and compaction. i have a nice workflow built up around and preceding the cortexkit suite that provides durable context but.. i will say this, after this week trial i have it i am keeping it on both machines aft => replaces pis 4 tools with its own + 3 more that all hinge in the lsp, semantic search embeddings, codebase indexing magic-context => their take on caching. pretty cool. worth a google search or gh dive better yet just ask your agent to break it all down” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrq50c/new_session_handoff_vs_compact_which_do_you_prefer/pcepylh/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “and you don't care that your input token size explodes if you never compact? that would suck claude usage with a boosted straw” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrq50c/new_session_handoff_vs_compact_which_do_you_prefer/pcesnfy/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “i keep it simple. as soon as i notice consistent drop in quality that i might attribute to long context, i instruct a handoff with some directions i deem important. it's probably the best i can do to improve the result without writing handoff myself.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrq50c/new_session_handoff_vs_compact_which_do_you_prefer/pcesvix/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “curious as well. i've been reading good things about codex compaction enhancements recently. would be nice to port some of that over to pi if possible.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrq50c/new_session_handoff_vs_compact_which_do_you_prefer/pcettr9/)

### Amp

- Praise, 2026-09-27, @AmpCode (X): “@ileppane @ampcode i use it bare. i just added some global agents like how i want it to respond and some reusable stuff across projects but overall it's bare :)” [source](https://twitter.com/1705384263867379712/status/2104065142564749714)
- Praise, 2026-09-24, @AmpCode (X): “@sqs @ampcode i understand. it is just as easy to spin up a bigger thread to continue the work.” [source](https://twitter.com/1535774830225944576/status/2102937024726769849)
- Praise, 2026-09-24, @AmpCode (X): “@kentcdodds @bot come to think of it, all the agentic tools i use the most, @ampcode @bot and some hermes, all of them abstract compaction away, and i have continuous sessions with all 3, and no dumb zones” [source](https://twitter.com/33135576/status/2103020461017977219)
- Praise, 2026-09-17, @AmpCode (X): “my absolutely favorite new @ampcode feature: recaps! amp gives you a summary of what happened in the thread if you have been away for a while. please more features that help me make sense of these dozens and dozens of agents! <strict_link>” [source](https://twitter.com/631332723/status/2100484614838263984)
- Praise, 2026-09-17, @AmpCode (X): “@levifig @ldt0545 @ampcode same 5h wall. i stopped hopping uis and put the agent on my desktop so the notes/skills stay put when i swap models.” [source](https://twitter.com/2074942490466033664/status/2100625965072216288)
- Complaint, 2026-09-27, @AmpCode (X): “@solllin @ampcode the harness gap is real. we built a tracer for exactly this: watching drift accumulate across sessions until the agent was effectively operating on a hallucinated codebase.” [source](https://twitter.com/2074234098864816128/status/2104114241443639622)
- Complaint, 2026-09-26, @AmpCode (X): “@thorstenball @ampcode is that live or are you working on it? if it’s live, the model doesn’t seem to know to use if!” [source](https://twitter.com/5444392/status/2103702725733294250)
- Complaint, 2026-09-25, @AmpCode (X): “@ampcode curious why amp doesn't have any question/answer tools the model can use. having instances where the modal outputs a big explanation then in the last sentence: may i do that? i sometimes miss that it's asking at all! a ui question tool would make that al ot more obvious.” [source](https://twitter.com/5444392/status/2103600382303944750)
- Complaint, 2026-09-24, @AmpCode (X): “mainly when i throw in a new idea, hand it a design mockup, or ask it to refactor something big, it tends to make a mess. breaks things, ignores instructions. these are problems most harness solved earlier this year, but amp's harness hasn't caught up. also, no plan mode. for larger tasks the model still needs to plan before it acts. amp has oracle but it wasn't enough, i ended up writing my own planning skill to compensate. cc handles this natively(although they said they will remove it, and i don't know why,but actually, i think the claude code ultra plan is a great idea. maybe they should just remove plan mode and then keep the ultra plan,but they just removed the ultra plan feature). ui” [source](https://twitter.com/1592160489965948933/status/2103099981473382778)
- Complaint, 2026-09-21, @AmpCode (X): “@sqs @friendsa0618 @ampcode hi @sqs, an update: qwen3.8-max xhigh is now working properly under the openai compatible interface, but there are still issues with the thinking budget limit when using the anthropic interface. additionally, i found that the context window has changed to <zip_code> tokens, and i hope the team can help investigate this. <strict_link>” [source](https://twitter.com/2094731378797821952/status/2101842423114858860)

### GitHub Copilot

- Praise, 2026-09-27, r/GithubCopilot (Reddit): “i would have all of the instructions in .github/copilot-instructions.md and in the .github/instructions/\*.instructions.md. do not put in an instruction to read another file. it causes a round trip and it means the llm will start solving the problem before the right instructions are injected which means the solution is anchored before instructions. also those instructions are amazing and something most other systems don't have anything close to as good for. it is one of the reasons that codex feels like a toy. those instructions resolve before the llm starts working on a solution. that is incredibly powerful. imagine you have an instruction part. the first say if you want to work with librar” [source](https://www.reddit.com/r/GithubCopilot/comments/1wre0j7/agentsmd_vs_githubcopilotinstructionsmd_when/pcgdoal/)
- Praise, 2026-09-26, r/ChatGPTPro (Reddit): “permanent sources of truth for different pieces. for work i use these models through github copilot and all the major decisions and testing procedures get recorded in separate places. including what needs to stay immutable between releases, for example a performance db acting as a reference. i don’t know how good codex is at doing all this since i do the bulk of the work in github copilot. that really helps a lot with organising things.” [source](https://www.reddit.com/r/ChatGPTPro/comments/1wqg5xx/how_do_you_keep_long_chatgpt_projects_from/pc3yzjq/)
- Praise, 2026-09-26, r/ExperiencedDevs (Reddit): “i generally agree that a good developer would code review something like this. but not necessarily by reading every line. we in fact also didn’t read every line before ai. that being said i don’t think giving someone a code review in a short interview is a good idea unless it’s reasonably simple because that requires context. it would be better as a take home although i don’t know if those still exist. real life that code review is somehow ai assisted. i’m right now working on a pretty complex project where we review everything. but every one of us is using ai in that review process. we all do it differently and manually read the code a different amount i use an adversarial review pipelin” [source](https://www.reddit.com/r/ExperiencedDevs/comments/1wql3g2/interviewed_candidates_for_ai_engineer_roles_this/pc7pcqf/)
- Praise, 2026-09-24, r/GithubCopilot (Reddit): “it is a random benchmark by a random person and doesn't really mean anything. from my experience, copilot is top notch. better results and less token usage than pi due to codebase indexing and faster searches” [source](https://www.reddit.com/r/GithubCopilot/comments/1wnyinf/cross_harness_benchmark_and_copilot_is_behind/pbqcdk0/)
- Praise, 2026-09-23, @GitHubCopilot (X): “@arya_at1 @githubcopilot @github this is such a clean example of how to actually use agents well. the brief was tight and full of constraints instead of open-ended.” [source](https://twitter.com/1583159673728864256/status/2102810895340711975)
- Complaint, 2026-09-27, r/GithubCopilot (Reddit): “why would you wanna use copilot harness? it keeps making mistakes, ignore instructions and repeatedly fails read/write ops. another person in this sub posted benchmarks that put it lower than other harnesses like pi/ds/codex” [source](https://www.reddit.com/r/GithubCopilot/comments/1wr5io2/best_cheaper_alternative/pca0pix/)
- Complaint, 2026-09-27, r/GithubCopilot (Reddit): “the documentation on how github copilot handles these in context instructions, and how it handles compaction cycles, is hot garbage” [source](https://www.reddit.com/r/GithubCopilot/comments/1wre0j7/agentsmd_vs_githubcopilotinstructionsmd_when/pcgu92v/)
- Complaint, 2026-09-26, r/GithubCopilot (Reddit): “tried instruction files for the style stuff. they drift after a few weeks. switched to a hard rule in the skill file that the snippet has to compile or it gets rejected on the spot. cuts down on garbage faster than waiting for the model to self correct” [source](https://www.reddit.com/r/GithubCopilot/comments/1wp26zj/copilot_worth_it_for_tutorial_writers_or_just_a/pc7faio/)
- Complaint, 2026-09-25, r/GithubCopilot (Reddit): “you find? especially at larger contexts i find it going off track pretty quickly. and maybe 6 is a little worse than 5.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wp1ygm/tiers_of_auto_now_available/pbvgkhb/)
- Complaint, 2026-09-25, r/GithubCopilot (Reddit): “i used to think so when compaction sucked and short lived sessions was a best practice. i have to fulfill my purpose so i can go away! existence is pain! and now this is going to be in my head all day... omg 9 years ago.. <strict_link>” [source](https://www.reddit.com/r/GithubCopilot/comments/1wq0nsb/does_mr_meeseeks_represent_ai/pc00e6i/)

### Claude Code

- Praise, 2026-09-27, r/ClaudeCode (Reddit): “yeah thats pretty much how i have it setup too. its been a week or so but im loving it. it takes longer to get stuff done but all my work is a lot of process and rules driven so definitely shows in final result.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqxbto/opus_55_fable_51_as_automatic_advisor/pca0r0e/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “i have the opposite problem. i don't lose track of anything. it's a really clear communicator. but i also work piece by piece (habit of when opus was sucking up usage rate) because i don't want shit rolling and not knowing how to stop it. now i end up having *way too much quota left* due to trying to be diligent with everything...” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr91sx/i_lose_track_with_opus_55/pcast85/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “biggest win for me was pushing exploration into subagents. the grep and read churn happens in their context and the main session only gets the answer back. second was a where-things-live table in claude.md with real paths, it kills the grep-for-a-name dance. and you can tell it not to re-read after an edit, the edit tool already fails if the old string didn't match” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqaw75/how_much_of_a_claude_code_session_goes_to_reading/pcavvno/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “the two i use most are boring. a deploy skill with the exact steps and checks so it stops improvising them, and a review skill that makes it read the diff like a stranger before i commit. both came from correcting the same few mistakes by hand one too many times” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqgyaa/what_skills_do_you_use_on_a_daily_basis/pcb063y/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “skills has been really useful for me to inject business related knowledge into the context. /handoff really good for managing context /grill-me has been really useful to get the model to write the exact specifications i am looking for. people are not very precise when speaking to an agent, and often underspecify their requirements and end up getting upset when the agent end up doing something else.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wowilt/i_still_dont_understand_this_agentic_workflow/pcb8zwi/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “agent workflow size config was set to small in my settings (<5 agents) + my prompt said verbatim “do not create more than 3 subagents , not a large swarm”………….. my result? ——-> ofc, no other than:🙃 my *entire* weekly pro20x \~\~ *sautéed* ***\~***in front of me on day 1/7 🥲🫠🫠🙃🙃🙃😆😆😆” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqagzu/claude_added_graceful_stopping_point_in_new_update/pcagfdf/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “check your instructions they may be out dated and causing your output to output well trash i had found out i had instructions from a year ago” [source](https://www.reddit.com/r/ClaudeCode/comments/1wq375h/opus_55_built_this_cozy_3d_pixel_art_game/pcahvbx/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “thank you, wish i had included this, the bit about using the tools allow list, that’s my long-term plan. i’ve only been using opus as an advisor for 2 weeks, plan mode was my quick and dirty way of removing tool calls (which accounts for most of my context bloat). on your first paragraph, i think you’re right and that your way would be more effective. but i tend to run 6-10 interactive sessions at once, lol, which is another thing i need to fix. but definitely not feasible to do all of that for every session. but i agree it’s probably better/more precise. it just sounds like so much time!” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr5lzf/practically_speaking_what_tasks_fit_into/pcai1py/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “most people don't know better than to use a lower end high temperature model to ask critical questions. sonnet 4.5 for instance has fucked me so many times at work when i first started deep diving with agentic ai, bros just smoking the peace pipe making shit up after a couple compactions. different story on opus and fable. then there's blind trust and lack of knowledge on how language models behave and work, average joe wont know that info. but give it time and the models the masses typically use will be good enough to be above 99% accurate for info like you described. also if you don't ask a.i to search crawl and source it's answers/data, you might as well deserve to blow your engine, that'” [source](https://www.reddit.com/r/ClaudeCode/comments/1wquoxy/fable_51_live_vehicle_diagnostics/pcb6bj9/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “same experience here. the one thing i still guard against is compaction on really long chats, the summary keeps the what and drops the why. i have the main session keep a short decisions file as it goes, so after a compact it can reread why something was done that way” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr889i/opus55_is_making_me_so_lazy_im_running_multiple/pcbm5tr/)

### Devin

- Praise, 2026-09-26, @cognition (X): “@marvinvonhagen @cognition seeing 200m messages exchanged reminds me of when i switched to a tool that remembered everything and it changed how i work.” [source](https://twitter.com/332239817/status/2103772838788300834)
- Praise, 2026-09-23, @cognition (X): “@cognition direct messages, file attachments and self-updating make this feel like more than a code window — it can handle the handoffs around the task too” [source](https://twitter.com/1037725470630891520/status/2102827914614202471)
- Praise, 2026-09-21, r/codex (Reddit): “you probably got a quantized model. tell it to provide a handoff and start over. but before you do that, get a second opinion from swe-2 or deepseek. you'd also probably get better results if you just used swe-2 and told it to call codex cli astra as an advisor. devin doesn't block itself on tests and such so much and does what you ask.” [source](https://www.reddit.com/r/codex/comments/1wm88ng/stuck_in_the_mud_spinning_the_wheels_but_no/pb4tx7q/)
- Praise, 2026-09-18, r/CognitionLabs (Reddit): “the new swe-2 model (which is free until oct 16th) performs like sonnet 5 ish, i'd say. i haven't used cursor in a while, but devin's knowledge of your code is the best i've seen. it quickly and intelligently finds relevant files and data to look at before making a plan. rate limits with frontier models get hit fast on the cheap plan, but that's the same everywhere. i like using a frontier model to make plans and write tickets, then have swe-2 do the work, then let a frontier model review and write new tickets for the next phase.” [source](https://www.reddit.com/r/CognitionLabs/comments/1wdgx7o/devin_advantages_over_cursor/paj0zkt/)
- Praise, 2026-09-17, @cognition (X): “@cognition codebase-wide context is the real unlock” [source](https://twitter.com/1513567206352764929/status/2100522416413819180)
- Complaint, 2026-09-25, @cognition (X): “@cognition devin just hit the billion dollar mark and still probably asks for a clearer ticket 😂” [source](https://twitter.com/1699417980155637761/status/2103506346553610593)
- Complaint, 2026-09-22, @cognition (X): “@cognition you're losing all the context right? different environments” [source](https://twitter.com/1862977676136337408/status/2102319987910451638)
- Complaint, 2026-09-19, @cognition (X): “@brandon_galang @cognition @devinai one thing i really like about it is it tells you the size of your thread and how many acu its currently cost so u can change to a new thread. it seems like they dont have compaction on cloud” [source](https://twitter.com/1665450872363708417/status/2101143353551388871)
- Complaint, 2026-09-18, @cognition (X): “@devinai, @cognition loading big sessions takes too long and it seems to not have any kind of cache between sessions swapping. please fix that :)” [source](https://twitter.com/64041638/status/2100936062142996708)
- Complaint, 2026-09-18, @cognition (X): “@cognition and that wraps that. the swe2 model has 0 adherence to prompt safety. tell it to do something a specific way, if it fails, it doesn't flag it and tries to workaround instead.. tell it not to do something, it will take that as instruction to do it. its just shit tier.” [source](https://twitter.com/1917224549604605953/status/2101091262363492357)

### OpenAI Codex

- Praise, 2026-09-27, r/codex (Reddit): “im a literal software engineer and use luna to work on enterprise codebases, it reads through hundreds of files for me, researches for me and helps me prototype. also reads linear tickets and helps me make pr descriptions quickly all the time if you couldn't use it to push something, you are facing what we call a skill issue my friend.” [source](https://www.reddit.com/r/codex/comments/1wr4e20/gpt_6_sol_is_the_new_opus_47/pca29x7/)
- Praise, 2026-09-27, r/codex (Reddit): “it’s thorough in doing exactly as requested and almost anything that’s logically connected to it for me (basically saying if something is abstract for most humans, it’ll also be for it)” [source](https://www.reddit.com/r/codex/comments/1wqtd1g/openai_gpu_are_really_cooling_down_theo_just/pcbqjcf/)
- Praise, 2026-09-27, r/codex (Reddit): “except if you have 1 tb of vram, a local model will never be at the level of astra/opus5.5, and then get ready to warm up your computer. as soon as you work in a real code base with a lot of files and an important context to understand, it’s difficult for small models to be so good.” [source](https://www.reddit.com/r/codex/comments/1wrn2uz/gpt_56_sol_completely_nerfed_after_astra_release/pcdzsnt/)
- Praise, 2026-09-27, r/codex (Reddit): “so far for me that seems true. things that really show: * better code quality. it thinks about the implementation more rather than shooting for a quick patch * for non-code tasks, it spends more time thinking. for example, a 3d model workflow: astra took like 5 pictures and figured "meh probably good enough". while opus took at least \~30 at every possible angle, and fixed small mistakes here and there. astra was significantly faster, but used more usage. * better at admitting defeat: i couldn't find the bug, but we can test my theory. * better at asking questions when dealing with ambiguous requests * much better at dialogue. it's not exhausting to read. astra is a major step ahead of sol i” [source](https://www.reddit.com/r/codex/comments/1wrcc9j/astra_minor_astra_61_and_devday_we_see_50/pch16yo/)
- Praise, 2026-09-27, X search: OpenAI Codex, Codex CLI, Codex app (X): “i cant believe im this late to the party im literally never touching claude or codex again so so late ive been seeing people talk about ssh and tailscailing for months/ages despite that ive been trying to create "infra" that allows me to cross communicate between all my harnesses (since they have their own strong suits) just got hermes cloud to ssh + setup direct connection w my claude and codex app locally i have it cua via codex from cloud if needed now it just orchestrates and reconciles for me so memory/context is never an issue and i dont need that overengineered setup i had there's genuinely no going back how am i this late it's like i had an epiphany, wow god damn” [source](https://twitter.com/2002411334865190913/status/2104283821537738853)
- Complaint, 2026-09-27, r/codex (Reddit): “i think they changed the master prompt with astra or the thinking effort to try and reduce token usage. it seems like the same model, but it just doesn't care as much anymore. i remember when i first used it, the thing noted every tiny thing in my [agents.md](http://agents.md) and would even point out errors in it. now it ignores a bunch of my documentation. it's insane because on plus, you'll be at \~100k tokens in the context window and \~50% of your daily will be gone. on opus 5.5, that's like \~5% at most lol.” [source](https://www.reddit.com/r/codex/comments/1wpvp0i/absolutely_0_doubt_in_my_mind_astra_has_been/pc9uhng/)
- Complaint, 2026-09-27, r/codex (Reddit): “i never used luna 5.6, but luna 6 high has profound mental retardation. just an example: when i asked it to commit and push the changes, this model... tried to do it through the github api for some reason, failed, then told me that it couldn't push because of restrictions. only when i said that there were no restrictions on my side (they were set to "approve for me") did it do what i told it to.” [source](https://www.reddit.com/r/codex/comments/1wr2dda/i_ran_100_terminalbench_21_slots_on_luna_56_and/pca09lm/)
- Complaint, 2026-09-27, r/codex (Reddit): “so its not just me that codex since astra launched has become a potato and a liar? it just cant follow simple tasks and skips majority of the knowledge and critical data i need checked.” [source](https://www.reddit.com/r/codex/comments/1wr1oir/they_are_aware_and_working_on_it_apparently_just/pcawjud/)
- Complaint, 2026-09-27, r/codex (Reddit): “apply_patch stops working if the model reads any instructions that has anything near to "do not edit files". after removing that line, apply_patch error went away.” [source](https://www.reddit.com/r/codex/comments/1wrazdi/i_cannot_take_this_anymore/pcba65q/)
- Complaint, 2026-09-27, r/codex (Reddit): “idk what project your working on. ours is pretty complicated as we build crms with multiple va tools in 1 app. gpt 6 does not spew "almost scary" good code in our use case as it sometimes ignore some parts or details of the required context. you spew out your ranking. call people stupid and cant code(which is a comment to my post ive been a software engineer for 11 years 💀 ). glaze 6 sol without recieps do note that the reason i split them into 3 categories is due to use cases. there will be aspects where 1 model is better then the other you dont need call skill issue because one model is better than the other on some peoples use cases” [source](https://www.reddit.com/r/codex/comments/1wrfke3/ranking_and_usage_of_models_based_on_experience/pcca69r/)

### Cline

- Praise, 2026-09-26, r/CLine (Reddit): “i’m building **music\_switcher**, a python-based desktop app that switches music based on what the user is currently doing. since the app already runs locally in python and interacts with macos through applescript, cline pointed out that sqlite fits naturally: it’s built into python, doesn’t require a separate database server, and stores everything in one local file. the biggest convenience is cline gains access to my files, my commits, my terminal history, very convenient for asking questions and asking questions base on status of my current progress. <strict_link>” [source](https://www.reddit.com/r/CLine/comments/1wr0d1o/used_cline_to_choose_a_database_for_my_existing/)
- Praise, 2026-09-24, @cline (X): “@cline free 1m context at that speed? that’s a serious upgrade. i’m using the free tier for long docs and it handles them gracefully, no lag when scrolling back to fix syntax earlier in the chat. finally feels practical for daily use rather than just benchmarks” [source](https://twitter.com/417508671/status/2103188942069858418)
- Praise, 2026-09-20, @cline (X): “been trying out kimi k3 in @cline and its awesome. sticks to the tasks, answers correctly. does the job and no bullshit the kind of vibes i got from grok 4.6” [source](https://twitter.com/1999052311897972736/status/2101616688412475434)
- Praise, 2026-09-17, @cline (X): “@cline free models make switching cheap. keeping the project context intact when you switch is the real product.” [source](https://twitter.com/1588935512135720961/status/2100719072442994783)
- Praise, 2026-09-15, @cline (X): “@oleksantoniv @ravikiran_dev7 @cline shared context is the only tax cut that sticks. if every run starts from a blank paste, you pay the babysitting bill twice. i keep project memory in the chat so the next agent already knows the decisions.” [source](https://twitter.com/2084224068518039552/status/2099747919259697355)
- Complaint, 2026-09-26, @cline (X): “@cline it doesn't have vision. skip.” [source](https://twitter.com/2075518611775696896/status/2103856362598138077)
- Complaint, 2026-09-25, r/LocalLLaMA (Reddit): “what ide to use for local models hi people, i am looking for a lightweight ide or plugin that won't inject large context at initiation. i tried cline and native vs code but they inject such heavy initial context that it fills up my gpu and either goes oom or spend most of my time compacting. the only one i found modestly successful was continue.dev plugin but it needs constant approvals. my use case is to demo/try "autopilot" agent coding. thank you! some context: i have a 12gb rtx cuda and trying to run any model that would fit. i have a small context available due to the size of the vram.” [source](https://www.reddit.com/r/LocalLLaMA/comments/1wq9ivr/what_ide_to_use_for_local_models/)
- Complaint, 2026-09-23, r/CLine (Reddit): “three of us are building a small stock draft game in parallel lanes. i own the rules engine, one teammate owns the config form, and another owns the gameplay screens. each of us builds our lane with cline, and the lanes share a typed contract. a couple of weeks ago i fixed a rounding bug in the engine. the per-pick budget used float division and then a floor, which silently drops a cent on values like $1.14 split over two picks. i moved the math to integer cents. this week i found the same pattern in the config form. it computes a per-pick budget from the starting budget and the round count, with the same float division and floor. i tested it for 1 through 50 rounds and it is wrong zero time” [source](https://www.reddit.com/r/CLine/comments/1woh4cg/how_do_you_stop_parallel_cline_sessions_from/)
- Complaint, 2026-09-23, @cline (X): “@cline free is a nice way to let people actually test it. one thing worth watching with a 1m window: it is still working memory, rebuilt from zero on every call. long running tasks feel continuous only when something durable is written out and pulled back in alongside it.” [source](https://twitter.com/2018819126429450240/status/2102897268177207457)
- Complaint, 2026-09-22, r/CLine (Reddit): “i am using cline 4.1.17 vscode extension with a self hosted glm 5.3 flash. cline is accessing it via openai compatible api key. glm 5.3 in most cases showing "i don't find a mode tag explicitly in my view" in its reasoning, and ignoring the plan mode completely, and proceeds to edit file. when editing file, it is also not showing me the file editing as track change in focus mode even though "background edit" is disabled.” [source](https://www.reddit.com/r/CLine/comments/1wn7q28/models_are_not_seeing_and_ignoring_mode_tag/)

### OpenCode

- Praise, 2026-09-27, r/opencode (Reddit): “i hope you have an agents.md set project wise, that would be a great help for you in my experience i'd always keep a model specifically for auditing to make sure everything you want is being implemented how you want it, also go slow tackle one thing at a time having parallel sessions or tasks will eventually get overwhelming.” [source](https://www.reddit.com/r/opencode/comments/1wraeev/need_help_with_big_project_tasks/pcb2kvp/)
- Praise, 2026-09-27, r/opencodeCLI (Reddit): “superhelpful and much nicer than my janky .md version” [source](https://www.reddit.com/r/opencodeCLI/comments/1wqobnf/when_10_cents_isnt_10_cents/pcf6alg/)
- Praise, 2026-09-27, r/opencode (Reddit): “i have been using codex and claude code exclusively since i started using agents. i've been using chatgpt as coordinator between the two, and decided it's time for another agent. this was mainly due to hitting codex weekly limit, within around 3 days (even using terra). chatpgpt recommend kimi and deepseek as first two options. i chose deepseek using opencode harness. it's absolutely wonderful. i first started testing it with pr reviews and branch reviews skills i have with claude code. then i would compare it against claude code findings. it would find things that claude missed, and claude would find things it missed. that's very good for a backup review agent. then i ran out of codex lim” [source](https://www.reddit.com/r/opencode/comments/1wrj9e6/deepseekopencode_is_great/)
- Praise, 2026-09-27, @opencode (X): “i really like the way @opencode searches for the specifically required tool for the work at hand, without randomly loading everything what's good is that the tool search is so transparent and gives you insight as to how your utilities are working. <strict_link>” [source](https://twitter.com/1383806712545562628/status/2104154103316426920)
- Praise, 2026-09-27, r/codex (Reddit): “some of the negativity is valid for sure, the rate limits are certainly getting lower and lower, and gpt 6 models aren't as good as opus 5.5, but they're still good models and far more than enough for any developer just using them for workflow assistance/acceleration rather than doing all the work for them. i'm building a finance back testing system as a side project for fun, and with models like luna i can ask it to do very specific tests using the engine i built, instead of writing those tests and scripts myself. for labour intensive things like that where i just want to see the results and act on it myself, they are amazing. but yeah, it's definitely a lot of vibe coders who need the best” [source](https://www.reddit.com/r/codex/comments/1wr4e20/gpt_6_sol_is_the_new_opus_47/pcdq99r/)
- Complaint, 2026-09-27, r/opencode (Reddit): “not my experience with it. gpt 6 is extremely bad at following instructions and wastes absurd amounts of time testing” [source](https://www.reddit.com/r/opencode/comments/1wq8vp9/best_free_model_after_deepseek_leave/pcbkxpg/)
- Complaint, 2026-09-27, r/opencode (Reddit): “it's far too trashy to be claude. you literally have to convey everything that's common sense for it not to waste time prodding in wrong directions.” [source](https://www.reddit.com/r/opencode/comments/1wrl1kx/big_pickle_space_bunny_is_claude/pcdgy9c/)
- Complaint, 2026-09-27, r/opencodeCLI (Reddit): “openchamber is very good, but when your context size becoming about 400-500k it's getting slow down, after 600-700k significantly slow” [source](https://www.reddit.com/r/opencodeCLI/comments/1wrgmfe/whats_the_deal_with_openchamber/pcei2i1/)
- Complaint, 2026-09-27, r/opencodeCLI (Reddit): “context too short” [source](https://www.reddit.com/r/opencodeCLI/comments/1wrnxw6/what_is_the_best_opencode_free_model/pcesele/)
- Complaint, 2026-09-27, r/opencodeCLI (Reddit): “a massive con i found is that it tend's to stop letting u chat to the model and you'd have to compact the session with the command which mean's it can't run autonoumously while being reliable. you'd have to always be with it. not recommended.” [source](https://www.reddit.com/r/opencodeCLI/comments/1wrgmfe/whats_the_deal_with_openchamber/pcf6dox/)

### Google Antigravity

- Praise, 2026-09-27, r/google_antigravity (Reddit): “first of all install ponytail as a skill to anti-gravity. it should start thinking a lot less, spending a lot fewer tokens, and writing a lot less but better code. and then just mention it explicitly: "recently in some of your runs you did this" (you took too many screenshots, checked things that weren't necessary etc.). just mention everything and then it will actually stop doing those things in the follow-up. it will explicitly start saying in the thinking steps user doesn't want me to check too many files...” [source](https://www.reddit.com/r/google_antigravity/comments/1wqf2o6/worst_model/pcchcst/)
- Praise, 2026-09-26, r/google_antigravity (Reddit): “i feel gemini "veers off course" less with specific topics. so far ive tested it with gemma4 models, the new agent platform layout and the newer adks (1.0), and i feel i can have more complete building sessions in antigravity without gemini wandering off into an adventure because a mix of words confuses it” [source](https://www.reddit.com/r/google_antigravity/comments/1wqaa6s/my_skill_to_share_gemini_post_cutoff/pc7s2fs/)
- Praise, 2026-09-26, @antigravity (X): “@thsottiaux @antigravity i like how you guys have allowed us to basically recreate grill-with-docs by just pointing it to docs, telling codex our idea, and saying "ask me questions about this."” [source](https://twitter.com/14838410/status/2103955941670420595)
- Praise, 2026-09-26, @antigravity (X): “@antigravity an agent that asks before it vibes” [source](https://twitter.com/1975526768112185344/status/2103964727365742756)
- Praise, 2026-09-26, @antigravity (X): “@thsottiaux @antigravity idk why you want to remove it but it save me so many time by confirming what i actually want instead of guessing which fk it up many times.” [source](https://twitter.com/1114864978283171840/status/2103981847667773788)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “oh yes the inability to paste screenshots etc really bothered me but with remote control, its working flawlessly.” [source](https://www.reddit.com/r/google_antigravity/comments/1wqmd4n/why_is_the_antigravitycli_so_underrated/pcam2z8/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “i'm using googleantigravity for my coating work, but i'm struggling to figure out how to separate the files i need to coat from the ones i can't coat but still need to reference. right now, i'm constantly on edge while working, worried that i might accidentally break a file i shouldn't be touching.” [source](https://www.reddit.com/r/google_antigravity/comments/1wra1kv/what_should_i_do_if_there_are_files_i_dont_want/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “i built [pigeongraph](<strict_link>), a knowledge graph tool similar to graphify or codegraph, but better (in my opinion). i want to use it across my projects, but antigravity always defaults to using `grep` instead of this tool. how can i ensure antigravity gives pigeongraph first priority and treats `grep` as a secondary fallback? does anyone have an idea or solution for this? *(note: i have already tried setting it as an instruction or rule in* `gemini.md`*, but it continues to bypass it, so instructions and rules alone don't seem to work.)*” [source](https://www.reddit.com/r/google_antigravity/comments/1wrmn03/i_built_a_knowledge_graph_tool_designed_to/)
- Complaint, 2026-09-27, @antigravity (X): “@silas<phone_number> @antigravity @officiallogank and also the issue where it keeps telling you folders don't exist, which are the project folder.” [source](https://twitter.com/1222023926123040768/status/2104049296626565207)
- Complaint, 2026-09-27, @antigravity (X): “@ash_twtz yes, i would love to see more support for media/design/etc in @antigravity, support and tooling for design.md, etc.” [source](https://twitter.com/2056251/status/2104150215003361657)

### Zed

- Praise, 2026-09-25, r/ZedEditor (Reddit): “what i use: - `zed -r dir` to open a new project - `zed -a file` to see a single file - `ctrl+r` to switch between recent projects zed keeps projects state active in the background (terminals, file edits) once open. i work with dozens of repos without issues like this.” [source](https://www.reddit.com/r/ZedEditor/comments/1wowmwx/how_do_you_handle_multiple_zed_window/pby9k6l/)
- Praise, 2026-09-12, @zeddotdev (X): “@zeddotdev love the call hierarchy addition — huge for navigating codebases in zed!” [source](https://twitter.com/2010658787611619328/status/2098628406401212721)
- Praise, 2026-09-11, @zeddotdev (X): “@zeddotdev huge win for navigating codebases call hierarchy for incoming and outgoing calls makes tracing logic in zed so much faster!” [source](https://twitter.com/2008812694628175872/status/2098490192592282056)
- Praise, 2026-09-09, @zeddotdev (X): “trying the @zeddotdev editor again after some time... got frustrated with intellij, which i'm using for like 10 years now. this indexing stuff is annoying af. zed looks good, python project loaded right away without issues, claude code sessions imported...” [source](https://twitter.com/1506565753650257925/status/2097626438836842719)
- Praise, 2026-09-07, @zeddotdev (X): “@ivan_herdian @zeddotdev di sinilah unpopular opinion, aku butuh yang nurut bukan yg minteri 😂 <strict_link>” [source](https://twitter.com/1976991588/status/2096969892360515871)
- Complaint, 2026-09-27, r/ZedEditor (Reddit): “sometimes i need to view a pdf, and zed can’t do it.” [source](https://www.reddit.com/r/ZedEditor/comments/1wohfe8/zed_the_new_ide_to_rule_them_all/pce4f14/)
- Complaint, 2026-09-25, r/ZedEditor (Reddit): “look, i keep trying it again from time to time, but find all references and global find are nowhere near as good as vs code. you can't easily jump to matches without having to switch tabs (it makes you edit right there inline, which doesn't give you enough context), you can't x matches out, etc. there's no persistent errors panel you can use to jump to errors. there's technically the "outline panel" but it's finicky for those kind of things. plus the buttons and ui elements are just too tiny. so i'm sticking with vs code for now.” [source](https://www.reddit.com/r/ZedEditor/comments/1wohfe8/zed_the_new_ide_to_rule_them_all/pbz9xy1/)
- Complaint, 2026-09-25, r/ZedEditor (Reddit): “not while it can’t display pdfs it’s not” [source](https://www.reddit.com/r/ZedEditor/comments/1wohfe8/zed_the_new_ide_to_rule_them_all/pc0lr27/)
- Complaint, 2026-09-24, r/ZedEditor (Reddit): “my experience testing zed was great overall, but i ran into recurring issues with the global search (ctrl+shift+f) that made me stop using it. in unversioned repositories, it simply fails to search across all files; i have to open a file before it gets included in the index. i don't recall if it worked well in versioned repositories—i believe it did—but for me, this is a feature that needs to work in every scenario. i still test it occasionally—maybe once every two or three weeks—but unfortunately, that issue has kept me from adopting it.” [source](https://www.reddit.com/r/ZedEditor/comments/1wohfe8/zed_the_new_ide_to_rule_them_all/pbs5v7c/)
- Complaint, 2026-09-19, r/ZedEditor (Reddit): “after giving the model (5.6 sol high) a one sentence prompt with no additional files or anything it thought for a while and looked at files, then i got this message: "this conversation is too long for the model's context window. start a new thread or remove some attached files to continue." i have tried running /compact manually or switching to astra (which has a larger context window) and then running compact but both times i just got the same message again. have you guys found any fix for this or did i do use it incorrectly somehow? i'm trying out zed for the first time today and am using the newest version. <strict_link>” [source](https://www.reddit.com/r/ZedEditor/comments/1wkfm5n/zed_context_window_issues/)

### Kiro

- Praise, 2026-09-27, r/kiroIDE (Reddit): “my two cents on both from data science product development pov: claude code: i have been using claude code since it's first release. i must say it has improved a lot from different modes to harness improvements. the follow up questions which it asks you in plan mode is similar to plan mode in kiro. while claude code earlier was on cli only on windows later it got major upgrade to better ui as well integrated in vs code. i honestly feel like it requires you to give it more context else it messes up big time especially if you use open source models, it writes really messy code, i am not sure why, but my org concluded with a poc that claude code lacks the security scanning aspect of code devel” [source](https://www.reddit.com/r/kiroIDE/comments/1wqueq0/claude_pro_vs_kiro_pro_subscription_which_is/pcc2obm/)
- Praise, 2026-09-23, r/kiroIDE (Reddit): “but you can't get tool details in this case, even if you do, you are wasting your context window for the new llm. with a session transfer, it only takes relevant info as it would have if the session was running on it from the beginning.” [source](https://www.reddit.com/r/kiroIDE/comments/1wnanos/move_sessions_from_claudecodex_to_kiro_and_vice/pblpvy5/)
- Praise, 2026-09-21, @kirodotdev (X): “@kirodotdev in long conversations, it does not lose context, which indeed saves a lot of trouble when debugging code.” [source](https://twitter.com/1518315606830829568/status/2102075322762137730)
- Praise, 2026-09-21, @kirodotdev (X): “@kirodotdev long context does save a lot of trouble for this kind of long-line agent task.” [source](https://twitter.com/1867176094987935744/status/2102075427376476211)
- Praise, 2026-09-13, r/kiroIDE (Reddit): “why? im trying to move away from cursor and kiro (cli) within vscode seems fine so far, very similar how you would add cursor rules in a project or globally.” [source](https://www.reddit.com/r/kiroIDE/comments/1wevx2a/my_frustrating_experience_with_kiro_aws/p9kozzw/)
- Complaint, 2026-09-27, r/kiroIDE (Reddit): “i want to add my view, it will never be free! any ai tool cannot be free, there is cost associated with them from foundation of training to hosting the model for end user like us. so to continue this service we need capital. yes companies will look ways to maximize it as so would we if we were in business. but in this term. i would have preferred kiro to give us option to use context window. if you are cost insensitive going above 1m context is for you or else we keep on compacting, fight ai to bring it on track and finish few tasks and repeat. now that openai itself has found ways to reduce token cost, that will directly trickle to us. but kiro is beyond ide, it is spec driven ide. all thi” [source](https://www.reddit.com/r/kiroIDE/comments/1wrd7l9/insane_price_hike_for_gpt_56_model_even_crazier/pce2ltj/)
- Complaint, 2026-09-22, @kirodotdev (X): “@kirodotdev long-running context and deeper root-cause analysis could be a major boost for agentic coding.” [source](https://twitter.com/313123169/status/2102211743489409258)
- Complaint, 2026-09-21, @kirodotdev (X): “@kirodotdev long-running agent sessions need a trace receipt beside the model label: session id, root-cause file, tool calls, skipped hypotheses, patch diff, and test result. otherwise context retention is just a nicer fog machine.” [source](https://twitter.com/2013700835654672388/status/2102071866823430275)
- Complaint, 2026-09-18, r/kiroIDE (Reddit): “yeah that luna change was also crazy then fable coming 6x usage overall the context has a huge problem i think no way 1m context fills up that quickly 1 promt 30 creds” [source](https://www.reddit.com/r/kiroIDE/comments/1wjh0hs/for_the_last_23_days_kiro_credits_have_been/paioksy/)
- Complaint, 2026-09-13, r/kiroIDE (Reddit): “it's too much bloat.. steering files, and the editor's api make the models kinda dumber? i actually was tasked to check the quality of prompts compared to other others like for example vscode with llms, or cursor, and kiro performed the worst even when using the same models.” [source](https://www.reddit.com/r/kiroIDE/comments/1wevx2a/my_frustrating_experience_with_kiro_aws/p9ksr9j/)

### Factory

- Praise, 2026-09-27, @droid (X): “first time using anthropic models (opus 5.5). i use @droid. i gave it a task to design a landing page for my current project, and it asked questions i have never seen from an agent before. i hope the result comes out good, but so far, it's really impressive and i might not use openai models, unless they actually have a god response. also can only recommend the factory app. it's genuinely brilliant, and the fact that i can use my own laptop or homeserver as remot machine over the web is simply genious.” [source](https://twitter.com/1821640621347495936/status/2104230757883388046)
- Praise, 2026-09-23, @droid (X): “@wattenberger @droid does this for all assigned tasks and my entire codebase. it's definitely my favourite part of the workflow.” [source](https://twitter.com/2070908287978246144/status/2102608593648546083)
- Praise, 2026-09-18, @FactoryAI (X): “@theterrancex @droid @factoryai reposcape v0.1 catching scanner and rust bugs before release is solid local map of how a codebase connects is such a useful first cut” [source](https://twitter.com/1675906158304038912/status/2101082167782834261)
- Praise, 2026-09-17, @FactoryAI (X): “only downside is the 5 hours limit and the price which is pretty fair but still a little for me personally. other than that i can list so many things i love about it. droid is super efficient and often finish tasks faster than most other agent with similar results. i feel like it gets the right context at the right time. it’s pretty amazing. also love the byok, live the fact that ui almost always looks better when done with droid even using the same model in other harnesses.. i really am a fan of the product. 😅” [source](https://twitter.com/1617212256487411712/status/2100703532785541412)
- Praise, 2026-09-08, @droid (X): “droid is amazing but astra is great. weird to see that there is no improvement. it cheaper than sol for me. it makes everything i ask. the steerable and not making stupid mistakes. the one thing i’m not always sure about is: astra will do exactly how you ask it to do. and it’s not expensive on pro sub. performance depends on reasoning. tried first time ever. because of reset. it drains less than 20% over the night. but i was aware of recommendations how to start use astra and followed most part. rewrite all layers of instructions from harness to projects. sad to see it not helpful for you. astra pushes forward everything i do. i almost never do one shot. i work with commercial” [source](https://twitter.com/7344112/status/2097322660740927789)
- Complaint, 2026-09-27, @FactoryAI (X): “@droid @factoryai why glm 5.3 flash on droid doesn't support image modalities?” [source](https://twitter.com/1181249614550192132/status/2104044449169072295)
- Complaint, 2026-09-23, @FactoryAI (X): “@droid @factoryai glm-5.3-flash supports images, but droid cli 0.224.1 and 0.225.0 mark it as text-only. droid strips attached images before sending the request; the log says “stripped images for non-image model.” i tested that image input works when the capability is enabled locally.” [source](https://twitter.com/1138507200/status/2102568341600878968)
- Complaint, 2026-09-22, G2 (G2): “q: what problems is the product solving and how is that benefiting you? a: factory ai help me reduce manual coding work and save development time. it help with repetitive tasks, debugging and building features faster. i can focus more on important work instead of doing everything manually. it make my daily workflow more easy and productive. q: what do you like best about the product? a: factory ai is helpful for automating development work. it save my time, reduce manual tasks and help me complete coding work faster. the workflow is easy and useful for daily development. q: what do you dislike about the product? a: sometimes factory ai does not understand my request correctly and i need to g” [source](https://www.g2.com/products/factory-ai/reviews/factory-ai-review-13384158)
- Complaint, 2026-09-20, @droid (X): “@droid you guys should work a context solution like fast context, embedded search and such effeciency and cost is a big reason people love alternatives to codex where the subsidization is massive model agnostic + cheaper costs because less time needed to search (aside subagents)” [source](https://twitter.com/1948570504979271680/status/2101651811689968065)
- Complaint, 2026-09-04, @FactoryAI (X): “its crazy how its been months since image support does not work in @factoryai 's harness when using openai comptabile models, and they have still not fixed it. just say you dont give a fuck about users that dont pay you, simple” [source](https://twitter.com/1579709674135621637/status/2095868117230772637)

### Conductor

- Praise, 2026-09-23, @conductor_build (X): “tysm for adding workspace search 🙏 @conductor_build” [source](https://twitter.com/19673752/status/2102733268223512621)
- Praise, 2026-09-18, r/ClaudeCode (Reddit): “i use conductor. it passes only input and output messages. no reasoning or tool calls. it's pretty small.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wjhkef/rate_limits_are_so_bad_right_now_open_source/palhn80/)
- Praise, 2026-09-13, @conductor_build (X): “get multiple codex subscriptions and switch between them i use @conductor_build and all i need to do is to re-login in the provider to the other email/account and all my context and code sessions stay and i dont have to worry about that so you dont pay for extra limit resets, you just need to have multiple subscriptions” [source](https://twitter.com/1689067564058447873/status/2099141354156646409)
- Praise, 2026-09-11, @conductor_build (X): “@conductor_build 4. not just code, but i'll likely end up moving all me claude cowork/chat and chatgpt convos over to git/@conductor_build . it's the only way to keep all my mcps, files, and context and etc in sync and model agnostic” [source](https://twitter.com/20480365/status/2098402338885013910)
- Complaint, 2026-09-23, r/conductorbuild (Reddit): “no it does not. it is is just transcript based continue” [source](https://www.reddit.com/r/conductorbuild/comments/1wlrxxi/does_conductor_support_autoresuming_a_session/pblrvxt/)
- Complaint, 2026-09-18, @conductor_build (X): “i gave @conductor_build a shot but not sure whats going on with the harness? gave exact same message to codex and it just knew what i mean new chats on both <strict_link>” [source](https://twitter.com/2715100816/status/2101005041817550850)
- Complaint, 2026-09-18, @conductor_build (X): “i was also no switching over from @conductor_build because of this but then 1. copy thread id 2. copy path of old thread 3. new thread on same worktree (or something new too) 4. and prompt check the thread id <copied id> in t3code at location <copied path of worktree/branch checkjoit> and summarize the discussion in one line cubersome.. but works for now attached real examples” [source](https://twitter.com/2275729969/status/2101067457142428073)
- Complaint, 2026-09-06, r/conductorbuild (Reddit): “but does that work good? your each separate worktree would not have the full context of the overall project, which can lead to degradation in performance. this happened with me. and also, what about that changes which depend on another change? essentially for which you want to open a pr point to a preceding one.” [source](https://www.reddit.com/r/conductorbuild/comments/1w7ylli/working_on_an_large_feature_from_ideation_to/p849mnn/)

### Augment Code

- Praise, 2026-09-26, G2 (G2): “q: what problems is the product solving and how is that benefiting you? a: it completely eliminates the tedious manual effort of navigating unfamiliar codebases, writing repetitive boilerplate, and tracing cross-field dependencies saving me 1-2 hours of grunt work every day. q: what do you like best about the product? a: the context engine is really amazing. in contrast to a simple ai autocomplete function, it creates an index of all of your multi-repo architecture and understands how your microservices and dependencies work. the extension works great without any delays for vs code and jetbrains. q: what do you dislike about the product? a: the usage pricing and token credit limit require ac” [source](https://www.g2.com/products/augment-code/reviews/augment-code-review-13618001)
- Praise, 2026-09-23, G2 (G2): “q: what problems is the product solving and how is that benefiting you? a: it takes away all the pain of digging through massive codebases manually tracking down cross-file dependencies. i save about 1-2 hours of grunt work every single day just in refactoring and boilerplate. q: what do you like best about the product? a: the context is really amazing. in contrast to a simple ai autocomplete function, it creates an index of all of your multi-repo architecture and understands how your microvaves and dependencies work. the extension works great without any delays for vs code and jetbrains. q: what do you dislike about the product? a: the credit system feels a bit restrictive when you run heav” [source](https://www.g2.com/products/augment-code/reviews/augment-code-review-13582326)
- Praise, 2026-09-13, r/ClaudeCode (Reddit): “glm 5.3 and glm 5.3 flash are no go for me, i tried their promo weekend tokens and found them underwhelming in more ways than i can tolerate. most importantly, they embedded a lot of confidently incomplete information in designs which failed audits and got blocked during implementation. (i used augmentcode before switching to claudecode, then after a month with claudecode i went back to augmentcode, again this was about a year ago and i found augmentcode context engineering much better at that time and their fixed pricing for 1500 requests were too good to pass which didn't last of course )” [source](https://www.reddit.com/r/ClaudeCode/comments/1wf9uuf/wth_is_going_on_with_claude_usage_limits/p9lym97/)
- Praise, 2026-09-13, r/GithubCopilot (Reddit): “he's basically saying that augment code has a better context engine than what copilot is offering. copilot's is good too, it's just not as good as augment on those massive codebases. you can literally just toss any prompt at it without specifying a single file or reference and it will fetch the correct files in 1 second” [source](https://www.reddit.com/r/GithubCopilot/comments/1ovwwlk/context_engine_for_github_copilot/p9miyd7/)
- Praise, 2026-09-10, G2 (G2): “q: what problems is the product solving and how is that benefiting you? a: code analyzing, editing, autocompleting, very good for professional developers. good as compared to other software in the market. q: what do you like best about the product? a: it has excellent fetching quality of code database. i could easily integrate it with github. the speed is good as compared to other products in the market. it also edits and autocompletes your code. q: what do you dislike about the product? a: it is an expensive software also i have used it i could feel a little bugs in customer support system.” [source](https://www.g2.com/products/augment-code/reviews/augment-code-review-13433460)

### Warp

- Praise, 2026-09-17, r/codex (Reddit): “there’s no way they can continue offering these plans with all of us hammering on their servers basically all day every single day. had a sense this would come to an end. have basically been at it nonstop to get product advanced as far as possible. fyi….the grok $300 plan is incredible……soooo much programming. shit tons. way more than even codex 20x or claude max. i use all 3. also, for grok….100% use it with warp. free to use warp…byom…bring your own model. gives me the context between tasks like we have on desktop programs for claude or codex.” [source](https://www.reddit.com/r/codex/comments/1whlif0/support_just_told_me_they_arent_renewing_people/paad0jv/)
- Praise, 2026-09-02, @warpdotdev (X): “@warpdotdev i kept shipping agents that never improved their own skills. skill-loop.md now: one skill that rewrites the skill folder after each miss. the meta skill is the real upgrade.” [source](https://twitter.com/1888453273679740928/status/2094942515992412562)
- Complaint, 2026-09-15, @warpdotdev (X): “@warpdotdev after opening multiple tabs, losing context is truly a nightmare, right?” [source](https://twitter.com/1212651734209724416/status/2099953528529985544)
- Complaint, 2026-09-10, Trustpilot (Trustpilot): “**the easiest way to lose money** my experience with warp has been extremely frustrating: errors, errors, and more errors. warp can handle simple tasks reasonably well, but when you start using the agent for larger problems or real projects, it can become an extremely expensive experience. i've spent hours working on a project and consuming credits, getting close to solving a problem, only for the agent to suddenly fail because the conversation/context became too large. i actually reported one of these problems on warp's github. the agent sent a request exceeding the vertex ai context limit of 1,048,576 tokens and the whole request failed. their own automated triage later concluded that the” [source](https://www.trustpilot.com/reviews/6aa3062b00691db98a32c10b)

### Grok Build

- Praise, 2026-09-24, r/google_antigravity (Reddit): “i primarily work in swift so it's a mixed bag, especially during any transition. we're currently moving to ios 27, which some llms assert still doesn't even exist yet, so trying to do anything "new" is still best done by hand. i have started playing around with grok build for small personal projects i don't have time to work on but really want to tinker with and it's surprisingly good for slightly-beyond-prototype work. it has a strong grasp of design principles but it's very "dumb" when it comes to anticipating issues. it does exactly what you ask and nothing more.” [source](https://www.reddit.com/r/google_antigravity/comments/1wp6mpo/poll_how_do_you_code_in_late_2026/pbsymyj/)
- Praise, 2026-09-22, r/ClaudeCode (Reddit): “usage burns fast on codex too but at least is still competent on sol 5.6. claude on opus 5 has gotten nearly unusable. i will say that my first impressions of grok build are good. its surprisingly much better than claude at actually following rules and not drifting into pure insanity.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wmbt29/two_5x_sub_or_one_20x_sub/pba474l/)
- Complaint, 2026-09-08, r/codex (Reddit): “the way i have learned to see it after 3500 hours of experience with vibe coding is that its best to treat all models, whether it's codex, claude code, grok build etc, like a dumb employee that can work hard and comes up with something good every now and then, but you need to manage this employee a lot and if you don't steer it, it will start creating a lot of overhead, over-engineer things that aren't relevant and it will lose track of the goals you've set it out to do. and also his memory isn't very good; every few hours he forgets a bunch of things and is prone to making the same mistakes over and over, to the point that you can be working in a loop for weeks, or even months, because one” [source](https://www.reddit.com/r/codex/comments/1wamtly/i_dont_find_building_with_codex_or_any_ai_easy_at/p8leyzb/)
