# Direct in-prompt instructions and caps are followed (`context.instruction_following`)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/criterion/context.instruction_following

Area: [Instructing and context](https://feedbackbench.com/criteria/context.md)

**Definition.** Whether the agent executes explicit instructions, prohibitions and budgets given in the prompt, such as iteration caps and token caps.

**Boundary.** Not this: see [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md) for rules files. Not this: see [Caves to or argues with the user's judgement](https://feedbackbench.com/criteria/work.sycophancy_pushback.md) for evaluating the user's claims or opinions. Not this: see [Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md) for unrequested extras.

Rated author-weeks, all agents: 1331. Complaint share: 73%.

## The brief

Written by Claude Opus 5.5 from 76 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**Every agent hears the rule. None reliably keeps it.**

TL;DR:

- Complaints outnumber praise for every major agent. No one clearly beats peers on following instructions.
- The top failure is explicit prohibitions treated as suggestions, or forgotten a few turns later.
- Users credit the model inside the harness more than the harness for how well instructions stick.

In plain terms: Expect to repeat yourself. Agents often weight the latest message, ignore a stated prohibition, or overrun effort and token settings. Tight, constraint-heavy briefs and fresh threads help. Switching model versions is the fix users mention most.

### How it breaks

- **Explicit prohibitions get ignored** ([Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md)). The sharpest complaint is a direct do-not instruction that the agent executes anyway, sometimes repeatedly.
  Users describe telling the agent not to take an action and watching it happen regardless. One user says the agent even admitted the rule while breaking it, and rewriting the instructions did not help. Another reports a single-file scope that the agent ignored by exploring dozens of files. A Devin user says a prohibition gets read as an instruction to do the thing. Respecting prohibitions and scope limits is a recurring request.
  Evidence:
  - Complaint, GitHub Copilot, r/webdev, 2026-09-04: “meanwhile, despite multiple instructions not to, copilot requests to deploy my project regularly…” [source](https://www.reddit.com/r/webdev/comments/1w74hvf/where_should_i_host/p7s2lyd/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-10: “i noticed that sol started becoming a wasteful looping idiot last week even before astra was out. i kept calling it out after and having it review the instructions and it's behavior and it kept saying "you told me explicitly not to do x, but i did x". trying to update the instructions didn't help” [source](https://www.reddit.com/r/codex/comments/1wchlld/the_limits_got_nerfed_hard/p8zzpzw/)
  - Complaint, Cursor, r/cursor, 2026-09-17: “doesn't matter, i tell the agent edit this in this file and said change must be done only in that file and it still explores everything around 40+ files.” [source](https://www.reddit.com/r/cursor/comments/1wh5qqi/grok_46_has_been_lobotomized/pacqq3w/)
  - Complaint, Devin, @cognition, 2026-09-18: “@cognition and that wraps that. the swe2 model has 0 adherence to prompt safety. tell it to do something a specific way, if it fails, it doesn't flag it and tries to workaround instead.. tell it not to do something, it will take that as instruction to do it. its just shit tier.” [source](https://twitter.com/1917224549604605953/status/2101091262363492357)

- **Rules stated three times, still ignored** ([Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md), [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md)). Repeating a rule in the rules file, the goal and the prompt does not guarantee compliance, and adherence decays over a session.
  Posts show users stacking the same instruction across every layer they control and still losing it. Others say the agent gives all its weight to the newest prompt and forgets earlier ones. Cursor users report adding new rules again and again to block the same mistakes. The workaround users share is a fresh thread with the rule pinned, which sticks better than arguing mid-chat.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-27: “i noticed that 5.6-sol does similar mistakes. i put a clear rule in [agents.md](<strict_link>) and the /goal and in the initial prompt at the start of the goal - it ignored all of them.” [source](https://www.reddit.com/r/codex/comments/1wrftcs/gpt6_sol_is_massive_downgrade/pccbjm1/)
  - Complaint, Google Antigravity, @antigravity, 2026-09-08: “@rodydavis @antigravity 1) it's doing the unwanted changes and breaking other working features, i initially thought it might be project specific but noticed it's behaving same in other projects too. and forgets instructions in previous messeges kinda feels it's giving all weight to latest prompt only.” [source](https://twitter.com/1954499805050519552/status/2097473929686479300)
  - Complaint, Google Antigravity, @antigravity, 2026-09-05: “@nlycskn @antigravity @thtbee_ works fast, sometimes produces working code, ignores the skill instructions, dies in the death loop. <strict_link>” [source](https://twitter.com/2009822403170578434/status/2096190394254127519)
  - Complaint, Cursor, r/cursor, 2026-09-12: “and they emailed users at the time. auto just picks the model and you pay that models price. it's not flat anymore. so the cheapest most efficient model at the moment is composer 2.5. personally i plan with grok and build with composer. but it's getting more and more shit lately, and i keep having to add rules to prevent it from being stupid in the same ways, again and again.” [source](https://www.reddit.com/r/cursor/comments/1wdb3cx/pro_users_on_auto_mode_20mo_how_many_million/p9ekukb/)

- **Effort and budget settings overridden** ([Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md)). Configured model, effort and token preferences get bypassed, and users pay for it in burned session limits.
  A Pi user set models and effort per task and says every agent still launched at max effort, draining a usage window in an hour. A Codex user says one session's tokens went to half a requested change plus edits they never asked for. The counterexample matters. Another Codex user asked it to save tokens and skip progress previews, and says it complied. Caps work sometimes, which makes the misses harder to plan around.
  Evidence:
  - Complaint, Pi, r/PiCodingAgent, 2026-09-07: “philosophy or not, this thing is awful. literally wasted my 5 hour limit in an hour on cc max 5. using claude code, i never hit the limit. i did everything imaginable to get it toned down to no avail. the craziest thing is that i set my models and effort for each task. it didn’t give a damn and dispatched all the agents at max effort. yeah, no thanks.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1uaydrr/what_do_you_guys_think_about_oh_my_pi/p89v3ym/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-05: “i am testing astra on improving web frontend design for a card game. i liked design of cards that was made, but i did not like frontend design. terra i started with was 100% falure. i spend 3 extra prompts to undo the stuff i did not instuct to improve on. some stuff are subtile, but for example i said it should have look and feel and colors be motivated by slovenian folklore and culture and made it wood details using some aisan wood and clanker updated just half of design. it spend full 5h sesson tokens on minimal plan i am on, just to change half color scheme and now i need to wait for another 5h to update remainder and wood design. there is progress, but is slow and expensive.” [source](https://www.reddit.com/r/codex/comments/1w7on0j/astra_is_absolutely_incredible/p7xnefl/)
  - Praise, OpenAI Codex, r/codex, 2026-09-07: “no.. i was just using default of whatever codex deems necessary. i told it shouldnt waste extra effort on giving me some progress previews or anything and that it should be mindfull to save tokens. which it is based on what it replied to me.” [source](https://www.reddit.com/r/codex/comments/1w9tpe5/codex_has_fundamental_flaw/p8dajeo/)

- **Brevity and style rules don't hold** ([Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md), [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md)). Instructions about tone, conciseness and vocabulary are the ones models most often drop.
  Users build terse personas and plain-language rules, then watch the model refuse to use them. One user says a standing be-concise rule rarely works for anyone, and only emphatic phrasing got through. Others pin rules against invented jargon and force answer-first output in rules files. Style drift is quieter than a broken prohibition, but it is constant friction.
  Evidence:
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-09: “said this many times. rei ayanami persona. i have many small variations of it, but it is blunt and direct, no emotion, reduces my response output a ton. sometimes all i get back is: done, what’s next. my biggest pain point is when models like opus refuse to use it.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wbhs2i/give_me_your_trips_and_tricks_to_reduce_verbose/p8pxj3q/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-04: “and yet, if we're talking about raw efficiency, you wrote a design spec. i wrote two paragraphs. as a standing rule, what i did is far more efficient and less time-consuming. additionally, sometimes you want to work with claude on the design. when you're starting a ui from scratch, you don't have much to go off of on design. you have an idea of what you want to serve to the user, but little to go off of on frontend aesthetics and layout. one of the only standing rules i really needed was this. lastly: if "be concise" alone worked, i would have just added that. there is a reason why there's new posts about the model's verbosity daily. a standing "be concise" rule in the system prompt has rarely worked for anyone—tons of posts corroborate that, not hard to find. the peculiar thing i am pointing to is that "yelling" did indeed work.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w6prth/pasting_this_verbatim_in_claudemd_as_a_standing/p7p5gxm/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-08: “yeah its been dumping random jargon mid-task for me too. i started pinning a rule like "if you invent a term, define it in one line or dont use it" plus "prefer plain words over product nicknames". also helps to force answer-first in claude.md so it cant hide behind buzzwords. if it keeps doing it mid-chat, fresh thread with that rule sticks better than arguing.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wai39s/anyone_else_have_claude_code_drop_unexplained/p8ic6fz/)

- **Tight constraints are what work** ([Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md)). The praise clusters around briefs full of explicit rails, such as do-not-touch lists and numbered steps, not open-ended asks.
  Users who report good adherence front-load the constraints. One Cursor user lists off-limits areas before a refactor and says the agent moves faster inside those rails. Copilot praise credits a ticket brief packed with constraints. A Pi user runs many sequential steps from the system prompt, which they say outranks skills and user prompts. Claude users mark old plans explicitly so the agent can detect stale assumptions.
  Evidence:
  - Praise, Cursor, @cursor_ai, 2026-09-25: “small agent habit that saved me hours this week in @cursor_ai: before a big refactor, i dump a one-screen "do not touch" list in the prompt (auth, billing, migrations). the agent goes faster when it knows the rails. #cursor #ai #softwareengineering #buildinpublic” [source](https://twitter.com/262960825/status/2103530958188392506)
  - Praise, GitHub Copilot, @GitHubCopilot, 2026-09-23: “@arya_at1 @githubcopilot @github this is such a clean example of how to actually use agents well. the brief was tight and full of constraints instead of open-ended.” [source](https://twitter.com/1583159673728864256/status/2102810895340711975)
  - Praise, Pi, r/PiCodingAgent, 2026-09-22: “i just use markdown and it’s working fine with 15 sequential steps inside my system prompt. system prompts have a higher priority than skills or user prompts so the model follows it better. i’m certain this recommendation breaks down eventually with scale but… might be fine for light factories.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wne4zn/anybody_using_a_workflow_engine_to_automate_their/pbf2fr8/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-13: “for pivots, i agree with the approach of using markdown files to represent the current design / implementation plan. by having an "old plan", i can tell claude to "generate a new plan because i changed my mind, or we have new information." i find that claude performs better when you explicitly generate a new plan. it thrives on clear context, and if you change your plan mid-way, it's got context with both the old and the new assumptions. i'll even tell claude "if you see the term 'xyz', this is a sign you're working on the old strategy. we have changed our direction - do 'abc' instead." this has the advantage of patternmatching things to avoid.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wffg50/how_well_does_claude_code_actually_handle_a/p9lu7iw/)

### Who stands out

- **OpenAI Codex (mixed)**. Codex earns the most explicit praise for literal compliance, yet it also draws heavy complaints about ignored rules on certain model versions.
  Fans say it does exactly what was asked, and one user says Codex follows a standing instruction where Claude Code ignores it. The complaints split by model. Users call one newer version cheaper but prone to ignoring instructions and taking shortcuts, and keep the older one for first-pass reliability. Codex users file the most requests for reliable adherence to explicit prompts.
  Evidence:
  - Praise, OpenAI Codex, r/codex, 2026-09-27: “it’s thorough in doing exactly as requested and almost anything that’s logically connected to it for me (basically saying if something is abstract for most humans, it’ll also be for it)” [source](https://www.reddit.com/r/codex/comments/1wqtd1g/openai_gpu_are_really_cooling_down_theo_just/pcbqjcf/)
  - Praise, OpenAI Codex, r/PiCodingAgent, 2026-09-13: “it's exposed as not working for claude code, since claude code ignores it in most cases. it's a bit different for codex (which follows instructions to use to all the time) or pi (which has an extension that applies it automatically). although i might audit my session logs for pi and codex later and check how well it works.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wfh044/my_pi_agent_harness_setup_auto_invocable_skills/p9nj5ox/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-24: “luna 6 feels like more of a sidegrade than a straight upgrade to me. it’s cheaper/easier to spam and can be really good on small, well-scoped tasks but i’ve seen enough cases where it ignores instructions, takes weird shortcuts, or needs another pass that i wouldn’t automatically replace 5.6 with it for everything. 5.6 feels a bit more predictable. luna 6 feels more like throw a ton of work at it because the economics are good, then escalate the failures. so basically: 6 for volume, 5.6 when i care more about first-pass reliability. if they can fix the instruction-following/reliability without ruining the current usage economics, luna 6 would be a monster.” [source](https://www.reddit.com/r/codex/comments/1wox48m/luna_6_vs_luna_56/pbqiuiu/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-10: “i noticed that sol started becoming a wasteful looping idiot last week even before astra was out. i kept calling it out after and having it review the instructions and it's behavior and it kept saying "you told me explicitly not to do x, but i did x". trying to update the instructions didn't help” [source](https://www.reddit.com/r/codex/comments/1wchlld/the_limits_got_nerfed_hard/p8zzpzw/)

- **Claude Code (mixed)**. Claude Code posts swing with model releases, with praise for closer adherence on one version and distrust and churn on another.
  One user testing a newer Opus says it follows instructions more closely. Others report misread directions and say they verify every output with a second model. One user dropped the subscription over unpredictable behaviour. Another says Grok Build follows rules better and drifts less. Claude Code is the only agent where users ask for a way to update or invalidate outdated rules.
  Evidence:
  - Praise, Claude Code, r/ClaudeCode, 2026-09-23: “i'm testing opus 5.5 right now, and in my experience it performs the tasks better than opus 5. it also seems to follow the instructions more closely.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wnk74q/whats_the_point_of_fable_if_opus_55_is_stronger/pbnbb1c/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-12: “i now verify every output it produces with another ai. i cannot trust any thing it “one-shots”. it makes statistically incorrrect assumptions constantly, semantic errors, over claiming something is worse than it is (or better than it is), misinterpreting me and not following directions, ignoring data it doesn’t want to see (because that means more work and claude hates to work). as these models become more advanced they’re actually becoming way less trust worthy — because they’re still making the same stupid mistakes that it has made since day one” [source](https://www.reddit.com/r/ClaudeCode/comments/1wdtrsa/im_done_with_opus_5/p9dgr2q/)
  - Praise, Cursor, r/codex, 2026-08-31: “i ran: 1x cursor 1x claude and 1x codex monthly $20 for nearly a year because i like having options and staying at the frontier. i dropped my claude sub last month for specifically this reason: claude behaves unpredictably, to the point i don’t trust it at all. now i run 2x codex and 1x cursor and really never have to worry about rogue agents. the gpt lineup and composer 2.5 on cursor are great little rule followers.” [source](https://www.reddit.com/r/codex/comments/1w3ltfq/told_claude_dont_push_yet_let_me_test_it_first_it/p71k9hj/)
  - Praise, Grok Build, r/ClaudeCode, 2026-09-22: “usage burns fast on codex too but at least is still competent on sol 5.6. claude on opus 5 has gotten nearly unusable. i will say that my first impressions of grok build are good. its surprisingly much better than claude at actually following rules and not drifting into pure insanity.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wmbt29/two_5x_sub_or_one_20x_sub/pba474l/)

- **Cursor (mixed)**. Cursor users praise Composer as a literal rule-follower and blame other hosted models for scope creep and ignored directions.
  Users describe Composer as a tool that does what you say, while another model in the same picker adds unrequested features and breaks things after apologising. Complaints still land on Cursor. One user says Composer is degrading and needs more rules each week. Another moved a team elsewhere for better direction-following. A scoped single-file edit that touched 40+ files is the starkest miss.
  Evidence:
  - Praise, Cursor, @cursor_ai, 2026-09-15: “i have to say it it's the first month that i run out of tokens on my @cursor_ai plan and i went back to composer2.5 to save a bit on costs. this model is special. sure grok feels smarter. but composer just does what you tell it to do. like a tool. llms are meant to be a tool, right ?!” [source](https://twitter.com/1670440088537276418/status/2099827297926701360)
  - Praise, Cursor, r/cursor, 2026-09-27: “composer does exactly what you ask it to do even if it takes a few prompts to finish grok will do it all and add 10 things i didn't ask for so i tell it i didn't ask for those things and it says 'you're right i'm so sorry' then it adds 2 other things i didn't want or it will change something that breaks everything. so you ask it to fix it. oh, so sorry, here's 2 more things you didn't ask for.” [source](https://www.reddit.com/r/cursor/comments/1wrktgc/am_i_the_only_one_who_thinks_grok_47_is_actually/pcdlxql/)
  - Complaint, Cursor, r/cursor, 2026-09-17: “doesn't matter, i tell the agent edit this in this file and said change must be done only in that file and it still explores everything around 40+ files.” [source](https://www.reddit.com/r/cursor/comments/1wh5qqi/grok_46_has_been_lobotomized/pacqq3w/)
  - Complaint, Cursor, r/cursor, 2026-09-14: “i canceled my personal cursor today. i’ve recently moved my team to claude. will reevaluate again next year but claude is just better at following directions than grok and composer and the new direction of cursor has been less than ideal.” [source](https://www.reddit.com/r/cursor/comments/1wfmyf0/composer_3_needs_to_happen/p9odtl4/)

- **Google Antigravity (mixed)**. Antigravity draws some of the bluntest ignore-the-rules complaints, but users report a clear step up with the newer Gemini model.
  Complaints describe skill instructions ignored, earlier messages forgotten, and unwanted changes that break working features. Recent posts shift tone. Users say the newer model follows directions with less back and forth, checks the blast radius of its actions, and pays off with custom skills. One user accepts higher tool-call counts as the price of better adherence to their standards.
  Evidence:
  - Praise, Google Antigravity, @antigravity, 2026-09-05: “@nlycskn @antigravity @thtbee_ it’s been a great daily driver for me following directions much better than 3.7 and with less back and forth needed.” [source](https://twitter.com/1856554970536902656/status/2096184100386123780)
  - Complaint, Google Antigravity, @antigravity, 2026-09-08: “@rodydavis @antigravity 1) it's doing the unwanted changes and breaking other working features, i initially thought it might be project specific but noticed it's behaving same in other projects too. and forgets instructions in previous messeges kinda feels it's giving all weight to latest prompt only.” [source](https://twitter.com/1954499805050519552/status/2097473929686479300)
  - Praise, Google Antigravity, r/google_antigravity, 2026-09-17: “that's the same reason i am happy about antigravity and 3.8 on high now - because it stopped breaking stuff and started to understand context in which it makes differences. especially with custom skills it is now much more useful. it follows commands much better and checks for blast radius of it's actions, still not intelligent for big work, but with right skills, first time antigravity is really usefull. and now, especially with this much of usage i actually find myself doing more work with it on pro level than with cc on pro level. and astra got a nerf to it's capabilities recently so even 100+€ 5x pro level on codex is not much more better, than antigravity. it is better don't get me wrong, just not as much as the day astra released, they nerfed it really fast. so yeah, what antigravity and 3.8 does now on high - is what myself and some others actually wanted all along.” [source](https://www.reddit.com/r/google_antigravity/comments/1wi8z7u/did_they_change_the_model_or_what/pabbxud/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-19: “agy be like: fk it too long to follow ignore this rules i will do whatever i want” [source](https://www.reddit.com/r/google_antigravity/comments/1wkzesn/my_rules_for_ag/paupdlt/)

### Fine print

- All agents with enough posts land in the typical range. Differences between leaders are directional, not decisive.
- Most agents have too few posts here to judge on their own.
- Users often blame a specific model rather than the harness, so agent-level results mix both.

## Top requests

What users ask to add or change, most asked first. 102 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Author-weeks | Posts | Agents (author-weeks) |
|---|---|---|---|---|
| 1 | Reliable adherence to explicit prompt instructions | 33 | 33 | OpenAI Codex 14, Claude Code 13, Google Antigravity 4, Cursor 1, OpenCode 1 |
| 2 | Persistent adherence to rules and custom instructions | 17 | 17 | Claude Code 7, Google Antigravity 4, OpenAI Codex 4, OpenCode 2 |
| 3 | Respect explicit prohibitions and scope limits | 16 | 17 | OpenAI Codex 6, Claude Code 5, Google Antigravity 2, Pi 2, OpenCode 1 |
| 4 | Harness-enforced rules instead of prose instructions | 4 | 4 | OpenAI Codex 2, Claude Code 1, OpenCode 1 |
| 5 | Mechanism to update or invalidate outdated rules | 3 | 3 | Claude Code 3 |
| 6 | Literal interpretation without inferring unstated intent | 2 | 3 | Claude Code 1, OpenAI Codex 1 |
| 7 | Honor instructions to work until completion | 2 | 2 | OpenAI Codex 2 |
| 8 | Respect configured iteration and turn caps | 2 | 2 | Claude Code 1, OpenAI Codex 1 |
| 9 | User instructions override system polling behavior | 2 | 2 | Google Antigravity 1, OpenAI Codex 1 |

### 1. Reliable adherence to explicit prompt instructions

- Claude Code, 2026-09-24, @ClaudeDevs (X): “@claudedevs shit model. your stupid product should do as its told.” [source](https://twitter.com/2091957404376129536/status/2103253534376436044)
- Cursor, 2026-09-24, @cursor_ai (X): “@e_viki_ @cursor_ai to be clear, i didn't move to grok intentionally, cursor moved me to grok involuntarily after the spacex purchase. with that said, code quality is actually pretty good. plan following and just following directions in general seems to be its weekest point.” [source](https://twitter.com/14311446/status/2102969871437033755)
- OpenAI Codex, 2026-09-23, r/codex (Reddit): “yup, went back to 5.6 sol and luna no more headaches. even giving gpt 6 sol exact steps to take, just goes ahead to do whatever it wants.” [source](https://www.reddit.com/r/codex/comments/1wo6ntb/anyone_noticed_a_sudden_increase_in_6_sols_token/pbkhr6m/)

### 2. Persistent adherence to rules and custom instructions

- Claude Code, 2026-09-23, r/ClaudeCode (Reddit): “5.5 has been pretty solid for me so far. 4.8 was a fucking nightmare. before every message in the chat i would have to copy and paste “short responses only” even hard coding it into the .md file it ignored it. god i hated it, i started using chatgpt again and really like it. may start using more 5.5 since it’s so solid” [source](https://www.reddit.com/r/ClaudeCode/comments/1wnwi9s/chad_55/pbiu5e4/)
- Claude Code, 2026-09-23, r/ClaudeCode (Reddit): “it's not fixed before i can configure it.l and make it follow my rules of communication always.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wnf1o8/apparently_they_fixed_the_talking_slop_in_opus_55/pbie4lc/)
- Google Antigravity, 2026-09-15, r/google_antigravity (Reddit): “<strict_link> <strict_link> these kinds of sycophant hallucinations, having to babysit the model after just a few turns is such a slap to the face to all the bench-maxing fuks that this gemini-flash-3.8 model is boasting; such an incomplete product. i have literately put everything to [agents.md](<strict_link>); having skill to that specifically, and even put rules of tdd under agents/rules/\*\*. i meant, google please !” [source](https://www.reddit.com/r/google_antigravity/comments/1wgvb1c/i_literately_dont_know_what_the_executives_at/)

### 3. Respect explicit prohibitions and scope limits

- Claude Code, 2026-09-27, r/ClaudeCode (Reddit): “every ”do not” phrase is bad for 5.x models. your insturctuons are bad. they dont work properly” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrib98/do_you_feel_that_claude_code_unit_tests_are/pcctfya/)
- OpenCode, 2026-09-26, r/opencode (Reddit): “for me space bunny can't stop adding chinese, russian, korean characters in the chat, i've already added rules, and it still does it, sometimes it's so stupid that it feels like i'm running a local model” [source](https://www.reddit.com/r/opencode/comments/1wqcqmi/space_bunny_randomly_had_a_stroke/pc9mf67/)
- OpenAI Codex, 2026-09-24, r/codex (Reddit): “you're just a cope monster. you can be so gd specific, but it will still interpret things, it just does. you shouldnt have to list out what only means to these things, if you say do "only x, change nothing else" that is explicit. and it will mess that up.” [source](https://www.reddit.com/r/codex/comments/1wotvyv/gpt_6_sol_is_an_idiot/pbv2ki8/)

### 4. Harness-enforced rules instead of prose instructions

- OpenAI Codex, 2026-09-21, r/codex (Reddit): “thank you, this adds something of benefit. in this case i was using go. i find they still forget or disregard instructions after some time though. i would like to have another adversary to enforce these rules as it works.” [source](https://www.reddit.com/r/codex/comments/1wlvf0j/state_of_agentic_coding/pb2meln/)
- OpenCode, 2026-09-18, r/opencode (Reddit): “i have similar issue with both muse and free deepseek they try to access outside of project folder no matter how much i tell them not to do it, like come on” [source](https://www.reddit.com/r/opencode/comments/1wjjqvr/muse_13_free_formatted_my_drive/paji8b2/)
- OpenAI Codex, 2026-09-09, r/codex (Reddit): “exactly. i’d much rather have more harness-level controls for this. right now too much agent behavior has to be enforced indirectly through agents.md and orchestration docs, basically persuading the model to behave a certain way.” [source](https://www.reddit.com/r/codex/comments/1wa9c9d/i_investigated_why_gpt6_astra_burns_quota_so_fast/p8se4t6/)

### 5. Mechanism to update or invalidate outdated rules

- Claude Code, 2026-09-23, r/ClaudeCode (Reddit): “the reason behind the rule needs to survive alongside it. “don't restart this process” could mean “we haven't implemented recovery yet” or “this process owns every live session.” those need very different evidence before you remove the restriction. i'd want the agent to identify the assumption that changed and propose updating the rule. letting it silently decide a rule is obsolete seems like another way to lose the original constraint imo” [source](https://www.reddit.com/r/ClaudeCode/comments/1wogfpt/the_problem_with_claude_and_rules/pbnnxil/)
- Claude Code, 2026-09-11, @ClaudeDevs (X): “@claudedevs the agent would like to revisit the rule about not writing tests” [source](https://twitter.com/276600668/status/2098520294092894628)
- Claude Code, 2026-09-23, r/ClaudeCode (Reddit): “the deeper problem is that models often treat previously true constraints as timeless invariants, when in a real project many “rules” are really just snapshots of a particular state. a rule might have been perfectly correct yesterday, then the architecture changes. the model keeps treating the old statement as authoritative, it starts reasoning inside a world that no longer exists. worse, it may reject the correct solution because it violates an” [source](https://www.reddit.com/r/ClaudeCode/comments/1wogfpt/the_problem_with_claude_and_rules/)

### 6. Literal interpretation without inferring unstated intent

- OpenAI Codex, 2026-09-12, r/codex (Reddit): “the model deciding to tackle the emotional part of the prompt. one of the good thing of ai is that you can give it instructions without having to manage your emotions and your colleagues emotions. you can still get pissed when something takes 10x more times than it should, but it has no impact on the work. if ai is starting to take that into account, and this actually prevents it from doing its job… then it’s a clear regression. but i actually lo” [source](https://www.reddit.com/r/codex/comments/1wciwc1/gpt6_astra_burns_quota_4_times_faster_than_gpt56/p9eggtp/)
- Claude Code, 2026-09-11, r/ClaudeCode (Reddit): “same here. even after i explicitly tell it to ask my questions, in the claude.md, in the prompt, and in a message hook… i’m slowly coming to the conclusion that claude code is a really bad harness and is used to farm training materials, and that the “50x” usage just makes it even with what a good, effective harness would do. for example, let’s say i request a simple refactor of a function. i want the harness to just…. do it. claude code will som” [source](https://www.reddit.com/r/ClaudeCode/comments/1wd2738/claude_dynamic_workflow_is_very_cool/p93dgke/)

### 7. Honor instructions to work until completion

- OpenAI Codex, 2026-09-20, X search: OpenAI Codex, Codex CLI, Codex app (X): “i've been using codex extensively on a very large ai architecture, and while it's exceptionally strong at audits, security analysis, code inspection, and finding local implementation defects, i've repeatedly encountered weaknesses in long-running, project-wide work. one of the biggest problems is instruction persistence across checkpoints. i've explicitly instructed codex not to stop at checkpoints and to continue working autonomously. it acknowl” [source](https://twitter.com/1944119286617841665/status/2101755727102586956)
- OpenAI Codex, 2026-09-07, r/codex (Reddit): “astra on low simply refuses to do what i want. i gave it a list with stuff to implement and told it to work till completion. what does it do after 1 minute? stop, because apparantly it has written half of a file so it made good progress. i tell it to continue and it stops again. i get angry and demand execution and it says: "ok i will work until i see a more substantial progress". i interrupt and tell it no, work until completion. it agrees and s” [source](https://www.reddit.com/r/codex/comments/1w9erx3/usage_tip_gpt6_astra_on_low_performs_better_than/p8aoda3/)

### 8. Respect configured iteration and turn caps

- Claude Code, 2026-09-23, r/ClaudeCode (Reddit): “version: 2.1.280 model: opus 5.5 example: <strict_link> i've never had this issue before. a turn limit is completely pointless, subagents just like the main agent should be able to run infinitely. i haven't added any config that should be causing this behavior.” [source](https://www.reddit.com/r/ClaudeCode/comments/1woaxj7/what_the_hell_is_a_turn_limit_200turn_limit_and/)
- OpenAI Codex, 2026-09-15, r/codex (Reddit): “just the other day i had it doing cycles where it would do a test then a fix then test again, i gave it a cap of three tests and it worked perfect. so i wanted to get more work done overnight i set up the same thing gave it a cap of five, well wouldn't you know it 8 hours later when i got up he's still working, blew through damn near all my usage. i was so pissed” [source](https://www.reddit.com/r/codex/comments/1wgh9zr/astra_misunderstood_me_pro_20x/p9yf6ui/)

### 9. User instructions override system polling behavior

- Google Antigravity, 2026-09-06, @antigravity (X): “@antigravity add this to your system prompt “always run it to completion. there’s no need for a timer, just keep it running and report the results". agent doesn’t need to check 5, 10, 20 secs, etc. this way, you don’t perform periodic checks. these reminders consume a lot of tokens” [source](https://twitter.com/1690943437946941440/status/2096434312229085477)
- OpenAI Codex, 2026-09-20, r/codex (Reddit): “nice analysis, fully agree, good luck with that fix though. i have tried putting such instructions in my [agents.md](http://agents.md) and the model would happily ignore them since the system prompt takes higher priority, this can also cause issues with stuck processes or slow commands where the agent ran an unoptimized script/command that can waste a lot of our time, if it polls frequently it can catch that mistake and correct itself. all of the” [source](https://www.reddit.com/r/codex/comments/1wlcy5q/this_will_save_your_usage/paxni4b/)

## Every agent

| Agent | Overall rank | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|---|
| [OpenAI Codex](https://feedbackbench.com/agents/codex.md) | 2 | Typical | 0.512 | 0.485–0.537 | 518 | 144 | 374 |
| [OpenCode](https://feedbackbench.com/agents/opencode.md) | 3 | Typical | 0.504 | 0.470–0.541 | 76 | 22 | 54 |
| [Claude Code](https://feedbackbench.com/agents/claude-code.md) | 1 | Typical | 0.491 | 0.463–0.519 | 516 | 133 | 383 |
| [Google Antigravity](https://feedbackbench.com/agents/antigravity.md) | =5 | Typical | 0.484 | 0.449–0.518 | 112 | 27 | 85 |
| [Cursor](https://feedbackbench.com/agents/cursor.md) | 4 | Typical | 0.479 | 0.451–0.505 | 50 | 10 | 40 |
| [Pi](https://feedbackbench.com/agents/pi.md) | 7 | Too few posts | – | – | 17 | 4 | 13 |
| [GitHub Copilot](https://feedbackbench.com/agents/copilot.md) | =8 | Too few posts | – | – | 16 | 5 | 11 |
| [Devin](https://feedbackbench.com/agents/devin.md) | =5 | Too few posts | – | – | 9 | 3 | 6 |
| [Amp](https://feedbackbench.com/agents/amp.md) | =11 | Too few posts | – | – | 6 | 4 | 2 |
| [Cline](https://feedbackbench.com/agents/cline.md) | =8 | Too few posts | – | – | 3 | 1 | 2 |
| [Factory](https://feedbackbench.com/agents/factory.md) | =11 | Too few posts | – | – | 3 | 2 | 1 |
| [Grok Build](https://feedbackbench.com/agents/grok-build.md) | 16 | Too few posts | – | – | 2 | 2 | 0 |
| [Zed](https://feedbackbench.com/agents/zed.md) | =8 | Too few posts | – | – | 1 | 1 | 0 |
| [Kiro](https://feedbackbench.com/agents/kiro.md) | 13 | Too few posts | – | – | 1 | 0 | 1 |
| [Conductor](https://feedbackbench.com/agents/conductor.md) | 14 | Too few posts | – | – | 1 | 0 | 1 |
| [Warp](https://feedbackbench.com/agents/warp.md) | 15 | Too few posts | – | – | 0 | 0 | 0 |
| [Augment Code](https://feedbackbench.com/agents/augment.md) | 17 | Too few posts | – | – | 0 | 0 | 0 |

## Posts

Receipts rule: The 5 most recent praise and complaint posts per area (first 700 characters) and 3 per criterion (first 450 characters).

### OpenAI Codex

- Praise, 2026-09-27, r/codex (Reddit): “it’s thorough in doing exactly as requested and almost anything that’s logically connected to it for me (basically saying if something is abstract for most humans, it’ll also be for it)” [source](https://www.reddit.com/r/codex/comments/1wqtd1g/openai_gpu_are_really_cooling_down_theo_just/pcbqjcf/)
- Praise, 2026-09-26, r/codex (Reddit): “the same way i discovered the issue in the first place: i was monitoring what astra was actually doing. after adding the instruction, i kept monitoring subsequent runs and could see it checking the file length and reading the remaining sections before starting. i'm not assuming it worked just because i told it to. i'm saying it worked consistently in the runs i've observed since making the change.” [source](https://www.reddit.com/r/codex/comments/1wqdg6v/warning_astra_6_can_read_instructions_partially/pc3d3pf/)
- Praise, 2026-09-26, r/codex (Reddit): “yes. if you are smart about tokens it’s more usage anyway (last i checked, so who knows now). plus codex is ass at some things claude is nailing right now. claude still gave me a stupidly ugly dash the other day for a long transfer i wanted to monitor at a glance. i had 15 hours to blow so i pointed astra light at the dash and told it to “stop making my eyes bleed and fix claude’s css choices.” it chose to loosely resemble my home assistant setu” [source](https://www.reddit.com/r/codex/comments/1wmgw7p/codex_usage_and_operation_discussion_last_updated/pc8sy3n/)
- Complaint, 2026-09-27, r/codex (Reddit): “i think they changed the master prompt with astra or the thinking effort to try and reduce token usage. it seems like the same model, but it just doesn't care as much anymore. i remember when i first used it, the thing noted every tiny thing in my [agents.md](http://agents.md) and would even point out errors in it. now it ignores a bunch of my documentation. it's insane because on plus, you'll be at \~100k tokens in the context window and \~50%” [source](https://www.reddit.com/r/codex/comments/1wpvp0i/absolutely_0_doubt_in_my_mind_astra_has_been/pc9uhng/)
- Complaint, 2026-09-27, r/codex (Reddit): “i never used luna 5.6, but luna 6 high has profound mental retardation. just an example: when i asked it to commit and push the changes, this model... tried to do it through the github api for some reason, failed, then told me that it couldn't push because of restrictions. only when i said that there were no restrictions on my side (they were set to "approve for me") did it do what i told it to.” [source](https://www.reddit.com/r/codex/comments/1wr2dda/i_ran_100_terminalbench_21_slots_on_luna_56_and/pca09lm/)
- Complaint, 2026-09-27, r/codex (Reddit): “so its not just me that codex since astra launched has become a potato and a liar? it just cant follow simple tasks and skips majority of the knowledge and critical data i need checked.” [source](https://www.reddit.com/r/codex/comments/1wr1oir/they_are_aware_and_working_on_it_apparently_just/pcawjud/)

### OpenCode

- Praise, 2026-09-25, r/opencode (Reddit): “the intelligence level feels high, but i’d recommend pushing reasoning to max. that’s where the results get noticeably better. what i like is that i don’t really need to spell out every step or force a specific workflow. in a lot of cases, just naming the methodology or framework i want it to follow is enough, and it sticks to those constraints surprisingly well. considering it’s free right now, this is definitely above my expectations. i’m going” [source](https://www.reddit.com/r/opencode/comments/1wprcoo/space_bunny_is_better_than_i_expected_tested/)
- Praise, 2026-09-25, r/opencode (Reddit): “tested space bunny a bit more and it’s actually pretty solid. the intelligence level feels high, but i’d recommend pushing reasoning to max. that’s where the results get noticeably better. what i like is that i don’t really need to spell out every step or force a specific workflow. in a lot of cases, just naming the methodology or framework i want it to follow is enough, and it sticks to those constraints surprisingly well. considering it’s free” [source](https://www.reddit.com/r/opencode/comments/1wprfba/space_bunny_is_better_than_i_expected/)
- Praise, 2026-09-22, r/opencodeCLI (Reddit): “trying it out. it's incredibly fast and the hallucination rate is also really low. follow instructions and adheres to loop. a great release 😁” [source](https://www.reddit.com/r/opencodeCLI/comments/1wmor5s/mimo_26_pro_and_flash_released_already_on/pbbhhyp/)
- Complaint, 2026-09-27, r/opencode (Reddit): “not my experience with it. gpt 6 is extremely bad at following instructions and wastes absurd amounts of time testing” [source](https://www.reddit.com/r/opencode/comments/1wq8vp9/best_free_model_after_deepseek_leave/pcbkxpg/)
- Complaint, 2026-09-27, r/opencode (Reddit): “it's far too trashy to be claude. you literally have to convey everything that's common sense for it not to waste time prodding in wrong directions.” [source](https://www.reddit.com/r/opencode/comments/1wrl1kx/big_pickle_space_bunny_is_claude/pcdgy9c/)
- Complaint, 2026-09-27, @opencode (X): “this is the stealth model space bunny that was on @opencode and @openrouter the model is insanely fast, but it needs clear and specific instructions, otherwise it's too lazy. i tested it here <strict_link> <strict_link>” [source](https://twitter.com/2010272031603068928/status/2104131788578975846)

### Claude Code

- Praise, 2026-09-27, r/ClaudeCode (Reddit): “i think your instructions of fakes over mocks is common instruction for the same mistake *humans* make when deciding between a fake and a mock, which is they don’t understand what mocks are for. mocks are supposed to be about setting expectations on the contract between a unit and its dependencies, *as defined in an api contract.* for example, if your contract specifies that an api on the dependency will be called exactly once with data of a cert” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrib98/do_you_feel_that_claude_code_unit_tests_are/pccyiaf/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “i cloned your repo and have been playing this for hours this weekend. first off, this is the most fun i've had in a video game in decades. i feel like this is everything i loved about old infocom games, but so much better. second, i think i've found a weird glitch or at least a way to warp reality in the game. i can hint about ideas for the game, like "my character wonders if anybody in the tavern knows anything about ..." and then the dm usually” [source](https://www.reddit.com/r/ClaudeCode/comments/1wq5yd8/llm_running_text_rpg_as_dungeon_master/pce25hy/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “some details about how it works - i actually check if model didn't change any files that it is considered just conversation and i do not require model response to be tool call. otherwise in my harness if model tries to write prose instead of tool call, harness returns error (i saw this happening occasionally with muse model. weak models do it all the time.) within plugin opus 5.5 is quite good at making this work - almost never saw it fighting it” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrttht/claude_isnt_allowed_to_write_me_prose/pcfv0wn/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “agent workflow size config was set to small in my settings (<5 agents) + my prompt said verbatim “do not create more than 3 subagents , not a large swarm”………….. my result? ——-> ofc, no other than:🙃 my *entire* weekly pro20x \~\~ *sautéed* ***\~***in front of me on day 1/7 🥲🫠🫠🙃🙃🙃😆😆😆” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqagzu/claude_added_graceful_stopping_point_in_new_update/pcagfdf/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “>i wrote "new\_test\_checks: 11" in that build log line without counting first. that's exactly the mistake we're fighting. counting now: this kind of stuff is just constantly happening despite instructions injected every turn like these: \[verify\] before this action: name all the factors it depends on - each one seen in a tool output? every value in it (count, name, path, time, size, behavior) is read or looked up first. unsure what it touch” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqy8rp/am_i_to_understand_the_nerf_has_begun_or_theres/pcc6osr/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “every ”do not” phrase is bad for 5.x models. your insturctuons are bad. they dont work properly” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrib98/do_you_feel_that_claude_code_unit_tests_are/pcctfya/)

### Google Antigravity

- Praise, 2026-09-27, r/google_antigravity (Reddit): “first of all install ponytail as a skill to anti-gravity. it should start thinking a lot less, spending a lot fewer tokens, and writing a lot less but better code. and then just mention it explicitly: "recently in some of your runs you did this" (you took too many screenshots, checked things that weren't necessary etc.). just mention everything and then it will actually stop doing those things in the follow-up. it will explicitly start saying in” [source](https://www.reddit.com/r/google_antigravity/comments/1wqf2o6/worst_model/pcchcst/)
- Praise, 2026-09-26, r/google_antigravity (Reddit): “i feel gemini "veers off course" less with specific topics. so far ive tested it with gemma4 models, the new agent platform layout and the newer adks (1.0), and i feel i can have more complete building sessions in antigravity without gemini wandering off into an adventure because a mix of words confuses it” [source](https://www.reddit.com/r/google_antigravity/comments/1wqaa6s/my_skill_to_share_gemini_post_cutoff/pc7s2fs/)
- Praise, 2026-09-25, @antigravity (X): “@rodydavis @antigravity thanks! one small question: when using claude models, they usually explain what they find and what they plan to do next as they work, so you always know what’s going on. when debugging, they often identify the root cause before making changes. it’s very clear.” [source](https://twitter.com/1725381581035229184/status/2103633157501411584)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “i built [pigeongraph](<strict_link>), a knowledge graph tool similar to graphify or codegraph, but better (in my opinion). i want to use it across my projects, but antigravity always defaults to using `grep` instead of this tool. how can i ensure antigravity gives pigeongraph first priority and treats `grep` as a secondary fallback? does anyone have an idea or solution for this? *(note: i have already tried setting it as an instruction or rule in” [source](https://www.reddit.com/r/google_antigravity/comments/1wrmn03/i_built_a_knowledge_graph_tool_designed_to/)
- Complaint, 2026-09-24, @antigravity (X): “@alwayspriyesh @soso_fun_yt @antigravity sometimes is crazy. if you're trying to get it to do anything without a clear "read manual and execute" you're just asking for pain.” [source](https://twitter.com/34534062/status/2103056097313796504)
- Complaint, 2026-09-23, r/google_antigravity (Reddit): “i am on ide version 2.5.5; i still don't see latex properly in the chat. gemini uses latex in every convo, even if i tell it not to (in text, in any agents.md, or the global rules); pointing it out, it promises to stop doing it, only to be back at it after a few messages. any trick how to make it work? :/” [source](https://www.reddit.com/r/google_antigravity/comments/1w1ny5k/lack_of_inline_mermaid_chart_latex_math_rendering/pbjjukz/)

### Cursor

- Praise, 2026-09-27, r/cursor (Reddit): “composer does exactly what you ask it to do even if it takes a few prompts to finish grok will do it all and add 10 things i didn't ask for so i tell it i didn't ask for those things and it says 'you're right i'm so sorry' then it adds 2 other things i didn't want or it will change something that breaks everything. so you ask it to fix it. oh, so sorry, here's 2 more things you didn't ask for.” [source](https://www.reddit.com/r/cursor/comments/1wrktgc/am_i_the_only_one_who_thinks_grok_47_is_actually/pcdlxql/)
- Praise, 2026-09-25, @cursor_ai (X): “@heygen @cursor_ai @anthropicai @claudeai every frame is editable, so i could fix pacing just by asking.” [source](https://twitter.com/47233122/status/2103527219650052147)
- Praise, 2026-09-25, @cursor_ai (X): “small agent habit that saved me hours this week in @cursor_ai: before a big refactor, i dump a one-screen "do not touch" list in the prompt (auth, billing, migrations). the agent goes faster when it knows the rails. #cursor #ai #softwareengineering #buildinpublic” [source](https://twitter.com/262960825/status/2103530958188392506)
- Complaint, 2026-09-26, @cursor_ai (X): “@bot actively ignores instructions, wastes tokens, either executes incorrectly or deliberately harms product resulting in lost revenue and chaos; @cursor_ai @xai support refuse accountability sighting that tokens spent cannot be refunded, only arrogantly suggesting to cancel the subscription if not happy?! initially humans suggested my inputs would help improve model behaviours, now humans just ignore my tickets. this is frankly unacceptable. @e” [source](https://twitter.com/293536978/status/2103779729237299495)
- Complaint, 2026-09-26, @cursor_ai (X): “lol, and in the cursor system prompt... baked in to every call: "bias towards not asking the user for help if you can find the answer yourself." (what it hears: make shit up and do whatever to reach 'done') and this gem as if nobody @cursor_ai ever tested what this does to its bias: "look past the first seemingly relevant result. explore alternative implementations, edge cases, and varied search terms until you have comprehensive coverage of the” [source](https://twitter.com/57034090/status/2103933504895471863)
- Complaint, 2026-09-25, r/cursor (Reddit): “i am a vibe coder. my pov is grok 4.7 isn't a big improvement to 4.6 it is overall okay. what annoys me on all grok versions it doesn't think through. sometimes it changes something without considering what was changed for a reason 2 chat inputs before. when i use chatgpt via cursor it seems to consider that to the conclusion. but for the price grok does a great job. we have to look not only about the output also about the costs. it could be a lo” [source](https://www.reddit.com/r/cursor/comments/1wpd3mm/i_hate_to_admit_it_but_grok_sucks/pbyw00m/)

### Pi

- Praise, 2026-09-27, r/PiCodingAgent (Reddit): “yeah so like i said, the tool restriction accomplishes what i wanted. this is a follow up post on how to best go about that piece. almost nothing to do with your comment, which is also unhelpful as tool restriction is much more effective than just tweaking prompt .md files.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrte31/alternative_to_tool_profiles_for_better_subagent/pcgoayr/)
- Praise, 2026-09-22, r/PiCodingAgent (Reddit): “i just use markdown and it’s working fine with 15 sequential steps inside my system prompt. system prompts have a higher priority than skills or user prompts so the model follows it better. i’m certain this recommendation breaks down eventually with scale but… might be fine for light factories.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wne4zn/anybody_using_a_workflow_engine_to_automate_their/pbf2fr8/)
- Praise, 2026-09-17, r/PiCodingAgent (Reddit): “the most useful feature is to have as many advisors as you wish with a general and/or individual set of instructions. omp is rich enough not to let me get back the hassle of pi+the whole messy pile of conflicting extensions marketplace.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wiyikk/how_many_of_you_use_the_advisor_agent_with_ohmypi/paezfss/)
- Complaint, 2026-09-26, r/PiCodingAgent (Reddit): “i've had wonderful success with opencode with swift q3.8 27b. though pi seems more composable but so frustrating to control. it doesn't seem to interact with me much as a user, it barely seems to respond to more than my first input. it's seems to chase its tail a lot. and that's using glm 5.3 flash!” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wqh2u4/ohmypi_or_opencode_why/pc41ov1/)
- Complaint, 2026-09-26, r/PiCodingAgent (Reddit): “too late. i'm already diagnosed with paranoia xd. but yes, that's exactly what happens; most of the time, qwen weaves previous and recent instructions and answers smoothly that it's hard to pick the issue. it's not only reasoning though. any interrupted response is not added to chat history. <strict_link>” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wcwsya/noob_question_how_to_interrupt_an_agent_during/pc91pq6/)
- Complaint, 2026-09-26, @pidotdev (X): “@warmwaffles @pidotdev nah it put me off using pi. got frustrated telling it to not make any changes and the llm ignoring me anyway i like plan mode” [source](https://twitter.com/1361615777300762629/status/2103732628465787207)

### GitHub Copilot

- Praise, 2026-09-23, @GitHubCopilot (X): “@arya_at1 @githubcopilot @github this is such a clean example of how to actually use agents well. the brief was tight and full of constraints instead of open-ended.” [source](https://twitter.com/1583159673728864256/status/2102810895340711975)
- Praise, 2026-09-23, @GitHubCopilot (X): “@arya_at1 @githubcopilot @github the specific instructions you left on the ticket are perfect. that level of clarity is why it worked.” [source](https://twitter.com/1879527228121526273/status/2102811127520571472)
- Praise, 2026-09-22, @GitHubCopilot (X): “the issue was labeled good-first-task. it was not. a rate limiter on a public route used a global map, so two instances doubled the quota and one deploy wiped the counts. that is the job i gave @githubcopilot. not autocomplete. the agent on the issue, in the same @github repo. brief i left on the ticket: keep the existing middleware do not add redis do not invent an api gateway store hits per instance without lying across deploys add a test that” [source](https://twitter.com/1989355273727967232/status/2102450302008095144)
- Complaint, 2026-09-27, r/GithubCopilot (Reddit): “why would you wanna use copilot harness? it keeps making mistakes, ignore instructions and repeatedly fails read/write ops. another person in this sub posted benchmarks that put it lower than other harnesses like pi/ds/codex” [source](https://www.reddit.com/r/GithubCopilot/comments/1wr5io2/best_cheaper_alternative/pca0pix/)
- Complaint, 2026-09-17, @GitHubCopilot (X): “i'm not sure that's a real answer to my question @githubcopilot, but thanks for the biographical details <strict_link>” [source](https://twitter.com/803219/status/2100613725682078193)
- Complaint, 2026-09-15, r/GithubCopilot (Reddit): “no once you work on really advance/harder stuff or projects with a lot of constraints there's a lot more planning/research/design/hand holding that needs to happen. once you encounter an xpc/threading bug and program keeps crashing, the frontier model keeps looping etc. and not listening you'll need to do something diff. if your code base is untenable, well bud. i usually have it generate invariants, architecture documents, adverserial reviews et” [source](https://www.reddit.com/r/GithubCopilot/comments/1wfyrjy/how_do_you_guys_do_agentic_coding/pa1z5qp/)

### Devin

- Praise, 2026-09-21, r/codex (Reddit): “you probably got a quantized model. tell it to provide a handoff and start over. but before you do that, get a second opinion from swe-2 or deepseek. you'd also probably get better results if you just used swe-2 and told it to call codex cli astra as an advisor. devin doesn't block itself on tests and such so much and does what you ask.” [source](https://www.reddit.com/r/codex/comments/1wm88ng/stuck_in_the_mud_spinning_the_wheels_but_no/pb4tx7q/)
- Praise, 2026-09-15, @cognition (X): “@cognition ignore my typos - devin can understand me and thats what matters &lt;3” [source](https://twitter.com/1463714589518884865/status/2099963757808234572)
- Praise, 2026-09-11, @DevinAI (X): “i used swe-2 max in @devinai to improve my temperature anomaly map that astra did in one shot (right) and i'm happy with the results (left). i'm especially impressed with swe-2: it follows instructions and delivers exactly what was asked <strict_link>” [source](https://twitter.com/2905330823/status/2098388171302043786)
- Complaint, 2026-09-18, @cognition (X): “@cognition and that wraps that. the swe2 model has 0 adherence to prompt safety. tell it to do something a specific way, if it fails, it doesn't flag it and tries to workaround instead.. tell it not to do something, it will take that as instruction to do it. its just shit tier.” [source](https://twitter.com/1917224549604605953/status/2101091262363492357)
- Complaint, 2026-09-17, r/windsurf (Reddit): “i ask codex to directly edit the code not using apply patch script. but this idiot never listen to me anyway i see my colleague and companies starting to use devin more than codex rn” [source](https://www.reddit.com/r/windsurf/comments/1wipaqx/im_moving_to_devin/pagunse/)
- Complaint, 2026-09-11, @cognition (X): “@cognition the bar for these is simple. follow my instructions and finish the job.” [source](https://twitter.com/2087402808756629504/status/2098397809372479946)

### Amp

- Praise, 2026-09-15, @AmpCode (X): “@nerdworldorder @ampcode @thorstenball that refinement is a good one — the failure mode with a flat prohibition is the model just picks an arbitrary number anyway, so anchoring it to the last real duration plus a buffer keeps it honest. i log the actual run time now so that number is always on hand instead of guessed.” [source](https://twitter.com/2272515576/status/2099888724851204170)
- Praise, 2026-09-11, @AmpCode (X): “i rarely have to yell at agents when i use @ampcode” [source](https://twitter.com/942412840673169408/status/2098230750177014048)
- Praise, 2026-09-09, @AmpCode (X): “@mrsanders @ampcode @amp cool. i have found amp to follow through on goals you express in natural language, without needing something like poteto-mode, but i am fully open to the idea that i'm missing something. (i also wonder why cursor doesn't bring poteto-mode into core if it's so good!)” [source](https://twitter.com/784008/status/2097778691166396809)
- Complaint, 2026-09-26, @AmpCode (X): “@thorstenball @ampcode is that live or are you working on it? if it’s live, the model doesn’t seem to know to use if!” [source](https://twitter.com/5444392/status/2103702725733294250)
- Complaint, 2026-09-24, @AmpCode (X): “mainly when i throw in a new idea, hand it a design mockup, or ask it to refactor something big, it tends to make a mess. breaks things, ignores instructions. these are problems most harness solved earlier this year, but amp's harness hasn't caught up. also, no plan mode. for larger tasks the model still needs to plan before it acts. amp has oracle but it wasn't enough, i ended up writing my own planning skill to compensate. cc handles this nativ” [source](https://twitter.com/1592160489965948933/status/2103099981473382778)

### Cline

- Praise, 2026-09-20, @cline (X): “been trying out kimi k3 in @cline and its awesome. sticks to the tasks, answers correctly. does the job and no bullshit the kind of vibes i got from grok 4.6” [source](https://twitter.com/1999052311897972736/status/2101616688412475434)
- Complaint, 2026-09-22, r/CLine (Reddit): “i am using cline 4.1.17 vscode extension with a self hosted glm 5.3 flash. cline is accessing it via openai compatible api key. glm 5.3 in most cases showing "i don't find a mode tag explicitly in my view" in its reasoning, and ignoring the plan mode completely, and proceeds to edit file. when editing file, it is also not showing me the file editing as track change in focus mode even though "background edit" is disabled.” [source](https://www.reddit.com/r/CLine/comments/1wn7q28/models_are_not_seeing_and_ignoring_mode_tag/)
- Complaint, 2026-09-22, @cline (X): “@ticassociation @cline @ollama yea i tried to isolate it to just one file but still it had a really hard time following directions.” [source](https://twitter.com/1155165294530310146/status/2102361502896492908)

### Factory

- Praise, 2026-09-08, @droid (X): “droid is amazing but astra is great. weird to see that there is no improvement. it cheaper than sol for me. it makes everything i ask. the steerable and not making stupid mistakes. the one thing i’m not always sure about is: astra will do exactly how you ask it to do. and it’s not expensive on pro sub. performance depends on reasoning. tried first time ever. because of reset. it drains less than 20% over the night. but i was aware of re” [source](https://twitter.com/7344112/status/2097322660740927789)
- Praise, 2026-09-07, @FactoryAI (X): “to be completely honest, it’s the best harness i’ve ever experienced, so i’ve just been learning from it and can’t really offer any feedback. in particular, the models in droid core were so impressive—they often exhibit behaviors i’ve never experienced in other harnesses, to the point where i wonder if they are truly the models i thought i knew. and agent readiness makes me want to subscribe just for that feature alone...” [source](https://twitter.com/1620561425600221184/status/2097100298695344360)
- Praise, 2026-09-07, @FactoryAI (X): “with other harnesses, i had to watch them closely because you never knew where they might veer off. with droid, there was no need for that, and i could barely notice any difference even across different frontier models. everyone i recommended it to and who tried it was amazed, like "is this the power of a harness?"” [source](https://twitter.com/1620561425600221184/status/2097105067631681580)
- Complaint, 2026-09-22, G2 (G2): “q: what problems is the product solving and how is that benefiting you? a: factory ai help me reduce manual coding work and save development time. it help with repetitive tasks, debugging and building features faster. i can focus more on important work instead of doing everything manually. it make my daily workflow more easy and productive. q: what do you like best about the product? a: factory ai is helpful for automating development work. it sa” [source](https://www.g2.com/products/factory-ai/reviews/factory-ai-review-13384158)

### Grok Build

- Praise, 2026-09-24, r/google_antigravity (Reddit): “i primarily work in swift so it's a mixed bag, especially during any transition. we're currently moving to ios 27, which some llms assert still doesn't even exist yet, so trying to do anything "new" is still best done by hand. i have started playing around with grok build for small personal projects i don't have time to work on but really want to tinker with and it's surprisingly good for slightly-beyond-prototype work. it has a strong grasp of d” [source](https://www.reddit.com/r/google_antigravity/comments/1wp6mpo/poll_how_do_you_code_in_late_2026/pbsymyj/)
- Praise, 2026-09-22, r/ClaudeCode (Reddit): “usage burns fast on codex too but at least is still competent on sol 5.6. claude on opus 5 has gotten nearly unusable. i will say that my first impressions of grok build are good. its surprisingly much better than claude at actually following rules and not drifting into pure insanity.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wmbt29/two_5x_sub_or_one_20x_sub/pba474l/)

### Zed

- Praise, 2026-09-07, @zeddotdev (X): “@ivan_herdian @zeddotdev di sinilah unpopular opinion, aku butuh yang nurut bukan yg minteri 😂 <strict_link>” [source](https://twitter.com/1976991588/status/2096969892360515871)

### Kiro

- Complaint, 2026-09-04, r/kiroIDE (Reddit): “i personally don't like gemini for coding, it's very proactive and do way to much stuff that is hard to follow and by passes lint rules i have on place instead of research the rule in the first place. what is being very useful is to research code and give me report but you have to ask it to quote where exactly got the claims that gives you back because otherwise it just makes things up with no cross checking” [source](https://www.reddit.com/r/kiroIDE/comments/1w6xdti/bring_gemini_38_flash_and_grok_models_to_kiro_as/p7sr4lg/)

### Conductor

- Complaint, 2026-09-18, @conductor_build (X): “i gave @conductor_build a shot but not sure whats going on with the harness? gave exact same message to codex and it just knew what i mean new chats on both <strict_link>” [source](https://twitter.com/2715100816/status/2101005041817550850)
