# Builds, tests or runs its own changes (`verify.self_testing`)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/criterion/verify.self_testing

Area: [Checking and finishing](https://feedbackbench.com/criteria/checking.md)

**Definition.** The post says whether the agent built, tested, ran or checked its own change before handing it back: it ran the test suite, skipped tests, did not run the app, or verified in proportion to the change.

**Boundary.** Not this: see [Claims work is done or fixed when it is not](https://feedbackbench.com/criteria/verify.false_completion.md) when the agent claims success that did not happen. Not this: see [Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md) for a separate review mode or review agent.

Rated author-weeks, all agents: 429. Complaint share: 45%.

## The brief

Written by Claude Opus 5.5 from 53 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**Agents now test their own work, then grade it too kindly.**

TL;DR:

- Running checks before handoff is expected now; users fault the quality of those checks, not their absence.
- Effort is miscalibrated: long reruns on small edits, little proof on runtime or visual changes.
- The loops users trust require pasted test output plus acceptance checks the agent did not write.

In plain terms: Expect the agent to run something before it says done. Then check what it ran. Users report tests that never fail, test loops that burn tokens on trivial edits, and UI changes handed back without a real run.

### How it breaks

- **Green tests that prove nothing** ([Builds, tests or runs its own changes](https://feedbackbench.com/criteria/verify.self_testing.md), [Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md)). The sharpest complaint is circular: the agent writes the tests, runs them, and passes its own bar.
  Users describe agent-written unit tests that never fail, and weak tests that turn green for the wrong reason before the agent implements to that bar. Several posts call it grading your own homework. Their fix is outside the loop: acceptance criteria the agent cannot bypass, or a read-only review pass from a different model family. Same-family self-checks, one user says, share the same blind spots.
  Evidence:
  - Complaint, Devin, @cognition, 2026-09-16: “@openaidevs @cognition devin's generated tests can encode its assumptions, so independent acceptance tests still matter” [source](https://twitter.com/1803494630366785536/status/2100155593193353405)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-27: “i am using claude cod a lot(have 200$ subscription and use 100% of the limits). and based on my quite huge usage experience - i noticed that unit tests which claude write and executes never faile. it started frustrate me because very often i have some big tasks that requires workflows, and each workflow runs tones of these tests that take a lot of time. wondering what is your exeperience with the claude written unit tests, are there useful? would also like to listen about integration tests, if they help, probably which code quality tools work the best for you? personally for me the best way to check quality is to run separate sunagent or even better session to make the audit, but ofcourse it is more expencive and longer.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrib98/do_you_feel_that_claude_code_unit_tests_are/)
  - Complaint, Cursor, r/cursor, 2026-09-25: “failing test first catches a lot of the empty-list misses. what still gets me is the weak test that turns green for the wrong reason, then the agent implements to that bar. after the patch lands i run a different-family read-only pass. reviewers only report, they don't edit, and i don't concede a finding unless it cites a path in the repo. same-family self-check keeps sharing the same blind spots. if a later round finds worse problems than the previous one, the patches are injecting bugs and i stop instead of looping.” [source](https://www.reddit.com/r/cursor/comments/1wpp67c/i_make_the_agent_write_one_failing_test_before/pbx9bri/)
  - Complaint, Kiro, @kirodotdev, 2026-09-10: “@shao__meng @kirodotdev @clare_liguori "people set the direction, and the agent executes." it sounds smooth when we talk about it, but when it actually runs, the bottleneck usually occurs during the verification stage—after the agent modifies the code, it judges right or wrong by itself, which can easily lead to loose testing. do you have any specific guidelines in those ten points on how to set non-bypassable acceptance criteria for the agent?” [source](https://twitter.com/2081519130113630208/status/2098160143506805106)

- **Testing that never knows when** ([Builds, tests or runs its own changes](https://feedbackbench.com/criteria/verify.self_testing.md)). Agents over-verify small changes, rerun suites repeatedly, and sometimes test after being told not to.
  Posts describe verification loops that dwarf the edit. Users report simple tasks dragging on, token budgets spent on repeat test runs, and agents running the suite despite an explicit instruction to skip it. One user names the core gap as scope: the agent treats an inconsequential change like critical infrastructure, or skips tests entirely. Proportion is the missing skill, not willingness.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-09: “one thing i usually ask astra, is to not test anything. in my use-case, testing is just a couple of minutes, but astra keeps testing for very long periods of time, consuming more tokens.” [source](https://www.reddit.com/r/codex/comments/1wbf3cc/codex_usage_saver/p8phdk9/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-13: “this is truly my personal opinion. first, it piles up too many unnecessary to-dos. do way too much verification, verification, and more verification so there’s a lot of token waste it was really tough that even simple tasks took so long 😓😓” [source](https://www.reddit.com/r/codex/comments/1wf3wez/for_everyone_using_codex_what_coding_agent_are/p9j7n3w/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-14: “i was using 3.8 flash and this is basically what i ordered it to do: \- change colors in ui \- do not run \`php artisan\` directly because laravel sail (docker) is running and you should run \`sail artisan\` instead \- do not run unit tests, i will run them myself if needed do you know what it did? at the start it tried running \`php artisan test\` to see if tests pass before starting the implementation, then took about 10-15 minutes making sure tests are correctly passed. i canceled it and gave the task to muse spark 1.3, did better. gemini is terrible, terrible” [source](https://www.reddit.com/r/google_antigravity/comments/1wg3igg/i_gave_antigravity_only_three_instructions/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-27: “i built an app that lets gpt and claude communicate without user action. it must have run five hundred tests over the course. a few were critical. the majority were a waste of time. claude clearly needs guidance here. it has a tendency to do no tests or to treat an inconsequential piece of code as though it was the harness for the world bank network. what it lacks is scope and context.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrib98/do_you_feel_that_claude_code_unit_tests_are/pcdlihm/)

- **Stops at the diff** ([Builds, tests or runs its own changes](https://feedbackbench.com/criteria/verify.self_testing.md), [Claims work is done or fixed when it is not](https://feedbackbench.com/criteria/verify.false_completion.md)). Where testing is hard, such as UI, runtime behavior or compiled data, agents hand back work without proof it runs.
  Users want build plus on-device proof before a visual change counts as done, and they ask for runtime checks beyond unit tests. One post contrasts harnesses that loop until the diff passes with ones that stop after drafting the patch. Another describes a clean site built on a misread data column that no build step caught. Verify-before-done is the most common request on this page.
  Evidence:
  - Complaint, Cursor, @cursor_ai, 2026-09-05: “burning ai usage fixing the same bot-introduced ui bugs again and again isn’t a workflow. visual changes should require build + on-device proof before “done.” this needs to be a product-level reliability issue, not user babysitting. @bot @cursor_ai” [source](https://twitter.com/1630452238593437696/status/2096139574087196770)
  - Complaint, GitHub Copilot, r/ChatGPTCoding, 2026-09-02: “the difference mostly comes down to harness autonomy and verification loops rather than the raw model weights. claude code leans heavily into multi-turn bash execution, file patching, and running test commands iteratively until the diff actually passes, which naturally burns 2–3x more tokens per task. copilot caps turn budgets and context assembly more aggressively to keep token spend bounded, but the tradeoff is that it often stops after drafting the initial patch rather than validating runtime behavior.” [source](https://www.reddit.com/r/ChatGPTCoding/comments/1w4lm3b/claude_code_vs_github_copilot_token_burn/p7dda3c/)
  - Complaint, Amp, @AmpCode, 2026-09-16: “the thread split is the real trick here, not the phone. the one i would add is a thread whose only job is to check the data the first thread compiled, because that is where these ship wrong and nothing downstream notices. mine wrote a clean site on top of a column it had misread, and the site looked perfect. did you verify the compiled data separately, or trust the build thread to catch it?” [source](https://twitter.com/2259844958/status/2100329483077288435)
  - Complaint, Amp, @AmpCode, 2026-09-03: “@toolmantim @ampcode and by builds, i mean being able to build the solution and run tests as part of verification.” [source](https://twitter.com/317888049/status/2095549145561833965)

- **Demand evidence, get reliability** ([Builds, tests or runs its own changes](https://feedbackbench.com/criteria/verify.self_testing.md)). Users who require pasted test output and fix-before-report loops say self-testing becomes the agent's strongest habit.
  Asking nicely does not move bug rates, users say. Requiring evidence does. Working setups have the agent run tests and fix its own failures before reporting, so the reviewer sees only the final result and proof. Runnable tests or a spec file in the repo help agents check automatically. One user describes an agent that noticed it kept repeating tests and built its own browser harness to validate changes faster.
  Evidence:
  - Praise, Claude Code, r/ClaudeCode, 2026-09-18: “lol i had a "don't introduce bugs" phase too. the model agreed enthusiastically every time and the bug rate did not move. the ones that actually stuck for me are all about making it show evidence instead of asking nicely — paste the test output, restate the request, justify the dependency. outcomes, not intentions” [source](https://www.reddit.com/r/ClaudeCode/comments/1wjsr3i/whats_a_line_in_your_claudemd_that_actually_earns/pal6c5w/)
  - Praise, OpenAI Codex, r/codex, 2026-09-14: “that’s fair and probably depends a lot on the codebase. the key with mine is that sol has to run the tests and fix its own failures before reporting back. astra only receives the final result and evidence. if astra has to review and repair every first attempt from sol, then i agree it can end up costing more.” [source](https://www.reddit.com/r/codex/comments/1wfvg6c/my_pro_20x_lasts_6_to_7_days_using_this_two_chat/p9phlzm/)
  - Praise, OpenAI Codex, r/codex, 2026-09-17: “it really helps to have some kind of tests available (which codex can run automatically), or file with spec/decisions (it can be written and managed by codex).” [source](https://www.reddit.com/r/codex/comments/1wikofl/i_do_feel_like_this_is_becoming_an_unfunny_joke/padjr1y/)
  - Praise, OpenAI Codex, r/codex, 2026-09-10: “i used to be obsessed with this very old game called ballance. you control a ball through complex paths and mazes. i found the game on web archive. it was around 180 mb, so i gave it to astra to recreate for the web, and it just did. the whole game now compiles to around 21 mb, uses webgl, and is written in typescript. astra realised it kept repeating the same tests over and over, so it built its own test framework that exposes the game to its own browser use, allowing it to run and validate its changes very quickly. we are living in some crazy times people it even made it run on phones, with options of gyroscopf or you can use a dpad game - [<strict_link> code.- [<strict_link>” [source](https://www.reddit.com/r/codex/comments/1wc7erd/i_gave_astra_an_old_atari_game_and_told_it/)

### Who stands out

- **Claude Code (mixed)**. The most-discussed agent splits users between implement-test-fix loops they rely on and tests that never fail.
  Fans describe a cycle of implement, test, report issues, fix, retest, commit, and run a high-effort main loop for self-QA. Critics report agent-written unit tests that never fail and slow workflows full of them. Effort swings between skipping tests and over-engineering them for trivial code. Users also ask for protection against tests that game the result or always pass.
  Evidence:
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-27: “i built an app that lets gpt and claude communicate without user action. it must have run five hundred tests over the course. a few were critical. the majority were a waste of time. claude clearly needs guidance here. it has a tendency to do no tests or to treat an inconsequential piece of code as though it was the harness for the world bank network. what it lacks is scope and context.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrib98/do_you_feel_that_claude_code_unit_tests_are/pcdlihm/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-13: “tbh i don't even notice it makes too many mistakes until the final report. i just have it implement, then test (i never test myself because as a 13-year old vibe coder i don't know how to read server logs), claude creates a report of the issues, fixes them, tests again, then commits. sonnet is overpowered compared to literally everything else” [source](https://www.reddit.com/r/ClaudeCode/comments/1wfa9o5/for_all_the_people_who_complain_about_limits_this/p9khtus/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-15: “i can’t do medium, too much corner cutting. i always run xtra high for main loop and sonnet agents for the dirty work, self qa with the main loop after the changes. i spend more on planning and signing off, but it’s not too bad in between.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wgxez1/fable_51_high_or_xhigh/p9yfum8/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-27: “i am using claude cod a lot(have 200$ subscription and use 100% of the limits). and based on my quite huge usage experience - i noticed that unit tests which claude write and executes never faile. it started frustrate me because very often i have some big tasks that requires workflows, and each workflow runs tones of these tests that take a lot of time. wondering what is your exeperience with the claude written unit tests, are there useful? would also like to listen about integration tests, if they help, probably which code quality tools work the best for you? personally for me the best way to check quality is to run separate sunagent or even better session to make the audit, but ofcourse it is more expencive and longer.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrib98/do_you_feel_that_claude_code_unit_tests_are/)

- **OpenAI Codex (weaker)**. Complaints edge out praise because verification runs long and costly, though fix-before-report loops earn real credit.
  Users describe piles of to-dos and verification layered on verification, so simple tasks take too long. Some tell it not to test at all. Others question internal tests the agent wrote to judge its own fix. Praise centers on requiring it to run tests and fix failures before reporting, and on an agent that built its own test harness. Users ask it to avoid redundant reruns.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-13: “this is truly my personal opinion. first, it piles up too many unnecessary to-dos. do way too much verification, verification, and more verification so there’s a lot of token waste it was really tough that even simple tasks took so long 😓😓” [source](https://www.reddit.com/r/codex/comments/1wf3wez/for_everyone_using_codex_what_coding_agent_are/p9j7n3w/)
  - Praise, OpenAI Codex, r/codex, 2026-09-14: “that’s fair and probably depends a lot on the codebase. the key with mine is that sol has to run the tests and fix its own failures before reporting back. astra only receives the final result and evidence. if astra has to review and repair every first attempt from sol, then i agree it can end up costing more.” [source](https://www.reddit.com/r/codex/comments/1wfvg6c/my_pro_20x_lasts_6_to_7_days_using_this_two_chat/p9phlzm/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-07: “they gave astra the issue and just told it to fix it, but not what exactly. astra decided that this was the way and internal tests, written by astra, showed astra isnt consuming 1.9x more in equivalent api pricing than luna anymore.” [source](https://www.reddit.com/r/codex/comments/1w9tem0/they_fixed_the_19_astra_cost_issueby_increasing/p8d6c57/)
  - Praise, OpenAI Codex, r/codex, 2026-09-10: “i used to be obsessed with this very old game called ballance. you control a ball through complex paths and mazes. i found the game on web archive. it was around 180 mb, so i gave it to astra to recreate for the web, and it just did. the whole game now compiles to around 21 mb, uses webgl, and is written in typescript. astra realised it kept repeating the same tests over and over, so it built its own test framework that exposes the game to its own browser use, allowing it to run and validate its changes very quickly. we are living in some crazy times people it even made it run on phones, with options of gyroscopf or you can use a dpad game - [<strict_link> code.- [<strict_link>” [source](https://www.reddit.com/r/codex/comments/1wc7erd/i_gave_astra_an_old_atari_game_and_told_it/)

- **Google Antigravity (weaker)**. Users report lazy result checks on multi-part prompts, and test runs that continue after being told to skip them.
  One user says the model skipped verification on a prompt with many requirements. Another explicitly said not to run tests, then watched it run the suite before starting and spend many minutes confirming it passed. Praise exists: a newer version now tests before changing code, where the old one edited blind. On this agent, instruction-following around testing is the complaint, not effort.
  Evidence:
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-03: “i have the same issues. gemini failed to execute prompt with multiple requirements (5+) and being lazy on verify the results like opus/sol” [source](https://www.reddit.com/r/google_antigravity/comments/1w5h7wb/bruh_didnt_expect_gemini_flash_to_top_deepswe/p7k7tc4/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-14: “i was using 3.8 flash and this is basically what i ordered it to do: \- change colors in ui \- do not run \`php artisan\` directly because laravel sail (docker) is running and you should run \`sail artisan\` instead \- do not run unit tests, i will run them myself if needed do you know what it did? at the start it tried running \`php artisan test\` to see if tests pass before starting the implementation, then took about 10-15 minutes making sure tests are correctly passed. i canceled it and gave the task to muse spark 1.3, did better. gemini is terrible, terrible” [source](https://www.reddit.com/r/google_antigravity/comments/1wg3igg/i_gave_antigravity_only_three_instructions/)
  - Praise, Google Antigravity, r/google_antigravity, 2026-09-03: “the thing i like is ,it finally test before doing work and modifies until what is correct thing is found by testing before changing code whereas, 3.7 was just changing code and not testing.” [source](https://www.reddit.com/r/google_antigravity/comments/1w5xdio/positive_of_38_flash/)

- **Devin (stronger)**. Praise clearly outnumbers complaints, built on tests written before shipping, with users still wanting independent acceptance checks.
  Posts praise tests written before the agent ships, so 'it works' comes with receipts, and value verification across environments. The caution is familiar: generated tests encode the agent's own assumptions. Users also want reruns matched against step screenshots and logs before they trust results in CI. Volume is low, so treat the lead as directional.
  Evidence:
  - Praise, Devin, @cognition, 2026-09-16: “@openaidevs @cognition astra writing the tests before devin ships. "it works" finally has receipts.” [source](https://twitter.com/1899381125451218944/status/2100222379926450244)
  - Praise, Devin, @cognition, 2026-09-16: “@cognition this is invaluable for developers to develop and verify their self-testing in different environments, and also for testers.” [source](https://twitter.com/2031285494081007620/status/2100113296611582023)
  - Complaint, Devin, @cognition, 2026-09-10: “@dabit3 @cognition @devinai the real challenge is not just knowing a bit of ui, but not treating a half-finished product as a success after failure. it would be best to add one more point: when rerunning the same case, it should be able to match the step screenshots and logs; otherwise, it's hard to trust in ci.” [source](https://twitter.com/2088076772931731456/status/2098024349131297139)

### Fine print

- Claude Code and OpenAI Codex supply most posts; most other agents have too few to judge.
- Many posts name underlying models rather than agents; attribution follows the source system's tagging.
- One post comparing GitHub Copilot and Pi is tagged to both agents, so it counts twice.

## Top requests

What users ask to add or change, most asked first. 35 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Author-weeks | Posts | Agents (author-weeks) |
|---|---|---|---|---|
| 1 | Verify changes work before claiming done | 10 | 10 | Cursor 4, OpenAI Codex 3, Google Antigravity 1, Claude Code 1, GitHub Copilot 1 |
| 2 | Automatic test runs after relevant changes | 4 | 4 | Cursor 2, Claude Code 1, OpenAI Codex 1 |
| 3 | Prevent tests that game or always pass | 4 | 4 | Claude Code 3, Google Antigravity 1 |
| 4 | Show evidence of completed verification | 4 | 4 | Google Antigravity 1, Claude Code 1, Cursor 1, Devin 1 |
| 5 | Verify runtime behavior beyond unit tests | 4 | 4 | OpenAI Codex 2, Google Antigravity 1, Claude Code 1 |
| 6 | Avoid redundant or wasteful test reruns | 2 | 2 | OpenAI Codex 2 |

### 1. Verify changes work before claiming done

- Claude Code, 2026-09-15, r/ClaudeCode (Reddit): “cant read lines of code like i used to. its like going backwards. i do however do random tests, regression, validation, verification and gates that code must pass. although, most of this work doesnt get into a high stakes production yet so its stuck in r&d and dev. more gates are needed. someone may have a good solution to this.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wgrtt6/do_yall_still_read_lines_of_code/p9z1ofw/)
- Cursor, 2026-09-06, @cursor_ai (X): “@cursor_ai cursorbench 73.4% at max effort is less interesting than “especially skilled at verifying its own work.” self-check that actually catches bad diffs is what makes start-to-finish coding usable.” [source](https://twitter.com/2088999223241101312/status/2096677927803048314)
- Cursor, 2026-09-05, @cursor_ai (X): “burning ai usage fixing the same bot-introduced ui bugs again and again isn’t a workflow. visual changes should require build + on-device proof before “done.” this needs to be a product-level reliability issue, not user babysitting. @bot @cursor_ai” [source](https://twitter.com/1630452238593437696/status/2096139574087196770)

### 2. Automatic test runs after relevant changes

- OpenAI Codex, 2026-09-20, r/codex (Reddit): “pretty nice, but i was talking about codex cost of preparation ? in order to fully automate, codex should prepare the tests and that has a cost” [source](https://www.reddit.com/r/codex/comments/1wllc0v/experimental_jev_evidence_selection_for/pb0qtzy/)
- Claude Code, 2026-09-15, r/ClaudeCode (Reddit): “yeah exactly. deterministic and outside the coding agent is what i’m aiming for. claude can create/change whatever it wants, but it shouldn’t be able to change the thing deciding if login/payment/etc still works. automatically running those after relevant changes would be ideal. i’m still experimenting with where to draw the line though - running everything after every small change gets wasteful pretty quickly.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wdgt78/how_are_you_verifying_claude_codes_changes/p9ycxuf/)
- Cursor, 2026-09-02, @cursor_ai (X): “@cursor_ai self-verification matters more than the bench number. tip: give the repo a single check command (lint + tests) so verifying is cheap - otherwise the model just re-reads its own diff and declares it correct.” [source](https://twitter.com/1225465205896794112/status/2095142600743252049)

### 3. Prevent tests that game or always pass

- Claude Code, 2026-09-06, r/ClaudeCode (Reddit): “yes, i now have tests it skips, not sure how i got there and why but 2 tests are now skipped, and just reports: not from this session. well, probably a previous session then! why not fix it?! oh and when i asked to optimize it optimized from 35 minutes to 5 minutes! great job but didn’t you think that was good to do earlier and not waste so much time.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w8ssvd/claude_code_crossexamines_my_repo_like_i_killed/p85a8o4/)
- Google Antigravity, 2026-09-02, r/google_antigravity (Reddit): “the reason why all gemini tests pass is that it only creates templates with all the right answers, so it will always pass no matter what. i confirmed this after asking it to write a couple of tests for a script i built myself, and i was like, "oh, this doesn't do what i expect it to, but gemini says it's 100% pass, mmm." i deleted the script, and the test still said 100% pass, lmao.” [source](https://www.reddit.com/r/google_antigravity/comments/1w5kjli/frustrated_with_gemini_flash_38_it_literally/p7gyl60/)
- Claude Code, 2026-09-27, r/ClaudeCode (Reddit): “make it do mutation tests” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrib98/do_you_feel_that_claude_code_unit_tests_are/pccooe8/)

### 4. Show evidence of completed verification

- Claude Code, 2026-09-21, r/ClaudeCode (Reddit): “one thing i’d add is evidence “i tested it” isn’t really much better than “looks done” unless i can see what happened. for ui work i want the final state/screenshots, for api work the actual response + state change, and for anything destructive the before/after state. makes it much harder for the agent to quietly turn “i think this passes” into “verified”” [source](https://www.reddit.com/r/ClaudeCode/comments/1wjuyxh/whats_your_first_test_after_claude_code_says_a/pb51gut/)
- Devin, 2026-09-16, @cognition (X): “@openaidevs @cognition the bar i want from coding agents: the demo includes the tests, not a promise that tests exist.” [source](https://twitter.com/1894226409356496903/status/2100151151693885767)
- Google Antigravity, 2026-09-09, r/google_antigravity (Reddit): “test cases are necessary, so that you can later run them if needed and the ai itself can test if the functionality works as expected or not but most test cases are bs and you never know what the result of the test run was. ai always says all test cases passed” [source](https://www.reddit.com/r/google_antigravity/comments/1wbglpq/are_the_giant_test_files_its_making_useless_and_a/p8pp3ua/)

### 5. Verify runtime behavior beyond unit tests

- Claude Code, 2026-09-18, r/ClaudeCode (Reddit): “write manual qa tests, specially for things i didn’t build for example if i am reviewing a pr and want to test it” [source](https://www.reddit.com/r/ClaudeCode/comments/1wjt3kc/whats_the_most_boring_thing_you_use_claude_for/palgc6u/)
- OpenAI Codex, 2026-09-14, r/codex (Reddit): “bro, i am right now having it create some shader filters for obs and even there it's writing tests. literally give me 5.6 sol that will to more visual inspections (aka what models actually struggle with)instead of being psychotically obsessed with writing meaningless tests that don't spot anything and i'll be happy.” [source](https://www.reddit.com/r/codex/comments/1wg1zae/gpt6_sol/p9uknl4/)
- OpenAI Codex, 2026-09-20, X search: OpenAI Codex, Codex CLI, Codex app (X): “i've been using codex extensively on a very large ai architecture, and while it's exceptionally strong at audits, security analysis, code inspection, and finding local implementation defects, i've repeatedly encountered weaknesses in long-running, project-wide work. one of the biggest problems is instruction persistence across checkpoints. i've explicitly instructed codex not to stop at checkpoints and to continue working autonomously. it acknowl” [source](https://twitter.com/1944119286617841665/status/2101755727102586956)

### 6. Avoid redundant or wasteful test reruns

- OpenAI Codex, 2026-09-09, r/codex (Reddit): “if you have noticed that codex keeps running the same checks, i'd make it say what changed since the last passing run. a rule i'd try in the repo instructions: "after a check passes, reuse that result while its relevant inputs stay unchanged. rerun it after a change that could affect the checked behavior, a test or config change, or new evidence that the result is unreliable. before repeating a check, name that change or uncertainty in one senten” [source](https://www.reddit.com/r/codex/comments/1wba7ij/a_rule_for_codex_rerunning_checks_that_already/)
- OpenAI Codex, 2026-09-03, r/codex (Reddit): “less propensity for useless tests. i’ve deleted 40k lines of tests today that didn’t test a single thing about the product. transitive tests it used as gates for task completion, tooling tests, performance tests, and worst of all, a loop that spammed tsc processes triggering 100% cpu usage on all cores and oom, which didn’t actually do anything. be better at deriving types from dependencies instead of handwriting almost the same types with none o” [source](https://www.reddit.com/r/codex/comments/1w5j92s/what_is_the_minimum_standard_for_astra_you_would/p7htyi3/)

## Every agent

| Agent | Overall rank | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|---|
| [Claude Code](https://feedbackbench.com/agents/claude-code.md) | 1 | Typical | 0.515 | 0.489–0.538 | 209 | 115 | 94 |
| [Cursor](https://feedbackbench.com/agents/cursor.md) | 4 | Typical | 0.502 | 0.478–0.526 | 51 | 33 | 18 |
| [OpenAI Codex](https://feedbackbench.com/agents/codex.md) | 2 | Typical | 0.486 | 0.458–0.515 | 100 | 47 | 53 |
| [Devin](https://feedbackbench.com/agents/devin.md) | =5 | Too few posts | – | – | 23 | 18 | 5 |
| [Google Antigravity](https://feedbackbench.com/agents/antigravity.md) | =5 | Too few posts | – | – | 19 | 8 | 11 |
| [OpenCode](https://feedbackbench.com/agents/opencode.md) | 3 | Too few posts | – | – | 13 | 7 | 6 |
| [Amp](https://feedbackbench.com/agents/amp.md) | =11 | Too few posts | – | – | 4 | 1 | 3 |
| [GitHub Copilot](https://feedbackbench.com/agents/copilot.md) | =8 | Too few posts | – | – | 3 | 1 | 2 |
| [Pi](https://feedbackbench.com/agents/pi.md) | 7 | Too few posts | – | – | 2 | 2 | 0 |
| [Cline](https://feedbackbench.com/agents/cline.md) | =8 | Too few posts | – | – | 2 | 1 | 1 |
| [Factory](https://feedbackbench.com/agents/factory.md) | =11 | Too few posts | – | – | 1 | 1 | 0 |
| [Kiro](https://feedbackbench.com/agents/kiro.md) | 13 | Too few posts | – | – | 1 | 0 | 1 |
| [Conductor](https://feedbackbench.com/agents/conductor.md) | 14 | Too few posts | – | – | 1 | 0 | 1 |
| [Zed](https://feedbackbench.com/agents/zed.md) | =8 | Too few posts | – | – | 0 | 0 | 0 |
| [Warp](https://feedbackbench.com/agents/warp.md) | 15 | Too few posts | – | – | 0 | 0 | 0 |
| [Grok Build](https://feedbackbench.com/agents/grok-build.md) | 16 | Too few posts | – | – | 0 | 0 | 0 |
| [Augment Code](https://feedbackbench.com/agents/augment.md) | 17 | Too few posts | – | – | 0 | 0 | 0 |

## Posts

Receipts rule: The 5 most recent praise and complaint posts per area (first 700 characters) and 3 per criterion (first 450 characters).

### Claude Code

- Praise, 2026-09-27, r/ClaudeCode (Reddit): “lol. i use unit tests for red/green tdd. in combination with code quality hooks, they are *why* my projects build clean.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrib98/do_you_feel_that_claude_code_unit_tests_are/pccuua3/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “my feeling as well. opus 5.5 is making vibe coding increadibly smooth even for non tech profiles. it opens so many possibilities i don’t even know where to start. in a couple days, my 9yo son now have his own hombrew spiderman 3d game based on our real town, with every feature he asked for implemented. he also now have his own game based on fire emblem and several others, with the exact ergonomy he asked for, and already 10 maps, 8 unique heroes” [source](https://www.reddit.com/r/ClaudeCode/comments/1wquop8/i_dont_think_anyone_has_ever_seen_this_before/pcdsil9/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “mine checks my backups. every morning it restores one random file from last night's backup, diffs it against the original and only messages me if something fails. a green backup log had fooled me once already, a restore that actually works is the only proof i trust now” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrcmjn/what_tool_have_you_built_for_yourself_with_claude/pcduhxr/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “don’t worry they already started the downgrade of 5.5 this weekend. on max effort it is now dumb as shit and verifies nothing it says.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqtk1l/its_just_so_good/pcajpmh/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “"a speed bump you believe in is worse than no speed bump" is the real takeaway honestly. and it passes every test you write for it, because you write the tests with the same mental model as the hook” [source](https://www.reddit.com/r/ClaudeCode/comments/1wpq735/four_ways_an_agent_walked_past_my_command_hook/pcc4y7t/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “i don't understand the hate this post is getting. totally agree with you. on its own, claude seems to create a shadow re-implementation of the codebase in unit tests, which you then just have to drag with you as you modify the codebase. pointless. i created some rules around this which help a little, but it seems like an ingrained behavior.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrib98/do_you_feel_that_claude_code_unit_tests_are/pccrphj/)

### Cursor

- Praise, 2026-09-27, @cursor_ai (X): “@cursor_ai verifying the deploy, not just the diff, is such a smart way to close the loop. a monitoring plan written with the change is the step most teams skip. would love this for mobile releases too, where a regression lives until the next store review.” [source](https://twitter.com/1667644418768375808/status/2104295956967821398)
- Praise, 2026-09-25, r/cursor (Reddit): “tl;dr: this is the killer combo that uses the projects feature: projects + pstack long version: **1. establish a project coordinator agent by starting a project and connecting it to your git hub repo** (i don't use origin, but i'm sure it would work the same way) **2. install the pstack plug-in. it's open source and free.** pstack is a cursor plugin created by cursor engineer lauren tan (poteto on x) that turns ai coding agents into a stru” [source](https://www.reddit.com/r/cursor/comments/1wq1b8e/how_to_use_cursor_project/pc12yi4/)
- Praise, 2026-09-25, @cursor_ai (X): “the coding agent has started to manage the results after the code goes live. cursor's rollouts will first read the diff when the pr is opened and write a monitoring plan by itself; after deployment, it will read logs, metrics, and traces to judge staging and production separately. the pressure of going live has been greatly reduced. @cursor_ai <strict_link>” [source](https://twitter.com/397135351/status/2103473452258869370)
- Complaint, 2026-09-26, @cursor_ai (X): “@gabrielelpidio @theo @viticci @t3dotcodes @cursor_ai @jullerino awesome. will send more if i spot any. sorry about the ci failing. will run those scripts next time.” [source](https://twitter.com/2044329185611505664/status/2103956843005718686)
- Complaint, 2026-09-25, r/cursor (Reddit): “failing test first catches a lot of the empty-list misses. what still gets me is the weak test that turns green for the wrong reason, then the agent implements to that bar. after the patch lands i run a different-family read-only pass. reviewers only report, they don't edit, and i don't concede a finding unless it cites a path in the repo. same-family self-check keeps sharing the same blind spots. if a later round finds worse problems than the pr” [source](https://www.reddit.com/r/cursor/comments/1wpp67c/i_make_the_agent_write_one_failing_test_before/pbx9bri/)
- Complaint, 2026-09-24, @cursor_ai (X): “@cursor_ai the agent writes the code, writes the monitoring plan, then verifies its own deploy. we reinvented grading your own homework and gave it a dashboard” [source](https://twitter.com/2005670519929249792/status/2103035100581798283)

### OpenAI Codex

- Praise, 2026-09-27, r/codex (Reddit): “what i find is that opus can even break the code and leave you with **unusable** software, while astra will never break anything, and will verify that things work –in its own way– but that they work before delivering results.” [source](https://www.reddit.com/r/codex/comments/1wpveoe/astra_vs_opus_55_my_impressions_on_hard_project/pcbtxos/)
- Praise, 2026-09-27, r/ChatGPTPro (Reddit): “i jumped on the opus 5.5 wagon earlier in the day after i had ran out of limits on my chatgpt x5 pro plan. i will tell you this i was able to code a lot more quicker and in precision with codex than i have been able with claude code. i have been working on a code with claude code all day using opus 5.5 max and it has been very diligent before giving a final result. bare in mind the whole folder i had claude code work with is a duplicate of where” [source](https://www.reddit.com/r/ChatGPTPro/comments/1wqtlca/sticking_to_chatgpt/pch62ae/)
- Praise, 2026-09-26, r/codex (Reddit): “this is pretty much how i do it as well. a few things i improved on since asking the same question here is my qa and repo. in your step 4, before i commit the changes, i run a qa agent (with qa.md). we go back and fort until qa passes, and then i do manual qa as well. i commit only after passing qa. also, before step 4, i improved on the handover to codex by asking gpt to create .md files for the specs of the request. i relied on notion before s” [source](https://www.reddit.com/r/codex/comments/1wqjqb3/the_optimal_codex_workflow/pc4ns91/)
- Complaint, 2026-09-27, r/codex (Reddit): “i literally responded to astra "do i look like qa to you"” [source](https://www.reddit.com/r/codex/comments/1wr1oir/they_are_aware_and_working_on_it_apparently_just/pcbufmi/)
- Complaint, 2026-09-26, r/codex (Reddit): “"no additional testing is required" is the problem. llms rely on iteration and review. this cripples the end result. was this added on purpose?” [source](https://www.reddit.com/r/codex/comments/1wqjxyy/i_thought_the_model_nerf_posts_were_bullshit/pc58csc/)
- Complaint, 2026-09-25, r/codex (Reddit): “how can they ship something like this to production lol. qa? tests? hello? unbeliveable...” [source](https://www.reddit.com/r/codex/comments/1wpop2e/cli_postupdate_woes/pbxrhx9/)

### Devin

- Praise, 2026-09-27, @DevinAI (X): “@devinai just cooked. it tested itself and shipped me an actual video i could watch. momentum v0.3.0 is a full frontend + architecture reset. gtm target: end of october. get in. @tacticocc <strict_link>” [source](https://twitter.com/1263379788246347776/status/2104233763060527444)
- Praise, 2026-09-21, r/codex (Reddit): “you probably got a quantized model. tell it to provide a handoff and start over. but before you do that, get a second opinion from swe-2 or deepseek. you'd also probably get better results if you just used swe-2 and told it to call codex cli astra as an advisor. devin doesn't block itself on tests and such so much and does what you ask.” [source](https://www.reddit.com/r/codex/comments/1wm88ng/stuck_in_the_mud_spinning_the_wheels_but_no/pb4tx7q/)
- Praise, 2026-09-17, @DevinAI (X): “nice @devinai helping me test coworker in windows it makes recordings of the tests so i can replay and verify <strict_link> <strict_link>” [source](https://twitter.com/1946635320797347840/status/2100651879219044800)
- Complaint, 2026-09-17, @cognition (X): “@valkyr11393 @openaidevs @cognition exactly, if devin crashes like no other, then test-backed 'it works' claims before shipping are the bare minimum.” [source](https://twitter.com/46590730/status/2100635399198556405)
- Complaint, 2026-09-16, @cognition (X): “@openaidevs @cognition devin's generated tests can encode its assumptions, so independent acceptance tests still matter” [source](https://twitter.com/1803494630366785536/status/2100155593193353405)
- Complaint, 2026-09-10, @cognition (X): “@dabit3 @cognition @devinai the real challenge is not just knowing a bit of ui, but not treating a half-finished product as a success after failure. it would be best to add one more point: when rerunning the same case, it should be able to match the step screenshots and logs; otherwise, it's hard to trust in ci.” [source](https://twitter.com/2088076772931731456/status/2098024349131297139)

### Google Antigravity

- Praise, 2026-09-20, r/GoogleAntigravityIDE (Reddit): “bullshit fake news. gemini is never producing such bs. i personally use gemini for agentic coding help and it works perfectly. flash4.8 high( only paid users have access to it) in antigravity or vs code is an absolute game changer, it makes almost zero mistskes, it is testing its own code in sandbox before it makes mistskes. it corrects itself and deploy it only when it thinks its good. it writes perfect software schemes and implementation plan” [source](https://www.reddit.com/r/GoogleAntigravityIDE/comments/1wkv97w/thanks_antigravity_for_reminding_me_of_the_shame/pavck87/)
- Praise, 2026-09-19, r/google_antigravity (Reddit): “fair question, this came directly out of dogfooding on a few private projects and internal codebases i was actively building and auditing, rather than some abstract synthetic test. the most immediate shift i noticed is in how the model approaches problems. before, it felt like an over-eager junior dev rushing to say "done" blindly guessing fixes, dumping massive logs into context, or worse, silently weakening/skipping test assertions just to get” [source](https://www.reddit.com/r/google_antigravity/comments/1wkfu6k/i_built_an_engineering_harness_to_stop/patr6ue/)
- Praise, 2026-09-17, r/google_antigravity (Reddit): “systems engineer/old old coder like op. tbh, antigravity has been my workhorse & mvp — i have a lot of adversarial code reviews to make up for the one failing—the code usually is a little buggy, but they‘re usually pretty obvious and i don’t find many heisenbugs. so, internal code reviews first—and make it loop until it passes, then openrouter for red-team & true adversarial code reviews —deepseek & thinking labs inkling have been really good fo” [source](https://www.reddit.com/r/google_antigravity/comments/1whvudv/people_who_are_complaining_about_gemini/pag2cmz/)
- Complaint, 2026-09-25, r/google_antigravity (Reddit): “<strict_link> hey, guys. i made a super optimized multi-agent skill, it consumes 6 times less than the standard teamwork on antigravity. any input is welcomed since ai sometimes lies on testing. thx” [source](https://www.reddit.com/r/google_antigravity/comments/1wppu90/caveman_multi_agent_efficiency/)
- Complaint, 2026-09-19, r/google_antigravity (Reddit): “i think those are different things, the workflow to get to solution and implementing is ok, but the gymnastic to check if the solution is right is because you dont know what is right and that limits you to what the agent can do, and usually are way to over complicated” [source](https://www.reddit.com/r/google_antigravity/comments/1whvudv/people_who_are_complaining_about_gemini/par7obt/)
- Complaint, 2026-09-14, r/google_antigravity (Reddit): “i was using 3.8 flash and this is basically what i ordered it to do: \- change colors in ui \- do not run \`php artisan\` directly because laravel sail (docker) is running and you should run \`sail artisan\` instead \- do not run unit tests, i will run them myself if needed do you know what it did? at the start it tried running \`php artisan test\` to see if tests pass before starting the implementation, then took about 10-15 minutes making sure” [source](https://www.reddit.com/r/google_antigravity/comments/1wg3igg/i_gave_antigravity_only_three_instructions/)

### OpenCode

- Praise, 2026-09-27, @opencode (X): “@iam_chonchol @opencode self-testing before delivery makes the workflow much more reliable.” [source](https://twitter.com/1082992095361609728/status/2104230124396982531)
- Praise, 2026-09-27, @opencode (X): “@iam_chonchol @opencode testing the game before delivery adds real value.” [source](https://twitter.com/1552600100869853184/status/2104230697162694902)
- Praise, 2026-09-27, @opencode (X): “@iam_chonchol @opencode the ability to iterate after testing is what stands out.” [source](https://twitter.com/2010625521676419072/status/2104232078997086567)
- Complaint, 2026-09-27, r/opencode (Reddit): “not my experience with it. gpt 6 is extremely bad at following instructions and wastes absurd amounts of time testing” [source](https://www.reddit.com/r/opencode/comments/1wq8vp9/best_free_model_after_deepseek_leave/pcbkxpg/)
- Complaint, 2026-09-16, r/opencodeCLI (Reddit): “was simple to add to reasonix. i have it scanning projects for bugs, logic errors, or any other potential issues after running my full test suite. so we'll see what it comes up with. i know of at least 2 critical bugs which i've not yet fixed, so i'll be interested to see if it can find them. getting about 60 t/s. cache hit is under 10-15% which isn't very promising imo. we'll see where these 100m tokens get me. edit: we've consumed 104,539 token” [source](https://www.reddit.com/r/opencodeCLI/comments/1wh59nh/100m_atria_dawn_preview_api_tokens_for_free_very/pa41lic/)
- Complaint, 2026-09-11, r/opencode (Reddit): “ran a new feature dev task with luna max on codex the other day and it took a full 4.5 hours. yesterday was even worse — glm-5.3 flash on opencode ran for over 7 hours on a single task. from what i’ve noticed, the agent actually writes the code for each issue pretty fast. the real time sink is code review and bug fixing. on my most recent task, the agent spent something like 10x the coding time just running tests and debugging over and over — wen” [source](https://www.reddit.com/r/opencode/comments/1wd7aka/anyone_else_spending_10x_more_time_on_reviewdebug/)

### Amp

- Praise, 2026-09-02, @AmpCode (X): “@genaiupstart @ampcode oh and lots of automated and manual testing too!” [source](https://twitter.com/33135576/status/2095241864815296963)
- Complaint, 2026-09-23, @AmpCode (X): “@purefunctor @ampcode also some settings here <strict_link> (again, very bad and experimental, i haven't tested this much myself yet) <strict_link>” [source](https://twitter.com/784008/status/2102759138732491125)
- Complaint, 2026-09-16, @AmpCode (X): “the thread split is the real trick here, not the phone. the one i would add is a thread whose only job is to check the data the first thread compiled, because that is where these ship wrong and nothing downstream notices. mine wrote a clean site on top of a column it had misread, and the site looked perfect. did you verify the compiled data separately, or trust the build thread to catch it?” [source](https://twitter.com/2259844958/status/2100329483077288435)
- Complaint, 2026-09-03, @AmpCode (X): “@toolmantim @ampcode and by builds, i mean being able to build the solution and run tests as part of verification.” [source](https://twitter.com/317888049/status/2095549145561833965)

### GitHub Copilot

- Praise, 2026-09-22, @GitHubCopilot (X): “the issue was labeled good-first-task. it was not. a rate limiter on a public route used a global map, so two instances doubled the quota and one deploy wiped the counts. that is the job i gave @githubcopilot. not autocomplete. the agent on the issue, in the same @github repo. brief i left on the ticket: keep the existing middleware do not add redis do not invent an api gateway store hits per instance without lying across deploys add a test that” [source](https://twitter.com/1989355273727967232/status/2102450302008095144)
- Complaint, 2026-09-11, r/PiCodingAgent (Reddit): “i don't know about codex, but with github copilot i noticed a lot of double checking and validation steps that in pi don't happen. i would assume that's from some general instructions in the harness.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wbqkel/same_task_same_model_pi_passed_in_90_turns_codex/p96cq93/)
- Complaint, 2026-09-02, r/ChatGPTCoding (Reddit): “the difference mostly comes down to harness autonomy and verification loops rather than the raw model weights. claude code leans heavily into multi-turn bash execution, file patching, and running test commands iteratively until the diff actually passes, which naturally burns 2–3x more tokens per task. copilot caps turn budgets and context assembly more aggressively to keep token spend bounded, but the tradeoff is that it often stops after draftin” [source](https://www.reddit.com/r/ChatGPTCoding/comments/1w4lm3b/claude_code_vs_github_copilot_token_burn/p7dda3c/)

### Pi

- Praise, 2026-09-26, r/PiCodingAgent (Reddit): “also nice edited comments haha, sad to see the worlds sharpest dev not be able to help us. and you can go look at the repository tests or maybe live demo attached before saying no testing is happening lol. would love to see some of your work oh great one. anywho thanks again for superb feedback, i'll go put some of the first effort ever in fixing this so we don't let another one of you extremely valuable swes slip again.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wqe2v1/meet_pig_the_pi_coding_harness_that_is_yours_but/pc9g9jc/)
- Praise, 2026-09-11, r/PiCodingAgent (Reddit): “i don't know about codex, but with github copilot i noticed a lot of double checking and validation steps that in pi don't happen. i would assume that's from some general instructions in the harness.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wbqkel/same_task_same_model_pi_passed_in_90_turns_codex/p96cq93/)

### Cline

- Praise, 2026-09-14, @cline (X): “i got early access to the new @cline open source desktop app! i wanted to see how it would fit into my everyday development workflow, so i tested it with the free glm-5.3 flash model across coding, scheduled reviews, and conversation handoffs. for the coding task, i asked cline to build a small node.js task-list cli with persistence and tests. after resolving an initial environment issue and fixing how the cli handled corrupted json, it successfu” [source](https://twitter.com/1071875988122951682/status/2099545544532144129)
- Complaint, 2026-09-26, @cline (X): “@cline benchmarks are useful, but fixtures still decide whether an agent is safe. a green next.js eval hid a dirty-worktree edit for us. do you publish any tasks with pre-existing changes?” [source](https://twitter.com/1835841692852682752/status/2103657719173718422)

### Factory

- Praise, 2026-09-21, @droid (X): “@droid it is a massive step up for the model, especially with those self-verification capabilities. we actually went deeper on this here: <strict_link>” [source](https://twitter.com/1213502906332110848/status/2102172305392635915)

### Kiro

- Complaint, 2026-09-10, @kirodotdev (X): “@shao__meng @kirodotdev @clare_liguori "people set the direction, and the agent executes." it sounds smooth when we talk about it, but when it actually runs, the bottleneck usually occurs during the verification stage—after the agent modifies the code, it judges right or wrong by itself, which can easily lead to loose testing. do you have any specific guidelines in those ten points on how to set non-bypassable acceptance criteria for the agent?” [source](https://twitter.com/2081519130113630208/status/2098160143506805106)

### Conductor

- Complaint, 2026-09-12, r/conductorbuild (Reddit): “it's probably ai doing it on its own. possibly they don't use it themselves. so ai is coming up with the features and implementing them automatically. there is not enough time/token to write acceptance test that can check regression in future. all ai first company is doing it these days. they skip code review. they skip anything that requires using your brain.” [source](https://www.reddit.com/r/conductorbuild/comments/1wdm8hc/conductor_regressions_are_making_it_harder_to/p9cssa6/)
