# Finding the right files in the codebase (`context.codebase_retrieval`)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/criterion/context.codebase_retrieval

Area: [Instructing and context](https://feedbackbench.com/criteria/context.md)

**Definition.** How the agent searches and indexes a repository and pulls in relevant files. Covers over-reading on trivial edits and missing files in large or multi-repo codebases.

**Boundary.** Not this: see [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md) for consumption not tied to file reading.

Rated author-weeks, all agents: 483. Complaint share: 56%.

## The brief

Written by Claude Opus 5.5 from 78 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**Agents find code by grepping, and users pay for every reread.**

TL;DR:

- Rereading files and broad grep loops are the top complaint, and they show up across nearly every agent.
- Cursor and Claude Code earn the most goodwill. Codex and Google Antigravity draw the sharpest complaints.
- Power users bolt on code graphs and feature maps because native retrieval falls short.

In plain terms: Agents often open far more files than a small edit needs. They reread what they already saw and still miss pieces in big repos. Users who fix this usually add their own map or graph.

### How it breaks

- **Rereading files already in context** ([Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md)). Agents reopen the same files and lines repeatedly, burning requests and tokens without gaining information.
  The pattern repeats across harnesses. One Antigravity user compares it to a copier rescanning the same pages. A Cursor user saw nine searches and fourteen files to locate one variable. Zed users report a file deliberately added to context still has to be read again by a tool before an edit is allowed. Some models in OpenCode reread full files by habit. The common thread is unranked, unremembered search, not a lack of model intelligence.
  Evidence:
  - Complaint, Zed, @zeddotdev, 2026-08-31: “@zeddotdev have you guys fixed the problem where if i deliberately include a file in an agent's context it still has to read the file again with a tool (wasting an extra request) before the edit tool is allowed to work?” [source](https://twitter.com/1600030777130962944/status/2094319263154880990)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-03: “if you look at the example [wentallout](<strict_link>) provided you will see that the issue is that it just keeps re-reading the same files, and the exact same file lines, over and over again. this isn't a human, it doesn't need to re-read something to further analyze it - reading twice is just a waste of time and tokens. this is like a copier machine scanning the same pages of a documents over and over again - the document won't change, the quality of the scan won't change, and the copier already has the information.” [source](https://www.reddit.com/r/google_antigravity/comments/1w62rr4/gemini_38_flash_goes_to_cycle_way_too_often/p7kplbh/)
  - Complaint, Cursor, r/cursor, 2026-09-01: “this tracks with what i've seen too - it's less paranoia and more that the search itself isn't ranked. nine searches and fourteen files for one variable means the agent is grepping broadly and re-reading matches instead of narrowing to the two or three files that own that constant. model choice matters less than how narrow that first pass is.” [source](https://www.reddit.com/r/cursor/comments/1vqj76f/cursor_desperately_seems_to_want_to_spend_my/p75ibkm/)
  - Complaint, OpenCode, @opencode, 2026-09-22: “@supejixi @opencode can you try v2? a lot of stuff is improved, use the install here: <strict_link> regarding mimo, yes it does tend to reread the full file” [source](https://twitter.com/14615235/status/2102403703764332720)

- **Missing files in large repos** ([Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md)). As repositories grow, agents return partial results and users stop trusting the search.
  Users describe searches that miss a few points, forcing a second pass with a different model. One Claude Code user ran a scan that found only 10 of 183 markdown files. Copilot sometimes fails to find a file until the user opens it by hand. OpenCode users say every model degrades faster once a mega repo is attached. Split worktrees in Conductor lose the overall project context. The cost is less the miss itself than the doubt it plants about every later answer.
  Evidence:
  - Complaint, GitHub Copilot, r/GithubCopilot, 2026-09-10: “there are quite some issues with this.... another one i hate: if you link something, continue writing and then go back change normal text **before that linked content**, the link will simply disappear. at other times, it just doesnt find stuff. i need to manually open the file, then it´s suddenly there. i mean, how hard is it to add at least the files? like, ctrl + p is super fast for this.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wcc1c4/i_m_losing_12hrs_of_life_expectancy_everytime_i/p8xgn0o/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-08: “yeah, gemini is by no means a bad model. but comparing it to sol, or opus 5 is insanity. even grok 4.6 performs way better than this. when i ask gemini to search something for me in codebase for me to work in a task, it always misses few points which i again need sol to do it for me. so, if i can't trust the output of the model, then there always be anxiety whether it performed the search well or it didn't hallacunite some result. actually i found glm 5.3 flash better than gemini 3.8 is day to day task.” [source](https://www.reddit.com/r/google_antigravity/comments/1w5e5zn/38_flash_is_13_performance_boost_at_the_same/p8ko242/)
  - Complaint, OpenCode, r/opencode, 2026-09-26: “reminds me of composer but turns but this is better than composer. but also all models need a fresh repo/folder. once you add mega repos all models start to degrade significantly faster. not a guarantee for results but gives you the cleanest shot to not run into errors. the other failure mode is not translating requirements and not actually knowing the stack yourself.” [source](https://www.reddit.com/r/opencode/comments/1won61w/real_life_performance_of_muse_spark_13/pc85m70/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-02: “i tried this on my own repo but it found only 10 of the 183 .md files. claude doesn't add the co-authored line consistently in my repo. however, i noticed the edits came in clusters, several hitting on the same day. a clear batch, and given my habits most likely done by agents. some of these haven't been edited in a while, and i'm not sure if that's because they are stable or are stale. this is a good signal for further investigation, thanks for the tip.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w3iv8i/opus_5_struggled_until_i_realized_my_nested_md/p7ek59d/)

- **Over-reading irrelevant code** ([Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md)). Agents wander into dependencies, build folders and whole workspaces when the task needs two files.
  OpenCode users report models reading node_modules and bin folders to understand a library. Those users now write explicit allow and deny rules to cap the spend. Cline tried to load a 2MB file whole and failed on context. Pi, in the same post, grepped for the relevant slice instead. Copilot users worry they cannot see what context is sent, or whether it is needed at all. Agents read broadly by default, and only manual guardrails narrow them.
  Evidence:
  - Complaint, GitHub Copilot, r/GithubCopilot, 2026-09-11: “we never know what context is given to the agent about the codebase. they can make profit by just having the agent read more and more context that isn't ultimately necessary to perform the task.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wcj4ce/i_cancelled_my_github_copilot_subscription_are/p96nsy2/)
  - Complaint, OpenCode, r/opencode, 2026-09-14: “glm-5.3-flash as planner and orchestrator and ds-v4.1-flash as main subagents model. i have debugger subagents that can try to fix bugs and escalate to the next one if cannot fix, all cheap models. i added strict instruction including opencode.json how to scan the codebase and what allowed read and not to read so any models use less token. i observe that some models read packages in node_modules or bin or .packages folder to understand the library which is a total waste of token. if you dont do this, expect lots of token usage.” [source](https://www.reddit.com/r/opencode/comments/1wb45yp/which_opencode_go_model_do_you_think_is_the_best/p9pdww5/)
  - Complaint, Cline, r/LocalLLaMA, 2026-09-02: “i switched to [pi.dev](<strict_link>) when i started to use qwen 3.8 27b. before that i used qwen 3.6 35-a3b with full 256k context so it wasn't a big deal but now there is no way... another difference is that i tried to give cline a big 2,1mb javascript file asking to analyze one part and it started with "ok, let me read the file" and tried to read the whole thing failing immediatly because 2mb > max context. with [pi.dev](<strict_link>) i didn't even give the file, i told him "there is this url, find how one part works" and it found the js by itself, used grep and other tools to search the relevant part without reading the whole thing.” [source](https://www.reddit.com/r/LocalLLaMA/comments/1vyt117/are_the_best_settings_for_single_3090_just/p7drsf1/)
  - Complaint, OpenCode, r/opencode, 2026-09-17: “one hour with 3 sessions attempting a review that would normally take deepseek flash 10 minutes, a series of failures, attempts to read a workspace it doesn't need (it's nosiness reminds me of astra), failure to write a file, failure to compress context, endless upstream errors. one completed, i told the one failing to write a file to just give me the summary and it spewed stuff in chinese and i cut it off after 10 minutes of that, the other was still ongoing and i just cut it off.” [source](https://www.reddit.com/r/opencode/comments/1wijqfv/union_alpha_is_not_a_model_at_all_it_is_a_2_tier/pac2aoa/)

- **Users build their own maps** ([Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md)). Experienced users replace grep with code graphs, feature maps and cheap scout models, then report large savings.
  This is the clearest sign that native retrieval is not enough. One Claude Code user found the agent reading 18 files per file written. After adding a feature map, it read two. Others wire in code graphs, compiler-backed symbol lookup or structural models over MCP, so the agent queries locations instead of rereading content. Native indexing is among the most common requests. The fixes work, but they are user-built plumbing on top of the agent.
  Evidence:
  - Praise, Claude Code, r/ClaudeCode, 2026-09-09: “the scout/researcher split is the underrated part here. most quota burn comes from big models re-reading whole files, so making cheap models report locations instead of content is a really good contract.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wbc03f/how_i_use_subagents_without_burning_through_fable/p8u1ts5/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-02: “i've been building with opentelemetry to improve my harness setup and skill files. it's crazy how dead spots in your skill files can waste so much time and tokens. i built this little package that records all your sessions locally and lets you query them from claude. one example: it found a part of my code where claude was reading 18 files for every file it wrote. the session analysis caught this, updated some of the skills, added a feature map, and it dropped down to 2 file reads for every write. saving a significant amount of time and tokens. [<strict_link>” [source](https://www.reddit.com/r/ClaudeCode/comments/1w3b38v/weekly_showcase_thread_what_are_you_building_with/p7cnn5r/)
  - Praise, Google Antigravity, r/google_antigravity, 2026-09-22: “one massive thing that's saved me a bunch of time (and tokens) is <strict_link> instead of it searching through a bunch of files it thinks are relevant, or searching for bits of text across everything, it just queries graphify (which has a pre-constructed graph of your codebase), which tells it where all the relevant bits of code are, what it's dependant/depended on, etc someone at my old job showed me, and i'm genuinely so grateful” [source](https://www.reddit.com/r/google_antigravity/comments/1wnfcoo/has_antigravity_started_aggressively_scanning/pbelnb8/)
  - Complaint, Cursor, r/cursor, 2026-09-26: “i’m building enola to make cursor faster on real codebases. cursor is a great tool! but quite some time is still spent understanding the codebase (tracing dependencies, finding the right files, figuring out how services and modules interact). i tackled this problem by building enola. enola builds a structural model and exposes it to cursor through mcp. \[open-source, deterministic\] instead of spending time and input tokens, cursor can ask the existing architectural model. enola also helps cursor validate changes against deterministic rules after the code is written. so it creates a loop: >understand the repository → make the change → check the architectural impact i’m curious how this compares with other people’s cursor workflows? what's your harness? open source: [<strict_link>” [source](https://www.reddit.com/r/cursor/comments/1wqhdb1/enola_i_built_an_opensource_architecture_layer/)

- **Indexing that stalls or is absent** ([Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md)). When indexing exists it sometimes never finishes, and some tools cannot point at a local repo at all.
  A Pi user saw indexing hang before the first turn even started. An Amp user filed a bug after indexing apparently never ran. Cursor users ask why they cannot select a local repository. Cline's codebase search breaks unless a target folder is chosen inside the workspace. Antigravity's desktop app shows no file system, by one user's account. These are basic plumbing failures that block retrieval before any model choice matters.
  Evidence:
  - Complaint, Pi, r/PiCodingAgent, 2026-09-04: “i've used it and it got stuck indexing my codebase. turn didn't even start. i don't need my harness to be yet another effin electron app.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1w72ajc/free_glm_53_flash_for_your_agents/p7sqahm/)
  - Complaint, Amp, @AmpCode, 2026-09-01: “@ampcode hey @ampcode team. i submitted a bug on this. tried what @0xbrettj suggested but it looks like the indexing isn't running? report id: amp_bug_2ase32dwhlnkfqxct84net status: new affected thread: create personal orbstack plugin” [source](https://twitter.com/1535774830225944576/status/2094907022450004165)
  - Complaint, Cursor, @cursor_ai, 2026-09-11: “@cursor_ai @bot why doesn't it support selecting a local repository?” [source](https://twitter.com/1242308237661102080/status/2098381361086357684)
  - Complaint, Cline, @cline, 2026-09-19: “@cline hey @cline, v0.0.32 is great, but workspace folder selection has a critical issue: without selecting a target folder inside the workspace, search_codebase becomes useless and forces absolute paths, breaking context workflow.” [source](https://twitter.com/4161108994/status/2101306814574813242)

### Who stands out

- **Cursor (stronger)**. Users credit Cursor's semantic indexing and context graph for finding its way around repos with less flailing than Claude Code.
  Several posts frame Cursor as the better harness for retrieval. One user runs Claude Code at work and Cursor at home, and says Cursor navigates repos better while Claude overdoes things. Another praises cross-codebase pulls across multiple projects. The goodwill comes with friction. Users ask for local repo selection and multi-repo support, and one reports pricing pushing enterprise teams to build their own graphs.
  Evidence:
  - Praise, Cursor, r/cursor, 2026-09-18: “i agree and it’s because of the semantic indexing and contex graph under the hood but they’re now upcharging enterprise users 70% on anthropic and openai models so it’s no longer feasible for enterprise use. and suddenly we’re all building context graphs and bailing out.” [source](https://www.reddit.com/r/cursor/comments/1wjcyla/got_cursor_start_just_to_see_if_cursor_can_fit_my/pajd0yt/)
  - Praise, Cursor, r/cursor, 2026-09-13: “using claude code at work, cursor at home. i'd go with cursor - i prefer seamless integration with ide. also it seems to be finding its way around the repos much better than claude which is overdoing things a lot. you can "directly" compare them by trying working side by side. use claude as a cursor model (in cursor's native agent window) vs using it as a cli or cc vs code extension. for some reason extension is not using all cursor's capabilities and is much worse than the native cursor's agent/chat. that's why i'm saying model is not as important as the harness and cursor seems to be a good harness simply.” [source](https://www.reddit.com/r/cursor/comments/1wfeim1/claude_code_vs_codex_vs_cursorwhat_are_you/p9mgaze/)
  - Praise, Cursor, r/cursor, 2026-09-03: “tried it for a couple of days and i’m pretty much sold. have to admit, it’s been a lot better for me than i expected. the thing that really sold me was being able to have the agent work across my other codebases and pull relevant code/components into the current one. that’s been surprisingly useful, especially when i’m working across multiple projects. it’s also really nice for workflows where you’re trying to tie a bunch of different tools together. for example, i’ve been using it to bridge comfyui into another laravel project and have it generate the assets/materials i need as part of the workflow. honestly, damn awesome if you ask me.” [source](https://www.reddit.com/r/cursor/comments/1w5nl3p/stop_forcing_me_to_the_agent_view_by_default/p7je2bg/)
  - Complaint, Cursor, r/cursor, 2026-09-01: “this tracks with what i've seen too - it's less paranoia and more that the search itself isn't ranked. nine searches and fourteen files for one variable means the agent is grepping broadly and re-reading matches instead of narrowing to the two or three files that own that constant. model choice matters less than how narrow that first pass is.” [source](https://www.reddit.com/r/cursor/comments/1vqj76f/cursor_desperately_seems_to_want_to_spend_my/p75ibkm/)

- **Claude Code (mixed)**. Claude Code can pinpoint code in huge repos when steered, but users say it greps and rederives by default.
  Praise centres on precision when the user points it well. One user locates specific code in a 2.5GB repo. Others credit features that avoid reading every file. Complaints describe repeated grepping, rederiving the same facts across sessions, and too many reads per write until a feature map is added. Its users file the most requests for built-in native indexing.
  Evidence:
  - Praise, Claude Code, r/ClaudeCode, 2026-09-09: “are you using haiku to save tokens? actually i always talk to claude like a person … no problem for me to pinpoint a specific part in a 2,5gb code base 🤔” [source](https://www.reddit.com/r/ClaudeCode/comments/1wbaxjh/im_sick_of_claude_acting_like_its_a_person/p8s4p6b/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-14: “claude is grepping and rederiving. i can recommend a code graph like graphify for the basis. but a code atlas that knows where the symbols are and how things hold together on a logical level is a nice addendum to that. don't forget to make it autoupdating. making claude do a sweep through sessions and ask what is being rederived over and over again has opened my eyes and resulted in some serious fixes.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wg80fa/does_claude_code_need_some_kind_of_persistent_map/p9s4jn1/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-02: “i've been building with opentelemetry to improve my harness setup and skill files. it's crazy how dead spots in your skill files can waste so much time and tokens. i built this little package that records all your sessions locally and lets you query them from claude. one example: it found a part of my code where claude was reading 18 files for every file it wrote. the session analysis caught this, updated some of the skills, added a feature map, and it dropped down to 2 file reads for every write. saving a significant amount of time and tokens. [<strict_link>” [source](https://www.reddit.com/r/ClaudeCode/comments/1w3b38v/weekly_showcase_thread_what_are_you_building_with/p7cnn5r/)
  - Praise, Claude Code, r/codex, 2026-09-15: “thanks! if it’s okay, can you tell me how many projects you’ve already worked on with codex, how big they were, and which plan you’re using? i’d also like to know how well codex can follow instructions compared to claude code. honestly, cc has many features that help save a huge amount of tokens instead of having to read all the files manually, so i’m interested in knowing how codex handles this as well. i use ai coding tools for both **small and large projects**, so this is pretty important to me.” [source](https://www.reddit.com/r/codex/comments/1wh8kgn/switching_from_claude_code_max_to_codex_advice/pa0ds4n/)

- **OpenAI Codex (weaker)**. Codex users praise deep dives into messy code but say lighter models get lost in bigger repos.
  One user had it find a specific spot in an undocumented 50k-line C++ repo within a minute. Complaints are about consistency. Some models lose their way in large repos. One user says newer models stopped tracing schema relationships the way an earlier one did. Another notes that high-thinking runs sweep the whole repo when asked for a small check. Scale and model choice decide the experience.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-16: “what's wrong with terra? luna gets lost in bigger repos in my experience, terra-medium has been a really good sweet spot for efficiency and speed for me.” [source](https://www.reddit.com/r/codex/comments/1widk1m/tier_list_according_to_me/pa9nox3/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-25: “they’re so bad. gpt 5.5 would start a session my discovering how my postgres schemas were related, it would actually trace pk/fk relationships and then use that discovery to recognize intent, build plans, find opportunities to tighten, etc.” [source](https://www.reddit.com/r/codex/comments/1wpu2b5/this_didnt_age_too_well/pbyte0q/)
  - Praise, OpenAI Codex, r/codex, 2026-09-01: “irrelevant. these tools when adding entire repo's or context will miss things. you as a human can subtract everything and just focus on "the icon". sol has been trained to be a thinker and to keep context. so when you ask it to "simply qa this" sol looks at the entire repo and upstream and downstream consequences of it, and it can get lost etc. that doesn't mean the model is bad. you can easily turn the model thinking down and ask directly to just move it and be damned with the consequences and it will do it. the entire premise of high thinking agents is to consider many things. all commercial ai is now in a proper harness, it's a buzzword people have latched onto. the harness isn't the issue here.” [source](https://www.reddit.com/r/codex/comments/1w42glk/this_is_how_we_know_astra_is_coming_on_thursday/p78wtkl/)
  - Praise, OpenAI Codex, r/codex, 2026-09-14: “llms especially astra are extremely good at navigating shit codebases and explaining to you how it works last day it went through an undocumented uncommented repo of 50k+ lines of c++ code and found the very specific part that i was looking for in a minute you can also just ask it to avoid super long one liners or add a few comments for tricky parts” [source](https://www.reddit.com/r/codex/comments/1wfw93b/i_am_sorry_but_why_does_astra_sometimes_writes/p9poo6k/)

- **Google Antigravity (weaker)**. Antigravity draws complaints that it greps everything, rereads the same lines and still misses results.
  Users describe the model running grep across everything. They report repeated reads of identical file lines and searches that miss points they then rerun elsewhere. Some defend the broad reading as diligent context gathering. Others report big gains only after adding an external code graph. Requests to stop scanning the entire repository cluster here.
  Evidence:
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-07: “thanks for this, hadn’t heard of code graph, 3.8 is running grep crazy on everything so i tried having claude make a map file to see if would help, but this seems like something better i should be using.” [source](https://www.reddit.com/r/google_antigravity/comments/1w9574o/gemini_flash_38_is_wasting_all_my_tokens/p89cwbu/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-03: “if you look at the example [wentallout](<strict_link>) provided you will see that the issue is that it just keeps re-reading the same files, and the exact same file lines, over and over again. this isn't a human, it doesn't need to re-read something to further analyze it - reading twice is just a waste of time and tokens. this is like a copier machine scanning the same pages of a documents over and over again - the document won't change, the quality of the scan won't change, and the copier already has the information.” [source](https://www.reddit.com/r/google_antigravity/comments/1w62rr4/gemini_38_flash_goes_to_cycle_way_too_often/p7kplbh/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-08: “yeah, gemini is by no means a bad model. but comparing it to sol, or opus 5 is insanity. even grok 4.6 performs way better than this. when i ask gemini to search something for me in codebase for me to work in a task, it always misses few points which i again need sol to do it for me. so, if i can't trust the output of the model, then there always be anxiety whether it performed the search well or it didn't hallacunite some result. actually i found glm 5.3 flash better than gemini 3.8 is day to day task.” [source](https://www.reddit.com/r/google_antigravity/comments/1w5e5zn/38_flash_is_13_performance_boost_at_the_same/p8ko242/)
  - Praise, Google Antigravity, r/google_antigravity, 2026-09-03: “i dont see the problem here it mean the model is trying to gather context of the project and not guess this is actually a good thing if u ask me” [source](https://www.reddit.com/r/google_antigravity/comments/1w62rr4/gemini_38_flash_goes_to_cycle_way_too_often/p7jy9dn/)

- **Augment Code (stronger)**. Every Augment Code post here is praise, centred on its context engine and code fetching quality.
  One G2 reviewer singles out its fetching quality across the code database and its speed against competitors. Another user asks Anthropic to build a comparable context engine into its own models. The sample is small. Still, it is the only agent whose retrieval is cited as the reason people want it.
  Evidence:
  - Praise, Augment Code, @augmentcode, 2026-09-08: “@bcherny @addyosmani @anthropicai boris please implement a context engine to your models. better yet take over @augmentcode and become unstoppable. i fucking beg you” [source](https://twitter.com/1795256572102209537/status/2097359764678214092)
  - Praise, Augment Code, G2, 2026-09-10: “q: what problems is the product solving and how is that benefiting you? a: code analyzing, editing, autocompleting, very good for professional developers. good as compared to other software in the market. q: what do you like best about the product? a: it has excellent fetching quality of code database. i could easily integrate it with github. the speed is good as compared to other products in the market. it also edits and autocompletes your code. q: what do you dislike about the product? a: it is an expensive software also i have used it i could feel a little bugs in customer support system.” [source](https://www.g2.com/products/augment-code/reviews/augment-code-review-13433460)

### Fine print

- Augment Code, Devin and smaller agents have too few posts for firm conclusions. Their direction is suggestive only.
- Many posts blame the underlying model rather than the harness, so agent and model effects are partly entangled here.

## Top requests

What users ask to add or change, most asked first. 56 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Author-weeks | Posts | Agents (author-weeks) |
|---|---|---|---|---|
| 1 | Built-in native codebase indexing | 6 | 6 | Claude Code 3, Augment Code 1, Cursor 1, Factory 1 |
| 2 | Multi-repo project context support | 6 | 6 | Cursor 3, Claude Code 2, OpenCode 1 |
| 3 | Faster and better codebase search | 5 | 5 | Claude Code 1, Cline 1, OpenAI Codex 1, Pi 1, Zed 1 |
| 4 | Avoid scanning entire repository | 4 | 4 | Google Antigravity 2, Claude Code 1, Cursor 1 |
| 5 | Projects support for local repositories and files | 4 | 4 | Cursor 2, Google Antigravity 1, Claude Code 1 |
| 6 | Stop rereading files already in context | 4 | 4 | Claude Code 1, OpenAI Codex 1, GitHub Copilot 1, Zed 1 |
| 7 | Automatic codebase understanding without rules files | 3 | 3 | Claude Code 1, OpenAI Codex 1, Cursor 1 |
| 8 | Codebase visualization and documentation tools | 2 | 2 | GitHub Copilot 2 |
| 9 | Fix file and codebase search bugs | 2 | 2 | Cline 1, Cursor 1 |
| 10 | Respect configured code search tools | 2 | 2 | GitHub Copilot 1, OpenCode 1 |
| 11 | Respect gitignore and exclude sensitive files | 2 | 2 | OpenAI Codex 1, Devin 1 |

### 1. Built-in native codebase indexing

- Cursor, 2026-09-23, r/cursor (Reddit): “agree cursor has the best overall integrated agentic workflow. even more so now with grok bot. if only they could make origin codebase easy to use, improve the mobile app, and get a proper frontier model” [source](https://www.reddit.com/r/cursor/comments/1woad6b/grok_47_performance_in_cursor/pbo2fhm/)
- Factory, 2026-09-20, @droid (X): “@droid you guys should work a context solution like fast context, embedded search and such effeciency and cost is a big reason people love alternatives to codex where the subsidization is massive model agnostic + cheaper costs because less time needed to search (aside subagents)” [source](https://twitter.com/1948570504979271680/status/2101651811689968065)
- Augment Code, 2026-09-08, @augmentcode (X): “@bcherny @addyosmani @anthropicai boris please implement a context engine to your models. better yet take over @augmentcode and become unstoppable. i fucking beg you” [source](https://twitter.com/1795256572102209537/status/2097359764678214092)

### 2. Multi-repo project context support

- OpenCode, 2026-09-25, @opencode (X): “@opencode really needs "add dir" or "add to workspace" feature to add folders it can work at once. references is fine, but having the harness view repos across different paths without having to manually make opencode.json is much better.” [source](https://twitter.com/1383806712545562628/status/2103403834756473288)
- Cursor, 2026-09-13, @cursor_ai (X): “@cursor_ai ideas: can we assign multiple repos as context for project, show it as workspace in mobile app and finally, can we please please pleaseee see the usage in mobile app like grok bot?” [source](https://twitter.com/193344639/status/2098982467478950243)
- Cursor, 2026-09-11, @cursor_ai (X): “@cursor_ai @bot it works great! some feedback / ideas: 1. allow collaboration between projects 2. allow me to add new repos to a project after creating it 3. allow me to connect to other systems” [source](https://twitter.com/1459892524999454722/status/2098472575537971596)

### 3. Faster and better codebase search

- Pi, 2026-09-25, r/PiCodingAgent (Reddit): “very large codebase without my filters and sandboxing, tool use his context quickly i.e. one grep on the repo will return 1000s of results” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wps5ha/i_use_ai_models_at_least_4_hours_per_day_and_i/pc0kazw/)
- Zed, 2026-09-15, r/ZedEditor (Reddit): “this is a really high effort post! i bet you’re a really wonderful teammate. 🙂 i’d also love the search behavior you’re describing.” [source](https://www.reddit.com/r/ZedEditor/comments/1t9818k/loving_zed_so_far_heres_what_i_still_miss/p9za30a/)
- OpenAI Codex, 2026-09-01, r/codex (Reddit): “sol thinks way too much and it is hard to change the course if it goes into wrong way. so it takes 2x more time and 2x more tokens to actually get things done than claude…. claude at least is very fast at searching codebase. codex sub agents take forever.” [source](https://www.reddit.com/r/codex/comments/1w4g47s/remember_when_gpt_would_just_do_the_task/p77j6oz/)

### 4. Avoid scanning entire repository

- Cursor, 2026-09-24, @cursor_ai (X): “@spolen23 @cursor_ai selective tools help. overnight bleed is still the full repo in context. persist outside the chat and pull the matching slice per turn.” [source](https://twitter.com/2068373412322451457/status/2103143044354543972)
- Claude Code, 2026-09-20, r/ClaudeCode (Reddit): “i’m on the pro plan. plugin is currently shipped to an exclusive early access list for bug finding but will become publicly available next week. most of the difficult work is done, but i want to streamline updates and new features without like, parsing the entire repository unnecessarily for example” [source](https://www.reddit.com/r/ClaudeCode/comments/1wl5ab7/built_an_audio_plugin_and_shipped_curious_how_to/paw3ca7/)
- Google Antigravity, 2026-09-16, r/google_antigravity (Reddit): “the high load doesn't justify the model checking out every file in the code. slowness is accepted but not correctly looking into correct file could be the cause of high traffic. it's kind of looking into every file of the repo right now” [source](https://www.reddit.com/r/google_antigravity/comments/1wh8h5u/psa_gemini_38_flash_slowerrors/pa7rukj/)

### 5. Projects support for local repositories and files

- Claude Code, 2026-09-23, @ClaudeDevs (X): “@claudedevs i am really like the new projects feature. could you share a timeline for when it might extend to local repositories and files?” [source](https://twitter.com/307326974/status/2102741547011739678)
- Cursor, 2026-09-13, @cursor_ai (X): “@cursor_ai it cannot access local files, it's worthless if it cannot achieve that. i'll keep using grokbot” [source](https://twitter.com/10045342/status/2099061292426285394)
- Cursor, 2026-09-12, @cursor_ai (X): “@fatih @saastrash @cursor_ai @fredrikalindh would love it to be able to see my local repositories/workspaces and skills” [source](https://twitter.com/2213148498/status/2098568285109293365)

### 6. Stop rereading files already in context

- Claude Code, 2026-09-17, r/ClaudeCode (Reddit): “not saying this isn't a viable method, but seems rather tedious and the model in theory should be trained to run a diff on a handful of files if you ask me!” [source](https://www.reddit.com/r/ClaudeCode/comments/1wj601j/help_improving_opus_5_performance/pagv0af/)
- Zed, 2026-08-31, @zeddotdev (X): “@zeddotdev have you guys fixed the problem where if i deliberately include a file in an agent's context it still has to read the file again with a tool (wasting an extra request) before the edit tool is allowed to work?” [source](https://twitter.com/1600030777130962944/status/2094319263154880990)
- OpenAI Codex, 2026-09-14, r/codex (Reddit): “yeah, i think you’re right in the broader sense. my config tweak is really just a workaround for a weak part of the codex harness. i’ve actually been considering trying omp instead of forking codex. one of my projects is a large telegram android client rewrite, so omp is especially interesting because of its tighter context management plus built-in lsp/ast support for java/kotlin. that could cut down a lot of repeated grepping and rereading of hu” [source](https://www.reddit.com/r/codex/comments/1wffvur/codex_system_prompt_still_forces_agents_to_wake/p9oddmk/)

### 7. Automatic codebase understanding without rules files

- OpenAI Codex, 2026-09-01, r/codex (Reddit): “yes i was considering documenting it in .md files but i was hoping there would be a better alternative or ways to 'discover' and 'know' a codebase properly.” [source](https://www.reddit.com/r/codex/comments/1w4rcqr/how_do_you_stop_known_coding_mistakes/p79jvqm/)
- Cursor, 2026-09-07, r/ClaudeCode (Reddit): “cursor *does* support multiple nowadays. i don't like to have to write how some virtual developer should behave or write these "hooks" which now seems almost like some legal form or law text that the ai needs to read an go through, only to fins some smart gap and don't obey in the end. it only slows me down cause it's often not actually needed to run in all context and it's super slow. the ai should figure out how to work by by analyzing the cod” [source](https://www.reddit.com/r/ClaudeCode/comments/1s93y8z/why_do_you_all_prefer_claude_code_over_cursor/p8dtzcw/)

### 8. Codebase visualization and documentation tools

- GitHub Copilot, 2026-09-05, r/GithubCopilot (Reddit): “yes something like cognitions deepwiki would be tremendous. there are a ton of middling products out there but nothing as good as it, and nothing in copilot that works as well. would love that hosted in the github infrastructure too.” [source](https://www.reddit.com/r/GithubCopilot/comments/1w7s69c/feature_request_openwiki/p7ykegv/)
- GitHub Copilot, 2026-09-03, r/GithubCopilot (Reddit): “perhaps some kind of ai-focused graph plotter like graphify (or similar) would help the ai to find these things in such a large codebase?” [source](https://www.reddit.com/r/GithubCopilot/comments/1w6a1ui/the_vscode_harness_has_a_major_flaw_on_large/p7leftq/)

### 9. Fix file and codebase search bugs

- Cursor, 2026-09-21, r/cursor (Reddit): “so the @ file search is broken for me. i want to look up a file in my working directory, a pdf. half of the time it works, and the other half of the time it bugs out and cursor just does not react any more or the desired file simply does not show up in the selection menu. please fix this bug!” [source](https://www.reddit.com/r/cursor/comments/1wm87gz/file_search_broken/)
- Cline, 2026-09-19, @cline (X): “@cline hey @cline, v0.0.32 is great, but workspace folder selection has a critical issue: without selecting a target folder inside the workspace, search_codebase becomes useless and forces absolute paths, breaking context workflow.” [source](https://twitter.com/4161108994/status/2101306814574813242)

### 10. Respect configured code search tools

- OpenCode, 2026-09-25, r/google_antigravity (Reddit): “to reduce token usage, i try to have every agent use codegraph. it has worked really well with opencode and copilot so far, but it seems antigravity isn't using it (it analyzes a lot of files, uses grep to find them,...). i already have the mcp configured, and i've added instructions in my \`agents.md\` file telling \`agy\` to use it and to sync the database every time it updates files: ## codegraph - **mcp usage**: use the codegraph mcp too” [source](https://www.reddit.com/r/google_antigravity/comments/1wppaje/codegraph_antigravity_integration_issues/)

### 11. Respect gitignore and exclude sensitive files

- OpenAI Codex, 2026-09-11, X search: OpenAI Codex, Codex CLI, Codex app (X): “your ai coding assistant might be reading your secrets files, and there's no reliable way to stop it. openai codex still can't exclude sensitive files from context. that's not a bug — it's a governance failure. read more: <strict_link>” [source](https://twitter.com/242644600/status/2098460339687813211)
- Devin, 2026-08-31, @cognition (X): “@da7_tech @cognition why can't devin read gitignore files normally? it would be nice if this was optimized” [source](https://twitter.com/2026153157953519617/status/2094291002429460513)

## Every agent

| Agent | Overall rank | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|---|
| [Cursor](https://feedbackbench.com/agents/cursor.md) | 4 | Typical | 0.523 | 0.496–0.551 | 64 | 36 | 28 |
| [Claude Code](https://feedbackbench.com/agents/claude-code.md) | 1 | Typical | 0.520 | 0.491–0.549 | 149 | 72 | 77 |
| [OpenCode](https://feedbackbench.com/agents/opencode.md) | 3 | Typical | 0.486 | 0.464–0.510 | 34 | 12 | 22 |
| [OpenAI Codex](https://feedbackbench.com/agents/codex.md) | 2 | Typical | 0.481 | 0.451–0.510 | 115 | 42 | 73 |
| [Google Antigravity](https://feedbackbench.com/agents/antigravity.md) | =5 | Typical | 0.477 | 0.451–0.503 | 47 | 15 | 32 |
| [Pi](https://feedbackbench.com/agents/pi.md) | 7 | Too few posts | – | – | 17 | 6 | 11 |
| [Devin](https://feedbackbench.com/agents/devin.md) | =5 | Too few posts | – | – | 13 | 8 | 5 |
| [GitHub Copilot](https://feedbackbench.com/agents/copilot.md) | =8 | Too few posts | – | – | 12 | 4 | 8 |
| [Cline](https://feedbackbench.com/agents/cline.md) | =8 | Too few posts | – | – | 9 | 3 | 6 |
| [Augment Code](https://feedbackbench.com/agents/augment.md) | 17 | Too few posts | – | – | 8 | 8 | 0 |
| [Zed](https://feedbackbench.com/agents/zed.md) | =8 | Too few posts | – | – | 7 | 3 | 4 |
| [Factory](https://feedbackbench.com/agents/factory.md) | =11 | Too few posts | – | – | 4 | 3 | 1 |
| [Amp](https://feedbackbench.com/agents/amp.md) | =11 | Too few posts | – | – | 2 | 1 | 1 |
| [Conductor](https://feedbackbench.com/agents/conductor.md) | 14 | Too few posts | – | – | 2 | 1 | 1 |
| [Kiro](https://feedbackbench.com/agents/kiro.md) | 13 | Too few posts | – | – | 0 | 0 | 0 |
| [Warp](https://feedbackbench.com/agents/warp.md) | 15 | Too few posts | – | – | 0 | 0 | 0 |
| [Grok Build](https://feedbackbench.com/agents/grok-build.md) | 16 | Too few posts | – | – | 0 | 0 | 0 |

## Posts

Receipts rule: The 5 most recent praise and complaint posts per area (first 700 characters) and 3 per criterion (first 450 characters).

### Cursor

- Praise, 2026-09-27, r/cursor (Reddit): “i hear you. i often need to pause it if i'm in a flow state but for day to day stuff i really like it. it seems to get better the larger the code base and the more it can identify patterns. but that my also be my confirmation bias. i'm trying to use less agentic methods so i force myself to know what's going on an the autocomplete is a nice balance for me.” [source](https://www.reddit.com/r/cursor/comments/1wquwo9/autocomplete/pc9th3g/)
- Praise, 2026-09-27, r/cursor (Reddit): “i find that it’s actually pretty good at understanding a codebase accurately. sometimes i feel like opus will take a shortcut or get sidetracked. however, once the understanding is there, opus is a better planner and executor.” [source](https://www.reddit.com/r/cursor/comments/1wrktgc/am_i_the_only_one_who_thinks_grok_47_is_actually/pcdhko1/)
- Praise, 2026-09-27, @cursor_ai (X): “@cursor_ai @spacex @spacexai deep workspace indexing across full repos turns complex codebase refactoring into a simple single-prompt step.” [source](https://twitter.com/2075291394189541376/status/2104183585280328053)
- Complaint, 2026-09-26, r/cursor (Reddit): “i’m building enola to make cursor faster on real codebases. cursor is a great tool! but quite some time is still spent understanding the codebase (tracing dependencies, finding the right files, figuring out how services and modules interact). i tackled this problem by building enola. enola builds a structural model and exposes it to cursor through mcp. \[open-source, deterministic\] instead of spending time and input tokens, cursor can ask the ex” [source](https://www.reddit.com/r/cursor/comments/1wqhdb1/enola_i_built_an_opensource_architecture_layer/)
- Complaint, 2026-09-24, r/cursor (Reddit): “i mean... for 20 bucks, it's just a git app at this point. $100 dollar sub with cc is unlimited opus 5.5 basically. on max. now that webstorm and rider are free for non commercial use, i just use those with cc. i literally have no use for cursor anymore. the features that cursor sold, are no longer features. the whole name "cursor" is behind the actual cursor on screen and the ai behind that. now that we don't really hand write code anymore, and” [source](https://www.reddit.com/r/cursor/comments/1wp0ext/cursor_is_good_as_a_second_hand/pbtjw4g/)
- Complaint, 2026-09-23, r/cursor (Reddit): “agree cursor has the best overall integrated agentic workflow. even more so now with grok bot. if only they could make origin codebase easy to use, improve the mobile app, and get a proper frontier model” [source](https://www.reddit.com/r/cursor/comments/1woad6b/grok_47_performance_in_cursor/pbo2fhm/)

### Claude Code

- Praise, 2026-09-27, r/ClaudeCode (Reddit): “most of that is native to claude code already, but the project-map idea is a good one. going to measure how many searches a session does before spending the 2k tokens on it."” [source](https://www.reddit.com/r/ClaudeCode/comments/1wropen/i_keep_hitting_usage_limits_on_20x_plan_have/pcekif4/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “yeah, derived project-map is one of the biggest savers, i must admit. i analyzed hunderds of sessions across different projects and in 95% of them, each session by default in claude code started with "where am i, what is it about, let's check readme's around" and that was much more expensive than appending budgeted map. i don't see most of these mechanisms being part of claude code (if that was the case, i wouldn't add it, because [intentic.dev](” [source](https://www.reddit.com/r/ClaudeCode/comments/1wropen/i_keep_hitting_usage_limits_on_20x_plan_have/pcenhhu/)
- Praise, 2026-09-25, r/ClaudeCode (Reddit): “got it now. i never hit either the 5h or the weekly limit, and i usually end the week at about 10% to 20% of weekly. i'm at 12% today, switched to the new model yesterday so i'm not noticing significant increase in usage or reduction of my total available usage. i'd still say it would help to understand what kind of work you are doing. i run about 4 agents in parallel on average, hardly ever using teams unless it is really big tasks and then i ca” [source](https://www.reddit.com/r/ClaudeCode/comments/1wps8yx/limits_nerfed_drastically/pbynq7u/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “you need to have good project management skills, [agents.md](<strict_link>), [architecture.md](<strict_link>), etc. most of your tokens are probably being burnt from claude reading the codebase over and over again. i have been using the 10x plan all day on 3 codebases and i maybe get 15% of my weekly limit per day.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrjhyp/are_the_limits_that_good/pcd0ejx/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “it depends what you are doing but you can use claude on the big island too. there are issues but my current approach: i am letting codex burn its full week of tokens this morning. i have codex running now 21 tasks (chats) on behalf of my claude coordinator chat. the claude chat makes a markdown with instructions and codex does the work. then i have an opus 4.6 chat looping continuously doing work. i have to answer questions and tell claude what l” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrnol4/are_multiple_20x_max_plans_allowed/pceu84n/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “i have moved from opencode to claude desktop for a while and i want something plugins to use that would remove the redundant reads and old calls and doesn't let the context flow , is there any good plugins for this in here” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrfag9/is_there_any_pluginsextensions_that_maybe_work/)

### OpenCode

- Praise, 2026-09-27, @opencode (X): “i really like the way @opencode searches for the specifically required tool for the work at hand, without randomly loading everything what's good is that the tool search is so transparent and gives you insight as to how your utilities are working. <strict_link>” [source](https://twitter.com/1383806712545562628/status/2104154103316426920)
- Praise, 2026-09-24, r/opencode (Reddit): “i noticed is that is very good at tool calling something that i noticed in m3 aswell when it first came out(for example, i was making a tool for building, formatting and giving a resume my compilling my c# projects, it was only on the tools folder on omp and opencode, and i didnt even mentioned in any [agents.md](http://agents.md), docs, nothing - m3 found it and used it perfectly, and the tool was there for a while and no model even touched/aske” [source](https://www.reddit.com/r/opencode/comments/1wp4kmi/m31_is_spacebunny/pbtyu3g/)
- Praise, 2026-09-20, r/codex (Reddit): “i just paid $5 directly to deepseek via their api. some people use a 3rd party services but i think those have limitations. the tool i use is opencode which gives a method to work with your repo directly. there are alternatives.” [source](https://www.reddit.com/r/codex/comments/1wld40n/i_spent_last_week_livid_at_how_quickly_i_burnt/pb0mg8h/)
- Complaint, 2026-09-27, r/opencode (Reddit): “i’m building enola to make opencode agents complete tasks faster. opencode is a great tool! but agents still spend quite some time understanding the codebase (tracing dependencies, finding the right files, figuring out how services and modules interact). i tackled this problem by building enola. enola builds a structural model and exposes it to opencode through mcp. \[open-source, deterministic\] instead of spending time and input tokens, agents” [source](https://www.reddit.com/r/opencode/comments/1wrcet3/enola_i_built_an_opensource_architecture_layer/)
- Complaint, 2026-09-26, r/opencode (Reddit): “its definitely not that "fast" for me... the context fills up so fast on this model because it reads too many files. it easily uses up the full 1m context and ends up compacting and drops relevant information. i have never had a model that reads that many files for almost every run ever. it feels like its scanning through every single file in the directory for data collection or something, not saying it is but the kind of behavior felt like it. i” [source](https://www.reddit.com/r/opencode/comments/1wqi4a8/my_honest_opinion_about_spacebunny/pc4kx5h/)
- Complaint, 2026-09-26, r/opencode (Reddit): “reminds me of composer but turns but this is better than composer. but also all models need a fresh repo/folder. once you add mega repos all models start to degrade significantly faster. not a guarantee for results but gives you the cleanest shot to not run into errors. the other failure mode is not translating requirements and not actually knowing the stack yourself.” [source](https://www.reddit.com/r/opencode/comments/1won61w/real_life_performance_of_muse_spark_13/pc85m70/)

### OpenAI Codex

- Praise, 2026-09-27, r/codex (Reddit): “im a literal software engineer and use luna to work on enterprise codebases, it reads through hundreds of files for me, researches for me and helps me prototype. also reads linear tickets and helps me make pr descriptions quickly all the time if you couldn't use it to push something, you are facing what we call a skill issue my friend.” [source](https://www.reddit.com/r/codex/comments/1wr4e20/gpt_6_sol_is_the_new_opus_47/pca29x7/)
- Praise, 2026-09-27, r/codex (Reddit): “except if you have 1 tb of vram, a local model will never be at the level of astra/opus5.5, and then get ready to warm up your computer. as soon as you work in a real code base with a lot of files and an important context to understand, it’s difficult for small models to be so good.” [source](https://www.reddit.com/r/codex/comments/1wrn2uz/gpt_56_sol_completely_nerfed_after_astra_release/pcdzsnt/)
- Praise, 2026-09-26, r/codex (Reddit): “i tried opus 5.5 on claude code and asked it some questions on my application . told me multiple things that were incorrect and barely even looked at the files. codex at the least actually goes and reads what’s going on . maybe i’m using claude code wrong” [source](https://www.reddit.com/r/codex/comments/1wq55sw/opus_wipes_the_floor_with_sol/pc2tzoa/)
- Complaint, 2026-09-27, r/codex (Reddit): “it's not as good at understanding things besides what it is immediately given in the prompt. opus 5.5 and astra are much better at this. but if you know what you're doing and know how to properly frame the technical aspects of what you want it to build, it is quite good and cost effective.” [source](https://www.reddit.com/r/codex/comments/1wrfcmw/sol_6_aint_that_bad/pcd905p/)
- Complaint, 2026-09-26, r/codex (Reddit): “a lot of models love to skim through docs after they feel like they got what they needed. it’s a real problem. glad you found a work around!!” [source](https://www.reddit.com/r/codex/comments/1wqdg6v/warning_astra_6_can_read_instructions_partially/pc3dcpz/)
- Complaint, 2026-09-26, r/codex (Reddit): “not new to astra 6. chatgpt has been using sed to read limited numbers of files and not getting into the rest for a while. it's so frustrating! maybe it needs to run into a custom \`sed\` tool which gives it the result but at the beginning or end warns it about how many other lines are in the file and tells it to consider reading the rest.” [source](https://www.reddit.com/r/codex/comments/1wqdg6v/warning_astra_6_can_read_instructions_partially/pc3egoj/)

### Google Antigravity

- Praise, 2026-09-25, @antigravity (X): “@androidstudio @antigravity byoa is the right abstraction—acp over model lock-in. the win is studio handing agents the project graph, build graph, and emulator context so they stop guessing modules.” [source](https://twitter.com/928838353784512517/status/2103354594017546261)
- Praise, 2026-09-24, r/codex (Reddit): “so what i've been doing is using chatgpt for the planning work and breaking the project into phases. i've found i just need to make sure every phase has its own chat, and i keep a master handoff document that gets updated at the end of each phase. then i use that master document to start the next chat so the context doesn't get completely lost. for the actual repo work i've been using gemini 3.8 flash high through antigravity quite a bit lately.” [source](https://www.reddit.com/r/codex/comments/1wnfc94/is_gemini_38_flash_actually_outperforming_gpt_56/pbrmax4/)
- Praise, 2026-09-22, r/google_antigravity (Reddit): “one massive thing that's saved me a bunch of time (and tokens) is <strict_link> instead of it searching through a bunch of files it thinks are relevant, or searching for bits of text across everything, it just queries graphify (which has a pre-constructed graph of your codebase), which tells it where all the relevant bits of code are, what it's dependant/depended on, etc someone at my old job showed me, and i'm genuinely so grateful” [source](https://www.reddit.com/r/google_antigravity/comments/1wnfcoo/has_antigravity_started_aggressively_scanning/pbelnb8/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “i'm using googleantigravity for my coating work, but i'm struggling to figure out how to separate the files i need to coat from the ones i can't coat but still need to reference. right now, i'm constantly on edge while working, worried that i might accidentally break a file i shouldn't be touching.” [source](https://www.reddit.com/r/google_antigravity/comments/1wra1kv/what_should_i_do_if_there_are_files_i_dont_want/)
- Complaint, 2026-09-27, @antigravity (X): “@silas<phone_number> @antigravity @officiallogank and also the issue where it keeps telling you folders don't exist, which are the project folder.” [source](https://twitter.com/1222023926123040768/status/2104049296626565207)
- Complaint, 2026-09-23, @antigravity (X): “@bulstherock @antigravity a review for the 2.0 - there is a bottleneck that is missing from you unlike the @cluadeai in vscode, feature - been able open/parse/extract html file(not just the code...)(/weblink data on the background and then use the data inside the agent, 2.0 is missing that” [source](https://twitter.com/2050594673530462208/status/2102765451444986157)

### Pi

- Praise, 2026-09-21, @pidotdev (X): “@eliaslumer @pidotdev working very well for me, but i suspect i can optimize which quest docs the model chooses to read. but i think the principle is sound. i'm trying to store cross-quest knowledge inside the project itself. sometimes i do split quests, or ask the agent to look into another quest.” [source](https://twitter.com/1196480401876963334/status/2102031724934783481)
- Praise, 2026-09-15, r/PiCodingAgent (Reddit): “it reads its own code just fine” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wh7i11/simple_orchestrator_for_pi/pa0fxia/)
- Praise, 2026-09-12, @pidotdev (X): “@juan_miqueo tras muchos meses “retorciendo” claude code para trabajar con modelos locales @pidotdev ha sido todo un descubrimiento: optimizas solo las herramientas que necesitas y eso *reduce el contexto*, que ahí está la clave. ⚡️pi+qwen3.8-flash-next{medium|xhigh} en 1x asus gx10” [source](https://twitter.com/108603561/status/2098676958510932200)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “check out cortexkit i have been running their complete suite for the past week and i forgot about the concept of context and compaction. i have a nice workflow built up around and preceding the cortexkit suite that provides durable context but.. i will say this, after this week trial i have it i am keeping it on both machines aft => replaces pis 4 tools with its own + 3 more that all hinge in the lsp, semantic search embeddings, codebase index” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrq50c/new_session_handoff_vs_compact_which_do_you_prefer/pcepylh/)
- Complaint, 2026-09-25, r/PiCodingAgent (Reddit): “very large codebase without my filters and sandboxing, tool use his context quickly i.e. one grep on the repo will return 1000s of results” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wps5ha/i_use_ai_models_at_least_4_hours_per_day_and_i/pc0kazw/)
- Complaint, 2026-09-24, r/PiCodingAgent (Reddit): “hi all, basically what the title says.. i like using @ to add files/folders to the context but for any of these added to .gitignore it stops working.. any way around this?” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wp8cos/cant_reference_gitignore_filesfolders_with/)

### Devin

- Praise, 2026-09-18, r/CognitionLabs (Reddit): “the new swe-2 model (which is free until oct 16th) performs like sonnet 5 ish, i'd say. i haven't used cursor in a while, but devin's knowledge of your code is the best i've seen. it quickly and intelligently finds relevant files and data to look at before making a plan. rate limits with frontier models get hit fast on the cheap plan, but that's the same everywhere. i like using a frontier model to make plans and write tickets, then have swe-2 d” [source](https://www.reddit.com/r/CognitionLabs/comments/1wdgx7o/devin_advantages_over_cursor/paj0zkt/)
- Praise, 2026-09-17, @cognition (X): “@cognition codebase-wide context is the real unlock” [source](https://twitter.com/1513567206352764929/status/2100522416413819180)
- Praise, 2026-09-16, @DevinAI (X): “@jaredpalmer @devinai i like that you can just tell devin what you want it to look for and let it go through the whole codebase” [source](https://twitter.com/1448193839424884739/status/2100290634276078043)
- Complaint, 2026-09-16, @DevinAI (X): “@househackerjon @muse @devinai rad thanks, i’ll dive into it. but can it not see all my files on my computer and obsidian base etc? i tried to get it to and it couldn’t but then i didn’t have time to troubleshoot” [source](https://twitter.com/1309631099551666176/status/2100051028125286910)
- Complaint, 2026-09-06, G2 (G2): “q: what problems is the product solving and how is that benefiting you? a: its integration with the code saves time, and it provides references to help solve issues. it can also update the code directly, and eventually this saves time, which reduces the overall development pricing so it is worth the money spent. also the pricing makes it affordable. q: what do you like best about the product? a: it provides better solutions and helps me resolve b” [source](https://www.g2.com/products/devin-ai/reviews/devin-ai-review-13416427)
- Complaint, 2026-09-04, r/CognitionLabs (Reddit): “one problem i’ve noticed is that there can be stale or outdated information inside the repository itself — for example old test fixtures, mock data, constants, comments, or tests that still reflect previous behaviour. this infuriates me as it wastes token and makes me argue with my agent non stop. has anyone dealt with this problem in a large codebase? how do you make coding agents distinguish between: current production behaviour authoritativ” [source](https://www.reddit.com/r/CognitionLabs/comments/1w751u9/hi_peeps_how_do_you_prevent_stale_test_datacode/)

### GitHub Copilot

- Praise, 2026-09-26, r/ExperiencedDevs (Reddit): “i generally agree that a good developer would code review something like this. but not necessarily by reading every line. we in fact also didn’t read every line before ai. that being said i don’t think giving someone a code review in a short interview is a good idea unless it’s reasonably simple because that requires context. it would be better as a take home although i don’t know if those still exist. real life that code review is somehow ai a” [source](https://www.reddit.com/r/ExperiencedDevs/comments/1wql3g2/interviewed_candidates_for_ai_engineer_roles_this/pc7pcqf/)
- Praise, 2026-09-24, r/GithubCopilot (Reddit): “it is a random benchmark by a random person and doesn't really mean anything. from my experience, copilot is top notch. better results and less token usage than pi due to codebase indexing and faster searches” [source](https://www.reddit.com/r/GithubCopilot/comments/1wnyinf/cross_harness_benchmark_and_copilot_is_behind/pbqcdk0/)
- Praise, 2026-09-08, r/GithubCopilot (Reddit): “i use claude code, codex as well as ghcp, all in vscode and cli. here's my personal comparison: * user experience vscode: ghcp > claude code > codex * user experience cli : claude code > ghcp > codex * overall performance of the biggest frontier models: doesn't matter, they all do well everywhere * overall performance of the cheap models: codex > ghcp > claude code * multi-root repositories work: basically only ghcp handles this well * exp” [source](https://www.reddit.com/r/GithubCopilot/comments/1wambt6/switched_to_claude_code/p8jjdng/)
- Complaint, 2026-09-17, r/AI_Agents (Reddit): “i'm not going to say that you're dumb, but they are not wrong. their code is likely much cleaner, better, far less mistakes or security flaws, than yours is if you're using just one pass with a single agent. multi-agent workflow has increased my efficiency and quality dramatically. sol (chat): the supervisor/orchestrator. planning and strategy. claude sonnet 5 (chat): adversarial review as needed cc (sonnet 5 on extra): main coding agent code” [source](https://www.reddit.com/r/AI_Agents/comments/1wiyhpz/dont_understand_why_everyone_want_to_have/pafzguj/)
- Complaint, 2026-09-13, r/GithubCopilot (Reddit): “it doesn’t scan “everything” but it is still doing too much. i’ll have to look at agents.md. i have claude.md set up.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wf8lkg/observations_on_claude_code_vs_ghcp/p9l7yll/)
- Complaint, 2026-09-12, r/cursor (Reddit): “i've been looking for an alternative for some time (really tried to walk away from elon...). i've tried vs code + various cli's or extensions, to get anything comparable to cursor flow: \- claude code \- cline \- copilot + codex + claude + omniroute (to several providers) \- opencode \- devin \- trae \- orca \- kilo code \- zoo code \- hermes \- codex i've no idea how cursor does it but all the other agents are like children in a for” [source](https://www.reddit.com/r/cursor/comments/1wc1hxp/any_alternative_to_cursor/p9c0wl3/)

### Cline

- Praise, 2026-09-26, r/CLine (Reddit): “i’m building **music\_switcher**, a python-based desktop app that switches music based on what the user is currently doing. since the app already runs locally in python and interacts with macos through applescript, cline pointed out that sqlite fits naturally: it’s built into python, doesn’t require a separate database server, and stores everything in one local file. the biggest convenience is cline gains access to my files, my commits, my term” [source](https://www.reddit.com/r/CLine/comments/1wr0d1o/used_cline_to_choose_a_database_for_my_existing/)
- Praise, 2026-09-14, @cline (X): “@cline open source coding agent @yashwant_eren it understood the entire codebase” [source](https://twitter.com/362498492/status/2099545376156229643)
- Praise, 2026-08-31, r/CLine (Reddit): “i liked it. it was the reason i used cline. it is also way better than claude at actually going through and understanding the code instead of making surface level assumptions that are often wrong. i like knowing exactly what coding agents do and follow every step. i catch changes that shouldn’t be in the code this way. it is why others are ai coding wrecking balls and i’m typically not. i’ve had so many issues with the updates. i want to scream.” [source](https://www.reddit.com/r/CLine/comments/1vrpayp/we_are_deprecating_focus_chain_in_cline_heres_why/p6yr1ww/)
- Complaint, 2026-09-25, r/LocalLLaMA (Reddit): “what ide to use for local models hi people, i am looking for a lightweight ide or plugin that won't inject large context at initiation. i tried cline and native vs code but they inject such heavy initial context that it fills up my gpu and either goes oom or spend most of my time compacting. the only one i found modestly successful was continue.dev plugin but it needs constant approvals. my use case is to demo/try "autopilot" agent coding. than” [source](https://www.reddit.com/r/LocalLLaMA/comments/1wq9ivr/what_ide_to_use_for_local_models/)
- Complaint, 2026-09-23, r/CLine (Reddit): “three of us are building a small stock draft game in parallel lanes. i own the rules engine, one teammate owns the config form, and another owns the gameplay screens. each of us builds our lane with cline, and the lanes share a typed contract. a couple of weeks ago i fixed a rounding bug in the engine. the per-pick budget used float division and then a floor, which silently drops a cent on values like $1.14 split over two picks. i moved the math” [source](https://www.reddit.com/r/CLine/comments/1woh4cg/how_do_you_stop_parallel_cline_sessions_from/)
- Complaint, 2026-09-19, @cline (X): “@cline hey @cline, v0.0.32 is great, but workspace folder selection has a critical issue: without selecting a target folder inside the workspace, search_codebase becomes useless and forces absolute paths, breaking context workflow.” [source](https://twitter.com/4161108994/status/2101306814574813242)

### Augment Code

- Praise, 2026-09-26, G2 (G2): “q: what problems is the product solving and how is that benefiting you? a: it completely eliminates the tedious manual effort of navigating unfamiliar codebases, writing repetitive boilerplate, and tracing cross-field dependencies saving me 1-2 hours of grunt work every day. q: what do you like best about the product? a: the context engine is really amazing. in contrast to a simple ai autocomplete function, it creates an index of all of your mult” [source](https://www.g2.com/products/augment-code/reviews/augment-code-review-13618001)
- Praise, 2026-09-23, G2 (G2): “q: what problems is the product solving and how is that benefiting you? a: it takes away all the pain of digging through massive codebases manually tracking down cross-file dependencies. i save about 1-2 hours of grunt work every single day just in refactoring and boilerplate. q: what do you like best about the product? a: the context is really amazing. in contrast to a simple ai autocomplete function, it creates an index of all of your multi-rep” [source](https://www.g2.com/products/augment-code/reviews/augment-code-review-13582326)
- Praise, 2026-09-13, r/ClaudeCode (Reddit): “glm 5.3 and glm 5.3 flash are no go for me, i tried their promo weekend tokens and found them underwhelming in more ways than i can tolerate. most importantly, they embedded a lot of confidently incomplete information in designs which failed audits and got blocked during implementation. (i used augmentcode before switching to claudecode, then after a month with claudecode i went back to augmentcode, again this was about a year ago and i found au” [source](https://www.reddit.com/r/ClaudeCode/comments/1wf9uuf/wth_is_going_on_with_claude_usage_limits/p9lym97/)

### Zed

- Praise, 2026-09-12, @zeddotdev (X): “@zeddotdev love the call hierarchy addition — huge for navigating codebases in zed!” [source](https://twitter.com/2010658787611619328/status/2098628406401212721)
- Praise, 2026-09-11, @zeddotdev (X): “@zeddotdev huge win for navigating codebases call hierarchy for incoming and outgoing calls makes tracing logic in zed so much faster!” [source](https://twitter.com/2008812694628175872/status/2098490192592282056)
- Praise, 2026-09-06, @zeddotdev (X): “i use scatchpad for a fast ide @zeddotdev and i set it to my dev root folder. whenever working with the agent apps/clis now when i want to look at the code files it’s a simple super + s and fast zed lookup search for the files or folders or whatever component i want to read. optimal dev ux imo!” [source](https://twitter.com/1729375010/status/2096603934555074581)
- Complaint, 2026-09-25, r/ZedEditor (Reddit): “look, i keep trying it again from time to time, but find all references and global find are nowhere near as good as vs code. you can't easily jump to matches without having to switch tabs (it makes you edit right there inline, which doesn't give you enough context), you can't x matches out, etc. there's no persistent errors panel you can use to jump to errors. there's technically the "outline panel" but it's finicky for those kind of things. plus” [source](https://www.reddit.com/r/ZedEditor/comments/1wohfe8/zed_the_new_ide_to_rule_them_all/pbz9xy1/)
- Complaint, 2026-09-24, r/ZedEditor (Reddit): “my experience testing zed was great overall, but i ran into recurring issues with the global search (ctrl+shift+f) that made me stop using it. in unversioned repositories, it simply fails to search across all files; i have to open a file before it gets included in the index. i don't recall if it worked well in versioned repositories—i believe it did—but for me, this is a feature that needs to work in every scenario. i still test it occasionally—m” [source](https://www.reddit.com/r/ZedEditor/comments/1wohfe8/zed_the_new_ide_to_rule_them_all/pbs5v7c/)
- Complaint, 2026-08-31, r/ZedEditor (Reddit): “so, i have been using zed for my day-to-day tasks for the past one year but i wanted to check the context usage by different code editors for fixing a low effort issue within the same codebase, i compared vscode context usage with zed keeping in mind that they both follow the: \- same user prompt \- same directory access \- same context files (2 same files, each 700 lines approx.) used the same model for each iteration (opencode zen/mimo 2.5)” [source](https://www.reddit.com/r/ZedEditor/comments/1w36jk8/why_context_usage_is_so_high_in_zed/)

### Factory

- Praise, 2026-09-23, @droid (X): “@wattenberger @droid does this for all assigned tasks and my entire codebase. it's definitely my favourite part of the workflow.” [source](https://twitter.com/2070908287978246144/status/2102608593648546083)
- Praise, 2026-09-18, @FactoryAI (X): “@theterrancex @droid @factoryai reposcape v0.1 catching scanner and rust bugs before release is solid local map of how a codebase connects is such a useful first cut” [source](https://twitter.com/1675906158304038912/status/2101082167782834261)
- Praise, 2026-09-17, @FactoryAI (X): “only downside is the 5 hours limit and the price which is pretty fair but still a little for me personally. other than that i can list so many things i love about it. droid is super efficient and often finish tasks faster than most other agent with similar results. i feel like it gets the right context at the right time. it’s pretty amazing. also love the byok, live the fact that ui almost always looks better when done with droid even using th” [source](https://twitter.com/1617212256487411712/status/2100703532785541412)
- Complaint, 2026-09-20, @droid (X): “@droid you guys should work a context solution like fast context, embedded search and such effeciency and cost is a big reason people love alternatives to codex where the subsidization is massive model agnostic + cheaper costs because less time needed to search (aside subagents)” [source](https://twitter.com/1948570504979271680/status/2101651811689968065)

### Amp

- Praise, 2026-09-03, @AmpCode (X): “@mrsanders @ampcode nope, we haven't found it needed with the models, just say the file basename or describe it and it works (and is less keystrokes/words)” [source](https://twitter.com/784008/status/2095594125496406466)
- Complaint, 2026-09-01, @AmpCode (X): “@ampcode hey @ampcode team. i submitted a bug on this. tried what @0xbrettj suggested but it looks like the indexing isn't running? report id: amp_bug_2ase32dwhlnkfqxct84net status: new affected thread: create personal orbstack plugin” [source](https://twitter.com/1535774830225944576/status/2094907022450004165)

### Conductor

- Praise, 2026-09-23, @conductor_build (X): “tysm for adding workspace search 🙏 @conductor_build” [source](https://twitter.com/19673752/status/2102733268223512621)
- Complaint, 2026-09-06, r/conductorbuild (Reddit): “but does that work good? your each separate worktree would not have the full context of the overall project, which can lead to degradation in performance. this happened with me. and also, what about that changes which depend on another change? essentially for which you want to open a pr point to a preceding one.” [source](https://www.reddit.com/r/conductorbuild/comments/1w7ylli/working_on_an_large_feature_from_ideation_to/p849mnn/)
