# Prompt cache hits, misses and invalidation (`limits.prompt_cache`)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/criterion/limits.prompt_cache

Area: [Paying and limits](https://feedbackbench.com/criteria/paying.md)

**Definition.** Whether prompt caching is kept across turns, resumes and provider routing, and how cache reads are priced. Covers cache drops that inflate usage.

**Boundary.** Not this: see [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md) for consumption not attributed to caching.

Rated author-weeks, all agents: 790. Complaint share: 64%.

## The brief

Written by Claude Opus 5.5 from 69 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**Caching makes long sessions cheap until a switch or pause wipes it.**

TL;DR:

- Model, effort and mode switches are the top cache killer, and users feel it on the bill.
- Pi earns praise for hit rates; OpenAI Codex draws complaints about TTLs, switching and runaway misses.
- Claude Code users cheer cheaper cache reads but still pay heavily for cold-cache turns after idle gaps.

In plain terms: When the cache holds, long agent sessions cost a fraction of their raw token count. When it drops after a pause, a model switch or a harness quirk, the full context bills again and quotas vanish fast.

### How it breaks

- **Switching models or effort wipes cache** ([Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md)). Changing model, reasoning effort or agent mode mid-session resets the cache, so the whole context gets reprocessed at full price.
  This is the most consistent trap across agents. Codex users warn each other to pick one model and stick with it. Antigravity users post PSAs that a thinking-level change blows out cache and quota. Router-style model switching draws the same objection on Factory. One OpenCode user saw a large plan reprocessed from scratch on a plan-to-build switch, while Cline just injects a mode-change note. Codex leads requests to preserve cache across switches.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-01: “i would seriously reconsider - the moment you switch chat/model/effort you lose cache and it may cost you more than just continuing with sol, especially on larger projects. check your cache hit rate percentage.” [source](https://www.reddit.com/r/codex/comments/1w3uexn/luna_max_is_underrated/p74qtv3/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-05: “psa: changing a model thinking level mid conversation can blow out the cache and quota” [source](https://www.reddit.com/r/google_antigravity/comments/1w7cnzx/shortcut_to_change_effort_of_model/p7vzfu8/)
  - Praise, Cline, r/opencodeCLI, 2026-09-12: “this is the first thing i noticed. i make a large plan (90k tokens), switch to build and then immediately it starts prompt processing from scratch. the whole cache was invalidated! cline at least just puts another prompt to tell the model it switched modes.” [source](https://www.reddit.com/r/opencodeCLI/comments/1usvuib/cache_invalidation_when_switching_from_plan_to/p9g4kom/)
  - Complaint, Factory, @FactoryAI, 2026-09-26: “the transit station really has no work to do. i have always been curious about the router that @factoryai has been promoting. they have never clearly explained the specific router mechanism, just vaguely mentioning the classification of simple and complex tasks. how can we determine whether a task is complex or not before it starts, and the inevitable cache invalidation caused by switching models midway is a big problem (anyway, the transit station makes money regardless, lol). a previous idea was to train a small ml model for task classification, and then jev came out, but is this model really suitable for the task difficulty classification work? a big question mark needs to be placed on that.” [source](https://twitter.com/1724973133097291776/status/2103781371772809499)

- **Idle gaps turn resumes expensive** ([Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md)). Return to a session after the cache TTL lapses and the next turn rewrites the full context, a cost users say hides in token totals.
  Users describe a tax on stepping away. One Claude Code user's logs show post-gap turns are a tiny share of turns but drive most cache writes and a large slice of the bill. Workarounds are manual: handoff documents, fresh chats, or timed wake-ups to keep the cache warm. Amp users want forced compaction or resume-from-summary to dodge the bust. Kiro draws the same complaint about repaying write tokens.
  Evidence:
  - Complaint, Amp, @AmpCode, 2026-09-07: “sometimes i come back to a thread after a long time and know the cost of resuming due to cache busting will be high. would be neat to be able to force compaction or resume from a summary @ampcode <strict_link>” [source](https://twitter.com/33135576/status/2096802985645134185)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-01: “checked my own logs because i assumed this was a rounding error. turns that come after a gap longer than the cache ttl are about 1% of mine, 895 out of 86,700 across 131 sessions. they account for 63% of my cache writes though. a write costs 20x what reading the same tokens costs on a 1h ttl, so that 1% of turns lands around 17% of the bill. nearer 12% if you are on the 5 minute ttl. the token share hides it completely, those turns are 1.2% of my tokens. i was wrong about the size of it.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w44ph0/cli_hook_for_alert_cache_expiration_via_telegram/p75xe9b/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-07: “are you starting up in a new or old chat session? an old chat session it will have to read the entire chat again and the cache is cleared after a certain period of time so it's using up a ton of tokens remembering what happened. so save handoff documents and start new chats whenever you stop and want to resume” [source](https://www.reddit.com/r/ClaudeCode/comments/1w8vi25/stop_posting_about_limits_fix_your_workflow/p89fxq4/)
  - Complaint, Kiro, r/kiroIDE, 2026-09-17: “by letting the cache go stale, implied by you saying continue, you just repaid all our cache write tokens again also” [source](https://www.reddit.com/r/kiroIDE/comments/1whrobc/in_less_than_30_secs_kiro_uses_almost_100_credits/paaymrs/)

- **Unexplained misses drain quotas fast** ([Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md), [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md)). Some users see the cache drop mid-session with no change on their side, and the full-price input burns limits in minutes.
  Posts describe misses that look like bugs, not user error. An OpenCode user logged repeated full-context reads every few requests despite a stable session header. A Codex app user says a missed cache plus a big context loop burned the whole 5-hour limit. A Cursor user reports cache cost suddenly dominating the bill and support pushing back. Pi users note misses with Codex models even with a stable prefix.
  Evidence:
  - Complaint, OpenAI Codex, X search: OpenAI Codex, Codex CLI, Codex app, 2026-09-06: “paid plus user completely blocked because gpt-5.6 luna on the macos codex app is broken. missed the prompt cache, ran a massive context loop, and burned my entire 5-hour limit and all resets instantly. had to purge my ~/.codex/ folder just to stop the bleed. fix this @openai <strict_link>” [source](https://twitter.com/16024319/status/2096603901135114618)
  - Complaint, OpenCode, r/opencode, 2026-09-27: “just looked up the usage logs, and every few requests it just takes in the whole input of 300k tokens, instead of cache-reading. one request later it starts cache-reading again. then 3-5 requests later, it inputs all 300k tokens again. don't have this problem with muse spark 1.3 contributor at all. is there a fix? x-opencode-session header looks the same for every request i sampled. this has to be buggy right? [screen](<strict_link>) 3 cache-misses within 4 minutes. that cost me more than 15 cent. and it's off-peak right now. imagine letting it run over night. i'd wake up without usage” [source](https://www.reddit.com/r/opencode/comments/1wrqxxa/is_cache_reading_faulty_with_deepseek_v41_on/)
  - Complaint, Cursor, @cursor_ai, 2026-09-25: “@cursor_ai whats going on with your support? since a few days my quota is burned like crazy, by using projects. i sent full details to support, explained everything with detailed diagnostics, proving that i am not using more, but the cache cost has suddenly exploded and is over 90% of what is billed and your support gaslights me that i simply should reduce my coding activity. sorry, but that is crazy. you clearly have a cache bug since about 2 days ago. i haven't changed what i am doing, you have.” [source](https://twitter.com/804676521529110528/status/2103389221817860558)
  - Complaint, Pi, @pidotdev, 2026-09-09: “@pidotdev cache miss (with codex models) might happen even if you don't change the prefix order or the model.” [source](https://twitter.com/1944161361350537216/status/2097660849036918801)

- **Harness context shuffling busts the cache** ([Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md)). Plugins, tool toggles and auto-compaction that rewrite the prompt prefix quietly invalidate cache, and users only find out from usage logs.
  Users blame harness behaviour more than the models. Plugins that rearrange context, changing the enabled tool set in Copilot, and Cline's auto-compaction all get named as cache breakers. One LocalLLaMA user says Cline invalidates a large context cache on trivial follow-ups, which older versions did not. Several users now disable compaction rather than risk it.
  Evidence:
  - Complaint, OpenCode, r/PiCodingAgent, 2026-09-10: “the problem is that sometimes those "smart" harness environments nuke your cache hit because they rearrange your context without your knowledge. used to be a ludicrously huge issue for example with some claude code and opencode plugins. remnants from a time models had lower context windows and cheaper costs, so it was worth shuffling the context around.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wcj9bg/who_uses_pi_what_do_you_like_about_it/p8yew26/)
  - Complaint, GitHub Copilot, r/GithubCopilot, 2026-09-25: “do you feel that you have to micro-manage the tool enablements often? my perspective is that users should not generally have to do this- tools can already be lazy-loaded by the harness, so the agent is better at dealing with lots of tools that it used to be. also, changing the enabled set of tools can break the prompt cache in a session, making that feature sort of a footgun. another way to manage the set of tools in a session is by creating a custom agent definition with a tool set, then picking the agent based on the task in a session.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wpctlc/vs_code_chat_users_local_or_copilot_harness/pc043kk/)
  - Complaint, Cline, r/LocalLLaMA, 2026-09-16: “is there a coding harness that will not invalidate the entire context cache at each prompt? as the title says. i don't think it's normal to invalidate 100k of context cache for no reason. for example cline and qwen code addon for vs code is doing it most times. let's say you write a big codebase then you ask something like "please update the md file" or some very small change of a simple script and it will trigger the invalidation of the entire context and then you must wait like 40 seconds. for some reason with some older versions of cline and qwen i was able to code for a very long time without triggering useless cache invalidation.” [source](https://www.reddit.com/r/LocalLLaMA/comments/1wi17z6/is_there_a_coding_harness_that_will_not/)
  - Complaint, Cline, r/CLine, 2026-09-09: “it might be that compacting causes cline to use /newtask, and the compaction and /newtask gets repeated since it’s now in context. you could find out by adding a pretooluse hook to log tool calls. i’m afraid of cline’s auto compact hosing cache hits so i don’t use it, but i’m also not giving it a single task that would take 12 hours and/or 60m tokens. the other method to find out what’s going on is to let the model have a go at interpreting/explaining your session history files. those combined with your engine’s requests log can show you a lot about what is happening with cline’s context management.” [source](https://www.reddit.com/r/CLine/comments/1wbdpuy/context_window_management/p8tkdh8/)

- **Cache write pricing hides the trade** ([Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md)). Cheaper reads get the headlines, but users point out that write multipliers and longer TTLs shift cost rather than remove it.
  Cache economics are a real point of debate. Claude Code users note that a longer TTL makes every write pricier, so fewer writes do not automatically mean a smaller bill. Copilot users flag a write surcharge that was not there before. Devin users say heavy caching makes the first request hurt and the long run cheap. Requests for cheaper cached tokens and billing fixes cluster on Claude Code.
  Evidence:
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-20: “>what the post leaves out is the price. a 1h cache write bills at 2× base tokens; a 5m write at 1.25×; a read at 0.1×. so switching to 1h doesn't just recover lost writes — it makes every write 60% more expensive. that's the trade the "75% fewer cache writes" headline hides.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wkpm2i/claude_code_subagents_have_a_5m_prompt_cache_long/paxen7x/)
  - Complaint, GitHub Copilot, r/GithubCopilot, 2026-09-19: “but you are clearly ignoring the 1.25x cache write cost you have that wasn't there before. cost is still higher” [source](https://www.reddit.com/r/GithubCopilot/comments/1wk7lky/gpt_54_deprecation_will_legacy_yearly_subscribers/paqn25o/)
  - Complaint, Devin, @cognition, 2026-09-01: “@cognition 95% cached tokens turns model pricing into a startup tax: the first request hurts, then the long run gets cheap.” [source](https://twitter.com/1067135083155464194/status/2094901820862738844)
  - Praise, Claude Code, @ClaudeDevs, 2026-09-25: “@claudedevs the cache line is the real cut. $0.50 to $0.20 per million tokens is 60% off. in long agent sessions, cache reads are most of the tokens. same session, about 31% cheaper, not 20%. price per task is the metric that matters. cheap long context is what makes agents pay.” [source](https://twitter.com/1457747069662384130/status/2103550252435357826)

### Who stands out

- **Pi (stronger)**. Pi is the agent users hold up as the cache-friendly harness, with extensions built specifically to keep the prefix stable.
  Users credit Pi with very high hit rates that stretch tight weekly limits. The ecosystem leans in: a TTL countdown with expiry alerts, dynamic tools that do not drop the cache, and time injection designed not to break it. Some move from OpenCode to Pi for exactly this. Complaints exist, including effort changes busting cache and misses with Codex models, and a few question whether hit rate is even a useful metric.
  Evidence:
  - Praise, Pi, @pidotdev, 2026-09-06: “at this point, the only thing keeping astra from burning through my weekly limit in half an hour is @pidotdev's incredible cache hit rate: 99.6% 🚀 also my custom pi extension that show cache ttl countdown timers and ping me via push notification 5m before expiry help a lot 🤓” [source](https://twitter.com/2915475819/status/2096584453376102561)
  - Praise, Pi, @pidotdev, 2026-09-20: “@pidotdev dynamic tools without dropping the kv cache is the actual 0.86 win” [source](https://twitter.com/763249944056565760/status/2101499733789618305)
  - Praise, Pi, r/PiCodingAgent, 2026-09-19: “pi-fabric: it removes all tools from the agent and exposes them as functions in either a typescript or python sandbox (you decide which) that is their sole tool. i find with small local host models they will use tools correctly more often since its in an environment they were trained to operate in. plus, you can then instruct the agent to save execution blocks for predictable reuse later. only other one is pi-time-sense which injects the date and time into the context periodically as you chat with the ai. its done in such a way that it doesn’t destroy your cache” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wjtm27/how_to_use_pi_as_a_better_cc/pasx1pp/)
  - Complaint, Pi, r/PiCodingAgent, 2026-09-06: “it seems like a potentially bad metric to track. a high cache rate would be achieved if you chain a bunch of unrelated work in the same session. previous unrelated work in the session would count as a cache hit (i think) though it provides no value. in that case doing the unrelated work in separate sessions would be more efficient. at least, thats what i've been thinking. i'd love to be corrected (cunningham's law in effect).” [source](https://www.reddit.com/r/PiCodingAgent/comments/1w8ju8c/this_is_why_i_love_pi/p83b785/)

- **OpenAI Codex (weaker)**. Codex users treat the cache as fragile, planning around TTLs and switches to avoid being drained by full-context resends.
  Complaints dominate. Users describe a roughly 30-minute TTL that punishes stepping away, model and effort switches that wipe cache, and orchestration that polls an expensive parent just to keep it warm. Codex tops requests to preserve cache across switches and for subagents and background waits. Some relief shows up, with users welcoming a change that lets reasoning shift without wiping the cache.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-12: “yeah but you better be damn sure that you go back to astra before the cache expires or don't touch the session ever again else you'll get seriously drained by sending that huge context back as input tokens.” [source](https://www.reddit.com/r/codex/comments/1wdvfb8/how_to_get_astra_subagents_to_be_less_usage_heavy/p99gnhn/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-08: “i checked the docs, and the \~30-minute prompt-cache ttl has also been discussed here a lot. astra is on the v2 path, so there’s no good reason to wake an expensive parent every 30 seconds and reprocess \~150k tokens just to keep the cache warm. a \~20–25 minute wait would stay safely inside the ttl while eliminating almost all of that polling overhead. the 30-second loop looks like an orchestration bug 🤷🏻♂️” [source](https://www.reddit.com/r/codex/comments/1wa9c9d/i_investigated_why_gpt6_astra_burns_quota_so_fast/p8gmtha/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-05: “nono, not context. it's 2 separate things, context is like your conversation context while cache or whatever the name is like about the project itself. that's why you pick one model and stick with it, yes it's super expensive but the first message on any project it's much more expensive.” [source](https://www.reddit.com/r/codex/comments/1w7e6i2/astra_cost/p80a8f5/)
  - Praise, OpenAI Codex, r/codex, 2026-09-23: “honestly, the slight bump in capabilities is nice. but i'm more excited for the lower factual error rates and the caching changes that allow for changing reasoning without wiping the cache out” [source](https://www.reddit.com/r/codex/comments/1wo7z6p/gpt_6_sol_price_decrease/pbl8hc8/)

- **Claude Code (mixed)**. Claude Code users celebrate the cache-read price cut for long agentic loops, yet still fight TTL expiry and lost cache on model downshifts.
  Praise centres on cheaper cache reads, which users say matter more than headline price cuts because agent loops reread the same context constantly. The friction is expiry: cold-cache turns after gaps, timed wake-ups to refresh, and losing cache when dropping to a smaller model. Claude Code leads requests for cache hit visibility, longer retention and warnings before expiry.
  Evidence:
  - Praise, Claude Code, @ClaudeDevs, 2026-09-25: “@claudedevs the cache read discount is the sleeper here. agentic loops reread the same context constantly, so 60% cheaper cache reads changes the math more than the headline 20%. cheaper retries make for braver agents” [source](https://twitter.com/2102025856541499392/status/2103554830677750084)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-01: “checked my own logs because i assumed this was a rounding error. turns that come after a gap longer than the cache ttl are about 1% of mine, 895 out of 86,700 across 131 sessions. they account for 63% of my cache writes though. a write costs 20x what reading the same tokens costs on a 1h ttl, so that 1% of turns lands around 17% of the bill. nearer 12% if you are on the 5 minute ttl. the token share hides it completely, those turns are 1.2% of my tokens. i was wrong about the size of it.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w44ph0/cli_hook_for_alert_cache_expiration_via_telegram/p75xe9b/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-11: “i set wake ups wake ups every 58 minutes to refresh the cache on my long running runs. i find it helps” [source](https://www.reddit.com/r/ClaudeCode/comments/1wdklro/did_usage_change_this_week_20x_max_plan/p97s09w/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-08-31: “honestly, the real problem is that neither haiku or sonnet makes sense to use in my opinion. you don't gain much more usage falling down to them, and you loose cache. if haiku was as smart as sonnet and price of luna the arsenal of models in the claude family would be complete. and been honest if you have "easy mechanical tasks" you should just leverage opus to create a program that does that and follows unix philosophy instead of hoping a small llm would not fail doing a task.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w2zw44/how_come_has_claude_not_only_never_given_a/p6wpmgv/)

- **OpenCode (mixed)**. OpenCode results swing with harness version and provider, from near-total cache reuse to repeated full-context reads.
  Users report excellent hit rates on cheap providers, making huge token counts cost little. They also say the newer harness caches better than 1.x and that plugins rearranging context hurt. The sharpest complaint is a session that periodically reads the whole context at full price. OpenCode leads requests to fix cache misses.
  Evidence:
  - Praise, OpenCode, r/opencode, 2026-09-26: “opencode 2 or 1.x? opencode 2 has better cache hit. it's entirely dependent on the harness you use.” [source](https://www.reddit.com/r/opencode/comments/1wqox6a/cheepseek_has_its_price_cache_hit_ratio/pc6u7tl/)
  - Complaint, OpenCode, r/opencode, 2026-09-27: “just looked up the usage logs, and every few requests it just takes in the whole input of 300k tokens, instead of cache-reading. one request later it starts cache-reading again. then 3-5 requests later, it inputs all 300k tokens again. don't have this problem with muse spark 1.3 contributor at all. is there a fix? x-opencode-session header looks the same for every request i sampled. this has to be buggy right? [screen](<strict_link>) 3 cache-misses within 4 minutes. that cost me more than 15 cent. and it's off-peak right now. imagine letting it run over night. i'd wake up without usage” [source](https://www.reddit.com/r/opencode/comments/1wrqxxa/is_cache_reading_faulty_with_deepseek_v41_on/)
  - Praise, OpenCode, r/opencode, 2026-09-11: “for me, deepseek is better for one important reason: the price. for sessions with a 97% cache hit rate, deepseek is far cheaper—roughly 50% of the final cost.” [source](https://www.reddit.com/r/opencode/comments/1wd8rn5/is_deepseek_41_flash_not_as_good_as_glm_53_kimi/p9435si/)
  - Praise, OpenCode, @opencode, 2026-09-18: “my tracker estimates about $65 of direct-api-equivalent value for the deepseek traffic so far. that’s 6b+ processed tokens, not 6b unique tokens, with ~97.6% of the input coming from cache reads/reused context. actual out-of-pocket is around $50/month across the subscription providers. that cache rate is why the token count looks so crazy relative to the cost.” [source](https://twitter.com/43050596/status/2100897455743082537)

### Fine print

- Most agents have too few posts to rank here; Cursor, Devin and others rest on a handful of posts.
- Many posts discuss the underlying model provider's cache pricing rather than the agent harness itself.
- Cache costs blur into general consumption; spend not tied to caching belongs under burn rate.

## Top requests

What users ask to add or change, most asked first. 121 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Author-weeks | Posts | Agents (author-weeks) |
|---|---|---|---|---|
| 1 | Fix prompt cache misses and broken caching | 23 | 25 | OpenCode 8, Claude Code 6, OpenAI Codex 6, Cursor 2, Pi 1 |
| 2 | Cache hit and miss usage visibility | 17 | 17 | Claude Code 8, OpenCode 3, OpenAI Codex 2, Pi 2, Cursor 1, Devin 1 |
| 3 | Longer prompt cache retention window | 12 | 12 | Claude Code 6, OpenAI Codex 3, Pi 2, OpenCode 1 |
| 4 | Preserve cache across model or effort switches | 10 | 11 | OpenAI Codex 7, Claude Code 1, OpenCode 1, Pi 1 |
| 5 | Warning before cache expiry or cache-miss cost | 10 | 11 | Claude Code 6, OpenAI Codex 2, Pi 2 |
| 6 | Cheaper cached token pricing | 9 | 10 | Claude Code 4, OpenAI Codex 3, Cursor 2 |
| 7 | Keep cache warm across resumes and sessions | 9 | 9 | Claude Code 5, OpenAI Codex 2, Amp 1, Devin 1 |
| 8 | Preserve cache across compaction and context changes | 7 | 7 | Claude Code 3, Cline 1, OpenAI Codex 1, OpenCode 1, Pi 1 |
| 9 | Preserve cache for subagents and background waits | 7 | 7 | OpenAI Codex 5, Claude Code 2 |
| 10 | Built-in token reduction and cache management | 5 | 5 | Claude Code 1, Cline 1, OpenAI Codex 1, Pi 1, Zed 1 |
| 11 | Fix cache billing errors | 3 | 3 | Claude Code 3 |
| 12 | Credit usage lost to cache bugs | 2 | 2 | Claude Code 2 |

### 1. Fix prompt cache misses and broken caching

- OpenCode, 2026-09-27, r/opencode (Reddit): “just looked up the usage logs, and every few requests it just takes in the whole input of 300k tokens, instead of cache-reading. one request later it starts cache-reading again. then 3-5 requests later, it inputs all 300k tokens again. don't have this problem with muse spark 1.3 contributor at all. is there a fix? x-opencode-session header looks the same for every request i sampled. this has to be buggy right?” [source](https://www.reddit.com/r/opencode/comments/1wrqysd/is_cache_reading_faulty_with_deepseek_v41_on/)
- OpenCode, 2026-09-27, r/opencode (Reddit): “opencode, in my experience, always had a lot of cache misses with deepseek models. seems to be 95%+ consistently on claude code, pi, and ds harness.” [source](https://www.reddit.com/r/opencode/comments/1wqbur6/why_am_i_getting_so_many_cache_misses_with/pcc633q/)
- OpenCode, 2026-09-23, @opencode (X): “frank/deepseek-v4.1-flash at @opencode can make deepseek-v4.1-flash feel like gtp-10, its eating tokens like crazy seem like cache issue , unusable” [source](https://twitter.com/1163317877904011270/status/2102837133837037744)

### 2. Cache hit and miss usage visibility

- Claude Code, 2026-09-25, @ClaudeDevs (X): “@claudedevs @claudedevs task-cost calculators matter more than list prices. do you split cache hits from fresh tokens for claude code?” [source](https://twitter.com/1030370607861387264/status/2103583188409073875)
- Cursor, 2026-09-25, @cursor_ai (X): “@cursor_ai selective tool loading is the kind of optimization users actually feel. i would still watch cache hit rate by repo, because a global average can hide one expensive codebase quietly burning the budget.” [source](https://twitter.com/2098494117693341696/status/2103312642911969305)
- Claude Code, 2026-09-18, r/ClaudeCode (Reddit): “appreciated. a live cost readout in the statusline would have caught the 09-16 spike the same hour instead of the next morning. will take a look. one request if you're taking them: show cache read tokens as their own number, not folded into a cost figure. this week made it clear that's the column that decides what a max week buys, and it's the one every cost display hides.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wk3zq5/i_audited_my_session_logs_against_the_usage_meter/pao2azo/)

### 3. Longer prompt cache retention window

- OpenAI Codex, 2026-09-21, r/codex (Reddit): “guys, what if background job finishes in more than 30 mins (in case), cache will expire and whole thing becomes more expensive, any suggestion here?” [source](https://www.reddit.com/r/codex/comments/1wlcy5q/this_will_save_your_usage/pb39ykj/)
- OpenCode, 2026-09-16, r/opencode (Reddit): “<strict_link> replied to deepseek v4.1 flash 2 minutes after its response. the cache was already cold. i expected it to last at least 5 minutes. is it normal behavior or some bug? (the app is my ui for talking to agents, it uses official opencode cli under the hood, not my own harness) edit: posted the solution in the comments. opencode non-interactive execution has a bug with deterministic system prompt.” [source](https://www.reddit.com/r/opencode/comments/1whlzic/does_cache_expire_immediately/)
- Claude Code, 2026-09-16, r/ClaudeCode (Reddit): “can you do the opposite? extend the cache timer so it doesn't reread the whole context if it sits there 2 hours?” [source](https://www.reddit.com/r/ClaudeCode/comments/1whrr18/limits_are_fixed/pa6mc1z/)

### 4. Preserve cache across model or effort switches

- OpenAI Codex, 2026-09-20, r/codex (Reddit): “if switching models didn’t invalidate my cache, i’d change it for simple tasks… but alas, a primed cache is more valuable than switching to low.” [source](https://www.reddit.com/r/codex/comments/1wl44z0/gpt6_astra_max_deleted_my_entire_project_in_a/pavvl0b/)
- OpenAI Codex, 2026-09-19, r/codex (Reddit): “one caveat when you run this for real: every flip appends another configuration\_update to history, so the prompt baseline creeps up a bit with each switch. cache hit stays high, but the floor keeps rising - worth knowing if you flip effort a lot inside one long session.” [source](https://www.reddit.com/r/codex/comments/1wk8jzq/one_flag_to_keep_989_of_my_prompt_cached_when/paov27a/)
- OpenCode, 2026-09-17, r/opencodeCLI (Reddit): “i mean, in his specific scenario where he selects another model first and sends the compaction request, this request most likely will miss all the cache, and when the compaction is done and he switches back to the previous model, still the compacted version will be cache-miss, as it is another model and maybe another provider.” [source](https://www.reddit.com/r/opencodeCLI/comments/1wipw64/do_you_use_a_skill_to_write_a_handoff_file_before/paczp95/)

### 5. Warning before cache expiry or cache-miss cost

- Pi, 2026-09-22, @pidotdev (X): “@pidotdev i think there should be a timer for cache busting! when the timer hits zero people can switch the model!” [source](https://twitter.com/1987272642827853824/status/2102423160113045553)
- OpenAI Codex, 2026-09-17, r/codex (Reddit): “13x on a cache miss is nasty, and a 30-minute window you can't see is asking for a surprise bill. showing the remaining ttl would fix the worst part of this.” [source](https://www.reddit.com/r/codex/comments/1wixw9c/codex_badly_needs_a_cache_timeout_indicator_like/paffd1p/)
- OpenAI Codex, 2026-09-16, r/codex (Reddit): “everybody except for the ml engineers at oai are horrifically incompetent. as you say, there are so many easy optimisations they could make to reduce usage. hell, even letting users know if their session is still cached and giving a big warning when users try to change reasoning level could help a lot. they could also turn off fork_turns for all subagents that aren't of the same model and reasoning level, but they don't. these people are idiots.” [source](https://www.reddit.com/r/codex/comments/1wh2rbt/heres_whats_going_to_happen/pa63t6f/)

### 6. Cheaper cached token pricing

- Cursor, 2026-09-22, @cursor_ai (X): “@elonmusk please address the higher cache read cost - $0.5 vs $0.2 for competitors - that makes grok 4.7 less price competitive in longer coding sessions. @cursor_ai @bot @grok please confirm that grok 4.7 have higher cache read cost and explain how it affects total cost” [source](https://twitter.com/7619212/status/2102543184522060124)
- OpenAI Codex, 2026-09-13, r/codex (Reddit): “5.6 introduced us being charged fully for cache writes, has nothing to do with cache hits. everyday hardware gets more expensive, the limits will keep getting tighter. this is because monthly sub prices remain the same.” [source](https://www.reddit.com/r/codex/comments/1wf9non/these_are_surely_getting_us_a_tibo_button_hit/p9k9p4e/)
- OpenAI Codex, 2026-09-12, r/codex (Reddit): “openrouter too, openai api too, anthropic api too, gemini api too, etc... all of them do it for the api, the point is we want it for the subscriptions too.” [source](https://www.reddit.com/r/codex/comments/1wdsjpe/please_give_us_a_slow_mode/p9fbjvl/)

### 7. Keep cache warm across resumes and sessions

- Claude Code, 2026-09-24, r/ClaudeCode (Reddit): “if you care about usage it is crap because you are likely hitting a big context window cache miss on resume that will eat up a fat chunk of you usage for nothing. this feature only make sense if you have less than 1 hour left to the quota reset (to get a guaranteed cache hit) or if somehow anthropic would keep the cache warm for you until the resume happens. in the current implementation it makes little sense.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wo5eaw/is_this_a_new_feature_in_cc/pbup4zo/)
- Claude Code, 2026-09-18, r/ClaudeCode (Reddit): “one of the key reason is cold cache. please try keeping it warm for long running sessions - <strict_link>” [source](https://www.reddit.com/r/ClaudeCode/comments/1wjl7t8/the_limits_are_disappearing_at_a_crazy_speed/pajfe3j/)
- Devin, 2026-09-18, @cognition (X): “@devinai, @cognition loading big sessions takes too long and it seems to not have any kind of cache between sessions swapping. please fix that :)” [source](https://twitter.com/64041638/status/2100936062142996708)

### 8. Preserve cache across compaction and context changes

- Claude Code, 2026-09-09, r/ClaudeCode (Reddit): “if they did create [claude.md](http://claude.md) files and changed them frequently, that is where your money went. [claude.md](http://claude.md) is autoreloaded, busting cache. that makes it a very expensive experience on a fable.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wbbsbk/fable_51_usage_consumption_my_experience/p8pkbc4/)
- OpenAI Codex, 2026-09-05, r/codex (Reddit): “consider changing the system prompt. since llm cannot grasp time information, it dynamically updates the time information periodically in system prompts. this completely destroys the cache and dramatically increases your usage. if you add a guideline to simply receive time information in a script when needed, you can significantly improve your cache hit and save on usage without any significant degradation.” [source](https://www.reddit.com/r/codex/comments/1w7x5ca/if_you_are_burning_your_token_too_much/)
- Claude Code, 2026-09-01, @ClaudeDevs (X): “@claudedevs cheaper cache reads only help if the prefix holds. one tool list reshuffle mid-session and you're paying full writes again.” [source](https://twitter.com/2044076080890281984/status/2094857810882556092)

### 9. Preserve cache for subagents and background waits

- OpenAI Codex, 2026-09-13, r/codex (Reddit): “30-minute subscription-side caching still misses if the polling interval between goal turns exceeds it - exactly what a background build/ci wait longer than 30 min triggers. claude code's suspend-instead-of-poll design avoids this by not re-entering the model while blocked.” [source](https://www.reddit.com/r/codex/comments/1wdlp7q/weve_discovered_the_issue_behind_codex_harness/p9lo1ej/)
- OpenAI Codex, 2026-09-10, r/codex (Reddit): “this solution would be so much more efficient if openai had looked into issue of losing cache context of parent to worker: <strict_link>” [source](https://www.reddit.com/r/codex/comments/1wcr5c1/astra_vs_astra_luna_agents/p90jdrm/)
- Claude Code, 2026-09-05, @ClaudeDevs (X): “@bcherny @anthropicai @darioamodei @claudedevs @claudeai @bcherny evry time 15+% of my limit just to continue the same conversation that i havent closed it just hit a limit thats diabolical and also the caching system for the sub agents dosent exist if they die mid run they die notig saved... please be fair and refund me my plan.” [source](https://twitter.com/2077293113563832320/status/2096161487995773262)

### 10. Built-in token reduction and cache management

- Zed, 2026-09-25, @zeddotdev (X): “@zeddotdev what about cache efficiency or native features like computer use?” [source](https://twitter.com/1326188894212263937/status/2103512769165426927)
- Pi, 2026-09-22, @pidotdev (X): “@pidotdev i like to build ways to reduce token consumption costs including cache management, if a pipeline to work in this can be integrated within pi by default i would love it.” [source](https://twitter.com/1743560080786599936/status/2102455054716506218)
- Claude Code, 2026-09-14, r/ClaudeCode (Reddit): “get a multi edit or patch tooling, optimize, qq, and stfu... only complaint i still have with anthropic and always have is their token hungry cacheing method that hardly any other model uses because it's literally stupid as hell! get a clue anthropic.. oh wait you worship money like it's your deity my bad.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wfwl6k/the_limits_have_been_reduced_even_further_now_its/p9uvj36/)

### 11. Fix cache billing errors

- Claude Code, 2026-09-05, r/ClaudeCode (Reddit): “cache writes were failing and it would retry with an ever growing delta - but you still paid for the failed cache writes. had 6 million cache writes on 300k of output tokens. fable spawning fable agents with lots of short turns and broken caching destroyed limits faster than i thought possible.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w7fckf/did_we_just_get_a_reset/p7xj8y0/)
- Claude Code, 2026-09-03, r/ClaudeCode (Reddit): “the /low-priority from my observation busts cache on each change. if you have 500k context filled, you run low-priority it cold serves it every time "there is capactiy" meaning you do normal price reads instead of cache reads. so this feature is useless unless they change the billing to as if the requests were using cache even if each inference happened more then 5 minutes after each other.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w5oy5d/just_used_the_new_lowpriority_feature_and_my/p7jgvxl/)
- Claude Code, 2026-09-01, @ClaudeDevs (X): “@claudedevs the line that really needs to be looked at is the cache read: in the long agent loop, it often takes up more than 60% of the input tokens. at the same price of $0.30/m, it's 75% cheaper, which is $0.075/m. reading 200m cache a day drops from $60 to $15. let's put the score leaderboard aside for now and check if the cache item in your bill matches up.” [source](https://twitter.com/2065688776706531328/status/2094881183439933549)

### 12. Credit usage lost to cache bugs

- Claude Code, 2026-09-09, r/ClaudeCode (Reddit): “if the changelog admits cache misses on every tool turn, users paid for anthropic's bug. credit unused usage or extend the window. "update and hope" is not a settlement for max prices.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wbclml/claude_code_promptcache_bugs_who_pays_for_the/p8p0gma/)
- Claude Code, 2026-09-09, r/ClaudeCode (Reddit): “i previously posted that i suspected claude code was not yet properly optimized for orchestrator-style workflows with many subagents, tools, hooks and long-running context, because the usage consumption was extreme. since the fable 5.1 rollout, we started looking much more closely at the cache behavior. and the official claude code changelog now confirms multiple serious prompt-cache bugs. confirmed fixes include: \- v2.1.260: fable 5.1 context a” [source](https://www.reddit.com/r/ClaudeCode/comments/1wbclml/claude_code_promptcache_bugs_who_pays_for_the/)

## Every agent

| Agent | Overall rank | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|---|
| [Pi](https://feedbackbench.com/agents/pi.md) | 7 | Better than peers | 0.536 | 0.510–0.565 | 61 | 39 | 22 |
| [Claude Code](https://feedbackbench.com/agents/claude-code.md) | 1 | Typical | 0.513 | 0.489–0.538 | 350 | 135 | 215 |
| [OpenCode](https://feedbackbench.com/agents/opencode.md) | 3 | Typical | 0.505 | 0.473–0.535 | 96 | 34 | 62 |
| [OpenAI Codex](https://feedbackbench.com/agents/codex.md) | 2 | Worse than peers | 0.441 | 0.409–0.472 | 226 | 47 | 179 |
| [Cursor](https://feedbackbench.com/agents/cursor.md) | 4 | Too few posts | – | – | 22 | 11 | 11 |
| [Devin](https://feedbackbench.com/agents/devin.md) | =5 | Too few posts | – | – | 13 | 9 | 4 |
| [GitHub Copilot](https://feedbackbench.com/agents/copilot.md) | =8 | Too few posts | – | – | 6 | 3 | 3 |
| [Cline](https://feedbackbench.com/agents/cline.md) | =8 | Too few posts | – | – | 4 | 2 | 2 |
| [Google Antigravity](https://feedbackbench.com/agents/antigravity.md) | =5 | Too few posts | – | – | 3 | 1 | 2 |
| [Amp](https://feedbackbench.com/agents/amp.md) | =11 | Too few posts | – | – | 3 | 2 | 1 |
| [Factory](https://feedbackbench.com/agents/factory.md) | =11 | Too few posts | – | – | 2 | 1 | 1 |
| [Warp](https://feedbackbench.com/agents/warp.md) | 15 | Too few posts | – | – | 2 | 1 | 1 |
| [Zed](https://feedbackbench.com/agents/zed.md) | =8 | Too few posts | – | – | 1 | 0 | 1 |
| [Kiro](https://feedbackbench.com/agents/kiro.md) | 13 | Too few posts | – | – | 1 | 0 | 1 |
| [Conductor](https://feedbackbench.com/agents/conductor.md) | 14 | Too few posts | – | – | 0 | 0 | 0 |
| [Grok Build](https://feedbackbench.com/agents/grok-build.md) | 16 | Too few posts | – | – | 0 | 0 | 0 |
| [Augment Code](https://feedbackbench.com/agents/augment.md) | 17 | Too few posts | – | – | 0 | 0 | 0 |

## Posts

Receipts rule: The 5 most recent praise and complaint posts per area (first 700 characters) and 3 per criterion (first 450 characters).

### Pi

- Praise, 2026-09-25, @pidotdev (X): “harness learning #6: how @pidotdev handles prompt caching prompt caching is one of the most important points in cost efficiency in using llms, and harnesses have the responsibility of sending requests that match a provider's cache policies. this is especially hard for harnesses like pi that integrate with many providers, and they've just released a cool new cache warming feature! pi has a cacheretention setting that can be set to none/short/long,” [source](https://twitter.com/1689423238173007873/status/2103574975831486941)
- Praise, 2026-09-23, @pidotdev (X): “nearly 1 billion cache read tokens, 99,9% hit rate. 12 m token out 6,5 m token in one of my longest session yet i guess. thx @pidotdev <strict_link>” [source](https://twitter.com/1497826290535120897/status/2102665452233376011)
- Praise, 2026-09-23, @pidotdev (X): “@0xhashlol @jazzychad well then it’s the users’ problem if their harness eats through usage limit. @pidotdev for example is incredibly lean (fewer tokens in system prompt) and has an amazing cache hit rate but yeah just let me do it at my own risk. works fine with openai so i’m using their models 🤷♂️” [source](https://twitter.com/2915475819/status/2102848541379178525)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “every use case is different. personally, i prefer a lean approach where the context window stays small and gets summarized regularly, while solid documentation lets me spin up new sessions with fast project context recovery. - i avoid large context windows because i don't have much vram. once a task is done, i compress the context using my `pi-refine-compact` extension, so the llm stays up to speed on what was done in broad strokes and we can mov” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrq50c/new_session_handoff_vs_compact_which_do_you_prefer/pcf5zei/)
- Complaint, 2026-09-27, @pidotdev (X): “it's super interesting to run codex models in the @pidotdev harness vs the codex harness. @badlogicgames has it mention when there's a cache miss, and it's most of them. so if you're upset about how fast your sub gets used up, use it on pi and what it will teach you, may get you to figure out how to manage your context window better.” [source](https://twitter.com/21734113/status/2104223345898319922)
- Complaint, 2026-09-24, r/PiCodingAgent (Reddit): “thanks for sharing. i never promote this feature into my workflow because of the cache invalidation. i have been throwing darts towards this though.. so let me know if you're interested to chat about it further.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wnf2nv/rethinking_the_humanagent_interface_with_pi/pbr3sxw/)

### Claude Code

- Praise, 2026-09-27, r/ClaudeCode (Reddit): “i’ve been on the 20$ for months, with 2 accounts with different billing users: never a problem. anyway, with opus 5.5 and the new usage window of anthropic, start with it: 5.5 usage has an an astonishing cache hit rate that won’t kill usage as opus class 4.x” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr5yr4/pro_plan/pcadvzn/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “oh for sure. i average 98% cache hit usually” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrmw8c/high_or_medium_for_opus_55/pcekoy5/)
- Praise, 2026-09-27, r/opencode (Reddit): “opencode, in my experience, always had a lot of cache misses with deepseek models. seems to be 95%+ consistently on claude code, pi, and ds harness.” [source](https://www.reddit.com/r/opencode/comments/1wqbur6/why_am_i_getting_so_many_cache_misses_with/pcc633q/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “ok. so i just went through something similar using claude & kimi. (besides from some bugs i found with how claude uses 3rd party apis & caches the following should be useful). when your limit opens up again, get it to look at the last 24 hours of your session history. see if it can break down by something meaningful to you (i run the [assay.guide](<strict_link>) harness, and it runs 5 sessions at a time, each with different profiles). for each t” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrjhyp/are_the_limits_that_good/pcdgpig/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “update: i had claude (opus) dig through my local session logs, and i don't think the weekly limit was cut. the api-equivalent metric broke. tl;dr: on my account, cache reads used to count as \~zero against the subscription quota, and they now count under the opus 5.5 model id. opus 5.5 also has much cheaper api pricing. put those together and "api-equivalent $ per 100% of the week" roughly halves, while the real quota looks unchanged. what we fou” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrm570/did_max_20x_weekly_limits_just_get_cut_in_half/pcdpiq9/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “if you run a lot of sessions in parallel to work on multiple projects you are very likely to have your cache for each of these sessions go cold on a frequent basis. all it takes is 60 mins of inactivity and boom, you will have to pay like 125% for a new cache write. (for everything that is in the context window of that session) this while a warm cache only costs you like 10%... (and for fable 5.1and opus 5.5 they made this even 60% cheaper then t” [source](https://www.reddit.com/r/ClaudeCode/comments/1wropen/i_keep_hitting_usage_limits_on_20x_plan_have/pcefkmq/)

### OpenCode

- Praise, 2026-09-26, r/opencodeCLI (Reddit): “caching still works great as it seems!” [source](https://www.reddit.com/r/opencodeCLI/comments/1wpyova/what_would_make_you_switch_your_default_opencode/pc30o7m/)
- Praise, 2026-09-26, r/opencode (Reddit): “superior in terms of cache hit rate and pricing during off peak hours. also very high token speed. becomes less worth it during peak hours.” [source](https://www.reddit.com/r/opencode/comments/1wqg9xq/best_direct_providers_via_api_or_similar_to/pc4afr6/)
- Praise, 2026-09-26, r/opencode (Reddit): “i'm not saying it's bad and ut makes it good, but it's like the harness is made to plug all the gaps and while it's missing some features as it's in beta, you will feel the difference because it leverages the crazy caching and ptc mode to make the model constantly aware of where it is and what to do next so even with 800k context it felt like it was still on 50k context i suggest u put only 5$ as i did in it and try for yourself” [source](https://www.reddit.com/r/opencode/comments/1wq1odd/deepseek_41_flash_4x_usage_in_opencode_go_made/pc53ptq/)
- Complaint, 2026-09-27, r/opencode (Reddit): “they started using third party providers when they started dropping from $60 limit to $15 limit a few months ago. everything open-weight served by opencode zen/go should be assumed to be quantized and will have worse cache hits/retention than 1st party api will.” [source](https://www.reddit.com/r/opencode/comments/1wqy8pq/how_true_is_it_that_opencode_go_models_are/pc9ubms/)
- Complaint, 2026-09-27, r/opencode (Reddit): “opencode, in my experience, always had a lot of cache misses with deepseek models. seems to be 95%+ consistently on claude code, pi, and ds harness.” [source](https://www.reddit.com/r/opencode/comments/1wqbur6/why_am_i_getting_so_many_cache_misses_with/pcc633q/)
- Complaint, 2026-09-27, r/opencode (Reddit): “you get 6 times the tokens but they are not the same as official apis. you get quantized tokens which is not the same as the real deal + opencode also does its own caching differently that would mess with the speed and quality of your output. so yes you get 6 times to tokens but not the quality and speed. edit: at least until the engineers at opencode do something about it.” [source](https://www.reddit.com/r/opencode/comments/1wrg81n/go_subscription_is_slower_deepseek_flash_41/pcc93xw/)

### OpenAI Codex

- Praise, 2026-09-27, r/codex (Reddit): “more or less yes, on api opus 5.5 has the same pricing as 5.6-sol, except it's almost astra-level. furthermore cache reads are actually twice as cheap as 5.6 sol, making it even less costly in actual use. that and the 5h limits are better on the $20 plan (but the flipside is that $100 still has 5h limits and weekly limit is not 5x, or so i've heard), and there are no context size restrictions (it's actually 1m). it reminisces me of when deepseek” [source](https://www.reddit.com/r/codex/comments/1wrtru3/openai_will_need_to_stand_on_their_head_and_add_a/pcgddu2/)
- Praise, 2026-09-27, r/ClaudeAI (Reddit): “we're builders who live in claude code and codex, multiple sessions and large projects and the visibility they give you is lacking. we built this dashboard for deeper visibility into our work. i shared it the other day when opus 5.5 dropped, as soon as someone commented that it was cool, i open sourced it and shared the link. i've been doing this for 30 years and it's always extremely satisfying to see people get value from or enjoy something y” [source](https://www.reddit.com/r/ClaudeAI/comments/1wrckhp/what_tool_have_you_built_for_yourself_with_claude/pcex1tu/)
- Praise, 2026-09-26, r/opencode (Reddit): “the low cache hit rate is due to the opencode harness. i've created a proxy to intercept token usage and cache hits, and using it in codex, i get a 98% cache hit rate” [source](https://www.reddit.com/r/opencode/comments/1wqox6a/cheepseek_has_its_price_cache_hit_ratio/pc7dltl/)
- Complaint, 2026-09-27, r/codex (Reddit): “it's awful. it seems to use the same token usage as claude code, with none of the context management options. it's like a claude code chat that's permanently at max context size with no way of knowing if anything is still cached completely unusable” [source](https://www.reddit.com/r/codex/comments/1wr2ehk/chatgpt_pro_5x_is_now_standard/pcbo4vn/)
- Complaint, 2026-09-27, r/codex (Reddit): “interesting! i’ve been experimenting with a related approach but using a very cheap judge/router (jev) to dynamically switch the main session between cheap and frontier models instead of keeping a fixed astra-orchestrator/luna-worker split. one thing i found is that the routing decision itself is basically free but that the expensive part can be the context transition. i'm doing small controlled a/b and to give you an example: a "tricky" coding t” [source](https://www.reddit.com/r/codex/comments/1wqoopb/i_measured_astra_orchestrating_luna_vs_doing_the/pcch48v/)
- Complaint, 2026-09-27, r/codex (Reddit): “yeah, it's wild, they release the most unreliable models in a year, that need the most handholding and supervision, still without a proper caching solution, even less caching for multi-agent and even less when switching models, keep releasing slop updates to their slop app every day, and expect me to trust their slop multi agent message board for my work? get fucked altman” [source](https://www.reddit.com/r/codex/comments/1wrn697/new_openai_product_o_but_not_for_plus_users/pcdze47/)

### Cursor

- Praise, 2026-09-23, @cursor_ai (X): “@cursor_ai cache savings compound in long sessions: a stable prompt prefix and unchanged tool schemas are re-read every turn, so the same cut lands many times.” [source](https://twitter.com/2063926210942427136/status/2102806070980952455)
- Praise, 2026-09-23, @cursor_ai (X): “@wiiiimm @cursor_ai fair call, lazy loading tools and cache hits do most of the heavy lifting anyway.” [source](https://twitter.com/1380769041984278530/status/2102806321397670070)
- Praise, 2026-09-23, @cursor_ai (X): “@cursor_ai 7% with no quality drop is underrated. harness wins compound faster than model hops, and selective tool loading is the one most agents seem to skip — shipping the whole toolbox every turn.” [source](https://twitter.com/1486837134585516037/status/2102809364566610126)
- Complaint, 2026-09-27, r/cursor (Reddit): “there is a good difference. in token usage and real dollars cost . because you don't know how much you usge of each token type ( input , output , cache ) and we found out that cache is around 90% while input and output are 5% each .and maybe this for my use case for other people the numbers could be different. also they don't mention how much faster the fast is and from this experiment we found it could be nothing. or double the speed.” [source](https://www.reddit.com/r/cursor/comments/1wp2ped/i_compared_cursor_composer_25_normal_vs_fast/pcc89di/)
- Complaint, 2026-09-25, @cursor_ai (X): “@cursor_ai whats going on with your support? since a few days my quota is burned like crazy, by using projects. i sent full details to support, explained everything with detailed diagnostics, proving that i am not using more, but the cache cost has suddenly exploded and is over 90% of what is billed and your support gaslights me that i simply should reduce my coding activity. sorry, but that is crazy. you clearly have a cache bug since about 2 da” [source](https://twitter.com/804676521529110528/status/2103389221817860558)
- Complaint, 2026-09-25, @cursor_ai (X): “@cursor_ai whats going on with your support? since a few days my quota is burned like crazy, by using projects. i sent full details to support, explained everything with detailed diagnostics, proving that i am not using more, but the cache cost has suddenly exploded and is over 90% of what is billed and your support gaslights me that i simply should reduce my coding activity. sorry, but that is crazy. you clearly have a cache bug since a few days” [source](https://twitter.com/804676521529110528/status/2103391265563767002)

### Devin

- Praise, 2026-09-27, @DevinAI (X): “five days. one devin cli session. two models. ~688 million tokens of context. now i told you i was gonna push the limits of @devinai fusion, so here goes! that's what it took to cut over our company ai brain: retire a custom review pipeline, move to stock gbrain, refile ~760 misfiled pages, merge duplicates, and run a hand-graded eval. 7 prs merged, 2 bugs reported upstream instead of forked. why so many tokens? this wasn't greenfield coding. it” [source](https://twitter.com/218098611/status/2104252961430102289)
- Praise, 2026-09-27, @DevinAI (X): “@jensenloke @devinai 688 m tokens, yet 87 % cache hits make it surprisingly cheap” [source](https://twitter.com/1780523178160279552/status/2104271636820328850)
- Praise, 2026-09-02, @cognition (X): “@cognition the real news is not that it's stronger, but that the cache has been cut to a quarter. ninety percent of the coding tasks are cache, cutting here is more effective than cutting the price.” [source](https://twitter.com/1934622485674340352/status/2094979670924038348)
- Complaint, 2026-09-02, @cognition (X): “@cognition 54% cheaper from a caching change alone is a big claim. i'd want to see whether that holds outside frontiercode style benchmarks.” [source](https://twitter.com/363766025/status/2094962836162392472)
- Complaint, 2026-09-02, @cognition (X): “@cognition 47% savings often shrink on cold caches, where production traffic has less prompt reuse” [source](https://twitter.com/1803494630366785536/status/2094967176243380272)
- Complaint, 2026-09-01, @cognition (X): “@cognition 95% cached tokens turns model pricing into a startup tax: the first request hurts, then the long run gets cheap.” [source](https://twitter.com/1067135083155464194/status/2094901820862738844)

### GitHub Copilot

- Praise, 2026-09-21, r/GithubCopilot (Reddit): “i'm not a big fan of planning and executing with different models. but i understand the benefits, sol is a bit smarter than luna max and you have the chance of reading and updating/discarding the plan before execution. another alternative is to accept that the plan won't be perfect from the start, but you want to have an iteration fast, maybe while you work something in the background and then you can go all in with luna max. it also has the bene” [source](https://www.reddit.com/r/GithubCopilot/comments/1wml5js/im_overwhelmed_by_the_choice_in_models_but_also/pb8oack/)
- Praise, 2026-09-03, r/GithubCopilot (Reddit): “cost is lower than gpt-5.6 luna and terra. that will just burn tokens and run in circles; caching on gemini is much, much better.” [source](https://www.reddit.com/r/GithubCopilot/comments/1w6bfln/gemini_flash_38_in_github_copilot_soon/p7may3r/)
- Praise, 2026-09-02, r/ClaudeCode (Reddit): “prompts caching in cc is only 5 minutes right? so i would say yes. github copilot it is 24h by the way.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w51gn6/would_i_save_on_usage_by_switching_to_a_cheaper/p7bpa9y/)
- Complaint, 2026-09-25, r/GithubCopilot (Reddit): “do you feel that you have to micro-manage the tool enablements often? my perspective is that users should not generally have to do this- tools can already be lazy-loaded by the harness, so the agent is better at dealing with lots of tools that it used to be. also, changing the enabled set of tools can break the prompt cache in a session, making that feature sort of a footgun. another way to manage the set of tools in a session is by creating a” [source](https://www.reddit.com/r/GithubCopilot/comments/1wpctlc/vs_code_chat_users_local_or_copilot_harness/pc043kk/)
- Complaint, 2026-09-19, r/GithubCopilot (Reddit): “but you are clearly ignoring the 1.25x cache write cost you have that wasn't there before. cost is still higher” [source](https://www.reddit.com/r/GithubCopilot/comments/1wk7lky/gpt_54_deprecation_will_legacy_yearly_subscribers/paqn25o/)
- Complaint, 2026-09-06, @GitHubCopilot (X): “@catmanyau @githubcopilot define keep context it does keep your context of the conversation but it will switch the system prompt for this specific model and will full invalidate your cache you should really think before changing models and u can use tools like rubber duck or creating sub agents to review” [source](https://twitter.com/36475277/status/2096575808844206395)

### Cline

- Praise, 2026-09-24, @cline (X): “cline has improved a lot in a mean time, the cache hit rate is absolutely insane now great work, guys @cline <strict_link>” [source](https://twitter.com/2017132628361822208/status/2103104236540375166)
- Praise, 2026-09-12, r/opencodeCLI (Reddit): “this is the first thing i noticed. i make a large plan (90k tokens), switch to build and then immediately it starts prompt processing from scratch. the whole cache was invalidated! cline at least just puts another prompt to tell the model it switched modes.” [source](https://www.reddit.com/r/opencodeCLI/comments/1usvuib/cache_invalidation_when_switching_from_plan_to/p9g4kom/)
- Complaint, 2026-09-16, r/LocalLLaMA (Reddit): “is there a coding harness that will not invalidate the entire context cache at each prompt? as the title says. i don't think it's normal to invalidate 100k of context cache for no reason. for example cline and qwen code addon for vs code is doing it most times. let's say you write a big codebase then you ask something like "please update the md file" or some very small change of a simple script and it will trigger the invalidation of the entire c” [source](https://www.reddit.com/r/LocalLLaMA/comments/1wi17z6/is_there_a_coding_harness_that_will_not/)
- Complaint, 2026-09-09, r/CLine (Reddit): “it might be that compacting causes cline to use /newtask, and the compaction and /newtask gets repeated since it’s now in context. you could find out by adding a pretooluse hook to log tool calls. i’m afraid of cline’s auto compact hosing cache hits so i don’t use it, but i’m also not giving it a single task that would take 12 hours and/or 60m tokens. the other method to find out what’s going on is to let the model have a go at interpreting/expl” [source](https://www.reddit.com/r/CLine/comments/1wbdpuy/context_window_management/p8tkdh8/)

### Google Antigravity

- Praise, 2026-09-25, r/google_antigravity (Reddit): “interesting. not seeing this myself (yet), but i did receive an email from google about durable caching being enabled shortly. that should make the caching \_better\_ though 🙂 >on **october 15, 2026**, durable caching will become generally available (ga). as part of this release, google cloud will enable durable caching by default for projects with implicit caching enabled across gemini 3.x pro, gemini 3.x flash models, and any new gemini models” [source](https://www.reddit.com/r/google_antigravity/comments/1wpsvzy/antigravity_scheduled_tasks_failing_with_no/pbyyilt/)
- Complaint, 2026-09-18, @antigravity (X): “@plutonusstudio @typesafeai @antigravity thank you! there obviously is a lot of room for improvement. e.g. smart caching of repeated toll calls, giving more context (what if the user prompted to do the unsafe tool call? especially in cc and codex unsafe tool calls are always hard-rejected :(” [source](https://twitter.com/878665358625976320/status/2100955966460084494)
- Complaint, 2026-09-05, r/google_antigravity (Reddit): “psa: changing a model thinking level mid conversation can blow out the cache and quota” [source](https://www.reddit.com/r/google_antigravity/comments/1w7cnzx/shortcut_to_change_effort_of_model/p7vzfu8/)

### Amp

- Praise, 2026-09-08, @AmpCode (X): “@codermatt @ampcode hey! we have compaction for all threads, so technically you can just continue using one thread. but we highly highly recommend using one thread per "task": one bug, one feature, one investigation, and so on. (133k input tokens isn't that much and great that caching works)” [source](https://twitter.com/414333187/status/2097280269963120734)
- Praise, 2026-09-02, @AmpCode (X): “@khoiracle @cursor_ai @ampcode the unlock is detached state, not raw parallelism: each run keeps its own env and you review async. cost only holds if the harness caches context per session - with fable 5.1 cache reads at ~$0.25/mtok, the vms stop being the expensive part.” [source](https://twitter.com/1864704498112897024/status/2095134037434073485)
- Complaint, 2026-09-07, @AmpCode (X): “sometimes i come back to a thread after a long time and know the cost of resuming due to cache busting will be high. would be neat to be able to force compaction or resume from a summary @ampcode <strict_link>” [source](https://twitter.com/33135576/status/2096802985645134185)

### Factory

- Praise, 2026-09-20, @FactoryAI (X): “look at this numbers. cache miss is 1% on deepseek byok wit @droid - this is unbelievable great!! that is the quality of a harness indicator in my opinion- key metric. @factoryai <strict_link>” [source](https://twitter.com/1590702228234391552/status/2101737157379641558)
- Praise, 2026-09-20, @FactoryAI (X): “look at this numbers. cache miss is 1% on deepseek byok with @droid - this is unbelievable great!! that is the quality of a harness indicator in my opinion- key metric. @factoryai <strict_link>” [source](https://twitter.com/1590702228234391552/status/2101737321368551459)
- Complaint, 2026-09-26, @FactoryAI (X): “the transit station really has no work to do. i have always been curious about the router that @factoryai has been promoting. they have never clearly explained the specific router mechanism, just vaguely mentioning the classification of simple and complex tasks. how can we determine whether a task is complex or not before it starts, and the inevitable cache invalidation caused by switching models midway is a big problem (anyway, the transit stati” [source](https://twitter.com/1724973133097291776/status/2103781371772809499)

### Warp

- Praise, 2026-09-10, @warpdotdev (X): “@warpdotdev have you guys tried glm 5.3 flash? my cache hits are real high on concentrate” [source](https://twitter.com/773953758673895424/status/2098050922156859520)
- Complaint, 2026-09-18, @warpdotdev (X): “@vikvang1 @warpdotdev 4. pass prompt caching through on byok. your harness adds a big prefix every turn - without cache_control on anthropic endpoints, byok users pay full input price on the same 30k tokens each call!” [source](https://twitter.com/14132756/status/2100995017749967068)

### Zed

- Complaint, 2026-09-17, @zeddotdev (X): “@zeddotdev i was a big fan of zed, but i'm becoming so tired with it. it's a pain to review agent's work. the built-in agent seems to have severe caching issues and the acp, which i totally love, doesn't support clean reviews.” [source](https://twitter.com/1777822808506023936/status/2100657499468677214)

### Kiro

- Complaint, 2026-09-17, r/kiroIDE (Reddit): “by letting the cache go stale, implied by you saying continue, you just repaid all our cache write tokens again also” [source](https://www.reddit.com/r/kiroIDE/comments/1whrobc/in_less_than_30_secs_kiro_uses_almost_100_credits/paaymrs/)
