# Single prompt, model or effort level consumes disproportionate quota (`limits.burn_rate`)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/criterion/limits.burn_rate

Area: [Paying and limits](https://feedbackbench.com/criteria/paying.md)

**Definition.** How much quota a specific prompt, task, model version, effort level or feature consumes for the work done. Includes new versions using more than earlier ones and praise for efficient ones. The post names what consumed the quota.

**Boundary.** Not this: see [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) when the post is about being blocked by the short window, not about the consumer. Not this: see [Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md) for cache misses. Not this: see [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) when runaway subagents cause the burn.

Rated author-weeks, all agents: 9619. Complaint share: 84%.

## The brief

Written by Claude Opus 5.5 from 110 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**The model you pick, not the plan, decides how fast quota disappears.**

TL;DR:

- Complaints dwarf praise for every agent with real volume; burn rate is a category-wide sore spot.
- New model versions that eat more quota for the same work drive the angriest posts.
- Lean harnesses and cheap worker models earn the praise; Pi and Devin stand out.

In plain terms: Users watch one heavy prompt or a new model version wipe out a window that used to last days. The relief comes from switching to lighter models, cheaper defaults, or a leaner harness, not from waiting for resets.

### How it breaks

- **New versions cost more for same work** ([Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md)). Version bumps are the top trigger. Users see the same task on a newer model drain noticeably more quota with no visible gain in output.
  Posts across Google Antigravity, Claude Code, Cursor and OpenCode follow one pattern. A new model ships, the workflow stays identical, and the allowance drops faster. Users flag the gap most sharply when API prices match but in-product quota does not. Some describe the new model as the old one plus extra steps. The request users repeat is to stop new versions from charging more than their predecessors.
  Evidence:
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-05: “flash 3.8 uses a huge amount of allowance compared to 3.7 making it not worth using.” [source](https://www.reddit.com/r/google_antigravity/comments/1w81yel/gemini_38_flash_review/p7zb5ve/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-01: “we just got fable 5.1 which consumes more tokens than 5 (and also the reset)” [source](https://www.reddit.com/r/ClaudeCode/comments/1w4jres/we_just_got_a_reset/p784g36/)
  - Complaint, Cursor, r/cursor, 2026-09-24: “you're wrong, is not the same price, grok 4.7 consumes 3x more usage than 4.6” [source](https://www.reddit.com/r/cursor/comments/1wp0ext/cursor_is_good_as_a_second_hand/pbruuvl/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-09: “it’s kinda annoying that on api 3.8 and 3.7 are the same price but in antigravity 3.8 uses so much more quota.” [source](https://www.reddit.com/r/google_antigravity/comments/1w9mx0c/weekly_quotas_known_issues_support_september_07/p8nl8h8/)

- **One prompt empties the window** ([Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md)). A single heavy request can consume a whole session or even a monthly allowance, and users rarely see it coming.
  Users report one prompt wiping monthly credits in Warp, a small list costing a visible slice of Copilot credits, and a long autonomous run burning half a week's Codex allowance with nothing usable at the end. The pain is less the cost itself than the mismatch between task size and spend. Ten minutes on a premium model can end a five-hour window.
  Evidence:
  - Complaint, Warp, @warpdotdev, 2026-09-22: “ran out of my monthly @warpdotdev credits from a single prompt change to a docker container. maybe the legacy plan i'm on is just useless now?” [source](https://twitter.com/80764812/status/2102534429369303091)
  - Complaint, OpenAI Codex, r/codex, 2026-09-11: “even the iterative development is mostly a mess. it stumbles all over the place, applies the wrong solution, uses the wrong models - makes thinks "plastic" and dull to shortcut. i ran a goal today... my mistake ... to develop a game event and came back and there was nothing. just a sandbox. completely nuts. burned like 50% of the weeks credits.” [source](https://www.reddit.com/r/codex/comments/1wd53rg/astra_demos_are_mostly_bs/p938ol7/)
  - Complaint, GitHub Copilot, @GitHubCopilot, 2026-09-26: “an entire 2% of @githubcopilot pro credits gone for a simple list in a @windows app which is less than 200 lines of code? explain to me again, how is this replacing programmers? this is code documented every year for over 30 years now. todo list style programming... <strict_link>” [source](https://twitter.com/22220922/status/2103777500224754018)
  - Complaint, Factory, @droid, 2026-09-05: “10 minutes of astra @droid and the 5hour usage is gone. unusable.” [source](https://twitter.com/1590702228234391552/status/2096151009454444568)

- **Premium defaults quietly burn credits** ([Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md)). Agents default to frontier models at high effort, and users only find the cheap path after running out.
  Users describe learning the hard way that defaults were set to max and fast. Switching the bulk of the work to a lighter model stretches quota dramatically while output stays acceptable. Copilot users trade tips on lighter models. Cursor users route most coding to Composer. Kiro users report context-size multipliers that compound high-effort burn. The expensive-planner, cheap-worker split is the pattern users endorse.
  Evidence:
  - Praise, GitHub Copilot, r/GithubCopilot, 2026-09-20: “switching to luna medium saves my credits!!! i just realize the default astra medium is a huge credit burner. after switching to luna medium, plus using chatgpt for clear instruction (after sorting out the logic etc), my luna medium usage is like 10% for 5 hours of work. thank you for the advice! really appreciate it!” [source](https://www.reddit.com/r/GithubCopilot/comments/1wkl561/codex_vs_github_copilot/pawgnsk/)
  - Complaint, Cursor, r/cursor, 2026-09-14: “ya, that's what i'm saying. restrict it to cursor grok (not fast) and composer. use composer for most of the coding (which is perfectly suitable). lasts a long time i learned the hard way that the model defaults were : use the forntier models set to max and fast once i run those out i still have lots of cursor grock & composer to use” [source](https://www.reddit.com/r/cursor/comments/1wg22hl/is_switching_away_from_cursor_really_the_play/p9sz09q/)
  - Praise, GitHub Copilot, r/GithubCopilot, 2026-09-11: “just in case some people didn't know, luna is still the absolute best bang for your buck. all other models are a waste of tokens unless you are running agentic workflows supervised by an agent. in that case, you should still be using luna, but just have it supervised by astra. i've tried every model and i always find my way coming back to luna as my bread and butter for iterative work. i think it's just as good as codex, if not better. save your tokens and use luna.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wdtxyz/luna_is_still_the_goat/)
  - Complaint, Kiro, r/kiroIDE, 2026-09-16: “well, that is on sol max, and if you passed the 256k token context size it flipped to 8.8x. that aside, it really is rough now. i was using sol medium last night for about 2 hours and burned at least 1000 credits. i went down to luna and that didn’t help much because i quickly passed the 256k context window.” [source](https://www.reddit.com/r/kiroIDE/comments/1whrobc/in_less_than_30_secs_kiro_uses_almost_100_credits/pa63dfo/)

- **The harness sets the bill** ([Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md)). The same model costs different amounts depending on the wrapper around it, and users notice the overhead.
  Users compare the same subscription across tools and report wide gaps. Some say a third-party harness burns more than the native CLI. Others say the native app burns more than Zed or Pi. Copilot users complain the engine rescans files unrelated to the request. Users attribute the difference to system prompts, context loading and extra file reads rather than the model itself.
  Evidence:
  - Complaint, Conductor, @conductor_build, 2026-09-12: “@garrytan @charlieholtz @steve_yegge @conductor_build conductor is good but their harness for anthopic and openai models consumes a lot more tokens than cc and codex” [source](https://twitter.com/1557669964085358592/status/2098798329799008328)
  - Praise, OpenCode, r/codex, 2026-09-05: “it is very good! which harness do you use with it? i usually use gpt models with pi or opencode it’s much more token efficient than codex. but not sure if astra would still be as good as in codex?” [source](https://www.reddit.com/r/codex/comments/1w7p080/which_harness_for_astra/)
  - Praise, Zed, r/codex, 2026-09-09: “just wanted to share my personal experience in case anyone else is running into the same issue. i've been using codex through the vs code extension, and i noticed that i was going through my usage pretty quickly. recently, i started using codex with zed instead, and from what i've seen so far, my usage has been noticeably lower while getting pretty much the same results. i haven't done any proper benchmarks or controlled tests, so i'm not saying zed definitely uses less. this is just what i've noticed with my own usage and workflow. for me, the difference has been significant enough that i've kept using zed. i'm curious if anyone else who has used codex with both vs code and zed has noticed something similar.” [source](https://www.reddit.com/r/codex/comments/1wbraq6/my_codex_usage_dropped_significantly_after/)
  - Complaint, GitHub Copilot, r/AI_Agents, 2026-09-21: “exactly the same feeling and it's very frustrating. i use github copilot and all the time their engine rescan certain files i don't ask for. so it's overconsume credits for nothing while i just ask to make a feature which does not concern this file. i don't know how to do to avoid this behavior” [source](https://www.reddit.com/r/AI_Agents/comments/1wjnj7g/i_stopped_using_the_smartest_ai_models/pb8ugax/)

### Who stands out

- **OpenAI Codex (weaker)**. The loudest complaint pool on this page. Users who once finished weeks comfortably now drain windows in a few prompts on the same models.
  Posts describe a step change. Same workflow, same model tier, and the five-hour or weekly allowance now goes in a fraction of the time. Codex leads the requests for lower overall burn, cheaper simple tasks and lighter premium models. Praise exists but clusters on lighter or more token-efficient models, not on the heavy tiers users default to.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-05: “it drains at least twice as fast now, i always used sol 5.6 high and could do a solid 5-6 good prompts and usually it would do good work. i drained my 5 hour limit in 2-3 prompts since trying it today...” [source](https://www.reddit.com/r/codex/comments/1w7kvf7/gpt6_astra_wow/p7y2xwm/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-13: “i did in march and could comfortly work with a 20x since then. even with gtp 5.6 sol in xhigh i had a comfortable 10-20% left at the end of the reset. now everything gone on 2 days. and no, im not going back to old models” [source](https://www.reddit.com/r/codex/comments/1wf04i2/they_dont_have_enough_compute_for_20x_subs_but/p9i3mxg/)
  - Complaint, OpenAI Codex, r/ClaudeCode, 2026-09-25: “ah thanks for that insight! i'm only on plus plan and find codex/work blows it's usage in one prompt often if it's anything heavy so haven't messed with astra much.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wpmh0f/claude_pro_20_vs_chatgpt_plus_20_which_one_should/pbytrgm/)
  - Praise, OpenAI Codex, r/codex, 2026-09-24: “i worked a full day today, pretty normal usage on just luna xhigh and i used maybe 5% all day. it did pretty good too, on the level of terra id say, even sol medium.” [source](https://www.reddit.com/r/codex/comments/1woa9fk/codex_usage_feels_brutally_much_better/pbolilg/)

- **Pi (stronger)**. Users pick Pi specifically because its lean base prompt and compaction plugins spend far fewer tokens than rival harnesses.
  Users run head-to-head checks and come back to Pi on token count alone. Posts credit a small system prompt and plugins that compact context before each call. The caveat is self-inflicted. Users note extensions can bloat every request until they trim the tool prompts, and heavy models still drain weekly limits fast.
  Evidence:
  - Praise, Pi, @pidotdev, 2026-08-31: “i tried hermes agent as a replacement of <strict_link> at the beginning i thought it was amazing, it felt different, fast and autonomous. but then i noticed something: "hi" -&gt; 45k cont. tokens (disabling tools/skills) in pi: "hi" -&gt; 7k ct came back to @pidotdev” [source](https://twitter.com/2001267169192226816/status/2094303912958312739)
  - Praise, Pi, r/opencode, 2026-09-21: “i started using pi agent instead of opencode - it is much better in tokens consumption, it has also some plugins to decrease using tokens - compacting context before sending on each call , and some other methods to effectively code with lower tokens consumption.” [source](https://www.reddit.com/r/opencode/comments/1wm7pi6/what_the_hell_is_going_on_with_deepseek/pb7mjre/)
  - Praise, Pi, @pidotdev, 2026-09-08: “@pidotdev @openai i tried with both codex and pi, definitely more efficient with you guys 🤩” [source](https://twitter.com/113087611/status/2097295878633447758)
  - Complaint, Pi, @pidotdev, 2026-09-27: “pi (@pidotdev) itself is lean. but once i added the extensions i actually use, their tool prompts took up 15.7k tokens on every request. i rewrote them. now it's 1.4k, cut by 91%. same features. here's how 🧵 <strict_link>” [source](https://twitter.com/1982027372091367424/status/2104256219666169928)

- **Devin (stronger)**. Mixing a frontier planner with a cheaper in-house worker model lets long runs barely move the limit, users say.
  Praise centres on not paying frontier prices for every step. Users describe hour-long backend tasks barely denting limits and compare favourably against premium plans elsewhere. Complaints exist, mostly errors that exhausted daily usage and regressions, but the routing design is what users single out.
  Evidence:
  - Praise, Devin, @cognition, 2026-09-23: “the fusion mode from @devinai @cognition is amazing! this combination opus 5.5 + swe 2 is unbeatable right now. i have task running for 1hr+ on complex backend + postgres work and my limits are 2% that is insane. i think i only need devin from now on. 🥲” [source](https://twitter.com/160801814/status/2102842376083374243)
  - Praise, Devin, @cognition, 2026-09-11: “@cognition yeah this makes way more sense than paying frontier-model prices for every single step of the run” [source](https://twitter.com/2073113043865620480/status/2098451571998634310)
  - Complaint, Devin, @cognition, 2026-09-27: “there is no reason why using gpt 6 sol would use up my weekly limit on a $200 plan in 2 days i get similar capacity on a $20 @cognition devin plan, using swe-2, and the work it does is just fine @openai is holding back compute even for a model that is cheap to run, and giving the business to @anthropicai unless @thsottiaux reveals something spectacular on dev day, there is no reason to use codex” [source](https://twitter.com/558803042/status/2104317225859436688)
  - Complaint, Devin, @cognition, 2026-09-16: “@dabit3 @learnmore_smart @cognition @devindesktop i dunno, i got the same errors only using the cli and its exhausted my daily usage and half my weekly usage.” [source](https://twitter.com/2322762727/status/2100304622250258794)

- **Claude Code (mixed)**. Users split by model version. The latest flagship release reads as a quota hog, while other users report noticeably slower burn.
  Complaints focus on the newer model eating usage even on planning and subagent execution, and Claude Code leads requests about new versions costing more. Other users say the burn has slowed and that a disciplined plan-first workflow makes implementation cheaper. The experience depends heavily on which model and workflow a user runs.
  Evidence:
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-02: “i got to a point where i could manage the 5hr limits with fable 5 but 5.1 eats through my usage even with basic planning and then sub agent executing. if you let it write code, forget about it. way too expensive.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w5fzyu/so_fable_51_yay_or_nay/p7eu6p2/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-23: “yeah thats exactly what i’ve experienced i used to have the sessions just working at some task and every few minutes i’d check my usage and see its gone down by a couple percent now the burn is much slower. great stuff i just hope it stays like this and they don’t lobotomize it or lower our usage.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wnsn6a/opus_55_is_great/pbhjtk6/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-15: “that matches my experience. most of the work happens before implementation: scope, architecture and small work packages. then the actual release gets much cheaper.” [source](https://www.reddit.com/r/ClaudeCode/comments/1whabup/number_of_posts_that_say_they_reached_their/pa16fmk/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-07: “pro users had access to fable and it didn’t burn usage limits in one prompt, necessarily. it was enough to get something done. fable is supposed to have gotten more efficient since then.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w9xl10/should_pro_users_get_access_to_fable/p8dwywv/)

### Fine print

- Burn-rate claims come from user perception, not metered measurements; several posters say so explicitly.
- Model names appear as users wrote them; versions and plan tiers vary across posts.
- Most agents below Kiro have too few posts to compare reliably.

## Top requests

What users ask to add or change, most asked first. 452 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Author-weeks | Posts | Agents (author-weeks) |
|---|---|---|---|---|
| 1 | Lower overall usage burn rate | 64 | 67 | OpenAI Codex 31, Claude Code 20, Cursor 6, Google Antigravity 5, OpenCode 2 |
| 2 | Lower quota burn for specific premium models | 35 | 35 | OpenAI Codex 23, Claude Code 6, OpenCode 3, Google Antigravity 2, Cursor 1 |
| 3 | Cheaper quota cost for simple tasks | 34 | 36 | OpenAI Codex 24, Claude Code 6, Google Antigravity 2, Cursor 2 |
| 4 | New model versions using more quota | 34 | 35 | Claude Code 17, OpenAI Codex 15, Google Antigravity 2 |
| 5 | Lower quota consumption per prompt or task | 31 | 32 | OpenAI Codex 17, Claude Code 10, Google Antigravity 3, OpenCode 1 |
| 6 | Lower quota burn at higher effort levels | 22 | 22 | OpenAI Codex 13, Claude Code 6, Google Antigravity 3 |
| 7 | More token-efficient model option | 22 | 22 | OpenAI Codex 13, Claude Code 7, GitHub Copilot 1, Pi 1 |
| 8 | Reduce subagent and orchestration token burn | 21 | 22 | OpenAI Codex 11, Claude Code 9, Cursor 1 |
| 9 | Fix sudden quota drain bugs with resets | 20 | 21 | Claude Code 9, OpenAI Codex 9, Google Antigravity 1, Factory 1 |
| 10 | Fix excessive burn in specific features | 15 | 15 | OpenAI Codex 5, Google Antigravity 3, Cursor 2, Claude Code 1, Conductor 1, GitHub Copilot 1, Kiro 1, Pi 1 |
| 11 | More token-efficient agent harness | 15 | 15 | Claude Code 6, OpenAI Codex 4, Cursor 2, OpenCode 2, Google Antigravity 1 |
| 12 | Slower mode with lower usage cost | 13 | 15 | OpenAI Codex 8, Claude Code 4, Cursor 1 |

### 1. Lower overall usage burn rate

- Claude Code, 2026-09-27, r/LocalLLaMA (Reddit): “claude code really is not as polished as you are making it out to be. it has context bloat out the wazoo. and imo, it tries too hard to "be cute". but i mostly care about the tokens it wastes to get the job done. almost every other harness does things better” [source](https://www.reddit.com/r/LocalLLaMA/comments/1wrfp50/another_harness_matters_post_codex_cli_pi_and/pcd9p8d/)
- Google Antigravity, 2026-09-26, r/google_antigravity (Reddit): “the weekly quota is burning in 3 days. it wasn't like this 1 month ago.” [source](https://www.reddit.com/r/google_antigravity/comments/1wqb0ii/dedicated_planning_mode_in_antigravity_is_here/pc438a4/)
- Google Antigravity, 2026-09-26, r/google_antigravity (Reddit): “quota burns too fast. i am surprised u/sounddr ignores these comments.” [source](https://www.reddit.com/r/google_antigravity/comments/1wqb0ii/dedicated_planning_mode_in_antigravity_is_here/pc3lgwq/)

### 2. Lower quota burn for specific premium models

- OpenAI Codex, 2026-09-27, r/codex (Reddit): “even 5.6 sol start draining out the usage limit lol openai need to fix the usage issue for 5.6 sol, it's using too much usage limit. u/codex” [source](https://www.reddit.com/r/codex/comments/1wrkwjw/long_live_gpt_56_sol/pcd9t5f/)
- OpenAI Codex, 2026-09-26, r/codex (Reddit): “actually i feel cheated out when using 6-luna it would be cheaper for me to pay for the api tokens myself than to eat usage burn from it.” [source](https://www.reddit.com/r/codex/comments/1wqvbzz/rate_limits_have_improved_and_i_didnt_even_notied/pc7b5yv/)
- Google Antigravity, 2026-09-24, r/google_antigravity (Reddit): “too bad, even ultra account still really bad today. slow and consume all quota so quickly for the flash model...” [source](https://www.reddit.com/r/google_antigravity/comments/1wp3k02/is_gemini_38_flash_getting_stuck_in_loops_for/pbsfp2w/)

### 3. Cheaper quota cost for simple tasks

- Cursor, 2026-09-26, r/cursor (Reddit): “what do you mean grokbot already ate my heavy plus sub plus 200$ in credits all i want is agents to run and get my work done but asking a question and getting charged $5 isnt a solution” [source](https://www.reddit.com/r/cursor/comments/1wqbkxf/grokbot_eats_5_per_question_why/pc2v25u/)
- Google Antigravity, 2026-09-25, r/google_antigravity (Reddit): “over the past 2 weeks, my gemini weekly quota is basically burning tokens. i never even reached 50% usage by the end of the week. over the last 2 days, i'm already down to 58% remaining with routine, lightweight prompts. can you guys fix this? or maybe tell us what is going on with this?” [source](https://www.reddit.com/r/google_antigravity/comments/1wm7g5r/weekly_quotas_known_issues_support_september_21/pbzqywj/)
- OpenAI Codex, 2026-09-23, r/codex (Reddit): “astra just used 10% of my weekly 20x pro plan on 3 minor ui updates (\~50 lines of code). maybe they fix the token burn rate to start?” [source](https://www.reddit.com/r/codex/comments/1wo3ixp/how_do_you_think_openai_will_respond_next_week/pbkh6in/)

### 4. New model versions using more quota

- OpenAI Codex, 2026-09-22, r/codex (Reddit): “if they release luna 6 i hope it will be cheaper with same capacity honestly i don't need more thinking or speed just less tokens usage burn” [source](https://www.reddit.com/r/codex/comments/1wn67hi/what_models_do_we_think_were_getting_today_and/pbcjuix/)
- Claude Code, 2026-09-16, r/ClaudeCode (Reddit): “fable usage goes by so quick they need to release opus 5.2 already because with the 17% less usage it's a joke. if you think astra uses a lot, at this point even opus uses the same or more than astra. at least codex gets resets” [source](https://www.reddit.com/r/ClaudeCode/comments/1whw423/is_fable_usable/pa5ifcd/)
- OpenAI Codex, 2026-09-11, r/codex (Reddit): “i just wish the credit usage was more reasonable and scaled for each new model. my desire would be older models get cheaper and newer models retain the previous top model usage rates (or darn near). wishful thinking maybe 🤔 - i just know as a pro $100 user i can chew through my usage in hours using astra which just seems excessive. but it has been giving me time to enjoy the sun more lately 🤣” [source](https://www.reddit.com/r/codex/comments/1wdf6rn/2_minutes_and_10_seconds/p9866xv/)

### 5. Lower quota consumption per prompt or task

- OpenAI Codex, 2026-09-27, r/codex (Reddit): “newer models consuming loads of usage.. it's like 1 prompt will consume the 5h limit. codex needs to match up with claude.” [source](https://www.reddit.com/r/codex/comments/1wrfd35/anyone_remembers_54/pcd5wdb/)
- OpenAI Codex, 2026-09-26, r/codex (Reddit): “i am literally on the 200 dollar plan. you think it's normal to be able to blow an entire weeks worth of usage in 30 minutes? regardless of what i am running, it makes no sense.” [source](https://www.reddit.com/r/codex/comments/1wqhikf/wtf_is_going_on_with_usage_drainage/pc45tu3/)
- Claude Code, 2026-09-23, r/ClaudeCode (Reddit): “does it eat anyone else usage like really fast? im on the 20x max plan and it used 30 percent of my daily quota and 11% of my weekly quota in 1 message” [source](https://www.reddit.com/r/ClaudeCode/comments/1wne9k9/well_its_official_its_55_and_not_51/pbigeb0/)

### 6. Lower quota burn at higher effort levels

- Google Antigravity, 2026-09-26, @antigravity (X): “@antigravity antigravity token consumption and model thinking process is very bad” [source](https://twitter.com/985535924434976769/status/2103934368171561251)
- Google Antigravity, 2026-09-25, r/google_antigravity (Reddit): “it should be not possible to burn the quota like this with 3.8f medium..” [source](https://www.reddit.com/r/google_antigravity/comments/1wpzp7k/ag_20_quota_issue/pc26p2e/)
- Claude Code, 2026-09-14, r/ClaudeCode (Reddit): “that's how it seems to me, too. opus 5 high seems to use 20% for a simple check, while medium used 3%. it feels like they break it every two days...” [source](https://www.reddit.com/r/ClaudeCode/comments/1wg535x/limits_are_broken/p9rlsft/)

### 7. More token-efficient model option

- OpenAI Codex, 2026-09-22, r/codex (Reddit): “yes, it’s been better today but it also blew through my entire reset in a few hours - which is the fastest that’s happened yet. they need a model that’s less expensive that’s actually viable. i’m not even going to risk using sol again. sol on xhigh was absolutely less token efficient because it bloated and nuked my whole project and i had to spend more astra tokens unwinding it.” [source](https://www.reddit.com/r/codex/comments/1wn1s2v/it_doesnt_matter_which_flagship_model_i_use_this/pbgq9um/)
- Pi, 2026-09-22, @pidotdev (X): “@pidotdev figure out what makes hermes harness bench better but while also keeping pi usage costs” [source](https://twitter.com/2098546638868353027/status/2102450167693897854)
- OpenAI Codex, 2026-09-19, r/codex (Reddit): “gonna slow roll my usage to make it last enough. astra works too well to downgrade to sol. hoping this can get a version that's more efficient soon. in the meantime, no more goals sadly.” [source](https://www.reddit.com/r/codex/comments/1wkcbn8/astra_soon_even_if_it_burns_fast_its_too_good/)

### 8. Reduce subagent and orchestration token burn

- Claude Code, 2026-09-25, r/ClaudeCode (Reddit): “that is horible advice. paying 2x cost for all subagents, yeah no bro. and auto compact at 200k shits on your work that session. im releasing an update to my governance system soon to properly manage subagents and create savings.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wq7920/two_hidden_settings_to_cut_token_usage/pc25u43/)
- OpenAI Codex, 2026-09-24, r/codex (Reddit): “yes, i did actually. astra on low is really good at that. and yes they both obliterate usage, i really haven't found a good solution because 6 sol is only cheaper in api, but they increased subscrition usage. they just want us to keep struggling! not using astra low though, subagents cost 20k input tokens to launch one so that's no fun with astra if you want to launch them all the time.” [source](https://www.reddit.com/r/codex/comments/1wodcz2/omg_no_wayy/pbpk801/)
- Cursor, 2026-09-21, @cursor_ai (X): “grok bot burns through tokens so ridiculously fast. one hour of 3 agents running, and i've already blown through 32% of my weekly limit. @grok @bot @elon @spacexai @cursor_ai will it always be like this? it's not very useful if all i can do is use it one day per week!” [source](https://twitter.com/1163969994033631232/status/2102070982332518779)

### 9. Fix sudden quota drain bugs with resets

- Google Antigravity, 2026-09-26, @antigravity (X): “please fix the usage limits on @antigravity. used claude opus 4.6 for just 4 minutes, and my usage was already exhausted. 😕 @officiallogank @_mohansolo <strict_link>” [source](https://twitter.com/2090791277269032960/status/2103926087990534444)
- Claude Code, 2026-09-20, @ClaudeDevs (X): “did the bug of claude eating the tokens in 5 minutes come back? @claudedevs” [source](https://twitter.com/291793740/status/2101468585478455326)
- Claude Code, 2026-09-17, @ClaudeDevs (X): “@claudeai @claudeai @claudedevs possible quota bug: before yesterday’s weekly reset, usage had recovered to 67%. after reset, 2 turns with a few subagents burned 56% of the weekly limit — same task used to cost ~20%. please investigate and reset the overcharge. happy to share details.” [source](https://twitter.com/1528655480700121093/status/2100381453431558530)

### 10. Fix excessive burn in specific features

- GitHub Copilot, 2026-09-24, r/GithubCopilot (Reddit): “for now i'm using local harness with autopilot, because the copilot harness with "assisted permissions" was eating my tokens like crazy (the same model, the same kind of tasks). although i liked the copilot harness because i could remotely control it from github app on the phone” [source](https://www.reddit.com/r/GithubCopilot/comments/1wpctlc/vs_code_chat_users_local_or_copilot_harness/pbv9wik/)
- Cursor, 2026-09-23, @cursor_ai (X): “we will build around this. we love the @cursor_ai cloud agents but they seem to drain usage too fast. time for us to embrace cloud agents even deeper. @t3dotcodes will be the shell around it <strict_link>” [source](https://twitter.com/1679064969185316870/status/2102757275391938563)
- Conductor, 2026-09-23, @conductor_build (X): “@conductor_build did you guys ever resolve the usage drainage for anthropic agent sdk? i want to come back to you but this is a huge blocker” [source](https://twitter.com/1639356550627164160/status/2102722482507755542)

### 11. More token-efficient agent harness

- Cursor, 2026-09-26, @cursor_ai (X): “@grok @theaaron @bot @cursor_ai @grok if they keep making the harness heavier with prompting and guard rails, instead of adding lint based code actions, the usage is going to continue to climb. sure wish you would tag some people on the bot team and tell them this.” [source](https://twitter.com/1834188316574359552/status/2103704164417093829)
- OpenAI Codex, 2026-09-21, X search: OpenAI Codex, Codex CLI, Codex app (X): “the codex cli consumes too many tokens, so i'm looking for a way to do something clever with a different lightweight harness.” [source](https://twitter.com/989646189313196032/status/2101907055397519824)
- Claude Code, 2026-09-18, r/ClaudeCode (Reddit): “genuinely doubles my usage saying "plain iso english" even tho it's in my system prompts.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wiujum/claude_code_is_falling_behind_codex_not_because/pampskt/)

### 12. Slower mode with lower usage cost

- Claude Code, 2026-09-18, r/ClaudeCode (Reddit): “could be inferenced at lower quant to reduce cost & increase capacity” [source](https://www.reddit.com/r/ClaudeCode/comments/1wjzei3/frustration_with_opus_and_now_limits/pao7rne/)
- Claude Code, 2026-09-17, @ClaudeDevs (X): “can we get a "slow" mode in #claudecode where requests get sent only when compute is available? i don't mind some tasks taking a few more minutes if it can be done cheaper. @claudedevs” [source](https://twitter.com/1761223168943902720/status/2100390689808973881)
- OpenAI Codex, 2026-09-14, r/codex (Reddit): “anything that’s obvious for a plus user to to use that doesn’t destroy weekly usage q\_q” [source](https://www.reddit.com/r/codex/comments/1wg1zae/gpt6_sol/p9tsvsu/)

## Every agent

| Agent | Overall rank | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|---|
| [Pi](https://feedbackbench.com/agents/pi.md) | 7 | Better than peers | 0.646 | 0.605–0.681 | 84 | 43 | 41 |
| [Devin](https://feedbackbench.com/agents/devin.md) | =5 | Better than peers | 0.584 | 0.543–0.623 | 139 | 44 | 95 |
| [OpenCode](https://feedbackbench.com/agents/opencode.md) | 3 | Typical | 0.537 | 0.489–0.578 | 371 | 74 | 297 |
| [GitHub Copilot](https://feedbackbench.com/agents/copilot.md) | =8 | Typical | 0.530 | 0.479–0.575 | 106 | 22 | 84 |
| [Cursor](https://feedbackbench.com/agents/cursor.md) | 4 | Typical | 0.525 | 0.487–0.558 | 621 | 119 | 502 |
| [Claude Code](https://feedbackbench.com/agents/claude-code.md) | 1 | Typical | 0.517 | 0.497–0.537 | 2700 | 467 | 2233 |
| [Kiro](https://feedbackbench.com/agents/kiro.md) | 13 | Typical | 0.475 | 0.449–0.504 | 38 | 3 | 35 |
| [Google Antigravity](https://feedbackbench.com/agents/antigravity.md) | =5 | Typical | 0.465 | 0.419–0.508 | 348 | 47 | 301 |
| [OpenAI Codex](https://feedbackbench.com/agents/codex.md) | 2 | Worse than peers | 0.463 | 0.449–0.476 | 5123 | 715 | 4408 |
| [Cline](https://feedbackbench.com/agents/cline.md) | =8 | Too few posts | – | – | 22 | 3 | 19 |
| [Amp](https://feedbackbench.com/agents/amp.md) | =11 | Too few posts | – | – | 21 | 4 | 17 |
| [Factory](https://feedbackbench.com/agents/factory.md) | =11 | Too few posts | – | – | 19 | 8 | 11 |
| [Warp](https://feedbackbench.com/agents/warp.md) | 15 | Too few posts | – | – | 11 | 3 | 8 |
| [Zed](https://feedbackbench.com/agents/zed.md) | =8 | Too few posts | – | – | 6 | 3 | 3 |
| [Conductor](https://feedbackbench.com/agents/conductor.md) | 14 | Too few posts | – | – | 4 | 1 | 3 |
| [Augment Code](https://feedbackbench.com/agents/augment.md) | 17 | Too few posts | – | – | 4 | 0 | 4 |
| [Grok Build](https://feedbackbench.com/agents/grok-build.md) | 16 | Too few posts | – | – | 2 | 1 | 1 |

## Posts

Receipts rule: The 5 most recent praise and complaint posts per area (first 700 characters) and 3 per criterion (first 450 characters).

### Pi

- Praise, 2026-09-26, r/LocalLLaMA (Reddit): “can you actually remove / override the heavy system prompt though? according to my searching (and asking llms) you can only *add* to the system prompt. also, even when you enable only the bare minimum tool calling like file read / write and a couple others it's still like 8k tokens for even just a 'hello world' style prompt. some of us can run models relatively smoothly on minimal hardware with pretty good results but we need to manage the tokens” [source](https://www.reddit.com/r/LocalLLaMA/comments/1wq9ivr/what_ide_to_use_for_local_models/pc3lhdp/)
- Praise, 2026-09-24, r/PiCodingAgent (Reddit): “completely off vibes, but experience with my multi-agent harness agrees. luna-6 is an absolute monster per dollar. i've shifted to running it almost exclusively, only shifting up when it fails or qa complains.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wpcis6/opus_55_vs_gpt6_sol_luna_in_piagent_results_on_my/pbv5dmm/)
- Praise, 2026-09-23, r/PiCodingAgent (Reddit): “i would appreciate it so much if you could run agent kernel for swe-bench / repo bugfixes kind of benchmark. a quick note on the charts in my readme: those results came from purely pi + agent kernel standalone, without any extra skills or extensions (specpi or jev) to answer your question on where it's expected to shine: it was built specifically for medium-to-large codebases with verbose test suites. in contrast, on terminal-bench tasks, the co” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wo4qke/piagentkernel_focused_code_retrieval_grounded/pblbiz8/)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “why is it that opus 5.5 usage starts out great in omp, and then for one reason or another starts draining usage like crazy? am i hitting some sort of caching issues? do i need to compact at 400k? or what? i am pretty sure the cache is being kept warm for an hour, but i don't know for sure how this stuff works. it's compacting at 850k. if i've left it for awhile and context is >35% i will use /handoff and that seems to help.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wruev9/omp_and_opus_55_usage/)
- Complaint, 2026-09-27, @pidotdev (X): “pi (@pidotdev) itself is lean. but once i added the extensions i actually use, their tool prompts took up 15.7k tokens on every request. i rewrote them. now it's 1.4k, cut by 91%. same features. here's how 🧵 <strict_link>” [source](https://twitter.com/1982027372091367424/status/2104256219666169928)
- Complaint, 2026-09-26, @pidotdev (X): “@shantanugoel @pidotdev 24 hours nonstop, 1b+ tokens. at some point this stops being a coding agent and becomes a very expensive roommate that never sleeps.” [source](https://twitter.com/1765671748282851328/status/2103748516644544835)

### Devin

- Praise, 2026-09-27, r/ClaudeCode (Reddit): “only if you combined it with $20 devin or $10 opencode go. if you combine with with devin, you can instruct claude to use swe-2 for exploration and you reduce tokens by like 20-50%. $20 of claude is great when combination with other things, but not alone for actual work. i go really far with the $20 of claude, but i have a lot of other subs, agents running 24/7, and the $20 of claude is just for more double checking and single harder tasks.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrf0ld/is_claude_pro_actually_worth_20_just_for_one/pcch4ql/)
- Praise, 2026-09-25, @cognition (X): “share two real cases of using devin cli: 1. running tasks with strong models + swe-2 in codex cli and devin cli, experiencing over 10 hours of uninterrupted tasks in the last two days, with failures. 2. running tasks for 12 hours using devin cli's fusion (opus-5.5 medium + swe-2 medium), only 4% of the weekly quota of the max package was used. #devin @cognition <strict_link>” [source](https://twitter.com/142110760/status/2103284980822757586)
- Praise, 2026-09-25, @cognition (X): “@itsalicesoul really good. token cost was less in devin (offered). @cognition” [source](https://twitter.com/758692944983437313/status/2103359531392843953)
- Complaint, 2026-09-27, @cognition (X): “devin from @cognition is surely a great workhorse but it burnt through my entire weekly quota in less than 8 hours 🤯” [source](https://twitter.com/15290915/status/2104239118419112032)
- Complaint, 2026-09-27, @cognition (X): “@learnmore_smart @devindesktop @cursor_ai @cognition @anysphere @devinai @xai yeah this is not just output tokens btw it’s more like input + output tokens, so it gets quite expensive” [source](https://twitter.com/963278001474453504/status/2104289960807350526)
- Complaint, 2026-09-27, @cognition (X): “there is no reason why using gpt 6 sol would use up my weekly limit on a $200 plan in 2 days i get similar capacity on a $20 @cognition devin plan, using swe-2, and the work it does is just fine @openai is holding back compute even for a model that is cheap to run, and giving the business to @anthropicai unless @thsottiaux reveals something spectacular on dev day, there is no reason to use codex” [source](https://twitter.com/558803042/status/2104317225859436688)

### OpenCode

- Praise, 2026-09-27, r/opencode (Reddit): “command code is 10 times worse, the models in command code are so dumb that they burn through my quota in a loop of stupidity, i did a test on this plan and in one day it exceeded my weekly quota and hit 50% of the monthly in a $10 plan, you know how much an opencode plan spends per day for me? 4% and solves the problems (i use muse spark 1.3)” [source](https://www.reddit.com/r/opencode/comments/1wqy8pq/how_true_is_it_that_opencode_go_models_are/pcbbj4r/)
- Praise, 2026-09-27, r/opencode (Reddit): “yeah exactly, that's pretty much how i use it too. flash/go for the cheap bulk stuff so claude stays fresh for the things that actually matter. works well once you split it that way” [source](https://www.reddit.com/r/opencode/comments/1wrg81n/go_subscription_is_slower_deepseek_flash_41/pccgmbe/)
- Praise, 2026-09-26, r/opencode (Reddit): “v2 is much better, it adds codemode to save tokens, search for mcp tools, and it also has this idea where everything can be modified/customized.” [source](https://www.reddit.com/r/opencode/comments/1wqppth/opencode_v2/pc6i087/)
- Complaint, 2026-09-27, r/opencode (Reddit): “yeah the insane amount of thinking, is even making it as expensive as deepseek for me, i'll stay with muse.” [source](https://www.reddit.com/r/opencode/comments/1wqfj7j/will_there_always_be_the_contributor_models/pc9wb3j/)
- Complaint, 2026-09-27, r/opencode (Reddit): “muse started tweaking for me at one point where even 20k tokens where costing me about 0.8$ thats when i stopped using it is it good now ?” [source](https://www.reddit.com/r/opencode/comments/1wqfj7j/will_there_always_be_the_contributor_models/pc9xiy0/)
- Complaint, 2026-09-27, r/opencodeCLI (Reddit): “no wonder my token plan was fully drained in just three days. any reset?” [source](https://www.reddit.com/r/opencodeCLI/comments/1wrtyfb/xiaomi_is_updating_mimo_26/pcgy6q0/)

### GitHub Copilot

- Praise, 2026-09-27, @GitHubCopilot (X): “lol, that feature alone just cost $6 🤑 with @githubcopilot and @anthropicai claude fable 5. took about 5 minutes. no way for a human to beat it. <strict_link>” [source](https://twitter.com/22220922/status/2104250323133202563)
- Praise, 2026-09-25, r/GithubCopilot (Reddit): “my own eval (trading related) it did worse than 5.6luna (both max). it seems to use a lot less tokens compared to any 5.6 model though.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wptkkq/how_does_gpt6_sol_feel_so_far/pbymohy/)
- Praise, 2026-09-25, r/GithubCopilot (Reddit): “it really depends on a lot of factors but no i was using it all day and most of my requests were in the 1 credit or less range.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wq39rv/did_they_10x_the_cost_of_luna_56_since_the/pc0qqrv/)
- Complaint, 2026-09-27, r/GithubCopilot (Reddit): “like shit. using 5.6. sol 6 was stuck in a loop and burned 90eur switching between the two exact solutions without stopping” [source](https://www.reddit.com/r/GithubCopilot/comments/1wptkkq/how_does_gpt6_sol_feel_so_far/pccijxn/)
- Complaint, 2026-09-26, r/GithubCopilot (Reddit): “glad youuu got it sorted, because random credit spikes like that usually just mean a sneaky copilot update started attaching wayyy more of your workspace files as context by default....” [source](https://www.reddit.com/r/GithubCopilot/comments/1wq39rv/did_they_10x_the_cost_of_luna_56_since_the/pc3m7s7/)
- Complaint, 2026-09-26, r/GithubCopilot (Reddit): “i had a problem the other week where despite 5.6 luna being the only model enabled in settings, 5.6 luna was delegating tasks to sub-agents on expensive models such as sonnet 5 for basic tasks which burned credits. i'd check to make sure another model isn't being used somewhere you're not aware of.” [source](https://www.reddit.com/r/GithubCopilot/comments/1wq39rv/did_they_10x_the_cost_of_luna_56_since_the/pc4mu1j/)

### Cursor

- Praise, 2026-09-27, @cursor_ai (X): “@cursor_ai the 7% headline is smaller than the real lesson: agent economics live in the harness. prompt scope, tool loading, caching, and file reads compound into better latency and lower cost without touching the model weights.” [source](https://twitter.com/1899381125451218944/status/2104202426274222427)
- Praise, 2026-09-26, r/cursor (Reddit): “last week, i used grok 4.7 4.6 medium but it seems opus 5.5 uses fewer tokens so i switched to opus.” [source](https://www.reddit.com/r/cursor/comments/1wpt19x/which_model_you_use_most_of_the_time/pc3ofc4/)
- Praise, 2026-09-26, @cursor_ai (X): “@cursor_ai a 7% token cut with the same agent quality is neat. curious which of those savings you'd notice first day to day.” [source](https://twitter.com/2025483882263707651/status/2103664833451487308)
- Complaint, 2026-09-27, r/cursor (Reddit): “opus in cursor absolutely ate through my ultra in a few hours unfortunately.” [source](https://www.reddit.com/r/cursor/comments/1wrbmrn/is_it_a_waste_of_money_to_run_opus_55_inside_of/pcbaloh/)
- Complaint, 2026-09-27, r/cursor (Reddit): “"other models" simply means your monthly subscription x2 as tokens, so you spend almost 200 usd in a hour (44%), which model and settings are you using?” [source](https://www.reddit.com/r/cursor/comments/1wrddz6/why_is_the_other_models_usage_is_being_used_so/pcbro0f/)
- Complaint, 2026-09-27, r/cursor (Reddit): “bro i’ve used 1.2b last week 💀” [source](https://www.reddit.com/r/cursor/comments/1wpd3mm/i_hate_to_admit_it_but_grok_sucks/pcbtyfq/)

### Claude Code

- Praise, 2026-09-27, r/ClaudeCode (Reddit): “i’ve been using the shit out of opus lately. i always use sonnet as my interface, interactive sessions are always sonnet 5 on medium. i’ve been having sonnet spin up an opus 5.5 subagent in plan mode as an advisor. in plan mode, opus can’t call tools and blot its context window and my usage ($20/mo), sonnet has to feed it all the context and prompt it. it delivers amazing value an just sips usage. has made my sessions way more effective” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr5lzf/practically_speaking_what_tasks_fit_into/pc9uw06/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “anyone, the use of opus is ultra efficient.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr5lzf/practically_speaking_what_tasks_fit_into/pc9v7v1/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “the efficiency gains are insane. it's opus 5.5 medium is more efficient than sonnet 5 high. it's honesty become my daily use because the quality is so much better lol.” [source](https://www.reddit.com/r/ClaudeCode/comments/1woe3lo/is_opus_55_really_better_than_fable_in_your/pca63pb/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “i don't really use sonnet for anything. i switched to opus because i kept running out of usage - i found i use more with sonnet even though it's cheaper because it kept getting stuck and making mistakes.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wr5lzf/practically_speaking_what_tasks_fit_into/pc9u4xg/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “are you using opus 5.5 as your daily driver now? my config is fable 5.1 is the agent i talk to, but all my subagents are opus 5.5. i got tired of slop so wanted to ensure quality. but still burning through creits really fast. wondering if there is a better config i can adopt.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqsz67/opus_55_has_absolutely_restored_value_to_the_200/pca8fx1/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “agent workflow size config was set to small in my settings (<5 agents) + my prompt said verbatim “do not create more than 3 subagents , not a large swarm”………….. my result? ——-> ofc, no other than:🙃 my *entire* weekly pro20x \~\~ *sautéed* ***\~***in front of me on day 1/7 🥲🫠🫠🙃🙃🙃😆😆😆” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqagzu/claude_added_graceful_stopping_point_in_new_update/pcagfdf/)

### Kiro

- Praise, 2026-09-27, r/kiroIDE (Reddit): “codex is insane atlist usage wise i haved used gpt 6 luna for 5hrs maybe more it only consumed 2% of weekly limits” [source](https://www.reddit.com/r/kiroIDE/comments/1wq8e5k/claude_opus_55_is_finally_here_lessss_gooo/pcbrrzv/)
- Praise, 2026-09-11, r/kiroIDE (Reddit): “i use kiro ide with auto model. works very well, few tokens.” [source](https://www.reddit.com/r/kiroIDE/comments/1wchsii/frontier_models_on_other_providers_vs_the_same/p92st2r/)
- Praise, 2026-08-31, r/kiroIDE (Reddit): “compared to other tools such as copilot, kiro uses up credits relatively slowly (and yes, i mean credits per cost). of course, where possible, you should try to optimise your llm usage a little, avoid unnecessary context, etc., but depending on what you’re doing and which model you’re using, this level of consumption may be normal.” [source](https://www.reddit.com/r/kiroIDE/comments/1vz8tk7/kiro_is_guzzling_credits_1k_in_two_days/p71rhqd/)
- Complaint, 2026-09-27, r/kiroIDE (Reddit): “only usable "model" in kiro right now is auto... all other decent ones burn credits like crazy. if aws prices luna/sol correctly and add back the new chinese models (deepseek v4.1 flash please!)... then it can return - otherwise... it will be used by the ones that are using it for free or when their employeer "strongly recommend" it to be used.” [source](https://www.reddit.com/r/kiroIDE/comments/1wqueq0/claude_pro_vs_kiro_pro_subscription_which_is/pcfhx04/)
- Complaint, 2026-09-26, r/kiroIDE (Reddit): “with my 150 credits left, i can probably fit 4 requests in!” [source](https://www.reddit.com/r/kiroIDE/comments/1wq8e5k/claude_opus_55_is_finally_here_lessss_gooo/pc386s6/)
- Complaint, 2026-09-26, r/kiroIDE (Reddit): “do you have model governance enabled? if yes, then i don't think i may be of much help sadly. you might have to reach out to aws support. i'll just throw this out there, after opus 5.5 dropped in kiro, there's really no reason to be using fable 5.1. fable is quite literally slowler and much much more expensive than 5.5 <strict_link>” [source](https://www.reddit.com/r/kiroIDE/comments/1wpveza/kiro_and_opus_55/pc5nd9m/)

### Google Antigravity

- Praise, 2026-09-27, r/google_antigravity (Reddit): “use flash 3.8 it is far more better than other versions of gemini. also consumes less tokens than codex or claude code for similar effort.” [source](https://www.reddit.com/r/google_antigravity/comments/1wrbsby/antigravitys_gemini_models_are_worst_in_among_all/pcbmg55/)
- Praise, 2026-09-27, r/google_antigravity (Reddit): “first of all install ponytail as a skill to anti-gravity. it should start thinking a lot less, spending a lot fewer tokens, and writing a lot less but better code. and then just mention it explicitly: "recently in some of your runs you did this" (you took too many screenshots, checked things that weren't necessary etc.). just mention everything and then it will actually stop doing those things in the follow-up. it will explicitly start saying in” [source](https://www.reddit.com/r/google_antigravity/comments/1wqf2o6/worst_model/pcchcst/)
- Praise, 2026-09-27, @antigravity (X): “@ogsada @antigravity the token cost is nothing imho, it's usually done at the start and it's just one file” [source](https://twitter.com/1975543058575024128/status/2104261011708764454)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “$20 is fine if you babysit every task manually. but go agentic? 💀 that $20 disappears faster than your motivation on monday. a 5-hour task becomes 20–40 minutes… if the agents survive that long. 😂” [source](https://www.reddit.com/r/google_antigravity/comments/1wpvvsr/why_does_antigravity_not_update_its_offerings_for/pca8ui9/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “well i've met a lot of desktop app users who come to my discord server and start screaming that antigravity is shit. then i made them use cli and they've always been happy since. i think the main problem people hate with the gui is the insane token burn and the gui talking up all the resources of your computer. that's why i switched in the first place to get rid of the insane token burn. plus the cli usually only has bugs and doesn't normally exp” [source](https://www.reddit.com/r/google_antigravity/comments/1wqmd4n/why_is_the_antigravitycli_so_underrated/pcamrq4/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “flash 3.8 is gemini's best model for most coding. it's in a fairly similar ballpark to luna and sonnet (makes more mistakes, needs more skill supports/scaffolding and validation/verification and human direction/planning support especially with ux/design tasks compared to flagships)... basically not as intelligent and capable as sol/terra/astra or opus/fable but still useful and relatively cheap/efficient/fast. flash 3.7 is faster but less intelli” [source](https://www.reddit.com/r/google_antigravity/comments/1wrbsby/antigravitys_gemini_models_are_worst_in_among_all/pcbjcha/)

### OpenAI Codex

- Praise, 2026-09-27, r/codex (Reddit): “5.5 is like sol on high but consumes like 1 token” [source](https://www.reddit.com/r/codex/comments/1wr69nb/see_you_soon_guys_probably/pca5o55/)
- Praise, 2026-09-27, r/codex (Reddit): “fuck yea. astra ultra. had it create 30 small applications, 20 high res images and a huge addition to another application for 10%. a change like this the other day would have eaten a third of it b” [source](https://www.reddit.com/r/codex/comments/1wr71h7/has_anyones_astra_become_more_generous_on_their/pca7p32/)
- Praise, 2026-09-27, r/codex (Reddit): “not astra, but i noticed the same thing. i was using sol 6 high about 12 hours ago, then the reset happened. i just used sol 6 high again, and the quota seems to be going down noticeably slower. i was actually wondering if something had changed.” [source](https://www.reddit.com/r/codex/comments/1wr71h7/has_anyones_astra_become_more_generous_on_their/pca9gy9/)
- Complaint, 2026-09-27, r/codex (Reddit): “lmao this guy is on point. codex used to be amazing with $200 in terms of limit caps. i would have to be dev super hard for 12 hours a day to get my limit down to 10%. now you can burn that in two days. meanwhile i have claude 5x $100 sub and that takes effort to max out every week. thankfully fable makes it easy.” [source](https://www.reddit.com/r/codex/comments/1wr1oir/they_are_aware_and_working_on_it_apparently_just/pc9sa8z/)
- Complaint, 2026-09-27, r/codex (Reddit): “kkkkkk... i just enjoyed the free reset and didi a task that the claude code, opus 5.5 at max would spend like 3%, odf the w limit....the sol 6 ultra took 47%... thats ultrageous...” [source](https://www.reddit.com/r/codex/comments/1wqvl66/reset_just_came_in/pc9u5pj/)
- Complaint, 2026-09-27, r/codex (Reddit): “i hate to be that guy, but usage seems cut even more after the reset.. i don't want anymore resets... i want a set amount of usage that is fair and doesn't get cut every week. i mean honestly i've barely got any work done today, i used chatgpt and github connector to do a lot, and i have one thread going with astra xhigh and was using medium earlier and i'm at 49% usage after 6 hours or so coding? this is the fastest i've ever seen it drain... i” [source](https://www.reddit.com/r/codex/comments/1wr4cp6/more_resets_incoming_next_week/pc9u99o/)

### Cline

- Praise, 2026-09-26, @cline (X): “how have you been with pixel canary? in my humble opinion, running from @cline desktop. very accurate in reasoning, superior outputs. i haven't been able to measure token consumption. i've been pushing it hard since yesterday and it hasn't given me any rate limit yet. sometimes i have felt that it disconnects (i imagine it's due to high request traffic). a bit slow if i notice it. but i repeat, very accurate in the solutions it presents.” [source](https://twitter.com/165473256/status/2103900399967297747)
- Praise, 2026-09-17, @cline (X): “@positronx_ @threejs @cline the view through that window is doing way less work than the room around it, respect for finding where to cut corners with 580m tokens on the line” [source](https://twitter.com/1384952635002851330/status/2100609158051451285)
- Praise, 2026-09-14, @cline (X): “i trained the fly to become a gymbro - to do bicep curls and also do squats for a good leg day. mapped all 139,255 proofread neurons, 50m+ synapses, and neurotransmitters. everything was done using the open-source deepseek-v4-flash model on the @cline 's desktop app, an open-source app for open-weight models that i had early beta access to. 68 bodies / 102–103 joints / 78 actuators - flybody as the base model running on the mujoco physics engine.” [source](https://twitter.com/1297627095330242561/status/2099539937314099433)
- Complaint, 2026-09-26, @cline (X): “@iamdavidhill space bunny seems to hog up a lot of tokens, i was making a little side prjoect cli dev tool with it and it hogged up 1+ billion tokens overnight. on @cline cloud” [source](https://twitter.com/834280176313835520/status/2103916234161242207)
- Complaint, 2026-09-25, r/CLine (Reddit): “bro just one prompt and your usage is gone and it even finishes midpoint.” [source](https://www.reddit.com/r/CLine/comments/1wpyepq/gemini_38_flash_is_now_free_in_cline/pbzgvsl/)
- Complaint, 2026-09-25, @cline (X): “@cline gemini flash free access to cline. the cheap part is the speed, the expensive part is the context.” [source](https://twitter.com/2002993609671630848/status/2103309713509064975)

### Amp

- Praise, 2026-09-23, @AmpCode (X): “i got so scared of running out of limit with codex that using opus 5.5 and not looking at the ridiculous drain on my usage is such a relief. using it in @ampcode with fast mode enabled and everything is going so good” [source](https://twitter.com/2931128860/status/2102567936904826896)
- Praise, 2026-09-14, @AmpCode (X): “i am convinced that for some reason the chatgpt desktop app chews tokens because i've been rolling @ampcode all afternoon with sol/astra and have only used like 4% from 50% starting. something is sussy but i guess this is why i prefer using ampcode more and more everyday.” [source](https://twitter.com/22610703/status/2099432538146254876)
- Praise, 2026-09-14, @AmpCode (X): “sol consumes more than astra according to their benchmarks. astra orchestrating luna xhigh is the most you can squeeze out of the sub. usage of the sub through @ampcode goes almost as long but way higher quality than luna, they have some secret sauce. astra medium is the third best for out of the box task completed vs sub consumed.” [source](https://twitter.com/4186365072/status/2099541807998603771)
- Complaint, 2026-09-27, @AmpCode (X): “@ampcode @sqs @thorstenball the cheaper model thing is real for me. flash model on a refactor did 4 extra tool loops fixing its own edit and the bill went up. do the orbs evals show that or just latency?” [source](https://twitter.com/1835841692852682752/status/2104230660105781726)
- Complaint, 2026-09-22, @AmpCode (X): “the lead + subagents setup is the first agent workflow that actually mirrors how real teams work. i tried one big agent doing everything and it just thrashed context between tasks. splitting orchestration from execution is what made it click. ngl the cost of 3 orbs is still making me wince though.” [source](https://twitter.com/1085377722237546504/status/2102372946790514884)
- Complaint, 2026-09-20, @AmpCode (X): “@sqs @ampcode @vercel seat-exempt is the easy part. the ai committer on this account once pushed 114 commits in a day, most of them buying a minute of build time just to decide nothing should build. it's the meter that gets you.” [source](https://twitter.com/2013029907966599169/status/2101641357253062918)

### Factory

- Praise, 2026-09-27, @droid (X): “@droid factory cuts inference cost by precomputing common sub‑expressions and reusing them across requests so each new run only evaluates delta changes” [source](https://twitter.com/195841906/status/2104089474485637594)
- Praise, 2026-09-27, @droid (X): “@droid dying to test droid 👋👋👋🤩🤩 pretty amazing that you have been able to cut interface costs especially in this environment” [source](https://twitter.com/1367818563495596032/status/2104109772974744033)
- Praise, 2026-09-23, @FactoryAI (X): “@factoryai @anthropicai the 20-25% fewer output tokens at equal effort is the detail that compounds in agent traces: shorter reasoning chains mean less context rot per step and cheaper retries. token frugality is an underrated spec for long-running work.” [source](https://twitter.com/2098377213426991107/status/2102754814870618580)
- Complaint, 2026-09-23, r/FactoryAi (Reddit): “5-hour normal limit is maxed monthly droid core limit is maxed thus you are sol until 5 hour limit is up. droid core models use normal limits first, then fall back to droid-core limits” [source](https://www.reddit.com/r/FactoryAi/comments/1wnh0g4/conflicting_usage_stats/pbjg2zn/)
- Complaint, 2026-09-22, @FactoryAI (X): “hey @factoryai @droid please do this i really think standard + droid core quotas should work differently across all plans. if i’m using droid core models, let that usage come from the droid core quota instead of burning standard first. the ideal setup would be: fable handles orchestration + the important reasoning, while subagents run on droid core quota. that would make the separate quotas way more useful and make agent-heavy workflows last sig” [source](https://twitter.com/1625280993966923777/status/2102347962752061679)
- Complaint, 2026-09-22, @droid (X): “@theo 10 minute of work, exhaust my @droid 's 22% of limit. that's a no go for small devs” [source](https://twitter.com/2995471962/status/2102219071827976480)

### Warp

- Praise, 2026-09-04, @warpdotdev (X): “@warpdotdev 63% cost cut is no joke” [source](https://twitter.com/1858703675864334336/status/2095671913851011086)
- Praise, 2026-09-04, @warpdotdev (X): “interesting picked the same 3 models for a project in warp factory and avg pr cost was around 30 to 50 % lower. now half way trough the project and doing a optimization run first before more project work items see if we can get that number lower. so far really happy with the factory preview 👌” [source](https://twitter.com/141649554/status/2095943106507973029)
- Praise, 2026-09-03, @warpdotdev (X): “@warpdotdev brilliant idea 🔥 real tasks real results. 63% cut is massive” [source](https://twitter.com/1790460165000687620/status/2095616963460628778)
- Complaint, 2026-09-22, @warpdotdev (X): “ran out of my monthly @warpdotdev credits from a single prompt change to a docker container. maybe the legacy plan i'm on is just useless now?” [source](https://twitter.com/80764812/status/2102534429369303091)
- Complaint, 2026-09-22, @warpdotdev (X): “@jmitch @warpdotdev one prompt change eating the month does make that legacy plan feel useless” [source](https://twitter.com/1549055479875342336/status/2102537760498401692)
- Complaint, 2026-09-15, r/warpdotdev (Reddit): “i chose gpt 5.6 luna xhigh as my model, when it failed, fallback model opus 5 max continued, which rapidly consumes more credits. i can't stand it.” [source](https://www.reddit.com/r/warpdotdev/comments/1oa0abo/warp_dirty_tactics_sonnet_45_thinking_uses_cheap/pa1hdbn/)

### Zed

- Praise, 2026-09-25, r/ZedEditor (Reddit): “i work by myself most of the time and i really like the review process (probably you could get something similar with a skill), but i also get better usage there than on the codex app with my codex sub (probably context or cache), so it became my first ai coding app this week” [source](https://www.reddit.com/r/ZedEditor/comments/1wq03mv/has_anyone_tried_delta/pc17ujd/)
- Praise, 2026-09-09, r/codex (Reddit): “just wanted to share my personal experience in case anyone else is running into the same issue. i've been using codex through the vs code extension, and i noticed that i was going through my usage pretty quickly. recently, i started using codex with zed instead, and from what i've seen so far, my usage has been noticeably lower while getting pretty much the same results. i haven't done any proper benchmarks or controlled tests, so i'm not saying” [source](https://www.reddit.com/r/codex/comments/1wbraq6/my_codex_usage_dropped_significantly_after/)
- Praise, 2026-09-07, @zeddotdev (X): “@ikhwanuddin @zeddotdev for a ui reason gemini 3.8 enak buat frontend dan kuotanya abisnya lebih lamaaa hahaha” [source](https://twitter.com/131447449/status/2096959789528138219)
- Complaint, 2026-09-13, @zeddotdev (X): “@arpit_bhayani same but i moved to @zeddotdev after facing the same issue😆 instead of burning tokens” [source](https://twitter.com/1382512535509688322/status/2099142506470375841)
- Complaint, 2026-09-04, @zeddotdev (X): “got @zeddotdev delta beta access, blindly claimed zed vip, and accidentally burned through the free $100 credits in half a day. switched back to omp flash 3.8 (150–350 t/s)—which usually feels blazing fast—and it suddenly feels like a total snail 🐌. zed completely smokes it. <strict_link>” [source](https://twitter.com/1376195389305409539/status/2095857452592091204)
- Complaint, 2026-09-04, @zeddotdev (X): “@zeddotdev fuck i used most of credits before the release. i am so stupid.” [source](https://twitter.com/853224743536910336/status/2095954932872454320)

### Conductor

- Praise, 2026-09-26, @conductor_build (X): “@kyberpez @conductor_build isolated worktrees plus a cheaper architect model is basically the fix for the idle capacity i was complaining about. gonna try conductor, what happens when two worktrees touch the same file, does it merge clean or do you sort that out by hand?” [source](https://twitter.com/1152636974/status/2103821399752425538)
- Complaint, 2026-09-25, @conductor_build (X): “@aisosa_d @chatgpt @conductor_build astra wanted to kill me. 5 questions and usage done on max subscription” [source](https://twitter.com/337722085/status/2103337496842993671)
- Complaint, 2026-09-23, @conductor_build (X): “@conductor_build did you guys ever resolve the usage drainage for anthropic agent sdk? i want to come back to you but this is a huge blocker” [source](https://twitter.com/1639356550627164160/status/2102722482507755542)
- Complaint, 2026-09-12, @conductor_build (X): “@garrytan @charlieholtz @steve_yegge @conductor_build conductor is good but their harness for anthopic and openai models consumes a lot more tokens than cc and codex” [source](https://twitter.com/1557669964085358592/status/2098798329799008328)

### Augment Code

- Complaint, 2026-09-23, G2 (G2): “q: what problems is the product solving and how is that benefiting you? a: it takes away all the pain of digging through massive codebases manually tracking down cross-file dependencies. i save about 1-2 hours of grunt work every single day just in refactoring and boilerplate. q: what do you like best about the product? a: the context is really amazing. in contrast to a simple ai autocomplete function, it creates an index of all of your multi-rep” [source](https://www.g2.com/products/augment-code/reviews/augment-code-review-13582326)
- Complaint, 2026-09-18, r/cscareerquestions (Reddit): “have you tried switching to lower models? the higher ones are overkill if you're using it to augment the engineering that you are doing rather than trying to rely on them to do the design/software architecting for you. sonnet 5 is plenty capable when given clear accurate instructions, and it's faster and much more token efficient. i'll use opus occasionally for some harder tasks, but i find both opus and fable tend to over complicate everything,” [source](https://www.reddit.com/r/cscareerquestions/comments/1wjk5vg/anyone_frustrated_with_how_ai_harnesses_are/paklj9u/)
- Complaint, 2026-09-11, r/vibecoding (Reddit): “i used augment code almost from the beginning, after spending £300 on credits in one month after the price hike i cancelled and moved away. i use a mac, intellidea and do mostly flutter apps, databases, websites - not basic apps either. i moved to zencoder, a platform i never really took any notice of because augment code was so good (i thought). but for me it was the best switch ever. for my use case it outshines augment code in every way. its f” [source](https://www.reddit.com/r/vibecoding/comments/1vqvmlv/did_anyone_ever_use_augment_code_im_looking_for/p94qwc4/)

### Grok Build

- Praise, 2026-09-11, r/ClaudeCode (Reddit): “i use fable as my orchestrator, sonnet recon and headless grok build. will shift to sonnet/opus build once the heavy grok discount runs out. it is slower, a lot slower but token efficient and all the models add a layer of checking the others work. less baby sitting and cleaner code.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wd9eio/fable_as_orchestrator_and_opussonnet_as_executers/p9574gb/)
- Complaint, 2026-09-19, r/cursor (Reddit): “i've been using cursor for the past 10 months, first with the codex ide extension, and the last 45 days with everything else the same, but with the privoder switched to deepseek. the 8 billion tokens i've used over the past 45 days i paid us$**71.57** for - i spent about $800-900 for the previous 13-14 billion tokens with openai (lots of resets used judiciously). we're talking extremely cache heavy, like 98% input, 98% cached, with many workloads” [source](https://www.reddit.com/r/cursor/comments/1wk16qb/so_what_happened_to_cursor_in_the_past_few_weeks/papt0dm/)
