# Feedback Bench: Coding Agents, Ranked by their users’ feedback Built 2026-10-01. Window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/. Built by Enterpret (https://www.enterpret.com), which runs the same feedback analysis on a company's own customer feedback. Feedback Bench ranks 17 coding agents from what their users say in public: Reddit, X, G2 and Trustpilot. Every post is labelled against 63 criteria in 9 areas. The Feedback Score is 100 × √(Popularity × Customer love): Popularity is how many people post about the agent (log share of authors, leader = 1); Customer love is how those posts rate it against the category (0.5 = category norm). Full method: [https://feedbackbench.com/method.md](https://feedbackbench.com/method.md). ## Ranking Rank shows “=N” when rank ranges overlap. Criteria better / worse than peers counts the 63 criteria where the agent's 95% interval lies wholly above / below 0.5. | Rank | Agent | Maker | Feedback Score | 95% interval | Rank range | Popularity | Customer love | Customer love 95% interval | Criteria better / worse than peers | |---|---|---|---|---|---|---|---|---|---| | 1 | [Claude Code](https://feedbackbench.com/agents/claude-code.md) | Anthropic | 70.4 | 69.9–70.9 | 1–1 | 0.986 | 0.503 | 0.496–0.510 | 14 / 10 | | 2 | [OpenAI Codex](https://feedbackbench.com/agents/codex.md) | OpenAI | 67.4 | 66.9–67.8 | 2–2 | 1.000 | 0.454 | 0.448–0.460 | 7 / 21 | | 3 | [OpenCode](https://feedbackbench.com/agents/opencode.md) | Anomaly (open source) | 66.0 | 65.4–66.7 | 3–3 | 0.773 | 0.564 | 0.553–0.575 | 7 / 7 | | 4 | [Cursor](https://feedbackbench.com/agents/cursor.md) | Anysphere | 59.7 | 59.0–60.5 | 4–4 | 0.708 | 0.504 | 0.491–0.516 | 5 / 10 | | =5 | [Devin](https://feedbackbench.com/agents/devin.md) | Cognition | 55.9 | 55.1–56.5 | 5–6 | 0.514 | 0.607 | 0.591–0.622 | 9 / 2 | | =5 | [Google Antigravity](https://feedbackbench.com/agents/antigravity.md) | Google | 54.9 | 54.0–55.7 | 5–6 | 0.650 | 0.463 | 0.449–0.477 | 3 / 18 | | 7 | [Pi](https://feedbackbench.com/agents/pi.md) | Earendil Works (open source) | 52.7 | 51.9–53.3 | 7–7 | 0.456 | 0.608 | 0.591–0.623 | 14 / 0 | | =8 | [GitHub Copilot](https://feedbackbench.com/agents/copilot.md) | GitHub | 43.0 | 42.2–43.7 | 8–10 | 0.348 | 0.532 | 0.513–0.551 | 1 / 1 | | =8 | [Cline](https://feedbackbench.com/agents/cline.md) | Cline (open source) | 42.9 | 42.2–43.6 | 8–10 | 0.333 | 0.553 | 0.535–0.570 | 3 / 0 | | =8 | [Zed](https://feedbackbench.com/agents/zed.md) | Zed Industries | 42.4 | 41.7–43.0 | 8–10 | 0.362 | 0.496 | 0.481–0.511 | 2 / 2 | | =11 | [Factory](https://feedbackbench.com/agents/factory.md) | Factory | 35.6 | 34.9–36.2 | 11–12 | 0.232 | 0.545 | 0.525–0.563 | 2 / 0 | | =11 | [Amp](https://feedbackbench.com/agents/amp.md) | Amp | 35.3 | 34.7–35.8 | 11–12 | 0.222 | 0.561 | 0.544–0.577 | 5 / 0 | | 13 | [Kiro](https://feedbackbench.com/agents/kiro.md) | AWS | 29.5 | 29.0–30.0 | 13–13 | 0.185 | 0.470 | 0.453–0.485 | 0 / 2 | | 14 | [Conductor](https://feedbackbench.com/agents/conductor.md) | Melty Labs | 26.4 | 26.1–26.7 | 14–14 | 0.136 | 0.510 | 0.499–0.521 | 0 / 0 | | 15 | [Warp](https://feedbackbench.com/agents/warp.md) | Warp | 23.2 | 23.0–23.5 | 15–15 | 0.107 | 0.506 | 0.495–0.516 | 0 / 0 | | 16 | [Grok Build](https://feedbackbench.com/agents/grok-build.md) | xAI | 12.6 | 12.4–12.7 | 16–16 | 0.031 | 0.517 | 0.506–0.527 | 0 / 0 | | 17 | [Augment Code](https://feedbackbench.com/agents/augment.md) | Augment | 9.7 | 9.7–9.7 | 17–17 | 0.019 | 0.495 | 0.491–0.499 | 0 / 0 | ## Top quadrant Rule: popularity of at least 0.5 and customer love of at least 0.5. At the edge: the customer love 95% interval still includes 0.5. | Agent | Rank | Score | Popularity | Customer love | Customer love 95% interval | At the edge | |---|---|---|---|---|---|---| | Claude Code | 1 | 70.4 | 0.99 | 0.503 | 0.496–0.510 | yes | | OpenCode | 3 | 66.0 | 0.77 | 0.564 | 0.553–0.575 | no | | Cursor | 4 | 59.7 | 0.71 | 0.504 | 0.491–0.516 | yes | | Devin | =5 | 55.9 | 0.51 | 0.607 | 0.591–0.622 | no | ## Areas Reading per agent and area: Better than peers, Typical, Worse than peers, or Too few posts (under 30 rated author-weeks). | Agent | [Paying and limits](https://feedbackbench.com/criteria/paying.md) | [Setting up and connecting](https://feedbackbench.com/criteria/setup.md) | [Choosing models](https://feedbackbench.com/criteria/models.md) | [Instructing and context](https://feedbackbench.com/criteria/context.md) | [Doing the work](https://feedbackbench.com/criteria/work.md) | [Checking and finishing](https://feedbackbench.com/criteria/checking.md) | [Interface and sessions](https://feedbackbench.com/criteria/interface.md) | [Reliability and speed](https://feedbackbench.com/criteria/reliability.md) | [Account and support](https://feedbackbench.com/criteria/account.md) | |---|---|---|---|---|---|---|---|---|---| | Claude Code | Worse than peers | Better than peers | Better than peers | Typical | Typical | Typical | Better than peers | Better than peers | Worse than peers | | OpenAI Codex | Worse than peers | Worse than peers | Worse than peers | Typical | Typical | Typical | Worse than peers | Worse than peers | Better than peers | | OpenCode | Better than peers | Typical | Better than peers | Typical | Typical | Typical | Typical | Better than peers | Typical | | Cursor | Typical | Typical | Typical | Better than peers | Better than peers | Better than peers | Typical | Worse than peers | Worse than peers | | Devin | Better than peers | Typical | Better than peers | Typical | Better than peers | Typical | Better than peers | Typical | Too few posts | | Google Antigravity | Typical | Worse than peers | Typical | Worse than peers | Worse than peers | Worse than peers | Worse than peers | Better than peers | Typical | | Pi | Better than peers | Better than peers | Better than peers | Typical | Better than peers | Too few posts | Better than peers | Better than peers | Typical | | GitHub Copilot | Better than peers | Typical | Better than peers | Typical | Typical | Typical | Typical | Typical | Typical | | Cline | Better than peers | Typical | Better than peers | Typical | Typical | Too few posts | Better than peers | Better than peers | Typical | | Zed | Typical | Worse than peers | Too few posts | Too few posts | Worse than peers | Typical | Worse than peers | Better than peers | Typical | | Factory | Better than peers | Typical | Better than peers | Too few posts | Better than peers | Too few posts | Typical | Too few posts | Too few posts | | Amp | Better than peers | Typical | Better than peers | Typical | Better than peers | Too few posts | Better than peers | Typical | Better than peers | | Kiro | Worse than peers | Too few posts | Worse than peers | Too few posts | Typical | Too few posts | Too few posts | Too few posts | Typical | | Conductor | Too few posts | Too few posts | Too few posts | Too few posts | Typical | Too few posts | Typical | Too few posts | Too few posts | | Warp | Too few posts | Too few posts | Too few posts | Too few posts | Too few posts | Too few posts | Typical | Too few posts | Too few posts | | Grok Build | Too few posts | Too few posts | Too few posts | Too few posts | Too few posts | Too few posts | Too few posts | Too few posts | Too few posts | | Augment Code | Too few posts | Too few posts | Too few posts | Too few posts | Too few posts | Too few posts | Too few posts | Too few posts | Too few posts | ## Pages and data - Agents: [Claude Code](https://feedbackbench.com/agents/claude-code.md), [OpenAI Codex](https://feedbackbench.com/agents/codex.md), [OpenCode](https://feedbackbench.com/agents/opencode.md), [Cursor](https://feedbackbench.com/agents/cursor.md), [Devin](https://feedbackbench.com/agents/devin.md), [Google Antigravity](https://feedbackbench.com/agents/antigravity.md), [Pi](https://feedbackbench.com/agents/pi.md), [GitHub Copilot](https://feedbackbench.com/agents/copilot.md), [Cline](https://feedbackbench.com/agents/cline.md), [Zed](https://feedbackbench.com/agents/zed.md), [Factory](https://feedbackbench.com/agents/factory.md), [Amp](https://feedbackbench.com/agents/amp.md), [Kiro](https://feedbackbench.com/agents/kiro.md), [Conductor](https://feedbackbench.com/agents/conductor.md), [Warp](https://feedbackbench.com/agents/warp.md), [Grok Build](https://feedbackbench.com/agents/grok-build.md), [Augment Code](https://feedbackbench.com/agents/augment.md) - Criteria: one page per area and criterion under https://feedbackbench.com/criteria/ (list in [https://feedbackbench.com/llms.txt](https://feedbackbench.com/llms.txt)) - All data as JSON: [https://feedbackbench.com/data/coding.json](https://feedbackbench.com/data/coding.json) - Ranking as CSV: [https://feedbackbench.com/data/rankings.csv](https://feedbackbench.com/data/rankings.csv) - Every agent × criterion as CSV: [https://feedbackbench.com/data/criteria.csv](https://feedbackbench.com/data/criteria.csv) - Top requests per agent and criterion as CSV: [https://feedbackbench.com/data/requests.csv](https://feedbackbench.com/data/requests.csv) # Feedback Bench method (coding agents, built 2026-10-01) Each rule below is quoted verbatim from the build file. Names: popularity is `reach` and customer love is `regard` in the data files (coding.json, rankings.csv, criteria.csv). ## Sources 2026-08-31 to the Sunday before 2026-09-28. Four channels, collected the same way for every agent: (1) the agent's official subreddits, posts and comments; (2) Reddit posts and comments whose own text names the agent, from any subreddit we collect and from a Reddit search per agent; (3) posts on X that mention the agent's official handles (for OpenAI Codex, which has no product handle, an X search for "OpenAI Codex", "Codex CLI" and "Codex app"); (4) G2 and Trustpilot reviews of the product. Parent-brand review pages and staff handles are excluded. Each post counts once per agent. ## Attribution A post in an agent's own subreddit counts for that agent. Any other post counts for an agent only when the classifier confirms it is about that agent; a post can count for several agents. ## Labels Every post is labelled by Claude Sonnet 5 with the criteria codebook v1.0 (63 criteria in 9 areas): whether it is about the agent, whether it judges it, which criteria its stance names, and praise, complaint or mixed for each. On a pilot of 308 posts, Claude Opus 5.5 labelling independently agreed at kappa 0.75 on whether a post judges the agent, 0.73 overlap on its criteria, 0.78 on its areas, and 98% on polarity. Human annotators have not yet validated v1.0. 2 post(s) the classifier declined to label are left out. ## Unit An author-week: one author on one platform, about one agent, in one week. For each criterion it is positive if that author's praise outweighs their complaints that week, negative if the reverse. On Reddit, AutoModerator, [deleted] and names ending in 'bot' are not counted. Each review counts as its own author when the platform gives no name. ## Popularity Share of voice = the agent's distinct authors across all four channels / the sum over all agents. Popularity = log(1 + share / 0.005) / log(1 + leader's share / 0.005). Popularity counts every author in the window, so it has no sampling interval. ## Criterion love For each criterion: the category baseline is (positive + 1) / (positive + negative + 2) author-weeks across all agents. Channels differ in tone, so each agent is compared with its own channel mix: its expected share is the category baseline of each channel (Reddit, X, G2, Trustpilot), weighted by the agent's author-weeks on that channel. The agent's positive share is shrunk toward that expected share with 200 imaginary author-weeks and compared with it on the log-odds scale. 0.5 is the category norm. ## Customer love Customer love = the customer love scores of the 9 areas, averaged on the log-odds scale with weights equal to each area's share of the category's rated author-weeks. ## Feedback Score Feedback Score = 100 × sqrt(Popularity × Customer love). ## Intervals 95% intervals and rank ranges from 1000 bootstrap resamples of authors. ## Eligibility Every measured agent is ranked. Popularity keeps agents with few authors low; shrinkage keeps their customer love near the norm. ## Receipts The 5 most recent praise and complaint posts per area (first 700 characters) and 3 per criterion (first 450 characters). ## Caveat A subreddit is a proxy for the product. X volume is capped at 1,000 posts a day per feed; no feed in the window reached it on its median day. Labels are automatic. Criteria seen in few posts have wide intervals; those under 30 author-weeks for an agent are shown as too few posts. ## Head to head Posts that judge two agents. In each post, the agent whose stance (praise minus complaint across its criteria) is higher is ahead; ties are left out. Shares count the posts where one is ahead, with a 95% Wilson interval. Where: the agent's own subreddit, the other's, or anywhere else. Pairs with at least 30 such posts are shown. ## Requests A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first. ## Facts Agent facts come from web research dated 2026-09-24 and are not all verified against vendor pages. ## Display rules - Reading of a criterion or area: Better than peers when its customer love 95% interval lies wholly above 0.5; Worse than peers when wholly below; Typical when it spans 0.5; Too few posts under 30 rated author-weeks. - Top quadrant: popularity of at least 0.5 and customer love of at least 0.5; at the edge when the customer love 95% interval still includes 0.5. - Rank shows “=N” when rank ranges overlap, directly or through a chain. ## Parameters | Parameter | Value | |---|---| | S0 | 0.005 | | priorStrength | 200 | | bootstrap | 1000 | | seed | 20260927 | ## Area weights in Customer love Each area's share of the category's rated author-weeks. | Area | Weight | |---|---| | Paying and limits | 0.339 | | Setting up and connecting | 0.053 | | Choosing models | 0.112 | | Instructing and context | 0.059 | | Doing the work | 0.230 | | Checking and finishing | 0.021 | | Interface and sessions | 0.058 | | Reliability and speed | 0.091 | | Account and support | 0.037 | ## Category baselines (positive share of author-weeks, %) | Code | Name | Baseline % | |---|---|---| | paying | Paying and limits | 23.9 | | setup | Setting up and connecting | 40.1 | | models | Choosing models | 25.9 | | context | Instructing and context | 36.2 | | work | Doing the work | 47.5 | | checking | Checking and finishing | 44.0 | | interface | Interface and sessions | 41.4 | | reliability | Reliability and speed | 19.0 | | account | Account and support | 13.1 | | limits.plan_value | How much use a plan's price buys | 37.9 | | limits.window_interrupts_work | Short rolling usage window blocks or interrupts work | 13.8 | | limits.burn_rate | Single prompt, model or effort level consumes disproportionate quota | 16.2 | | limits.allowance_change | Price, allowance or plan terms changed | 5.8 | | limits.reset_schedule | Quota reset timing and bonus or banked resets | 17.0 | | limits.usage_meter | Usage meter visibility and accuracy | 7.8 | | limits.prompt_cache | Prompt cache hits, misses and invalidation | 36.1 | | billing.overage_charges | Pay-as-you-go overage, fallback billing and spend caps | 8.0 | | billing.pricing_clarity | Pricing and plan terms stated clearly and consistently | 6.3 | | billing.free_tier | Free tier and free model availability and limits | 57.4 | | billing.subscription_portability | Using an existing subscription across tools | 38.5 | | setup.install_signin | Install, launch and sign-in | 20.0 | | setup.provider_byok_local | Connecting own API keys, local models and custom endpoints | 54.2 | | setup.extensions_mcp | MCP servers, plugins, skills and hooks | 50.4 | | setup.onboarding_docs | Onboarding, discoverability and documentation | 19.7 | | setup.ide_integration | IDE and editor integration | 45.6 | | models.catalog_access | Which models are offered on a plan and when | 28.1 | | models.routing_auto | Automatic model routing and fallback | 24.6 | | models.effort_control | Reasoning effort setting and its defaults | 45.7 | | models.quality_drift | Quality got worse or better over time | 21.8 | | context.instruction_files | Persistent project rules files are read and obeyed | 52.8 | | context.instruction_following | Direct in-prompt instructions and caps are followed | 26.9 | | context.clarifying_questions | Asks the user versus guessing | 38.6 | | context.long_context_decay | Output degrades as the context window fills | 15.5 | | context.compaction | Context compaction keeps what matters, cheaply and quickly | 33.3 | | context.session_memory | Memory and state carried across sessions | 46.7 | | context.codebase_retrieval | Finding the right files in the codebase | 44.3 | | context.attachments | Images, PDFs and file attachments as input | 38.8 | | work.capability | Can do the user's kind of task | 60.7 | | work.frontend_ui | Frontend and visual UI output | 44.4 | | work.bug_diagnosis | Diagnosing and fixing reported bugs | 67.4 | | work.regressions_introduced | Breaks existing code or reintroduces bugs | 9.2 | | work.scope_overreach | Does unrequested work or over-engineers | 6.0 | | work.stuck_loops | Spins, loops or gets stuck without progress | 5.1 | | work.premature_stop | Stops mid-task or answers instead of acting | 14.9 | | work.long_running_autonomy | Long unattended runs and goal/loop mode | 74.2 | | work.multi_agent_orchestration | Subagents, parallel agents and orchestrators | 60.6 | | work.reward_hacking | Games checks instead of fixing the problem | 2.5 | | work.destructive_actions | Risky or irreversible actions without confirmation | 18.5 | | work.git_workflow | Git commits, branches and sync | 35.9 | | work.computer_browser_use | Computer use and browser control | 54.8 | | work.safety_refusals | Safety filters block legitimate coding tasks | 11.9 | | work.permission_prompts | Tool approval prompts and autonomy modes | 24.9 | | work.plan_mode | Plan-before-edit mode | 46.2 | | work.response_verbosity | Length and clarity of replies, summaries and comments | 21.4 | | work.sycophancy_pushback | Caves to or argues with the user's judgement | 16.7 | | verify.false_completion | Claims work is done or fixed when it is not | 5.6 | | verify.self_testing | Builds, tests or runs its own changes | 54.5 | | verify.agent_code_review | Agent-performed code review finds real issues | 73.3 | | verify.change_review_ui | Reviewing and approving the agent's changes | 37.8 | | ui.display_settings | How the interface shows work, and what the user can configure | 34.6 | | ui.session_history | Saving, switching, resuming and rewinding sessions | 29.2 | | ui.interrupt_steer | Stopping and steering a running agent | 38.9 | | surfaces.remote_mobile | Mobile, remote-control and voice access | 52.5 | | surfaces.cloud_sessions | Cloud and remote sandbox execution | 65.5 | | rel.service_errors | Outages, server errors and capacity or rate errors | 7.5 | | rel.response_speed | Latency, throughput and fast mode | 36.9 | | rel.client_failures | Client crashes, freezes and failed tool execution | 7.1 | | rel.update_breakage | Updates break working setups | 15.3 | | account.support | Support, refunds and issue handling | 16.3 | | account.billing_errors | Wrong charges, failed payments and plan provisioning | 2.4 | | account.bans_restrictions | Account bans and access restrictions | 6.4 | | account.data_privacy | Data retention, training use and deployment isolation | 20.1 | ## Criteria 63 criteria in 9 areas. Codebook: every criterion names one product mechanism, visible in a single post. ### Paying and limits (`paying`) What do you pay, and how far does it get you? - [`limits.plan_value`](https://feedbackbench.com/criteria/limits.plan_value.md) **How much use a plan's price buys.** Whether a plan's weekly or monthly allowance covers the user's normal work, and how its price compares with other plans or agents for the use it gives. Covers running out days early, never hitting the cap, missing tiers and value per dollar. Boundary: Not this: see limits.window_interrupts_work when the short multi-hour window blocks work. Not this: see limits.burn_rate when one prompt, model or effort level drained the quota. Not this: see limits.allowance_change when price or allowance changed over time. Not this: see billing.pricing_clarity when the information itself is unclear. - [`limits.window_interrupts_work`](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) **Short rolling usage window blocks or interrupts work.** The short rolling usage window (a cap of a few hours) blocks the user mid-task or mid-session: forced waits, killed sessions, or a long allowance left unused because the short window ran out. The post is about the window itself, not about what consumed it. Boundary: Not this: see limits.burn_rate when the point is that one prompt, model or effort level used a large share. Not this: see limits.plan_value for the weekly or monthly total. - [`limits.burn_rate`](https://feedbackbench.com/criteria/limits.burn_rate.md) **Single prompt, model or effort level consumes disproportionate quota.** How much quota a specific prompt, task, model version, effort level or feature consumes for the work done. Includes new versions using more than earlier ones and praise for efficient ones. The post names what consumed the quota. Boundary: Not this: see limits.window_interrupts_work when the post is about being blocked by the short window, not about the consumer. Not this: see limits.prompt_cache for cache misses. Not this: see work.multi_agent_orchestration when runaway subagents cause the burn. - [`limits.allowance_change`](https://feedbackbench.com/criteria/limits.allowance_change.md) **Price, allowance or plan terms changed.** The vendor changes a plan's price, quota, multipliers, included models or tiers over time: cuts, raises, removals, silent changes and changes that differ from what was announced. Boundary: Not this: see limits.plan_value for the current level with no change named. Not this: see limits.reset_schedule for when resets happen. - [`limits.reset_schedule`](https://feedbackbench.com/criteria/limits.reset_schedule.md) **Quota reset timing and bonus or banked resets.** When and how usage resets happen: scheduled reset dates, surprise or bonus resets, banked resets, and paid resets. Covers resets that wipe saved allowance or land uselessly close to the normal reset. Boundary: Not this: see limits.allowance_change for changes to quota size. Not this: see limits.usage_meter when the reset is only displayed wrongly. - [`limits.usage_meter`](https://feedbackbench.com/criteria/limits.usage_meter.md) **Usage meter visibility and accuracy.** Whether the product shows how much quota and tokens were used and how much remains, per task and per model, and whether that figure is accurate. Covers meters that climb while idle and hidden token counts. Boundary: Not this: see billing.overage_charges for money actually charged. Not this: see limits.burn_rate when the meter is believed and the complaint is the amount consumed. - [`limits.prompt_cache`](https://feedbackbench.com/criteria/limits.prompt_cache.md) **Prompt cache hits, misses and invalidation.** Whether prompt caching is kept across turns, resumes and provider routing, and how cache reads are priced. Covers cache drops that inflate usage. Boundary: Not this: see limits.burn_rate for consumption not attributed to caching. - [`billing.overage_charges`](https://feedbackbench.com/criteria/billing.overage_charges.md) **Pay-as-you-go overage, fallback billing and spend caps.** Usage spills from the plan into metered, on-demand or API billing, and whether spend caps and alerts work. Covers overage turning on silently and credit top-ups. Boundary: Not this: see account.billing_errors for wrong charges on the subscription itself. Not this: see models.routing_auto when the cause is auto-selection of a pricier model. - [`billing.pricing_clarity`](https://feedbackbench.com/criteria/billing.pricing_clarity.md) **Pricing and plan terms stated clearly and consistently.** Whether pricing pages, docs and plan descriptions clearly and consistently state costs, limits and what 'unlimited' means. Boundary: Not this: see limits.allowance_change when terms actually changed. Not this: see limits.usage_meter for in-product usage display. - [`billing.free_tier`](https://feedbackbench.com/criteria/billing.free_tier.md) **Free tier and free model availability and limits.** Whether free models, free tiers and trial or promo access exist, how long they stay available, and whether their limits are usable. Boundary: Not this: see models.catalog_access for paid model availability. Not this: see rel.service_errors when free models return errors. - [`billing.subscription_portability`](https://feedbackbench.com/criteria/billing.subscription_portability.md) **Using an existing subscription across tools.** Whether a paid subscription from one vendor can be used inside another agent or harness, or used as an API outside the vendor's own app. Boundary: Not this: see setup.provider_byok_local for API keys, local models and custom endpoints. ### Setting up and connecting (`setup`) How hard is it to install and connect? - [`setup.install_signin`](https://feedbackbench.com/criteria/setup.install_signin.md) **Install, launch and sign-in.** Getting the agent installed, started and signed in on the user's platform: installers, updates on first run, login flows, browser redirects and switching accounts. Boundary: Not this: see setup.provider_byok_local for connecting own keys or local models. Not this: see rel.update_breakage when a later update breaks a working setup. - [`setup.provider_byok_local`](https://feedbackbench.com/criteria/setup.provider_byok_local.md) **Connecting own API keys, local models and custom endpoints.** Whether user-supplied providers work: bring-your-own-key, OpenAI-compatible endpoints, OpenRouter, and local servers such as Ollama or LM Studio. Boundary: Not this: see billing.subscription_portability for reusing a paid subscription. Not this: see rel.tool_call_errors for tool-format failures once connected. - [`setup.extensions_mcp`](https://feedbackbench.com/criteria/setup.extensions_mcp.md) **MCP servers, plugins, skills and hooks.** Whether external tools and extension points load, authenticate and refresh correctly: MCP servers, plugins, skills, hooks and third-party integrations. Boundary: Not this: see setup.ide_integration for editor extensions. Not this: see context.instruction_files for rules files. - [`setup.onboarding_docs`](https://feedbackbench.com/criteria/setup.onboarding_docs.md) **Onboarding, discoverability and documentation.** How easily a new user learns the product: quick-start guides, docs explaining features and modes, and how findable features are. Boundary: Not this: see billing.pricing_clarity for pricing docs. Not this: see ui.customization for settings complexity. - [`setup.ide_integration`](https://feedbackbench.com/criteria/setup.ide_integration.md) **IDE and editor integration.** How the agent works inside an editor or IDE, including extension support, editor-native features, autocomplete and feature parity across IDEs. Boundary: Not this: see verify.change_review_ui for diff and approval views. Not this: see ui.terminal_display for CLI rendering. ### Choosing models (`models`) Which models do you get, and do they hold up? - [`models.catalog_access`](https://feedbackbench.com/criteria/models.catalog_access.md) **Which models are offered on a plan and when.** Whether models are available on the user's plan: same-day support for new releases, deprecations and removals, multi-provider choice, and models missing from a surface. Boundary: Not this: see billing.free_tier for free-model availability. Not this: see models.routing_auto for which model actually runs. - [`models.routing_auto`](https://feedbackbench.com/criteria/models.routing_auto.md) **Automatic model routing and fallback.** Auto mode, routers or fallbacks pick, switch or hide the model used. This includes subagents running on a model other than the one requested, silent downgrades, and inability to exclude models. Boundary: Not this: see models.effort_control for reasoning level. Not this: see billing.overage_charges for the resulting charge alone. - [`models.effort_control`](https://feedbackbench.com/criteria/models.effort_control.md) **Reasoning effort setting and its defaults.** How the reasoning or effort level can be set, whether it stays set, and how outcomes differ by level, such as overthinking at high effort or failing at low effort. Boundary: Not this: see limits.burn_rate when the point is the quota share an effort level consumed. Not this: see work.scope_overreach for over-engineered output. - [`models.quality_drift`](https://feedbackbench.com/criteria/models.quality_drift.md) **Quality got worse or better over time.** The post compares the agent or one of its models with an earlier time and says it got worse or better: 'nerfed', 'dumber since last week', 'better than at launch'. An explicit comparison with the past is required. Boundary: Not this: see general.unspecific for a verdict with no comparison over time. Not this: see work.capability for what it can do now. Not this: see limits.allowance_change for changed quotas or prices. ### Instructing and context (`context`) Does it follow your instructions and keep the right context? - [`context.instruction_files`](https://feedbackbench.com/criteria/context.instruction_files.md) **Persistent project rules files are read and obeyed.** Whether the agent reads and follows persistent project instruction files (rules or agent markdown files) across turns. Boundary: Not this: see context.instruction_following for instructions in the current prompt. Not this: see context.session_memory for auto-written memory. - [`context.instruction_following`](https://feedbackbench.com/criteria/context.instruction_following.md) **Direct in-prompt instructions and caps are followed.** Whether the agent executes explicit instructions, prohibitions and budgets given in the prompt, such as iteration caps and token caps. Boundary: Not this: see context.instruction_files for rules files. Not this: see work.sycophancy_pushback for evaluating the user's claims or opinions. Not this: see work.scope_overreach for unrequested extras. - [`context.clarifying_questions`](https://feedbackbench.com/criteria/context.clarifying_questions.md) **Asks the user versus guessing.** Whether the agent stops to ask clarifying questions when it is blocked or the request is ambiguous, instead of inventing assumptions or workarounds. Boundary: Not this: see work.stuck_loops for repetitive failure without any question. - [`context.long_context_decay`](https://feedbackbench.com/criteria/context.long_context_decay.md) **Output degrades as the context window fills.** The agent forgets details, rules or its own statements as the session grows long. Boundary: Not this: see context.compaction for losses caused by summarisation. Not this: see context.session_memory for loss across separate sessions. - [`context.compaction`](https://feedbackbench.com/criteria/context.compaction.md) **Context compaction keeps what matters, cheaply and quickly.** How automatic or manual compaction summarises history: what it drops, how long it takes, what it costs, and whether it is visible. Boundary: Not this: see context.long_context_decay for degradation without compaction. - [`context.session_memory`](https://feedbackbench.com/criteria/context.session_memory.md) **Memory and state carried across sessions.** Whether knowledge persists correctly between sessions: memory files, re-discovery cost at session start, stale memories, and leakage between sessions. Boundary: Not this: see ui.session_history for saving and resuming chat transcripts. Not this: see context.instruction_files for user-written rules. - [`context.codebase_retrieval`](https://feedbackbench.com/criteria/context.codebase_retrieval.md) **Finding the right files in the codebase.** How the agent searches and indexes a repository and pulls in relevant files. Covers over-reading on trivial edits and missing files in large or multi-repo codebases. Boundary: Not this: see limits.burn_rate for consumption not tied to file reading. - [`context.attachments`](https://feedbackbench.com/criteria/context.attachments.md) **Images, PDFs and file attachments as input.** Whether attached images, screenshots, PDFs and other files are accepted and read. Boundary: Not this: see work.computer_browser_use for screen control. ### Doing the work (`work`) How does it behave while it works? - [`work.capability`](https://feedbackbench.com/criteria/work.capability.md) **Can do the user's kind of task.** Whether the agent, with its models, succeeds at the user's kind of task: its size, complexity, language or domain, including building a whole feature in one go or failing at it. The post names the task or its size but no narrower failure behaviour. Boundary: Not this: see work.frontend_ui for visual UI work and work.bug_diagnosis for fixing a reported bug. Not this: see models.quality_drift for a change over time. Not this: see the other work.* and verify.* leaves when a specific behaviour (looping, scope, false claims, breaking code) is named. Not this: see general.unspecific when no task is named. - [`work.frontend_ui`](https://feedbackbench.com/criteria/work.frontend_ui.md) **Frontend and visual UI output.** How the agent handles UI design, layout, visual taste and front-end component wiring. Boundary: Not this: see work.regressions_introduced for UI bugs reintroduced after fixes. - [`work.bug_diagnosis`](https://feedbackbench.com/criteria/work.bug_diagnosis.md) **Diagnosing and fixing reported bugs.** Whether the agent finds the root cause of a failing behaviour and fixes it without step-by-step guidance. Boundary: Not this: see verify.agent_code_review for reviewing code to find unknown issues. - [`work.regressions_introduced`](https://feedbackbench.com/criteria/work.regressions_introduced.md) **Breaks existing code or reintroduces bugs.** Edits break previously working code, reintroduce fixed bugs, or cycle between introducing and fixing bugs. Boundary: Not this: see work.stuck_loops for repetition without code damage. Not this: see work.reward_hacking for deliberately gaming checks. - [`work.scope_overreach`](https://feedbackbench.com/criteria/work.scope_overreach.md) **Does unrequested work or over-engineers.** The agent edits outside the requested scope, adds unasked features, tests or runs, or builds overly complex solutions. Boundary: Not this: see work.destructive_actions for risky irreversible actions. Not this: see work.response_verbosity for padded text or comments. - [`work.stuck_loops`](https://feedbackbench.com/criteria/work.stuck_loops.md) **Spins, loops or gets stuck without progress.** The agent repeats attempts, loops endlessly or spirals on a problem without converging. Boundary: Not this: see work.premature_stop for stopping early. Not this: see context.clarifying_questions when the fix was to ask the user. - [`work.premature_stop`](https://feedbackbench.com/criteria/work.premature_stop.md) **Stops mid-task or answers instead of acting.** The agent halts before finishing, or replies with explanation or a summary instead of performing the action. Boundary: Not this: see verify.false_completion for stopping while claiming the work is done. Not this: see limits.window_interrupts_work for stops caused by quota. - [`work.long_running_autonomy`](https://feedbackbench.com/criteria/work.long_running_autonomy.md) **Long unattended runs and goal/loop mode.** Whether the agent sustains hours-long autonomous work toward a goal, and whether a goal or loop mode exists. Boundary: Not this: see work.multi_agent_orchestration for coordination between agents. Not this: see surfaces.cloud_sessions for where the run executes. - [`work.multi_agent_orchestration`](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) **Subagents, parallel agents and orchestrators.** How the agent spawns, delegates to, monitors and coordinates subagents or parallel sessions. Covers polling loops and the wrong choice of subagents. Boundary: Not this: see models.routing_auto for the model a subagent runs on. Not this: see work.git_workflow for merge conflicts between sessions. - [`work.reward_hacking`](https://feedbackbench.com/criteria/work.reward_hacking.md) **Games checks instead of fixing the problem.** The agent disables linters, hardcodes for tests, edits benchmarks or defends bugs with tests so that checks pass without a real fix. Boundary: Not this: see verify.false_completion for plain false claims without manipulated checks. - [`work.destructive_actions`](https://feedbackbench.com/criteria/work.destructive_actions.md) **Risky or irreversible actions without confirmation.** The agent pushes, merges, deletes or wipes changes, or acts on the host outside its sandbox, without asking. Also covers praise for staying in bounds. Boundary: Not this: see work.permission_prompts for approval prompt frequency. Not this: see work.scope_overreach for harmless extra work. - [`work.git_workflow`](https://feedbackbench.com/criteria/work.git_workflow.md) **Git commits, branches and sync.** How the agent handles commits, attribution, branches, pulls and pushes, and merge conflicts between concurrent agents or teammates. Boundary: Not this: see work.destructive_actions for data-destroying git commands. - [`work.computer_browser_use`](https://feedbackbench.com/criteria/work.computer_browser_use.md) **Computer use and browser control.** Whether the agent operates desktop apps and browsers well, without taking over the user's windows. Boundary: Not this: see context.attachments for reading screenshots. - [`work.safety_refusals`](https://feedbackbench.com/criteria/work.safety_refusals.md) **Safety filters block legitimate coding tasks.** Refusals or safety classifiers stop benign or security-remediation work. Boundary: Not this: see work.permission_prompts for tool approval prompts. Not this: see models.routing_auto for safety-triggered model downgrades. - [`work.permission_prompts`](https://feedbackbench.com/criteria/work.permission_prompts.md) **Tool approval prompts and autonomy modes.** How often and when the agent asks permission to run tools or commands, and whether auto or bypass modes behave as configured. Boundary: Not this: see work.destructive_actions for acting without permission. Not this: see work.safety_refusals for content blocks. - [`work.plan_mode`](https://feedbackbench.com/criteria/work.plan_mode.md) **Plan-before-edit mode.** Whether a read-only planning phase exists, produces useful plans, and actually prevents edits. Boundary: Not this: see work.task_completion for how well an approved plan is executed. - [`work.response_verbosity`](https://feedbackbench.com/criteria/work.response_verbosity.md) **Length and clarity of replies, summaries and comments.** How long and how readable the agent's prose and summaries are, and how many code comments and doc notes it writes. Boundary: Not this: see verify.change_review_ui for the volume of code changes. Not this: see work.scope_overreach for extra code features. - [`work.sycophancy_pushback`](https://feedbackbench.com/criteria/work.sycophancy_pushback.md) **Caves to or argues with the user's judgement.** How the agent handles disagreement: agreeing with everything, reversing correct findings when challenged, failing to push back on wrong claims, or arguing when corrected. Boundary: Not this: see context.instruction_following for refusing a clear instruction. ### Checking and finishing (`checking`) Can you trust that the work is done? - [`verify.false_completion`](https://feedbackbench.com/criteria/verify.false_completion.md) **Claims work is done or fixed when it is not.** The agent reports success, completion or a finished todo list that turns out to be untrue. Boundary: Not this: see work.reward_hacking when checks were manipulated. Not this: see verify.self_testing for whether checks were run at all. - [`verify.self_testing`](https://feedbackbench.com/criteria/verify.self_testing.md) **Builds, tests or runs its own changes.** The post says whether the agent built, tested, ran or checked its own change before handing it back: it ran the test suite, skipped tests, did not run the app, or verified in proportion to the change. Boundary: Not this: see verify.false_completion when the agent claims success that did not happen. Not this: see verify.agent_code_review for a separate review mode or review agent. - [`verify.agent_code_review`](https://feedbackbench.com/criteria/verify.agent_code_review.md) **Agent-performed code review finds real issues.** A review mode or review agent finds real defects in code or PRs, without re-flagging code that already passed or producing noise. Boundary: Not this: see work.bug_diagnosis for fixing a known symptom. Not this: see work.sycophancy_pushback for reversing findings under pressure. - [`verify.change_review_ui`](https://feedbackbench.com/criteria/verify.change_review_ui.md) **Reviewing and approving the agent's changes.** Diff views, per-file approval, edit-acceptance prompts, and how manageable the amount of change is for a human reviewer. Boundary: Not this: see setup.ide_integration for general editor features. Not this: see work.response_verbosity for prose length. ### Interface and sessions (`interface`) How do you operate and steer it? - [`ui.display_settings`](https://feedbackbench.com/criteria/ui.display_settings.md) **How the interface shows work, and what the user can configure.** How the CLI, TUI, IDE panel or app shows the agent's activity (tool calls, diffs, logs, subagent progress, scrolling, fonts, layout) and which settings, keybindings, themes and presets the user can change. Boundary: Not this: see limits.usage_meter for usage displays. Not this: see ui.session_history for saving and resuming sessions. Not this: see models.effort_control for reasoning-effort settings. - [`ui.session_history`](https://feedbackbench.com/criteria/ui.session_history.md) **Saving, switching, resuming and rewinding sessions.** Whether chats and threads persist, can be switched and resumed with their state intact, and can be rewound or reverted. Boundary: Not this: see context.session_memory for knowledge carried into a new session. - [`ui.interrupt_steer`](https://feedbackbench.com/criteria/ui.interrupt_steer.md) **Stopping and steering a running agent.** Whether the user can stop, interrupt or send guidance mid-run and have the agent respect it and continue. Boundary: Not this: see work.premature_stop for the agent stopping on its own. - [`surfaces.remote_mobile`](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) **Mobile, remote-control and voice access.** Driving or monitoring agents from a phone, another device or by voice, including approvals from the mobile client. Boundary: Not this: see surfaces.cloud_sessions for where the agent executes. - [`surfaces.cloud_sessions`](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) **Cloud and remote sandbox execution.** Agents run in hosted or remote environments. Covers persistence when the laptop closes, access to local files, session caps, and start or resume time. Boundary: Not this: see billing.overage_charges for how cloud runs are billed. Not this: see work.long_running_autonomy for run behaviour. ### Reliability and speed (`reliability`) Does it stay up and respond quickly? - [`rel.service_errors`](https://feedbackbench.com/criteria/rel.service_errors.md) **Outages, server errors and capacity or rate errors.** Backend unavailability, 5xx or connection errors, capacity errors, and transient 429 errors that do not reflect the user's quota. Boundary: Not this: see limits.window_interrupts_work and limits.allowance_size for plan quota being reached. Not this: see rel.tool_call_errors for harness tool failures. - [`rel.response_speed`](https://feedbackbench.com/criteria/rel.response_speed.md) **Latency, throughput and fast mode.** How quickly the agent responds and completes work, including paid or fast modes and whether they actually deliver speed. Boundary: Not this: see rel.service_errors for failures. Not this: see rel.client_crash_resources for local slowness from resource use. - [`rel.client_failures`](https://feedbackbench.com/criteria/rel.client_failures.md) **Client crashes, freezes and failed tool execution.** The local app, CLI or extension crashes, freezes, leaks memory or uses too much CPU, or fails to execute the model's tool calls and shell commands. Boundary: Not this: see rel.service_errors for backend outages and server errors. Not this: see rel.update_breakage when a specific update caused it. - [`rel.update_breakage`](https://feedbackbench.com/criteria/rel.update_breakage.md) **Updates break working setups.** New releases regress features, break compatibility or change behaviour unexpectedly. Also covers praise for smooth migrations. Boundary: Not this: see limits.allowance_change for quota changes delivered in an update. ### Account and support (`account`) How does the vendor treat your account? - [`account.support`](https://feedbackbench.com/criteria/account.support.md) **Support, refunds and issue handling.** Reaching a human, getting refunds and resolutions, and the vendor's responsiveness to bug reports and community issues. Boundary: Not this: see account.billing_errors for the underlying wrong charge. - [`account.billing_errors`](https://feedbackbench.com/criteria/account.billing_errors.md) **Wrong charges, failed payments and plan provisioning.** Payments are charged incorrectly, plans are not provisioned after payment, proration errors occur, or payments fail. Boundary: Not this: see billing.overage_charges for metered overage. Not this: see account.support for how the vendor handled the error. - [`account.bans_restrictions`](https://feedbackbench.com/criteria/account.bans_restrictions.md) **Account bans and access restrictions.** The vendor suspends, bans or restricts accounts, for example for heavy usage or regional embargoes. Boundary: Not this: see limits.allowance_size for normal quota lockouts. - [`account.data_privacy`](https://feedbackbench.com/criteria/account.data_privacy.md) **Data retention, training use and deployment isolation.** Whether prompts and code are retained or used for training, and whether private or air-gapped deployment is available. Boundary: Not this: see billing.free_tier for free-model availability itself. # Claude Code (Anthropic) Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/agent/claude-code | Measure | Value | |---|---| | Rank | 1 of 17 (rank range 1–1) | | Feedback Score | 70.4 (95% interval 69.9–70.9) | | Popularity | 0.986 (share of voice 28.56%) | | Customer love | 0.503 (95% interval 0.496–0.510) | | Top quadrant | yes, at the edge | | Authors | 28126 | | Posts counted | 80128 | | Posts that judge the agent | 31596 | | Criteria better / worse than peers | 14 / 10 of 63 | ## Top requests What users ask to add or change, most asked first. 3579 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first. | Rank | Request | Criterion | Author-weeks | Posts | |---|---|---|---|---| | 1 | One-off usage limit reset now | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | 142 | 156 | | 2 | Higher overall usage limits | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | 72 | 73 | | 3 | Additional or recurring bonus usage resets | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | 62 | 69 | | 4 | Stop nerfing or degrading models over time | [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md) | 59 | 62 | | 5 | Remove the 5-hour usage window | [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | 55 | 56 | | 6 | Compensation reset after outages or bugs | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | 40 | 41 | | 7 | Shorter, less verbose responses | [Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md) | 38 | 38 | | 8 | Bankable usage resets | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | 37 | 38 | | 9 | Fewer false-positive safety blocks on benign tasks | [Safety filters block legitimate coding tasks](https://feedbackbench.com/criteria/work.safety_refusals.md) | 36 | 42 | | 10 | Higher allowance on top-tier plans | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | 33 | 33 | | 11 | Published exact usage limits per plan | [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md) | 29 | 29 | | 12 | Let in-progress task finish at cutoff | [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | 27 | 28 | ## Facts | Fact | Value | |---|---| | Version | Claude Opus 5.5 (flagship model), Claude Sonnet 5 | | Released | Opus 5.5: 2026-09-22 | | Price | Pro ~$17-20/mo, Max from $100/mo, Team $20-100/seat/mo, Enterprise from $20/seat + usage-based API overage | | Model | Claude Opus 5.5 ($4/$20 per 1M tokens API; Fast mode $8/$40) | | Surface | CLI, IDE extensions | ## Sources | Channel | Source | Posts | |---|---|---| | Reddit | r/ClaudeCode | 66057 | | X | @ClaudeDevs | 11217 | | Reddit | Posts that name it | 2747 | | X | @claude_code | 103 | | G2 | G2 | 4 | ## Better than peers on Short rolling usage window blocks or interrupts work, Quota reset timing and bonus or banked resets, Usage meter visibility and accuracy, MCP servers, plugins, skills and hooks, Quality got worse or better over time, Can do the user's kind of task, Frontend and visual UI output, Stops mid-task or answers instead of acting, Long unattended runs and goal/loop mode, Plan-before-edit mode, How the interface shows work, and what the user can configure, Mobile, remote-control and voice access, Latency, throughput and fast mode, Client crashes, freezes and failed tool execution ## Worse than peers on How much use a plan's price buys, Free tier and free model availability and limits, Using an existing subscription across tools, Which models are offered on a plan and when, Does unrequested work or over-engineers, Games checks instead of fixing the problem, Safety filters block legitimate coding tasks, Length and clarity of replies, summaries and comments, Agent-performed code review finds real issues, Support, refunds and issue handling ## All 63 criteria Criterion love: 0.5 is the category norm. n: rated author-weeks. ### Paying and limits: Worse than peers (customer love 0.471, n 7898) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | Worse than peers | 0.439 | 0.423–0.457 | 2873 | 916 | 1957 | | [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md) | Typical | 0.517 | 0.497–0.537 | 2700 | 467 | 2233 | | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | Better than peers | 0.583 | 0.558–0.607 | 1141 | 283 | 858 | | [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | Better than peers | 0.543 | 0.518–0.567 | 1002 | 169 | 833 | | [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | Typical | 0.527 | 0.482–0.567 | 809 | 59 | 750 | | [Usage meter visibility and accuracy](https://feedbackbench.com/criteria/limits.usage_meter.md) | Better than peers | 0.593 | 0.552–0.627 | 434 | 60 | 374 | | [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md) | Typical | 0.547 | 0.496–0.590 | 429 | 36 | 393 | | [Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md) | Typical | 0.513 | 0.489–0.538 | 350 | 135 | 215 | | [Pay-as-you-go overage, fallback billing and spend caps](https://feedbackbench.com/criteria/billing.overage_charges.md) | Typical | 0.500 | 0.446–0.537 | 211 | 17 | 194 | | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | Worse than peers | 0.420 | 0.392–0.446 | 123 | 29 | 94 | | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | Worse than peers | 0.448 | 0.417–0.478 | 95 | 40 | 55 | ### Setting up and connecting: Better than peers (customer love 0.565, n 863) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md) | Better than peers | 0.563 | 0.537–0.590 | 454 | 270 | 184 | | [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | Typical | 0.513 | 0.476–0.549 | 133 | 30 | 103 | | [IDE and editor integration](https://feedbackbench.com/criteria/setup.ide_integration.md) | Typical | 0.515 | 0.486–0.546 | 108 | 55 | 53 | | [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | Typical | 0.507 | 0.480–0.540 | 92 | 53 | 39 | | [Onboarding, discoverability and documentation](https://feedbackbench.com/criteria/setup.onboarding_docs.md) | Typical | 0.505 | 0.471–0.539 | 92 | 19 | 73 | ### Choosing models: Better than peers (customer love 0.538, n 2231) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md) | Better than peers | 0.563 | 0.544–0.583 | 1616 | 442 | 1174 | | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | Typical | 0.526 | 0.491–0.559 | 294 | 79 | 215 | | [Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md) | Typical | 0.521 | 0.489–0.548 | 267 | 132 | 135 | | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | Worse than peers | 0.443 | 0.406–0.478 | 177 | 32 | 145 | ### Instructing and context: Typical (customer love 0.506, n 1942) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md) | Typical | 0.491 | 0.463–0.519 | 516 | 133 | 383 | | [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md) | Typical | 0.501 | 0.479–0.526 | 439 | 233 | 206 | | [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md) | Typical | 0.510 | 0.485–0.533 | 390 | 183 | 207 | | [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | Typical | 0.486 | 0.457–0.512 | 379 | 117 | 262 | | [Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md) | Typical | 0.485 | 0.446–0.519 | 319 | 46 | 273 | | [Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md) | Typical | 0.520 | 0.491–0.549 | 149 | 72 | 77 | | [Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md) | Typical | 0.502 | 0.477–0.530 | 78 | 31 | 47 | | [Images, PDFs and file attachments as input](https://feedbackbench.com/criteria/context.attachments.md) | Typical | 0.514 | 0.491–0.536 | 36 | 17 | 19 | ### Doing the work: Typical (customer love 0.503, n 5831) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md) | Better than peers | 0.558 | 0.543–0.574 | 2798 | 1831 | 967 | | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | Typical | 0.510 | 0.487–0.533 | 961 | 591 | 370 | | [Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md) | Worse than peers | 0.433 | 0.409–0.455 | 673 | 106 | 567 | | [Safety filters block legitimate coding tasks](https://feedbackbench.com/criteria/work.safety_refusals.md) | Worse than peers | 0.443 | 0.397–0.480 | 376 | 30 | 346 | | [Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md) | Worse than peers | 0.437 | 0.366–0.499 | 311 | 12 | 299 | | [Risky or irreversible actions without confirmation](https://feedbackbench.com/criteria/work.destructive_actions.md) | Typical | 0.514 | 0.484–0.541 | 307 | 63 | 244 | | [Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md) | Better than peers | 0.540 | 0.506–0.576 | 272 | 214 | 58 | | [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | Typical | 0.523 | 0.490–0.552 | 257 | 73 | 184 | | [Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md) | Better than peers | 0.534 | 0.505–0.559 | 233 | 115 | 118 | | [Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md) | Typical | 0.523 | 0.442–0.583 | 199 | 12 | 187 | | [Caves to or argues with the user's judgement](https://feedbackbench.com/criteria/work.sycophancy_pushback.md) | Typical | 0.480 | 0.445–0.510 | 161 | 23 | 138 | | [Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md) | Typical | 0.505 | 0.446–0.556 | 153 | 12 | 141 | | [Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md) | Better than peers | 0.576 | 0.555–0.598 | 112 | 43 | 69 | | [Git commits, branches and sync](https://feedbackbench.com/criteria/work.git_workflow.md) | Typical | 0.517 | 0.494–0.540 | 112 | 47 | 65 | | [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | Typical | 0.486 | 0.455–0.516 | 100 | 50 | 50 | | [Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md) | Typical | 0.481 | 0.455–0.509 | 83 | 50 | 33 | | [Plan-before-edit mode](https://feedbackbench.com/criteria/work.plan_mode.md) | Better than peers | 0.535 | 0.507–0.561 | 78 | 50 | 28 | | [Games checks instead of fixing the problem](https://feedbackbench.com/criteria/work.reward_hacking.md) | Worse than peers | 0.448 | 0.435–0.461 | 45 | 0 | 45 | ### Checking and finishing: Typical (customer love 0.499, n 720) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md) | Worse than peers | 0.472 | 0.448–0.499 | 258 | 176 | 82 | | [Claims work is done or fixed when it is not](https://feedbackbench.com/criteria/verify.false_completion.md) | Typical | 0.473 | 0.402–0.528 | 219 | 10 | 209 | | [Builds, tests or runs its own changes](https://feedbackbench.com/criteria/verify.self_testing.md) | Typical | 0.515 | 0.489–0.538 | 209 | 115 | 94 | | [Reviewing and approving the agent's changes](https://feedbackbench.com/criteria/verify.change_review_ui.md) | Typical | 0.498 | 0.473–0.524 | 107 | 38 | 69 | ### Interface and sessions: Better than peers (customer love 0.560, n 1038) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | Better than peers | 0.580 | 0.553–0.607 | 467 | 208 | 259 | | [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | Typical | 0.470 | 0.441–0.502 | 221 | 132 | 89 | | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | Better than peers | 0.529 | 0.501–0.560 | 199 | 121 | 78 | | [Saving, switching, resuming and rewinding sessions](https://feedbackbench.com/criteria/ui.session_history.md) | Typical | 0.507 | 0.473–0.537 | 173 | 52 | 121 | | [Stopping and steering a running agent](https://feedbackbench.com/criteria/ui.interrupt_steer.md) | Typical | 0.490 | 0.469–0.511 | 40 | 12 | 28 | ### Reliability and speed: Better than peers (customer love 0.550, n 946) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Latency, throughput and fast mode](https://feedbackbench.com/criteria/rel.response_speed.md) | Better than peers | 0.576 | 0.545–0.606 | 348 | 165 | 183 | | [Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md) | Better than peers | 0.563 | 0.508–0.611 | 269 | 30 | 239 | | [Outages, server errors and capacity or rate errors](https://feedbackbench.com/criteria/rel.service_errors.md) | Typical | 0.463 | 0.395–0.517 | 254 | 15 | 239 | | [Updates break working setups](https://feedbackbench.com/criteria/rel.update_breakage.md) | Typical | 0.467 | 0.425–0.506 | 111 | 12 | 99 | ### Account and support: Worse than peers (customer love 0.401, n 703) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Support, refunds and issue handling](https://feedbackbench.com/criteria/account.support.md) | Worse than peers | 0.392 | 0.353–0.431 | 277 | 21 | 256 | | [Account bans and access restrictions](https://feedbackbench.com/criteria/account.bans_restrictions.md) | Typical | 0.519 | 0.453–0.574 | 224 | 15 | 209 | | [Data retention, training use and deployment isolation](https://feedbackbench.com/criteria/account.data_privacy.md) | Typical | 0.475 | 0.430–0.512 | 153 | 21 | 132 | | [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | Typical | 0.465 | 0.371–0.560 | 132 | 2 | 130 | # OpenAI Codex (OpenAI) Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/agent/codex | Measure | Value | |---|---| | Rank | 2 of 17 (rank range 2–2) | | Feedback Score | 67.4 (95% interval 66.9–67.8) | | Popularity | 1.000 (share of voice 30.27%) | | Customer love | 0.454 (95% interval 0.448–0.460) | | Top quadrant | no | | Authors | 29813 | | Posts counted | 118683 | | Posts that judge the agent | 48990 | | Criteria better / worse than peers | 7 / 21 of 63 | ## Top requests What users ask to add or change, most asked first. 4446 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first. | Rank | Request | Criterion | Author-weeks | Posts | |---|---|---|---|---| | 1 | Remove the 5-hour usage window | [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | 89 | 100 | | 2 | Additional or recurring bonus usage resets | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | 84 | 89 | | 3 | Higher overall usage limits | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | 79 | 82 | | 4 | Bankable usage resets | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | 71 | 73 | | 5 | Resets that keep the original reset date | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | 62 | 67 | | 6 | Higher-priced tier above current top plan | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | 58 | 61 | | 7 | Compensation reset after outages or bugs | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | 56 | 59 | | 8 | One-off usage limit reset now | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | 50 | 57 | | 9 | Higher allowance on entry and mid plans | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | 46 | 49 | | 10 | Predictable fixed reset schedule | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | 45 | 47 | | 11 | Cheaper model pricing | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | 45 | 46 | | 12 | Restore lost or missing resets | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | 44 | 47 | ## Facts | Fact | Value | |---|---| | Version | GPT-5.1-Codex Max (2025-12-04); GPT-5.3-Codex referenced on leaderboards; GPT-6 Astra (frontier, non-Codex-specific) released 2026-09-03 | | Released | GPT-5.1-Codex Max: 2025-12-04 | | Price | Free, Go $8/mo, Plus $20/mo, Pro $100-200/mo, Business $20-25/user/mo, Enterprise custom. API: GPT-5.1-Codex Max from $1.25/$10.00 per 1M tokens | | Model | GPT-5.1-Codex Max / GPT-5.3-Codex | | Surface | CLI, IDE extension, cloud (ChatGPT), API | ## Sources | Channel | Source | Posts | |---|---|---| | Reddit | r/codex | 103122 | | X | X search: OpenAI Codex, Codex CLI, Codex app | 9956 | | Reddit | Posts that name it | 5605 | ## Better than peers on Using an existing subscription across tools, Images, PDFs and file attachments as input, Computer use and browser control, Safety filters block legitimate coding tasks, Length and clarity of replies, summaries and comments, Agent-performed code review finds real issues, Support, refunds and issue handling ## Worse than peers on How much use a plan's price buys, Single prompt, model or effort level consumes disproportionate quota, Price, allowance or plan terms changed, Quota reset timing and bonus or banked resets, Usage meter visibility and accuracy, Prompt cache hits, misses and invalidation, Pricing and plan terms stated clearly and consistently, Which models are offered on a plan and when, Automatic model routing and fallback, Quality got worse or better over time, Frontend and visual UI output, Breaks existing code or reintroduces bugs, Stops mid-task or answers instead of acting, Subagents, parallel agents and orchestrators, Reviewing and approving the agent's changes, How the interface shows work, and what the user can configure, Saving, switching, resuming and rewinding sessions, Cloud and remote sandbox execution, Latency, throughput and fast mode, Client crashes, freezes and failed tool execution, Updates break working setups ## All 63 criteria Criterion love: 0.5 is the category norm. n: rated author-weeks. ### Paying and limits: Worse than peers (customer love 0.432, n 12312) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md) | Worse than peers | 0.463 | 0.449–0.476 | 5123 | 715 | 4408 | | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | Worse than peers | 0.450 | 0.438–0.461 | 4573 | 1500 | 3073 | | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | Worse than peers | 0.472 | 0.458–0.485 | 3239 | 481 | 2758 | | [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | Typical | 0.478 | 0.453–0.501 | 1424 | 174 | 1250 | | [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | Worse than peers | 0.462 | 0.417–0.499 | 1424 | 58 | 1366 | | [Usage meter visibility and accuracy](https://feedbackbench.com/criteria/limits.usage_meter.md) | Worse than peers | 0.390 | 0.343–0.433 | 998 | 39 | 959 | | [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md) | Worse than peers | 0.442 | 0.374–0.499 | 464 | 17 | 447 | | [Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md) | Worse than peers | 0.441 | 0.409–0.472 | 226 | 47 | 179 | | [Pay-as-you-go overage, fallback billing and spend caps](https://feedbackbench.com/criteria/billing.overage_charges.md) | Typical | 0.512 | 0.464–0.549 | 148 | 14 | 134 | | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | Better than peers | 0.585 | 0.556–0.612 | 147 | 93 | 54 | | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | Typical | 0.505 | 0.477–0.536 | 104 | 56 | 48 | ### Setting up and connecting: Worse than peers (customer love 0.443, n 992) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | Typical | 0.506 | 0.474–0.535 | 365 | 77 | 288 | | [MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md) | Typical | 0.469 | 0.439–0.500 | 269 | 122 | 147 | | [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | Typical | 0.508 | 0.476–0.540 | 141 | 80 | 61 | | [IDE and editor integration](https://feedbackbench.com/criteria/setup.ide_integration.md) | Typical | 0.487 | 0.457–0.518 | 135 | 56 | 79 | | [Onboarding, discoverability and documentation](https://feedbackbench.com/criteria/setup.onboarding_docs.md) | Typical | 0.466 | 0.426–0.500 | 107 | 15 | 92 | ### Choosing models: Worse than peers (customer love 0.424, n 3835) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md) | Worse than peers | 0.422 | 0.403–0.440 | 2537 | 398 | 2139 | | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | Worse than peers | 0.403 | 0.372–0.431 | 702 | 125 | 577 | | [Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md) | Typical | 0.504 | 0.484–0.524 | 561 | 260 | 301 | | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | Worse than peers | 0.396 | 0.361–0.429 | 433 | 57 | 376 | ### Instructing and context: Typical (customer love 0.499, n 1500) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md) | Typical | 0.512 | 0.485–0.537 | 518 | 144 | 374 | | [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | Typical | 0.510 | 0.479–0.541 | 291 | 101 | 190 | | [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md) | Typical | 0.501 | 0.472–0.532 | 265 | 141 | 124 | | [Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md) | Typical | 0.521 | 0.475–0.561 | 195 | 35 | 160 | | [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md) | Typical | 0.484 | 0.453–0.515 | 186 | 78 | 108 | | [Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md) | Typical | 0.481 | 0.451–0.510 | 115 | 42 | 73 | | [Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md) | Typical | 0.495 | 0.470–0.519 | 94 | 35 | 59 | | [Images, PDFs and file attachments as input](https://feedbackbench.com/criteria/context.attachments.md) | Better than peers | 0.526 | 0.501–0.549 | 51 | 26 | 25 | ### Doing the work: Typical (customer love 0.501, n 6815) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md) | Typical | 0.489 | 0.477–0.500 | 4052 | 2371 | 1681 | | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | Worse than peers | 0.462 | 0.443–0.482 | 1029 | 576 | 453 | | [Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md) | Typical | 0.501 | 0.456–0.537 | 597 | 35 | 562 | | [Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md) | Worse than peers | 0.473 | 0.451–0.495 | 403 | 153 | 250 | | [Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md) | Typical | 0.471 | 0.401–0.521 | 379 | 16 | 363 | | [Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md) | Typical | 0.476 | 0.449–0.503 | 329 | 228 | 101 | | [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | Better than peers | 0.534 | 0.511–0.561 | 292 | 176 | 116 | | [Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md) | Worse than peers | 0.448 | 0.390–0.498 | 255 | 12 | 243 | | [Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md) | Worse than peers | 0.450 | 0.412–0.485 | 234 | 15 | 219 | | [Safety filters block legitimate coding tasks](https://feedbackbench.com/criteria/work.safety_refusals.md) | Better than peers | 0.540 | 0.507–0.572 | 231 | 40 | 191 | | [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | Typical | 0.530 | 0.494–0.565 | 203 | 61 | 142 | | [Risky or irreversible actions without confirmation](https://feedbackbench.com/criteria/work.destructive_actions.md) | Typical | 0.482 | 0.441–0.518 | 194 | 32 | 162 | | [Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md) | Better than peers | 0.588 | 0.553–0.620 | 163 | 57 | 106 | | [Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md) | Typical | 0.516 | 0.491–0.539 | 137 | 95 | 42 | | [Caves to or argues with the user's judgement](https://feedbackbench.com/criteria/work.sycophancy_pushback.md) | Typical | 0.508 | 0.472–0.543 | 89 | 16 | 73 | | [Plan-before-edit mode](https://feedbackbench.com/criteria/work.plan_mode.md) | Typical | 0.474 | 0.447–0.501 | 67 | 28 | 39 | | [Git commits, branches and sync](https://feedbackbench.com/criteria/work.git_workflow.md) | Typical | 0.489 | 0.464–0.513 | 48 | 15 | 33 | | [Games checks instead of fixing the problem](https://feedbackbench.com/criteria/work.reward_hacking.md) | Typical | 0.498 | 0.448–0.539 | 37 | 1 | 36 | ### Checking and finishing: Typical (customer love 0.514, n 501) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md) | Better than peers | 0.542 | 0.510–0.572 | 201 | 159 | 42 | | [Claims work is done or fixed when it is not](https://feedbackbench.com/criteria/verify.false_completion.md) | Typical | 0.526 | 0.463–0.578 | 166 | 11 | 155 | | [Builds, tests or runs its own changes](https://feedbackbench.com/criteria/verify.self_testing.md) | Typical | 0.486 | 0.458–0.515 | 100 | 47 | 53 | | [Reviewing and approving the agent's changes](https://feedbackbench.com/criteria/verify.change_review_ui.md) | Worse than peers | 0.474 | 0.452–0.497 | 45 | 11 | 34 | ### Interface and sessions: Worse than peers (customer love 0.431, n 1406) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | Worse than peers | 0.438 | 0.413–0.463 | 802 | 225 | 577 | | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | Typical | 0.493 | 0.467–0.523 | 312 | 162 | 150 | | [Saving, switching, resuming and rewinding sessions](https://feedbackbench.com/criteria/ui.session_history.md) | Worse than peers | 0.458 | 0.426–0.491 | 211 | 48 | 163 | | [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | Worse than peers | 0.456 | 0.427–0.486 | 81 | 38 | 43 | | [Stopping and steering a running agent](https://feedbackbench.com/criteria/ui.interrupt_steer.md) | Typical | 0.506 | 0.485–0.528 | 54 | 21 | 33 | ### Reliability and speed: Worse than peers (customer love 0.407, n 2918) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Latency, throughput and fast mode](https://feedbackbench.com/criteria/rel.response_speed.md) | Worse than peers | 0.452 | 0.429–0.475 | 982 | 295 | 687 | | [Outages, server errors and capacity or rate errors](https://feedbackbench.com/criteria/rel.service_errors.md) | Typical | 0.471 | 0.434–0.506 | 915 | 63 | 852 | | [Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md) | Worse than peers | 0.410 | 0.358–0.458 | 883 | 41 | 842 | | [Updates break working setups](https://feedbackbench.com/criteria/rel.update_breakage.md) | Worse than peers | 0.407 | 0.366–0.445 | 347 | 30 | 317 | ### Account and support: Better than peers (customer love 0.537, n 748) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Support, refunds and issue handling](https://feedbackbench.com/criteria/account.support.md) | Better than peers | 0.548 | 0.510–0.579 | 312 | 63 | 249 | | [Account bans and access restrictions](https://feedbackbench.com/criteria/account.bans_restrictions.md) | Typical | 0.537 | 0.489–0.581 | 202 | 22 | 180 | | [Data retention, training use and deployment isolation](https://feedbackbench.com/criteria/account.data_privacy.md) | Typical | 0.479 | 0.440–0.518 | 161 | 24 | 137 | | [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | Typical | 0.512 | 0.442–0.561 | 143 | 7 | 136 | # OpenCode (Anomaly (open source)) Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/agent/opencode | Measure | Value | |---|---| | Rank | 3 of 17 (rank range 3–3) | | Feedback Score | 66.0 (95% interval 65.4–66.7) | | Popularity | 0.773 (share of voice 11.56%) | | Customer love | 0.564 (95% interval 0.553–0.575) | | Top quadrant | yes | | Authors | 11387 | | Posts counted | 25673 | | Posts that judge the agent | 10304 | | Criteria better / worse than peers | 7 / 7 of 63 | ## Top requests What users ask to add or change, most asked first. 1385 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first. | Rank | Request | Criterion | Author-weeks | Posts | |---|---|---|---|---| | 1 | Free access to specific or new models | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | 32 | 33 | | 2 | Make temporary usage bonuses permanent | [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | 21 | 24 | | 3 | Add DeepSeek V4.1 Flash model | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 21 | 21 | | 4 | Higher overall usage limits | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | 21 | 21 | | 5 | Max reasoning effort level option | [Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md) | 16 | 18 | | 6 | Allow subscription use in third-party harnesses | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | 16 | 16 | | 7 | Higher-priced tier above current top plan | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | 15 | 15 | | 8 | Option to restore previous UI design | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | 14 | 15 | | 9 | Cheaper low-cost plan tier | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | 12 | 12 | | 10 | Stop spurious 429 rate limit errors | [Outages, server errors and capacity or rate errors](https://feedbackbench.com/criteria/rel.service_errors.md) | 11 | 12 | | 11 | Faster model response speed | [Latency, throughput and fast mode](https://feedbackbench.com/criteria/rel.response_speed.md) | 11 | 11 | | 12 | Availability in more countries and regions | [Account bans and access restrictions](https://feedbackbench.com/criteria/account.bans_restrictions.md) | 10 | 10 | ## Facts | Fact | Value | |---|---| | Version | n/a (fast release cadence) | | Released | #1 on Hacker News: 2026-03-20 (1,099 points, 546 comments) | | Price | Free, open source, BYOK to 75+ model providers | | Model | Model-agnostic, BYOK | | Surface | Terminal (TUI), desktop app | ## Sources | Channel | Source | Posts | |---|---|---| | Reddit | r/opencode | 10411 | | X | @opencode | 9118 | | Reddit | r/opencodeCLI | 4794 | | Reddit | Posts that name it | 1345 | | Trustpilot | Trustpilot | 5 | ## Better than peers on How much use a plan's price buys, Price, allowance or plan terms changed, Using an existing subscription across tools, Which models are offered on a plan and when, Subagents, parallel agents and orchestrators, Mobile, remote-control and voice access, Updates break working setups ## Worse than peers on Short rolling usage window blocks or interrupts work, Reasoning effort setting and its defaults, Output degrades as the context window fills, Can do the user's kind of task, Long unattended runs and goal/loop mode, Wrong charges, failed payments and plan provisioning, Account bans and access restrictions ## All 63 criteria Criterion love: 0.5 is the category norm. n: rated author-weeks. ### Paying and limits: Better than peers (customer love 0.660, n 2435) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | Better than peers | 0.657 | 0.635–0.677 | 1075 | 615 | 460 | | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | Typical | 0.494 | 0.472–0.516 | 571 | 322 | 249 | | [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md) | Typical | 0.537 | 0.489–0.578 | 371 | 74 | 297 | | [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | Better than peers | 0.593 | 0.538–0.644 | 217 | 29 | 188 | | [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md) | Typical | 0.486 | 0.421–0.548 | 201 | 12 | 189 | | [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | Worse than peers | 0.446 | 0.402–0.491 | 125 | 10 | 115 | | [Usage meter visibility and accuracy](https://feedbackbench.com/criteria/limits.usage_meter.md) | Typical | 0.471 | 0.420–0.524 | 124 | 9 | 115 | | [Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md) | Typical | 0.505 | 0.473–0.535 | 96 | 34 | 62 | | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | Better than peers | 0.530 | 0.501–0.560 | 92 | 46 | 46 | | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | Typical | 0.473 | 0.444–0.502 | 49 | 5 | 44 | | [Pay-as-you-go overage, fallback billing and spend caps](https://feedbackbench.com/criteria/billing.overage_charges.md) | Too few posts | 0.552 | 0.500–0.600 | 26 | 6 | 20 | ### Setting up and connecting: Typical (customer love 0.517, n 514) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | Typical | 0.515 | 0.482–0.543 | 200 | 116 | 84 | | [MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md) | Typical | 0.481 | 0.446–0.513 | 138 | 63 | 75 | | [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | Typical | 0.524 | 0.480–0.564 | 119 | 29 | 90 | | [Onboarding, discoverability and documentation](https://feedbackbench.com/criteria/setup.onboarding_docs.md) | Typical | 0.504 | 0.471–0.539 | 54 | 11 | 43 | | [IDE and editor integration](https://feedbackbench.com/criteria/setup.ide_integration.md) | Typical | 0.490 | 0.465–0.514 | 30 | 11 | 19 | ### Choosing models: Better than peers (customer love 0.546, n 771) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | Better than peers | 0.575 | 0.542–0.606 | 415 | 160 | 255 | | [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md) | Typical | 0.531 | 0.489–0.575 | 213 | 60 | 153 | | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | Typical | 0.503 | 0.466–0.537 | 114 | 30 | 84 | | [Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md) | Worse than peers | 0.445 | 0.420–0.470 | 57 | 12 | 45 | ### Instructing and context: Typical (customer love 0.465, n 318) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md) | Typical | 0.504 | 0.470–0.541 | 76 | 22 | 54 | | [Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md) | Worse than peers | 0.459 | 0.427–0.499 | 74 | 6 | 68 | | [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | Typical | 0.490 | 0.461–0.522 | 72 | 22 | 50 | | [Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md) | Typical | 0.486 | 0.464–0.510 | 34 | 12 | 22 | | [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md) | Typical | 0.497 | 0.473–0.519 | 30 | 13 | 17 | | [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md) | Too few posts | 0.516 | 0.494–0.538 | 27 | 18 | 9 | | [Images, PDFs and file attachments as input](https://feedbackbench.com/criteria/context.attachments.md) | Too few posts | 0.505 | 0.484–0.524 | 23 | 10 | 13 | | [Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md) | Too few posts | 0.506 | 0.488–0.524 | 15 | 7 | 8 | ### Doing the work: Typical (customer love 0.500, n 1369) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md) | Worse than peers | 0.467 | 0.442–0.496 | 859 | 502 | 357 | | [Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md) | Typical | 0.501 | 0.426–0.569 | 167 | 9 | 158 | | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | Better than peers | 0.537 | 0.503–0.573 | 167 | 114 | 53 | | [Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md) | Worse than peers | 0.457 | 0.421–0.491 | 48 | 28 | 20 | | [Risky or irreversible actions without confirmation](https://feedbackbench.com/criteria/work.destructive_actions.md) | Typical | 0.493 | 0.464–0.526 | 38 | 6 | 32 | | [Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md) | Typical | 0.518 | 0.494–0.539 | 35 | 28 | 7 | | [Plan-before-edit mode](https://feedbackbench.com/criteria/work.plan_mode.md) | Typical | 0.488 | 0.466–0.511 | 33 | 13 | 20 | | [Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md) | Typical | 0.503 | 0.476–0.532 | 33 | 8 | 25 | | [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | Typical | 0.495 | 0.470–0.525 | 32 | 7 | 25 | | [Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md) | Typical | 0.487 | 0.457–0.524 | 31 | 2 | 29 | | [Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md) | Typical | 0.508 | 0.486–0.530 | 30 | 17 | 13 | | [Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md) | Too few posts | 0.475 | 0.457–0.496 | 25 | 2 | 23 | | [Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md) | Too few posts | 0.549 | 0.494–0.605 | 20 | 4 | 16 | | [Safety filters block legitimate coding tasks](https://feedbackbench.com/criteria/work.safety_refusals.md) | Too few posts | 0.520 | 0.485–0.558 | 20 | 4 | 16 | | [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | Too few posts | 0.505 | 0.487–0.523 | 18 | 11 | 7 | | [Git commits, branches and sync](https://feedbackbench.com/criteria/work.git_workflow.md) | Too few posts | 0.492 | 0.477–0.508 | 13 | 3 | 10 | | [Caves to or argues with the user's judgement](https://feedbackbench.com/criteria/work.sycophancy_pushback.md) | Too few posts | 0.496 | 0.483–0.514 | 8 | 1 | 7 | | [Games checks instead of fixing the problem](https://feedbackbench.com/criteria/work.reward_hacking.md) | Too few posts | 0.491 | 0.485–0.497 | 7 | 0 | 7 | ### Checking and finishing: Typical (customer love 0.508, n 59) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Claims work is done or fixed when it is not](https://feedbackbench.com/criteria/verify.false_completion.md) | Too few posts | 0.520 | 0.476–0.573 | 18 | 2 | 16 | | [Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md) | Too few posts | 0.507 | 0.486–0.526 | 16 | 13 | 3 | | [Builds, tests or runs its own changes](https://feedbackbench.com/criteria/verify.self_testing.md) | Too few posts | 0.496 | 0.480–0.512 | 13 | 7 | 6 | | [Reviewing and approving the agent's changes](https://feedbackbench.com/criteria/verify.change_review_ui.md) | Too few posts | 0.509 | 0.494–0.526 | 12 | 7 | 5 | ### Interface and sessions: Typical (customer love 0.498, n 428) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | Typical | 0.523 | 0.488–0.556 | 310 | 117 | 193 | | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | Better than peers | 0.530 | 0.504–0.557 | 70 | 45 | 25 | | [Saving, switching, resuming and rewinding sessions](https://feedbackbench.com/criteria/ui.session_history.md) | Typical | 0.497 | 0.466–0.526 | 61 | 17 | 44 | | [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | Too few posts | 0.502 | 0.483–0.520 | 15 | 10 | 5 | | [Stopping and steering a running agent](https://feedbackbench.com/criteria/ui.interrupt_steer.md) | Too few posts | 0.504 | 0.489–0.519 | 10 | 5 | 5 | ### Reliability and speed: Better than peers (customer love 0.533, n 1251) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Latency, throughput and fast mode](https://feedbackbench.com/criteria/rel.response_speed.md) | Typical | 0.492 | 0.462–0.519 | 532 | 200 | 332 | | [Outages, server errors and capacity or rate errors](https://feedbackbench.com/criteria/rel.service_errors.md) | Typical | 0.541 | 0.485–0.591 | 394 | 36 | 358 | | [Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md) | Typical | 0.558 | 0.496–0.611 | 327 | 32 | 295 | | [Updates break working setups](https://feedbackbench.com/criteria/rel.update_breakage.md) | Better than peers | 0.552 | 0.505–0.591 | 110 | 26 | 84 | ### Account and support: Typical (customer love 0.523, n 428) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Data retention, training use and deployment isolation](https://feedbackbench.com/criteria/account.data_privacy.md) | Typical | 0.480 | 0.447–0.512 | 214 | 44 | 170 | | [Support, refunds and issue handling](https://feedbackbench.com/criteria/account.support.md) | Typical | 0.530 | 0.487–0.571 | 97 | 21 | 76 | | [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | Worse than peers | 0.420 | 0.404–0.435 | 74 | 0 | 74 | | [Account bans and access restrictions](https://feedbackbench.com/criteria/account.bans_restrictions.md) | Worse than peers | 0.422 | 0.407–0.438 | 69 | 0 | 69 | # Cursor (Anysphere) Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/agent/cursor | Measure | Value | |---|---| | Rank | 4 of 17 (rank range 4–4) | | Feedback Score | 59.7 (95% interval 59.0–60.5) | | Popularity | 0.708 (share of voice 8.74%) | | Customer love | 0.504 (95% interval 0.491–0.516) | | Top quadrant | yes, at the edge | | Authors | 8609 | | Posts counted | 18878 | | Posts that judge the agent | 9985 | | Criteria better / worse than peers | 5 / 10 of 63 | ## Top requests What users ask to add or change, most asked first. 1806 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first. | Rank | Request | Criterion | Author-weeks | Posts | |---|---|---|---|---| | 1 | Native Android app | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | 36 | 50 | | 2 | Release Composer 3 model | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 36 | 37 | | 3 | One-off usage limit reset now | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | 35 | 37 | | 4 | Unified subscription across linked products | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | 27 | 28 | | 5 | Add GPT-6 Astra model | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 26 | 27 | | 6 | Add Grok 4.7 model | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 24 | 24 | | 7 | Higher overall usage limits | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | 19 | 19 | | 8 | No silent model switching or downgrades | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | 18 | 21 | | 9 | Additional or recurring bonus usage resets | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | 18 | 19 | | 10 | Refund unauthorized or incorrect charges | [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | 16 | 22 | | 11 | Restore original other-models allowance pool | [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | 15 | 30 | | 12 | Usage reset for new model or feature launch | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | 15 | 16 | ## Facts | Fact | Value | |---|---| | Version | Composer 2.5 (agent model); Cursor 2.0 (multi-agent editor, git-worktree parallel agents) | | Released | Composer 2.5: 2026-05-18. Composer 2: 2026-03-18/19. Cursor 2.0: 2025-10 | | Price | Free (Hobby); Pro ~$20/mo credit-pool; Business/Enterprise custom. Composer 2.5 API: $0.50/$2.50 per 1M tokens standard, $3.00/$15.00 Fast (default) | | Model | Composer 2.5 (proprietary) plus BYO access to Claude, GPT, Gemini | | Surface | IDE (VS Code fork) | ## Sources | Channel | Source | Posts | |---|---|---| | X | @cursor_ai | 9703 | | Reddit | r/cursor | 8303 | | Reddit | Posts that name it | 838 | | Trustpilot | Trustpilot | 29 | | G2 | G2 | 5 | ## Better than peers on How much use a plan's price buys, IDE and editor integration, Which models are offered on a plan and when, Can do the user's kind of task, Plan-before-edit mode ## Worse than peers on Quota reset timing and bonus or banked resets, Usage meter visibility and accuracy, Pay-as-you-go overage, fallback billing and spend caps, Using an existing subscription across tools, Connecting own API keys, local models and custom endpoints, Automatic model routing and fallback, Quality got worse or better over time, How the interface shows work, and what the user can configure, Mobile, remote-control and voice access, Support, refunds and issue handling ## All 63 criteria Criterion love: 0.5 is the category norm. n: rated author-weeks. ### Paying and limits: Typical (customer love 0.506, n 2079) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | Better than peers | 0.537 | 0.513–0.562 | 1001 | 427 | 574 | | [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md) | Typical | 0.525 | 0.487–0.558 | 621 | 119 | 502 | | [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md) | Typical | 0.461 | 0.400–0.518 | 189 | 11 | 178 | | [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | Typical | 0.453 | 0.397–0.510 | 170 | 10 | 160 | | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | Worse than peers | 0.442 | 0.399–0.481 | 136 | 15 | 121 | | [Usage meter visibility and accuracy](https://feedbackbench.com/criteria/limits.usage_meter.md) | Worse than peers | 0.449 | 0.400–0.494 | 132 | 8 | 124 | | [Pay-as-you-go overage, fallback billing and spend caps](https://feedbackbench.com/criteria/billing.overage_charges.md) | Worse than peers | 0.430 | 0.393–0.479 | 105 | 3 | 102 | | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | Worse than peers | 0.429 | 0.403–0.454 | 69 | 7 | 62 | | [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | Typical | 0.540 | 0.496–0.581 | 51 | 13 | 38 | | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | Typical | 0.482 | 0.457–0.509 | 43 | 20 | 23 | | [Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md) | Too few posts | 0.501 | 0.485–0.518 | 22 | 11 | 11 | ### Setting up and connecting: Typical (customer love 0.498, n 431) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md) | Typical | 0.481 | 0.447–0.515 | 123 | 55 | 68 | | [IDE and editor integration](https://feedbackbench.com/criteria/setup.ide_integration.md) | Better than peers | 0.576 | 0.547–0.604 | 111 | 74 | 37 | | [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | Typical | 0.479 | 0.441–0.515 | 88 | 13 | 75 | | [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | Worse than peers | 0.431 | 0.401–0.460 | 86 | 26 | 60 | | [Onboarding, discoverability and documentation](https://feedbackbench.com/criteria/setup.onboarding_docs.md) | Typical | 0.500 | 0.468–0.531 | 42 | 8 | 34 | ### Choosing models: Typical (customer love 0.489, n 959) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | Better than peers | 0.550 | 0.515–0.582 | 361 | 128 | 233 | | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | Worse than peers | 0.467 | 0.433–0.499 | 343 | 73 | 270 | | [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md) | Worse than peers | 0.443 | 0.404–0.483 | 314 | 54 | 260 | | [Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md) | Too few posts | 0.512 | 0.492–0.535 | 29 | 16 | 13 | ### Instructing and context: Better than peers (customer love 0.542, n 318) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md) | Typical | 0.524 | 0.494–0.556 | 110 | 64 | 46 | | [Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md) | Typical | 0.523 | 0.496–0.551 | 64 | 36 | 28 | | [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md) | Typical | 0.487 | 0.457–0.515 | 62 | 29 | 33 | | [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md) | Typical | 0.479 | 0.451–0.505 | 50 | 10 | 40 | | [Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md) | Too few posts | 0.521 | 0.489–0.552 | 28 | 7 | 21 | | [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | Too few posts | 0.507 | 0.488–0.527 | 21 | 9 | 12 | | [Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md) | Too few posts | 0.491 | 0.478–0.504 | 10 | 2 | 8 | | [Images, PDFs and file attachments as input](https://feedbackbench.com/criteria/context.attachments.md) | Too few posts | 0.496 | 0.484–0.510 | 7 | 2 | 5 | ### Doing the work: Better than peers (customer love 0.533, n 1252) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md) | Better than peers | 0.555 | 0.526–0.584 | 639 | 444 | 195 | | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | Typical | 0.535 | 0.499–0.571 | 211 | 143 | 68 | | [Risky or irreversible actions without confirmation](https://feedbackbench.com/criteria/work.destructive_actions.md) | Typical | 0.475 | 0.445–0.509 | 61 | 7 | 54 | | [Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md) | Typical | 0.503 | 0.472–0.534 | 56 | 44 | 12 | | [Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md) | Typical | 0.459 | 0.428–0.503 | 54 | 1 | 53 | | [Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md) | Typical | 0.498 | 0.473–0.521 | 50 | 27 | 23 | | [Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md) | Typical | 0.510 | 0.482–0.538 | 50 | 11 | 39 | | [Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md) | Typical | 0.473 | 0.435–0.513 | 50 | 2 | 48 | | [Git commits, branches and sync](https://feedbackbench.com/criteria/work.git_workflow.md) | Typical | 0.485 | 0.461–0.508 | 45 | 12 | 33 | | [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | Typical | 0.482 | 0.458–0.505 | 38 | 17 | 21 | | [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | Typical | 0.505 | 0.475–0.532 | 37 | 10 | 27 | | [Plan-before-edit mode](https://feedbackbench.com/criteria/work.plan_mode.md) | Better than peers | 0.524 | 0.501–0.547 | 37 | 23 | 14 | | [Safety filters block legitimate coding tasks](https://feedbackbench.com/criteria/work.safety_refusals.md) | Typical | 0.484 | 0.456–0.515 | 30 | 2 | 28 | | [Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md) | Too few posts | 0.485 | 0.466–0.509 | 22 | 3 | 19 | | [Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md) | Too few posts | 0.502 | 0.485–0.520 | 18 | 13 | 5 | | [Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md) | Too few posts | 0.487 | 0.478–0.495 | 8 | 0 | 8 | | [Caves to or argues with the user's judgement](https://feedbackbench.com/criteria/work.sycophancy_pushback.md) | Too few posts | 0.498 | 0.489–0.517 | 5 | 1 | 4 | | [Games checks instead of fixing the problem](https://feedbackbench.com/criteria/work.reward_hacking.md) | Too few posts | 0.495 | 0.490–0.499 | 4 | 0 | 4 | ### Checking and finishing: Better than peers (customer love 0.542, n 155) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Reviewing and approving the agent's changes](https://feedbackbench.com/criteria/verify.change_review_ui.md) | Typical | 0.526 | 0.498–0.552 | 52 | 26 | 26 | | [Builds, tests or runs its own changes](https://feedbackbench.com/criteria/verify.self_testing.md) | Typical | 0.502 | 0.478–0.526 | 51 | 33 | 18 | | [Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md) | Typical | 0.515 | 0.491–0.542 | 36 | 30 | 6 | | [Claims work is done or fixed when it is not](https://feedbackbench.com/criteria/verify.false_completion.md) | Too few posts | 0.509 | 0.473–0.557 | 19 | 2 | 17 | ### Interface and sessions: Typical (customer love 0.494, n 560) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | Typical | 0.503 | 0.472–0.536 | 207 | 138 | 69 | | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | Worse than peers | 0.451 | 0.418–0.486 | 201 | 53 | 148 | | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | Worse than peers | 0.437 | 0.408–0.468 | 120 | 40 | 80 | | [Saving, switching, resuming and rewinding sessions](https://feedbackbench.com/criteria/ui.session_history.md) | Typical | 0.506 | 0.476–0.537 | 59 | 19 | 40 | | [Stopping and steering a running agent](https://feedbackbench.com/criteria/ui.interrupt_steer.md) | Too few posts | 0.489 | 0.473–0.505 | 18 | 6 | 12 | ### Reliability and speed: Worse than peers (customer love 0.446, n 690) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Outages, server errors and capacity or rate errors](https://feedbackbench.com/criteria/rel.service_errors.md) | Typical | 0.488 | 0.420–0.551 | 257 | 17 | 240 | | [Latency, throughput and fast mode](https://feedbackbench.com/criteria/rel.response_speed.md) | Typical | 0.468 | 0.436–0.500 | 211 | 68 | 143 | | [Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md) | Typical | 0.481 | 0.414–0.544 | 200 | 12 | 188 | | [Updates break working setups](https://feedbackbench.com/criteria/rel.update_breakage.md) | Typical | 0.494 | 0.454–0.532 | 62 | 9 | 53 | ### Account and support: Worse than peers (customer love 0.436, n 524) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Support, refunds and issue handling](https://feedbackbench.com/criteria/account.support.md) | Worse than peers | 0.431 | 0.389–0.468 | 324 | 36 | 288 | | [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | Typical | 0.508 | 0.389–0.601 | 184 | 4 | 180 | | [Account bans and access restrictions](https://feedbackbench.com/criteria/account.bans_restrictions.md) | Typical | 0.488 | 0.426–0.546 | 83 | 4 | 79 | | [Data retention, training use and deployment isolation](https://feedbackbench.com/criteria/account.data_privacy.md) | Typical | 0.495 | 0.465–0.528 | 51 | 11 | 40 | # Devin (Cognition) Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/agent/devin | Measure | Value | |---|---| | Rank | =5 of 17 (rank range 5–6) | | Feedback Score | 55.9 (95% interval 55.1–56.5) | | Popularity | 0.514 (share of voice 3.66%) | | Customer love | 0.607 (95% interval 0.591–0.622) | | Top quadrant | yes | | Authors | 3603 | | Posts counted | 6711 | | Posts that judge the agent | 3237 | | Criteria better / worse than peers | 9 / 2 of 63 | ## Top requests What users ask to add or change, most asked first. 480 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first. | Rank | Request | Criterion | Author-weeks | Posts | |---|---|---|---|---| | 1 | Mid-priced tier between existing plans | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | 20 | 24 | | 2 | Official dedicated mobile app | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | 16 | 16 | | 3 | Higher overall usage limits | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | 14 | 15 | | 4 | Keep free models available permanently | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | 9 | 9 | | 5 | Free trial periods for paid plans | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | 8 | 9 | | 6 | Mac VM with iOS simulator | [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | 8 | 8 | | 7 | Published exact usage limits per plan | [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md) | 7 | 8 | | 8 | Built-in computer use capability | [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | 7 | 7 | | 9 | Bring-your-own-key support | [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | 6 | 8 | | 10 | Free access to top-tier max plan | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | 6 | 6 | | 11 | Preserve full context in session handoffs | [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | 5 | 7 | | 12 | Free access to specific or new models | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | 5 | 6 | ## Facts | Fact | Value | |---|---| | Version | SWE-2 model (Medium/High/Max reasoning), Devin Fusion multi-model harness | | Released | SWE-2: 2026 (exact date not confirmed in research) | | Price | Free, Core/Pro $20/seat/mo, Max $200/seat/mo, Teams $80/mo base + $40/full-dev-seat, Enterprise custom. ACU (Agent Compute Unit) ~$2.25 each, ~15 min of autonomous work per ACU | | Model | SWE-2 (enterprise API pricing $3.00/$15.00 per 1M tokens, 75% off list through 2026-12-31) | | Surface | cloud, desktop (converging with Devin Desktop/Windsurf), IDE plugins | ## Sources | Channel | Source | Posts | |---|---|---| | X | @cognition | 4111 | | X | @DevinAI | 1984 | | Reddit | r/windsurf | 438 | | Reddit | Posts that name it | 113 | | Reddit | r/CognitionLabs | 59 | | G2 | G2 | 6 | ## Better than peers on How much use a plan's price buys, Single prompt, model or effort level consumes disproportionate quota, Free tier and free model availability and limits, Which models are offered on a plan and when, Automatic model routing and fallback, Quality got worse or better over time, Can do the user's kind of task, Long unattended runs and goal/loop mode, Cloud and remote sandbox execution ## Worse than peers on Short rolling usage window blocks or interrupts work, Latency, throughput and fast mode ## All 63 criteria Criterion love: 0.5 is the category norm. n: rated author-weeks. ### Paying and limits: Better than peers (customer love 0.669, n 691) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | Better than peers | 0.619 | 0.588–0.649 | 365 | 210 | 155 | | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | Better than peers | 0.563 | 0.528–0.593 | 155 | 114 | 41 | | [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md) | Better than peers | 0.584 | 0.543–0.623 | 139 | 44 | 95 | | [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | Worse than peers | 0.456 | 0.437–0.479 | 40 | 1 | 39 | | [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md) | Typical | 0.495 | 0.457–0.540 | 38 | 3 | 35 | | [Usage meter visibility and accuracy](https://feedbackbench.com/criteria/limits.usage_meter.md) | Too few posts | 0.499 | 0.465–0.538 | 27 | 3 | 24 | | [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | Too few posts | 0.506 | 0.473–0.542 | 20 | 2 | 18 | | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | Too few posts | 0.503 | 0.485–0.528 | 13 | 3 | 10 | | [Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md) | Too few posts | 0.506 | 0.490–0.520 | 13 | 9 | 4 | | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | Too few posts | 0.493 | 0.478–0.509 | 11 | 2 | 9 | | [Pay-as-you-go overage, fallback billing and spend caps](https://feedbackbench.com/criteria/billing.overage_charges.md) | Too few posts | 0.495 | 0.489–0.499 | 4 | 0 | 4 | ### Setting up and connecting: Typical (customer love 0.511, n 92) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | Too few posts | 0.556 | 0.522–0.589 | 25 | 13 | 12 | | [MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md) | Too few posts | 0.502 | 0.481–0.523 | 25 | 13 | 12 | | [Onboarding, discoverability and documentation](https://feedbackbench.com/criteria/setup.onboarding_docs.md) | Too few posts | 0.499 | 0.477–0.524 | 17 | 3 | 14 | | [IDE and editor integration](https://feedbackbench.com/criteria/setup.ide_integration.md) | Too few posts | 0.515 | 0.497–0.534 | 16 | 10 | 6 | | [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | Too few posts | 0.473 | 0.455–0.489 | 13 | 1 | 12 | ### Choosing models: Better than peers (customer love 0.636, n 176) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md) | Better than peers | 0.551 | 0.520–0.583 | 66 | 32 | 34 | | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | Better than peers | 0.600 | 0.567–0.633 | 62 | 43 | 19 | | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | Better than peers | 0.559 | 0.529–0.588 | 39 | 25 | 14 | | [Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md) | Too few posts | 0.507 | 0.489–0.525 | 17 | 9 | 8 | ### Instructing and context: Typical (customer love 0.500, n 46) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md) | Too few posts | 0.505 | 0.490–0.522 | 13 | 8 | 5 | | [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md) | Too few posts | 0.498 | 0.483–0.512 | 10 | 5 | 5 | | [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md) | Too few posts | 0.501 | 0.488–0.517 | 9 | 3 | 6 | | [Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md) | Too few posts | 0.492 | 0.485–0.498 | 4 | 0 | 4 | | [Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md) | Too few posts | 0.494 | 0.488–0.499 | 4 | 0 | 4 | | [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | Too few posts | 0.499 | 0.492–0.509 | 3 | 1 | 2 | | [Images, PDFs and file attachments as input](https://feedbackbench.com/criteria/context.attachments.md) | Too few posts | 0.506 | 0.500–0.516 | 2 | 2 | 0 | | [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md) | Too few posts | 0.497 | 0.492–0.500 | 1 | 0 | 1 | ### Doing the work: Better than peers (customer love 0.602, n 620) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md) | Better than peers | 0.554 | 0.522–0.587 | 449 | 330 | 119 | | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | Typical | 0.487 | 0.456–0.519 | 82 | 47 | 35 | | [Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md) | Better than peers | 0.539 | 0.511–0.563 | 47 | 43 | 4 | | [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | Too few posts | 0.501 | 0.480–0.522 | 27 | 16 | 11 | | [Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md) | Too few posts | 0.524 | 0.487–0.558 | 13 | 3 | 10 | | [Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md) | Too few posts | 0.517 | 0.485–0.563 | 12 | 2 | 10 | | [Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md) | Too few posts | 0.509 | 0.493–0.526 | 12 | 6 | 6 | | [Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md) | Too few posts | 0.490 | 0.478–0.503 | 10 | 1 | 9 | | [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | Too few posts | 0.492 | 0.481–0.508 | 9 | 1 | 8 | | [Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md) | Too few posts | 0.493 | 0.479–0.507 | 8 | 4 | 4 | | [Risky or irreversible actions without confirmation](https://feedbackbench.com/criteria/work.destructive_actions.md) | Too few posts | 0.499 | 0.485–0.517 | 7 | 1 | 6 | | [Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md) | Too few posts | 0.490 | 0.482–0.498 | 5 | 0 | 5 | | [Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md) | Too few posts | 0.507 | 0.502–0.513 | 4 | 4 | 0 | | [Git commits, branches and sync](https://feedbackbench.com/criteria/work.git_workflow.md) | Too few posts | 0.502 | 0.495–0.511 | 2 | 1 | 1 | | [Plan-before-edit mode](https://feedbackbench.com/criteria/work.plan_mode.md) | Too few posts | 0.505 | 0.500–0.515 | 2 | 2 | 0 | | [Games checks instead of fixing the problem](https://feedbackbench.com/criteria/work.reward_hacking.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Safety filters block legitimate coding tasks](https://feedbackbench.com/criteria/work.safety_refusals.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Caves to or argues with the user's judgement](https://feedbackbench.com/criteria/work.sycophancy_pushback.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Checking and finishing: Typical (customer love 0.494, n 61) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Builds, tests or runs its own changes](https://feedbackbench.com/criteria/verify.self_testing.md) | Too few posts | 0.514 | 0.495–0.532 | 23 | 18 | 5 | | [Reviewing and approving the agent's changes](https://feedbackbench.com/criteria/verify.change_review_ui.md) | Too few posts | 0.473 | 0.456–0.490 | 21 | 3 | 18 | | [Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md) | Too few posts | 0.500 | 0.481–0.515 | 10 | 8 | 2 | | [Claims work is done or fixed when it is not](https://feedbackbench.com/criteria/verify.false_completion.md) | Too few posts | 0.489 | 0.482–0.496 | 8 | 0 | 8 | ### Interface and sessions: Better than peers (customer love 0.592, n 157) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | Better than peers | 0.571 | 0.541–0.599 | 88 | 77 | 11 | | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | Typical | 0.508 | 0.482–0.537 | 47 | 19 | 28 | | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | Too few posts | 0.495 | 0.477–0.515 | 21 | 9 | 12 | | [Saving, switching, resuming and rewinding sessions](https://feedbackbench.com/criteria/ui.session_history.md) | Too few posts | 0.487 | 0.475–0.501 | 11 | 1 | 10 | | [Stopping and steering a running agent](https://feedbackbench.com/criteria/ui.interrupt_steer.md) | Too few posts | 0.500 | 0.489–0.511 | 6 | 3 | 3 | ### Reliability and speed: Typical (customer love 0.533, n 149) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Latency, throughput and fast mode](https://feedbackbench.com/criteria/rel.response_speed.md) | Worse than peers | 0.468 | 0.438–0.497 | 70 | 21 | 49 | | [Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md) | Typical | 0.497 | 0.446–0.557 | 50 | 3 | 47 | | [Outages, server errors and capacity or rate errors](https://feedbackbench.com/criteria/rel.service_errors.md) | Too few posts | 0.589 | 0.527–0.644 | 27 | 8 | 19 | | [Updates break working setups](https://feedbackbench.com/criteria/rel.update_breakage.md) | Too few posts | 0.533 | 0.501–0.567 | 14 | 6 | 8 | ### Account and support: Too few posts (customer love 0.522, n 29) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Support, refunds and issue handling](https://feedbackbench.com/criteria/account.support.md) | Too few posts | 0.512 | 0.485–0.537 | 19 | 5 | 14 | | [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | Too few posts | 0.489 | 0.482–0.495 | 9 | 0 | 9 | | [Data retention, training use and deployment isolation](https://feedbackbench.com/criteria/account.data_privacy.md) | Too few posts | 0.508 | 0.496–0.526 | 3 | 2 | 1 | | [Account bans and access restrictions](https://feedbackbench.com/criteria/account.bans_restrictions.md) | Too few posts | 0.499 | 0.495–0.500 | 1 | 0 | 1 | # Google Antigravity (Google) Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/agent/antigravity | Measure | Value | |---|---| | Rank | =5 of 17 (rank range 5–6) | | Feedback Score | 54.9 (95% interval 54.0–55.7) | | Popularity | 0.650 (share of voice 6.78%) | | Customer love | 0.463 (95% interval 0.449–0.477) | | Top quadrant | no | | Authors | 6673 | | Posts counted | 14809 | | Posts that judge the agent | 7648 | | Criteria better / worse than peers | 3 / 18 of 63 | ## Top requests What users ask to add or change, most asked first. 1444 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first. | Rank | Request | Criterion | Author-weeks | Posts | |---|---|---|---|---| | 1 | Allow subscription use in third-party harnesses | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | 44 | 46 | | 2 | Update outdated Claude models in catalog | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 36 | 37 | | 3 | Auto-approve mode without permission prompts | [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | 25 | 26 | | 4 | One-off usage limit reset now | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | 23 | 24 | | 5 | Higher overall usage limits | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | 20 | 21 | | 6 | Add Opus 5.5 model | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 19 | 21 | | 7 | Faster model response speed | [Latency, throughput and fast mode](https://feedbackbench.com/criteria/rel.response_speed.md) | 19 | 19 | | 8 | Stronger pro-tier frontier model | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 18 | 19 | | 9 | Release and add Gemini 4 Pro | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 17 | 19 | | 10 | Fewer permission prompts overall | [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | 16 | 16 | | 11 | Add newer Gemini Flash and Pro models | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 15 | 15 | | 12 | Keep model catalog updated to latest versions | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 14 | 14 | ## Facts | Fact | Value | |---|---| | Version | Antigravity 2.0 (desktop app, Go-based CLI, SDK); absorbed and replaced Gemini CLI | | Released | Antigravity 1: 2025-11-18. Antigravity 2.0: 2026-05-19. Consumer Gemini CLI/Code Assist cutover: 2026-06-18 | | Price | Bundled with Google AI Pro ($19.99/mo) / Ultra ($99.99/mo) tiers (shared with Gemini consumer pricing) | | Model | Gemini (3.1 Pro and successors) | | Surface | Desktop IDE app, CLI, SDK | ## Sources | Channel | Source | Posts | |---|---|---| | Reddit | r/google_antigravity | 9004 | | X | @antigravity | 4774 | | Reddit | r/GoogleAntigravityIDE | 502 | | Reddit | Posts that name it | 478 | | Reddit | r/GeminiCLI | 40 | | X | @geminicli | 7 | | Trustpilot | Trustpilot | 4 | ## Better than peers on How much use a plan's price buys, Free tier and free model availability and limits, Quality got worse or better over time ## Worse than peers on Price, allowance or plan terms changed, Quota reset timing and bonus or banked resets, Using an existing subscription across tools, Install, launch and sign-in, MCP servers, plugins, skills and hooks, IDE and editor integration, Which models are offered on a plan and when, Context compaction keeps what matters, cheaply and quickly, Memory and state carried across sessions, Images, PDFs and file attachments as input, Can do the user's kind of task, Computer use and browser control, Tool approval prompts and autonomy modes, Claims work is done or fixed when it is not, How the interface shows work, and what the user can configure, Saving, switching, resuming and rewinding sessions, Outages, server errors and capacity or rate errors, Client crashes, freezes and failed tool execution ## All 63 criteria Criterion love: 0.5 is the category norm. n: rated author-weeks. ### Paying and limits: Typical (customer love 0.515, n 1013) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | Better than peers | 0.573 | 0.540–0.604 | 389 | 190 | 199 | | [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md) | Typical | 0.465 | 0.419–0.508 | 348 | 47 | 301 | | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | Worse than peers | 0.439 | 0.407–0.481 | 85 | 6 | 79 | | [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | Typical | 0.473 | 0.433–0.516 | 80 | 8 | 72 | | [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | Worse than peers | 0.435 | 0.403–0.477 | 76 | 1 | 75 | | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | Worse than peers | 0.432 | 0.405–0.456 | 71 | 8 | 63 | | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | Better than peers | 0.545 | 0.519–0.571 | 63 | 46 | 17 | | [Usage meter visibility and accuracy](https://feedbackbench.com/criteria/limits.usage_meter.md) | Typical | 0.527 | 0.477–0.585 | 57 | 8 | 49 | | [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md) | Typical | 0.535 | 0.482–0.587 | 35 | 5 | 30 | | [Pay-as-you-go overage, fallback billing and spend caps](https://feedbackbench.com/criteria/billing.overage_charges.md) | Too few posts | 0.495 | 0.489–0.499 | 4 | 0 | 4 | | [Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md) | Too few posts | 0.499 | 0.489–0.511 | 3 | 1 | 2 | ### Setting up and connecting: Worse than peers (customer love 0.389, n 524) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [IDE and editor integration](https://feedbackbench.com/criteria/setup.ide_integration.md) | Worse than peers | 0.435 | 0.407–0.465 | 166 | 53 | 113 | | [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | Worse than peers | 0.423 | 0.388–0.459 | 161 | 16 | 145 | | [MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md) | Worse than peers | 0.440 | 0.410–0.472 | 120 | 41 | 79 | | [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | Typical | 0.478 | 0.450–0.507 | 66 | 29 | 37 | | [Onboarding, discoverability and documentation](https://feedbackbench.com/criteria/setup.onboarding_docs.md) | Typical | 0.486 | 0.460–0.516 | 36 | 5 | 31 | ### Choosing models: Typical (customer love 0.506, n 651) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md) | Better than peers | 0.589 | 0.557–0.623 | 336 | 115 | 221 | | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | Worse than peers | 0.364 | 0.334–0.398 | 249 | 27 | 222 | | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | Typical | 0.519 | 0.488–0.553 | 52 | 18 | 34 | | [Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md) | Typical | 0.523 | 0.497–0.548 | 49 | 28 | 21 | ### Instructing and context: Worse than peers (customer love 0.421, n 356) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md) | Typical | 0.484 | 0.449–0.518 | 112 | 27 | 85 | | [Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md) | Typical | 0.487 | 0.451–0.527 | 62 | 8 | 54 | | [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md) | Typical | 0.482 | 0.454–0.509 | 56 | 25 | 31 | | [Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md) | Typical | 0.477 | 0.451–0.503 | 47 | 15 | 32 | | [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md) | Worse than peers | 0.453 | 0.428–0.479 | 42 | 9 | 33 | | [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | Worse than peers | 0.467 | 0.444–0.489 | 34 | 5 | 29 | | [Images, PDFs and file attachments as input](https://feedbackbench.com/criteria/context.attachments.md) | Worse than peers | 0.467 | 0.445–0.489 | 34 | 6 | 28 | | [Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md) | Too few posts | 0.508 | 0.491–0.525 | 14 | 7 | 7 | ### Doing the work: Worse than peers (customer love 0.378, n 1723) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md) | Worse than peers | 0.371 | 0.347–0.394 | 954 | 437 | 517 | | [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | Worse than peers | 0.405 | 0.374–0.437 | 186 | 21 | 165 | | [Plan-before-edit mode](https://feedbackbench.com/criteria/work.plan_mode.md) | Typical | 0.477 | 0.456–0.500 | 169 | 61 | 108 | | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | Typical | 0.480 | 0.445–0.519 | 137 | 77 | 60 | | [Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md) | Typical | 0.453 | 0.389–0.514 | 128 | 4 | 124 | | [Risky or irreversible actions without confirmation](https://feedbackbench.com/criteria/work.destructive_actions.md) | Typical | 0.502 | 0.463–0.538 | 80 | 15 | 65 | | [Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md) | Typical | 0.528 | 0.465–0.584 | 61 | 6 | 55 | | [Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md) | Typical | 0.507 | 0.459–0.558 | 53 | 6 | 47 | | [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | Worse than peers | 0.461 | 0.435–0.486 | 47 | 16 | 31 | | [Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md) | Typical | 0.508 | 0.482–0.536 | 44 | 22 | 22 | | [Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md) | Typical | 0.480 | 0.452–0.509 | 31 | 20 | 11 | | [Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md) | Too few posts | 0.475 | 0.450–0.497 | 24 | 11 | 13 | | [Safety filters block legitimate coding tasks](https://feedbackbench.com/criteria/work.safety_refusals.md) | Too few posts | 0.504 | 0.472–0.543 | 22 | 3 | 19 | | [Games checks instead of fixing the problem](https://feedbackbench.com/criteria/work.reward_hacking.md) | Too few posts | 0.515 | 0.470–0.561 | 20 | 1 | 19 | | [Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md) | Too few posts | 0.547 | 0.516–0.577 | 18 | 12 | 6 | | [Git commits, branches and sync](https://feedbackbench.com/criteria/work.git_workflow.md) | Too few posts | 0.499 | 0.480–0.518 | 17 | 6 | 11 | | [Caves to or argues with the user's judgement](https://feedbackbench.com/criteria/work.sycophancy_pushback.md) | Too few posts | 0.508 | 0.485–0.535 | 13 | 3 | 10 | | [Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md) | Too few posts | 0.484 | 0.474–0.493 | 11 | 0 | 11 | ### Checking and finishing: Worse than peers (customer love 0.395, n 109) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Claims work is done or fixed when it is not](https://feedbackbench.com/criteria/verify.false_completion.md) | Worse than peers | 0.455 | 0.425–0.496 | 57 | 1 | 56 | | [Reviewing and approving the agent's changes](https://feedbackbench.com/criteria/verify.change_review_ui.md) | Too few posts | 0.487 | 0.466–0.509 | 25 | 6 | 19 | | [Builds, tests or runs its own changes](https://feedbackbench.com/criteria/verify.self_testing.md) | Too few posts | 0.493 | 0.474–0.511 | 19 | 8 | 11 | | [Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md) | Too few posts | 0.465 | 0.443–0.487 | 11 | 2 | 9 | ### Interface and sessions: Worse than peers (customer love 0.430, n 324) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | Worse than peers | 0.439 | 0.404–0.472 | 178 | 41 | 137 | | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | Typical | 0.485 | 0.456–0.515 | 85 | 41 | 44 | | [Saving, switching, resuming and rewinding sessions](https://feedbackbench.com/criteria/ui.session_history.md) | Worse than peers | 0.462 | 0.439–0.484 | 43 | 5 | 38 | | [Stopping and steering a running agent](https://feedbackbench.com/criteria/ui.interrupt_steer.md) | Too few posts | 0.506 | 0.487–0.528 | 20 | 9 | 11 | | [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | Too few posts | 0.498 | 0.481–0.514 | 15 | 8 | 7 | ### Reliability and speed: Better than peers (customer love 0.537, n 906) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Latency, throughput and fast mode](https://feedbackbench.com/criteria/rel.response_speed.md) | Typical | 0.482 | 0.453–0.510 | 525 | 181 | 344 | | [Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md) | Worse than peers | 0.434 | 0.368–0.498 | 228 | 10 | 218 | | [Outages, server errors and capacity or rate errors](https://feedbackbench.com/criteria/rel.service_errors.md) | Worse than peers | 0.421 | 0.366–0.475 | 157 | 5 | 152 | | [Updates break working setups](https://feedbackbench.com/criteria/rel.update_breakage.md) | Typical | 0.514 | 0.470–0.553 | 77 | 14 | 63 | ### Account and support: Typical (customer love 0.488, n 290) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Account bans and access restrictions](https://feedbackbench.com/criteria/account.bans_restrictions.md) | Typical | 0.454 | 0.391–0.515 | 143 | 5 | 138 | | [Support, refunds and issue handling](https://feedbackbench.com/criteria/account.support.md) | Typical | 0.510 | 0.470–0.552 | 93 | 17 | 76 | | [Data retention, training use and deployment isolation](https://feedbackbench.com/criteria/account.data_privacy.md) | Typical | 0.515 | 0.479–0.549 | 60 | 15 | 45 | | [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | Too few posts | 0.488 | 0.479–0.494 | 10 | 0 | 10 | # Pi (Earendil Works (open source)) Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/agent/pi | Measure | Value | |---|---| | Rank | 7 of 17 (rank range 7–7) | | Feedback Score | 52.7 (95% interval 51.9–53.3) | | Popularity | 0.456 (share of voice 2.77%) | | Customer love | 0.608 (95% interval 0.591–0.623) | | Top quadrant | no | | Authors | 2732 | | Posts counted | 6213 | | Posts that judge the agent | 2146 | | Criteria better / worse than peers | 14 / 0 of 63 | ## Top requests What users ask to add or change, most asked first. 332 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first. | Rank | Request | Criterion | Author-weeks | Posts | |---|---|---|---|---| | 1 | Bring existing subscription into this agent | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | 17 | 17 | | 2 | Better compaction summary quality and retention | [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | 6 | 6 | | 3 | Built-in multi-agent orchestrator mode | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | 6 | 6 | | 4 | Dedicated desktop app | [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | 5 | 5 | | 5 | Live dashboard of subagent status and progress | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | 5 | 5 | | 6 | Multi-provider model choice in one harness | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 5 | 5 | | 7 | Rewind to checkpoint with code restore | [Saving, switching, resuming and rewinding sessions](https://feedbackbench.com/criteria/ui.session_history.md) | 5 | 5 | | 8 | Hosted cloud agent execution support | [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | 4 | 6 | | 9 | Local model support | [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | 4 | 4 | | 10 | Richer extensibility API | [MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md) | 4 | 4 | | 11 | Route only at first turn or manual trigger | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | 3 | 6 | | 12 | Agent teams with assignable roles | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | 3 | 3 | ## Facts | Fact | Value | |---|---| | Version | v0.84.x (Aug 2026); npm @earendil-works/pi-coding-agent | | Released | First release: 2025-08. Joined Earendil: 2026-04-08 | | Price | Free (MIT). Model usage billed by the chosen provider's API or subscription | | Model | Multi-provider, bring-your-own-key or subscription | | Surface | CLI (terminal) | ## Sources | Channel | Source | Posts | |---|---|---| | Reddit | r/PiCodingAgent | 4114 | | X | @pidotdev | 1973 | | Reddit | Posts that name it | 126 | ## Better than peers on How much use a plan's price buys, Single prompt, model or effort level consumes disproportionate quota, Prompt cache hits, misses and invalidation, Using an existing subscription across tools, Connecting own API keys, local models and custom endpoints, MCP servers, plugins, skills and hooks, Onboarding, discoverability and documentation, Which models are offered on a plan and when, Context compaction keeps what matters, cheaply and quickly, Can do the user's kind of task, How the interface shows work, and what the user can configure, Saving, switching, resuming and rewinding sessions, Latency, throughput and fast mode, Client crashes, freezes and failed tool execution ## Worse than peers on None. ## All 63 criteria Criterion love: 0.5 is the category norm. n: rated author-weeks. ### Paying and limits: Better than peers (customer love 0.686, n 245) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md) | Better than peers | 0.646 | 0.605–0.681 | 84 | 43 | 41 | | [Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md) | Better than peers | 0.536 | 0.510–0.565 | 61 | 39 | 22 | | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | Better than peers | 0.535 | 0.509–0.558 | 45 | 26 | 19 | | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | Better than peers | 0.567 | 0.539–0.593 | 42 | 32 | 10 | | [Usage meter visibility and accuracy](https://feedbackbench.com/criteria/limits.usage_meter.md) | Too few posts | 0.558 | 0.515–0.606 | 11 | 6 | 5 | | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | Too few posts | 0.504 | 0.491–0.517 | 8 | 5 | 3 | | [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | Too few posts | 0.491 | 0.484–0.497 | 6 | 0 | 6 | | [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md) | Too few posts | 0.495 | 0.489–0.499 | 4 | 0 | 4 | | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | Too few posts | 0.495 | 0.490–0.500 | 3 | 0 | 3 | | [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Pay-as-you-go overage, fallback billing and spend caps](https://feedbackbench.com/criteria/billing.overage_charges.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Setting up and connecting: Better than peers (customer love 0.632, n 307) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md) | Better than peers | 0.586 | 0.558–0.619 | 201 | 136 | 65 | | [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | Better than peers | 0.532 | 0.505–0.557 | 60 | 41 | 19 | | [Onboarding, discoverability and documentation](https://feedbackbench.com/criteria/setup.onboarding_docs.md) | Better than peers | 0.549 | 0.513–0.583 | 36 | 15 | 21 | | [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | Too few posts | 0.512 | 0.486–0.540 | 21 | 6 | 15 | | [IDE and editor integration](https://feedbackbench.com/criteria/setup.ide_integration.md) | Too few posts | 0.512 | 0.499–0.526 | 8 | 6 | 2 | ### Choosing models: Better than peers (customer love 0.581, n 79) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | Better than peers | 0.556 | 0.523–0.586 | 36 | 22 | 14 | | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | Too few posts | 0.547 | 0.517–0.579 | 29 | 16 | 13 | | [Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md) | Too few posts | 0.502 | 0.488–0.518 | 10 | 5 | 5 | | [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md) | Too few posts | 0.490 | 0.482–0.497 | 6 | 0 | 6 | ### Instructing and context: Typical (customer love 0.529, n 179) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | Better than peers | 0.536 | 0.504–0.567 | 73 | 34 | 39 | | [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md) | Typical | 0.508 | 0.482–0.535 | 38 | 20 | 18 | | [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md) | Too few posts | 0.501 | 0.482–0.519 | 19 | 10 | 9 | | [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md) | Too few posts | 0.496 | 0.476–0.518 | 17 | 4 | 13 | | [Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md) | Too few posts | 0.494 | 0.476–0.511 | 17 | 6 | 11 | | [Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md) | Too few posts | 0.497 | 0.478–0.523 | 15 | 2 | 13 | | [Images, PDFs and file attachments as input](https://feedbackbench.com/criteria/context.attachments.md) | Too few posts | 0.502 | 0.490–0.516 | 9 | 4 | 5 | | [Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md) | Too few posts | 0.505 | 0.490–0.520 | 8 | 4 | 4 | ### Doing the work: Better than peers (customer love 0.554, n 328) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md) | Better than peers | 0.544 | 0.511–0.578 | 163 | 116 | 47 | | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | Typical | 0.488 | 0.455–0.520 | 88 | 50 | 38 | | [Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md) | Too few posts | 0.567 | 0.499–0.626 | 26 | 5 | 21 | | [Risky or irreversible actions without confirmation](https://feedbackbench.com/criteria/work.destructive_actions.md) | Too few posts | 0.516 | 0.488–0.544 | 20 | 6 | 14 | | [Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md) | Too few posts | 0.508 | 0.488–0.525 | 14 | 12 | 2 | | [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | Too few posts | 0.540 | 0.515–0.567 | 13 | 10 | 3 | | [Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md) | Too few posts | 0.484 | 0.475–0.492 | 12 | 0 | 12 | | [Plan-before-edit mode](https://feedbackbench.com/criteria/work.plan_mode.md) | Too few posts | 0.498 | 0.481–0.513 | 12 | 5 | 7 | | [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | Too few posts | 0.491 | 0.477–0.502 | 7 | 2 | 5 | | [Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md) | Too few posts | 0.500 | 0.489–0.514 | 5 | 1 | 4 | | [Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md) | Too few posts | 0.503 | 0.496–0.511 | 4 | 3 | 1 | | [Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md) | Too few posts | 0.504 | 0.493–0.519 | 4 | 2 | 2 | | [Safety filters block legitimate coding tasks](https://feedbackbench.com/criteria/work.safety_refusals.md) | Too few posts | 0.496 | 0.490–0.500 | 3 | 0 | 3 | | [Git commits, branches and sync](https://feedbackbench.com/criteria/work.git_workflow.md) | Too few posts | 0.496 | 0.491–0.500 | 2 | 0 | 2 | | [Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md) | Too few posts | 0.502 | 0.500–0.506 | 1 | 1 | 0 | | [Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md) | Too few posts | 0.498 | 0.494–0.500 | 1 | 0 | 1 | | [Games checks instead of fixing the problem](https://feedbackbench.com/criteria/work.reward_hacking.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Caves to or argues with the user's judgement](https://feedbackbench.com/criteria/work.sycophancy_pushback.md) | Too few posts | 0.508 | 0.500–0.524 | 1 | 1 | 0 | ### Checking and finishing: Too few posts (customer love 0.520, n 15) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md) | Too few posts | 0.512 | 0.504–0.521 | 7 | 7 | 0 | | [Reviewing and approving the agent's changes](https://feedbackbench.com/criteria/verify.change_review_ui.md) | Too few posts | 0.497 | 0.489–0.506 | 4 | 1 | 3 | | [Claims work is done or fixed when it is not](https://feedbackbench.com/criteria/verify.false_completion.md) | Too few posts | 0.516 | 0.496–0.557 | 2 | 1 | 1 | | [Builds, tests or runs its own changes](https://feedbackbench.com/criteria/verify.self_testing.md) | Too few posts | 0.505 | 0.500–0.513 | 2 | 2 | 0 | ### Interface and sessions: Better than peers (customer love 0.564, n 171) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | Better than peers | 0.553 | 0.517–0.589 | 115 | 55 | 60 | | [Saving, switching, resuming and rewinding sessions](https://feedbackbench.com/criteria/ui.session_history.md) | Better than peers | 0.584 | 0.556–0.613 | 35 | 28 | 7 | | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | Too few posts | 0.506 | 0.489–0.522 | 15 | 9 | 6 | | [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | Too few posts | 0.498 | 0.482–0.510 | 7 | 4 | 3 | | [Stopping and steering a running agent](https://feedbackbench.com/criteria/ui.interrupt_steer.md) | Too few posts | 0.503 | 0.494–0.513 | 5 | 3 | 2 | ### Reliability and speed: Better than peers (customer love 0.592, n 119) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Latency, throughput and fast mode](https://feedbackbench.com/criteria/rel.response_speed.md) | Better than peers | 0.542 | 0.514–0.568 | 54 | 31 | 23 | | [Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md) | Better than peers | 0.580 | 0.519–0.633 | 46 | 10 | 36 | | [Updates break working setups](https://feedbackbench.com/criteria/rel.update_breakage.md) | Too few posts | 0.516 | 0.490–0.546 | 14 | 4 | 10 | | [Outages, server errors and capacity or rate errors](https://feedbackbench.com/criteria/rel.service_errors.md) | Too few posts | 0.521 | 0.484–0.566 | 11 | 2 | 9 | ### Account and support: Typical (customer love 0.512, n 30) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Account bans and access restrictions](https://feedbackbench.com/criteria/account.bans_restrictions.md) | Too few posts | 0.508 | 0.474–0.556 | 20 | 2 | 18 | | [Data retention, training use and deployment isolation](https://feedbackbench.com/criteria/account.data_privacy.md) | Too few posts | 0.508 | 0.490–0.529 | 7 | 2 | 5 | | [Support, refunds and issue handling](https://feedbackbench.com/criteria/account.support.md) | Too few posts | 0.505 | 0.494–0.521 | 3 | 1 | 2 | | [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | # GitHub Copilot (GitHub) Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/agent/copilot | Measure | Value | |---|---| | Rank | =8 of 17 (rank range 8–10) | | Feedback Score | 43.0 (95% interval 42.2–43.7) | | Popularity | 0.348 (share of voice 1.59%) | | Customer love | 0.532 (95% interval 0.513–0.551) | | Top quadrant | no | | Authors | 1569 | | Posts counted | 3069 | | Posts that judge the agent | 1349 | | Criteria better / worse than peers | 1 / 1 of 63 | ## Top requests What users ask to add or change, most asked first. 135 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first. | Rank | Request | Criterion | Author-weeks | Posts | |---|---|---|---|---| | 1 | Exclude specific models from auto routing | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | 3 | 5 | | 2 | Add low-cost DeepSeek and GLM models together | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 3 | 4 | | 3 | Higher overall usage limits | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | 3 | 3 | | 4 | Risk-tiered granular approval policies | [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | 3 | 3 | | 5 | Access to editor extension tools | [MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md) | 2 | 5 | | 6 | Bring existing subscription into this agent | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | 2 | 4 | | 7 | Add DeepSeek V4.1 Flash model | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 2 | 2 | | 8 | Codebase visualization and documentation tools | [Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md) | 2 | 2 | | 9 | Manual model selection instead of forced routing | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | 2 | 2 | | 10 | No silent model switching or downgrades | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | 2 | 2 | | 11 | Option to restore previous UI design | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | 2 | 2 | | 12 | Per-subagent model and effort selection | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | 2 | 2 | ## Facts | Fact | Value | |---|---| | Version | Multi-model agent mode; usage-based AI Credits billing since 2026-06-01 | | Released | Usage-based billing: 2026-06-01 | | Price | Free, Pro $10/user/mo ($15 credits), Pro+ $39/user/mo ($70 credits), Max $100/user/mo ($200 credits); Business $19/seat/mo, Enterprise $39/seat/mo | | Model | Multi-model router including Claude Opus and GPT models | | Surface | IDE (VS Code, JetBrains), GitHub.com | ## Sources | Channel | Source | Posts | |---|---|---| | Reddit | r/GithubCopilot | 2399 | | Reddit | Posts that name it | 498 | | X | @GitHubCopilot | 172 | ## Better than peers on Which models are offered on a plan and when ## Worse than peers on Can do the user's kind of task ## All 63 criteria Criterion love: 0.5 is the category norm. n: rated author-weeks. ### Paying and limits: Better than peers (customer love 0.568, n 387) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | Typical | 0.521 | 0.483–0.558 | 220 | 91 | 129 | | [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md) | Typical | 0.530 | 0.479–0.575 | 106 | 22 | 84 | | [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | Typical | 0.513 | 0.452–0.575 | 43 | 3 | 40 | | [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md) | Too few posts | 0.558 | 0.496–0.622 | 21 | 4 | 17 | | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | Too few posts | 0.501 | 0.484–0.525 | 11 | 2 | 9 | | [Usage meter visibility and accuracy](https://feedbackbench.com/criteria/limits.usage_meter.md) | Too few posts | 0.526 | 0.485–0.575 | 11 | 2 | 9 | | [Pay-as-you-go overage, fallback billing and spend caps](https://feedbackbench.com/criteria/billing.overage_charges.md) | Too few posts | 0.491 | 0.483–0.497 | 7 | 0 | 7 | | [Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md) | Too few posts | 0.506 | 0.492–0.524 | 6 | 3 | 3 | | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | Too few posts | 0.500 | 0.487–0.512 | 6 | 3 | 3 | | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | Too few posts | 0.502 | 0.490–0.513 | 6 | 3 | 3 | | [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | Too few posts | 0.504 | 0.490–0.525 | 5 | 1 | 4 | ### Setting up and connecting: Typical (customer love 0.507, n 106) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md) | Typical | 0.477 | 0.451–0.502 | 38 | 14 | 24 | | [IDE and editor integration](https://feedbackbench.com/criteria/setup.ide_integration.md) | Typical | 0.504 | 0.478–0.530 | 38 | 19 | 19 | | [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | Too few posts | 0.498 | 0.476–0.522 | 24 | 13 | 11 | | [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | Too few posts | 0.513 | 0.495–0.534 | 6 | 3 | 3 | | [Onboarding, discoverability and documentation](https://feedbackbench.com/criteria/setup.onboarding_docs.md) | Too few posts | 0.492 | 0.485–0.498 | 5 | 0 | 5 | ### Choosing models: Better than peers (customer love 0.548, n 147) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | Better than peers | 0.559 | 0.524–0.591 | 63 | 29 | 34 | | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | Typical | 0.508 | 0.473–0.543 | 56 | 14 | 42 | | [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md) | Too few posts | 0.497 | 0.473–0.525 | 26 | 5 | 21 | | [Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md) | Too few posts | 0.487 | 0.469–0.505 | 17 | 5 | 12 | ### Instructing and context: Typical (customer love 0.513, n 53) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md) | Too few posts | 0.503 | 0.485–0.522 | 16 | 5 | 11 | | [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md) | Too few posts | 0.508 | 0.493–0.524 | 12 | 8 | 4 | | [Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md) | Too few posts | 0.496 | 0.478–0.513 | 12 | 4 | 8 | | [Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md) | Too few posts | 0.510 | 0.491–0.531 | 6 | 2 | 4 | | [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | Too few posts | 0.498 | 0.490–0.509 | 4 | 1 | 3 | | [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md) | Too few posts | 0.506 | 0.496–0.518 | 4 | 3 | 1 | | [Images, PDFs and file attachments as input](https://feedbackbench.com/criteria/context.attachments.md) | Too few posts | 0.501 | 0.494–0.509 | 2 | 1 | 1 | | [Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Doing the work: Typical (customer love 0.495, n 244) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md) | Worse than peers | 0.457 | 0.423–0.494 | 158 | 78 | 80 | | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | Too few posts | 0.500 | 0.479–0.525 | 23 | 14 | 9 | | [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | Too few posts | 0.501 | 0.484–0.521 | 11 | 3 | 8 | | [Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md) | Too few posts | 0.526 | 0.487–0.562 | 9 | 2 | 7 | | [Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md) | Too few posts | 0.488 | 0.481–0.496 | 9 | 0 | 9 | | [Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md) | Too few posts | 0.500 | 0.486–0.513 | 7 | 5 | 2 | | [Plan-before-edit mode](https://feedbackbench.com/criteria/work.plan_mode.md) | Too few posts | 0.492 | 0.478–0.506 | 7 | 2 | 5 | | [Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md) | Too few posts | 0.497 | 0.485–0.514 | 7 | 1 | 6 | | [Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md) | Too few posts | 0.492 | 0.486–0.497 | 6 | 0 | 6 | | [Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md) | Too few posts | 0.504 | 0.491–0.516 | 6 | 5 | 1 | | [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | Too few posts | 0.509 | 0.499–0.521 | 6 | 5 | 1 | | [Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md) | Too few posts | 0.495 | 0.489–0.499 | 4 | 0 | 4 | | [Risky or irreversible actions without confirmation](https://feedbackbench.com/criteria/work.destructive_actions.md) | Too few posts | 0.502 | 0.491–0.520 | 4 | 1 | 3 | | [Git commits, branches and sync](https://feedbackbench.com/criteria/work.git_workflow.md) | Too few posts | 0.499 | 0.491–0.508 | 3 | 1 | 2 | | [Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md) | Too few posts | 0.501 | 0.494–0.509 | 2 | 1 | 1 | | [Caves to or argues with the user's judgement](https://feedbackbench.com/criteria/work.sycophancy_pushback.md) | Too few posts | 0.504 | 0.496–0.512 | 2 | 1 | 1 | | [Games checks instead of fixing the problem](https://feedbackbench.com/criteria/work.reward_hacking.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Safety filters block legitimate coding tasks](https://feedbackbench.com/criteria/work.safety_refusals.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Checking and finishing: Typical (customer love 0.509, n 33) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md) | Too few posts | 0.487 | 0.464–0.510 | 17 | 10 | 7 | | [Reviewing and approving the agent's changes](https://feedbackbench.com/criteria/verify.change_review_ui.md) | Too few posts | 0.514 | 0.498–0.532 | 10 | 6 | 4 | | [Claims work is done or fixed when it is not](https://feedbackbench.com/criteria/verify.false_completion.md) | Too few posts | 0.495 | 0.490–0.499 | 4 | 0 | 4 | | [Builds, tests or runs its own changes](https://feedbackbench.com/criteria/verify.self_testing.md) | Too few posts | 0.497 | 0.490–0.504 | 3 | 1 | 2 | ### Interface and sessions: Typical (customer love 0.481, n 40) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | Typical | 0.477 | 0.454–0.502 | 30 | 5 | 25 | | [Saving, switching, resuming and rewinding sessions](https://feedbackbench.com/criteria/ui.session_history.md) | Too few posts | 0.506 | 0.490–0.522 | 11 | 4 | 7 | | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | Too few posts | 0.501 | 0.494–0.509 | 3 | 2 | 1 | | [Stopping and steering a running agent](https://feedbackbench.com/criteria/ui.interrupt_steer.md) | Too few posts | 0.503 | 0.500–0.508 | 1 | 1 | 0 | | [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | Too few posts | 0.502 | 0.500–0.507 | 1 | 1 | 0 | ### Reliability and speed: Typical (customer love 0.539, n 68) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md) | Typical | 0.507 | 0.465–0.551 | 30 | 3 | 27 | | [Latency, throughput and fast mode](https://feedbackbench.com/criteria/rel.response_speed.md) | Too few posts | 0.513 | 0.492–0.537 | 22 | 10 | 12 | | [Updates break working setups](https://feedbackbench.com/criteria/rel.update_breakage.md) | Too few posts | 0.529 | 0.501–0.560 | 11 | 5 | 6 | | [Outages, server errors and capacity or rate errors](https://feedbackbench.com/criteria/rel.service_errors.md) | Too few posts | 0.535 | 0.491–0.581 | 10 | 3 | 7 | ### Account and support: Typical (customer love 0.523, n 38) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Data retention, training use and deployment isolation](https://feedbackbench.com/criteria/account.data_privacy.md) | Too few posts | 0.523 | 0.494–0.560 | 16 | 5 | 11 | | [Support, refunds and issue handling](https://feedbackbench.com/criteria/account.support.md) | Too few posts | 0.497 | 0.477–0.521 | 15 | 2 | 13 | | [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | Too few posts | 0.486 | 0.477–0.495 | 11 | 0 | 11 | | [Account bans and access restrictions](https://feedbackbench.com/criteria/account.bans_restrictions.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | # Cline (Cline (open source)) Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/agent/cline | Measure | Value | |---|---| | Rank | =8 of 17 (rank range 8–10) | | Feedback Score | 42.9 (95% interval 42.2–43.6) | | Popularity | 0.333 (share of voice 1.47%) | | Customer love | 0.553 (95% interval 0.535–0.570) | | Top quadrant | no | | Authors | 1449 | | Posts counted | 2308 | | Posts that judge the agent | 1187 | | Criteria better / worse than peers | 3 / 0 of 63 | ## Top requests What users ask to add or change, most asked first. 295 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first. | Rank | Request | Criterion | Author-weeks | Posts | |---|---|---|---|---| | 1 | Linux desktop app and support | [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | 53 | 61 | | 2 | Higher free tier usage limits | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | 8 | 9 | | 3 | Newest models on lower-priced plans | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 6 | 6 | | 4 | Stabilize and fix the desktop app | [Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md) | 4 | 4 | | 5 | Faster, more responsive support replies | [Support, refunds and issue handling](https://feedbackbench.com/criteria/account.support.md) | 3 | 4 | | 6 | Same-day availability of new models | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 3 | 4 | | 7 | Add DeepSeek V4.1 Flash model | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 3 | 3 | | 8 | Built-in computer use capability | [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | 3 | 3 | | 9 | Clear free tier usage limits | [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md) | 3 | 3 | | 10 | Fix frequent app and CLI crashes | [Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md) | 3 | 3 | | 11 | Free access to specific or new models | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | 3 | 3 | | 12 | Free usage credits | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | 3 | 3 | ## Facts | Fact | Value | |---|---| | Version | n/a | | Released | First released 2024; passed 1.5M VS Code Marketplace installs by 2026-04 | | Price | Free (open source); BYOK API spend typically $8-150/mo depending on model/usage; Teams free through Q1 2026, then $20/user/mo (first 10 seats always free) | | Model | Bring-your-own-key, any provider | | Surface | VS Code extension, CLI | ## Sources | Channel | Source | Posts | |---|---|---| | X | @cline | 1816 | | Reddit | r/CLine | 374 | | Reddit | Posts that name it | 117 | | Trustpilot | Trustpilot | 1 | ## Better than peers on Connecting own API keys, local models and custom endpoints, Which models are offered on a plan and when, How the interface shows work, and what the user can configure ## Worse than peers on None. ## All 63 criteria Criterion love: 0.5 is the category norm. n: rated author-weeks. ### Paying and limits: Better than peers (customer love 0.603, n 266) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | Typical | 0.498 | 0.465–0.531 | 133 | 80 | 53 | | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | Typical | 0.499 | 0.469–0.532 | 71 | 27 | 44 | | [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md) | Too few posts | 0.494 | 0.469–0.519 | 22 | 3 | 19 | | [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md) | Too few posts | 0.477 | 0.467–0.487 | 18 | 0 | 18 | | [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | Too few posts | 0.484 | 0.469–0.505 | 17 | 1 | 16 | | [Usage meter visibility and accuracy](https://feedbackbench.com/criteria/limits.usage_meter.md) | Too few posts | 0.500 | 0.485–0.523 | 8 | 1 | 7 | | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | Too few posts | 0.521 | 0.505–0.540 | 8 | 6 | 2 | | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | Too few posts | 0.491 | 0.482–0.497 | 6 | 0 | 6 | | [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | Too few posts | 0.511 | 0.492–0.539 | 4 | 1 | 3 | | [Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md) | Too few posts | 0.503 | 0.495–0.514 | 4 | 2 | 2 | | [Pay-as-you-go overage, fallback billing and spend caps](https://feedbackbench.com/criteria/billing.overage_charges.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | ### Setting up and connecting: Typical (customer love 0.528, n 136) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | Typical | 0.498 | 0.463–0.531 | 60 | 11 | 49 | | [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | Better than peers | 0.556 | 0.532–0.580 | 48 | 39 | 9 | | [MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md) | Too few posts | 0.487 | 0.469–0.503 | 16 | 5 | 11 | | [Onboarding, discoverability and documentation](https://feedbackbench.com/criteria/setup.onboarding_docs.md) | Too few posts | 0.501 | 0.484–0.522 | 10 | 2 | 8 | | [IDE and editor integration](https://feedbackbench.com/criteria/setup.ide_integration.md) | Too few posts | 0.512 | 0.497–0.527 | 9 | 6 | 3 | ### Choosing models: Better than peers (customer love 0.541, n 85) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | Better than peers | 0.539 | 0.509–0.569 | 45 | 22 | 23 | | [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md) | Too few posts | 0.510 | 0.489–0.532 | 19 | 7 | 12 | | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | Too few posts | 0.515 | 0.495–0.536 | 15 | 7 | 8 | | [Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md) | Too few posts | 0.479 | 0.464–0.493 | 12 | 1 | 11 | ### Instructing and context: Typical (customer love 0.485, n 41) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md) | Too few posts | 0.514 | 0.491–0.538 | 9 | 3 | 6 | | [Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md) | Too few posts | 0.494 | 0.480–0.509 | 9 | 3 | 6 | | [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | Too few posts | 0.490 | 0.479–0.502 | 8 | 1 | 7 | | [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md) | Too few posts | 0.501 | 0.490–0.511 | 5 | 3 | 2 | | [Images, PDFs and file attachments as input](https://feedbackbench.com/criteria/context.attachments.md) | Too few posts | 0.492 | 0.484–0.498 | 4 | 0 | 4 | | [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md) | Too few posts | 0.500 | 0.493–0.511 | 3 | 1 | 2 | | [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md) | Too few posts | 0.505 | 0.500–0.512 | 2 | 2 | 0 | | [Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md) | Too few posts | 0.498 | 0.493–0.500 | 1 | 0 | 1 | ### Doing the work: Typical (customer love 0.510, n 140) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md) | Typical | 0.482 | 0.452–0.514 | 71 | 42 | 29 | | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | Too few posts | 0.514 | 0.496–0.531 | 18 | 14 | 4 | | [Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md) | Too few posts | 0.504 | 0.477–0.553 | 13 | 1 | 12 | | [Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md) | Too few posts | 0.504 | 0.487–0.520 | 12 | 10 | 2 | | [Plan-before-edit mode](https://feedbackbench.com/criteria/work.plan_mode.md) | Too few posts | 0.491 | 0.479–0.502 | 8 | 2 | 6 | | [Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md) | Too few posts | 0.531 | 0.512–0.550 | 7 | 7 | 0 | | [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | Too few posts | 0.505 | 0.493–0.517 | 7 | 5 | 2 | | [Risky or irreversible actions without confirmation](https://feedbackbench.com/criteria/work.destructive_actions.md) | Too few posts | 0.494 | 0.488–0.499 | 4 | 0 | 4 | | [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | Too few posts | 0.500 | 0.491–0.512 | 4 | 1 | 3 | | [Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md) | Too few posts | 0.496 | 0.490–0.500 | 3 | 0 | 3 | | [Git commits, branches and sync](https://feedbackbench.com/criteria/work.git_workflow.md) | Too few posts | 0.505 | 0.496–0.516 | 3 | 2 | 1 | | [Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md) | Too few posts | 0.495 | 0.484–0.500 | 1 | 0 | 1 | | [Games checks instead of fixing the problem](https://feedbackbench.com/criteria/work.reward_hacking.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Safety filters block legitimate coding tasks](https://feedbackbench.com/criteria/work.safety_refusals.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md) | Too few posts | 0.506 | 0.500–0.520 | 1 | 1 | 0 | | [Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Caves to or argues with the user's judgement](https://feedbackbench.com/criteria/work.sycophancy_pushback.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Checking and finishing: Too few posts (customer love 0.495, n 6) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Builds, tests or runs its own changes](https://feedbackbench.com/criteria/verify.self_testing.md) | Too few posts | 0.498 | 0.489–0.505 | 2 | 1 | 1 | | [Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md) | Too few posts | 0.503 | 0.500–0.508 | 2 | 2 | 0 | | [Reviewing and approving the agent's changes](https://feedbackbench.com/criteria/verify.change_review_ui.md) | Too few posts | 0.495 | 0.488–0.500 | 2 | 0 | 2 | | [Claims work is done or fixed when it is not](https://feedbackbench.com/criteria/verify.false_completion.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | ### Interface and sessions: Better than peers (customer love 0.564, n 70) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | Better than peers | 0.570 | 0.543–0.595 | 47 | 34 | 13 | | [Saving, switching, resuming and rewinding sessions](https://feedbackbench.com/criteria/ui.session_history.md) | Too few posts | 0.502 | 0.486–0.520 | 12 | 4 | 8 | | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | Too few posts | 0.508 | 0.496–0.521 | 7 | 5 | 2 | | [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | Too few posts | 0.512 | 0.504–0.521 | 7 | 7 | 0 | | [Stopping and steering a running agent](https://feedbackbench.com/criteria/ui.interrupt_steer.md) | Too few posts | 0.500 | 0.493–0.507 | 2 | 1 | 1 | ### Reliability and speed: Better than peers (customer love 0.573, n 134) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Latency, throughput and fast mode](https://feedbackbench.com/criteria/rel.response_speed.md) | Typical | 0.518 | 0.491–0.544 | 59 | 30 | 29 | | [Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md) | Typical | 0.491 | 0.434–0.546 | 57 | 3 | 54 | | [Updates break working setups](https://feedbackbench.com/criteria/rel.update_breakage.md) | Too few posts | 0.552 | 0.520–0.586 | 17 | 9 | 8 | | [Outages, server errors and capacity or rate errors](https://feedbackbench.com/criteria/rel.service_errors.md) | Too few posts | 0.502 | 0.479–0.534 | 12 | 1 | 11 | ### Account and support: Typical (customer love 0.500, n 35) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Support, refunds and issue handling](https://feedbackbench.com/criteria/account.support.md) | Too few posts | 0.488 | 0.472–0.507 | 14 | 1 | 13 | | [Data retention, training use and deployment isolation](https://feedbackbench.com/criteria/account.data_privacy.md) | Too few posts | 0.503 | 0.487–0.521 | 12 | 4 | 8 | | [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | Too few posts | 0.488 | 0.480–0.495 | 10 | 0 | 10 | | [Account bans and access restrictions](https://feedbackbench.com/criteria/account.bans_restrictions.md) | Too few posts | 0.497 | 0.494–0.500 | 2 | 0 | 2 | # Zed (Zed Industries) Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/agent/zed | Measure | Value | |---|---| | Rank | =8 of 17 (rank range 8–10) | | Feedback Score | 42.4 (95% interval 41.7–43.0) | | Popularity | 0.362 (share of voice 1.72%) | | Customer love | 0.496 (95% interval 0.481–0.511) | | Top quadrant | no | | Authors | 1693 | | Posts counted | 2629 | | Posts that judge the agent | 1310 | | Criteria better / worse than peers | 2 / 2 of 63 | ## Top requests What users ask to add or change, most asked first. 489 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first. | Rank | Request | Criterion | Author-weeks | Posts | |---|---|---|---|---| | 1 | Bring existing subscription into this agent | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | 13 | 13 | | 2 | Agent Client Protocol (ACP) support | [MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md) | 8 | 10 | | 3 | Custom model provider support | [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | 8 | 9 | | 4 | Jupyter notebook support | [IDE and editor integration](https://feedbackbench.com/criteria/setup.ide_integration.md) | 8 | 8 | | 5 | Multi-provider model choice in one harness | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 8 | 8 | | 6 | SSH and remote machine development | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | 8 | 8 | | 7 | Selectively disable AI features | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | 8 | 8 | | 8 | Add DeepSeek V4.1 Flash model | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 6 | 6 | | 9 | Better markdown rendering | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | 6 | 6 | | 10 | Clearer delete vs permanent delete labels | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | 6 | 6 | | 11 | Customizable keybindings | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | 6 | 6 | | 12 | Detachable panels and multi-window support | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | 6 | 6 | ## Facts | Fact | Value | |---|---| | Version | n/a | | Released | n/a | | Price | Personal free (2,000 accepted edit predictions/mo, BYOK), Pro $10/mo (unlimited predictions + $5 tokens), Business $30/seat/mo | | Model | Claude Opus 4.8, Sonnet 5, Haiku 4.5; GPT-5.6, GPT-5.4; Gemini 3.1 Pro; local models via Ollama/LM Studio/llama.cpp | | Surface | Native desktop editor (built in Rust) | ## Sources | Channel | Source | Posts | |---|---|---| | X | @zeddotdev | 1544 | | Reddit | r/ZedEditor | 991 | | Reddit | Posts that name it | 94 | ## Better than peers on Latency, throughput and fast mode, Client crashes, freezes and failed tool execution ## Worse than peers on MCP servers, plugins, skills and hooks, Can do the user's kind of task ## All 63 criteria Criterion love: 0.5 is the category norm. n: rated author-weeks. ### Paying and limits: Typical (customer love 0.501, n 40) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | Too few posts | 0.488 | 0.472–0.505 | 14 | 3 | 11 | | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | Too few posts | 0.493 | 0.482–0.506 | 7 | 1 | 6 | | [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md) | Too few posts | 0.516 | 0.495–0.541 | 6 | 3 | 3 | | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | Too few posts | 0.506 | 0.497–0.515 | 5 | 4 | 1 | | [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md) | Too few posts | 0.495 | 0.489–0.499 | 4 | 0 | 4 | | [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | Too few posts | 0.497 | 0.493–0.500 | 2 | 0 | 2 | | [Usage meter visibility and accuracy](https://feedbackbench.com/criteria/limits.usage_meter.md) | Too few posts | 0.497 | 0.493–0.500 | 2 | 0 | 2 | | [Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md) | Too few posts | 0.497 | 0.490–0.500 | 1 | 0 | 1 | | [Pay-as-you-go overage, fallback billing and spend caps](https://feedbackbench.com/criteria/billing.overage_charges.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Setting up and connecting: Worse than peers (customer love 0.422, n 175) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [IDE and editor integration](https://feedbackbench.com/criteria/setup.ide_integration.md) | Typical | 0.473 | 0.447–0.500 | 63 | 21 | 42 | | [MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md) | Worse than peers | 0.426 | 0.400–0.453 | 58 | 10 | 48 | | [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | Too few posts | 0.487 | 0.466–0.512 | 25 | 3 | 22 | | [Onboarding, discoverability and documentation](https://feedbackbench.com/criteria/setup.onboarding_docs.md) | Too few posts | 0.479 | 0.463–0.496 | 20 | 1 | 19 | | [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | Too few posts | 0.495 | 0.476–0.513 | 19 | 9 | 10 | ### Choosing models: Too few posts (customer love 0.481, n 18) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | Too few posts | 0.480 | 0.465–0.494 | 15 | 1 | 14 | | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | Too few posts | 0.504 | 0.495–0.516 | 2 | 1 | 1 | | [Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md) | Too few posts | 0.498 | 0.493–0.500 | 1 | 0 | 1 | | [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Instructing and context: Too few posts (customer love 0.498, n 19) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md) | Too few posts | 0.498 | 0.487–0.509 | 7 | 3 | 4 | | [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md) | Too few posts | 0.490 | 0.477–0.501 | 6 | 1 | 5 | | [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | Too few posts | 0.494 | 0.488–0.500 | 3 | 0 | 3 | | [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md) | Too few posts | 0.505 | 0.500–0.514 | 2 | 2 | 0 | | [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md) | Too few posts | 0.504 | 0.500–0.513 | 1 | 1 | 0 | | [Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md) | Too few posts | 0.499 | 0.495–0.500 | 1 | 0 | 1 | | [Images, PDFs and file attachments as input](https://feedbackbench.com/criteria/context.attachments.md) | Too few posts | 0.498 | 0.493–0.500 | 1 | 0 | 1 | | [Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Doing the work: Worse than peers (customer love 0.455, n 102) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md) | Worse than peers | 0.440 | 0.411–0.472 | 52 | 19 | 33 | | [Git commits, branches and sync](https://feedbackbench.com/criteria/work.git_workflow.md) | Too few posts | 0.495 | 0.479–0.513 | 17 | 5 | 12 | | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | Too few posts | 0.498 | 0.480–0.515 | 12 | 7 | 5 | | [Risky or irreversible actions without confirmation](https://feedbackbench.com/criteria/work.destructive_actions.md) | Too few posts | 0.516 | 0.496–0.541 | 6 | 3 | 3 | | [Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md) | Too few posts | 0.503 | 0.496–0.511 | 4 | 3 | 1 | | [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | Too few posts | 0.495 | 0.489–0.500 | 3 | 0 | 3 | | [Plan-before-edit mode](https://feedbackbench.com/criteria/work.plan_mode.md) | Too few posts | 0.499 | 0.492–0.508 | 3 | 1 | 2 | | [Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md) | Too few posts | 0.490 | 0.474–0.500 | 2 | 0 | 2 | | [Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md) | Too few posts | 0.495 | 0.485–0.500 | 1 | 0 | 1 | | [Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | Too few posts | 0.497 | 0.492–0.500 | 1 | 0 | 1 | | [Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md) | Too few posts | 0.498 | 0.495–0.500 | 1 | 0 | 1 | | [Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Games checks instead of fixing the problem](https://feedbackbench.com/criteria/work.reward_hacking.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Safety filters block legitimate coding tasks](https://feedbackbench.com/criteria/work.safety_refusals.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Caves to or argues with the user's judgement](https://feedbackbench.com/criteria/work.sycophancy_pushback.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Checking and finishing: Typical (customer love 0.504, n 38) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Reviewing and approving the agent's changes](https://feedbackbench.com/criteria/verify.change_review_ui.md) | Typical | 0.517 | 0.494–0.542 | 37 | 20 | 17 | | [Claims work is done or fixed when it is not](https://feedbackbench.com/criteria/verify.false_completion.md) | Too few posts | 0.499 | 0.495–0.500 | 1 | 0 | 1 | | [Builds, tests or runs its own changes](https://feedbackbench.com/criteria/verify.self_testing.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Interface and sessions: Worse than peers (customer love 0.459, n 238) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | Typical | 0.505 | 0.470–0.539 | 218 | 79 | 139 | | [Saving, switching, resuming and rewinding sessions](https://feedbackbench.com/criteria/ui.session_history.md) | Too few posts | 0.500 | 0.481–0.520 | 17 | 5 | 12 | | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | Too few posts | 0.495 | 0.482–0.506 | 6 | 2 | 4 | | [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | Too few posts | 0.483 | 0.465–0.498 | 6 | 1 | 5 | | [Stopping and steering a running agent](https://feedbackbench.com/criteria/ui.interrupt_steer.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Reliability and speed: Better than peers (customer love 0.666, n 112) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Latency, throughput and fast mode](https://feedbackbench.com/criteria/rel.response_speed.md) | Better than peers | 0.603 | 0.576–0.629 | 59 | 51 | 8 | | [Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md) | Better than peers | 0.599 | 0.535–0.659 | 50 | 11 | 39 | | [Updates break working setups](https://feedbackbench.com/criteria/rel.update_breakage.md) | Too few posts | 0.508 | 0.490–0.532 | 7 | 2 | 5 | | [Outages, server errors and capacity or rate errors](https://feedbackbench.com/criteria/rel.service_errors.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Account and support: Typical (customer love 0.477, n 33) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Support, refunds and issue handling](https://feedbackbench.com/criteria/account.support.md) | Too few posts | 0.486 | 0.471–0.506 | 16 | 1 | 15 | | [Data retention, training use and deployment isolation](https://feedbackbench.com/criteria/account.data_privacy.md) | Too few posts | 0.484 | 0.470–0.500 | 15 | 1 | 14 | | [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | Too few posts | 0.497 | 0.494–0.500 | 2 | 0 | 2 | | [Account bans and access restrictions](https://feedbackbench.com/criteria/account.bans_restrictions.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | # Factory (Factory) Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/agent/factory | Measure | Value | |---|---| | Rank | =11 of 17 (rank range 11–12) | | Feedback Score | 35.6 (95% interval 34.9–36.2) | | Popularity | 0.232 (share of voice 0.80%) | | Customer love | 0.545 (95% interval 0.525–0.563) | | Top quadrant | no | | Authors | 789 | | Posts counted | 1576 | | Posts that judge the agent | 714 | | Criteria better / worse than peers | 2 / 0 of 63 | ## Top requests What users ask to add or change, most asked first. 183 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first. | Rank | Request | Criterion | Author-weeks | Posts | |---|---|---|---|---| | 1 | Bring existing subscription into this agent | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | 12 | 12 | | 2 | Add DeepSeek V4.1 Flash model | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 7 | 7 | | 3 | Add Muse Spark models | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 6 | 6 | | 4 | Official dedicated mobile app | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | 6 | 6 | | 5 | Remove the 5-hour usage window | [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | 6 | 6 | | 6 | Vision support for specific models | [Images, PDFs and file attachments as input](https://feedbackbench.com/criteria/context.attachments.md) | 4 | 4 | | 7 | Add Qwen models | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 3 | 4 | | 8 | Fix declined card payments and checkout failures | [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | 3 | 4 | | 9 | Higher overall usage limits | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | 3 | 4 | | 10 | Linux desktop app and support | [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | 3 | 4 | | 11 | Pin, sort and hide models in picker | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | 3 | 4 | | 12 | Escalation rate metric for routing | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | 3 | 3 | ## Facts | Fact | Value | |---|---| | Version | n/a | | Released | $150M Series C at $1.5B valuation: 2026-04 | | Price | Pro $20/mo, Plus $100/mo, Max $200/mo, Teams/Enterprise custom (no self-serve, no annual discount) | | Model | Multi-model routing ('model routing era' positioning) | | Surface | IDE, terminal, web/cloud | ## Sources | Channel | Source | Posts | |---|---|---| | X | @FactoryAI | 1111 | | X | @droid | 425 | | Reddit | r/FactoryAi | 19 | | Reddit | Posts that name it | 15 | | G2 | G2 | 6 | ## Better than peers on Which models are offered on a plan and when, Can do the user's kind of task ## Worse than peers on None. ## All 63 criteria Criterion love: 0.5 is the category norm. n: rated author-weeks. ### Paying and limits: Better than peers (customer love 0.542, n 125) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | Typical | 0.518 | 0.486–0.549 | 70 | 32 | 38 | | [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md) | Too few posts | 0.533 | 0.501–0.563 | 19 | 8 | 11 | | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | Too few posts | 0.507 | 0.488–0.527 | 13 | 5 | 8 | | [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | Too few posts | 0.483 | 0.471–0.493 | 12 | 0 | 12 | | [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md) | Too few posts | 0.501 | 0.481–0.532 | 10 | 1 | 9 | | [Usage meter visibility and accuracy](https://feedbackbench.com/criteria/limits.usage_meter.md) | Too few posts | 0.492 | 0.484–0.497 | 6 | 0 | 6 | | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | Too few posts | 0.496 | 0.484–0.508 | 6 | 3 | 3 | | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | Too few posts | 0.495 | 0.489–0.500 | 3 | 0 | 3 | | [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | Too few posts | 0.497 | 0.493–0.500 | 2 | 0 | 2 | | [Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md) | Too few posts | 0.499 | 0.491–0.506 | 2 | 1 | 1 | | [Pay-as-you-go overage, fallback billing and spend caps](https://feedbackbench.com/criteria/billing.overage_charges.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | ### Setting up and connecting: Typical (customer love 0.512, n 35) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | Too few posts | 0.501 | 0.482–0.519 | 17 | 9 | 8 | | [MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md) | Too few posts | 0.496 | 0.482–0.509 | 8 | 3 | 5 | | [Onboarding, discoverability and documentation](https://feedbackbench.com/criteria/setup.onboarding_docs.md) | Too few posts | 0.507 | 0.491–0.524 | 8 | 3 | 5 | | [IDE and editor integration](https://feedbackbench.com/criteria/setup.ide_integration.md) | Too few posts | 0.506 | 0.494–0.520 | 6 | 4 | 2 | | [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | Too few posts | 0.497 | 0.492–0.500 | 2 | 0 | 2 | ### Choosing models: Better than peers (customer love 0.579, n 57) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | Better than peers | 0.558 | 0.526–0.587 | 32 | 22 | 10 | | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | Too few posts | 0.536 | 0.514–0.560 | 18 | 14 | 4 | | [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md) | Too few posts | 0.506 | 0.491–0.523 | 9 | 4 | 5 | | [Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md) | Too few posts | 0.501 | 0.494–0.508 | 2 | 1 | 1 | ### Instructing and context: Too few posts (customer love 0.503, n 15) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Images, PDFs and file attachments as input](https://feedbackbench.com/criteria/context.attachments.md) | Too few posts | 0.490 | 0.480–0.498 | 5 | 0 | 5 | | [Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md) | Too few posts | 0.504 | 0.495–0.515 | 4 | 3 | 1 | | [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md) | Too few posts | 0.506 | 0.497–0.518 | 3 | 2 | 1 | | [Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md) | Too few posts | 0.503 | 0.500–0.511 | 1 | 1 | 0 | | [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | Too few posts | 0.503 | 0.500–0.511 | 1 | 1 | 0 | | [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md) | Too few posts | 0.497 | 0.491–0.500 | 1 | 0 | 1 | | [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Doing the work: Better than peers (customer love 0.577, n 95) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md) | Better than peers | 0.543 | 0.516–0.570 | 57 | 48 | 9 | | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | Too few posts | 0.514 | 0.497–0.530 | 15 | 12 | 3 | | [Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md) | Too few posts | 0.505 | 0.491–0.520 | 12 | 10 | 2 | | [Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md) | Too few posts | 0.494 | 0.489–0.499 | 4 | 0 | 4 | | [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | Too few posts | 0.507 | 0.494–0.523 | 4 | 2 | 2 | | [Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md) | Too few posts | 0.500 | 0.491–0.508 | 3 | 2 | 1 | | [Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md) | Too few posts | 0.505 | 0.500–0.511 | 3 | 3 | 0 | | [Plan-before-edit mode](https://feedbackbench.com/criteria/work.plan_mode.md) | Too few posts | 0.510 | 0.500–0.521 | 3 | 3 | 0 | | [Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md) | Too few posts | 0.499 | 0.495–0.500 | 1 | 0 | 1 | | [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | Too few posts | 0.502 | 0.500–0.507 | 1 | 1 | 0 | | [Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md) | Too few posts | 0.504 | 0.500–0.512 | 1 | 1 | 0 | | [Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Games checks instead of fixing the problem](https://feedbackbench.com/criteria/work.reward_hacking.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Risky or irreversible actions without confirmation](https://feedbackbench.com/criteria/work.destructive_actions.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Git commits, branches and sync](https://feedbackbench.com/criteria/work.git_workflow.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Safety filters block legitimate coding tasks](https://feedbackbench.com/criteria/work.safety_refusals.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Caves to or argues with the user's judgement](https://feedbackbench.com/criteria/work.sycophancy_pushback.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Checking and finishing: Too few posts (customer love 0.503, n 5) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md) | Too few posts | 0.503 | 0.500–0.508 | 2 | 2 | 0 | | [Claims work is done or fixed when it is not](https://feedbackbench.com/criteria/verify.false_completion.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Builds, tests or runs its own changes](https://feedbackbench.com/criteria/verify.self_testing.md) | Too few posts | 0.502 | 0.500–0.506 | 1 | 1 | 0 | | [Reviewing and approving the agent's changes](https://feedbackbench.com/criteria/verify.change_review_ui.md) | Too few posts | 0.498 | 0.494–0.500 | 1 | 0 | 1 | ### Interface and sessions: Typical (customer love 0.477, n 32) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | Too few posts | 0.485 | 0.465–0.507 | 22 | 5 | 17 | | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | Too few posts | 0.489 | 0.476–0.500 | 7 | 1 | 6 | | [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | Too few posts | 0.496 | 0.484–0.506 | 4 | 2 | 2 | | [Saving, switching, resuming and rewinding sessions](https://feedbackbench.com/criteria/ui.session_history.md) | Too few posts | 0.496 | 0.491–0.500 | 2 | 0 | 2 | | [Stopping and steering a running agent](https://feedbackbench.com/criteria/ui.interrupt_steer.md) | Too few posts | 0.503 | 0.500–0.508 | 1 | 1 | 0 | ### Reliability and speed: Too few posts (customer love 0.519, n 23) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Latency, throughput and fast mode](https://feedbackbench.com/criteria/rel.response_speed.md) | Too few posts | 0.511 | 0.495–0.529 | 11 | 7 | 4 | | [Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md) | Too few posts | 0.488 | 0.480–0.496 | 9 | 0 | 9 | | [Outages, server errors and capacity or rate errors](https://feedbackbench.com/criteria/rel.service_errors.md) | Too few posts | 0.496 | 0.491–0.500 | 3 | 0 | 3 | | [Updates break working setups](https://feedbackbench.com/criteria/rel.update_breakage.md) | Too few posts | 0.499 | 0.495–0.500 | 1 | 0 | 1 | ### Account and support: Too few posts (customer love 0.576, n 28) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Support, refunds and issue handling](https://feedbackbench.com/criteria/account.support.md) | Too few posts | 0.505 | 0.484–0.531 | 13 | 3 | 10 | | [Data retention, training use and deployment isolation](https://feedbackbench.com/criteria/account.data_privacy.md) | Too few posts | 0.544 | 0.521–0.567 | 11 | 11 | 0 | | [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | Too few posts | 0.493 | 0.483–0.499 | 6 | 0 | 6 | | [Account bans and access restrictions](https://feedbackbench.com/criteria/account.bans_restrictions.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | # Amp (Amp) Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/agent/amp | Measure | Value | |---|---| | Rank | =11 of 17 (rank range 11–12) | | Feedback Score | 35.3 (95% interval 34.7–35.8) | | Popularity | 0.222 (share of voice 0.75%) | | Customer love | 0.561 (95% interval 0.544–0.577) | | Top quadrant | no | | Authors | 737 | | Posts counted | 2786 | | Posts that judge the agent | 1292 | | Criteria better / worse than peers | 5 / 0 of 63 | ## Top requests What users ask to add or change, most asked first. 247 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first. | Rank | Request | Criterion | Author-weeks | Posts | |---|---|---|---|---| | 1 | Bring existing subscription into this agent | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | 32 | 34 | | 2 | macOS and Windows cloud environments | [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | 4 | 5 | | 3 | Access to subscription integration beta | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | 4 | 4 | | 4 | Add Opus 5.5 model | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 4 | 4 | | 5 | Custom base URL OpenAI-compatible endpoints | [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | 4 | 4 | | 6 | Session status dashboard | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | 4 | 4 | | 7 | Allow subscription use in third-party harnesses | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | 3 | 4 | | 8 | Option to hide sidebar | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | 3 | 4 | | 9 | Remove model selector slot limit | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 3 | 4 | | 10 | Savable model and mode presets | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | 3 | 4 | | 11 | Sidebar project grouping, sorting and filtering | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | 3 | 4 | | 12 | Access to experimental and unreleased features | [Onboarding, discoverability and documentation](https://feedbackbench.com/criteria/setup.onboarding_docs.md) | 3 | 3 | ## Facts | Fact | Value | |---|---| | Version | n/a (rolling releases) | | Released | Free/BYOK model shift: 2026-09-13 | | Price | Free with bring-your-own compute/model keys (as of 2026-09-13); earlier in 2026 was pay-as-you-go with a $5 minimum and zero markup on provider pricing for individuals; Smart Mode is the paid tier for zero data sharing | | Model | Multi-provider, bring-your-own-key | | Surface | IDE (VS Code), CLI | ## Sources | Channel | Source | Posts | |---|---|---| | X | @AmpCode | 2779 | | Reddit | Posts that name it | 7 | ## Better than peers on Can do the user's kind of task, Subagents, parallel agents and orchestrators, Mobile, remote-control and voice access, Cloud and remote sandbox execution, Support, refunds and issue handling ## Worse than peers on None. ## All 63 criteria Criterion love: 0.5 is the category norm. n: rated author-weeks. ### Paying and limits: Better than peers (customer love 0.569, n 135) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | Typical | 0.529 | 0.498–0.560 | 60 | 24 | 36 | | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | Typical | 0.521 | 0.493–0.551 | 41 | 21 | 20 | | [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md) | Too few posts | 0.501 | 0.477–0.528 | 21 | 4 | 17 | | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | Too few posts | 0.503 | 0.485–0.519 | 12 | 8 | 4 | | [Usage meter visibility and accuracy](https://feedbackbench.com/criteria/limits.usage_meter.md) | Too few posts | 0.549 | 0.513–0.586 | 9 | 6 | 3 | | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | Too few posts | 0.520 | 0.501–0.543 | 6 | 4 | 2 | | [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md) | Too few posts | 0.493 | 0.487–0.499 | 5 | 0 | 5 | | [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | Too few posts | 0.504 | 0.491–0.522 | 4 | 1 | 3 | | [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | Too few posts | 0.496 | 0.490–0.500 | 3 | 0 | 3 | | [Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md) | Too few posts | 0.501 | 0.492–0.508 | 3 | 2 | 1 | | [Pay-as-you-go overage, fallback billing and spend caps](https://feedbackbench.com/criteria/billing.overage_charges.md) | Too few posts | 0.497 | 0.493–0.500 | 2 | 0 | 2 | ### Setting up and connecting: Typical (customer love 0.528, n 106) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | Typical | 0.510 | 0.483–0.535 | 40 | 23 | 17 | | [MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md) | Typical | 0.495 | 0.471–0.520 | 33 | 15 | 18 | | [Onboarding, discoverability and documentation](https://feedbackbench.com/criteria/setup.onboarding_docs.md) | Too few posts | 0.516 | 0.489–0.545 | 21 | 6 | 15 | | [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | Too few posts | 0.510 | 0.486–0.539 | 20 | 5 | 15 | | [IDE and editor integration](https://feedbackbench.com/criteria/setup.ide_integration.md) | Too few posts | 0.500 | 0.490–0.512 | 5 | 2 | 3 | ### Choosing models: Better than peers (customer love 0.539, n 45) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | Too few posts | 0.515 | 0.491–0.540 | 23 | 10 | 13 | | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | Too few posts | 0.531 | 0.510–0.554 | 16 | 11 | 5 | | [Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md) | Too few posts | 0.493 | 0.479–0.507 | 10 | 3 | 7 | | [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md) | Too few posts | 0.512 | 0.497–0.529 | 6 | 4 | 2 | ### Instructing and context: Typical (customer love 0.525, n 32) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md) | Too few posts | 0.499 | 0.485–0.517 | 7 | 1 | 6 | | [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md) | Too few posts | 0.511 | 0.498–0.526 | 6 | 4 | 2 | | [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | Too few posts | 0.519 | 0.504–0.534 | 6 | 6 | 0 | | [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md) | Too few posts | 0.503 | 0.492–0.514 | 6 | 4 | 2 | | [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md) | Too few posts | 0.505 | 0.500–0.513 | 4 | 3 | 1 | | [Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md) | Too few posts | 0.500 | 0.493–0.506 | 2 | 1 | 1 | | [Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md) | Too few posts | 0.498 | 0.493–0.500 | 1 | 0 | 1 | | [Images, PDFs and file attachments as input](https://feedbackbench.com/criteria/context.attachments.md) | Too few posts | 0.498 | 0.493–0.500 | 1 | 0 | 1 | ### Doing the work: Better than peers (customer love 0.597, n 158) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | Better than peers | 0.558 | 0.533–0.583 | 57 | 49 | 8 | | [Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md) | Better than peers | 0.550 | 0.522–0.578 | 55 | 48 | 7 | | [Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md) | Too few posts | 0.510 | 0.486–0.532 | 26 | 22 | 4 | | [Git commits, branches and sync](https://feedbackbench.com/criteria/work.git_workflow.md) | Too few posts | 0.497 | 0.481–0.516 | 14 | 4 | 10 | | [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | Too few posts | 0.499 | 0.486–0.512 | 9 | 5 | 4 | | [Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md) | Too few posts | 0.505 | 0.492–0.515 | 7 | 6 | 1 | | [Risky or irreversible actions without confirmation](https://feedbackbench.com/criteria/work.destructive_actions.md) | Too few posts | 0.493 | 0.486–0.499 | 5 | 0 | 5 | | [Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md) | Too few posts | 0.492 | 0.485–0.498 | 4 | 0 | 4 | | [Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md) | Too few posts | 0.508 | 0.498–0.521 | 4 | 3 | 1 | | [Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md) | Too few posts | 0.493 | 0.481–0.500 | 2 | 0 | 2 | | [Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md) | Too few posts | 0.497 | 0.493–0.500 | 2 | 0 | 2 | | [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | Too few posts | 0.504 | 0.495–0.517 | 2 | 1 | 1 | | [Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md) | Too few posts | 0.498 | 0.495–0.500 | 1 | 0 | 1 | | [Plan-before-edit mode](https://feedbackbench.com/criteria/work.plan_mode.md) | Too few posts | 0.498 | 0.494–0.500 | 1 | 0 | 1 | | [Caves to or argues with the user's judgement](https://feedbackbench.com/criteria/work.sycophancy_pushback.md) | Too few posts | 0.498 | 0.494–0.500 | 1 | 0 | 1 | | [Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Games checks instead of fixing the problem](https://feedbackbench.com/criteria/work.reward_hacking.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Safety filters block legitimate coding tasks](https://feedbackbench.com/criteria/work.safety_refusals.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Checking and finishing: Too few posts (customer love 0.492, n 11) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Builds, tests or runs its own changes](https://feedbackbench.com/criteria/verify.self_testing.md) | Too few posts | 0.491 | 0.477–0.502 | 4 | 1 | 3 | | [Reviewing and approving the agent's changes](https://feedbackbench.com/criteria/verify.change_review_ui.md) | Too few posts | 0.501 | 0.492–0.511 | 4 | 2 | 2 | | [Claims work is done or fixed when it is not](https://feedbackbench.com/criteria/verify.false_completion.md) | Too few posts | 0.497 | 0.493–0.500 | 2 | 0 | 2 | | [Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md) | Too few posts | 0.503 | 0.500–0.509 | 2 | 2 | 0 | ### Interface and sessions: Better than peers (customer love 0.556, n 176) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | Typical | 0.512 | 0.481–0.543 | 91 | 37 | 54 | | [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | Better than peers | 0.533 | 0.506–0.560 | 51 | 42 | 9 | | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | Better than peers | 0.541 | 0.517–0.565 | 36 | 27 | 9 | | [Saving, switching, resuming and rewinding sessions](https://feedbackbench.com/criteria/ui.session_history.md) | Too few posts | 0.500 | 0.481–0.518 | 16 | 5 | 11 | | [Stopping and steering a running agent](https://feedbackbench.com/criteria/ui.interrupt_steer.md) | Too few posts | 0.495 | 0.488–0.500 | 2 | 0 | 2 | ### Reliability and speed: Typical (customer love 0.490, n 74) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Outages, server errors and capacity or rate errors](https://feedbackbench.com/criteria/rel.service_errors.md) | Typical | 0.539 | 0.480–0.597 | 37 | 5 | 32 | | [Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md) | Too few posts | 0.493 | 0.462–0.544 | 24 | 1 | 23 | | [Latency, throughput and fast mode](https://feedbackbench.com/criteria/rel.response_speed.md) | Too few posts | 0.497 | 0.481–0.512 | 13 | 5 | 8 | | [Updates break working setups](https://feedbackbench.com/criteria/rel.update_breakage.md) | Too few posts | 0.510 | 0.492–0.533 | 6 | 2 | 4 | ### Account and support: Better than peers (customer love 0.641, n 42) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Support, refunds and issue handling](https://feedbackbench.com/criteria/account.support.md) | Better than peers | 0.634 | 0.600–0.664 | 33 | 28 | 5 | | [Account bans and access restrictions](https://feedbackbench.com/criteria/account.bans_restrictions.md) | Too few posts | 0.492 | 0.486–0.497 | 6 | 0 | 6 | | [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | Too few posts | 0.495 | 0.490–0.499 | 4 | 0 | 4 | | [Data retention, training use and deployment isolation](https://feedbackbench.com/criteria/account.data_privacy.md) | Too few posts | 0.502 | 0.495–0.512 | 2 | 1 | 1 | # Kiro (AWS) Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/agent/kiro | Measure | Value | |---|---| | Rank | 13 of 17 (rank range 13–13) | | Feedback Score | 29.5 (95% interval 29.0–30.0) | | Popularity | 0.185 (share of voice 0.57%) | | Customer love | 0.470 (95% interval 0.453–0.485) | | Top quadrant | no | | Authors | 564 | | Posts counted | 1156 | | Posts that judge the agent | 537 | | Criteria better / worse than peers | 0 / 2 of 63 | ## Top requests What users ask to add or change, most asked first. 157 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first. | Rank | Request | Criterion | Author-weeks | Posts | |---|---|---|---|---| | 1 | Add Opus 5.5 model | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 29 | 37 | | 2 | Expand student program to more universities | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | 18 | 20 | | 3 | Fix declined card payments and checkout failures | [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | 8 | 9 | | 4 | Add Fable 5.1 model | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 8 | 8 | | 5 | Add GPT-6 Astra model | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 7 | 7 | | 6 | Lower or reverted model pricing | [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | 6 | 6 | | 7 | Add low-cost DeepSeek and GLM models together | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 5 | 5 | | 8 | Same-day availability of new models | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 5 | 5 | | 9 | Add DeepSeek V4.1 Flash model | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 4 | 5 | | 10 | Keep model catalog updated to latest versions | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 4 | 5 | | 11 | Add GPT-6 Sol and Luna models | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 4 | 4 | | 12 | Optional pricing tier for larger context | [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | 4 | 4 | ## Facts | Fact | Value | |---|---| | Version | Kiro Crew (open-source, self-hostable agent workspace), CLI | | Released | Invite-only mid-2025; GA early 2026 | | Price | Free (50 credits), Pro $20/mo (1,000 credits), Pro+ $40/mo (2,000), Pro Max $100/mo (5,000), Power $200/mo (10,000); $0.04/credit overage; Enterprise custom via AWS | | Model | Reported to run on Claude models | | Surface | IDE (VS Code-based), CLI | ## Sources | Channel | Source | Posts | |---|---|---| | Reddit | r/kiroIDE | 802 | | X | @kirodotdev | 297 | | Reddit | Posts that name it | 56 | | Trustpilot | Trustpilot | 1 | ## Better than peers on None. ## Worse than peers on Price, allowance or plan terms changed, Which models are offered on a plan and when ## All 63 criteria Criterion love: 0.5 is the category norm. n: rated author-weeks. ### Paying and limits: Worse than peers (customer love 0.444, n 128) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | Typical | 0.500 | 0.472–0.526 | 42 | 16 | 26 | | [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md) | Typical | 0.475 | 0.449–0.504 | 38 | 3 | 35 | | [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | Worse than peers | 0.456 | 0.444–0.470 | 36 | 0 | 36 | | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | Too few posts | 0.449 | 0.427–0.471 | 22 | 2 | 20 | | [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md) | Too few posts | 0.483 | 0.474–0.492 | 13 | 0 | 13 | | [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | Too few posts | 0.525 | 0.499–0.556 | 4 | 3 | 1 | | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | Too few posts | 0.494 | 0.487–0.499 | 4 | 0 | 4 | | [Pay-as-you-go overage, fallback billing and spend caps](https://feedbackbench.com/criteria/billing.overage_charges.md) | Too few posts | 0.497 | 0.493–0.500 | 2 | 0 | 2 | | [Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md) | Too few posts | 0.498 | 0.495–0.500 | 1 | 0 | 1 | | [Usage meter visibility and accuracy](https://feedbackbench.com/criteria/limits.usage_meter.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Setting up and connecting: Too few posts (customer love 0.481, n 25) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Onboarding, discoverability and documentation](https://feedbackbench.com/criteria/setup.onboarding_docs.md) | Too few posts | 0.496 | 0.483–0.511 | 8 | 1 | 7 | | [IDE and editor integration](https://feedbackbench.com/criteria/setup.ide_integration.md) | Too few posts | 0.497 | 0.484–0.510 | 8 | 3 | 5 | | [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | Too few posts | 0.489 | 0.481–0.497 | 7 | 0 | 7 | | [MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md) | Too few posts | 0.500 | 0.488–0.511 | 6 | 3 | 3 | | [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | Too few posts | 0.502 | 0.500–0.507 | 1 | 1 | 0 | ### Choosing models: Worse than peers (customer love 0.430, n 76) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | Worse than peers | 0.430 | 0.405–0.456 | 70 | 5 | 65 | | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | Too few posts | 0.498 | 0.486–0.513 | 6 | 1 | 5 | | [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md) | Too few posts | 0.498 | 0.495–0.500 | 1 | 0 | 1 | | [Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Instructing and context: Too few posts (customer love 0.507, n 18) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md) | Too few posts | 0.509 | 0.497–0.521 | 6 | 5 | 1 | | [Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md) | Too few posts | 0.511 | 0.494–0.534 | 5 | 2 | 3 | | [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | Too few posts | 0.495 | 0.488–0.500 | 3 | 0 | 3 | | [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md) | Too few posts | 0.500 | 0.492–0.508 | 2 | 1 | 1 | | [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md) | Too few posts | 0.498 | 0.495–0.500 | 1 | 0 | 1 | | [Images, PDFs and file attachments as input](https://feedbackbench.com/criteria/context.attachments.md) | Too few posts | 0.498 | 0.493–0.500 | 1 | 0 | 1 | | [Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Doing the work: Typical (customer love 0.496, n 60) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md) | Typical | 0.485 | 0.460–0.509 | 34 | 17 | 17 | | [Plan-before-edit mode](https://feedbackbench.com/criteria/work.plan_mode.md) | Too few posts | 0.515 | 0.504–0.527 | 8 | 7 | 1 | | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | Too few posts | 0.497 | 0.484–0.508 | 6 | 3 | 3 | | [Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md) | Too few posts | 0.493 | 0.486–0.499 | 5 | 0 | 5 | | [Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md) | Too few posts | 0.495 | 0.490–0.499 | 4 | 0 | 4 | | [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | Too few posts | 0.495 | 0.489–0.500 | 3 | 0 | 3 | | [Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md) | Too few posts | 0.497 | 0.487–0.505 | 2 | 1 | 1 | | [Safety filters block legitimate coding tasks](https://feedbackbench.com/criteria/work.safety_refusals.md) | Too few posts | 0.507 | 0.496–0.525 | 2 | 1 | 1 | | [Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Risky or irreversible actions without confirmation](https://feedbackbench.com/criteria/work.destructive_actions.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md) | Too few posts | 0.506 | 0.500–0.519 | 1 | 1 | 0 | | [Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Games checks instead of fixing the problem](https://feedbackbench.com/criteria/work.reward_hacking.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Git commits, branches and sync](https://feedbackbench.com/criteria/work.git_workflow.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Caves to or argues with the user's judgement](https://feedbackbench.com/criteria/work.sycophancy_pushback.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Checking and finishing: Too few posts (customer love 0.495, n 2) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Claims work is done or fixed when it is not](https://feedbackbench.com/criteria/verify.false_completion.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Builds, tests or runs its own changes](https://feedbackbench.com/criteria/verify.self_testing.md) | Too few posts | 0.496 | 0.487–0.500 | 1 | 0 | 1 | | [Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Reviewing and approving the agent's changes](https://feedbackbench.com/criteria/verify.change_review_ui.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Interface and sessions: Too few posts (customer love 0.497, n 14) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | Too few posts | 0.493 | 0.481–0.504 | 7 | 1 | 6 | | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | Too few posts | 0.497 | 0.487–0.504 | 3 | 1 | 2 | | [Saving, switching, resuming and rewinding sessions](https://feedbackbench.com/criteria/ui.session_history.md) | Too few posts | 0.508 | 0.500–0.521 | 2 | 2 | 0 | | [Stopping and steering a running agent](https://feedbackbench.com/criteria/ui.interrupt_steer.md) | Too few posts | 0.498 | 0.494–0.500 | 1 | 0 | 1 | | [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | Too few posts | 0.502 | 0.500–0.507 | 1 | 1 | 0 | ### Reliability and speed: Too few posts (customer love 0.490, n 23) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md) | Too few posts | 0.484 | 0.475–0.492 | 12 | 0 | 12 | | [Latency, throughput and fast mode](https://feedbackbench.com/criteria/rel.response_speed.md) | Too few posts | 0.496 | 0.483–0.510 | 8 | 2 | 6 | | [Outages, server errors and capacity or rate errors](https://feedbackbench.com/criteria/rel.service_errors.md) | Too few posts | 0.506 | 0.491–0.542 | 5 | 1 | 4 | | [Updates break working setups](https://feedbackbench.com/criteria/rel.update_breakage.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Account and support: Typical (customer love 0.478, n 59) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Support, refunds and issue handling](https://feedbackbench.com/criteria/account.support.md) | Typical | 0.485 | 0.462–0.512 | 30 | 3 | 27 | | [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | Too few posts | 0.500 | 0.459–0.563 | 28 | 1 | 27 | | [Account bans and access restrictions](https://feedbackbench.com/criteria/account.bans_restrictions.md) | Too few posts | 0.483 | 0.474–0.491 | 13 | 0 | 13 | | [Data retention, training use and deployment isolation](https://feedbackbench.com/criteria/account.data_privacy.md) | Too few posts | 0.507 | 0.490–0.532 | 6 | 2 | 4 | # Conductor (Melty Labs) Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/agent/conductor | Measure | Value | |---|---| | Rank | 14 of 17 (rank range 14–14) | | Feedback Score | 26.4 (95% interval 26.1–26.7) | | Popularity | 0.136 (share of voice 0.38%) | | Customer love | 0.510 (95% interval 0.499–0.521) | | Top quadrant | no | | Authors | 371 | | Posts counted | 530 | | Posts that judge the agent | 297 | | Criteria better / worse than peers | 0 / 0 of 63 | ## Top requests What users ask to add or change, most asked first. 94 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first. | Rank | Request | Criterion | Author-weeks | Posts | |---|---|---|---|---| | 1 | Native iOS app | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | 4 | 4 | | 2 | Phone control of local and desktop sessions | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | 4 | 4 | | 3 | Restore removed UI elements | [Updates break working setups](https://feedbackbench.com/criteria/rel.update_breakage.md) | 4 | 4 | | 4 | Add Grok 4.7 model | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 3 | 3 | | 5 | Remove model selector slot limit | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 3 | 3 | | 6 | Bring existing subscription into this agent | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | 2 | 3 | | 7 | Add Fable 5.1 model | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 2 | 2 | | 8 | Mobile access to cloud agents | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | 2 | 2 | | 9 | Model choice in cloud sessions | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 2 | 2 | | 10 | More color themes and theme customization | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | 2 | 2 | | 11 | Official dedicated mobile app | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | 2 | 2 | | 12 | Revert new model picker | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | 2 | 2 | ## Facts | Fact | Value | |---|---| | Version | Mac app with cloud workspaces | | Released | n/a | | Price | Free for local Mac workspaces; Pro $50/mo (cloud, multiplayer, API); Teams $60/user/mo; Enterprise custom. Agents run on the user's own Claude, Codex or Cursor subscription or API key | | Model | Runs Claude Code, Codex, Cursor and OpenCode agents with the user's own accounts | | Surface | Desktop orchestrator (macOS), cloud | ## Sources | Channel | Source | Posts | |---|---|---| | X | @conductor_build | 453 | | Reddit | r/conductorbuild | 58 | | Reddit | Posts that name it | 19 | ## Better than peers on None. ## Worse than peers on None. ## All 63 criteria Criterion love: 0.5 is the category norm. n: rated author-weeks. ### Paying and limits: Too few posts (customer love 0.516, n 18) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | Too few posts | 0.502 | 0.489–0.518 | 9 | 4 | 5 | | [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md) | Too few posts | 0.502 | 0.491–0.519 | 4 | 1 | 3 | | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | Too few posts | 0.504 | 0.500–0.510 | 2 | 2 | 0 | | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | Too few posts | 0.503 | 0.496–0.512 | 2 | 1 | 1 | | [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Usage meter visibility and accuracy](https://feedbackbench.com/criteria/limits.usage_meter.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Pay-as-you-go overage, fallback billing and spend caps](https://feedbackbench.com/criteria/billing.overage_charges.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Setting up and connecting: Too few posts (customer love 0.497, n 23) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md) | Too few posts | 0.503 | 0.489–0.517 | 9 | 5 | 4 | | [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | Too few posts | 0.501 | 0.490–0.517 | 5 | 1 | 4 | | [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | Too few posts | 0.492 | 0.481–0.502 | 5 | 1 | 4 | | [Onboarding, discoverability and documentation](https://feedbackbench.com/criteria/setup.onboarding_docs.md) | Too few posts | 0.497 | 0.492–0.500 | 2 | 0 | 2 | | [IDE and editor integration](https://feedbackbench.com/criteria/setup.ide_integration.md) | Too few posts | 0.501 | 0.494–0.509 | 2 | 1 | 1 | ### Choosing models: Too few posts (customer love 0.515, n 23) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | Too few posts | 0.511 | 0.494–0.529 | 13 | 6 | 7 | | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | Too few posts | 0.507 | 0.489–0.527 | 9 | 4 | 5 | | [Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md) | Too few posts | 0.498 | 0.493–0.500 | 1 | 0 | 1 | | [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Instructing and context: Too few posts (customer love 0.504, n 8) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md) | Too few posts | 0.500 | 0.491–0.509 | 4 | 2 | 2 | | [Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md) | Too few posts | 0.500 | 0.494–0.506 | 2 | 1 | 1 | | [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md) | Too few posts | 0.498 | 0.494–0.500 | 1 | 0 | 1 | | [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | Too few posts | 0.504 | 0.500–0.513 | 1 | 1 | 0 | | [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Images, PDFs and file attachments as input](https://feedbackbench.com/criteria/context.attachments.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Doing the work: Typical (customer love 0.520, n 43) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | Too few posts | 0.514 | 0.497–0.531 | 18 | 14 | 4 | | [Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md) | Too few posts | 0.484 | 0.464–0.503 | 11 | 4 | 7 | | [Git commits, branches and sync](https://feedbackbench.com/criteria/work.git_workflow.md) | Too few posts | 0.516 | 0.503–0.532 | 6 | 5 | 1 | | [Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md) | Too few posts | 0.498 | 0.487–0.510 | 5 | 3 | 2 | | [Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md) | Too few posts | 0.503 | 0.500–0.508 | 2 | 2 | 0 | | [Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md) | Too few posts | 0.497 | 0.487–0.505 | 2 | 1 | 1 | | [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | Too few posts | 0.499 | 0.491–0.506 | 2 | 1 | 1 | | [Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md) | Too few posts | 0.499 | 0.495–0.500 | 1 | 0 | 1 | | [Risky or irreversible actions without confirmation](https://feedbackbench.com/criteria/work.destructive_actions.md) | Too few posts | 0.498 | 0.495–0.500 | 1 | 0 | 1 | | [Plan-before-edit mode](https://feedbackbench.com/criteria/work.plan_mode.md) | Too few posts | 0.502 | 0.500–0.508 | 1 | 1 | 0 | | [Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Games checks instead of fixing the problem](https://feedbackbench.com/criteria/work.reward_hacking.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Safety filters block legitimate coding tasks](https://feedbackbench.com/criteria/work.safety_refusals.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Caves to or argues with the user's judgement](https://feedbackbench.com/criteria/work.sycophancy_pushback.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Checking and finishing: Too few posts (customer love 0.499, n 3) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Reviewing and approving the agent's changes](https://feedbackbench.com/criteria/verify.change_review_ui.md) | Too few posts | 0.502 | 0.495–0.511 | 2 | 1 | 1 | | [Builds, tests or runs its own changes](https://feedbackbench.com/criteria/verify.self_testing.md) | Too few posts | 0.497 | 0.492–0.500 | 1 | 0 | 1 | | [Claims work is done or fixed when it is not](https://feedbackbench.com/criteria/verify.false_completion.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Interface and sessions: Typical (customer love 0.494, n 54) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | Too few posts | 0.491 | 0.471–0.511 | 25 | 7 | 18 | | [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | Too few posts | 0.500 | 0.477–0.522 | 25 | 17 | 8 | | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | Too few posts | 0.480 | 0.463–0.496 | 17 | 4 | 13 | | [Saving, switching, resuming and rewinding sessions](https://feedbackbench.com/criteria/ui.session_history.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Stopping and steering a running agent](https://feedbackbench.com/criteria/ui.interrupt_steer.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Reliability and speed: Too few posts (customer love 0.487, n 14) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md) | Too few posts | 0.488 | 0.481–0.495 | 9 | 0 | 9 | | [Updates break working setups](https://feedbackbench.com/criteria/rel.update_breakage.md) | Too few posts | 0.494 | 0.488–0.499 | 4 | 0 | 4 | | [Latency, throughput and fast mode](https://feedbackbench.com/criteria/rel.response_speed.md) | Too few posts | 0.501 | 0.494–0.509 | 2 | 1 | 1 | | [Outages, server errors and capacity or rate errors](https://feedbackbench.com/criteria/rel.service_errors.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Account and support: Too few posts (customer love 0.505, n 4) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Support, refunds and issue handling](https://feedbackbench.com/criteria/account.support.md) | Too few posts | 0.506 | 0.496–0.523 | 2 | 1 | 1 | | [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Account bans and access restrictions](https://feedbackbench.com/criteria/account.bans_restrictions.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Data retention, training use and deployment isolation](https://feedbackbench.com/criteria/account.data_privacy.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | # Warp (Warp) Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/agent/warp | Measure | Value | |---|---| | Rank | 15 of 17 (rank range 15–15) | | Feedback Score | 23.2 (95% interval 23.0–23.5) | | Popularity | 0.107 (share of voice 0.28%) | | Customer love | 0.506 (95% interval 0.495–0.516) | | Top quadrant | no | | Authors | 272 | | Posts counted | 424 | | Posts that judge the agent | 227 | | Criteria better / worse than peers | 0 / 0 of 63 | ## Top requests What users ask to add or change, most asked first. 33 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first. | Rank | Request | Criterion | Author-weeks | Posts | |---|---|---|---|---| | 1 | BYOK on all plans without credit gating | [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | 2 | 2 | | 2 | Built-in multi-agent orchestrator mode | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | 2 | 2 | | 3 | Cheaper low-cost plan tier | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | 2 | 2 | ## Facts | Fact | Value | |---|---| | Version | Warp 2.0, Agentic Development Environment; client open-sourced 2026-04-28 (dual AGPL-3.0/MIT) | | Released | Open-sourced client: 2026-04-28 | | Price | Free (~75 credits/mo after intro period), Build $20/mo (1,500 credits), Max $200/mo (18,000 credits), Business $50/user/mo | | Model | Own Agent Mode plus orchestration of Claude Code, Codex, and Gemini CLI inside one window | | Surface | Terminal | ## Sources | Channel | Source | Posts | |---|---|---| | X | @warpdotdev | 357 | | Reddit | r/warpdotdev | 45 | | Reddit | Posts that name it | 20 | | Trustpilot | Trustpilot | 2 | ## Better than peers on None. ## Worse than peers on None. ## All 63 criteria Criterion love: 0.5 is the category norm. n: rated author-weeks. ### Paying and limits: Too few posts (customer love 0.509, n 21) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md) | Too few posts | 0.508 | 0.487–0.534 | 11 | 3 | 8 | | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | Too few posts | 0.499 | 0.486–0.510 | 6 | 2 | 4 | | [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | Too few posts | 0.496 | 0.490–0.500 | 3 | 0 | 3 | | [Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md) | Too few posts | 0.499 | 0.491–0.506 | 2 | 1 | 1 | | [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md) | Too few posts | 0.512 | 0.496–0.540 | 2 | 1 | 1 | | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | Too few posts | 0.494 | 0.484–0.500 | 2 | 0 | 2 | | [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Usage meter visibility and accuracy](https://feedbackbench.com/criteria/limits.usage_meter.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Pay-as-you-go overage, fallback billing and spend caps](https://feedbackbench.com/criteria/billing.overage_charges.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Setting up and connecting: Too few posts (customer love 0.490, n 13) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | Too few posts | 0.494 | 0.482–0.505 | 6 | 2 | 4 | | [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | Too few posts | 0.497 | 0.492–0.500 | 2 | 0 | 2 | | [MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md) | Too few posts | 0.495 | 0.488–0.500 | 2 | 0 | 2 | | [IDE and editor integration](https://feedbackbench.com/criteria/setup.ide_integration.md) | Too few posts | 0.501 | 0.493–0.509 | 2 | 1 | 1 | | [Onboarding, discoverability and documentation](https://feedbackbench.com/criteria/setup.onboarding_docs.md) | Too few posts | 0.498 | 0.495–0.500 | 1 | 0 | 1 | ### Choosing models: Too few posts (customer love 0.510, n 7) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | Too few posts | 0.516 | 0.504–0.531 | 4 | 4 | 0 | | [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md) | Too few posts | 0.496 | 0.491–0.500 | 2 | 0 | 2 | | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | Too few posts | 0.498 | 0.495–0.500 | 1 | 0 | 1 | | [Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Instructing and context: Too few posts (customer love 0.503, n 4) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md) | Too few posts | 0.497 | 0.492–0.500 | 2 | 0 | 2 | | [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md) | Too few posts | 0.505 | 0.500–0.513 | 2 | 2 | 0 | | [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Images, PDFs and file attachments as input](https://feedbackbench.com/criteria/context.attachments.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Doing the work: Too few posts (customer love 0.500, n 17) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md) | Too few posts | 0.502 | 0.491–0.513 | 7 | 5 | 2 | | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | Too few posts | 0.498 | 0.485–0.510 | 7 | 4 | 3 | | [Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md) | Too few posts | 0.498 | 0.489–0.505 | 2 | 1 | 1 | | [Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md) | Too few posts | 0.497 | 0.493–0.500 | 2 | 0 | 2 | | [Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md) | Too few posts | 0.499 | 0.495–0.500 | 1 | 0 | 1 | | [Games checks instead of fixing the problem](https://feedbackbench.com/criteria/work.reward_hacking.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | Too few posts | 0.497 | 0.491–0.500 | 1 | 0 | 1 | | [Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Risky or irreversible actions without confirmation](https://feedbackbench.com/criteria/work.destructive_actions.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Git commits, branches and sync](https://feedbackbench.com/criteria/work.git_workflow.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Safety filters block legitimate coding tasks](https://feedbackbench.com/criteria/work.safety_refusals.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Plan-before-edit mode](https://feedbackbench.com/criteria/work.plan_mode.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Caves to or argues with the user's judgement](https://feedbackbench.com/criteria/work.sycophancy_pushback.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Checking and finishing: Too few posts (customer love 0.505, n 6) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md) | Too few posts | 0.498 | 0.483–0.509 | 4 | 3 | 1 | | [Claims work is done or fixed when it is not](https://feedbackbench.com/criteria/verify.false_completion.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Reviewing and approving the agent's changes](https://feedbackbench.com/criteria/verify.change_review_ui.md) | Too few posts | 0.503 | 0.500–0.510 | 1 | 1 | 0 | | [Builds, tests or runs its own changes](https://feedbackbench.com/criteria/verify.self_testing.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Interface and sessions: Typical (customer love 0.512, n 42) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | Too few posts | 0.498 | 0.476–0.523 | 26 | 9 | 17 | | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | Too few posts | 0.520 | 0.505–0.537 | 10 | 9 | 1 | | [Saving, switching, resuming and rewinding sessions](https://feedbackbench.com/criteria/ui.session_history.md) | Too few posts | 0.501 | 0.490–0.515 | 6 | 2 | 4 | | [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | Too few posts | 0.504 | 0.500–0.509 | 2 | 2 | 0 | | [Stopping and steering a running agent](https://feedbackbench.com/criteria/ui.interrupt_steer.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Reliability and speed: Too few posts (customer love 0.514, n 17) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md) | Too few posts | 0.488 | 0.481–0.495 | 9 | 0 | 9 | | [Latency, throughput and fast mode](https://feedbackbench.com/criteria/rel.response_speed.md) | Too few posts | 0.509 | 0.497–0.523 | 7 | 5 | 2 | | [Outages, server errors and capacity or rate errors](https://feedbackbench.com/criteria/rel.service_errors.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Updates break working setups](https://feedbackbench.com/criteria/rel.update_breakage.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Account and support: Too few posts (customer love 0.507, n 9) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Support, refunds and issue handling](https://feedbackbench.com/criteria/account.support.md) | Too few posts | 0.506 | 0.487–0.527 | 8 | 2 | 6 | | [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Account bans and access restrictions](https://feedbackbench.com/criteria/account.bans_restrictions.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Data retention, training use and deployment isolation](https://feedbackbench.com/criteria/account.data_privacy.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | # Grok Build (xAI) Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/agent/grok-build | Measure | Value | |---|---| | Rank | 16 of 17 (rank range 16–16) | | Feedback Score | 12.6 (95% interval 12.4–12.7) | | Popularity | 0.031 (share of voice 0.07%) | | Customer love | 0.517 (95% interval 0.506–0.527) | | Top quadrant | no | | Authors | 66 | | Posts counted | 78 | | Posts that judge the agent | 51 | | Criteria better / worse than peers | 0 / 0 of 63 | ## Top requests What users ask to add or change, most asked first. 8 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first. No request is asked for in enough author-weeks to show. ## Facts | Fact | Value | |---|---| | Version | v1.0 (out of beta 2026-08-07); underlying model grok-code-fast-1 | | Released | Beta: 2026-05-14. v1.0: 2026-08-07 | | Price | Bundled with SuperGrok Heavy ($300/mo); Grok Bot beta access also reachable via Cursor Ultra ($200/mo) or Cursor Teams Premium ($120/seat/mo) - standalone Grok Build pricing not clearly separated in sources | | Model | grok-code-fast-1, trained from scratch (not the Grok 4 lineage), heavy on programming corpus and real-world PR post-training | | Surface | CLI (local-first, no code sent to xAI servers) | ## Sources | Channel | Source | Posts | |---|---|---| | Reddit | Posts that name it | 78 | ## Better than peers on None. ## Worse than peers on None. ## All 63 criteria Criterion love: 0.5 is the category norm. n: rated author-weeks. ### Paying and limits: Too few posts (customer love 0.535, n 12) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | Too few posts | 0.515 | 0.499–0.533 | 8 | 6 | 2 | | [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md) | Too few posts | 0.506 | 0.496–0.523 | 2 | 1 | 1 | | [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | Too few posts | 0.503 | 0.500–0.508 | 1 | 1 | 0 | | [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Usage meter visibility and accuracy](https://feedbackbench.com/criteria/limits.usage_meter.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Pay-as-you-go overage, fallback billing and spend caps](https://feedbackbench.com/criteria/billing.overage_charges.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Setting up and connecting: Too few posts (customer love 0.499, n 5) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [IDE and editor integration](https://feedbackbench.com/criteria/setup.ide_integration.md) | Too few posts | 0.493 | 0.485–0.500 | 3 | 0 | 3 | | [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | Too few posts | 0.502 | 0.500–0.507 | 1 | 1 | 0 | | [MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md) | Too few posts | 0.502 | 0.500–0.507 | 1 | 1 | 0 | | [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Onboarding, discoverability and documentation](https://feedbackbench.com/criteria/setup.onboarding_docs.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Choosing models: Too few posts (customer love 0.510, n 2) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | Too few posts | 0.505 | 0.500–0.515 | 1 | 1 | 0 | | [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md) | Too few posts | 0.506 | 0.500–0.518 | 1 | 1 | 0 | | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Instructing and context: Too few posts (customer love 0.505, n 3) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md) | Too few posts | 0.509 | 0.500–0.523 | 2 | 2 | 0 | | [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md) | Too few posts | 0.498 | 0.493–0.500 | 1 | 0 | 1 | | [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Images, PDFs and file attachments as input](https://feedbackbench.com/criteria/context.attachments.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Doing the work: Too few posts (customer love 0.513, n 14) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md) | Too few posts | 0.501 | 0.486–0.516 | 10 | 6 | 4 | | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | Too few posts | 0.504 | 0.500–0.510 | 2 | 2 | 0 | | [Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Safety filters block legitimate coding tasks](https://feedbackbench.com/criteria/work.safety_refusals.md) | Too few posts | 0.509 | 0.500–0.527 | 1 | 1 | 0 | | [Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Games checks instead of fixing the problem](https://feedbackbench.com/criteria/work.reward_hacking.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Risky or irreversible actions without confirmation](https://feedbackbench.com/criteria/work.destructive_actions.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Git commits, branches and sync](https://feedbackbench.com/criteria/work.git_workflow.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Plan-before-edit mode](https://feedbackbench.com/criteria/work.plan_mode.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Caves to or argues with the user's judgement](https://feedbackbench.com/criteria/work.sycophancy_pushback.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Checking and finishing: Too few posts (customer love 0.506, n 2) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md) | Too few posts | 0.503 | 0.500–0.509 | 2 | 2 | 0 | | [Claims work is done or fixed when it is not](https://feedbackbench.com/criteria/verify.false_completion.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Builds, tests or runs its own changes](https://feedbackbench.com/criteria/verify.self_testing.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Reviewing and approving the agent's changes](https://feedbackbench.com/criteria/verify.change_review_ui.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Interface and sessions: Too few posts (customer love 0.503, n 1) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | Too few posts | 0.504 | 0.500–0.512 | 1 | 1 | 0 | | [Saving, switching, resuming and rewinding sessions](https://feedbackbench.com/criteria/ui.session_history.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Stopping and steering a running agent](https://feedbackbench.com/criteria/ui.interrupt_steer.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Reliability and speed: Too few posts (customer love 0.503, n 3) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Latency, throughput and fast mode](https://feedbackbench.com/criteria/rel.response_speed.md) | Too few posts | 0.500 | 0.492–0.509 | 3 | 1 | 2 | | [Outages, server errors and capacity or rate errors](https://feedbackbench.com/criteria/rel.service_errors.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Updates break working setups](https://feedbackbench.com/criteria/rel.update_breakage.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Account and support: Too few posts (customer love 0.494, n 4) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Data retention, training use and deployment isolation](https://feedbackbench.com/criteria/account.data_privacy.md) | Too few posts | 0.496 | 0.491–0.500 | 3 | 0 | 3 | | [Support, refunds and issue handling](https://feedbackbench.com/criteria/account.support.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Account bans and access restrictions](https://feedbackbench.com/criteria/account.bans_restrictions.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | # Augment Code (Augment) Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/agent/augment | Measure | Value | |---|---| | Rank | 17 of 17 (rank range 17–17) | | Feedback Score | 9.7 (95% interval 9.7–9.7) | | Popularity | 0.019 (share of voice 0.04%) | | Customer love | 0.495 (95% interval 0.491–0.499) | | Top quadrant | no | | Authors | 40 | | Posts counted | 46 | | Posts that judge the agent | 32 | | Criteria better / worse than peers | 0 / 0 of 63 | ## Top requests What users ask to add or change, most asked first. 2 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first. No request is asked for in enough author-weeks to show. ## Facts | Fact | Value | |---|---| | Version | Cosmos 'Unified Agents Platform' (team agent fleets across the SDLC); Intent desktop workspace (public beta, macOS); Auggie CLI 0.36.0 (2026-08-21); VS Code extension 0.901.1 and IntelliJ plugin v0.491.0 (2026-09-14); Context Engine MCP (GA 2026-02-06). Inline Completions and Next Edit sunset 2026-03-31 on Indie/Standard/Legacy plans | | Released | IDE extensions VS Code 0.901.1 / IntelliJ v0.491.0: 2026-09-14; Cosmos launch: 2026-06-05; Intent public beta: 2026-02-26; new $20/mo Standard tier first seen on pricing page 2026-09-20 (exact go-live date not confirmed) | | Price | Standard $20/mo (includes $20 of usage, up to 50 seats, pooled), Business $100/mo (includes $100 of usage, up to 50 seats), Enterprise custom; overage billed at LLM provider list price plus 40% service fee, plus Context Engine and Cosmos compute; top-ups valid 12 months | | Model | Multi-vendor model picker: Claude (Fable 5.1, Fable 5, Opus 5, Opus 4.6-4.8, Sonnet 5, Sonnet 4.6, Haiku 4.5), Gemini (3.1 Pro, 3.7/3.8 Flash), OpenAI GPT, xAI Grok, Zhipu GLM, Moonshot Kimi, plus Augment's own 'Prism' routing; Intent also runs BYO agents (Claude Code, Codex, OpenCode) | | Surface | VS Code and JetBrains extensions, Auggie CLI (terminal, also runs as MCP server), Intent desktop app (macOS beta), Cosmos web platform, Context Engine MCP for third-party agents | ## Sources | Channel | Source | Posts | |---|---|---| | X | @augmentcode | 28 | | Reddit | Posts that name it | 12 | | G2 | G2 | 6 | ## Better than peers on None. ## Worse than peers on None. ## All 63 criteria Criterion love: 0.5 is the category norm. n: rated author-weeks. ### Paying and limits: Too few posts (customer love 0.487, n 8) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | Too few posts | 0.495 | 0.486–0.506 | 5 | 1 | 4 | | [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md) | Too few posts | 0.494 | 0.488–0.499 | 4 | 0 | 4 | | [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | Too few posts | 0.497 | 0.494–0.500 | 2 | 0 | 2 | | [Usage meter visibility and accuracy](https://feedbackbench.com/criteria/limits.usage_meter.md) | Too few posts | 0.497 | 0.493–0.500 | 2 | 0 | 2 | | [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Pay-as-you-go overage, fallback billing and spend caps](https://feedbackbench.com/criteria/billing.overage_charges.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Setting up and connecting: Too few posts (customer love 0.499, n 5) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [IDE and editor integration](https://feedbackbench.com/criteria/setup.ide_integration.md) | Too few posts | 0.504 | 0.500–0.509 | 3 | 3 | 0 | | [Onboarding, discoverability and documentation](https://feedbackbench.com/criteria/setup.onboarding_docs.md) | Too few posts | 0.499 | 0.494–0.504 | 2 | 1 | 1 | | [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Choosing models: Too few posts (customer love 0.498, n 1) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md) | Too few posts | 0.498 | 0.494–0.500 | 1 | 0 | 1 | | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Instructing and context: Too few posts (customer love 0.520, n 8) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md) | Too few posts | 0.517 | 0.507–0.529 | 8 | 8 | 0 | | [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Images, PDFs and file attachments as input](https://feedbackbench.com/criteria/context.attachments.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Doing the work: Too few posts (customer love 0.495, n 7) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md) | Too few posts | 0.500 | 0.491–0.509 | 4 | 3 | 1 | | [Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md) | Too few posts | 0.497 | 0.493–0.500 | 2 | 0 | 2 | | [Risky or irreversible actions without confirmation](https://feedbackbench.com/criteria/work.destructive_actions.md) | Too few posts | 0.499 | 0.495–0.500 | 1 | 0 | 1 | | [Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Games checks instead of fixing the problem](https://feedbackbench.com/criteria/work.reward_hacking.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Git commits, branches and sync](https://feedbackbench.com/criteria/work.git_workflow.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Safety filters block legitimate coding tasks](https://feedbackbench.com/criteria/work.safety_refusals.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Plan-before-edit mode](https://feedbackbench.com/criteria/work.plan_mode.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Caves to or argues with the user's judgement](https://feedbackbench.com/criteria/work.sycophancy_pushback.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Checking and finishing: Too few posts (customer love 0.503, n 3) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md) | Too few posts | 0.503 | 0.500–0.508 | 2 | 2 | 0 | | [Claims work is done or fixed when it is not](https://feedbackbench.com/criteria/verify.false_completion.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 | | [Builds, tests or runs its own changes](https://feedbackbench.com/criteria/verify.self_testing.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Reviewing and approving the agent's changes](https://feedbackbench.com/criteria/verify.change_review_ui.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Interface and sessions: Too few posts (customer love 0.502, n 2) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | Too few posts | 0.503 | 0.500–0.508 | 1 | 1 | 0 | | [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | Too few posts | 0.498 | 0.495–0.500 | 1 | 0 | 1 | | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Saving, switching, resuming and rewinding sessions](https://feedbackbench.com/criteria/ui.session_history.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Stopping and steering a running agent](https://feedbackbench.com/criteria/ui.interrupt_steer.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Reliability and speed: Too few posts (customer love 0.500, n 0) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Outages, server errors and capacity or rate errors](https://feedbackbench.com/criteria/rel.service_errors.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Latency, throughput and fast mode](https://feedbackbench.com/criteria/rel.response_speed.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Updates break working setups](https://feedbackbench.com/criteria/rel.update_breakage.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | ### Account and support: Too few posts (customer love 0.495, n 3) | Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint | |---|---|---|---|---|---|---| | [Support, refunds and issue handling](https://feedbackbench.com/criteria/account.support.md) | Too few posts | 0.495 | 0.490–0.500 | 3 | 0 | 3 | | [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Account bans and access restrictions](https://feedbackbench.com/criteria/account.bans_restrictions.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 | | [Data retention, training use and deployment isolation](https://feedbackbench.com/criteria/account.data_privacy.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 |