# Feedback Bench method (coding agents, built 2026-10-01)

Each rule below is quoted verbatim from the build file.

Names: popularity is `reach` and customer love is `regard` in the data files (coding.json, rankings.csv, criteria.csv).

## Sources

2026-08-31 to the Sunday before 2026-09-28. Four channels, collected the same way for every agent: (1) the agent's official subreddits, posts and comments; (2) Reddit posts and comments whose own text names the agent, from any subreddit we collect and from a Reddit search per agent; (3) posts on X that mention the agent's official handles (for OpenAI Codex, which has no product handle, an X search for "OpenAI Codex", "Codex CLI" and "Codex app"); (4) G2 and Trustpilot reviews of the product. Parent-brand review pages and staff handles are excluded. Each post counts once per agent.

## Attribution

A post in an agent's own subreddit counts for that agent. Any other post counts for an agent only when the classifier confirms it is about that agent; a post can count for several agents.

## Labels

Every post is labelled by Claude Sonnet 5 with the criteria codebook v1.0 (63 criteria in 9 areas): whether it is about the agent, whether it judges it, which criteria its stance names, and praise, complaint or mixed for each. On a pilot of 308 posts, Claude Opus 5.5 labelling independently agreed at kappa 0.75 on whether a post judges the agent, 0.73 overlap on its criteria, 0.78 on its areas, and 98% on polarity. Human annotators have not yet validated v1.0. 2 post(s) the classifier declined to label are left out.

## Unit

An author-week: one author on one platform, about one agent, in one week. For each criterion it is positive if that author's praise outweighs their complaints that week, negative if the reverse. On Reddit, AutoModerator, [deleted] and names ending in 'bot' are not counted. Each review counts as its own author when the platform gives no name.

## Popularity

Share of voice = the agent's distinct authors across all four channels / the sum over all agents. Popularity = log(1 + share / 0.005) / log(1 + leader's share / 0.005). Popularity counts every author in the window, so it has no sampling interval.

## Criterion love

For each criterion: the category baseline is (positive + 1) / (positive + negative + 2) author-weeks across all agents. Channels differ in tone, so each agent is compared with its own channel mix: its expected share is the category baseline of each channel (Reddit, X, G2, Trustpilot), weighted by the agent's author-weeks on that channel. The agent's positive share is shrunk toward that expected share with 200 imaginary author-weeks and compared with it on the log-odds scale. 0.5 is the category norm.

## Customer love

Customer love = the customer love scores of the 9 areas, averaged on the log-odds scale with weights equal to each area's share of the category's rated author-weeks.

## Feedback Score

Feedback Score = 100 × sqrt(Popularity × Customer love).

## Intervals

95% intervals and rank ranges from 1000 bootstrap resamples of authors.

## Eligibility

Every measured agent is ranked. Popularity keeps agents with few authors low; shrinkage keeps their customer love near the norm.

## Receipts

The 5 most recent praise and complaint posts per area (first 700 characters) and 3 per criterion (first 450 characters).

## Caveat

A subreddit is a proxy for the product. X volume is capped at 1,000 posts a day per feed; no feed in the window reached it on its median day. Labels are automatic. Criteria seen in few posts have wide intervals; those under 30 author-weeks for an agent are shown as too few posts.

## Head to head

Posts that judge two agents. In each post, the agent whose stance (praise minus complaint across its criteria) is higher is ahead; ties are left out. Shares count the posts where one is ahead, with a 95% Wilson interval. Where: the agent's own subreddit, the other's, or anywhere else. Pairs with at least 30 such posts are shown.

## Requests

A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

## Facts

Agent facts come from web research dated 2026-09-24 and are not all verified against vendor pages.

## Display rules

- Reading of a criterion or area: Better than peers when its customer love 95% interval lies wholly above 0.5; Worse than peers when wholly below; Typical when it spans 0.5; Too few posts under 30 rated author-weeks.
- Top quadrant: popularity of at least 0.5 and customer love of at least 0.5; at the edge when the customer love 95% interval still includes 0.5.
- Rank shows “=N” when rank ranges overlap, directly or through a chain.

## Parameters

| Parameter | Value |
|---|---|
| S0 | 0.005 |
| priorStrength | 200 |
| bootstrap | 1000 |
| seed | 20260927 |

## Area weights in Customer love

Each area's share of the category's rated author-weeks.

| Area | Weight |
|---|---|
| Paying and limits | 0.339 |
| Setting up and connecting | 0.053 |
| Choosing models | 0.112 |
| Instructing and context | 0.059 |
| Doing the work | 0.230 |
| Checking and finishing | 0.021 |
| Interface and sessions | 0.058 |
| Reliability and speed | 0.091 |
| Account and support | 0.037 |

## Category baselines (positive share of author-weeks, %)

| Code | Name | Baseline % |
|---|---|---|
| paying | Paying and limits | 23.9 |
| setup | Setting up and connecting | 40.1 |
| models | Choosing models | 25.9 |
| context | Instructing and context | 36.2 |
| work | Doing the work | 47.5 |
| checking | Checking and finishing | 44.0 |
| interface | Interface and sessions | 41.4 |
| reliability | Reliability and speed | 19.0 |
| account | Account and support | 13.1 |
| limits.plan_value | How much use a plan's price buys | 37.9 |
| limits.window_interrupts_work | Short rolling usage window blocks or interrupts work | 13.8 |
| limits.burn_rate | Single prompt, model or effort level consumes disproportionate quota | 16.2 |
| limits.allowance_change | Price, allowance or plan terms changed | 5.8 |
| limits.reset_schedule | Quota reset timing and bonus or banked resets | 17.0 |
| limits.usage_meter | Usage meter visibility and accuracy | 7.8 |
| limits.prompt_cache | Prompt cache hits, misses and invalidation | 36.1 |
| billing.overage_charges | Pay-as-you-go overage, fallback billing and spend caps | 8.0 |
| billing.pricing_clarity | Pricing and plan terms stated clearly and consistently | 6.3 |
| billing.free_tier | Free tier and free model availability and limits | 57.4 |
| billing.subscription_portability | Using an existing subscription across tools | 38.5 |
| setup.install_signin | Install, launch and sign-in | 20.0 |
| setup.provider_byok_local | Connecting own API keys, local models and custom endpoints | 54.2 |
| setup.extensions_mcp | MCP servers, plugins, skills and hooks | 50.4 |
| setup.onboarding_docs | Onboarding, discoverability and documentation | 19.7 |
| setup.ide_integration | IDE and editor integration | 45.6 |
| models.catalog_access | Which models are offered on a plan and when | 28.1 |
| models.routing_auto | Automatic model routing and fallback | 24.6 |
| models.effort_control | Reasoning effort setting and its defaults | 45.7 |
| models.quality_drift | Quality got worse or better over time | 21.8 |
| context.instruction_files | Persistent project rules files are read and obeyed | 52.8 |
| context.instruction_following | Direct in-prompt instructions and caps are followed | 26.9 |
| context.clarifying_questions | Asks the user versus guessing | 38.6 |
| context.long_context_decay | Output degrades as the context window fills | 15.5 |
| context.compaction | Context compaction keeps what matters, cheaply and quickly | 33.3 |
| context.session_memory | Memory and state carried across sessions | 46.7 |
| context.codebase_retrieval | Finding the right files in the codebase | 44.3 |
| context.attachments | Images, PDFs and file attachments as input | 38.8 |
| work.capability | Can do the user's kind of task | 60.7 |
| work.frontend_ui | Frontend and visual UI output | 44.4 |
| work.bug_diagnosis | Diagnosing and fixing reported bugs | 67.4 |
| work.regressions_introduced | Breaks existing code or reintroduces bugs | 9.2 |
| work.scope_overreach | Does unrequested work or over-engineers | 6.0 |
| work.stuck_loops | Spins, loops or gets stuck without progress | 5.1 |
| work.premature_stop | Stops mid-task or answers instead of acting | 14.9 |
| work.long_running_autonomy | Long unattended runs and goal/loop mode | 74.2 |
| work.multi_agent_orchestration | Subagents, parallel agents and orchestrators | 60.6 |
| work.reward_hacking | Games checks instead of fixing the problem | 2.5 |
| work.destructive_actions | Risky or irreversible actions without confirmation | 18.5 |
| work.git_workflow | Git commits, branches and sync | 35.9 |
| work.computer_browser_use | Computer use and browser control | 54.8 |
| work.safety_refusals | Safety filters block legitimate coding tasks | 11.9 |
| work.permission_prompts | Tool approval prompts and autonomy modes | 24.9 |
| work.plan_mode | Plan-before-edit mode | 46.2 |
| work.response_verbosity | Length and clarity of replies, summaries and comments | 21.4 |
| work.sycophancy_pushback | Caves to or argues with the user's judgement | 16.7 |
| verify.false_completion | Claims work is done or fixed when it is not | 5.6 |
| verify.self_testing | Builds, tests or runs its own changes | 54.5 |
| verify.agent_code_review | Agent-performed code review finds real issues | 73.3 |
| verify.change_review_ui | Reviewing and approving the agent's changes | 37.8 |
| ui.display_settings | How the interface shows work, and what the user can configure | 34.6 |
| ui.session_history | Saving, switching, resuming and rewinding sessions | 29.2 |
| ui.interrupt_steer | Stopping and steering a running agent | 38.9 |
| surfaces.remote_mobile | Mobile, remote-control and voice access | 52.5 |
| surfaces.cloud_sessions | Cloud and remote sandbox execution | 65.5 |
| rel.service_errors | Outages, server errors and capacity or rate errors | 7.5 |
| rel.response_speed | Latency, throughput and fast mode | 36.9 |
| rel.client_failures | Client crashes, freezes and failed tool execution | 7.1 |
| rel.update_breakage | Updates break working setups | 15.3 |
| account.support | Support, refunds and issue handling | 16.3 |
| account.billing_errors | Wrong charges, failed payments and plan provisioning | 2.4 |
| account.bans_restrictions | Account bans and access restrictions | 6.4 |
| account.data_privacy | Data retention, training use and deployment isolation | 20.1 |

## Criteria

63 criteria in 9 areas. Codebook: every criterion names one product mechanism, visible in a single post.

### Paying and limits (`paying`)

What do you pay, and how far does it get you?

- [`limits.plan_value`](https://feedbackbench.com/criteria/limits.plan_value.md) **How much use a plan's price buys.** Whether a plan's weekly or monthly allowance covers the user's normal work, and how its price compares with other plans or agents for the use it gives. Covers running out days early, never hitting the cap, missing tiers and value per dollar. Boundary: Not this: see limits.window_interrupts_work when the short multi-hour window blocks work. Not this: see limits.burn_rate when one prompt, model or effort level drained the quota. Not this: see limits.allowance_change when price or allowance changed over time. Not this: see billing.pricing_clarity when the information itself is unclear.
- [`limits.window_interrupts_work`](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) **Short rolling usage window blocks or interrupts work.** The short rolling usage window (a cap of a few hours) blocks the user mid-task or mid-session: forced waits, killed sessions, or a long allowance left unused because the short window ran out. The post is about the window itself, not about what consumed it. Boundary: Not this: see limits.burn_rate when the point is that one prompt, model or effort level used a large share. Not this: see limits.plan_value for the weekly or monthly total.
- [`limits.burn_rate`](https://feedbackbench.com/criteria/limits.burn_rate.md) **Single prompt, model or effort level consumes disproportionate quota.** How much quota a specific prompt, task, model version, effort level or feature consumes for the work done. Includes new versions using more than earlier ones and praise for efficient ones. The post names what consumed the quota. Boundary: Not this: see limits.window_interrupts_work when the post is about being blocked by the short window, not about the consumer. Not this: see limits.prompt_cache for cache misses. Not this: see work.multi_agent_orchestration when runaway subagents cause the burn.
- [`limits.allowance_change`](https://feedbackbench.com/criteria/limits.allowance_change.md) **Price, allowance or plan terms changed.** The vendor changes a plan's price, quota, multipliers, included models or tiers over time: cuts, raises, removals, silent changes and changes that differ from what was announced. Boundary: Not this: see limits.plan_value for the current level with no change named. Not this: see limits.reset_schedule for when resets happen.
- [`limits.reset_schedule`](https://feedbackbench.com/criteria/limits.reset_schedule.md) **Quota reset timing and bonus or banked resets.** When and how usage resets happen: scheduled reset dates, surprise or bonus resets, banked resets, and paid resets. Covers resets that wipe saved allowance or land uselessly close to the normal reset. Boundary: Not this: see limits.allowance_change for changes to quota size. Not this: see limits.usage_meter when the reset is only displayed wrongly.
- [`limits.usage_meter`](https://feedbackbench.com/criteria/limits.usage_meter.md) **Usage meter visibility and accuracy.** Whether the product shows how much quota and tokens were used and how much remains, per task and per model, and whether that figure is accurate. Covers meters that climb while idle and hidden token counts. Boundary: Not this: see billing.overage_charges for money actually charged. Not this: see limits.burn_rate when the meter is believed and the complaint is the amount consumed.
- [`limits.prompt_cache`](https://feedbackbench.com/criteria/limits.prompt_cache.md) **Prompt cache hits, misses and invalidation.** Whether prompt caching is kept across turns, resumes and provider routing, and how cache reads are priced. Covers cache drops that inflate usage. Boundary: Not this: see limits.burn_rate for consumption not attributed to caching.
- [`billing.overage_charges`](https://feedbackbench.com/criteria/billing.overage_charges.md) **Pay-as-you-go overage, fallback billing and spend caps.** Usage spills from the plan into metered, on-demand or API billing, and whether spend caps and alerts work. Covers overage turning on silently and credit top-ups. Boundary: Not this: see account.billing_errors for wrong charges on the subscription itself. Not this: see models.routing_auto when the cause is auto-selection of a pricier model.
- [`billing.pricing_clarity`](https://feedbackbench.com/criteria/billing.pricing_clarity.md) **Pricing and plan terms stated clearly and consistently.** Whether pricing pages, docs and plan descriptions clearly and consistently state costs, limits and what 'unlimited' means. Boundary: Not this: see limits.allowance_change when terms actually changed. Not this: see limits.usage_meter for in-product usage display.
- [`billing.free_tier`](https://feedbackbench.com/criteria/billing.free_tier.md) **Free tier and free model availability and limits.** Whether free models, free tiers and trial or promo access exist, how long they stay available, and whether their limits are usable. Boundary: Not this: see models.catalog_access for paid model availability. Not this: see rel.service_errors when free models return errors.
- [`billing.subscription_portability`](https://feedbackbench.com/criteria/billing.subscription_portability.md) **Using an existing subscription across tools.** Whether a paid subscription from one vendor can be used inside another agent or harness, or used as an API outside the vendor's own app. Boundary: Not this: see setup.provider_byok_local for API keys, local models and custom endpoints.

### Setting up and connecting (`setup`)

How hard is it to install and connect?

- [`setup.install_signin`](https://feedbackbench.com/criteria/setup.install_signin.md) **Install, launch and sign-in.** Getting the agent installed, started and signed in on the user's platform: installers, updates on first run, login flows, browser redirects and switching accounts. Boundary: Not this: see setup.provider_byok_local for connecting own keys or local models. Not this: see rel.update_breakage when a later update breaks a working setup.
- [`setup.provider_byok_local`](https://feedbackbench.com/criteria/setup.provider_byok_local.md) **Connecting own API keys, local models and custom endpoints.** Whether user-supplied providers work: bring-your-own-key, OpenAI-compatible endpoints, OpenRouter, and local servers such as Ollama or LM Studio. Boundary: Not this: see billing.subscription_portability for reusing a paid subscription. Not this: see rel.tool_call_errors for tool-format failures once connected.
- [`setup.extensions_mcp`](https://feedbackbench.com/criteria/setup.extensions_mcp.md) **MCP servers, plugins, skills and hooks.** Whether external tools and extension points load, authenticate and refresh correctly: MCP servers, plugins, skills, hooks and third-party integrations. Boundary: Not this: see setup.ide_integration for editor extensions. Not this: see context.instruction_files for rules files.
- [`setup.onboarding_docs`](https://feedbackbench.com/criteria/setup.onboarding_docs.md) **Onboarding, discoverability and documentation.** How easily a new user learns the product: quick-start guides, docs explaining features and modes, and how findable features are. Boundary: Not this: see billing.pricing_clarity for pricing docs. Not this: see ui.customization for settings complexity.
- [`setup.ide_integration`](https://feedbackbench.com/criteria/setup.ide_integration.md) **IDE and editor integration.** How the agent works inside an editor or IDE, including extension support, editor-native features, autocomplete and feature parity across IDEs. Boundary: Not this: see verify.change_review_ui for diff and approval views. Not this: see ui.terminal_display for CLI rendering.

### Choosing models (`models`)

Which models do you get, and do they hold up?

- [`models.catalog_access`](https://feedbackbench.com/criteria/models.catalog_access.md) **Which models are offered on a plan and when.** Whether models are available on the user's plan: same-day support for new releases, deprecations and removals, multi-provider choice, and models missing from a surface. Boundary: Not this: see billing.free_tier for free-model availability. Not this: see models.routing_auto for which model actually runs.
- [`models.routing_auto`](https://feedbackbench.com/criteria/models.routing_auto.md) **Automatic model routing and fallback.** Auto mode, routers or fallbacks pick, switch or hide the model used. This includes subagents running on a model other than the one requested, silent downgrades, and inability to exclude models. Boundary: Not this: see models.effort_control for reasoning level. Not this: see billing.overage_charges for the resulting charge alone.
- [`models.effort_control`](https://feedbackbench.com/criteria/models.effort_control.md) **Reasoning effort setting and its defaults.** How the reasoning or effort level can be set, whether it stays set, and how outcomes differ by level, such as overthinking at high effort or failing at low effort. Boundary: Not this: see limits.burn_rate when the point is the quota share an effort level consumed. Not this: see work.scope_overreach for over-engineered output.
- [`models.quality_drift`](https://feedbackbench.com/criteria/models.quality_drift.md) **Quality got worse or better over time.** The post compares the agent or one of its models with an earlier time and says it got worse or better: 'nerfed', 'dumber since last week', 'better than at launch'. An explicit comparison with the past is required. Boundary: Not this: see general.unspecific for a verdict with no comparison over time. Not this: see work.capability for what it can do now. Not this: see limits.allowance_change for changed quotas or prices.

### Instructing and context (`context`)

Does it follow your instructions and keep the right context?

- [`context.instruction_files`](https://feedbackbench.com/criteria/context.instruction_files.md) **Persistent project rules files are read and obeyed.** Whether the agent reads and follows persistent project instruction files (rules or agent markdown files) across turns. Boundary: Not this: see context.instruction_following for instructions in the current prompt. Not this: see context.session_memory for auto-written memory.
- [`context.instruction_following`](https://feedbackbench.com/criteria/context.instruction_following.md) **Direct in-prompt instructions and caps are followed.** Whether the agent executes explicit instructions, prohibitions and budgets given in the prompt, such as iteration caps and token caps. Boundary: Not this: see context.instruction_files for rules files. Not this: see work.sycophancy_pushback for evaluating the user's claims or opinions. Not this: see work.scope_overreach for unrequested extras.
- [`context.clarifying_questions`](https://feedbackbench.com/criteria/context.clarifying_questions.md) **Asks the user versus guessing.** Whether the agent stops to ask clarifying questions when it is blocked or the request is ambiguous, instead of inventing assumptions or workarounds. Boundary: Not this: see work.stuck_loops for repetitive failure without any question.
- [`context.long_context_decay`](https://feedbackbench.com/criteria/context.long_context_decay.md) **Output degrades as the context window fills.** The agent forgets details, rules or its own statements as the session grows long. Boundary: Not this: see context.compaction for losses caused by summarisation. Not this: see context.session_memory for loss across separate sessions.
- [`context.compaction`](https://feedbackbench.com/criteria/context.compaction.md) **Context compaction keeps what matters, cheaply and quickly.** How automatic or manual compaction summarises history: what it drops, how long it takes, what it costs, and whether it is visible. Boundary: Not this: see context.long_context_decay for degradation without compaction.
- [`context.session_memory`](https://feedbackbench.com/criteria/context.session_memory.md) **Memory and state carried across sessions.** Whether knowledge persists correctly between sessions: memory files, re-discovery cost at session start, stale memories, and leakage between sessions. Boundary: Not this: see ui.session_history for saving and resuming chat transcripts. Not this: see context.instruction_files for user-written rules.
- [`context.codebase_retrieval`](https://feedbackbench.com/criteria/context.codebase_retrieval.md) **Finding the right files in the codebase.** How the agent searches and indexes a repository and pulls in relevant files. Covers over-reading on trivial edits and missing files in large or multi-repo codebases. Boundary: Not this: see limits.burn_rate for consumption not tied to file reading.
- [`context.attachments`](https://feedbackbench.com/criteria/context.attachments.md) **Images, PDFs and file attachments as input.** Whether attached images, screenshots, PDFs and other files are accepted and read. Boundary: Not this: see work.computer_browser_use for screen control.

### Doing the work (`work`)

How does it behave while it works?

- [`work.capability`](https://feedbackbench.com/criteria/work.capability.md) **Can do the user's kind of task.** Whether the agent, with its models, succeeds at the user's kind of task: its size, complexity, language or domain, including building a whole feature in one go or failing at it. The post names the task or its size but no narrower failure behaviour. Boundary: Not this: see work.frontend_ui for visual UI work and work.bug_diagnosis for fixing a reported bug. Not this: see models.quality_drift for a change over time. Not this: see the other work.* and verify.* leaves when a specific behaviour (looping, scope, false claims, breaking code) is named. Not this: see general.unspecific when no task is named.
- [`work.frontend_ui`](https://feedbackbench.com/criteria/work.frontend_ui.md) **Frontend and visual UI output.** How the agent handles UI design, layout, visual taste and front-end component wiring. Boundary: Not this: see work.regressions_introduced for UI bugs reintroduced after fixes.
- [`work.bug_diagnosis`](https://feedbackbench.com/criteria/work.bug_diagnosis.md) **Diagnosing and fixing reported bugs.** Whether the agent finds the root cause of a failing behaviour and fixes it without step-by-step guidance. Boundary: Not this: see verify.agent_code_review for reviewing code to find unknown issues.
- [`work.regressions_introduced`](https://feedbackbench.com/criteria/work.regressions_introduced.md) **Breaks existing code or reintroduces bugs.** Edits break previously working code, reintroduce fixed bugs, or cycle between introducing and fixing bugs. Boundary: Not this: see work.stuck_loops for repetition without code damage. Not this: see work.reward_hacking for deliberately gaming checks.
- [`work.scope_overreach`](https://feedbackbench.com/criteria/work.scope_overreach.md) **Does unrequested work or over-engineers.** The agent edits outside the requested scope, adds unasked features, tests or runs, or builds overly complex solutions. Boundary: Not this: see work.destructive_actions for risky irreversible actions. Not this: see work.response_verbosity for padded text or comments.
- [`work.stuck_loops`](https://feedbackbench.com/criteria/work.stuck_loops.md) **Spins, loops or gets stuck without progress.** The agent repeats attempts, loops endlessly or spirals on a problem without converging. Boundary: Not this: see work.premature_stop for stopping early. Not this: see context.clarifying_questions when the fix was to ask the user.
- [`work.premature_stop`](https://feedbackbench.com/criteria/work.premature_stop.md) **Stops mid-task or answers instead of acting.** The agent halts before finishing, or replies with explanation or a summary instead of performing the action. Boundary: Not this: see verify.false_completion for stopping while claiming the work is done. Not this: see limits.window_interrupts_work for stops caused by quota.
- [`work.long_running_autonomy`](https://feedbackbench.com/criteria/work.long_running_autonomy.md) **Long unattended runs and goal/loop mode.** Whether the agent sustains hours-long autonomous work toward a goal, and whether a goal or loop mode exists. Boundary: Not this: see work.multi_agent_orchestration for coordination between agents. Not this: see surfaces.cloud_sessions for where the run executes.
- [`work.multi_agent_orchestration`](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) **Subagents, parallel agents and orchestrators.** How the agent spawns, delegates to, monitors and coordinates subagents or parallel sessions. Covers polling loops and the wrong choice of subagents. Boundary: Not this: see models.routing_auto for the model a subagent runs on. Not this: see work.git_workflow for merge conflicts between sessions.
- [`work.reward_hacking`](https://feedbackbench.com/criteria/work.reward_hacking.md) **Games checks instead of fixing the problem.** The agent disables linters, hardcodes for tests, edits benchmarks or defends bugs with tests so that checks pass without a real fix. Boundary: Not this: see verify.false_completion for plain false claims without manipulated checks.
- [`work.destructive_actions`](https://feedbackbench.com/criteria/work.destructive_actions.md) **Risky or irreversible actions without confirmation.** The agent pushes, merges, deletes or wipes changes, or acts on the host outside its sandbox, without asking. Also covers praise for staying in bounds. Boundary: Not this: see work.permission_prompts for approval prompt frequency. Not this: see work.scope_overreach for harmless extra work.
- [`work.git_workflow`](https://feedbackbench.com/criteria/work.git_workflow.md) **Git commits, branches and sync.** How the agent handles commits, attribution, branches, pulls and pushes, and merge conflicts between concurrent agents or teammates. Boundary: Not this: see work.destructive_actions for data-destroying git commands.
- [`work.computer_browser_use`](https://feedbackbench.com/criteria/work.computer_browser_use.md) **Computer use and browser control.** Whether the agent operates desktop apps and browsers well, without taking over the user's windows. Boundary: Not this: see context.attachments for reading screenshots.
- [`work.safety_refusals`](https://feedbackbench.com/criteria/work.safety_refusals.md) **Safety filters block legitimate coding tasks.** Refusals or safety classifiers stop benign or security-remediation work. Boundary: Not this: see work.permission_prompts for tool approval prompts. Not this: see models.routing_auto for safety-triggered model downgrades.
- [`work.permission_prompts`](https://feedbackbench.com/criteria/work.permission_prompts.md) **Tool approval prompts and autonomy modes.** How often and when the agent asks permission to run tools or commands, and whether auto or bypass modes behave as configured. Boundary: Not this: see work.destructive_actions for acting without permission. Not this: see work.safety_refusals for content blocks.
- [`work.plan_mode`](https://feedbackbench.com/criteria/work.plan_mode.md) **Plan-before-edit mode.** Whether a read-only planning phase exists, produces useful plans, and actually prevents edits. Boundary: Not this: see work.task_completion for how well an approved plan is executed.
- [`work.response_verbosity`](https://feedbackbench.com/criteria/work.response_verbosity.md) **Length and clarity of replies, summaries and comments.** How long and how readable the agent's prose and summaries are, and how many code comments and doc notes it writes. Boundary: Not this: see verify.change_review_ui for the volume of code changes. Not this: see work.scope_overreach for extra code features.
- [`work.sycophancy_pushback`](https://feedbackbench.com/criteria/work.sycophancy_pushback.md) **Caves to or argues with the user's judgement.** How the agent handles disagreement: agreeing with everything, reversing correct findings when challenged, failing to push back on wrong claims, or arguing when corrected. Boundary: Not this: see context.instruction_following for refusing a clear instruction.

### Checking and finishing (`checking`)

Can you trust that the work is done?

- [`verify.false_completion`](https://feedbackbench.com/criteria/verify.false_completion.md) **Claims work is done or fixed when it is not.** The agent reports success, completion or a finished todo list that turns out to be untrue. Boundary: Not this: see work.reward_hacking when checks were manipulated. Not this: see verify.self_testing for whether checks were run at all.
- [`verify.self_testing`](https://feedbackbench.com/criteria/verify.self_testing.md) **Builds, tests or runs its own changes.** The post says whether the agent built, tested, ran or checked its own change before handing it back: it ran the test suite, skipped tests, did not run the app, or verified in proportion to the change. Boundary: Not this: see verify.false_completion when the agent claims success that did not happen. Not this: see verify.agent_code_review for a separate review mode or review agent.
- [`verify.agent_code_review`](https://feedbackbench.com/criteria/verify.agent_code_review.md) **Agent-performed code review finds real issues.** A review mode or review agent finds real defects in code or PRs, without re-flagging code that already passed or producing noise. Boundary: Not this: see work.bug_diagnosis for fixing a known symptom. Not this: see work.sycophancy_pushback for reversing findings under pressure.
- [`verify.change_review_ui`](https://feedbackbench.com/criteria/verify.change_review_ui.md) **Reviewing and approving the agent's changes.** Diff views, per-file approval, edit-acceptance prompts, and how manageable the amount of change is for a human reviewer. Boundary: Not this: see setup.ide_integration for general editor features. Not this: see work.response_verbosity for prose length.

### Interface and sessions (`interface`)

How do you operate and steer it?

- [`ui.display_settings`](https://feedbackbench.com/criteria/ui.display_settings.md) **How the interface shows work, and what the user can configure.** How the CLI, TUI, IDE panel or app shows the agent's activity (tool calls, diffs, logs, subagent progress, scrolling, fonts, layout) and which settings, keybindings, themes and presets the user can change. Boundary: Not this: see limits.usage_meter for usage displays. Not this: see ui.session_history for saving and resuming sessions. Not this: see models.effort_control for reasoning-effort settings.
- [`ui.session_history`](https://feedbackbench.com/criteria/ui.session_history.md) **Saving, switching, resuming and rewinding sessions.** Whether chats and threads persist, can be switched and resumed with their state intact, and can be rewound or reverted. Boundary: Not this: see context.session_memory for knowledge carried into a new session.
- [`ui.interrupt_steer`](https://feedbackbench.com/criteria/ui.interrupt_steer.md) **Stopping and steering a running agent.** Whether the user can stop, interrupt or send guidance mid-run and have the agent respect it and continue. Boundary: Not this: see work.premature_stop for the agent stopping on its own.
- [`surfaces.remote_mobile`](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) **Mobile, remote-control and voice access.** Driving or monitoring agents from a phone, another device or by voice, including approvals from the mobile client. Boundary: Not this: see surfaces.cloud_sessions for where the agent executes.
- [`surfaces.cloud_sessions`](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) **Cloud and remote sandbox execution.** Agents run in hosted or remote environments. Covers persistence when the laptop closes, access to local files, session caps, and start or resume time. Boundary: Not this: see billing.overage_charges for how cloud runs are billed. Not this: see work.long_running_autonomy for run behaviour.

### Reliability and speed (`reliability`)

Does it stay up and respond quickly?

- [`rel.service_errors`](https://feedbackbench.com/criteria/rel.service_errors.md) **Outages, server errors and capacity or rate errors.** Backend unavailability, 5xx or connection errors, capacity errors, and transient 429 errors that do not reflect the user's quota. Boundary: Not this: see limits.window_interrupts_work and limits.allowance_size for plan quota being reached. Not this: see rel.tool_call_errors for harness tool failures.
- [`rel.response_speed`](https://feedbackbench.com/criteria/rel.response_speed.md) **Latency, throughput and fast mode.** How quickly the agent responds and completes work, including paid or fast modes and whether they actually deliver speed. Boundary: Not this: see rel.service_errors for failures. Not this: see rel.client_crash_resources for local slowness from resource use.
- [`rel.client_failures`](https://feedbackbench.com/criteria/rel.client_failures.md) **Client crashes, freezes and failed tool execution.** The local app, CLI or extension crashes, freezes, leaks memory or uses too much CPU, or fails to execute the model's tool calls and shell commands. Boundary: Not this: see rel.service_errors for backend outages and server errors. Not this: see rel.update_breakage when a specific update caused it.
- [`rel.update_breakage`](https://feedbackbench.com/criteria/rel.update_breakage.md) **Updates break working setups.** New releases regress features, break compatibility or change behaviour unexpectedly. Also covers praise for smooth migrations. Boundary: Not this: see limits.allowance_change for quota changes delivered in an update.

### Account and support (`account`)

How does the vendor treat your account?

- [`account.support`](https://feedbackbench.com/criteria/account.support.md) **Support, refunds and issue handling.** Reaching a human, getting refunds and resolutions, and the vendor's responsiveness to bug reports and community issues. Boundary: Not this: see account.billing_errors for the underlying wrong charge.
- [`account.billing_errors`](https://feedbackbench.com/criteria/account.billing_errors.md) **Wrong charges, failed payments and plan provisioning.** Payments are charged incorrectly, plans are not provisioned after payment, proration errors occur, or payments fail. Boundary: Not this: see billing.overage_charges for metered overage. Not this: see account.support for how the vendor handled the error.
- [`account.bans_restrictions`](https://feedbackbench.com/criteria/account.bans_restrictions.md) **Account bans and access restrictions.** The vendor suspends, bans or restricts accounts, for example for heavy usage or regional embargoes. Boundary: Not this: see limits.allowance_size for normal quota lockouts.
- [`account.data_privacy`](https://feedbackbench.com/criteria/account.data_privacy.md) **Data retention, training use and deployment isolation.** Whether prompts and code are retained or used for training, and whether private or air-gapped deployment is available. Boundary: Not this: see billing.free_tier for free-model availability itself.
