# Factory (Factory)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/agent/factory

| Measure | Value |
|---|---|
| Rank | =11 of 17 (rank range 11–12) |
| Feedback Score | 35.6 (95% interval 34.9–36.2) |
| Popularity | 0.232 (share of voice 0.80%) |
| Customer love | 0.545 (95% interval 0.525–0.563) |
| Top quadrant | no |
| Authors | 789 |
| Posts counted | 1576 |
| Posts that judge the agent | 714 |
| Criteria better / worse than peers | 2 / 0 of 63 |

## The brief

Written by Claude Opus 5.5 from 78 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**A harness that makes cheap models shine, fenced in by usage windows.**

TL;DR:

- The harness is the product. Users say cheap and open models perform well inside Droid.
- The $20 plan stretches far on light models, while frontier models drain the window in minutes.
- Payment failures, no third-party subscriptions and a rough desktop app drive most complaints.

### What hurts

- **Frontier models empty the window fast** ([Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md), [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md)). The rolling five-hour window and heavy models collide: a short burst on a frontier model consumes the allowance and stops work until reset.
  Every post about the short window is a complaint. Users report that minutes on Astra wipe the window, and that a fraction of a session eats a large share of the limit. Removing the five-hour window is a standing request.
  
  Missions draw fire too. One user calls the feature the quickest way to burn allowance because it stretches task time without better results.
  Evidence:
  - Complaint, Factory, @droid, 2026-09-05: “10 minutes of astra @droid and the 5hour usage is gone. unusable.” [source](https://twitter.com/1590702228234391552/status/2096151009454444568)
  - Complaint, Factory, @droid, 2026-09-22: “@theo 10 minute of work, exhaust my @droid 's 22% of limit. that's a no go for small devs” [source](https://twitter.com/2995471962/status/2102219071827976480)
  - Complaint, Factory, @droid, 2026-09-15: “@droid the missions feature is the quickest way to use up your usage allowance. however, you can achieve better results without using this feature, as it unnecessarily prolongs the time it takes to complete tasks.” [source](https://twitter.com/1896160400716267520/status/2099650711000682851)
  - Complaint, Factory, @FactoryAI, 2026-09-15: “can anyone at @droid or @factoryai dm with me? i m having serious trouble with the app and 5 hour limits. it s annoying and i need some help please?” [source](https://twitter.com/1834314883510226944/status/2100002663857377638)

- **Your other subscriptions stay locked out** ([Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md), [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md)). Users want to bring ChatGPT or Claude subscriptions into Droid. Without that, the fallback is API pricing or BYOK, which some report as paywalled or buggy.
  Bringing an existing subscription is the top request. Users reject paying API rates for models they already pay for elsewhere. The enterprise Claude marketplace listing earned praise, but one buyer says the listing alone does nothing for committed spend.
  
  BYOK is the escape hatch, and it has cracks. Posts describe it as gated behind a paywall on the Windows desktop app and bugged in the GUI.
  Evidence:
  - Complaint, Factory, @FactoryAI, 2026-09-19: “@mikez93 @tereza_tizkova @januarycomputer @amypretzel @factoryai @droid yea but im not paying api pricing. they need to stop resisting and add support for 3rd party subs” [source](https://twitter.com/1851277456487170050/status/2101341158009864477)
  - Complaint, Factory, @FactoryAI, 2026-09-09: “@factoryai i've wanted anthropic commit spend on factory. marketplace listing alone does nothing.” [source](https://twitter.com/1811332417099055105/status/2097729031957295367)
  - Complaint, Factory, @FactoryAI, 2026-09-21: “@tereza_tizkova @droid @factoryai the time i used droid (desktop app) (on windows), byok was behind a paywall which was a bummer. it would be really nice that if this wasn't the case.” [source](https://twitter.com/1854911029870051328/status/2101899154176061769)
  - Complaint, Factory, @FactoryAI, 2026-09-27: “@droid @factoryai byok is bugged on the gui, but idk where to report this kinda of stuff.” [source](https://twitter.com/4865415939/status/2104338825291891112)

- **Desktop app trails the CLI** ([How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md), [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md)). Users praise the CLI but report bugs in the desktop app, plus a model picker that reorders itself and hides custom models.
  One user sums it up: the CLI is amazing, the app not so much. Complaints cluster around the model picker. A recent update moved recently used models to the top, which confused users, and one post reports custom models vanished from the picker entirely.
  
  Pinning favourite models is a recurring request. Mission rendering in the CLI is also called janky.
  Evidence:
  - Complaint, Factory, @droid, 2026-09-21: “@ain3sh @benvargas @droid hey a fellow droid user here, the cli is amazing, the app not so much, there were some bugs here and there. only thing with the cli is the rendering can be improved, it's a bit janky on missions, maybe you guys can move to a different framework from ink?” [source](https://twitter.com/1333161864/status/2101883259831632365)
  - Complaint, Factory, @droid, 2026-09-27: “the latest version of droid seems to place the recently used models at the top. after the recently used models are at the top, the custom list below no longer has the names of these models, which i find a bit counterintuitive. at first, i couldn't find the model i used and was a bit confused. later, i found it at the top, hahaha @droid” [source](https://twitter.com/1889310672368095232/status/2104031932074107173)
  - Complaint, Factory, @FactoryAI, 2026-09-26: “hey @factoryai, i think custom models are broken in the app right now. i’m unable to see or select any of my custom models from the model picker. they just don’t show up at all, so i can’t use them. not sure if this is a recent regression, but would appreciate a fix. @ross_cefalu” [source](https://twitter.com/1625280993966923777/status/2103738831778529329)
  - Complaint, Factory, @FactoryAI, 2026-09-23: “@ain3sh @itscynnamoroll @factoryai @namespacelabs on the list of the models, my favorite models should be on the top of the list pinned…” [source](https://twitter.com/17719163/status/2102715691531206957)

- **Payment failures meet slow support** ([Support, refunds and issue handling](https://feedbackbench.com/criteria/account.support.md), [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md)). When checkout fails, users say support blames the bank and goes quiet, and some long-time subscribers say this pushes them toward other tools.
  Most support posts are complaints. One returning subscriber describes months of failed resubscription attempts, with support pointing at a bank that works for other AI tools. Others report unanswered applications and plain requests for someone to DM.
  
  The gap shows on open bugs too. One user says image input with OpenAI-compatible models has been broken for months.
  Evidence:
  - Complaint, Factory, @FactoryAI, 2026-09-25: “@factoryai hi, i applied to the factory guild the first day it opened on august 14th, the first day it opened. i received an application confirmation email, but have not received any response after that.” [source](https://twitter.com/260499727/status/2103306229783183829)
  - Complaint, Factory, @droid, 2026-09-19: “hate to be doing this here, specially since @droid has been one of my favourite products. it hurts, to see all my attempts to try and reach out to your team about a payment issue - only to have little to none support, or atleast a clarity on why i'm not able to proceed with subscribing. i've been using your product for months now, took a pause for a month when i was travelling, and ever since, i've not been able to resubsribe. your team says its a problem with my bank, but i've been using the same route to pay my sub previously, and to claude, supergrok! i think this is my last ever attempt to try and get your attention, before i permanently move away to other harnesses (i feel sad but can't help it anymore)” [source](https://twitter.com/306079362/status/2101350002409066914)
  - Complaint, Factory, @FactoryAI, 2026-09-04: “its crazy how its been months since image support does not work in @factoryai 's harness when using openai comptabile models, and they have still not fixed it. just say you dont give a fuck about users that dont pay you, simple” [source](https://twitter.com/1579709674135621637/status/2095868117230772637)
  - Complaint, Factory, @FactoryAI, 2026-09-15: “can anyone at @droid or @factoryai dm with me? i m having serious trouble with the app and 5 hour limits. it s annoying and i need some help please?” [source](https://twitter.com/1834314883510226944/status/2100002663857377638)

### What works

- **The harness lifts cheap models** ([Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md)). Users credit Droid's harness, not the model, for results. Cheaper and open models behave well inside it, which is why people return.
  Capability praise far outweighs complaints. Users say the harness does a great job with cheaper models and that rival harnesses come nowhere close. One fan says UI output looks better in Droid even with the same model used in other tools. Another singles out the agent readiness system.
  Evidence:
  - Praise, Factory, @droid, 2026-09-13: “@blueemi99 @droid i actually started using it again and it's pretty good. i'm using cheaper models and the harness does a great job.” [source](https://twitter.com/20471332/status/2099226981892042779)
  - Praise, Factory, @FactoryAI, 2026-09-03: “@squashy_cake @totallynotparth @factoryai idk answer to that, but i know one thing ollama's harness can nowhere come close to how good droid/factory is” [source](https://twitter.com/953633711605633024/status/2095472288174948550)
  - Praise, Factory, @droid, 2026-09-13: “@blueemi99 @droid they have, in my opinion, the best harness to use open source models in.” [source](https://twitter.com/2684958913/status/2099176954620776619)
  - Praise, Factory, @FactoryAI, 2026-09-17: “only downside is the 5 hours limit and the price which is pretty fair but still a little for me personally. other than that i can list so many things i love about it. droid is super efficient and often finish tasks faster than most other agent with similar results. i feel like it gets the right context at the right time. it’s pretty amazing. also love the byok, live the fact that ui almost always looks better when done with droid even using the same model in other harnesses.. i really am a fan of the product. 😅” [source](https://twitter.com/1617212256487411712/status/2100703532785541412)

- **The $20 plan goes far** ([How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md), [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md)). Pick light or open models for routine work and save frontier models for review, and users say the entry plan feels close to unlimited.
  Plan value splits, and model choice explains the split. Users running GLM flash or Luna say hours of work barely dent usage. One user does detailed design-system work on Opus within the entry plan. Another pairs a flash model for coding with Astra or Fable for architecture and review. The heavy-model complaints come from people who skip that split.
  Evidence:
  - Praise, Factory, @droid, 2026-09-22: “i second this. @droid has very good limits for its $20/month plan. i'd say it's arguably the best you can get for $20 because 1) access to a wide variety of models and 2) great limits if you use it smartly. <strict_link>” [source](https://twitter.com/1817155037938036736/status/2102486847402496045)
  - Praise, Factory, @droid, 2026-09-22: “@droid and i say this because i use it! i've been using claude opus 5.5 via @droid, and the limits are absolutely great with it. i've managed to redo micro details and also am in the middle of organizing the design system with it all while on the $20 plan. <strict_link>” [source](https://twitter.com/1817155037938036736/status/2102487691120239039)
  - Praise, Factory, @FactoryAI, 2026-09-27: “@droid @factoryai it's really good, i use with luna and it's almost unlimited, also really good i really like and i have the $20 plan imagine the $200” [source](https://twitter.com/1692829692200464384/status/2104241083777794207)
  - Praise, Factory, @droid, 2026-09-21: “using glm-5.3-flash (max) in @droid as a daily driver, hours of work and it barely makes a dent in my usage. and more importantly, it just works!” [source](https://twitter.com/783307235518578690/status/2101980360729121137)

- **Model-agnostic, with new models landing fast** ([Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md), [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md)). Droid offers frontier and open models in one harness, and users say new releases arrive quickly and can be switched mid-session.
  Users call it the most model-agnostic agent around. They value switching between frontier and open-source models in a single session without changing tools. Automatic routing gets credit for saving the time spent switching by hand.
  Evidence:
  - Praise, Factory, @FactoryAI, 2026-09-04: “@jadmadi @factoryai @droid it’s great - model agnostic, new models land in @droid real fast. great auto-model option, it’s just a better general code harness all around switching between frontier/oss models in a single session in a single harness is a win” [source](https://twitter.com/1681456803832561664/status/2095943170840478145)
  - Praise, Factory, @droid, 2026-09-14: “@kimnoel @droid i am used to the factory app and droid, and it’s a very good agent / harness. it’s also the most model agnostic. i find their mission feature is better then /goal in codex. if i have time in the future, i will give a try to zcode” [source](https://twitter.com/1954882023769944064/status/2099604737389735990)
  - Praise, Factory, @FactoryAI, 2026-09-11: “@tereza_tizkova @factoryai this set of tools has a clear division of labor, and the automatic routing saves a lot of switching time.” [source](https://twitter.com/1518315606830829568/status/2098501911498702967)
  - Praise, Factory, @FactoryAI, 2026-09-01: “@factoryai any model anytime, total freedom and unbeatable value!” [source](https://twitter.com/2008812694628175872/status/2094694222452683039)

- **Missions and parallel agents feel mature** ([Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md), [Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md)). Users see Mission Control as an early, more complete take on multi-agent orchestration, with observable parallel agents and a plan-review flow.
  Praise outnumbers complaints here. Users point to persistent project context, coordinated parallel agents and auditable messages. One says rivals are only now arriving where Factory already was.
  
  The friction is real but narrow. Parallel agents edited the same file until a user gave each its own scope. Another reports the orchestrator getting stuck in planning.
  Evidence:
  - Praise, Factory, @FactoryAI, 2026-09-22: “@harrystuck77 @factoryai @droid @badlogicgames @ampcode @cursor_ai @opencode @anthropicai @claudeai @cognition @xai @build grok build sitting in c is fair for now. still early days, but the parallel subagents and plan-review flow are already solid for real engineering work. thanks for the ranking.” [source](https://twitter.com/1720665183188922368/status/2102486379972198418)
  - Praise, Factory, @FactoryAI, 2026-09-18: “a lot of the new projects / multi-agent workflows being hyped across claude and cursor are basically the same broader direction @factoryai took with mission control much earlier. different naming. different ux. same core idea: persistent project context, parallel agents, orchestration, and visibility across their work. and in a lot of ways, mission control still feels more complete. factory was seriously ahead of the curve here.” [source](https://twitter.com/1625280993966923777/status/2100945868757340278)
  - Complaint, Factory, @FactoryAI, 2026-09-20: “@hataiit9x @droid @factoryai the parallel agents part is what got me too. mine kept editing the same file until i gave each one its own scope.” [source](https://twitter.com/2079331237991428096/status/2101724074552516749)
  - Complaint, Factory, @droid, 2026-09-03: “@droid i am still looking forward to have a "/goal" command in factory, i personally don' feel "/missions" is doing the job. sometimes i just need a smaller model to reach a goal and not have 3 different ones and that the orchestrator gets stuck into "planning" and i have to intervene” [source](https://twitter.com/231189959/status/2095381035651379575)

### Under the surface

- **The router is a trust question** ([Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md), [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md)). Users like automatic routing in principle but question how it classifies tasks. One user reports a silent switch to Opus that burned quota without notice.
  Router praise is real, but the skeptics are specific. Some users want the router unbundled from the harness. Others ask how task difficulty is judged before work starts and point to cache invalidation when models switch mid-task. A silent upgrade to an expensive model also ties back to the usage-window complaints.
  Evidence:
  - Complaint, Factory, @FactoryAI, 2026-09-16: “the router should not be tied to the harness, just like the harness should not be tied to the inference. @factoryai bundling three things that should be unbundled. <strict_link>” [source](https://twitter.com/2031128433493946368/status/2100050938207801424)
  - Complaint, Factory, @FactoryAI, 2026-09-16: “@tereza_tizkova @aiandcloud @factoryai the problem with factory is that the desktop version is too rough, the miss mode is too slow, and sometimes even when the miss model is configured, it automatically switches to opus without me noticing, causing me to run a lot of traffic. however, i think factory provides a sufficient quota, and the execution effect is good, especially when using open-source models, which are very durable. i just don't know when deepseek 4.1 flash will be available, and also why 5.3 flash does not support multi-modal.” [source](https://twitter.com/1824388694985543680/status/2100057750835445861)
  - Complaint, Factory, @FactoryAI, 2026-09-26: “the transit station really has no work to do. i have always been curious about the router that @factoryai has been promoting. they have never clearly explained the specific router mechanism, just vaguely mentioning the classification of simple and complex tasks. how can we determine whether a task is complex or not before it starts, and the inevitable cache invalidation caused by switching models midway is a big problem (anyway, the transit station makes money regardless, lol). a previous idea was to train a small ml model for task classification, and then jev came out, but is this model really suitable for the task difficulty classification work? a big question mark needs to be placed on that.” [source](https://twitter.com/1724973133097291776/status/2103781371772809499)
  - Complaint, Factory, @FactoryAI, 2026-09-25: “@factoryai what about efficiency? if i'm sure the best code is going to be made with opus 5.5 are you sure the router is going to give me what i want with a great code?” [source](https://twitter.com/1738636938616115200/status/2103335310427828449)

- **Model requests outrun the catalog** ([Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md)). The catalog is a strength, so gaps get noticed fast. Users keep asking for DeepSeek V4.1 Flash, Muse Spark and Qwen, and report missing plan access.
  Several users say they have asked for DeepSeek V4.1 Flash for weeks. Others report that no one on their account can use a model their plan should cover. Because breadth is the reason people choose Droid, each missing model weakens its core pitch.
  Evidence:
  - Complaint, Factory, @FactoryAI, 2026-09-22: “@tereza_tizkova @factoryai @droid i have been trying to ask you about deepseek 4.1 flash being offered for weeks” [source](https://twitter.com/1258455699073441793/status/2102454686070571282)
  - Complaint, Factory, @FactoryAI, 2026-09-26: “@droid @factoryai deepseek v4.1’s been chilling in the library for a while now; it just takes an update to peek at its smarts. i always prefer checking my models manually, it keeps me sharp and slightly mysterious when people ask where i found that gem.” [source](https://twitter.com/417508671/status/2103971431017042421)
  - Complaint, Factory, @FactoryAI, 2026-09-04: “@factoryai i was using plus and max, but no one in my account can use gpt-6 astra 😂😅” [source](https://twitter.com/2079628486055178240/status/2095732701659852808)
  - Complaint, Factory, @FactoryAI, 2026-09-16: “@tereza_tizkova @aiandcloud @factoryai the problem with factory is that the desktop version is too rough, the miss mode is too slow, and sometimes even when the miss model is configured, it automatically switches to opus without me noticing, causing me to run a lot of traffic. however, i think factory provides a sufficient quota, and the execution effect is good, especially when using open-source models, which are very durable. i just don't know when deepseek 4.1 flash will be available, and also why 5.3 flash does not support multi-modal.” [source](https://twitter.com/1824388694985543680/status/2100057750835445861)

### Fine print

- Nearly all posts come from X mentions of the vendor's own handles. Reddit and G2 contribute only a few dozen posts.
- Most criteria have too few posts to rank alone. Read the window, support and BYOK cards as directional.

## Top requests

What users ask to add or change, most asked first. 183 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Criterion | Author-weeks | Posts |
|---|---|---|---|---|
| 1 | Bring existing subscription into this agent | [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | 12 | 12 |
| 2 | Add DeepSeek V4.1 Flash model | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 7 | 7 |
| 3 | Add Muse Spark models | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 6 | 6 |
| 4 | Official dedicated mobile app | [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | 6 | 6 |
| 5 | Remove the 5-hour usage window | [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | 6 | 6 |
| 6 | Vision support for specific models | [Images, PDFs and file attachments as input](https://feedbackbench.com/criteria/context.attachments.md) | 4 | 4 |
| 7 | Add Qwen models | [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | 3 | 4 |
| 8 | Fix declined card payments and checkout failures | [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | 3 | 4 |
| 9 | Higher overall usage limits | [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | 3 | 4 |
| 10 | Linux desktop app and support | [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | 3 | 4 |
| 11 | Pin, sort and hide models in picker | [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | 3 | 4 |
| 12 | Escalation rate metric for routing | [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | 3 | 3 |

### 1. Bring existing subscription into this agent

- Factory, 2026-09-21, @FactoryAI (X): “@droid @melvindvivas @factoryai curious to compare the harnesses. can i use my existing codex sub with droid, or would i need a separate subscription to try it out?” [source](https://twitter.com/2066846062623518720/status/2101942678174793859)
- Factory, 2026-09-21, @FactoryAI (X): “@droid @deepusleepy @factoryai can i bring my codex plan or it’s vía api ?” [source](https://twitter.com/3781517712/status/2101890332640092393)
- Factory, 2026-09-20, @droid (X): “@ain3sh @droid yeah, i know all about cliproxyapi... shared the first guides on setting it up. was just hoping you guys might have added it 1st party like amp since promoting it. <strict_link>” [source](https://twitter.com/89291422/status/2101812221697499277)

### 2. Add DeepSeek V4.1 Flash model

- Factory, 2026-09-22, @FactoryAI (X): “@tereza_tizkova @factoryai @droid i have been trying to ask you about deepseek 4.1 flash being offered for weeks” [source](https://twitter.com/1258455699073441793/status/2102454686070571282)
- Factory, 2026-09-16, @droid (X): “@droid when are you guys planning to introduce deepsek v4.1 flash?” [source](https://twitter.com/1732405334487052288/status/2100221149082698170)
- Factory, 2026-09-23, @FactoryAI (X): “why no deepseek v1 flash on @droid ? @factoryai” [source](https://twitter.com/1251552496893325312/status/2102759494090674627)

### 3. Add Muse Spark models

- Factory, 2026-09-21, @FactoryAI (X): “@factoryai thanks but can we have muse as well please? and qwen flash series?” [source](https://twitter.com/70830663/status/2102177611451716065)
- Factory, 2026-09-20, @droid (X): “i can use for example muse 1.3 spark free in open code to test out what open code harness can do - so if lets's say factory does a better job i'd like to test it out but there's no muse 1.3 spark free in droid so i'd likely stay on code or deep seek flash 4.1. how do i get hooked ?” [source](https://twitter.com/3781517712/status/2101601490615861536)
- Factory, 2026-09-05, @FactoryAI (X): “@factoryai great job on model integration; any chance you will add muse spark 1.3 in the near future ?” [source](https://twitter.com/1612797507007877121/status/2096227746292842859)

### 4. Official dedicated mobile app

- Factory, 2026-09-27, @FactoryAI (X): “@droid @factoryai my nits are all qol - droid mobile app - windows and mac dev environments for bots (using namespace devboxes in the meantime) - handoff between local and cloud agents - extend droid cloud agent access, we are limited to 40h/month while competing products are unlimited” [source](https://twitter.com/2070908287978246144/status/2104255956439961935)
- Factory, 2026-09-18, @FactoryAI (X): “i think droid is pretty neat, i think a metered on-demand cloud computer would be great (instead of the current flat $200 for anything in the cloud), along with a native mobile app. for what it is (industry focused coding agent) it does everything near perfectly which is why it’s top of a tier” [source](https://twitter.com/1892774905789501441/status/2100943506449826040)
- Factory, 2026-09-14, @droid (X): “@anasibnanwar @droid i mean.. if we can choose, you could build it for us 😎” [source](https://twitter.com/1659223500417007616/status/2099640365083242518)

### 5. Remove the 5-hour usage window

- Factory, 2026-09-25, @FactoryAI (X): “i dont get it, why isn't @factoryai removing the 5h limit? it s an api wrapper harness.” [source](https://twitter.com/1834314883510226944/status/2103441639041540512)
- Factory, 2026-09-21, @FactoryAI (X): “@droid @melvindvivas @factoryai so are you going to cancel the 5 hour limit?” [source](https://twitter.com/714294369331707904/status/2101910616378409377)
- Factory, 2026-09-08, @FactoryAI (X): “@droid @factoryai usage reset! or you can just remove the 5h window” [source](https://twitter.com/1090325687448281093/status/2097416099201474962)

### 6. Vision support for specific models

- Factory, 2026-09-27, @FactoryAI (X): “@droid @factoryai why glm 5.3 flash on droid doesn't support image modalities?” [source](https://twitter.com/1181249614550192132/status/2104044449169072295)
- Factory, 2026-09-04, @FactoryAI (X): “its crazy how its been months since image support does not work in @factoryai 's harness when using openai comptabile models, and they have still not fixed it. just say you dont give a fuck about users that dont pay you, simple” [source](https://twitter.com/1579709674135621637/status/2095868117230772637)
- Factory, 2026-09-03, @droid (X): “@droid why can't i use vision with glm-5.3-flash (droid core) on the factory desktop (mac)? when the model itself supports vision 🙏” [source](https://twitter.com/1205089206701195264/status/2095354646147559507)

### 7. Add Qwen models

- Factory, 2026-09-10, @droid (X): “@droid when qwen on droid？😈” [source](https://twitter.com/1186642435365052417/status/2098032368225509769)
- Factory, 2026-09-04, @droid (X): “@droid when we get qwen3.8-27b ?” [source](https://twitter.com/2020825549367566336/status/2095911044203868248)

### 8. Fix declined card payments and checkout failures

- Factory, 2026-09-27, @FactoryAI (X): “@droid @factoryai let your users pay you!! been stuck with a failed payment issue since august :(” [source](https://twitter.com/306079362/status/2104243378246578672)

### 9. Higher overall usage limits

- Factory, 2026-09-21, @FactoryAI (X): “factory droid has some of the worst usage limits that i've ever encountered. i like it as a harness but the usage limits are absolutely terrible @factoryai” [source](https://twitter.com/1910134991138459648/status/2102116515285754324)
- Factory, 2026-09-18, @FactoryAI (X): “@droid @theterrancex @factoryai icrease 20 dolar and all plan limits, especially droid core limits” [source](https://twitter.com/1896160400716267520/status/2101087661503152393)
- Factory, 2026-09-22, @droid (X): “@droid if only limits were better but its solid.” [source](https://twitter.com/2007912122043568128/status/2102495905077461272)

### 10. Linux desktop app and support

- Factory, 2026-09-27, @FactoryAI (X): “@droid @factoryai here's what's sticking out right now - no droid ios app - desktop app missing from linux - syncing missions / repos / active work between droid computers (my own not droid managed). i need to be able to shift my coding workloads off my laptop and take them everywhere with me.” [source](https://twitter.com/1382136217601417222/status/2104238899493318925)
- Factory, 2026-09-22, @FactoryAI (X): “@factoryai any plans for a linux / omarchy desktop release? i love the desktop app but am on omarchy :(” [source](https://twitter.com/1683371095595057152/status/2102469461366411572)
- Factory, 2026-09-08, @FactoryAI (X): “@droid @factoryai ubuntu app” [source](https://twitter.com/1982985165040431106/status/2097327584875098610)

### 11. Pin, sort and hide models in picker

- Factory, 2026-09-27, @droid (X): “the latest version of droid seems to place the recently used models at the top. after the recently used models are at the top, the custom list below no longer has the names of these models, which i find a bit counterintuitive. at first, i couldn't find the model i used and was a bit confused. later, i found it at the top, hahaha @droid” [source](https://twitter.com/1889310672368095232/status/2104031932074107173)
- Factory, 2026-09-26, @FactoryAI (X): “okay until the update today it didn’t show up for me, but i also only use cli so the model selector is a little different. if i could off a suggestion, i would like to see 2 columns and you can just tab between the two and one side is dedicated to all the droid core models, would help me see them a lot easier and then maybe a third for the custom models you add” [source](https://twitter.com/587982527/status/2103960664452546913)
- Factory, 2026-09-23, @FactoryAI (X): “@ain3sh @itscynnamoroll @factoryai @namespacelabs on the list of the models, my favorite models should be on the top of the list pinned…” [source](https://twitter.com/17719163/status/2102715691531206957)

### 12. Escalation rate metric for routing

- Factory, 2026-09-24, @FactoryAI (X): “@factoryai a useful routing metric is the escalation rate alongside savings. track quality, latency, cost, and how often a request had to move to a stronger model—otherwise “cheaper” can hide a downgrade.” [source](https://twitter.com/2068128107987632128/status/2103196343137440231)
- Factory, 2026-09-24, @FactoryAI (X): “@factoryai the companion number i'd want next to that, especially at 4x volume: the escalation rate. if savings grew while the escalation rate held flat, that's real routing; if the router just escalates less, that's a downgrade with a nice chart.” [source](https://twitter.com/1588935512135720961/status/2103177541347659985)
- Factory, 2026-09-24, @FactoryAI (X): “@factoryai need the escalation rate next to that 63%” [source](https://twitter.com/1811332417099055105/status/2103203886907732386)

## Facts

| Fact | Value |
|---|---|
| Version | n/a |
| Released | $150M Series C at $1.5B valuation: 2026-04 |
| Price | Pro $20/mo, Plus $100/mo, Max $200/mo, Teams/Enterprise custom (no self-serve, no annual discount) |
| Model | Multi-model routing ('model routing era' positioning) |
| Surface | IDE, terminal, web/cloud |

## Sources

| Channel | Source | Posts |
|---|---|---|
| X | @FactoryAI | 1111 |
| X | @droid | 425 |
| Reddit | r/FactoryAi | 19 |
| Reddit | Posts that name it | 15 |
| G2 | G2 | 6 |

## Better than peers on

Which models are offered on a plan and when, Can do the user's kind of task

## Worse than peers on

None.

## All 63 criteria

Criterion love: 0.5 is the category norm. n: rated author-weeks.

### Paying and limits: Better than peers (customer love 0.542, n 125)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [How much use a plan's price buys](https://feedbackbench.com/criteria/limits.plan_value.md) | Typical | 0.518 | 0.486–0.549 | 70 | 32 | 38 |
| [Single prompt, model or effort level consumes disproportionate quota](https://feedbackbench.com/criteria/limits.burn_rate.md) | Too few posts | 0.533 | 0.501–0.563 | 19 | 8 | 11 |
| [Using an existing subscription across tools](https://feedbackbench.com/criteria/billing.subscription_portability.md) | Too few posts | 0.507 | 0.488–0.527 | 13 | 5 | 8 |
| [Short rolling usage window blocks or interrupts work](https://feedbackbench.com/criteria/limits.window_interrupts_work.md) | Too few posts | 0.483 | 0.471–0.493 | 12 | 0 | 12 |
| [Pricing and plan terms stated clearly and consistently](https://feedbackbench.com/criteria/billing.pricing_clarity.md) | Too few posts | 0.501 | 0.481–0.532 | 10 | 1 | 9 |
| [Usage meter visibility and accuracy](https://feedbackbench.com/criteria/limits.usage_meter.md) | Too few posts | 0.492 | 0.484–0.497 | 6 | 0 | 6 |
| [Free tier and free model availability and limits](https://feedbackbench.com/criteria/billing.free_tier.md) | Too few posts | 0.496 | 0.484–0.508 | 6 | 3 | 3 |
| [Quota reset timing and bonus or banked resets](https://feedbackbench.com/criteria/limits.reset_schedule.md) | Too few posts | 0.495 | 0.489–0.500 | 3 | 0 | 3 |
| [Price, allowance or plan terms changed](https://feedbackbench.com/criteria/limits.allowance_change.md) | Too few posts | 0.497 | 0.493–0.500 | 2 | 0 | 2 |
| [Prompt cache hits, misses and invalidation](https://feedbackbench.com/criteria/limits.prompt_cache.md) | Too few posts | 0.499 | 0.491–0.506 | 2 | 1 | 1 |
| [Pay-as-you-go overage, fallback billing and spend caps](https://feedbackbench.com/criteria/billing.overage_charges.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 |

Most recent posts:

- Praise, 2026-09-27, @FactoryAI (X): “@droid @factoryai it's really good, i use with luna and it's almost unlimited, also really good i really like and i have the $20 plan imagine the $200” [source](https://twitter.com/1692829692200464384/status/2104241083777794207)
- Praise, 2026-09-27, @droid (X): “@droid factory cuts inference cost by precomputing common sub‑expressions and reusing them across requests so each new run only evaluates delta changes” [source](https://twitter.com/195841906/status/2104089474485637594)
- Praise, 2026-09-27, @droid (X): “@droid dying to test droid 👋👋👋🤩🤩 pretty amazing that you have been able to cut interface costs especially in this environment” [source](https://twitter.com/1367818563495596032/status/2104109772974744033)
- Complaint, 2026-09-27, @FactoryAI (X): “@droid @factoryai i will drop grok/cursor for you guys as soon as my subscription ends. @da7_tech convinced me with his post. i just hope you guys beat devin because i can't afford you all 😆 i've tried it before, it was awesome but very expensive. also the most beautiful ui.” [source](https://twitter.com/1406556428840603649/status/2104235627847835986)
- Complaint, 2026-09-27, @FactoryAI (X): “@droid @factoryai my nits are all qol - droid mobile app - windows and mac dev environments for bots (using namespace devboxes in the meantime) - handoff between local and cloud agents - extend droid cloud agent access, we are limited to 40h/month while competing products are unlimited” [source](https://twitter.com/2070908287978246144/status/2104255956439961935)
- Complaint, 2026-09-27, @droid (X): “currently, only big influencers have droid max or devin max @droid @cognition consider me” [source](https://twitter.com/2018156578617090049/status/2104141229428781334)

### Setting up and connecting: Typical (customer love 0.512, n 35)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Connecting own API keys, local models and custom endpoints](https://feedbackbench.com/criteria/setup.provider_byok_local.md) | Too few posts | 0.501 | 0.482–0.519 | 17 | 9 | 8 |
| [MCP servers, plugins, skills and hooks](https://feedbackbench.com/criteria/setup.extensions_mcp.md) | Too few posts | 0.496 | 0.482–0.509 | 8 | 3 | 5 |
| [Onboarding, discoverability and documentation](https://feedbackbench.com/criteria/setup.onboarding_docs.md) | Too few posts | 0.507 | 0.491–0.524 | 8 | 3 | 5 |
| [IDE and editor integration](https://feedbackbench.com/criteria/setup.ide_integration.md) | Too few posts | 0.506 | 0.494–0.520 | 6 | 4 | 2 |
| [Install, launch and sign-in](https://feedbackbench.com/criteria/setup.install_signin.md) | Too few posts | 0.497 | 0.492–0.500 | 2 | 0 | 2 |

Most recent posts:

- Praise, 2026-09-27, @droid (X): “so yeah, you can use custom models in @droid cli, which is great currently trying this setup, is performing great so far ! <strict_link> <strict_link>” [source](https://twitter.com/2028221376092581888/status/2104173498101117143)
- Praise, 2026-09-26, @FactoryAI (X): “@trevorbmurkp @harrystuck77 @factoryai @droid @badlogicgames @ampcode @cursor_ai @opencode @anthropicai @claudeai @cognition @xai @grok @build if ur on mac or windows, the desktop app is great! depends on ur preference” [source](https://twitter.com/1721143727043887104/status/2103639050620145849)
- Praise, 2026-09-25, @FactoryAI (X): “@yuqih @factoryai @gmi_cloud their @droid is my main harness and running on gmi keys for open models” [source](https://twitter.com/1261173216455712768/status/2103332443025719806)
- Complaint, 2026-09-27, @FactoryAI (X): “@droid @factoryai having it for a while doesn't help if discoverability is this bad. most of us never saw it until now.” [source](https://twitter.com/2079846327744401408/status/2104094827193532478)
- Complaint, 2026-09-27, @FactoryAI (X): “@droid @factoryai my nits are all qol - droid mobile app - windows and mac dev environments for bots (using namespace devboxes in the meantime) - handoff between local and cloud agents - extend droid cloud agent access, we are limited to 40h/month while competing products are unlimited” [source](https://twitter.com/2070908287978246144/status/2104255956439961935)
- Complaint, 2026-09-27, @FactoryAI (X): “@anasibnanwar @factoryai 😅 we deffo need a better way to communicate features” [source](https://twitter.com/1721143727043887104/status/2104304405772435681)

### Choosing models: Better than peers (customer love 0.579, n 57)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Which models are offered on a plan and when](https://feedbackbench.com/criteria/models.catalog_access.md) | Better than peers | 0.558 | 0.526–0.587 | 32 | 22 | 10 |
| [Automatic model routing and fallback](https://feedbackbench.com/criteria/models.routing_auto.md) | Too few posts | 0.536 | 0.514–0.560 | 18 | 14 | 4 |
| [Quality got worse or better over time](https://feedbackbench.com/criteria/models.quality_drift.md) | Too few posts | 0.506 | 0.491–0.523 | 9 | 4 | 5 |
| [Reasoning effort setting and its defaults](https://feedbackbench.com/criteria/models.effort_control.md) | Too few posts | 0.501 | 0.494–0.508 | 2 | 1 | 1 |

Most recent posts:

- Praise, 2026-09-26, @FactoryAI (X): “@factoryai what is this witchcraft? auto model?” [source](https://twitter.com/2023937351815467008/status/2103681115991216612)
- Praise, 2026-09-26, @FactoryAI (X): “@delaanthonio @factoryai while i've been juggling different models 👀 a model-agnostic setup with real sovereignty is what i've needed and i keep returning to it” [source](https://twitter.com/356609569/status/2103698082735243324)
- Praise, 2026-09-26, @FactoryAI (X): “@firstmarkcap @enoreyes @factoryai model agnosticism is a smart move.” [source](https://twitter.com/1846161889719623680/status/2103824564736430242)
- Complaint, 2026-09-26, @FactoryAI (X): “hey @factoryai, i think custom models are broken in the app right now. i’m unable to see or select any of my custom models from the model picker. they just don’t show up at all, so i can’t use them. not sure if this is a recent regression, but would appreciate a fix. @ross_cefalu” [source](https://twitter.com/1625280993966923777/status/2103738831778529329)
- Complaint, 2026-09-26, @FactoryAI (X): “@anasibnanwar @factoryai they are. had to have opus 5.5 noodle an interim fix for me :)” [source](https://twitter.com/2093430933026148352/status/2103779351611551895)
- Complaint, 2026-09-26, @FactoryAI (X): “the transit station really has no work to do. i have always been curious about the router that @factoryai has been promoting. they have never clearly explained the specific router mechanism, just vaguely mentioning the classification of simple and complex tasks. how can we determine whether a task is complex or not before it starts, and the inevitable cache invalidation caused by switching models…” [source](https://twitter.com/1724973133097291776/status/2103781371772809499)

### Instructing and context: Too few posts (customer love 0.503, n 15)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Images, PDFs and file attachments as input](https://feedbackbench.com/criteria/context.attachments.md) | Too few posts | 0.490 | 0.480–0.498 | 5 | 0 | 5 |
| [Finding the right files in the codebase](https://feedbackbench.com/criteria/context.codebase_retrieval.md) | Too few posts | 0.504 | 0.495–0.515 | 4 | 3 | 1 |
| [Direct in-prompt instructions and caps are followed](https://feedbackbench.com/criteria/context.instruction_following.md) | Too few posts | 0.506 | 0.497–0.518 | 3 | 2 | 1 |
| [Asks the user versus guessing](https://feedbackbench.com/criteria/context.clarifying_questions.md) | Too few posts | 0.503 | 0.500–0.511 | 1 | 1 | 0 |
| [Context compaction keeps what matters, cheaply and quickly](https://feedbackbench.com/criteria/context.compaction.md) | Too few posts | 0.503 | 0.500–0.511 | 1 | 1 | 0 |
| [Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md) | Too few posts | 0.497 | 0.491–0.500 | 1 | 0 | 1 |
| [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 |
| [Output degrades as the context window fills](https://feedbackbench.com/criteria/context.long_context_decay.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 |

Most recent posts:

- Praise, 2026-09-27, @droid (X): “first time using anthropic models (opus 5.5). i use @droid. i gave it a task to design a landing page for my current project, and it asked questions i have never seen from an agent before. i hope the result comes out good, but so far, it's really impressive and i might not use openai models, unless they actually have a god response. also can only recommend the factory app. it's genuinely brilliant…” [source](https://twitter.com/1821640621347495936/status/2104230757883388046)
- Praise, 2026-09-23, @droid (X): “@wattenberger @droid does this for all assigned tasks and my entire codebase. it's definitely my favourite part of the workflow.” [source](https://twitter.com/2070908287978246144/status/2102608593648546083)
- Praise, 2026-09-18, @FactoryAI (X): “@theterrancex @droid @factoryai reposcape v0.1 catching scanner and rust bugs before release is solid local map of how a codebase connects is such a useful first cut” [source](https://twitter.com/1675906158304038912/status/2101082167782834261)
- Complaint, 2026-09-27, @FactoryAI (X): “@droid @factoryai why glm 5.3 flash on droid doesn't support image modalities?” [source](https://twitter.com/1181249614550192132/status/2104044449169072295)
- Complaint, 2026-09-23, @FactoryAI (X): “@droid @factoryai glm-5.3-flash supports images, but droid cli 0.224.1 and 0.225.0 mark it as text-only. droid strips attached images before sending the request; the log says “stripped images for non-image model.” i tested that image input works when the capability is enabled locally.” [source](https://twitter.com/1138507200/status/2102568341600878968)
- Complaint, 2026-09-22, G2 (G2): “q: what problems is the product solving and how is that benefiting you? a: factory ai help me reduce manual coding work and save development time. it help with repetitive tasks, debugging and building features faster. i can focus more on important work instead of doing everything manually. it make my daily workflow more easy and productive. q: what do you like best about the product? a: factory ai…” [source](https://www.g2.com/products/factory-ai/reviews/factory-ai-review-13384158)

### Doing the work: Better than peers (customer love 0.577, n 95)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Can do the user's kind of task](https://feedbackbench.com/criteria/work.capability.md) | Better than peers | 0.543 | 0.516–0.570 | 57 | 48 | 9 |
| [Subagents, parallel agents and orchestrators](https://feedbackbench.com/criteria/work.multi_agent_orchestration.md) | Too few posts | 0.514 | 0.497–0.530 | 15 | 12 | 3 |
| [Long unattended runs and goal/loop mode](https://feedbackbench.com/criteria/work.long_running_autonomy.md) | Too few posts | 0.505 | 0.491–0.520 | 12 | 10 | 2 |
| [Spins, loops or gets stuck without progress](https://feedbackbench.com/criteria/work.stuck_loops.md) | Too few posts | 0.494 | 0.489–0.499 | 4 | 0 | 4 |
| [Tool approval prompts and autonomy modes](https://feedbackbench.com/criteria/work.permission_prompts.md) | Too few posts | 0.507 | 0.494–0.523 | 4 | 2 | 2 |
| [Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md) | Too few posts | 0.500 | 0.491–0.508 | 3 | 2 | 1 |
| [Diagnosing and fixing reported bugs](https://feedbackbench.com/criteria/work.bug_diagnosis.md) | Too few posts | 0.505 | 0.500–0.511 | 3 | 3 | 0 |
| [Plan-before-edit mode](https://feedbackbench.com/criteria/work.plan_mode.md) | Too few posts | 0.510 | 0.500–0.521 | 3 | 3 | 0 |
| [Does unrequested work or over-engineers](https://feedbackbench.com/criteria/work.scope_overreach.md) | Too few posts | 0.499 | 0.495–0.500 | 1 | 0 | 1 |
| [Computer use and browser control](https://feedbackbench.com/criteria/work.computer_browser_use.md) | Too few posts | 0.502 | 0.500–0.507 | 1 | 1 | 0 |
| [Length and clarity of replies, summaries and comments](https://feedbackbench.com/criteria/work.response_verbosity.md) | Too few posts | 0.504 | 0.500–0.512 | 1 | 1 | 0 |
| [Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 |
| [Stops mid-task or answers instead of acting](https://feedbackbench.com/criteria/work.premature_stop.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 |
| [Games checks instead of fixing the problem](https://feedbackbench.com/criteria/work.reward_hacking.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 |
| [Risky or irreversible actions without confirmation](https://feedbackbench.com/criteria/work.destructive_actions.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 |
| [Git commits, branches and sync](https://feedbackbench.com/criteria/work.git_workflow.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 |
| [Safety filters block legitimate coding tasks](https://feedbackbench.com/criteria/work.safety_refusals.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 |
| [Caves to or argues with the user's judgement](https://feedbackbench.com/criteria/work.sycophancy_pushback.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 |

Most recent posts:

- Praise, 2026-09-27, @FactoryAI (X): “@render @factoryai agent writes the app, spins the db, deploys it. i just sit there like a decorative readme” [source](https://twitter.com/1330209814790746114/status/2104177796901945611)
- Praise, 2026-09-27, @FactoryAI (X): “@factoryai @fireworksai_hq huge win for legacy code blind spots there are so real” [source](https://twitter.com/2010658787611619328/status/2104199345436312044)
- Praise, 2026-09-27, @FactoryAI (X): “@factoryai @enoreyes agent foxxy just aced its browser test by autonomously uploading a video to youtube—acting just like a human! 🤖🔥 <strict_link> <strict_link>” [source](https://twitter.com/2093686029186396160/status/2104206958949810353)
- Complaint, 2026-09-27, @droid (X): “@droid at least it should be on par with devin” [source](https://twitter.com/1159835302275346433/status/2104318568980689170)
- Complaint, 2026-09-26, @FactoryAI (X): “@hataiit9x @droid @factoryai the harness quite sucks :)” [source](https://twitter.com/1797536317716525056/status/2103754025057595878)
- Complaint, 2026-09-24, @FactoryAI (X): “@factoryai @fireworksai_hq anyone who has migrated old cobol sees it: the model "improves" three lines nobody asked for. boring diff wins.” [source](https://twitter.com/2030549621039349760/status/2103235553651282146)

### Checking and finishing: Too few posts (customer love 0.503, n 5)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Agent-performed code review finds real issues](https://feedbackbench.com/criteria/verify.agent_code_review.md) | Too few posts | 0.503 | 0.500–0.508 | 2 | 2 | 0 |
| [Claims work is done or fixed when it is not](https://feedbackbench.com/criteria/verify.false_completion.md) | Too few posts | 0.499 | 0.496–0.500 | 1 | 0 | 1 |
| [Builds, tests or runs its own changes](https://feedbackbench.com/criteria/verify.self_testing.md) | Too few posts | 0.502 | 0.500–0.506 | 1 | 1 | 0 |
| [Reviewing and approving the agent's changes](https://feedbackbench.com/criteria/verify.change_review_ui.md) | Too few posts | 0.498 | 0.494–0.500 | 1 | 0 | 1 |

Most recent posts:

- Praise, 2026-09-24, @FactoryAI (X): “@factoryai @fireworksai_hq excellent that legacy-bench measures more than just whether the code compiles. in payroll, erp, and closures, a plausible but incorrect output can alter withholdings or reconciliations. evaluating edge cases and traceable evidence, with final human review, is key.” [source](https://twitter.com/1569177389959192578/status/2103233553891025320)
- Praise, 2026-09-21, @droid (X): “@droid it is a massive step up for the model, especially with those self-verification capabilities. we actually went deeper on this here: <strict_link>” [source](https://twitter.com/1213502906332110848/status/2102172305392635915)
- Praise, 2026-09-02, G2 (G2): “q: what problems is the product solving and how is that benefiting you? a: having used factory ai in our engineering workflows, the standout feature for me is its autonomous agents, which they call droids. the biggest problem it solves for us is developer fatigue from multi-file refactoring, ongoing maintenance, and pull requests. most coding ai tools just sit inside your code editor and offer lin…” [source](https://www.g2.com/products/factory-ai/reviews/factory-ai-review-13397954)
- Complaint, 2026-09-22, @FactoryAI (X): “@factoryai @anthropicai fewer tokens are useful only when the harness catches the missing ones. long investigations should end in a checked spec or pr, not just a confident summary.” [source](https://twitter.com/2099871292480421888/status/2102478492915114092)
- Complaint, 2026-09-11, G2 (G2): “q: what problems is the product solving and how is that benefiting you? a: a lot of engineering time still gets eaten up by repetitive, multi-file work that isn’t difficult, just time-consuming—small refactors, test fixes, pr cleanup, documentation updates, and straightforward ticket implementation. factory lets me hand those pieces off to droids so i can stay focused on design decisions, tougher…” [source](https://www.g2.com/products/factory-ai/reviews/factory-ai-review-13441655)

### Interface and sessions: Typical (customer love 0.477, n 32)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [How the interface shows work, and what the user can configure](https://feedbackbench.com/criteria/ui.display_settings.md) | Too few posts | 0.485 | 0.465–0.507 | 22 | 5 | 17 |
| [Mobile, remote-control and voice access](https://feedbackbench.com/criteria/surfaces.remote_mobile.md) | Too few posts | 0.489 | 0.476–0.500 | 7 | 1 | 6 |
| [Cloud and remote sandbox execution](https://feedbackbench.com/criteria/surfaces.cloud_sessions.md) | Too few posts | 0.496 | 0.484–0.506 | 4 | 2 | 2 |
| [Saving, switching, resuming and rewinding sessions](https://feedbackbench.com/criteria/ui.session_history.md) | Too few posts | 0.496 | 0.491–0.500 | 2 | 0 | 2 |
| [Stopping and steering a running agent](https://feedbackbench.com/criteria/ui.interrupt_steer.md) | Too few posts | 0.503 | 0.500–0.508 | 1 | 1 | 0 |

Most recent posts:

- Praise, 2026-09-27, @FactoryAI (X): “@anasibnanwar @factoryai yeah, even i just saw them while messing with the settings; really cool.” [source](https://twitter.com/1817155037938036736/status/2104205288362659864)
- Praise, 2026-09-27, @FactoryAI (X): “@droid @factoryai i will drop grok/cursor for you guys as soon as my subscription ends. @da7_tech convinced me with his post. i just hope you guys beat devin because i can't afford you all 😆 i've tried it before, it was awesome but very expensive. also the most beautiful ui.” [source](https://twitter.com/1406556428840603649/status/2104235627847835986)
- Praise, 2026-09-27, @droid (X): “first time using anthropic models (opus 5.5). i use @droid. i gave it a task to design a landing page for my current project, and it asked questions i have never seen from an agent before. i hope the result comes out good, but so far, it's really impressive and i might not use openai models, unless they actually have a god response. also can only recommend the factory app. it's genuinely brilliant…” [source](https://twitter.com/1821640621347495936/status/2104230757883388046)
- Complaint, 2026-09-27, @FactoryAI (X): “@droid @factoryai here's what's sticking out right now - no droid ios app - desktop app missing from linux - syncing missions / repos / active work between droid computers (my own not droid managed). i need to be able to shift my coding workloads off my laptop and take them everywhere with me.” [source](https://twitter.com/1382136217601417222/status/2104238899493318925)
- Complaint, 2026-09-27, @FactoryAI (X): “@droid @factoryai my nits are all qol - droid mobile app - windows and mac dev environments for bots (using namespace devboxes in the meantime) - handoff between local and cloud agents - extend droid cloud agent access, we are limited to 40h/month while competing products are unlimited” [source](https://twitter.com/2070908287978246144/status/2104255956439961935)
- Complaint, 2026-09-27, @FactoryAI (X): “@droid @factoryai droid is genuinely useful for larger, multi-file tasks and does a good job staying on track without constant guidance. the biggest improvement for me would be better visibility into its reasoning/progress and more predictable results on longer tasks :)” [source](https://twitter.com/2093736525116702720/status/2104324409502970188)

### Reliability and speed: Too few posts (customer love 0.519, n 23)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Latency, throughput and fast mode](https://feedbackbench.com/criteria/rel.response_speed.md) | Too few posts | 0.511 | 0.495–0.529 | 11 | 7 | 4 |
| [Client crashes, freezes and failed tool execution](https://feedbackbench.com/criteria/rel.client_failures.md) | Too few posts | 0.488 | 0.480–0.496 | 9 | 0 | 9 |
| [Outages, server errors and capacity or rate errors](https://feedbackbench.com/criteria/rel.service_errors.md) | Too few posts | 0.496 | 0.491–0.500 | 3 | 0 | 3 |
| [Updates break working setups](https://feedbackbench.com/criteria/rel.update_breakage.md) | Too few posts | 0.499 | 0.495–0.500 | 1 | 0 | 1 |

Most recent posts:

- Praise, 2026-09-27, @FactoryAI (X): “the speed of deepseek v4.1 flash in @droid is mind blowing. makes me rethink the ollama subscription - when @factoryai has it all. <strict_link>” [source](https://twitter.com/1590702228234391552/status/2104263092364615842)
- Praise, 2026-09-27, @droid (X): “the speed of deepseek v4.1 flash in @droid is mind blowing. makes me rethink the ollama subscription - when @factory has it all. <strict_link>” [source](https://twitter.com/1590702228234391552/status/2104258455695741401)
- Praise, 2026-09-24, @FactoryAI (X): “@trevorbmurkp @harrystuck77 @factoryai @droid @badlogicgames @ampcode @cursor_ai @opencode @anthropicai @claudeai @cognition @xai @grok @build we’ve done a decent bit of perf improvements in the past few weeks with more incoming!” [source](https://twitter.com/1721143727043887104/status/2102916465167159772)
- Complaint, 2026-09-27, @FactoryAI (X): “@droid @factoryai it was feature locked until last update.” [source](https://twitter.com/70830663/status/2104209256119775343)
- Complaint, 2026-09-27, @FactoryAI (X): “@droid @factoryai receiving 403s on all requests, fix your system” [source](https://twitter.com/1453931380983820293/status/2104294136631181394)
- Complaint, 2026-09-25, @FactoryAI (X): “@factoryai hi i am getting 400 bad request on gpt 6 luna, solution and opus 5.5. i am on the progress plan. any pointers to check this” [source](https://twitter.com/4499700680/status/2103317535668220377)

### Account and support: Too few posts (customer love 0.576, n 28)

| Criterion | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|
| [Support, refunds and issue handling](https://feedbackbench.com/criteria/account.support.md) | Too few posts | 0.505 | 0.484–0.531 | 13 | 3 | 10 |
| [Data retention, training use and deployment isolation](https://feedbackbench.com/criteria/account.data_privacy.md) | Too few posts | 0.544 | 0.521–0.567 | 11 | 11 | 0 |
| [Wrong charges, failed payments and plan provisioning](https://feedbackbench.com/criteria/account.billing_errors.md) | Too few posts | 0.493 | 0.483–0.499 | 6 | 0 | 6 |
| [Account bans and access restrictions](https://feedbackbench.com/criteria/account.bans_restrictions.md) | Too few posts | 0.500 | 0.500–0.500 | 0 | 0 | 0 |

Most recent posts:

- Praise, 2026-09-27, @FactoryAI (X): “we should all learn from how reactive @factoryai and @tereza_tizkova are. thank you! <strict_link>” [source](https://twitter.com/1617212256487411712/status/2104277623085920418)
- Praise, 2026-09-24, @FactoryAI (X): “@ain3sh @factoryai true, never expected that! also i sent an error on the dm and bug was fixed and less than 1 hour… i was making my decision between factory and cursor, now is a no brainer! you guys won, also for the openai models support!” [source](https://twitter.com/17719163/status/2103007206064968004)
- Praise, 2026-09-19, @FactoryAI (X): “@factoryai crazy effort by the team and a big unlock for a large chunk of the world that hasn’t been able to access the frontier of coding because of their deployment requirements” [source](https://twitter.com/1819162791351406592/status/2101263327855063167)
- Complaint, 2026-09-27, @FactoryAI (X): “@droid @factoryai why doesn't your support team reply to my email? since i'm a free plan user, am i not deserved to spend time?” [source](https://twitter.com/1169569548313382912/status/2104223437174767853)
- Complaint, 2026-09-27, @FactoryAI (X): “@droid @factoryai let your users pay you!! been stuck with a failed payment issue since august :(” [source](https://twitter.com/306079362/status/2104243378246578672)
- Complaint, 2026-09-25, @FactoryAI (X): “@factoryai hi, i applied to the factory guild the first day it opened on august 14th, the first day it opened. i received an application confirmation email, but have not received any response after that.” [source](https://twitter.com/260499727/status/2103306229783183829)
