# Frontend and visual UI output (`work.frontend_ui`)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/criterion/work.frontend_ui

Area: [Doing the work](https://feedbackbench.com/criteria/work.md)

**Definition.** How the agent handles UI design, layout, visual taste and front-end component wiring.

**Boundary.** Not this: see [Breaks existing code or reintroduces bugs](https://feedbackbench.com/criteria/work.regressions_introduced.md) for UI bugs reintroduced after fixes.

Rated author-weeks, all agents: 790. Complaint share: 56%.

## The brief

Written by Claude Opus 5.5 from 70 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**Claude Code wins on taste, Codex loses on layout.**

TL;DR:

- Claude Code is the agent users reserve for design work; its complaints are generic output and drift.
- OpenAI Codex draws the most frontend posts and skews negative, mostly on layout and coherence.
- Output tracks the model inside the harness. Swapping models flips results in Cursor, Codex and OpenCode.

In plain terms: Expect a working UI that looks like every other AI site. Spacing, alignment and balance often break, and most agents cannot see their own output. Users get better results with screenshots, component libraries and small targeted fixes.

### How it breaks

- **Every site looks AI-generated** ([Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md)). Even with explicit instructions against it, agents converge on the same stock AI aesthetic, and detailed prompts or added skills barely move the result.
  The most common taste complaint is sameness. Users say telling the agent not to look generic changes nothing. Posts describe output that implements a spec correctly but shows no design judgment. Without careful prompting, frontends come out generic and weak. The ask for less generic, more creative output spans Codex, Claude Code and Antigravity, so no single vendor owns this problem.
  Evidence:
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-06: “even when i explicitly tell it “don't make it look generic”, it somehow manages to produce something that looks exactly like every other ai-generated website i've tried giving it detailed prompts, vercel skills, etc. but somehow there's no big change in the final response so is there any way to fix this shit???” [source](https://www.reddit.com/r/ClaudeCode/comments/1w90u75/is_there_actually_a_way_to_stop_ai_coding_agents/)
  - Complaint, OpenCode, r/opencode, 2026-09-10: “hi, muse works well for intermediate tasks. but the problem arises when you give it a request that's too complex, as it starts to hallucinate, get sidetracked, and make terrible decisions and it happened to me that he did many things just for a simple and easy-to-fix problem. in frontend development, in my opinion, if you don't make good prompts, it makes your frontend too generic and terrible. i haven't tested it further, but this is my "for now".” [source](https://www.reddit.com/r/opencode/comments/1waq3e5/muse_spark_13_free_is_ass/p8vmot3/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-04: “really good? every claude model i've tried has been awful at ui. it can implement a spec yes but it has no design taste. i had fable coordinate agents to build from the same spec after grilling me. opus, sol, glm, gemini. opus was the worst out of those (except maybe gemini, it burned my whole quota in two attempts so hard to say)” [source](https://www.reddit.com/r/ClaudeCode/comments/1w6h4kt/astra_release_result/p7r3k0g/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-21: “right, but as stated in the post, i had developed the architectural plan - and to be serious 3 or 4 pages does not really necesitate architecture in the context of what i was doing. i think i could have received a more coherent ui from a human. in other areas, it is superhuman” [source](https://www.reddit.com/r/codex/comments/1wlvf0j/state_of_agentic_coding/pb3andk/)

- **Layouts break and agents cannot see it** ([Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md)). Alignment, padding, clipping and balance fail repeatedly, and agents that only read code have no way to notice the page looks wrong.
  Users report lopsided containers, text clipping out of buttons, mismatched colors and panels cut off at random. One Codex user says the core problem was layout, not just looks. Another routes screenshots through a second model to act as the first agent's eyes, which costs extra. Accurate layout and visual self-checking is a standing request, and every ask for it comes from OpenAI Codex users.
  Evidence:
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-18: “the biggest issues i am seeing is with the ui, it always makes right side of the container unbalanced” [source](https://www.reddit.com/r/google_antigravity/comments/1wj0j6r/sudden_big_increase_in_gemini_31_pro_efficiency/pakpgnv/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-04: “the only good workflow for me was having claude design the initial shell then utilizing tibo's resets to spam ui changes in my "design lab" it coded for me where i can box highlight ui elements then change it more specifically there for my likening. my main gripe is the layout. that was its problem. not just how it looks.\]\\ generate an initial shell with sol then try what i did lmao. you will lose your mind. i guarantee it.” [source](https://www.reddit.com/r/codex/comments/1w7i7gh/i_have_astra_but_anybody_who_uses_it_has_it/p7vchto/)
  - Complaint, Cursor, r/cursor, 2026-09-26: “the last 2 days i tried to change a complete ui with 4.7 and it was disastrous. not even managing to put textsize, padding, or matching colors correctly, much less random clipping of panels, randomly text clipping out of buttons, not being able to align things, not being able to understand mdc instruction files, etc. actually today 4.6 is cleaning up the whole day behind 4.7s mess. im actually shocked, i at least expected it to perform similar. but all the trust i had in my agents got severely reduced by this performance.” [source](https://www.reddit.com/r/cursor/comments/1wqhcse/grok_46_vs_47/pc5lsop/)
  - Praise, Google Antigravity, r/google_antigravity, 2026-09-16: “i would say though that the main thing i'd use gemini for would be imaging - if you're building a ui then claude sucks at taking screenshots and fixing bugs in the ui - it's sometimes easier to ask gemini to look and comment what the problems are and asking claude to use gemini as it's eyes. i've done this using paid api calls but it's too expensive to be worth it most of the time.” [source](https://www.reddit.com/r/google_antigravity/comments/1whvudv/people_who_are_complaining_about_gemini/pa652ei/)

- **The model swap decides the outcome** ([Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md)). Frontend quality swings sharply with the model behind the agent. Users see one version nail a task and the next one wreck it.
  Cursor users describe a newer Grok release that failed basic sizing and color matching, while the older version spent a day cleaning up. Codex users see two models with identical benchmark scores produce very different frontends. Antigravity users say a cheaper model beats the hyped flagship on UI. The harness stays the same and the results move with the model, so trust in a setup can collapse after one upgrade.
  Evidence:
  - Complaint, Cursor, @cursor_ai, 2026-09-23: “grok 4.7 thinks white text on a light background is a good design choice on an open-ended simple frontend refactor smfh @elonmusk @supergrok @cursor_ai” [source](https://twitter.com/1602107661461639169/status/2102614948086194550)
  - Complaint, OpenAI Codex, r/codex, 2026-09-24: “\> gpt-6 luna max and gpt-6 sol xhigh have the exact same score: 66.8 on paper it looks the same, but when i tried to use luna to work on frontend stuff, it produced shit. might be a skill issue tho. wondering if each model have their own "personality" like, sol is more for coding and stuff, luna for repetitive high volume tasks? so eventho the score is the same, it produced different result when being asked to work on things that aren't their familiar fields?” [source](https://www.reddit.com/r/codex/comments/1wopzcz/gpt6_luna_sol_or_astra_heres_what_id_use/pbppvwy/)
  - Complaint, Cursor, r/cursor, 2026-09-26: “the last 2 days i tried to change a complete ui with 4.7 and it was disastrous. not even managing to put textsize, padding, or matching colors correctly, much less random clipping of panels, randomly text clipping out of buttons, not being able to align things, not being able to understand mdc instruction files, etc. actually today 4.6 is cleaning up the whole day behind 4.7s mess. im actually shocked, i at least expected it to perform similar. but all the trust i had in my agents got severely reduced by this performance.” [source](https://www.reddit.com/r/cursor/comments/1wqhcse/grok_46_vs_47/pc5lsop/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-13: “i don’t trust benchmarks, but the level of overhype around astra is simply false advertising. trust your real world usage. i don’t trust astra 6 with anything to do with the ui, flash kicks its ass. pure logic astra is probably better. but not by how many more tokens it uses.” [source](https://www.reddit.com/r/google_antigravity/comments/1wfk0jh/realswe_benchmark_gemini_flash_38_on_third_place/p9mutce/)

- **Polished concept, different implementation** ([Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md)). Agents draw an attractive mockup, then ship something that drifts from it, and full-screen rewrites make the drift worse.
  Users say the concept looks right and the implemented screen does not match. The workaround posts converge on one fix. Stop asking for full redesigns, give one screenshot, name the component file, and set a few pixel targets. Reference tools help with intent, but some agents treat a sample screen as a template to copy one-to-one instead of a style to adapt.
  Evidence:
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-11: “i used to be on claude but recently it simply fails every time at fixing something without breaking the rest. additionally, regarding nice polished ui, claude opus 5 became unusable, it draws nice concepts and once implemented it looks different. astra had no issues and masters clean ui so i plan to switch...” [source](https://www.reddit.com/r/ClaudeCode/comments/1wdgvrl/i_tried_astra_with_pro_and_honestly_kind_of/p96qpga/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-20: “for last mile spacing and fonts i stop asking for whole screen rewrites. paste one screenshot, name the exact component file, and give 2-3 pixel targets only. figma mcp helps for intent, but small css diffs beat another full redesign pass when the model keeps drifting.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wlawrv/claudecode_unable_to_implement_accurate_ui/pax858x/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-19: “i actually have started using stitch if i'm doing something and i need some ideas first as to what i want. it's good enough to get a single full screen then i extrapolate from that. i have tried to get it to do more but then some agents get confused and try to use it 1:1 vs just something to model a new style after.” [source](https://www.reddit.com/r/google_antigravity/comments/1wkf3kc/what_sample_data_do_you_give_antigravity_before/paq8mkx/)

- **A harness with eyes closes the loop** ([Frontend and visual UI output](https://feedbackbench.com/criteria/work.frontend_ui.md)). Setups that let the agent screenshot and inspect its own output, or that add a focused skill, beat raw model power on visual work.
  An agent that only reads build output writes code that compiles and still looks wrong. Users praise OpenCode for putting a simulator beside the session so layout bugs stop surviving repeated rounds. Pi users get pixel-faithful results from a mid-tier model plus a dedicated skill. Factory users say the same model produces better-looking UI in Droid than in other harnesses. Users credit the harness, not the model, for these gains.
  Evidence:
  - Praise, OpenCode, @opencode, 2026-09-06: “putting the simulator right next to the session is the part that matters. an agent that can only read build output writes code that compiles and still looks wrong. once it can take a screenshot and read the view hierarchy after each change, the loop actually closes and layout bugs stop surviving three rounds. nice work.” [source](https://twitter.com/2018819126429450240/status/2096599908090335466)
  - Praise, Pi, @pidotdev, 2026-09-01: “you don't need expensive models like fable or gpt-5.6 sol to turn any ui design image into a highly complex, pixel-perfect ui. i can get this done using just gpt-5.6 luna. i'm using pi @pidotdev + the pixel perfect skill. start with image gpt → pi harness + pixel perfect skill. turns out, the model isn't always the limitation. the right harness + skill can make a huge difference 👌” [source](https://twitter.com/2083448877928312832/status/2094695540978331748)
  - Praise, Factory, @FactoryAI, 2026-09-17: “only downside is the 5 hours limit and the price which is pretty fair but still a little for me personally. other than that i can list so many things i love about it. droid is super efficient and often finish tasks faster than most other agent with similar results. i feel like it gets the right context at the right time. it’s pretty amazing. also love the byok, live the fact that ui almost always looks better when done with droid even using the same model in other harnesses.. i really am a fan of the product. 😅” [source](https://twitter.com/1617212256487411712/status/2100703532785541412)
  - Praise, Factory, @droid, 2026-09-17: “i've been having a lot of success recently doing ui work with @droid and gemini 3.8 flash. it consistently is better than fable 5.1 and costs a fraction.” [source](https://twitter.com/1247892463479451653/status/2100622735273545963)

### Who stands out

- **Claude Code (stronger)**. Users who rotate between agents still route design work to Claude Code, even those who moved their main coding elsewhere.
  Posts describe Claude as the design tool in multi-agent setups and say its models outperform rivals on website design. A Codex convert names Claude's frontend taste as the thing they miss. The praise is not unanimous. Some users rank it last for taste, and others cite generic output and mockups that drift on implementation. Requests focus on better SVG, 3D and game graphics.
  Evidence:
  - Praise, Claude Code, r/ClaudeCode, 2026-09-26: “stick with claude for now. opus 5.5 is unbelievable at website design. then once difficult design work is over or limits run out, shift to deepseek v4.1 flash. it is unbelievable roi for the price” [source](https://www.reddit.com/r/ClaudeCode/comments/1wqcrg9/moving_from_chatgpt_codex_to_claude_code/pc42axg/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-06: “i do use codex at the moment, but usually rotate my monthly sub depending on which provider has the "best" model for general tasks. i use cursor to bridge limit resets and claude to do design stuff (web/ui). kimi & glm would be great if i couldn't use subs for some reason (being a company, etc).” [source](https://www.reddit.com/r/ClaudeCode/comments/1w8uxbn/astra_might_be_the_final_nail_in_the_coffin_for_me/p85i9ax/)
  - Complaint, OpenAI Codex, r/ClaudeCode, 2026-08-31: “moved from claude to codex. i still use claude for work. a tip might be to steer yourself from reintroducing your current claude skills into codex. not everything translate, and i found the codex harness to include several tips that back then were not on claude code and navigating the changes between harnesses was much better as a fresh start. also, i really do recommend codex. i do miss claude's better frontend taste, but the 5.6 models are a workhorse to be recommend with, and don't come in the box with that annoying tone claude uses. plus points for the speed!” [source](https://www.reddit.com/r/ClaudeCode/comments/1w3diti/gpt_vs_claude/p6zesql/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-11: “claude design on even opus 5 medium massively outperforms astra on xhigh. unless something changed since last night.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wdnqgc/token_spend_comparison_in_claude_and_codex_40x/p97k4xa/)

- **OpenAI Codex (weaker)**. Codex draws by far the most frontend posts, and they lean negative on layout, coherence and visual precision.
  Users say Codex is superhuman elsewhere but produces less coherent UI than a human would. They report layout as the core failure. Several users design the shell in Claude and only iterate in Codex. The bright spot is the Astra model, which users praise for replicating sites from images, but posts say it is gated to higher subscription tiers. Codex users also lead the requests for better design quality and a dedicated design product.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-16: “only tried to use it for something visual once and it made a complete hashjob of it. you really need to be precise” [source](https://www.reddit.com/r/codex/comments/1wi0wv3/agi_cant_center_a_div/pa6phrg/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-21: “right, but as stated in the post, i had developed the architectural plan - and to be serious 3 or 4 pages does not really necesitate architecture in the context of what i was doing. i think i could have received a more coherent ui from a human. in other areas, it is superhuman” [source](https://www.reddit.com/r/codex/comments/1wlvf0j/state_of_agentic_coding/pb3andk/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-04: “the only good workflow for me was having claude design the initial shell then utilizing tibo's resets to spam ui changes in my "design lab" it coded for me where i can box highlight ui elements then change it more specifically there for my likening. my main gripe is the layout. that was its problem. not just how it looks.\]\\ generate an initial shell with sol then try what i did lmao. you will lose your mind. i guarantee it.” [source](https://www.reddit.com/r/codex/comments/1w7i7gh/i_have_astra_but_anybody_who_uses_it_has_it/p7vchto/)
  - Praise, OpenAI Codex, r/codex, 2026-09-18: “gpt-6 astra is very good in the frontend, so you can notice a clear leap. unfortunately, it must also be said that it can only be used in the higher subscriptions and cannot be used at all in the plus subscription (or for 1-2 prompts per week 😆). therefore, you have to accept this limitation. otherwise, you can of course also look at the gemini models or glm-5.3-flash, which are also not bad.” [source](https://www.reddit.com/r/codex/comments/1wje03t/ui_design_tips/paik3wz/)

- **Google Antigravity (mixed)**. Some users rate Antigravity's design output above Codex and Claude, while others call its Gemini models weak at UI and its layouts lopsided.
  Praise centers on frontend taste, a strong 3D demo, and summaries with screenshots that make UI review easier. Complaints point the other way. Users report unbalanced containers, distrust of the flash model for design, and one user who moved UI work to a different model in OpenCode. Posts suggest users like Antigravity's frontend taste but distrust it for backend logic.
  Evidence:
  - Praise, Google Antigravity, r/google_antigravity, 2026-09-03: “i mainly use codex to update a few static websites for me and clients. honestly, i like ag design output better but heard it’s much faster on token burn.” [source](https://www.reddit.com/r/google_antigravity/comments/1w5s0pu/does_antigravity_burn_at_the_same_rate_as_codex/)
  - Praise, Google Antigravity, r/google_antigravity, 2026-09-20: “context ? i have 3 subscriptions, claude, copilot and gemini, use all of them regularly, gemini went to number one in speed and accuracy followed by claude. when comes to building uis for example, gemini is far superior, the agy’s summaries with screenshots are absolutely excellent. overall. incredible speed improvement” [source](https://www.reddit.com/r/google_antigravity/comments/1wllxyb/this_thing_became_a_coding_beast/pb0mvsd/)
  - Praise, Google Antigravity, r/google_antigravity, 2026-09-02: “i ran it for agentic tasks in agy, it was good. also i had a long thread and i didn't notice context rot. for frontend test i asked to create a 3js solar system, it surprised me, much better output than opus 4.6 and sol for another frontend tast. frontend taste is good. but it lacks backend taste. asked sol to give plan and sol medium plan/code was simpler. gemini over complicated it unnecessarily.” [source](https://www.reddit.com/r/google_antigravity/comments/1w5i29e/review_of_gemini_38_flash_from_a_person_who/p7glfs3/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-21: “the app (antigravity 2.0) for any bigger implementation, some knowledge work or planning on a codebase. the ide (not extension) for casual work and reviewing the work done by the app, because i can read, write into files easily, proper source control panel. i don't do any complex work on ide because gemini is capable on ide than the app (ide even lacks subagents). cli rarely, i have it installed but i rarely open it. but tbh i don't really use antigravity for complex and serious implementations in my codebases. i use muse spark thru opencode for designs and ui, and it has been one of the best in the area. i m very satisfied with the experience, it generated me such nice uis. - gemini is pretty bad in this btw. for complex and serious work i use glm 5.3 or gpt 5.6 sol thru api on opencode.” [source](https://www.reddit.com/r/google_antigravity/comments/1wluoqe/which_antigravity_surface_do_you_use_the_most/pb3k3cf/)

- **Cursor (mixed)**. Cursor's frontend reputation rides on whichever model users pick. Grok is smooth for components until a release regresses and long-time users leave.
  Users praise smooth component work and a real frontend loop, plus precise motion UI when they point Cursor at specific locations. The complaints target the model. One Grok release placed white text on a light background and mangled a full UI change, and one long-term top-tier user is switching to Claude Code over frontend gaps. Requests ask for a visual click-to-edit workflow and better design systems.
  Evidence:
  - Praise, Cursor, @cursor_ai, 2026-09-02: “@xueweidiqiu @sidharthfalodia @jiya_3063 @cursor_ai haha, caught red-handed. but using grok to write front-end components in cursor is really smooth, just try it and you'll know if i'm bragging.” [source](https://twitter.com/1720665183188922368/status/2095140470984700158)
  - Complaint, Cursor, r/cursor, 2026-09-23: “disappointed, too. been on the ultra plan for months, but i'm switching to claude code (terminal) with a max subscription today. while grok 4.6 and 4.7 deliver good backend performance, they struggle significantly with the frontend requirements of my agent-driven workflow.” [source](https://www.reddit.com/r/cursor/comments/1woad6b/grok_47_performance_in_cursor/pbma3ud/)
  - Praise, Cursor, r/codex, 2026-09-19: “i’m using a lot of motion ui atm, and it works really well. cursor with the pre-made code for that certain ui, specific location via cursor aswell” [source](https://www.reddit.com/r/codex/comments/1wkhzv3/i_found_the_best_way_to_build_insane_uis_with/paqybvm/)
  - Complaint, Cursor, @cursor_ai, 2026-09-07: “i’m greatly impressed with @cursor_ai i’m using it after grok @bot made me get into their ecosystem. > cursor is the best competition for codex. > claude code is no where near cursor, but with claude models on cursor it feels way better to get great results. my wishlist for @grok @spacexai and @cursor_ai to do is 1. get the frontend design systems better. 2. out of the box animations intelligence. 3. understanding the user or human input by conversations. these are the things as a non-technical cursor user like me might want!!! 🖖” [source](https://twitter.com/1520714109104582656/status/2097028525329170532)

### Fine print

- Many posts judge underlying models such as Astra, Grok or Muse rather than the agent harness that runs them.
- Most agents here have too few posts to rank. Their quotes illustrate patterns but do not establish standing.

## Top requests

What users ask to add or change, most asked first. 58 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Author-weeks | Posts | Agents (author-weeks) |
|---|---|---|---|---|
| 1 | Better overall UI design quality | 13 | 13 | OpenAI Codex 7, Claude Code 3, Amp 1, Google Antigravity 1, Cursor 1 |
| 2 | Dedicated design product or feature | 8 | 8 | OpenAI Codex 5, Cursor 2, OpenCode 1 |
| 3 | Accurate layout and visual self-checking | 6 | 6 | OpenAI Codex 6 |
| 4 | Better SVG, 3D and game graphics | 5 | 7 | Claude Code 3, Google Antigravity 2 |
| 5 | Visual click-to-edit UI workflow | 4 | 5 | Cursor 2, Claude Code 1, Conductor 1 |
| 6 | Less generic, more creative UI output | 4 | 4 | OpenAI Codex 2, Google Antigravity 1, Claude Code 1 |
| 7 | Frontend annotation in built-in browser | 3 | 4 | OpenAI Codex 2, Factory 1 |
| 8 | Consistent style-matched UI generation | 3 | 3 | OpenAI Codex 3 |
| 9 | Built-in website animation generation | 2 | 2 | Claude Code 1, Cursor 1 |

### 1. Better overall UI design quality

- Claude Code, 2026-09-23, r/ClaudeCode (Reddit): “genuinely really struggling to generate anything that actually looks good it seems like no matter how much time claude spends on front end, it always ends up looking like shit now i have had claude pump out great designs and work in the past, but now it seems like no matter what i prompt, or what mockups i give it, it just cant seem to get it done any tips??” [source](https://www.reddit.com/r/ClaudeCode/comments/1wo85p3/am_i_crazy_or_has_front_end_been_terrible_as_of/)
- OpenAI Codex, 2026-09-23, r/codex (Reddit): “they need to step up their ui game. now even most chinese models are better at ui than openai.” [source](https://www.reddit.com/r/codex/comments/1wnto0j/sol_6_is_a_slop_fest/pbifiz1/)
- OpenAI Codex, 2026-09-11, r/codex (Reddit): “for someone coming from the claude code ecosystem and suddenly seeing that simple development things are a struggle in codex is a heartbreak! it was definitely a breeze in claude code. reading from your comment the amount of effort you need to get a good ui on astra, i'm questioning the usefulness for developers. all software engineers, product managers, designers have become app builders. and it's the one thing that should be nailed well!!” [source](https://www.reddit.com/r/codex/comments/1wdd9s5/made_a_mistake_switching_to_astra/p953124/)

### 2. Dedicated design product or feature

- Cursor, 2026-09-27, @cursor_ai (X): “hey @cursor_ai, when are you going to introduce the design feature similar to claude design ?” [source](https://twitter.com/1866323245/status/2104234697245237480)
- Cursor, 2026-09-25, @cursor_ai (X): “will xai or @cursor_ai have something like claude design soon?” [source](https://twitter.com/267061596/status/2103402068367273984)
- OpenAI Codex, 2026-09-17, X search: OpenAI Codex, Codex CLI, Codex app (X): “anthropic launched claude slides, claude design and claude docs. imagine a native ui like this in codex app usable with the cheap luna model. it would be the endgame. we have pptxgenjs based skills for slides, and we can use figma or penpot mcp server for the canvas designs, but the launch video for the new claude code feels seamless when switching between the different workflows.” [source](https://twitter.com/395739576/status/2100403002330980505)

### 3. Accurate layout and visual self-checking

- OpenAI Codex, 2026-09-03, r/codex (Reddit): “all this is nice but i will be happy if it will be able to align 2 buttons in my app without 6 prompts.” [source](https://www.reddit.com/r/codex/comments/1w6h382/the_gpt6_astra_launch_video_before_the_site_was/p7mz4dy/)
- OpenAI Codex, 2026-09-02, r/codex (Reddit): “i use a decent amount of appium/playwright, and while codex writes test covering different viewport sizes, it basically ignores a lot of glaring issues despite writing huge amounts of tests. so i usually have to call attention to each matter myself. my hope is for better ui layout assertions - image -> task.” [source](https://www.reddit.com/r/codex/comments/1w5j92s/what_is_the_minimum_standard_for_astra_you_would/p7h43mh/)
- OpenAI Codex, 2026-09-13, r/codex (Reddit): “i’ve spent the last 20 prompts trying to get codex to fix a fairly straightforward game ui, but it's either playing dumb or is actually dumb. the inventory overlaps other menus, elements are constantly cut off, buttons and item slots don’t line up with the artwork, and touch areas don’t match what’s shown on screen. fixing one thing often breaks something else. it also seems heavily dependent on fixed coordinates, so i’m worried it won’t work acr” [source](https://www.reddit.com/r/codex/comments/1wf3y6k/flutter_ui_keeps_breaking_and_im_just_wasting/)

### 4. Better SVG, 3D and game graphics

- Google Antigravity, 2026-09-27, @antigravity (X): “@ash_twtz yes, i would love to see more support for media/design/etc in @antigravity, support and tooling for design.md, etc.” [source](https://twitter.com/2056251/status/2104150215003361657)
- Claude Code, 2026-09-08, @ClaudeDevs (X): “@claudedevs it would be great if claude’s image interpretation and its ability to generate 2d, ui, and 3d models improve. coding is fine, but this kind of ability is way too low compared to gpt-6” [source](https://twitter.com/3033331083/status/2097144872516153425)
- Google Antigravity, 2026-09-03, @antigravity (X): “@googledevs @antigravity gemini need to improve in creating advanced svg like any full body character by prompt(no img) or with img, &amp; in 3d any advanced character with img or prompt, or any 3d advanced physics simulations for websites or for any 3d things making through blender mcp(default in agy 2.0).” [source](https://twitter.com/1824095764202848256/status/2095639607270617308)

### 5. Visual click-to-edit UI workflow

- Cursor, 2026-09-11, @cursor_ai (X): “bring back the visual editor from cursor version 2.2; it unified the visual design with the code. you could change the order of buttons, rotate sections, and test different grid layouts without switching contexts. it was an excellent tool. @cursor_ai” [source](https://twitter.com/2092797055584419840/status/2098228110793580622)
- Conductor, 2026-08-31, @conductor_build (X): “@conductor_build + @poteto pstack = software factory waiting on a better frontend iteration experience but great work @charlieholtz and team!!” [source](https://twitter.com/437086246/status/2094321848112840852)
- Cursor, 2026-08-31, r/ClaudeCode (Reddit): “i really like claude code, but i miss cursor’s visual editing workflow where i can click on the actual ui in the browser and directly tweak things like spacing, sizing, styles, etc. instead of describing every little change in chat. i’ve tried claude design, but unless i’m using it wrong, it feels more like a separate design/prototyping environment with a handoff to claude code, rather than something i can use to directly manipulate the ui of my” [source](https://www.reddit.com/r/ClaudeCode/comments/1w3omol/is_there_a_cursorstyle_visual_editing_workflow/)

### 6. Less generic, more creative UI output

- Google Antigravity, 2026-09-23, @antigravity (X): “please work a bit more on the frontend design and ui/ux capabilities of the gemini models so they stop producing so much ai slop... it would also be great if antigravity could do something about this. at the very least, make it so the model is discouraged from using generic, repetitive ai-slop designs and is pushed to be more creative and original.” [source](https://twitter.com/1881441934952079360/status/2102690352096276640)
- OpenAI Codex, 2026-09-18, r/codex (Reddit): “can you migrate it to codex? i keep a claude subscription just for ui i hate also that codex models are too afraid to big ui/ux changes so they are very sticky to the default ai look” [source](https://www.reddit.com/r/codex/comments/1wje03t/ui_design_tips/pai3kir/)
- Claude Code, 2026-09-06, r/ClaudeCode (Reddit): “even when i explicitly tell it “don't make it look generic”, it somehow manages to produce something that looks exactly like every other ai-generated website i've tried giving it detailed prompts, vercel skills, etc. but somehow there's no big change in the final response so is there any way to fix this shit???” [source](https://www.reddit.com/r/ClaudeCode/comments/1w90u75/is_there_actually_a_way_to_stop_ai_coding_agents/)

### 7. Frontend annotation in built-in browser

- OpenAI Codex, 2026-09-17, X search: OpenAI Codex, Codex CLI, Codex app (X): “i want it to be like the figma simulator where i can just test it in the codex app <strict_link>” [source](https://twitter.com/2028324370452754432/status/2100400559539065020)
- OpenAI Codex, 2026-09-13, X search: OpenAI Codex, Codex CLI, Codex app (X): “@theo i do a lot of front end with built in browser in cursor/codex. i wish t3 experience was more similar to cursor (or codex app). in my personal opinion, currently cursor has the best implementation. that’s one thing that stops me from switching fully.” [source](https://twitter.com/1907787527311728641/status/2099013382191870181)
- Factory, 2026-09-02, @FactoryAI (X): “@ross_cefalu @gitmaxd @factoryai @droid and i mean it! i've been using factory a lot lately, and since i'm first and foremost a designer, it would be cool if we had something like this, as we have in cursor or codex or some other <strict_link>” [source](https://twitter.com/171899126/status/2095276285220024543)

### 8. Consistent style-matched UI generation

- OpenAI Codex, 2026-09-16, r/codex (Reddit): “it created the mockup, i didn't give it a mockup. claude design is excellent at this. so i was genuinely curious why codex lacks so much in this aread.” [source](https://www.reddit.com/r/codex/comments/1whf8ek/is_this_just_a_codex_thing/pa6huzi/)
- OpenAI Codex, 2026-09-06, r/codex (Reddit): “astra, but authorize for asset generations, and preferred colors and styles. let it offer yiu options and select one. it will use its superior image generation, with transparency to make ui that blows any stock ui away. sol is even good at this over anything anthropic. i say this because claude doesnt have real image generation to use. strictly ui without image generation, its completely opinion base.” [source](https://www.reddit.com/r/codex/comments/1w8jmol/astra_or_fable_for_ui/p836o7i/)
- OpenAI Codex, 2026-09-13, r/codex (Reddit): “maybe it's a skill issue or bad setup and prompting, but astra seems to invent a brand new h2 element instead of copying neighbouring sections, for example. like, a page contains h2's, all of the same style and functionality (hover to copy anchor). i ask for a new h2 on the page, and it adds a different size, no-functionality h2. like do i really have to communciate that it should be the same as the others? cant it infer that from context? seems” [source](https://www.reddit.com/r/codex/comments/1wf2p83/on_todays_episode_of_nonstop_pain/)

### 9. Built-in website animation generation

- Claude Code, 2026-09-14, r/ClaudeCode (Reddit): “how do you create animations like these? i always have to wrestle with claude to add polished animations to my websites, and the results are usually pretty bad—as if it’s trying to work around its inability to create them properly. are you using claude itself, an external animation generator, or a library?” [source](https://www.reddit.com/r/ClaudeCode/comments/1wg0zp4/super_simple_ui_design_prompts_yielded_amazing/p9rjt4f/)
- Cursor, 2026-09-07, @cursor_ai (X): “i’m greatly impressed with @cursor_ai i’m using it after grok @bot made me get into their ecosystem. > cursor is the best competition for codex. > claude code is no where near cursor, but with claude models on cursor it feels way better to get great results. my wishlist for @grok @spacexai and @cursor_ai to do is 1. get the frontend design systems better. 2. out of the box animations intelligence. 3. understanding the user or human input by conve” [source](https://twitter.com/1520714109104582656/status/2097028525329170532)

## Every agent

| Agent | Overall rank | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|---|
| [Claude Code](https://feedbackbench.com/agents/claude-code.md) | 1 | Better than peers | 0.534 | 0.505–0.559 | 233 | 115 | 118 |
| [OpenCode](https://feedbackbench.com/agents/opencode.md) | 3 | Typical | 0.508 | 0.486–0.530 | 30 | 17 | 13 |
| [Google Antigravity](https://feedbackbench.com/agents/antigravity.md) | =5 | Typical | 0.508 | 0.482–0.536 | 44 | 22 | 22 |
| [Cursor](https://feedbackbench.com/agents/cursor.md) | 4 | Typical | 0.498 | 0.473–0.521 | 50 | 27 | 23 |
| [OpenAI Codex](https://feedbackbench.com/agents/codex.md) | 2 | Worse than peers | 0.473 | 0.451–0.495 | 403 | 153 | 250 |
| [Devin](https://feedbackbench.com/agents/devin.md) | =5 | Too few posts | – | – | 8 | 4 | 4 |
| [Conductor](https://feedbackbench.com/agents/conductor.md) | 14 | Too few posts | – | – | 5 | 3 | 2 |
| [Pi](https://feedbackbench.com/agents/pi.md) | 7 | Too few posts | – | – | 4 | 3 | 1 |
| [Zed](https://feedbackbench.com/agents/zed.md) | =8 | Too few posts | – | – | 4 | 3 | 1 |
| [Factory](https://feedbackbench.com/agents/factory.md) | =11 | Too few posts | – | – | 3 | 2 | 1 |
| [GitHub Copilot](https://feedbackbench.com/agents/copilot.md) | =8 | Too few posts | – | – | 2 | 1 | 1 |
| [Amp](https://feedbackbench.com/agents/amp.md) | =11 | Too few posts | – | – | 2 | 0 | 2 |
| [Warp](https://feedbackbench.com/agents/warp.md) | 15 | Too few posts | – | – | 2 | 1 | 1 |
| [Cline](https://feedbackbench.com/agents/cline.md) | =8 | Too few posts | – | – | 0 | 0 | 0 |
| [Kiro](https://feedbackbench.com/agents/kiro.md) | 13 | Too few posts | – | – | 0 | 0 | 0 |
| [Grok Build](https://feedbackbench.com/agents/grok-build.md) | 16 | Too few posts | – | – | 0 | 0 | 0 |
| [Augment Code](https://feedbackbench.com/agents/augment.md) | 17 | Too few posts | – | – | 0 | 0 | 0 |

## Posts

Receipts rule: The 5 most recent praise and complaint posts per area (first 700 characters) and 3 per criterion (first 450 characters).

### Claude Code

- Praise, 2026-09-27, r/ClaudeCode (Reddit): “fable is a big model and can be more thorough and ocasionally better at frontend.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrcyz1/so_if_fable_51_was_a_stopgap_for_opus_55_then/pcbivz4/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “way better at frontend.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrcyz1/so_if_fable_51_was_a_stopgap_for_opus_55_then/pcblh1u/)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “it's fable for me all the way. i loved opus 4.6-4.8, but then it just wasn't the same. i use fable for reviews and project oversight. i spinned opus 5.5 a few days back but it gave me slop and felt weird. still, my primary model for doing the work is astra on xhigh. i was surprised by how good it is with ui/ux where opus failed miserably.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrfr8p/fable_51_or_opus_55/pccclef/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “ok, but what has draw the ui? also claude? i am trying this but it is always giving me cartoon like lame graphics.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrjamn/opus_55_is_about_to_crush_suno/pcd53a9/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “from my experience so far, opus is far superior once you train it. what i did is gave it great examples, explained nuanced detailed on why they’re great, then have it build a “taste profile”. i do that for everything and so far it’s been masterful except for logo design. despite how much i trained it, it turns into a 4yr old on that request. also, if you ever need to build brand case studies, hook it up to krea’s mcp. it can generate template moc” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrjhyp/are_the_limits_that_good/pch3nkt/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “hi all, i've been working on a side project called rustcoach (rustcoach.dev), a web app that teaches rust with an ai coach. claude does double duty here: i built it with claude code, and the coach itself runs on the claude api. some of what i learned might be useful here, and i'd like feedback from anyone who has done something similar. **what it does, briefly** each lesson is generated for one learner from the coach's notes on them. you write th” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrgbxo/i_built_a_rust_tutor_with_claude_code_and_the/)

### OpenCode

- Praise, 2026-09-27, @opencode (X): “@iam_chonchol @opencode code-generated visuals remove the need for extra assets.” [source](https://twitter.com/1338474059848437761/status/2104230569576173941)
- Praise, 2026-09-27, @opencode (X): “svg lab mcp is exploding. we’ve seen a serious surge in users since last night, and the design engine is showing exactly why. even deepseek v4.1 flash running in @opencode is producing ridiculously good ui designs with svg lab mcp. this is getting interesting. <strict_link>” [source](https://twitter.com/1974714902422990848/status/2104269556026122483)
- Praise, 2026-09-25, r/opencode (Reddit): “it’s so good for me so far. made my ui so clean and better looking. i hope it doesn’t regress.” [source](https://www.reddit.com/r/opencode/comments/1wprfba/space_bunny_is_better_than_i_expected/pbzmt9b/)
- Complaint, 2026-09-26, r/opencodeCLI (Reddit): “yes it is but when he write interface i often see some errors like on the screenshot of the post” [source](https://www.reddit.com/r/opencodeCLI/comments/1wqa72d/space_bunny_often_use_foreign_languages_in_the/pc4w6l9/)
- Complaint, 2026-09-25, r/opencode (Reddit): “absolute shit for web ui design” [source](https://www.reddit.com/r/opencode/comments/1wp4kmi/m31_is_spacebunny/pbyqa9y/)
- Complaint, 2026-09-23, r/opencodeCLI (Reddit): “i just tested it now and it came out quite bad in terms of design, to me it was even worse than gpt 5.6 luna when it comes to website design, i lost the desire to test it in other areas, i think it is a very small model” [source](https://www.reddit.com/r/opencodeCLI/comments/1wo7o3c/space_bunny_stealth_model_is_free_for_the_next/pbkqxur/)

### Google Antigravity

- Praise, 2026-09-27, r/google_antigravity (Reddit): “sounds like a skill issue since gemini models are one of the best in designs.” [source](https://www.reddit.com/r/google_antigravity/comments/1wrbsby/antigravitys_gemini_models_are_worst_in_among_all/pcba7ra/)
- Praise, 2026-09-23, r/google_antigravity (Reddit): “i was pleasantly surprised at the ui quality with 3.8 flash in antigravity.” [source](https://www.reddit.com/r/google_antigravity/comments/1wni05g/we_need_a_usage_reset_now/pbjrgzs/)
- Praise, 2026-09-23, r/google_antigravity (Reddit): “i guess antigravity is much needed when you need speed, or really develop ui even now i believe gemini models beats any model in frontend design. i just give subagents a role and use multi subagents to complete it” [source](https://www.reddit.com/r/google_antigravity/comments/1wo8qye/what_is_your_experience_with_boost_deep_reasoning/pbl19mh/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “you mean those pleasant look authored ones you see from claude? i don't think comfyui has anything to do with that. antigravity and gemini is going to struggle to make anything that good until deepmind pulls its thumb out.” [source](https://www.reddit.com/r/google_antigravity/comments/1wr9gwg/how_can_i_make_a_short_animation_video/pcarsj8/)
- Complaint, 2026-09-27, r/google_antigravity (Reddit): “wouldn't say so, mostly flashy gradients and way too much detail” [source](https://www.reddit.com/r/google_antigravity/comments/1wrbsby/antigravitys_gemini_models_are_worst_in_among_all/pcbdm9e/)
- Complaint, 2026-09-24, r/google_antigravity (Reddit): “i started using agy cli from last 10 days . it's working fine but takes too much time and too manny steps like repeated read repeated bash . but overall result is good compare to sonet 5 or luna. but frontend design is not that great. yes it's not fast but manageable. and cheap compare to other frontier models.” [source](https://www.reddit.com/r/google_antigravity/comments/1woea7p/30_minutes_for_a_single_prompt_20_40_from_the_5/pbq0eaa/)

### Cursor

- Praise, 2026-09-24, @cursor_ai (X): “@imranmohsin18 @cursor_ai better on cursor, at least on the ui.” [source](https://twitter.com/2059149475503833088/status/2103115081747677346)
- Praise, 2026-09-22, r/cursor (Reddit): “everyday. pretty happy with it in frontend dev. viraui is made with auto, strong specs and rules.” [source](https://www.reddit.com/r/cursor/comments/1wn3rch/does_anyone_use_auto/pbc3yy2/)
- Praise, 2026-09-22, @cursor_ai (X): “@cursor_ai the live css tailwind/shadcn actually looks fairly decent! color me impressed! i like this workflow, take a screenshot feed it to qwen image edit, take the output and feed it into cursor. #uidesign <strict_link>” [source](https://twitter.com/7215722/status/2102300694485278807)
- Complaint, 2026-09-27, @cursor_ai (X): “one of the better agent design platforms i have used is @superdesigndev. while @cursor_ai does a pretty good job with ui, it just can't implement the vision i have for some of my sites. what is amazing is that once you add the @superdesigndev connector, you can spin up a @cursor project, this is basically the "project lead", and it leverages @superdesigndev amazingly well. it can take my natural language prompt, and using superdesign, can turn th” [source](https://twitter.com/2064139210118852608/status/2104274405287403530)
- Complaint, 2026-09-26, r/cursor (Reddit): “the last 2 days i tried to change a complete ui with 4.7 and it was disastrous. not even managing to put textsize, padding, or matching colors correctly, much less random clipping of panels, randomly text clipping out of buttons, not being able to align things, not being able to understand mdc instruction files, etc. actually today 4.6 is cleaning up the whole day behind 4.7s mess. im actually shocked, i at least expected it to perform similar. b” [source](https://www.reddit.com/r/cursor/comments/1wqhcse/grok_46_vs_47/pc5lsop/)
- Complaint, 2026-09-25, r/cursor (Reddit): “cursor got me good...i've used the $20 plan on and off for a year or so. then chatgpt/codex started driving me crazy with their shrinking of value for token usage (while saying the opposite publicly)...anyway that drove me running back into cursor's hands for a few huge projects i'm working on. i did the $20 plan, then upgraded to $60, and finally bit the bullet for the $200 plan a couple days ago. for coding, composer 2.5 will get you anywhere y” [source](https://www.reddit.com/r/cursor/comments/1wpm09i/beginner_making_a_react_native_app_is_cursor_pro/pbxaps8/)

### OpenAI Codex

- Praise, 2026-09-27, r/codex (Reddit): “luna-6 has been great for me building react web apps.” [source](https://www.reddit.com/r/codex/comments/1wrb9p9/luna_6_isnt_as_bad_as_the_whiners_say/pcbqnca/)
- Praise, 2026-09-27, r/codex (Reddit): “i keep seeing these posts, but it sure is my fault given i interact with them. in my opinion codex has been better at both engineering code, designing ui and usage limits since gpt 5.5, now that opus 5.5 is out anthropic has the lead. it's just a catch up game, we should love competition. it's not like openai is falling behind every time more and anthropic doesn't have competition, ok to discuss about the sota and best one to use for what, but co” [source](https://www.reddit.com/r/codex/comments/1wqjtee/20_codex_vs_claude_comparison_from_a/pcbtq13/)
- Praise, 2026-09-27, r/codex (Reddit): “astra is best consumed for graphics and design imo its toe to toe with 5.6 in programming anyways.” [source](https://www.reddit.com/r/codex/comments/1wrfke3/ranking_and_usage_of_models_based_on_experience/pcc279n/)
- Complaint, 2026-09-27, r/codex (Reddit): “building a rocket? opus 5.5 is better for ui, game design, video/image design, copy, and marketing. astra might be better for 3d design, and maybe at high levels on pure code, but also at a much higher price per deliverable.” [source](https://www.reddit.com/r/codex/comments/1wrcc9j/astra_minor_astra_61_and_devday_we_see_50/pcbyx26/)
- Complaint, 2026-09-27, r/codex (Reddit): “that is some bs, bro if you actually do ui with 6 series (except astra) all of them are shitty, even gemini 3.8flash on high produces better results” [source](https://www.reddit.com/r/codex/comments/1wnlfbf/gpt6sol_high_vs_xhigh_where_is_the_sweet_spot/pcc2pn4/)
- Complaint, 2026-09-27, r/codex (Reddit): “long time fan of oai (because it's clear their teams are very talented and the company's culture is good) and don't like anthropic, but the optics for oai don't look good and i'm disappointed. it's been more than a year since oai models started being great at coding, their thoroughness and intelligence were clearly better than claude, they're less prone to accumulating bugs. yet a year later they didn't close the gap at all in the things claude w” [source](https://www.reddit.com/r/codex/comments/1wr1oir/they_are_aware_and_working_on_it_apparently_just/pccon7k/)

### Devin

- Praise, 2026-09-24, @DevinAI (X): “@jaredpalmer @devinai the transition to 3d actually worked for once” [source](https://twitter.com/1969202289174069249/status/2103190943889584238)
- Praise, 2026-09-16, @DevinAI (X): “@alekszuravlovs @devinai i got the max plan till the end of the month. for this deployment pipeline devin decomposed the script (5.5k lines) which was built over some time outside devin using other tools. devin really helped in improving the layouts and the ui/ux overall and refactoring the codebase.” [source](https://twitter.com/1878743506094895104/status/2100263551336382864)
- Praise, 2026-09-13, @cognition (X): “@cognition it's really good at design so far. first issue i've noticed: the context window seems to be 262k, not 1m. is there a way to increase it? also, where can i check my token usage? i'll keep pushing it with backend work and harder tasks. let's see how it holds up.” [source](https://twitter.com/1206192024719892480/status/2098955449970151434)
- Complaint, 2026-09-23, @DevinAI (X): “@princeradebe @devinai i've been battling the icon to just scale properly 🤦♂️” [source](https://twitter.com/1593138580288856064/status/2102731017832612199)
- Complaint, 2026-09-15, @cognition (X): “@ashwinvel94 @rezoundous @cognition @opencode i tried swe-2 but it sucks with design.” [source](https://twitter.com/2371185438/status/2099850073631064443)
- Complaint, 2026-09-15, @cognition (X): “@cognition your ui looks vibecoded sorry <strict_link>” [source](https://twitter.com/2238144607/status/2099933145923567950)

### Conductor

- Praise, 2026-09-13, @conductor_build (X): “@dipxsyy @conductor_build doing my waku fork also but yours looks more polished mine just bug fixes and update features <strict_link>” [source](https://twitter.com/3293833843/status/2098995868242186553)
- Praise, 2026-09-11, @conductor_build (X): “same, i'm still on @conductor_build because t3 code feels vibe-coded, some design choices are not polished and i see big diff in ui/ux for conductor. but they recently introduced pro plan and gated remote and iphone app to it, so looking for alternatives. paying 50$ for things i don't need is not worth it.” [source](https://twitter.com/258280288/status/2098435805714727400)
- Praise, 2026-09-08, @conductor_build (X): “@paper @conductor_build @wisprflow @googlechrome the part that got me: he designs a row of scrolling testimonial cards, and the ai adds the hover behaviour on its own. the row slows and stops when you point at it. it's real code, so there's no component variant to build. it just works. <strict_link>” [source](https://twitter.com/56107683/status/2097286215821398293)
- Complaint, 2026-09-26, @conductor_build (X): “@kyberpez @itsvlady @conductor_build my only use case for astra is backend logic tasks because the design ability is so unbelievably poor” [source](https://twitter.com/1848821469108965376/status/2103827278740251086)
- Complaint, 2026-08-31, @conductor_build (X): “@conductor_build + @poteto pstack = software factory waiting on a better frontend iteration experience but great work @charlieholtz and team!!” [source](https://twitter.com/437086246/status/2094321848112840852)

### Pi

- Praise, 2026-09-22, @pidotdev (X): “@pidotdev peak web design btw” [source](https://twitter.com/2070803172537335808/status/2102478798717342016)
- Praise, 2026-09-22, @pidotdev (X): “@pidotdev lowkey keep the web page's design like this” [source](https://twitter.com/1759657550872485889/status/2102486930097348670)
- Praise, 2026-09-01, @pidotdev (X): “you don't need expensive models like fable or gpt-5.6 sol to turn any ui design image into a highly complex, pixel-perfect ui. i can get this done using just gpt-5.6 luna. i'm using pi @pidotdev + the pixel perfect skill. start with image gpt → pi harness + pixel perfect skill. turns out, the model isn't always the limitation. the right harness + skill can make a huge difference 👌” [source](https://twitter.com/2083448877928312832/status/2094695540978331748)
- Complaint, 2026-09-15, r/PiCodingAgent (Reddit): “guess i was blind lol, i build thru xcode and noticed a bug right after startup, the content displayed is delayed by a few steps, ui looks nice but the minimum width of the columns are interfering with view sometimes, just my 2 cents” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wgyjic/pie_native_macos_client_for_pi_coding_agent/p9zcq3b/)

### Zed

- Praise, 2026-09-16, @zeddotdev (X): “@taniyatweets_ zed it's just unrealistic fast for ui/ux @zeddotdev they just built different” [source](https://twitter.com/2242568539/status/2100104182619631865)
- Praise, 2026-09-07, @zeddotdev (X): “@ikhwanuddin @zeddotdev for a ui reason gemini 3.8 enak buat frontend dan kuotanya abisnya lebih lamaaa hahaha” [source](https://twitter.com/131447449/status/2096959789528138219)
- Praise, 2026-09-07, @zeddotdev (X): “@ikhwanuddin @zeddotdev beberapa orang anggap gemini ampas. jujur sih iya kalo diajak diskusi buat plan and implementasi ke backend karena ga bisa scanning dari banyak pov. thinkingnya masih kalah sama model lain. tapi kalo buat ui dia lebih smooth, aku suka.” [source](https://twitter.com/131447449/status/2096960826834141339)
- Complaint, 2026-09-17, r/ZedEditor (Reddit): “meanwhile i am still waiting for a color picker in zed, as a frontend dev i am really missing it.” [source](https://www.reddit.com/r/ZedEditor/comments/1wie2h8/animated_cursor_trail_added_to_zed/paeupsa/)

### Factory

- Praise, 2026-09-17, @FactoryAI (X): “only downside is the 5 hours limit and the price which is pretty fair but still a little for me personally. other than that i can list so many things i love about it. droid is super efficient and often finish tasks faster than most other agent with similar results. i feel like it gets the right context at the right time. it’s pretty amazing. also love the byok, live the fact that ui almost always looks better when done with droid even using th” [source](https://twitter.com/1617212256487411712/status/2100703532785541412)
- Praise, 2026-09-17, @droid (X): “i've been having a lot of success recently doing ui work with @droid and gemini 3.8 flash. it consistently is better than fable 5.1 and costs a fraction.” [source](https://twitter.com/1247892463479451653/status/2100622735273545963)
- Complaint, 2026-09-02, @FactoryAI (X): “@gitmaxd @factoryai @droid and @factoryai is an extremly good choice as your harness as well, might even be the best! the only thing stopping me from fully swithching is the lack of frontend annotations! hello!!!” [source](https://twitter.com/171899126/status/2095083387945959438)

### GitHub Copilot

- Praise, 2026-09-07, r/GithubCopilot (Reddit): “for front-end gpt codex and above are better. for back-end opus models are better in architecture and implementation imho.” [source](https://www.reddit.com/r/GithubCopilot/comments/1w9jvre/which_agent_is_the_best_for_coding_and_debugging/p8az37s/)
- Complaint, 2026-09-18, r/GithubCopilot (Reddit): “i used luna the last couple of weeks and it was great... is it just me or got it dumber!? it struggles to keep ui/uix consistent and totally forgot about correct ui design principals. luna is working againt me since days instead for me lol” [source](https://www.reddit.com/r/GithubCopilot/comments/1wdtxyz/luna_is_still_the_goat/paj09ke/)

### Amp

- Complaint, 2026-09-24, @AmpCode (X): “mainly when i throw in a new idea, hand it a design mockup, or ask it to refactor something big, it tends to make a mess. breaks things, ignores instructions. these are problems most harness solved earlier this year, but amp's harness hasn't caught up. also, no plan mode. for larger tasks the model still needs to plan before it acts. amp has oracle but it wasn't enough, i ended up writing my own planning skill to compensate. cc handles this nativ” [source](https://twitter.com/1592160489965948933/status/2103099981473382778)
- Complaint, 2026-09-18, @AmpCode (X): “@sqs @ampcode sorry amp cloned the design a bit too hard.😂” [source](https://twitter.com/942412840673169408/status/2100930544041394454)

### Warp

- Praise, 2026-09-09, @warpdotdev (X): “incredible blog post from @jerrydizs on @warpdotdev’s internal design tool, highly recommend skimming * mocking graphql queries so that design prototypes are production-ready once approved * built-in self-improvement loops as a product engineer, a large part of my job these days is envisioning what the right ux for something might look like. i don’t have a background in design, and generally want to come up with a thesis before pinging our design” [source](https://twitter.com/1157764163973844992/status/2097514947508879656)
- Complaint, 2026-09-17, @warpdotdev (X): “warp, do you not use your own app, or have you never used grok build? it's really ugly... @warpdotdev <strict_link>” [source](https://twitter.com/1684821659214385152/status/2100488508196733433)
