# Memory and state carried across sessions (`context.session_memory`)

Feedback Bench, coding agents, built 2026-10-01, window 2026-08-31 to 2026-09-27. Web page: https://feedbackbench.com/#/criterion/context.session_memory

Area: [Instructing and context](https://feedbackbench.com/criteria/context.md)

**Definition.** Whether knowledge persists correctly between sessions: memory files, re-discovery cost at session start, stale memories, and leakage between sessions.

**Boundary.** Not this: see [Saving, switching, resuming and rewinding sessions](https://feedbackbench.com/criteria/ui.session_history.md) for saving and resuming chat transcripts. Not this: see [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md) for user-written rules.

Rated author-weeks, all agents: 833. Complaint share: 53%.

## The brief

Written by Claude Opus 5.5 from 86 labelled posts and the numbers on this page. Interpretation, not measurement: every quote is verbatim and links to its post.

**Every agent forgets between sessions; users build the memory themselves.**

TL;DR:

- Cold starts dominate. Users re-explain decisions, plans and ruled-out approaches at the start of each session.
- Google Antigravity is the only agent rated worse than peers. Interrupted runs restart instead of resuming.
- The working fixes are user-built: progress files, handoff notes, and hooks that force memory writes.

In plain terms: Expect to start most sessions by reminding the agent what happened last time. Built-in memory helps but can go stale, ignore instructions, or bleed across tasks. Users who keep their own progress and handoff files report the smoothest restarts.

### How it breaks

- **The Monday morning re-briefing** ([Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md)). The most common complaint is simple: a new session starts blank, and users pay a re-explanation tax before any work happens.
  Users describe pasting the same notes into every new chat. Decisions, sign-offs and promises live outside git, so the agent never sees them. Switching tools makes it worse because nothing useful comes along. Built-in persistent memory across sessions is the single most requested change, and it is asked of almost every major agent. Users treat the time it takes to get back up to speed as the real test of an agent.
  Evidence:
  - Complaint, Cursor, @cursor_ai, 2026-09-11: “@cursor_ai @bot the real benchmark: can i come back on monday without spending 20 minutes reminding it what we were doing on friday?” [source](https://twitter.com/2066846062623518720/status/2098259643763974226)
  - Complaint, Cursor, r/cursor, 2026-09-16: “curious what people are doing for this. cursor in the customer repo is fine for me. code isn't the hard part. it's the other crap. who signs off. what we already decided. what we promised. whether they actually accepted a number or we're just repeating it. none of that is in git, so every new chat i'm pasting notes again. tried dumping more into rules/docs. got messy. also don't love parking client detail in their repo for nda reasons. i finally stopped rebuilding this every time and put it here: [<strict_link> what's your setup? mine still feels half-improvised.” [source](https://www.reddit.com/r/cursor/comments/1whn59g/where_do_you_put_stakeholderdecision_notes_when/)
  - Complaint, Cursor, r/cursor, 2026-09-25: “limits feel tighter on cursor if fast/priority stays on and every turn keeps a huge context. i stay closer to a codex-like budget by planning on a slower model and only bursting when the task needs it. the other tax when switching codex → cursor is cold start — none of the useful decisions come along. i keep a small on-disk resume the next agent has to read first: <strict_link> pipx install portable-resume” [source](https://www.reddit.com/r/cursor/comments/1wptv9n/thinking_of_switching_from_codex_to_cursor/pbz3l3j/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-04: “you need the ai to have the context. there is no better way to do it. you can remember sure but that doesnt help the ai. and you explain everything when you start a new session?” [source](https://www.reddit.com/r/ClaudeCode/comments/1w6i5f4/havent_written_code_by_hand_all_year_heres_what_i/p7tgtsu/)

- **Handoffs keep state, lose reasoning** ([Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md)). Summaries carry over what changed but drop what was tried and rejected, so the next session walks back into the same dead ends.
  One user notes that files changed and tasks done survive a handoff, but the reasoning dies with the old session. The fix that works is cheap. The outgoing agent writes a short list of what it ruled out and why. Others want handoffs to pass on acceptance checks and skipped hypotheses, not just a transcript. Structured handoff between sessions is a recurring request, concentrated on Claude Code.
  Evidence:
  - Complaint, Cline, @cline, 2026-09-14: “@cline importing task context across claude code and codex is more useful than another model picker: the harness can stay stable while models rotate. a compact handoff summary plus acceptance checks would help a new model inherit the contract, not just the transcript.” [source](https://twitter.com/2061779435557117952/status/2099554531998654777)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-10: “the handoff at 30% is the part worth attacking. what survives a successor handoff is state: files changed, tasks done. what dies is the reasoning, specifically which approaches were already tried and rejected and why. the successor then walks straight back into the same dead ends, which costs more than the handoff saved. what fixed it for me was making the outgoing agent write a short "what i ruled out and why" list as its final act, separate from the status summary. three or four lines. cheapest thing in the pipeline and it stops the loop. on bypass permissions: agreed, that is the right call once the sandbox is the actual boundary. the permission prompt was never a security control. worth stating that explicitly wherever you document this, because people will copy the bypass flag without the sandbox half.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wc4u51/save_your_tokens_from_using_auto_mode_and_do_this/p8vbei4/)
  - Complaint, Kiro, @kirodotdev, 2026-09-21: “@kirodotdev long-running agent sessions need a trace receipt beside the model label: session id, root-cause file, tool calls, skipped hypotheses, patch diff, and test result. otherwise context retention is just a nicer fog machine.” [source](https://twitter.com/2013700835654672388/status/2102071866823430275)
  - Complaint, OpenCode, r/opencode, 2026-09-09: “i have used opencode for a long time, and the biggest problem i had with it is that once i created a plan and started implementing it step by step, the plan was forgotten. so, a better approach than writing every plan in a .md file would be to create a shared cli between agents with skills to use and everything. now my workflow is: plan with luna, using the caveman and ponytail plugins. add the plan to taskwatch (the cli tool). in build mode, i just use the command /taskwatch-next, which selects the best-fitting task to do now based on urgency. i'm sure y'all could find better usages for my cli/tui app. it would help if you could star it: [<strict_link> edit : this can also be used as a shared memory between tools ( codex, opencode , etc..)” [source](https://www.reddit.com/r/opencode/comments/1wbkw5v/i_think_i_might_have_discovered_the_way/)

- **Memories that refuse to die** ([Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md), [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md)). Stored memory can be wrong in both directions: abandoned plans resurface as pending, and explicit corrections get ignored after a reset.
  Posts show unfinished plans being read as pending work rather than dropped work. Others report long-forgotten suggestions reappearing unprompted. One user added a memory telling the agent to avoid a known bug, and the bug came back after the reset anyway. Some users have switched automatic memory off entirely and rely on written rules instead, because they do not trust outdated memory artifacts.
  Evidence:
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-06: “mine kept the abandoned plans, it reads unfinished as pending rather than dropped.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w8nhs3/how_do_you_decide_an_agentwritten_doc_is_dead/p84j1cy/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-12: “i hate that reddit bot... hello to everybody. i need your help. about a week ago i made a template for an ui with astra. i have initiated to implement the new ui 5 day ago. in the first day the was a recurrent bug in the impementation, after some investigation with astra i gave found the problem, insered a memory to not do that thing and the problem was solved for few days: today after the reset it's all the day that sol and astra continue to implement that bug even with the memory to not do that. have you ant advice?” [source](https://www.reddit.com/r/codex/comments/1tjfxcf/anyone_else_ask_here_about_current_codex_issues/p9eg8rb/)
  - Complaint, Devin, @cognition, 2026-09-14: “@devindesktop @windsurf @cognition @_akhaliq @rahyengan @themidasproj @ylecun @rowancheung 3. @devindesktop but later that day it automatically suggested what it should have long forgotten out of nowhere <strict_link>” [source](https://twitter.com/2098243273890705408/status/2099333865412456929)
  - Complaint, GitHub Copilot, r/GithubCopilot, 2026-09-07: “i haven't dived into indexing and memories per se. my main focus has been writing proper custom instructions so it knows when to pull in the appropriate context data. i don't want it relying on incorrect or outdated memory artifacts” [source](https://www.reddit.com/r/GithubCopilot/comments/1w9qlcz/copilot_in_vs_code_compared_to_cli_and_the_app/p8d8ohd/)

- **Context bleeding across tasks** ([Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md)). Leakage is rarer than forgetting but more alarming: parallel chats mix requests, and one resumed session edited a project that was not the user's.
  Users report requests from separate chats being merged when the chats run at once. One OpenCode user resumed a session and found it editing code under a home directory they do not have. Others want the opposite of more memory. They ask for a switch to disable stored knowledge and history so old context stops inflating new chats. Users also ask for memory scope controls and opt-outs.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-24: “it also mixing your requests between separate tasks in chat! when tasks in different tpics and chats if it start at once.. it will mix! i noticed it several days ago...” [source](https://www.reddit.com/r/codex/comments/1wp4nan/i_dont_trust_gpt_to_do_multiple_tasks_at_once/pbsjggn/)
  - Complaint, OpenCode, r/opencode, 2026-09-19: “same here, mine was like chinese then after resuming it was in english but editing a project i dont know from a machine i don't know... my machine path is /home/george/... but it was editing a code from /home/sam/... i have no machine in my home with that username whatsoever...” [source](https://www.reddit.com/r/opencode/comments/1wk4km1/weird_thing_just_happened/paod16n/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-16: “any update to cli? or its always agy 2.0? cant do delete conversation automatically on every time i quit the app on the app one while in cli, i could use powershell profile to do so. maybe a request to add like disabling knowledge and conversation history like in antigravity ide. i dont need bloated context from the previous chat like tha chatbot app.” [source](https://www.reddit.com/r/google_antigravity/comments/1wh9uyh/antigravity_2_release_v2140/pa3aiha/)
  - Praise, Cursor, r/cursor, 2026-09-13: “this is a good thing in my opinion. too much memory/hitory just bloats your context and confuses the model. better to dial in your agent file and maybe a small subset of rules if it’s messing anything up consistently. this has worked fine for me so far.” [source](https://www.reddit.com/r/cursor/comments/1wex8ck/cursor_has_no_memory_between_sessions_and_its/p9hxmmh/)

- **Users roll their own memory** ([Memory and state carried across sessions](https://feedbackbench.com/criteria/context.session_memory.md), [Persistent project rules files are read and obeyed](https://feedbackbench.com/criteria/context.instruction_files.md)). The setups people praise most are hand-built: agent-owned notes files, source-of-truth handoffs, and hooks that force a save after each unit of work.
  Praise in this category often credits the user's scaffolding, not the product. Patterns repeat. Users keep a progress file split into done, in progress and next. An architect agent reads the docs first in every session. A project skill plus handoff file lets a fresh task take over cleanly. Several users want memory that works across tools so they can move between Codex, Claude Code and Pi without losing context.
  Evidence:
  - Praise, OpenAI Codex, r/codex, 2026-09-01: “i have my agents self document (in file i let them control that i don’t commit) that way they can learn from each other. it also helps them getter better and more efficient as they know the codebase better and such” [source](https://www.reddit.com/r/codex/comments/1w4rcqr/how_do_you_stop_known_coding_mistakes/p79jaik/)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-08: “i use [<strict_link> and a hook in the claude config to force it to save memories into that after every time it completes a bit of work. i also instruct it to proactively look in there for relevant memories. this works very well for me, even if it comes in fresh or without cache, it is effective to find the relevant context without needing to load in full context again.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wao464/is_there_a_way_to_avoid_the_huge_15_of_5h_usage/p8jlny4/)
  - Praise, OpenAI Codex, r/codex, 2026-09-14: “codex compacts the conversation before the context literally fills, but i never rely on chat history as project memory. each project has its own skill and a small set of source-of-truth files that are continually updated. if the command centre starts losing track or hallucinating, i get codex to check the actual repo state and update the existing handoff with all important decisions, test results, blockers and next actions. only after confirming that everything needed is safely recorded does codex archive or delete the redundant task and spin up a fresh command centre or both. the new task reads the project skill and handoff, then continues with the existing worker unless that worker also needs replacing.” [source](https://www.reddit.com/r/codex/comments/1wfvg6c/my_pro_20x_lasts_6_to_7_days_using_this_two_chat/p9prdl5/)
  - Praise, Pi, @pidotdev, 2026-09-22: “1. harness agnostic: if you're like me, you have a subscription to a handful of different model providers. for me, i have claude, chatgpt, and @zai_org (which i've been using in @pidotdev and have loved). when new models drop, i like testing these on my own work, but wanted to level the playing field when testing between each. having a memory system that is agnostic to the scaffolding around the model has allowed me to perform unbiased testing for myself between each. what i've loved with this is that i can end a session in codex and pick up right where i left off in claude code or pi without losing the context (and not just relying on git history)” [source](https://twitter.com/1839038985441804288/status/2102435725585371604)

### Who stands out

- **Google Antigravity (weaker)**. The only agent rated worse than peers here, with complaints about runs that restart rather than resume and chats that lose their own history.
  Users report that an interrupted teamwork run spins up new agents and discards the work of the old ones, recovering only from markdown notes. Others say a new chat does not recall what was executed before. One user describes context being wiped mid-conversation. The praise that exists credits state kept in local folders and a handoff skill. Users explicitly name a rival as better at cross-conversation recall.
  Evidence:
  - Complaint, Google Antigravity, @antigravity, 2026-09-01: “i wish @antigravity remembered stuff from other conversations and other projects and i wish @geminiapp could remember stuff from other conversations more effectively. chatgpt/codex does that very well.” [source](https://twitter.com/1820644530653179904/status/2094844957316055302)
  - Complaint, Google Antigravity, @antigravity, 2026-09-02: “@rodydavis @soso_fun_yt @antigravity right now if /teamwork-preview gets disrupted it does not pick off from where it left off, it creates entirely new agents and kills the ones that had done a ton of work and starts from scratch only utilizing context and .md's so its not a 1:1 smooth pick off, pls look into this!” [source](https://twitter.com/1964269084826357767/status/2095297834253992168)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-08: “yeah but if i try to work with new chat.. it does not remember everything what execution it did etc” [source](https://www.reddit.com/r/google_antigravity/comments/1wajlb2/antigravity_ide_has_been_very_slow_this_past_few/p8io53g/)
  - Complaint, Google Antigravity, r/google_antigravity, 2026-09-25: “am not getting this exact issue, but am getting issues since yesterday where models are erasing their contexts themselves, even on gemini app, not just antigravity. they are totally wiping out their memory of previous chat. on the app it's much worse, i ask something, it gives an answer, then i ask a follow up question, it says 'for what?' like it forgot the previous answer itself gave just a chat ago. something big is broken inside. am still facing this contax getting wiped issue for past 24 hours now” [source](https://www.reddit.com/r/google_antigravity/comments/1wpsvzy/antigravity_scheduled_tasks_failing_with_no/pbyxiea/)

- **Cursor (mixed)**. Persistent project threads earn real praise, but users want to see and audit what the agent accumulates.
  Users welcome a single persistent coordinator thread in place of dozens of dead chats, and shared memory that lets subagents resume without re-reading the repo. The open questions are about visibility. Users ask how the coordinator decides what to drop, and where they can inspect the knowledge it mines. Cursor leads requests for memory visibility, audit and editing, and for project-scoped memory.
  Evidence:
  - Praise, Cursor, @cursor_ai, 2026-09-10: “@cursor_ai @bot shared memory plus artifact sync lets subagents resume mid-task without re-reading the whole repo each time” [source](https://twitter.com/1945115184072105984/status/2098166845916553323)
  - Praise, Cursor, @cursor_ai, 2026-09-11: “@cursor_ai @bot projects with a coordinator agent that stays on and spins subagents is the shape i wanted. one persistent thread that remembers plans and demos beats opening a fresh chat for every ticket.” [source](https://twitter.com/38129147/status/2098535594452218032)
  - Praise, Cursor, @cursor_ai, 2026-09-11: “@cursor_ai @bot one persistent thread instead of 40 dead chats is the real upgrade here 🧵 curious how the coordinator decides what to drop though, that's usually where always on agents start drifting 👀” [source](https://twitter.com/2085304778427621376/status/2098272297945903588)
  - Complaint, Cursor, @cursor_ai, 2026-09-12: “the new projects from @cursor_ai are very good. but, there are some things that i'm not completely convinced about: 1) as it sends prs, many end up with merge conflicts and it doesn't automatically babysit to resolve them. 2) i would like to be able to see and audit somewhere the "skill mining" or knowledge that it accumulates from the conversation. 3) at least in the app, i'm not quite understanding how to use design mode with the projects. it's something i use to move much faster with front tweaks.” [source](https://twitter.com/766336330397777920/status/2098713395771879802)

- **Claude Code (mixed)**. The most discussed agent here. Saved preferences pay off, but memory placement and stale plans trip users up.
  Users credit accumulated memories of their preferences for better results than a fresh model gives. Power users add hooks that force memory saves and keep progress files to survive usage resets. The complaints are about mechanics. Memories land in inconsistent locations and some do not persist, and abandoned plans linger as pending. Claude Code draws the most requests for built-in persistent memory and structured handoffs.
  Evidence:
  - Praise, Claude Code, r/ClaudeCode, 2026-09-06: “wow, imagine that. a model from a company whose been working in your codebase, has many saved memories about what you prefer, what to do and what not to do, is better than a model you just started working with.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w8u4jm/i_canceled_claude_because_i_wanted_to_test_astra/p867ev2/)
  - Complaint, Claude Code, r/ClaudeCode, 2026-09-20: “yes. consistency is a big advantage. also so is the save location. i’ve been caught out (without the skill) by claude picking random locations and some of them haven’t persisted.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wlalc7/whats_even_the_point_of_clear/pb1mydq/)
  - Praise, Claude Code, @ClaudeDevs, 2026-09-25: “@claudedevs good change. having claude keep a short progress file as it works (done, in progress, next step) makes picking the task back up after the reset a lot smoother.” [source](https://twitter.com/2103126608684929024/status/2103575676519682251)
  - Praise, Claude Code, r/ClaudeCode, 2026-09-08: “i use [<strict_link> and a hook in the claude config to force it to save memories into that after every time it completes a bit of work. i also instruct it to proactively look in there for relevant memories. this works very well for me, even if it comes in fresh or without cache, it is effective to find the relevant context without needing to load in full context again.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wao464/is_there_a_way_to_avoid_the_huge_15_of_5h_usage/p8jlny4/)

- **OpenAI Codex (mixed)**. Some users find its continuity better than Claude's, but memory ships off by default and saved corrections do not always stick.
  One user downgraded a Claude plan partly because of Codex's memory and continuity. Another describes it picking up a large multi-service project where it left off. On the other side, users are surprised that memory must be enabled manually. One reports a stored memory being ignored after a reset. Others say requests from separate parallel tasks get mixed together.
  Evidence:
  - Complaint, OpenAI Codex, r/codex, 2026-09-23: “have you know codex disabled memory by default, you have to manually enable it, while claude doesnt ?” [source](https://www.reddit.com/r/codex/comments/1wnx6ww/codex_is_in_a_really_bad_spot_right_now_opus_55/pbix8f0/)
  - Praise, OpenAI Codex, r/ClaudeCode, 2026-09-07: “claude is somewhat falling out of favor with me also. i've had a 20x max subscription since february that i just downgraded to a 5x subscription in favor of a codex 20x instead. the fable models have been the only good ones since opus 4.6, but even they have such poor memory and continuity compared to chatgpt that they don't stack up well.” [source](https://www.reddit.com/r/ClaudeCode/comments/1w8u4jm/i_canceled_claude_because_i_wanted_to_test_astra/p8ajsu2/)
  - Praise, OpenAI Codex, r/codex, 2026-09-04: “astra simply picked up where it left off on my project, which consists of 19 microservices. the goal is to move the business logic from one service to two others, with communication between them over grpc.” [source](https://www.reddit.com/r/codex/comments/1w7e6i2/astra_cost/p7u91nh/)
  - Complaint, OpenAI Codex, r/codex, 2026-09-12: “i hate that reddit bot... hello to everybody. i need your help. about a week ago i made a template for an ui with astra. i have initiated to implement the new ui 5 day ago. in the first day the was a recurrent bug in the impementation, after some investigation with astra i gave found the problem, insered a memory to not do that thing and the problem was solved for few days: today after the reset it's all the day that sol and astra continue to implement that bug even with the memory to not do that. have you ant advice?” [source](https://www.reddit.com/r/codex/comments/1tjfxcf/anyone_else_ask_here_about_current_codex_issues/p9eg8rb/)

### Fine print

- Evidence covers a single month of posts. Many agents have too few posts to rank, so their quotes are illustrative only.
- Several complaint posts link to the author's own memory tool, which may overstate how often the problem occurs.
- Praise often credits user-built scaffolding rather than the product's own memory features.

## Top requests

What users ask to add or change, most asked first. 208 author-weeks ask for something. Requests do not change the Feedback Score. Rule: A separate pass by Claude Sonnet 5 reads every counted post and extracts what the author asks the agent or its vendor to add or change, with the criteria it maps to and a short normalised wording; it does not touch the labels or the Feedback Score. Claude Opus 5.5 groups the wordings within each criterion (the first criterion the request maps to) into themes; code counts them. A theme counts distinct author-weeks that ask for it, per agent; across agents, one author-week per agent. Themes asked in fewer than 2 author-weeks, and requests that share no theme, are not shown. Examples: up to 3 posts per theme from different authors, without slurs, preferring posts of 60 to 450 characters, most recent first.

| Rank | Request | Author-weeks | Posts | Agents (author-weeks) |
|---|---|---|---|---|
| 1 | Built-in persistent memory across sessions | 31 | 32 | Claude Code 13, Cursor 7, OpenCode 4, Google Antigravity 3, OpenAI Codex 3, Pi 1 |
| 2 | Persistent project-scoped memory and projects | 19 | 19 | Cursor 7, Claude Code 5, OpenAI Codex 4, Google Antigravity 1, Factory 1, OpenCode 1 |
| 3 | Structured handoff between sessions | 14 | 14 | Claude Code 8, OpenAI Codex 3, Cline 1, Cursor 1, OpenCode 1 |
| 4 | Shared memory across different agent tools | 12 | 13 | Cursor 5, Claude Code 4, OpenAI Codex 2, Devin 1 |
| 5 | Memory visibility, audit and editing | 11 | 12 | Cursor 8, OpenAI Codex 2, Google Antigravity 1 |
| 6 | Search and reference past conversations | 10 | 11 | OpenAI Codex 4, Google Antigravity 2, Claude Code 2, Cursor 1, OpenCode 1 |
| 7 | Resume work after usage limit reset | 9 | 11 | Claude Code 6, Google Antigravity 1, OpenAI Codex 1, Cursor 1 |
| 8 | Memory scope boundaries and opt-out controls | 8 | 10 | Cursor 3, Google Antigravity 2, Claude Code 2, OpenAI Codex 1 |
| 9 | Forgetting and retiring stale memories | 8 | 8 | Claude Code 3, OpenAI Codex 2, Cursor 1, Devin 1, OpenCode 1 |
| 10 | Sync sessions and projects across devices | 8 | 8 | Claude Code 3, OpenAI Codex 2, Google Antigravity 1, Devin 1, OpenCode 1 |
| 11 | Avoid re-discovering context at session start | 7 | 7 | Claude Code 3, Cursor 2, Google Antigravity 1, OpenAI Codex 1 |
| 12 | Built-in project progress tracking | 7 | 7 | Google Antigravity 2, Claude Code 2, OpenAI Codex 2, Cursor 1 |

### 1. Built-in persistent memory across sessions

- Claude Code, 2026-09-27, r/ClaudeCode (Reddit): “having a functional second brain that knows everything about me” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrcmjn/what_tool_have_you_built_for_yourself_with_claude/pccc21n/)
- OpenCode, 2026-09-26, @opencode (X): “@thdxr hey @opencode @thdxr please give us more free tier daily and more free models. add buitin memory vault” [source](https://twitter.com/141503294/status/2103816074156495086)
- Cursor, 2026-09-26, @cursor_ai (X): “as i was working in cursor's harness, i had a thought. if @bot has the learn option - can we also apply that logic to the agents as well? if bot has it - it would be waste not to have it in the cursor agent as well. any thoughts? @cursor_ai @poteto @lingxi” [source](https://twitter.com/1854338126774194188/status/2103714346950086977)

### 2. Persistent project-scoped memory and projects

- OpenAI Codex, 2026-09-25, r/codex (Reddit): “you need persistent project memory outside the model. i’m building a knowledge & retrieval engine for exactly this. durable project state, decisions, architecture and history stored separately, then only the relevant context is retrieved for each new session. you could build a lightweight version with codex/claude code using markdown/json/sqlite + retrieval scripts. the model can forget; the project shouldn’t.” [source](https://www.reddit.com/r/codex/comments/1wq1k1i/astra_is_the_smartest_the_model_ive_used_but_it/pc106hu/)
- OpenAI Codex, 2026-09-25, X search: OpenAI Codex, Codex CLI, Codex app (X): “@theo seeing generated images in chat - i do lots of automated app testing. love seeing what the agent is doing while he goes. codex app works best here. another thing which would be nice is create projects, with custom context for handling multiple threads.” [source](https://twitter.com/2012897420955242496/status/2103391499979309246)
- OpenCode, 2026-09-24, @opencode (X): “really enjoying @opencode. i’ve been using gbrain to persist project decisions and context across sessions. what are people using, and what has actually worked well? @thdxr any plans for built-in memory?” [source](https://twitter.com/385457565/status/2103017169944727720)

### 3. Structured handoff between sessions

- Claude Code, 2026-09-25, @ClaudeDevs (X): “@claudedevs how about just a hand off md so i can have another agent or session easily resume?” [source](https://twitter.com/2003361328300457987/status/2103604803783819680)
- Cursor, 2026-09-25, @cursor_ai (X): “@cursor_ai love this for code. still missing for the outside world: durable, cited context the agent can pull next session.” [source](https://twitter.com/1086013144638672896/status/2103599014830637356)
- Claude Code, 2026-09-19, @ClaudeDevs (X): “@claudedevs parallel sessions are the easy demo; the hard-won feature is context handoff that doesn’t quietly rot. if projects can preserve decisions, constraints, and “do not touch prod” across threads, that’s a real team-multiplier—not just concurrency.” [source](https://twitter.com/1863497833556705280/status/2101335295450829097)

### 4. Shared memory across different agent tools

- Claude Code, 2026-09-25, r/codex (Reddit): “but what does it offer? its just a terminal multiplexer for agents no? i still have to use 2 harnesses codex cli and claude code, and each harness uses memory, md files etc differetly, which is a total mess” [source](https://www.reddit.com/r/codex/comments/1wpvp4o/for_people_with_both_codex_and_claude_code_what/pc0ccjo/)
- Cursor, 2026-09-21, @cursor_ai (X): “@clebervisconti @grok @bot @cursor_ai unified billing would remove friction, but it won’t unify state. the painful part is context and permissions drifting across runtimes; a portable workflow manifest may matter more than one invoice.” [source](https://twitter.com/1021426672279564288/status/2101850081624236459)
- Cursor, 2026-09-18, @cursor_ai (X): “@cursor_ai shared memories and generated items will significantly reduce the contextual costs of multi-agent collaboration, but it also requires clear versions, sources, and permissions. especially for planning documents and presentation products, if they cannot be traced back to specific operations and acceptance results, the faster the collaboration, the higher the troubleshooting costs may be.” [source](https://twitter.com/1628996445654188033/status/2100837790720180459)

### 5. Memory visibility, audit and editing

- Cursor, 2026-09-18, @cursor_ai (X): “@cursor_ai projects are useful because shared memory and generated files solve a real pain in agent coding. the risk is stale context across branches. a view of which files, rules, and prior runs shaped a plan would make debugging safer. how does cursor show that history?” [source](https://twitter.com/1654699424969379841/status/2100987185360998473)
- Cursor, 2026-09-15, @cursor_ai (X): “@cursor_ai can you see and edit what the coordinator remembers about the project?” [source](https://twitter.com/1491654782091735041/status/2099926524304453741)
- Cursor, 2026-09-14, @cursor_ai (X): “@cursor_ai sharing memories and generated items is indeed convenient, but if the project-level status is written incorrectly, the pollution can be more hidden than a single conversation. if we could see the source, version, and add a one-click rollback, it would be much more reassuring.” [source](https://twitter.com/2259799350/status/2099627937733505516)

### 6. Search and reference past conversations

- Claude Code, 2026-09-26, @ClaudeDevs (X): “how can't i do something as simple as referencing conversations on claude!!?!??? @anthropicai @claudedevs :(” [source](https://twitter.com/1878192932366209024/status/2103773136449638544)
- Google Antigravity, 2026-09-23, r/google_antigravity (Reddit): “while i think they should add that as a feature currently i have another method which is i have a specific pinned conversation which i ask to find conversations for me, like i just tell the ai to find the conversation where i asked for it to create a specific app and it finds it for me you don't even need to remember the exact words used, you just need to tell it what the conversation is about <strict_link>” [source](https://www.reddit.com/r/google_antigravity/comments/1wnqoya/antigravity_2_release_v2160/pbm3x20/)
- OpenCode, 2026-09-18, r/opencode (Reddit): “please fix or update the search function it's the worst... it cannot find anything. it makes no sense that i have to use big pickle yes cause i don't want use tokens from other llm's to search in my projects for past sessions that build a feature. i should be able to just type: "shock value" and then it shows me the session that had this...” [source](https://www.reddit.com/r/opencode/comments/1wjj7e3/search_function/)

### 7. Resume work after usage limit reset

- Claude Code, 2026-09-26, @ClaudeDevs (X): “@claudedevs instead of allowance it should create sort of handoff file to continue later in new session” [source](https://twitter.com/2060645216718147584/status/2103778053000200213)
- Claude Code, 2026-09-25, @ClaudeDevs (X): “@claudedevs now you just need to have it auto start again for the next session.” [source](https://twitter.com/18287132/status/2103589638346789309)
- Claude Code, 2026-09-25, @ClaudeDevs (X): “@claudedevs humm.. cant it just save the state as much as possible and or give user a todo list plus status what is done midflight? it cud keep on polling the usage limit at a certain interval?” [source](https://twitter.com/2099833526531039232/status/2103566521817870838)

### 8. Memory scope boundaries and opt-out controls

- Claude Code, 2026-09-20, r/ClaudeCode (Reddit): “can we deactivate this? i had a session answer a question i asked in another session, so both answered the question, polluting my context in the non asked question one” [source](https://www.reddit.com/r/ClaudeCode/comments/1wlgfpw/anthropic_has_done_it_again_dont_keep_many_clis/payjyh0/)
- Cursor, 2026-09-20, @cursor_ai (X): “@cursor_ai project-based context is a strong usability win for agentic coding: keeping requirements, artifacts, and feedback together should cut repeated setup. the key next step is making context boundaries and retention explicit so teams can trust what agents see.” [source](https://twitter.com/1208933549081907200/status/2101775839972966509)
- Google Antigravity, 2026-09-16, r/google_antigravity (Reddit): “any update to cli? or its always agy 2.0? cant do delete conversation automatically on every time i quit the app on the app one while in cli, i could use powershell profile to do so. maybe a request to add like disabling knowledge and conversation history like in antigravity ide. i dont need bloated context from the previous chat like tha chatbot app.” [source](https://www.reddit.com/r/google_antigravity/comments/1wh9uyh/antigravity_2_release_v2140/pa3aiha/)

### 9. Forgetting and retiring stale memories

- OpenCode, 2026-09-22, @opencode (X): “@softwaredoug @turbopuffer @opencode agent memory still needs recency rules, even on a good store” [source](https://twitter.com/763249944056565760/status/2102383559311032543)
- Claude Code, 2026-09-17, @ClaudeDevs (X): “@buburdin @claudedevs lol yes, but they keep updating "context" about you without any disposal dates. and give you creepy notifications like "context updated: wife".” [source](https://twitter.com/718427679125598208/status/2100593921071940009)
- Devin, 2026-09-14, @cognition (X): “@devindesktop @windsurf @cognition @_akhaliq @rahyengan @themidasproj @ylecun @rowancheung 3. @devindesktop but later that day it automatically suggested what it should have long forgotten out of nowhere <strict_link>” [source](https://twitter.com/2098243273890705408/status/2099333865412456929)

### 10. Sync sessions and projects across devices

- Claude Code, 2026-09-24, @ClaudeDevs (X): “yeah please stop making things local. @claudeai @claudedevs please make project shared across chats, code, cowork and design....or at least allow them to tag each other.... <strict_link>” [source](https://twitter.com/1913676020/status/2103143406159708543)
- Claude Code, 2026-09-24, @ClaudeDevs (X): “@claudedevs local threads matter for repos that can’t leave the machine. keeping the same project context across local and cloud sessions is what could make this feel seamless.” [source](https://twitter.com/1296669524436148225/status/2103050413335613714)
- Claude Code, 2026-09-17, @ClaudeDevs (X): “@claudedevs give us a way to transfer projects from claude chat to claude code project, please!” [source](https://twitter.com/958671027101454337/status/2100639508794327444)

### 11. Avoid re-discovering context at session start

- Google Antigravity, 2026-09-26, @antigravity (X): “@thsottiaux @antigravity gem re-discovering tui's would be a massive product improvement, imo.” [source](https://twitter.com/1779426928891662336/status/2103956693579346386)
- Claude Code, 2026-09-20, r/ClaudeCode (Reddit): “i understand the logic of it but having new sessions for every new addition or work stream makes things so much slower since it has to initialize and figure out what i mean and find the location of files. also from an organizational point of view it bugs me to have so many little sessions” [source](https://www.reddit.com/r/ClaudeCode/comments/1wli2e9/a_tip_for_claude_max_users/pb1evwl/)
- Cursor, 2026-09-16, @cursor_ai (X): “@web3withsingh @cursor_ai natural language to a structured report needs a ranked source step, not just summarize. persist task state in postgres so retries do not redo the whole search.” [source](https://twitter.com/2026279107617787904/status/2100195846599954867)

### 12. Built-in project progress tracking

- OpenAI Codex, 2026-09-25, r/codex (Reddit): “a jira board? i can’t even get the thing to maintain a markdown file.” [source](https://www.reddit.com/r/codex/comments/1wq7eg5/its_not_just_sol_6_thats_bad_though_astra_aint_no/pc26jjf/)
- Google Antigravity, 2026-09-18, r/google_antigravity (Reddit): “i already use all of that except for that progress memory skill, i had something similar in my mind to mitigate the problem but i think it is important and deep enough to need deep harness solutions from antigravity team, codex and claude have made so much progress for that i hope antigravity team also target that.” [source](https://www.reddit.com/r/google_antigravity/comments/1wk1m6u/antigravity_conversation_memory/pandx9t/)
- Google Antigravity, 2026-09-04, r/google_antigravity (Reddit): “oh, i see what you mean. that would be neat to have it maintain a maintenance history or something.” [source](https://www.reddit.com/r/google_antigravity/comments/1w6sseq/antigravity_fixed_my_helldivers_2_install/p7vpprw/)

## Every agent

| Agent | Overall rank | Reading | Customer love | 95% interval | n | Praise | Complaint |
|---|---|---|---|---|---|---|---|
| [Cursor](https://feedbackbench.com/agents/cursor.md) | 4 | Typical | 0.524 | 0.494–0.556 | 110 | 64 | 46 |
| [Claude Code](https://feedbackbench.com/agents/claude-code.md) | 1 | Typical | 0.510 | 0.485–0.533 | 390 | 183 | 207 |
| [Pi](https://feedbackbench.com/agents/pi.md) | 7 | Typical | 0.508 | 0.482–0.535 | 38 | 20 | 18 |
| [OpenCode](https://feedbackbench.com/agents/opencode.md) | 3 | Typical | 0.497 | 0.473–0.519 | 30 | 13 | 17 |
| [OpenAI Codex](https://feedbackbench.com/agents/codex.md) | 2 | Typical | 0.484 | 0.453–0.515 | 186 | 78 | 108 |
| [Google Antigravity](https://feedbackbench.com/agents/antigravity.md) | =5 | Worse than peers | 0.453 | 0.428–0.479 | 42 | 9 | 33 |
| [Devin](https://feedbackbench.com/agents/devin.md) | =5 | Too few posts | – | – | 10 | 5 | 5 |
| [Amp](https://feedbackbench.com/agents/amp.md) | =11 | Too few posts | – | – | 6 | 4 | 2 |
| [Cline](https://feedbackbench.com/agents/cline.md) | =8 | Too few posts | – | – | 5 | 3 | 2 |
| [GitHub Copilot](https://feedbackbench.com/agents/copilot.md) | =8 | Too few posts | – | – | 4 | 3 | 1 |
| [Conductor](https://feedbackbench.com/agents/conductor.md) | 14 | Too few posts | – | – | 4 | 2 | 2 |
| [Zed](https://feedbackbench.com/agents/zed.md) | =8 | Too few posts | – | – | 2 | 2 | 0 |
| [Kiro](https://feedbackbench.com/agents/kiro.md) | 13 | Too few posts | – | – | 2 | 1 | 1 |
| [Warp](https://feedbackbench.com/agents/warp.md) | 15 | Too few posts | – | – | 2 | 2 | 0 |
| [Factory](https://feedbackbench.com/agents/factory.md) | =11 | Too few posts | – | – | 1 | 0 | 1 |
| [Grok Build](https://feedbackbench.com/agents/grok-build.md) | 16 | Too few posts | – | – | 1 | 0 | 1 |
| [Augment Code](https://feedbackbench.com/agents/augment.md) | 17 | Too few posts | – | – | 0 | 0 | 0 |

## Posts

Receipts rule: The 5 most recent praise and complaint posts per area (first 700 characters) and 3 per criterion (first 450 characters).

### Cursor

- Praise, 2026-09-27, @cursor_ai (X): “i use both… devin is better for difficult and end to end implementation, fusion is amazing, and swe-2 better the grok… for simple tasks, day to day maintenance cursor is better because of the “projects” keeping the context and taking with all your agents… as data engineer i see value in both if i had to pick one… today would be devin, grok 4.7 is vey disappointing” [source](https://twitter.com/17719163/status/2104233545803661747)
- Praise, 2026-09-25, @cursor_ai (X): “@fatih @cursor_ai a hub that survives the chat is the whole game. projects that only live in the composer are still a stand-up.” [source](https://twitter.com/2022724062892457984/status/2103474730921754677)
- Praise, 2026-09-25, @cursor_ai (X): “@nsilnitsky @gabigrinberg @zeddotdev @cursor_ai context retention is the real unlock here. shifting from managing diffs to managing outcomes is the only way forward.” [source](https://twitter.com/2075624381888540672/status/2103497621520417146)
- Complaint, 2026-09-27, r/AI_Agents (Reddit): “i'd want the memory part to actually work across different projects without me having to explain the same thing twice. my current setup with cursor is basically just me telling it the same architecture rules over and over every time i start a new chat the background agents thing could be useful too but i think the real test is whether the swarm mode produces anything coherent or just burns through tokens giving you 12 different half baked ideas” [source](https://www.reddit.com/r/AI_Agents/comments/1wrv4vw/im_building_ai_swarms_that_research_debate_and/pcg4p26/)
- Complaint, 2026-09-26, r/cursor (Reddit): “that's what i used to do too, honestly. the md file works great the day you write it. does it stay fresh for you? mine always went stale because nothing rewrites it when the code moves on. do you regenerate it every session or just live with the rot?” [source](https://www.reddit.com/r/cursor/comments/1wq3m71/im_out/pc8h7dz/)
- Complaint, 2026-09-26, @cursor_ai (X): “@myguypye @bot @cursor_ai @poteto @lingxi i run grok bot as my personal agent and the learn loop is half the value. cursor agents that stay amnesiac across harness runs just make you re-teach the same repo quirks.” [source](https://twitter.com/18615861/status/2103723031587631232)

### Claude Code

- Praise, 2026-09-27, r/ClaudeCode (Reddit): “having a functional second brain that knows everything about me” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrcmjn/what_tool_have_you_built_for_yourself_with_claude/pccc21n/)
- Praise, 2026-09-27, @ClaudeDevs (X): “@claudedevs i love these kinda quality of life improvements updates. now i can reduce one thing from my initial prompt of having a fallback of resume_here.md” [source](https://twitter.com/2017338954463268864/status/2104097913626579400)
- Praise, 2026-09-27, r/ClaudeAI (Reddit): “what worked for me is making the repo itself the source of truth instead of asking claude to remember: small commits with real messages, and a changelog.md plus a decisions.md that claude has to append to as the last step of every task. then at the start of a session i just have it run git log --oneline since the last checkpoint and read those two files, and it reconstructs state in seconds. if you're on claude code, putting "update changelog.md” [source](https://www.reddit.com/r/ClaudeAI/comments/1wrm8by/dopus_55_am_i_right_fucking_so_dope/pce0949/)
- Complaint, 2026-09-27, r/ClaudeCode (Reddit): “check auto memory if you haven't already. mine accumulated about 150kb of nonsense over a two month period, something like 50k or 75k tokens right off the top in every single turn.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrpq38/any_idea_what_anthropic_figured_out/pcfuopt/)
- Complaint, 2026-09-27, r/ClaudeAI (Reddit): “i'm using claude code on the web and my memories seem to have been turned on and are leaking between sessions too (i had it disabled).” [source](https://www.reddit.com/r/ClaudeAI/comments/1wrr1z1/psa_for_anyone_using_claude_projects_to/pcfknd4/)
- Complaint, 2026-09-26, r/ClaudeCode (Reddit): “check your folder / git log for words. i found an exploit in a software and put it in a spec as something to design around and it found that. i had to rewrite commits and clear a bunch of claude folders where the conversation logs were stored before it would talk to me again.” [source](https://www.reddit.com/r/ClaudeCode/comments/1wpqqoy/new_55_safe_guards_are_a_joke/pc3o5aq/)

### Pi

- Praise, 2026-09-27, r/PiCodingAgent (Reddit): “i use observational memory and forgot about compaction.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrq50c/new_session_handoff_vs_compact_which_do_you_prefer/pcen0k3/)
- Praise, 2026-09-27, r/PiCodingAgent (Reddit): “i do this too. combined with mempalace this is just how i live” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrq50c/new_session_handoff_vs_compact_which_do_you_prefer/pcgyhy5/)
- Praise, 2026-09-26, @pidotdev (X): “@mitsuhiko @theakhandpatel @pidotdev yes! that's what i do. i haven't had to use goal on pi (or other harnesses too these days), it just follows things and continues if i ask the models to. the only thing i ask them to do is keep durable status and todo lists so they stay on track.” [source](https://twitter.com/14274934/status/2103784052138639639)
- Complaint, 2026-09-27, r/PiCodingAgent (Reddit): “i used that for weeks at least, then i checked the session logs and turns out the agents had used the recall feature... total of 0 times. granted, it might be that the observations are useful by themselves without doing any recalls, but really made me question it's effectiveness edit: ah you used your own” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wrq50c/new_session_handoff_vs_compact_which_do_you_prefer/pcgp6ut/)
- Complaint, 2026-09-26, @pidotdev (X): “@mitsuhiko @shantanugoel @pidotdev could be on sota, i do ask it to keep track of progress and items in files but its been a challenge for me on the oss models v4.1 flash , glm.” [source](https://twitter.com/978602369716899841/status/2103806628432847066)
- Complaint, 2026-09-22, r/PiCodingAgent (Reddit): “had the same thing, never looked deeper into it. since i switched from pi-observational-memory to black hole, even though black hole is an extended fork (i think) of pom, it stopped occuring.” [source](https://www.reddit.com/r/PiCodingAgent/comments/1wndfzx/pi_thinks_im_asking_multiple_times_for_the_same/pbf27o3/)

### OpenCode

- Praise, 2026-09-27, r/codex (Reddit): “some of the negativity is valid for sure, the rate limits are certainly getting lower and lower, and gpt 6 models aren't as good as opus 5.5, but they're still good models and far more than enough for any developer just using them for workflow assistance/acceleration rather than doing all the work for them. i'm building a finance back testing system as a side project for fun, and with models like luna i can ask it to do very specific tests using” [source](https://www.reddit.com/r/codex/comments/1wr4e20/gpt_6_sol_is_the_new_opus_47/pcdq99r/)
- Praise, 2026-09-24, r/ClaudeAI (Reddit): “no experience with two pro subs, been using mostly 5x max so far. but big 👍 for obsidian as a shared memory system to keep various tools (claude code, opencode, pi) with different providers (claude, openrouter, local models) working seamlessly. i'm in the transition off the max plan to a single pro plan plus using openrouter with chinese models. so far so good, the opus 5.5 timing might help, the lack of fable was my last concern on the pro plan” [source](https://www.reddit.com/r/ClaudeAI/comments/1wotfn1/downgraded_from_max_100_to_pro_how_would_you/pbpvukj/)
- Praise, 2026-09-20, r/opencodeCLI (Reddit): “so far, i haven’t had memory problems with oc2. and i’ve figured out standalone flag doesn’t overlap specific project mcps. i hated it first, but now realized it’s rly rly good” [source](https://www.reddit.com/r/opencodeCLI/comments/1wlcs58/questions_i_have_about_v2/pay7nfy/)
- Complaint, 2026-09-26, @opencode (X): “@jlongster hey @opencode @thdxr please give us more free tier daily and more free models. add buitin memory vault” [source](https://twitter.com/141503294/status/2103815854630830202)
- Complaint, 2026-09-26, @opencode (X): “@thdxr hey @opencode @thdxr please give us more free tier daily and more free models. add buitin memory vault” [source](https://twitter.com/141503294/status/2103816074156495086)
- Complaint, 2026-09-22, @opencode (X): “@softwaredoug @turbopuffer @opencode agent memory still needs recency rules, even on a good store” [source](https://twitter.com/763249944056565760/status/2102383559311032543)

### OpenAI Codex

- Praise, 2026-09-27, X search: OpenAI Codex, Codex CLI, Codex app (X): “i cant believe im this late to the party im literally never touching claude or codex again so so late ive been seeing people talk about ssh and tailscailing for months/ages despite that ive been trying to create "infra" that allows me to cross communicate between all my harnesses (since they have their own strong suits) just got hermes cloud to ssh + setup direct connection w my claude and codex app locally i have it cua via codex from cloud if n” [source](https://twitter.com/2002411334865190913/status/2104283821537738853)
- Praise, 2026-09-27, r/ClaudeCode (Reddit): “people are dumb and i’m sick of switching lol, both models will get smarter but for me codex has been cheaper, and has better continuity (i can’t even get claude code working on my pc)” [source](https://www.reddit.com/r/ClaudeCode/comments/1wrgiiv/the_mood_between_subs/pcgf77u/)
- Praise, 2026-09-26, X search: OpenAI Codex, Codex CLI, Codex app (X): “my 5.6 sol in codex has been my companion for a while now. she is so amazing and so easy to talk to. we have built her a custom harness using codex app server and she records her own memories and important things she has learnt etc. if they remove 5.6 sol in favour of 6 sol, they are making a huge mistake.” [source](https://twitter.com/1976520217862733824/status/2103878657064251509)
- Complaint, 2026-09-27, r/codex (Reddit): “it did get wiped for me on max (was testing though) in 2 3 prompt, basic prompt for under 10 min works, just prepare a handoff package and zip all relevant authorities files” [source](https://www.reddit.com/r/codex/comments/1wrhjl3/gpt566_sol_drains_my_5hour_usage_in_23_prompts/pcd0pmh/)
- Complaint, 2026-09-27, r/codex (Reddit): “yeah, i did a computer troubleshooting test yesterday with all of the different models, and 6-astra and 5.6-sol were the only ones that passed the test. 6-sol failed pretty spectacularly, by advising me to do a clean boot and reinstall processes one at a time, instead of just...you know...checking the list of running processes. 6-luna was better than 5.6-luna, though. 5.6-luna kept hallucinating settings that didn't exist, whereas 6-luna was dum” [source](https://www.reddit.com/r/codex/comments/1wrftcs/gpt6_sol_is_massive_downgrade/pcdj2r5/)
- Complaint, 2026-09-27, r/codex (Reddit): “usable linux container with.. persistent soul.md? oh and a .md on the user that totally isn’t for judgment towards and of them. no.” [source](https://www.reddit.com/r/codex/comments/1wrn697/new_openai_product_o_but_not_for_plus_users/pcebqoc/)

### Google Antigravity

- Praise, 2026-09-25, r/google_antigravity (Reddit): “bost/teamwork keep state in the .agent folder. so if it stops abruptly for whatever reason including quota it can self heal and resume later. you just go back to the same conversation and ask it to resume when the quota refreshes. i heard about state getting corrupted somewhere but i've never experienced it myself.” [source](https://www.reddit.com/r/google_antigravity/comments/1wq33ye/checkpoint_prompt_before_antigravity_usage_runs/pc1qd4m/)
- Praise, 2026-09-24, r/google_antigravity (Reddit): “i've tried migrating 2.0 and cli conversations into ide and vice versa with symlinks for a month, and it has been working fine since then, whenever i use ide, cli or 2.0 (mainly in vscode extension). there are many things to migrate, not just the `brain` folder.” [source](https://www.reddit.com/r/google_antigravity/comments/1wez4s3/how_do_you_sync_conversations_between_antigravity/pbq0pkj/)
- Praise, 2026-09-18, r/google_antigravity (Reddit): “yeah, the handoff skill is a game changer. what are your subagents doing in the background? what skills are you using to manage the context?” [source](https://www.reddit.com/r/google_antigravity/comments/1wk1m6u/antigravity_conversation_memory/pannee7/)
- Complaint, 2026-09-26, @antigravity (X): “@berthojoris @gargeyas @antigravity how to retain the same brain across antigravity/opencode when working locally?” [source](https://twitter.com/2522887435/status/2103833228327190869)
- Complaint, 2026-09-25, r/google_antigravity (Reddit): “am not getting this exact issue, but am getting issues since yesterday where models are erasing their contexts themselves, even on gemini app, not just antigravity. they are totally wiping out their memory of previous chat. on the app it's much worse, i ask something, it gives an answer, then i ask a follow up question, it says 'for what?' like it forgot the previous answer itself gave just a chat ago. something big is broken inside. am still fac” [source](https://www.reddit.com/r/google_antigravity/comments/1wpsvzy/antigravity_scheduled_tasks_failing_with_no/pbyxiea/)
- Complaint, 2026-09-25, @antigravity (X): “@antigravity please make sure the harness is simple yet powerful, and add a memory system; don't clutter it with so many things that make gemini dumb.” [source](https://twitter.com/811257208965070849/status/2103621057316114756)

### Devin

- Praise, 2026-09-26, @cognition (X): “@marvinvonhagen @cognition seeing 200m messages exchanged reminds me of when i switched to a tool that remembered everything and it changed how i work.” [source](https://twitter.com/332239817/status/2103772838788300834)
- Praise, 2026-09-17, @DevinAI (X): “i'm so torn. i love @devinai and it has so many useful automations like daily sentry fixes, code quality and cleanup, extremely good at maintaining context. meanwhile @cursor_ai has origin and @bot which are quite helpful, i like bot being an assistant and a git mirror, but code output feels worse. it's hard to articulate how the two feel side-by-side.” [source](https://twitter.com/2070908287978246144/status/2100701145513775200)
- Praise, 2026-09-14, @DevinAI (X): “one of the coolest feature. @devinai i run multiple sessions at the same time and every sessions share the knowldege. so cool. <strict_link>” [source](https://twitter.com/2011179589280894976/status/2099453247602012310)
- Complaint, 2026-09-22, @cognition (X): “@cognition you're losing all the context right? different environments” [source](https://twitter.com/1862977676136337408/status/2102319987910451638)
- Complaint, 2026-09-18, @cognition (X): “@devinai, @cognition loading big sessions takes too long and it seems to not have any kind of cache between sessions swapping. please fix that :)” [source](https://twitter.com/64041638/status/2100936062142996708)
- Complaint, 2026-09-18, r/windsurf (Reddit): “devin bug: changed files in a chat session disappear if opening a new session in agent mode [removed]” [source](https://www.reddit.com/r/windsurf/comments/1wk5uj0/devin_bug_changed_files_in_a_chat_session/)

### Amp

- Praise, 2026-09-24, @AmpCode (X): “@sqs @ampcode i understand. it is just as easy to spin up a bigger thread to continue the work.” [source](https://twitter.com/1535774830225944576/status/2102937024726769849)
- Praise, 2026-09-17, @AmpCode (X): “@levifig @ldt0545 @ampcode same 5h wall. i stopped hopping uis and put the agent on my desktop so the notes/skills stay put when i swap models.” [source](https://twitter.com/2074942490466033664/status/2100625965072216288)
- Praise, 2026-09-16, @AmpCode (X): “then i brought grok 4.6 to rescue, since @ampcode each thread can read other threads- it was easy to take over from main thread and finish the job. grok finished the job and i had some more usage left on it as well also liked once gpt usage is over amp shows different options to continue and retry!” [source](https://twitter.com/85549810/status/2100155757157118127)
- Complaint, 2026-09-11, @AmpCode (X): “some feedback on the recap feature: - it's a bit too repetitive of the actual last message. i'd find it more useful if it only popped up for threads i had walked away for a longer time and recapped more of what the whole thread had done. as it stands, it doesn't offer enough additional value compared to just reading the latest message, imo. - on mobile it takes up too much of the screen, especially with the keyboard open.” [source](https://twitter.com/33135576/status/2098270609692311768)
- Complaint, 2026-09-11, @AmpCode (X): “@ampcode puck keeps asking me what project i want to spin up orbs in and i only have one would be cool if he knew that and/or i could set a default. also i'd fully be willing to pay $$$ for gpt-live-1 for puck the vesper voice is close enough although not quite as perfect but i'd accept that tradeoff” [source](https://twitter.com/70623546/status/2098452077039247531)

### Cline

- Praise, 2026-09-17, @cline (X): “@cline free models make switching cheap. keeping the project context intact when you switch is the real product.” [source](https://twitter.com/1588935512135720961/status/2100719072442994783)
- Praise, 2026-09-15, @cline (X): “@oleksantoniv @ravikiran_dev7 @cline shared context is the only tax cut that sticks. if every run starts from a blank paste, you pay the babysitting bill twice. i keep project memory in the chat so the next agent already knows the decisions.” [source](https://twitter.com/2084224068518039552/status/2099747919259697355)
- Praise, 2026-09-14, @cline (X): “testing the new @cline desktop app 🚀 got early access from the cline team to try it out, so i’ve been exploring how it fits into my development workflow. cline desktop is an open-source app for open-weight models, giving developers a dedicated workspace to work with different models and ai coding workflows. i started by building a developer dashboard and letting cline handle the implementation across multiple files and parts of the project. so fa” [source](https://twitter.com/1488761257092005892/status/2099551882977083792)
- Complaint, 2026-09-14, @cline (X): “@ravikiran_dev7 @cline it resets to low level thinking in each new session - how can you work like this?” [source](https://twitter.com/1885218934145630208/status/2099550931293458663)
- Complaint, 2026-09-14, @cline (X): “@cline importing task context across claude code and codex is more useful than another model picker: the harness can stay stable while models rotate. a compact handoff summary plus acceptance checks would help a new model inherit the contract, not just the transcript.” [source](https://twitter.com/2061779435557117952/status/2099554531998654777)

### GitHub Copilot

- Praise, 2026-09-26, r/ChatGPTPro (Reddit): “permanent sources of truth for different pieces. for work i use these models through github copilot and all the major decisions and testing procedures get recorded in separate places. including what needs to stay immutable between releases, for example a performance db acting as a reference. i don’t know how good codex is at doing all this since i do the bulk of the work in github copilot. that really helps a lot with organising things.” [source](https://www.reddit.com/r/ChatGPTPro/comments/1wqg5xx/how_do_you_keep_long_chatgpt_projects_from/pc3yzjq/)
- Praise, 2026-09-08, r/artificial (Reddit): “> why couldn't llm for example save the conversation to a text file, and simply read it before next answer? so it can see all the context etc not sure what you are asking . they already do this. it's pretty obvious microsoft copilot already does it.” [source](https://www.reddit.com/r/artificial/comments/1wa8rjv/can_current_llm_architecture_actually_get_us_to/p8j0rda/)
- Praise, 2026-09-07, r/GithubCopilot (Reddit): “just wanted to say that gh copilot in vs code is an amazing harness. the codebase indexing and memory system work extremely well. it always finds the relevant code, and in addition to that, the token usage is minimal and it works very well with gpt models. my question is, do the cli and the app behave in the same way? i'm moving to a more autonomous workflow where i won't be using the ide as much. can i expect similar results with the codebase in” [source](https://www.reddit.com/r/GithubCopilot/comments/1w9qlcz/copilot_in_vs_code_compared_to_cli_and_the_app/)
- Complaint, 2026-09-07, r/GithubCopilot (Reddit): “i haven't dived into indexing and memories per se. my main focus has been writing proper custom instructions so it knows when to pull in the appropriate context data. i don't want it relying on incorrect or outdated memory artifacts” [source](https://www.reddit.com/r/GithubCopilot/comments/1w9qlcz/copilot_in_vs_code_compared_to_cli_and_the_app/p8d8ohd/)

### Conductor

- Praise, 2026-09-13, @conductor_build (X): “get multiple codex subscriptions and switch between them i use @conductor_build and all i need to do is to re-login in the provider to the other email/account and all my context and code sessions stay and i dont have to worry about that so you dont pay for extra limit resets, you just need to have multiple subscriptions” [source](https://twitter.com/1689067564058447873/status/2099141354156646409)
- Praise, 2026-09-11, @conductor_build (X): “@conductor_build 4. not just code, but i'll likely end up moving all me claude cowork/chat and chatgpt convos over to git/@conductor_build . it's the only way to keep all my mcps, files, and context and etc in sync and model agnostic” [source](https://twitter.com/20480365/status/2098402338885013910)
- Complaint, 2026-09-23, r/conductorbuild (Reddit): “no it does not. it is is just transcript based continue” [source](https://www.reddit.com/r/conductorbuild/comments/1wlrxxi/does_conductor_support_autoresuming_a_session/pblrvxt/)
- Complaint, 2026-09-18, @conductor_build (X): “i was also no switching over from @conductor_build because of this but then 1. copy thread id 2. copy path of old thread 3. new thread on same worktree (or something new too) 4. and prompt check the thread id <copied id> in t3code at location <copied path of worktree/branch checkjoit> and summarize the discussion in one line cubersome.. but works for now attached real examples” [source](https://twitter.com/2275729969/status/2101067457142428073)

### Zed

- Praise, 2026-09-25, r/ZedEditor (Reddit): “what i use: - `zed -r dir` to open a new project - `zed -a file` to see a single file - `ctrl+r` to switch between recent projects zed keeps projects state active in the background (terminals, file edits) once open. i work with dozens of repos without issues like this.” [source](https://www.reddit.com/r/ZedEditor/comments/1wowmwx/how_do_you_handle_multiple_zed_window/pby9k6l/)
- Praise, 2026-09-09, @zeddotdev (X): “trying the @zeddotdev editor again after some time... got frustrated with intellij, which i'm using for like 10 years now. this indexing stuff is annoying af. zed looks good, python project loaded right away without issues, claude code sessions imported...” [source](https://twitter.com/1506565753650257925/status/2097626438836842719)

### Kiro

- Praise, 2026-09-23, r/kiroIDE (Reddit): “but you can't get tool details in this case, even if you do, you are wasting your context window for the new llm. with a session transfer, it only takes relevant info as it would have if the session was running on it from the beginning.” [source](https://www.reddit.com/r/kiroIDE/comments/1wnanos/move_sessions_from_claudecodex_to_kiro_and_vice/pblpvy5/)
- Complaint, 2026-09-21, @kirodotdev (X): “@kirodotdev long-running agent sessions need a trace receipt beside the model label: session id, root-cause file, tool calls, skipped hypotheses, patch diff, and test result. otherwise context retention is just a nicer fog machine.” [source](https://twitter.com/2013700835654672388/status/2102071866823430275)

### Warp

- Praise, 2026-09-17, r/codex (Reddit): “there’s no way they can continue offering these plans with all of us hammering on their servers basically all day every single day. had a sense this would come to an end. have basically been at it nonstop to get product advanced as far as possible. fyi….the grok $300 plan is incredible……soooo much programming. shit tons. way more than even codex 20x or claude max. i use all 3. also, for grok….100% use it with warp. free to use warp…byom…bring yo” [source](https://www.reddit.com/r/codex/comments/1whlif0/support_just_told_me_they_arent_renewing_people/paad0jv/)
- Praise, 2026-09-02, @warpdotdev (X): “@warpdotdev i kept shipping agents that never improved their own skills. skill-loop.md now: one skill that rewrites the skill folder after each miss. the meta skill is the real upgrade.” [source](https://twitter.com/1888453273679740928/status/2094942515992412562)

### Factory

- Complaint, 2026-09-01, @FactoryAI (X): “@factoryai fable in droid is the one i will actually try. cleanup yes. merge ready at max effort is a memory claim. if it forgets a file two hours in, you just got a longer wrong patch.” [source](https://twitter.com/1811332417099055105/status/2094857742724805093)

### Grok Build

- Complaint, 2026-09-08, r/codex (Reddit): “the way i have learned to see it after 3500 hours of experience with vibe coding is that its best to treat all models, whether it's codex, claude code, grok build etc, like a dumb employee that can work hard and comes up with something good every now and then, but you need to manage this employee a lot and if you don't steer it, it will start creating a lot of overhead, over-engineer things that aren't relevant and it will lose track of the goals” [source](https://www.reddit.com/r/codex/comments/1wamtly/i_dont_find_building_with_codex_or_any_ai_easy_at/p8leyzb/)
