Gbenro Dara
Writing
AI
8 min read

Building Smarter Agents: How Persistent Memory Solves Workflows

Agents have to remember to help users navigate complex workflows successfully

1 October 2026 · Updated 2 October 2026

Building Smarter Agents: How Persistent Memory Solves Workflows

The Case File Is the Product

Since I started building and fiddling with agents, my early observation was, They can draft. They can summarize. They can sound confident for five minutes. Then you come back a week later and have to explain the case again.

What happened? What did we try? Why did it fail? What are we waiting for?

I think this is where a lot of the conversation about agents goes wrong. People talk about autonomy as if the main question is how many steps an agent can take without supervision. That matters. But in the messy workflows I know best, I would first ask whether it can reliably carry the case from one interaction to the next.

The useful agent knows what has already happened and uses that history to move the work forward.

That sounds basic but in financial services and enterprise use cases like insurance, claims, appeals, and benefits navigation, it is a serious product requirement.

The problem is not reasoning alone

I have spent a lot of my career in markets where the official system is fragmented, slow, and hard to trust. Insurance claims are like this. So are many of the processes people go through to access benefits or resolve medical bills.

These are decision problems. But making the right decision often starts with reconstructing what happened.

Take a denied medical claim. The denial letter arrived on September 12. The insurer cited a missing pre-authorization. The hospital submitted the wrong code the first time. The patient has already called twice. A supervisor promised a callback within 72 hours. The plan document sets out the appeal process.

An agent that ignores that history may produce a perfectly written letter that does nothing useful. The next step might be to get the corrected submission, establish whether authorization was actually required, or escalate an unanswered request.

Remembering the chronology does not guarantee a successful appeal. The coverage rules still matter. So do the evidence and the quality of the reasoning. But losing the chronology makes it harder to know what to do next.

A larger context window does not automatically solve this. In Lost in the Middle, Nelson Liu and his co-authors found that the models they tested used long inputs unevenly: performance depended on where the relevant information appeared. The study does not tell us how every current model performs. It does establish an important distinction between having information available and using it reliably.

That is the distinction I care about. A pile of documents is useful only if the system can find the right evidence and understand what it means for the case now.

The case file is the product

One opportunity I keep coming back to is helping under-resourced US patients fight claim denials and resolve medical bills. It is repetitive work, full of handoffs, with real consequences when something gets missed.

KFF's analysis of 2024 ACA Marketplace claims found that consumers rarely appealed denied claims, and insurers usually upheld the denials that were appealed. That does not tell us how many outcomes better case management would change. It does show why we should take the difficulty of navigating these processes seriously.

The same product question comes up in hospital financial assistance and benefits enrollment. Can the system maintain a working record of the person's situation and keep it moving?

The institution may have records across several departments. The user has emails, PDFs, screenshots, call notes, and whatever they can remember. Neither necessarily has a complete picture. But the user is often the one expected to connect the pieces.

So the job is simple to state and hard to do: create a persistent case file that becomes more useful with every interaction.

It should capture the parties involved, the chronology, the relevant rules, missing documents, prior commitments, failed paths, deadlines, and who is responsible for the next step. It should link back to the evidence behind those facts.

It should survive a model change, a new communication channel, and a human handoff. A person stepping in should be able to understand where things stand without reading every chat.

That is what I mean when I say the case file is the product.

Carry the burden of memory

I learned in insurance that users care most about whether the system works when something goes wrong. That is why I wrote that speed is the product. The claim is the moment of truth.

Continuity faces a similar test. If a family has to restate its situation every time it calls, the service feels unreliable. If a document has already been submitted and the system keeps asking for it, the user has more work to do. If nobody remembers a promised callback, the case can sit there while everyone assumes someone else is handling it.

There is a useful body of research on this burden. Donald Moynihan, Pamela Herd, and Hope Harvey describe administrative burden through learning costs, compliance costs, and psychological costs in people's interactions with government. Understanding the process, supplying the paperwork, and dealing with the experience all impose costs.

I see a related pattern in insurance. A person dealing with a medical bill is already carrying enough. Requiring them to rebuild the case at every handoff adds to that burden.

This is why continuity matters to me. The agent should carry more of the work of remembering, checking, and following up. The user should be able to return to the case and find that it has a clear history and a clear next step.

Memory has to be shaped around the job

External memory for AI agents is not a new idea. MemGPT, for example, explored how a system could manage different memory tiers beyond a model's immediate context window.

The question for a founder is what the system needs to remember to do a particular job well.

A claims appeal agent needs denial reasons, correspondence, coverage terms, documents, deadlines, and escalation history. A benefits navigator needs household circumstances, eligibility requirements, submitted forms, missing proofs, and renewal dates.

Remembering that someone prefers concise answers is useful. Remembering that their household circumstances changed, and that this affects an unfinished application, is a different level of responsibility.

LongMemEval, developed by Di Wu and his co-authors, tests capabilities including reasoning across sessions, tracking time, updating knowledge, and recognizing when the available information is insufficient. Those are much closer to what a working case file requires than simply recalling a fact from an old conversation.

The system also has to remember correctly. A patient may initially give the wrong date. A hospital may correct a code. A new document may contradict an earlier account.

The file needs to preserve the source, record the correction, and distinguish confirmed facts from user reports and agent inference. It needs to know when a task is complete and when a rule or deadline has changed. Users should be able to inspect, correct, export, and delete their records.

Otherwise, we risk building a system that remembers its mistakes very well.

Useful memory is not automatically a moat

I still think there is a business opportunity here. But storing case history alone does not make a company defensible. Claims systems and case-management software already do that.

The harder question is whether an agent can use that history to deliver better work. Can it identify the real blocker? Can it prepare the right evidence? Can it make a human handoff faster? Can it follow through without repeating a failed approach?

There may be a compounding advantage if the system learns from verified outcomes across cases: which documents tend to be missing, which process failures recur, and which actions help under particular circumstances. But that learning has to be demonstrated, with appropriate permission to use the data. More stored information does not automatically mean better decisions.

Distribution, integrations, domain knowledge, and trust will still matter. So will model capability.

My bet is that reliable continuity will become an important part of the advantage in these workflows. It gives the model a better foundation to work from and gives the user a reason to keep relying on the product.

The test I would build to

I have two tests for this kind of agent.

First, if I change the model in six months, does the case survive? The history, evidence, decisions, and responsibilities should remain usable. That tells me whether I have built something durable around the model.

Second, does it improve the work compared with a transcript-based assistant or ordinary case-management software? Are users repeating themselves less? Are fewer deadlines missed? Are there fewer incorrect follow-ups? Can a human understand the case faster?

The first tests the architecture. The second tests the product. Neither, on its own, proves a moat.

I made a related argument in Designing an Operating System for Human-Agent Teams: memory, responsibility, and action need to work together. An agent that knows what happened but never moves the case forward has only done part of the job.

That is where I would focus. Give the agent a reliable case to reason from. Make it clear what needs to happen next. Measure whether it helps the user get there.

In the workflows that matter to me, a useful agent has to remember what the user cannot afford to have forgotten.

That is the standard I would build to.

memory
workflow
ai
claims

Occasional essays

On AI, insurance, building in Africa, and what I'm learning. No cadence promises. Only when I have something worth saying.