@akshay_pachaar: https://x.com/akshay_pachaar/status/2104938749679607902

X AI KOLs Following Tools

Summary

The article explores how splitting tasks across conversation turns reduces agent accuracy and proposes a memory system using context-specific retrieval types to maintain accuracy, with a hands-on example using the open-source Graphiti engine.

https://t.co/iEuE8KOGtg
Original Article
View Cached Full Text

Cached at: 09/29/26, 03:51 PM

Agent Memory Is Not a Search Problem

Across 15 models and more than 200,000 simulated conversations, researchers found that splitting a task across several turns cost 39% accuracy on average. Combining the same content into a single turn recovered most of that loss, even though nothing was added or removed.

So the answer depends on how context is arranged, and in an agent, the memory system makes that decision on every turn before the model runs.

Nearly every memory system makes it the same way. It stores every message, extracts claims from them, runs a similarity search, and hands back the passages that scored highest.

That works when the question is “what did they say about X,” but it breaks down when the question is “what changed and when.”

If a customer said in March that they were on the Standard plan and in May that they had moved to Enterprise, a similarity search may return both statements. Both are relevant, and neither says that one replaced the other.

The gap isn’t in how well the system retrieves, it’s that one kind of passage is being asked to answer every kind of question.

The fix is to retrieve the same memory in several forms, each built for a different kind of question.

Today, we’ll test this with a memory layer for agents built on the open-source Graphiti engine, and go through the six context types a modern agent memory system needs. Using one customer history throughout, we’ll see how context shaped for the question helps an agent use the current facts, keep exact details, and avoid confident answers built on outdated context.

It’s a hands-on article, and you can easily adapt it to your use case.

Bookmark it or drop it in the chat of your agent. It would understand what we are trying to do and adapt it to your own use case.

Let’s get into it.

Memory runs before the model does

A model has no memory. Every API call is independent. Whatever your agent knows about a person on turn twenty is there because something outside the model put it into the prompt first.

That something is the memory system, and its entire job finishes before the model is invoked.

It works with three constraints:

  • A fixed token budget for context

  • An unbounded history to draw from

  • One incoming message to work out what matters

Storage is the easy half. Writing every message to a database is a solved problem and every memory product does it well. The selection step is where systems differ from each other.

When a vendor says their memory is better, they almost always mean their selection is better. Nobody is losing because they failed to save the data.

Four generations of agent memory

Agent memory has been rebuilt four times. Each rebuild fixed the previous failure and introduced a new one.

Send the whole conversation

Paste the full history into every prompt. Simple, and correct as long as the conversation stays short.

It breaks on length and on cost. It also breaks on quality before either of those, because model accuracy drops as input grows, well before the context window is full.

Summarize as you go

Compress the history into a rolling summary. The token count stops growing.

The problem is what a summarizer discards. It keeps the gist and drops the specifics, and the specifics are usually what the agent needed to be correct. Dates, IDs, error codes, exact amounts.

Chunk and embed

Store every message as a vector. At query time, return the k most similar chunks.

This retrieves what resembles the question, which is not the same as what is true right now. A claim and the claim that replaced it look nearly identical to an embedding model. Both come back. Neither is marked as current.

Build a knowledge graph

Instead of storing text and searching over it, extract the things in the data and the relationships between them.

A customer and their pipeline become nodes. “Owns” becomes a connection between the two. A plan change becomes a connection carrying the dates it was true.

This is the shift that matters. The stored unit is no longer text. It is structure. And structure can be read in more than one way.

  • Follow a single connection and you get one precise claim.

  • Gather every connection touching a node and you get a description of that thing.

  • Look at how clusters of connections relate and you get a pattern nobody wrote down.

The first three generations gave you one way to read what you stored. A graph gives you several, from the same underlying data.

That is the property the rest of this article is about.

The labels the field borrowed

Before getting to what those several ways actually are, it is worth clearing something out of the way, because it comes up in almost every memory product’s documentation.

Once systems started extracting structure, they needed vocabulary for the kinds of thing they were storing. The field reached for a taxonomy from cognitive psychology. Three terms, worth defining plainly:

  • Semantic memory is facts you know. Paris is the capital of France.

  • Episodic memory is things that happened to you. You visited Paris last March.

  • Procedural memory is things you know how to do. Riding a bike.

These came out of research into human memory, decades before anyone was building agents.

Here is the problem with them.

This taxonomy sorts memories by what they are about. It tells you the subject matter of a stored item. It tells you nothing about how to get that item back.

In practice, these labels usually do nothing at all. A system lets you tag a memory as one of the three, stores the tag, and then retrieves everything the same way regardless. The tag was never a retrieval path.

Now look at what happens when people measure memory instead of describing it.

The standard benchmark for long-term agent memory evaluates five abilities:

  • Information extraction. Pulling a specific detail out of a long history.

  • Multi-session reasoning. Connecting things said in separate conversations.

  • Knowledge updates. Handling a fact that replaced an earlier fact.

  • Temporal reasoning. Answering questions about when things were true.

  • Abstention. Knowing when the history does not contain the answer.

Not one of those five is semantic, episodic, or procedural.

The people measuring memory organized it by what the system has to be able to do. The people describing memory organized it by what a memory is about.

Only the first produces an engineering decision and one which you can act on.

Six questions, six shapes

Here are some questions a support agent typically gets asked in a normal week.

“What changed, and when?” Needs a small dated claim that can be marked as no longer true without being deleted.

“Tell me about this thing.” Needs a maintained description of one thing, assembled from dozens of separate claims.

“What did the customer actually say?” Needs the original text, unaltered.

“How did that session end?” Needs an outcome belonging to a conversation as a whole, not to any single claim inside it.

“Why does this keep happening?” Needs a pattern spanning sessions that nobody ever wrote down.

“Hi.” First message of a new session. Nothing to search against. Needs something present before there is a query at all.

Compare that against the five benchmark abilities.

  • Knowledge updates and temporal reasoning both need dated claims with boundaries.

  • Multi-session reasoning needs the cross-session pattern.

  • Information extraction needs both the description and the raw text.

Where Zep fits

Zep is a memory layer for AI agents, built on their open-source temporal graph engine, Graphiti.

You send it conversations and business records from wherever your data lives. It builds a temporal knowledge graph out of both, tracking not only what is true but when each thing became true and when it stopped.

On every turn, it assembles the relevant part of that graph into a block of text that goes into your agent’s prompt.

The part that matters is that Zep does not return one shape. It describes six different context types, and they map onto the six questions above.

Retrieving them is one call with a different scope:

One customer history

We actually build out a context graph in Zep to test this.

Everything below comes from that same small synthetic history for a fictional data platform.

Priya Raman is a staff data engineer at Northwind Logistics. In March she sets up a pipeline called orders_ingest, reading shipment orders out of Postgres through a self-hosted connector.

In May she reports two things in one conversation: her team moved to the Enterprise plan on April 28, and orders_ingest failed overnight with a schema-drift error.

Deployment and helpdesk records add the connector version, the cluster, and a ticket status.

In August, Priya opens a new conversation with a single message: “hi”.

Six questions against that history. Six different answers.

Let’s go over each of the context types one by one

1. Facts

A fact in Zep is a claim connecting two entities, stored on a graph edge. It carries four timestamps:

  • valid_at and invalid_at mark when the claim started and stopped being true in the world

  • created_at and expired_at mark when Zep learned about each of those events

Here’s an example to understand this. A user gets married in March. Zep does not process the relevant data until April. The fact’s valid_at is March and its created_at is April.

Now ask the agent whether the user was married in March. It answers correctly, because the graph knows the difference between when something was true and when it found out.

Here is that behavior on our example. A customer moves from the Standard plan to Enterprise on April 28, but only mentions it in a message on May 8:

Two facts come back. The Enterprise claim opens on April 28. The Standard claim closes on April 28. No gap between them and no overlap.

The date of the conversation and the date of the change are different, and the retrieved context keeps both.

Notice what did not happen. Zep did not overwrite the Standard plan when it learned about the upgrade. It closed the interval. Both versions stay in the graph, so the agent can still answer what was true last quarter.

A similarity search can hold both claims perfectly well and nothing in it says that one replaced the other.

2. Entities

An entity is a named thing Zep pulled out of your data. A person, a company, a product, a pipeline. Each one carries a description that Zep regenerates as new facts arrive.

Question: “Tell me about the orders_ingest pipeline.”

Individual claims about the pipeline come back fine. But the history spans months and dozens of separate claims, and the agent would have to reassemble a coherent picture from fragments on every turn.

That description was assembled from messages sent weeks apart. Nobody wrote it. It covers what the pipeline does, how much it handles, and where it reads from.

One point worth noting here is that an entity description is written to read well as a paragraph, and paragraphs do not carry validity intervals. A fact is written to be checkable, and checkable claims do not read like a briefing.

An agent that only had entity summaries would be fluent and wrong. An agent that only had facts would be correct and incoherent, reassembling the customer’s history from scratch on every turn.

3. Episodes

Episodes are the original messages, text, and records you sent to Zep, kept alongside everything derived from them.

They are the source-of-truth layer underneath the rest of the graph.

Question: “What did the customer actually report?”

Priya’s message contains a field name, a type mismatch, an error code, and a byte offset. Someone filing an upstream ticket needs those exactly as typed.

An extracted fact can carry the error code too. What the episode adds is access to the source itself, with the wording around it, so the agent can quote rather than paraphrase.

Reach for episodes when the agent needs to show its evidence rather than summarize it.

4. Thread summaries

A thread summary covers one conversation rather than everything a customer has ever done. Zep builds it incrementally as messages arrive.

Question: “How was the schema drift issue resolved, and what was the outcome?”

That conversation had several steps. Priya reports the failure. Support suggests a cast. Priya applies it and confirms the pipeline is running. Then she says the root cause was never found.

Service recovered. The cause did not get resolved. Those are two different outcomes, and the next support conversation needs both.

Individual facts can describe pieces of the outcome. The summary packages the conversation at a level the agent can act on without reassembling it.

5. Observations

Observations describe patterns, decisions, and relationships across the graph rather than sitting on any single connection. They are evidence-backed and can be retrieved with a query much like the other search scopes.

We covered the clustering mechanism and how they are detected in our observations article…👇

Akshay 🚀@akshay_pachaar·Aug 9 ArticleYour Agent Remembers Everything and Understands Nothing Agent memory is where analytics was for years, returning what you asked for and nothing more. Then analytics started surfacing which number moved and why, without the query. Agent memory hasn’t made…1684516105K

Question: “What do these failures have in common?”

This one needs a longer history, so the example below uses a separate one containing three schema-drift incidents and the scheduled backfill runs that preceded each of them.

The observation brings three incidents and three scheduled jobs into one description. No message in that history states the connection. It exists only in how the records relate to each other.

The observation gives the agent a recurring sequence worth investigating. It does not establish that the backfill caused the failures.

6. User summaries

The user summary is attached to the customer’s central node in the graph and carries the gist of the user according to all the conversations they had in the past.

It is not filtered by the incoming message. It is background rather than an answer to a query.

Priya opens a new conversation with “hi”.

There is nothing in that message to search against. But Zep already holds her history.

That is what makes it work on turn one. The other five types need something to match against. This one is simply present.

For a returning customer, it means the agent starts with their role, their setup, and their open issues, instead of asking them to explain themselves again.

You don’t assemble them yourself

Six types could mean six searches, a decision about which scopes to hit, and a limit to tune on each. It doesn’t.

thread.get_user_context() returns a single context block that Zep builds for you:

What sits behind that call is Smart Context Assembly. It retrieves across all six types, ranks the candidates from the five searchable ones together rather than separately, and fills the block up to a character budget.

Smart Context Assembly lets the shape follow the question. A question about a plan change comes back weighted toward facts. A cold-start “hi” comes back weighted toward the user summary.

Smaller blocks also cost less and leave more of the context window for the rest of your prompt.

You will not need all six types for every agent either. Start with the default block, look at what your agent gets wrong, then map each failure to the shape that would have answered it.

  • An agent answering one-shot lookups may need facts and little else.

  • An agent with no history across sessions gets nothing from observations, because there is no pattern to find yet.

  • An agent whose users arrive mid-task, already identified, gains less from the user summary than one that greets people cold.

Splitting a task across turns cost those 15 models 39% accuracy while the content stayed identical. Arrangement was the only variable, and it decided the answer.

A memory layer makes that same decision on every turn, and it makes it for you.

Storing more history does not change what it hands over. Reranking harder does not either, because the answers to “what changed and when” and “what did they actually say” are not the same kind of object, and one passage of retrieved text cannot be both.

Better retrieval was never the fix. More shapes was.

Resources:

  • Zep Graphiti GitHub repo →

  • Zep context types documentation →

That’s a wrap.

Thanks for reading!

Similar Articles

@akshay_pachaar: 6 context types for agent memory. (retrieving facts alone is not memory) An agent may need the current truth, exact ori…

X AI KOLs Following

A breakdown of six context types for agent memory — facts, entities, episodes, thread summaries, observations, and user summaries — arguing that a good memory system stores history once but reads it in different forms. The post highlights Zep and its open-source temporal knowledge graph framework, Graphiti, which assembles context types into a fixed-budget block for agents.