Spots

The Problem With Making an Agent Remember Everything

Imagine you are a sales representative preparing for a call with the CFO of a company you have been pursuing for six months. You know you have spoken to them before. You remember there was a pricing objection. You remember another stakeholder liked the product. You vaguely remember a competitor being mentioned. But before the call, you still have to open the CRM, search through old notes, read meeting summaries, and look through emails to reconstruct what actually happened. The frustrating part is that the information is not missing. It is just scattered. A long sales cycle can produce hundreds of individual interactions. A customer can object to pricing in January, bring up integration in March, add a new decision-maker in May, and change their budget situation in June. By the next call, the salesperson is not simply looking for information.

They are trying to reconstruct how the deal

They are trying to reconstruct how the deal changed. That is the problem I wanted to solve with my Deal Intelligence Agent. I wanted a salesperson to be able to ask something as simple as: "Brief me for my next call with the CFO." and get an answer based on the history of that specific deal. My first instinct was to solve this by giving the model more context. That turned out to be the wrong abstraction. The obvious solution: give the model the whole history If the model needs to know what happened six months ago, why not just give it six months of history? I could take CRM notes, meeting summaries, emails, and previous conversations and put them into the prompt. The LLM would then have everything it needed to answer the question. At first, this sounds reasonable. But a longer context creates another problem: not everything in the history is relevant to the current question.

Consider a deal with this timeline: January CFO

Consider a deal with this timeline: January CFO: "Your pricing is too high." March CTO: "The product looks good, but integration is a concern." May Customer receives additional funding. June CFO: "Budget is no longer the main issue."

July IT team: "We need more information about

July IT team: "We need more information about integration." All five events are useful historical facts. But if the salesperson is preparing for the July call, they do not need five equally weighted facts. They need to understand that the original pricing objection has changed, the CTO is already positive, and integration is now the active concern. Giving the model more context does not automatically give it better context. I was treating the problem as a context-window problem when it was really a memory problem. I changed the architecture instead of making the prompt bigger Instead of continuously feeding the LLM more historical information, I moved the long-term state outside the model. I use Hindsight as the long-term memory layer.

The basic flow is: Customer interaction | v

The basic flow is: Customer interaction | v retain() | v Persistent memory | | later v recall() | v Relevant deal history | +------ Current query | v GPT-OSS-120B | v Deal-specific answer The LLM remains responsible for reasoning about the current question. Hindsight provides the persistent memory that allows information from previous interactions to become available later. That separation is important. I do not need the model itself to permanently remember every conversation. I need the application to preserve the experience and retrieve the relevant parts when a future question requires them. The memory layer therefore becomes the bridge between two otherwise separate interactions. The deal becomes the memory boundary A sales conversation only makes sense in context. The same sentence can mean very different things depending on the customer, stakeholder, and stage of the deal.

That is why the application is organized around

That is why the application is organized around individual deals. A salesperson selects the deal they are working on and asks questions about it. When a new interaction or call outcome is retained, it is associated with that deal. Conceptually, the operation looks like this: client.retain( content=call_outcome, context=f"deal:{deal_id}" ) The important part is not the number of lines of code. It is the fact that the interaction becomes part of persistent deal-specific memory instead of remaining only in the current conversation. Later, a query such as: "Brief me for the next call with the CFO" can trigger a recall operation for that deal. The model then receives the current question together with the relevant memories. The same question with and without memory I wanted to make the difference between the two approaches visible, so the application includes a Use Memory toggle.

With memory disabled, the query is sent directly

With memory disabled, the query is sent directly to the LLM. With memory enabled, the application first recalls relevant memories and then gives those memories to the LLM along with the query. The question stays the same. For example: "Brief me for my next call with the CFO." Without memory, the model can only provide general sales advice: Prepare to discuss pricing, ROI, implementation costs, and likely objections. That answer is not wrong. It is simply generic. Now suppose the memory layer retrieves: CFO rejected the previous 15% discount. CTO supports the product. Competitor X is offering a lower upfront price. CFO requested a three-year cost comparison.

News

The Problem With Making an Agent Remember Everything

Imagine you are a sales representative preparing for a call with the CFO of a company you have been pursuing for six months.

@spots #dev
Source: Dev.to
See more like this