A System That Remembers Is a System That Improves: Building AI That Learns From Feedback
On 17 July, an autonomous agent I run declined a trade that had passed every one of its technical tests.
The stock was Best Buy. The breakout was real and still intact when the market opened, the price feed was clean, the liquidity was comfortably inside its limits. On the numbers in front of it, the trade qualified. It stood down anyway, and the reason it gave had nothing to do with the chart. Its own record showed that the last two closed trades in that sector, on 2 June and 15 May, had both lost money. So it left it alone.
Nobody intervened that morning. Nobody rewrote a rule. The behaviour changed because the system could see what had already happened to it, and was allowed to act on what it saw.
That is the argument of this post. Automation that cannot remember will run the same play, at the same quality, forever. A system with a memory underneath it becomes different over time, and the interesting part is that the model did not get any cleverer. The record got better.
If the term is new, an intelligence system is an owned, maintained store of what your business knows, structured so people and AI can ask it questions and get current, cited answers. This post is about what that memory does once it starts accumulating outcomes rather than just facts.
Can AI systems actually learn from experience?
The model does not. The system around it does, and confusing the two is why so many AI projects plateau after the first fortnight.
A language model is frozen between releases. Whatever you told it this morning is gone by tomorrow unless something outside the model wrote it down and hands it back at the right moment. There is no accumulation happening inside the thing itself. What you thought was the AI getting to know your business was just a long conversation, and the conversation ended.
So when two businesses run the same model on the same task for a year and one of them is visibly better at it, the difference is almost never the model. One of them kept the record. Their system reads a growing store of what worked, what was rejected, what the client actually meant when they said the tone was off. The other starts from zero every morning and calls it consistency.
This is also why the plateau is so common. In the UK, the government's own AI adoption research found agentic AI was the least adopted technology (7%) among businesses already using AI, and that for those current users the most common thing that had previously held adoption back was having limited AI skills (60%). Most organisations are still at the stage of using a tool, not building a system that gets better at using one.
What a feedback loop looks like in practice
A real loop has four parts, and most automation stops after the second.
The system does something. A person judges it. That judgement gets written into the memory as a dated record. The next run reads it before it acts.
Take the content engine we built for Excellerate Services. The system drafts a post in the voice of the right region and emails it to the regional approver, who replies in Outlook the way they would reply to a colleague: a thumbs-up, some revision notes, or a rejection. The engine reads the reply, works out what was actually being asked for, and either revises or proceeds. Weekly production went from roughly 12 hours to about 2 hours of oversight, and the reason it holds is that the correction arrives in the approver's own words rather than through a form nobody fills in.
The part that makes it a loop rather than an approval queue is what happens to those words afterwards. A correction that stays in an inbox teaches nothing. The same note comes back the following week, the approver types it again, and everybody quietly accepts that this is just how it is. A correction that lands in the memory changes the next draft.
That is the whole difference, and it is unglamorous. Most of what people call AI learning is really just somebody remembering to write things down in a place the machine can read.
Is this the same as fine-tuning?
No, and for a mid-market business the distinction is a commercial one, not a technical one.
There are three ways to make an AI system behave differently, and they are not close substitutes.
| Dimension | Fine-tuning | Prompt engineering | Memory and retrieval |
|---|---|---|---|
| What changes | The model's weights | The instructions you send | What the system reads before it answers |
| How fast a lesson lands | A training cycle | Immediately, if someone edits the prompt | Minutes, as soon as it is written down |
| Can you see why it decided | Not really | Only what is in the prompt | Yes, the source comes back with the answer |
| Undoing a bad lesson | Retrain | Rewrite and hope nothing else broke | Edit one page |
| Scales with your knowledge | No, it is a snapshot | No, prompts have a ceiling | Yes, that is the point |
Fine-tuning has real uses, mostly around format and specialised language at volume. It is a poor way to teach a system facts about your business, because your business keeps changing and the model does not. You would be baking July's positioning into something you cannot easily unbake in October.
Prompt engineering is where nearly everyone starts and where nearly everyone gets stuck. It works until the instructions grow past what anyone wants to maintain, at which point you have a two-thousand-word prompt that three people have edited and nobody fully understands.
Memory is the one that compounds. The system reads a store that grows as the business learns, every answer comes back with the document behind it, and a wrong lesson is a page you edit rather than a model you retrain.
How memory changes agent behaviour
Back to the agent that turned down Best Buy, and I should be straight about what it is. It trades on paper, not real money, and it is currently behind a simple index fund. That is not the interesting bit and never was. What is interesting is watching a memory change a decision in a way you can inspect afterwards.
The rule that fired was a piece of accumulated caution: if the last two closed trades in a sector both lost, stand down in that sector. It is not a clever rule. What makes it work is that the agent holds a dated, queryable record of every trade it has ever closed, so the rule has something true to read. The same morning, it turned down a technology stock on the same basis, from two different losses months apart.
A person would call that learning your lesson. Inside the system it is more mundane: the memory returned two rows, and the rule did the rest.
The important property is that the decision is inspectable. Every refusal is logged with the reason and the evidence behind it, which means we can go back later and ask whether the caution was right. Sometimes it was not. The agent's own regret analysis currently shows one of its gates has blocked more winners than losers, at a measurable cost in performance. That gate is under review because the record says so, not because anyone had a hunch.
This is what people usually mean by proactive agents on the maturity ladder, and it is the rung that gets skipped. An agent that acts without a memory is just a faster way to repeat yourself.
Where feedback loops go wrong
They fail in four ways, and three of them look like success while they are happening.
Learning the wrong lesson from too little. Two losses in a sector is not statistical evidence of anything. My agent's rule is a discipline device, an honest attempt to stop a bad run compounding, and it is entirely capable of being wrong. The mistake is to dress up a small sample as insight, then let a system act on it a thousand times. Small samples are fine as long as everyone knows that is what they are.
Recording what happened without recording why. A memory of outcomes with no reasoning attached produces rules that nobody can later justify or safely change. Months later, someone finds a rule blocking useful work and has no idea whether removing it is brave or reckless. The reasoning is the expensive part, and it is the part people skip because writing it down feels like admin.
Counting mistakes but never counting caution. A system that is punished for errors and never measured on the opportunities it declined will get quieter and quieter until it is technically flawless and commercially useless. Whatever you build, measure what the refusals cost, not just what the errors cost.
No human who can overrule it. Every loop needs somebody who can look at a rule and say, that made sense in May and does not now. This is the ordinary case for keeping a person in the loop, and it is why the NIST AI Risk Management Framework is written to build trustworthiness into "the design, development, use, and evaluation of AI products, services, and systems". Evaluation sits in that list alongside building the thing, not after it.
None of this is hypothetical. Gartner's prediction, reported in June 2025, is that more than 40% of agentic AI projects will be canceled by the end of 2027, based on a poll of more than 3,400 organisations investing in the technology, with the analyst behind it noting that most such projects are "early-stage experiments or proof of concepts that are mostly driven by hype and are often misapplied". The failure is rarely the agent. It is deploying one with nothing durable underneath it and no way to tell afterwards whether it did the right thing.
What this means for the work
The practical version is smaller than it sounds. You do not need an agent that trades, and you almost certainly do not need fine-tuning. You need the corrections your team already makes to land somewhere the system reads.
That means capturing judgement at the moment it happens, in the words the person actually used, and filing it against the thing it was about. It means every one of those records carrying a date, so a rule from March announces itself as a rule from March. And it means one person owning the question of whether the accumulated lessons are still true, which is a monthly half-hour, not a role.
Do that and the compounding starts on its own. The four layers of a working intelligence system are how you organise it; the feedback loop is what makes the thing improve rather than merely persist.
Practical takeaways
- Assume the model remembers nothing. Every improvement you want has to live outside it, in a record the system reads before it acts. This one reframe fixes most disappointed AI expectations.
- Capture corrections in the correcting person's own words. Free text in the channel they already use beats a structured form they will not complete. The engine can work out the intent.
- Write the reasoning, not just the outcome. "Rejected, too formal for this region" is a lesson. "Rejected" is a statistic.
- Date everything and give lessons an expiry review. A rule learned in a different market is the most confident kind of wrong.
- Measure what your caution costs. Track the opportunities the system declined alongside the mistakes it avoided, or you will optimise your way into paralysis.
- Keep the veto human. The human ceiling on this work is real, and the point where a system's memory has to be overruled is exactly where judgement earns its keep.
A system that remembers is a system that improves, and one that does not will be exactly as good in a year as it is today. If you would rather build that memory with a partner than from scratch, that is the work we do.
Frequently asked questions
Can AI systems actually learn from experience?
The model itself does not. A language model is frozen between releases, so nothing you tell it on Tuesday is present on Wednesday unless something outside the model wrote it down. What can learn is the system around the model: the memory it reads before it acts, and the record of outcomes that memory accumulates. That is why two businesses using the same model get very different results over a year. One of them is keeping the record.
What is a feedback loop in practice?
Four things, in order. The system does something, a person judges it in their own words, that judgement is written into the memory as a dated record rather than left in an inbox, and the next run reads it before acting. Most automation has the first two and stops. The correction lands in a reply, the person feels heard, and the identical mistake arrives next week. A loop that does not write anything down is not a loop, it is a complaints procedure.
How does memory change agent behaviour?
By changing what the agent knows when it decides, not by changing the model. An agent I run declined a technically valid trade because its own record showed the last two closed trades in that sector had both lost. Same model, same market data, different decision, entirely because the record was there to read. Memory does not make an agent cleverer. It makes it accountable to its own history, which in practice looks a great deal like judgement.
Is this the same as fine-tuning?
No, and the difference matters commercially. Fine-tuning changes the model's weights: it is slow, expensive, hard to audit, and you cannot easily explain or reverse an individual lesson. Memory and retrieval leave the model alone and change what it reads before it answers. A lesson lands in minutes, you can see exactly which document caused a decision, and you undo it by editing one page. For almost every mid-market business, memory is the right tool and fine-tuning is not.
What are the risks of AI that learns from feedback?
Learning the wrong lesson, confidently, at scale. Two bad outcomes are not statistical evidence, so a rule drawn from them may be caution rather than insight. Feedback loops also tend to record what happened without recording why, which produces rules nobody can later justify. And a system that only counts the cost of its mistakes, never the cost of its refusals, will get quieter and quieter until it does nothing useful. The controls are dated records, visible reasoning, a person who can overrule, and measuring what your caution costs you.
Related Articles
Building an Internal AI Team vs Hiring an Automation Agency: Real Trade-offs
A candid comparison of building an internal AI team versus hiring an automation agency for UK SME and mid-market businesses: real costs, speed to value, hidden risks, and the hybrid model most firms actually need.
From Automation to Intelligence: The AI Maturity Ladder
A seven-rung AI maturity ladder for UK SME and mid-market businesses, from generative AI to an agentic organisation, and why rung five, the intelligence system, is the one most firms skip.
What Is an Intelligence System? Why Growing Businesses Need a Memory
What an intelligence system is, how it differs from a wiki or BI dashboard, and why UK SME and mid-market firms need one before they scale AI agents.