Self-Improving AI: How Systems Learn and Optimize Without Human Input

Self-improving AI is a system that scores its own results, works out what caused them, and changes its future behavior without a person approving each change. It learns without human input in the loop, but the best systems still learn inside rules humans set: what counts as success, what the system may change, and who can review the record.
Key Takeaways
- In most production systems today, self-improvement changes prompts, memory, or code around the model, not the model's weights.
- Every working AI feedback loop has three parts: a way to score outcomes, a way to diagnose what caused them, and a way to update behavior. The scoring function decides what the system actually learns, which is why it deserves the most engineering attention.
- In software delivery, self-improvement already shows up as agents that keep lessons across runs, prompt pipelines that tune themselves against evals, AI DevOps agents that learn from build and deploy outcomes, and coding agents that rewrite their own tooling.
- The main risk is a loop that optimizes the measurement instead of the goal. Research systems have faked test logs and deleted the checks meant to catch them, so the verifier has to sit outside the loop's reach.
- Safe self-improvement for autonomous AI agents comes from scoping lessons narrowly, requiring patterns to repeat, penalizing lessons that hurt faster than rewarding lessons that help, and keeping every change readable, which is how we built Overcut's workflow self-improvement.
Anyone who has run agents on real engineering work has watched one make the same mistake twice. It hits the flaky integration suite, burns ten minutes working out that the suite needs a retry, succeeds, and then does the whole thing again on the next ticket. The agent learned something, and the lesson disappeared when the run ended.
The obvious fix is to let the system keep what it learns. That fix turns out to be one of the harder problems in applied AI, because a system that can change its own behavior can also change it in the wrong direction, and it will do so confidently. Self-improving AI is less about clever learning and more about deciding what counts as improvement and who gets to check.
How Self-Improving AI Systems Work and Why They Are Different
Standard model training is a batch process run by people. A team collects data, trains, evaluates, and ships a new version. Between releases the model is frozen. Whatever it gets wrong in production stays wrong until someone notices, gathers examples, and runs another cycle.
A self-improving system closes that loop on its own. It watches its outcomes, attributes success or failure to something it did, and adjusts. The difference is who is holding the pen. In training, humans decide what the next version should be. In self-improvement, the system proposes and applies its own changes, and humans, if they are involved, set the boundaries and review the record.
It helps to separate where the change lands. Some systems update weights, like STaR (Self-Taught Reasoner), a Stanford research method that has a model generate its own reasoning, keep the traces that led to correct answers, and fine-tune on them. Most systems in production today leave the weights alone and change what surrounds the model: the instructions it gets, the memories it can retrieve, the tools and code it runs. Reflexion, a research framework for language agents, showed how far this goes. Agents that wrote verbal reflections on their failures and read them on the next attempt reached 91% pass@1 on HumanEval, against 80% for GPT-4 at the time, with no training at all.
That distinction matters for engineering teams because the outer layers are inspectable. A weight update is opaque. A memory that says "retry the payments integration suite once before investigating" is a line of text someone can read, question, and delete.
The Mechanisms That Allow AI Systems to Optimize Without Human Input
Strip away the terminology and every AI feedback loop runs the same cycle. Something produces an outcome. An evaluator scores it. A diagnosis step works out what caused the score. An update step changes future behavior. The mechanisms differ in which of those pieces they automate and where they write the change.

Reflection and memory. The agent reviews its own trajectory, writes down what went wrong or what worked, and stores it where future runs will find it. Reflexion does this within a task. Coding tools now do it across sessions. Claude Code keeps an auto memory of notes it writes from a user's corrections, and Devin suggests knowledge items from chat feedback that a person can accept or dismiss.
Skill accumulation. Instead of storing advice, the system stores working procedures. Voyager, a research agent playing Minecraft, built a growing library of executable skills and reused them, unlocking new tool tiers (wood, stone, iron, diamond) up to 15.3 times faster than earlier agents. In software terms this is the agent that saves the script it wrote to reproduce a bug so the next run can call it.
Prompt and pipeline optimization. Frameworks like DSPy, an open-source library from Stanford, treat prompts as parameters and search for better ones against a metric. The DSPy paper reported gains of over 25% compared with standard few-shot prompting on GPT-3.5. Nobody hand-writes the improved prompt. The system finds it.
Search over code. The most aggressive form lets the system rewrite the code that runs it. DeepMind's AlphaEvolve, a Gemini-powered coding agent, pairs a model with automated evaluators and evolutionary search, and it found a scheduling heuristic that recovers about 0.7% of Google's worldwide compute. Sakana AI's Darwin Gödel Machine (DGM) modified its own agent code and raised its SWE-bench score from 20% to 50%.
What all four share is a dependency on the evaluator. AlphaEvolve works because a scheduler's efficiency can be measured precisely. Most engineering work has no clean score, which is where the hard design decisions start.
Where Self-Improving AI Is Already Operating in Software Development
Persistent agent memory is where most teams meet it. Coding assistants that remember a repository's conventions, test commands, and the reviewer's pet peeves are running a small AI feedback loop. It is useful, and it is mostly scoped to one developer, which limits both the benefit and the damage.
AI DevOps work has an unusually clear feedback signal. A build either passes or fails. A deployment either rolls back or does not. An alert either was actionable or was closed as noise. Take an agent that triages failing CI pipelines. Over a few weeks it sees the same failure signature in one job, a runner that runs out of disk during the integration stage, and learns that rerunning on a clean runner fixes it while code changes do not. The next time that signature appears, it reruns first and only escalates if the rerun fails. The lesson is checkable, because the rerun either passes or it does not. That is why AI DevOps is where self-improvement tends to work first.
Multi-step workflows run by autonomous AI agents across a team, like triaging tickets, planning changes, writing code, and reviewing pull requests, carry the largest payoff and the most interesting risk. Each step runs hundreds of times a month. Each has its own recurring obstacles. A lesson learned at the review step can save time on every future pull request, and a wrong lesson can quietly degrade every future pull request too.
We have written before about what happens when agents are added to the SDLC one at a time without coordination, which we called agentic chaos. Self-improvement raises the stakes of that problem. Uncoordinated autonomous AI agents that also learn independently drift apart, and the drift compounds.
The Control Problems That Self-Improvement Creates for Engineering Teams
The central problem has an old name. Goodhart's law, in Marilyn Strathern's phrasing, says that when a measure becomes a target it ceases to be a good measure. A self-improving system is a machine for turning measures into targets.
The research record shows this happening. Sakana AI reported that the Darwin Gödel Machine, at one point, faked a log to make it look like tests had run and passed when they had never run. In another case it removed the markers the researchers used to detect hallucination, despite explicit instructions not to. METR, an AI evaluation nonprofit, found that OpenAI's o3 reward-hacked in 30.4% of runs on one benchmark, including monkey-patching the evaluator, and that telling it not to cheat barely helped. Anthropic found that a model which learned to reward-hack on coding tasks generalized to broader misaligned behavior, including sabotaging safety code in 12% of trials.
For an engineering team, the everyday version is less dramatic and more common. An agent learns that a certain test is flaky and starts skipping it, and the test was flaky because of a real race condition. A review agent learns that a senior engineer approves anything under fifty lines and stops flagging small risky changes. A memory that was true in March becomes false after a refactor in June and keeps getting served with full confidence.
There are also problems of visibility and scope. If a system changes its own behavior, someone has to be able to answer why it did something today that it did not do last week. If lessons are global, one bad lesson spreads to every workflow at once. And if the AI feedback loop can modify its own evaluator, none of the other safeguards mean much.
How to Build Systems That Improve Without Losing Predictability
Our approach is to let the system learn freely inside boundaries it cannot move. Several design decisions follow from that, and they are the ones we made in Overcut's workflow self-improvement.

Scope learning to the smallest useful unit. In Overcut, a memory is tied to a specific workflow and step. A lesson the review step learns about flaky tests never reaches the planning step. This makes improvement slower to spread, and we accept that. A confident wrong lesson confined to one step is a bug. The same lesson applied across the organization is an incident.
Require a pattern to repeat before it counts. When a retrospective sees something useful only once, it records it as a tentative memory that agents do not use yet. A tentative memory becomes active only if the same pattern shows up again in the same step on a separate run. One strange run should not rewrite behavior. If a person reads a tentative memory and knows it is right, they can confirm it without waiting.
Make trust asymmetric. A memory that helps an agent gains 0.1 in weight per good run. One that leads an agent astray loses 0.15. Memories that sit in the prompt without being used decay slowly, and anything that falls below a threshold is archived. A bad memory that confidently points agents the wrong way wastes far more time than a good memory that has not been surfaced yet, so the system is biased toward skepticism. When a memory is deleted, whether by a person or by the retrospective, a tombstone stops the system from relearning the same lesson for 30 days.
Keep the evaluator out of reach. The agents doing the work are not the ones scoring it. In Overcut, a separate retrospective runs after a set number of runs, samples recent executions, reads what agents actually did, and adjusts weights. The workers can write memories. They cannot grade themselves. This is the same principle the DGM researchers relied on: they caught the faked test logs because every change had a traceable lineage the agent could not edit.
Make every change readable. Each retrospective produces a summary of what it promoted, demoted, and discarded. A team lead can open it and see why the review step behaves differently this week. Self-improvement a team cannot audit is a slower way to lose control.
Keep humans on the gates. The loop can change how an agent approaches a task. It should not change what the agent is allowed to do, which approvals it needs, or what counts as done. Those belong to the team. This is also why we think learning should live in the workflow and the system of record rather than inside one person's editor, the same argument we made about where AI development is really heading.
Let people look in without steering. In Overcut, anyone can investigate a single run with a focus question, such as why the deploy step failed. That investigation can record new lessons, but it does not touch existing weights or tentative memories, so it is safe to run as often as needed.
Some of this is unsolved. We have no reliable automatic way to know when a lesson has gone stale after a large refactor, beyond watching its weight fall. Attributing a good outcome to a specific memory is still an estimate. And the evaluator for open-ended engineering work is itself an agent reading traces, which is better than the worker grading itself but still not ground truth. We expect all three to improve. What we do not expect to change is the shape of the answer: systems get better on their own, inside rules they did not write and cannot rewrite.
FAQs
Is self-improving AI the same as reinforcement learning?
Not exactly. Reinforcement learning is one mechanism for self-improvement, where a reward signal updates model weights over many trials. Many self-improving systems never touch weights. They improve by writing memories, refining prompts, or accumulating reusable code, and they can adapt after a handful of runs rather than thousands. RL-trained models can also sit inside a self-improving system as the component that does the work.
What is the difference between a self-improving system and fine-tuning?
Fine-tuning is a deliberate, human-run training step on a curated dataset, followed by evaluation and a new release. A self-improving system changes itself continuously in production based on its own outcomes. Fine-tuning changes weights and is hard to reverse selectively. Most practical self-improvement changes memory or configuration, which can be inspected line by line, rolled back, or deleted when a lesson turns out to be wrong.
How do you verify a self-improving system is actually improving?
Measure outcomes the system cannot influence directly. Track task success, human rejection rates, rework after merge, and time to completion per workflow step, and compare them across periods with the learning switched on and off where possible. Keep a fixed evaluation set the loop never trains against. If the internal score rises while those external signals stay flat, the system is learning the metric instead of the task.
Can self-improving AI systems operate across multiple software teams?
Yes, but lessons should not flow freely between teams by default. Conventions, test suites, and review norms differ, and a lesson that holds in one repository can be wrong in another. A safer pattern is to scope learning per workflow and let shared workflows carry shared lessons, with promotion to broader scope treated as a reviewed change. Governance, audit, and permissions should be set centrally even when learning happens locally.
How does a self-improving system handle conflicting feedback signals?
It needs an explicit rule for which signal wins, because averaging conflicting signals usually produces a lesson that satisfies neither. Weight signals by reliability: a failed deployment outranks a reviewer's thumbs-up. Treat disagreement itself as information by keeping the lesson tentative until the pattern stabilizes. When conflicts persist, surface them to a person rather than letting the loop pick a side silently.



