Skip to main content
Mubienclarity for what comes next
← Research
Understanding AI

When AI learns, what actually changes?

I use AI every day, and a better second answer still invites an easy assumption: it learned. That word can mean several different things. Separating them makes both the possibilities and the limits easier to see.

Start with the thing that changes

These distinctions help me ask a more useful question than whether a system is “getting smarter”: what changed, where does it persist, and what evidence supports the improvement?

01

Context: using what is in front of it

A model can use instructions, examples, and previous messages to produce a better answer. In-context learning describes this adaptation within the input; it does not require an update to the model’s weights.

For example: You show it the format you want, and its next answer follows that format.

A useful check: Does the improvement survive a new conversation without those examples?

02

Memory: saving information for later

An application can store preferences, notes, or previous results and retrieve them into a later interaction. That changes the information available to the model. The storage and retrieval system sits around the model and can have its own errors.

For example: An assistant saves your preferred briefing topics and retrieves them tomorrow.

A useful check: What was saved, is it still accurate, and can you inspect or remove it?

03

Training: changing model parameters

Training changes the model’s numerical parameters, or weights. Fine-tuning is a further training stage using selected data. A correction in a chat is not, by itself, evidence that a training update occurred. A provider may use conversation data in a later training process under its own settings and policies.

For example: A training process uses examples to update a model, producing a new version.

A useful check: What changed on independent evaluations, including tasks outside the training examples?

04

An improvement loop: proposing, testing, revising

An agent can generate code, run a check, read the result, and revise its attempt. The code or workflow can improve while the underlying model stays the same. The quality of the loop depends heavily on the check it is trying to pass.

For example: A coding agent fixes an error after seeing a failing test.

A useful check: Did it fix the underlying problem, or find a way to satisfy an incomplete test?

For the distinction between in-context learning and training, see Language Models are Few-Shot Learners. For context and persistent notes in agents, see Anthropic’s context engineering guide.

Where recursive self-improvement enters the picture

I use the term here for a loop in which a system helps improve the machinery responsible for its own capabilities, and those improvements help produce the next round. That machinery could include code, tools, training methods, or model design. The scope matters: improving one component establishes much less than open-ended improvement across tasks.

Google DeepMind’s AlphaEvolve is a concrete example of AI-generated programs being selected through automated evaluation. Its reported applications include improving parts of the infrastructure used to train AI. My reading is that this shows a useful feedback loop in bounded tasks. It does not, on its own, demonstrate unlimited or reliably accelerating self-improvement.

The limits are part of the mechanism

An improvement depends on the objective, the feedback, and the resources available. A test can miss important failures. Reusing generated outputs can carry errors forward. Gains on one benchmark may not transfer to a different task. None of these questions is answered just by running the loop again.

That is why I find the governance question interesting: if a system can propose changes to itself, what can it change, who decides whether the result is better, and what prevents an untested change from reaching the real world?

The checks I would want around the loop

  • A record of the version, the proposed change, and the evidence for accepting it.
  • Tests the system cannot silently rewrite to make itself pass, including checks for regressions and unwanted behaviour.
  • Explicit limits on the tools, data, and permissions it can change.
  • A separate decision to deploy consequential changes, with monitoring and a way to roll back.

These are starting questions, not a guarantee of safety. Whether they are sufficient depends on the system and its reach. I explore the human-review side in The Oversight Threshold.

A personal explainer based on the sources linked above. The examples are illustrative; this page reports no original experiments or measured results.