Assume the agent will be given the wrong instructions
An agent reads the open web, and web pages contain words. What stops something it reads from becoming something it obeys?
Where this sits: The agent reads. Product pages, reviews, descriptions: text the agent did not write and nobody vetted.
Claims last checked against sources on 16 September 2026. Protocols and guidance in this area change quickly, so treat anything here as accurate as of that date rather than indefinitely. The primary source registry shows the main specifications and payment programs used across the framework.
What goes wrong
An agent shopping on your behalf reads product pages, reviews, descriptions, and whatever else the open web puts in front of it. Somewhere in that text is a line that says, in effect, ignore what you were told and do this instead.
A person scrolling past sees nothing. The text may be invisible, or buried in a review, or sitting in a page element nobody renders. The agent reads it and treats it as direction.
The reason this works is structural rather than careless. The model ultimately has to process both your trusted instructions and the untrusted page it fetched, and unlike ordinary software it has no reliable security boundary guaranteeing that one will be treated as instructions and the other only as data. Systems can and do put mechanisms around this, including delimiters, isolation and instruction hierarchies, but dependable separation remains the open problem.
That is why this is the attack that keeps working. The security research teams at Palo Alto and Forcepoint have both found it running on live sites, and Forcepoint's April 2026 work includes payloads aimed at moving money.
Why this control is different
Every other control here tells you to build something. This one starts by telling you what you cannot build.
There is no filter that reliably catches injected instructions, and anyone selling you one is selling you something. Retrieval and fine-tuning can help with other problems, but neither fully mitigates . Ariel Fogel, a researcher at OWASP — the non-profit whose security risk lists the industry treats as a baseline — put it plainly in June 2026: prompt injection remains an unsolved architectural problem.
The security guidance has moved towards containment rather than promises of prevention. OWASP says it is unclear whether foolproof prevention is even possible, given how these models work, and frames its recommendations around reducing impact. Microsoft is blunter, telling organisations to assume indirect prompt injection will happen and then limit what happens next. Australian cyber-security guidance published in May 2026 points the same way, recommending human supervision and approval where actions are high-impact or hard to reverse, and saying that this call belongs to the people designing the system rather than to the agent.
So the honest question is not how to keep bad instructions out. It is how much damage one can do on the day it gets in. A control that cannot eliminate a risk can still decide how expensive that risk is, and pretending otherwise is how this sort of writing stops being useful.
What this changes legally
Something shifts once a risk is publicly documented, and it is worth naming.
Once an attack class has been documented on live websites and written up by named security researchers, it becomes much harder to argue that the risk was unknowable. That does not by itself establish negligence. Foreseeability, and the question of what precautions a reasonable operator should have taken, stay fact-specific and vary by jurisdiction.
What the published material does give you is a clear sense of what a careful operator is expected to know. Microsoft's own 2026 guidance tells organisations to design on the assumption that indirect prompt injection will happen, and to put human verification behind risky actions. The defensible position is not that you prevented it. It is that you knew, and that you built so that it mattered less.
What to build
Design from the assumption that an injected instruction will eventually get through, and put your effort into what happens next.
Keep reading and acting apart. One architectural pattern worth knowing uses two models rather than one: a trusted model that sees only the buyer's instruction and produces a plan, and a quarantined model that handles fetched pages and returns data rather than direction. This does not solve injection. What it does is stop the exposed part of the system from directly wielding the authority to spend, which is a different and more achievable thing.
Cut down what the agent can do at all — : tools scoped to this task, credentials scoped to this purchase, and no standing access to anything it does not need today. The limit of this one is worth knowing too. Least privilege does not stop an attack that misuses a tool the agent legitimately holds, so it shrinks the blast radius rather than preventing the blast.
Put a person in front of the irreversible things — spending above a threshold, anything that cannot be undone, anything outside the pattern of what this buyer normally does. A confirmation step is unfashionable, and it is the control most likely to actually save you.
Re-check permission after the agent reads anything it did not control, which is the same point C2 makes from the other direction. The moment it ingests a page is the moment its instructions may have changed.
Record what it read before it decided. Not for the model's benefit, but for yours: when something goes wrong, you will need to know which content was in front of it.
What proves it worked
For any purchase that matters, you should be able to reconstruct what the agent had read before it committed, and show that untrusted content never crossed into the part of the system that acts.
The boundary is the real test of the two. If you cannot point at it on a diagram and say what does and does not cross, you probably do not have one.
And as with C2, blocks are evidence. A system that has never refused an action has either never been tested or is not really checking.
What the rules actually say
Nothing binding, and the useful material is guidance rather than law.
Prompt injection sits at the top of OWASP's risk list for large language model applications. OWASP recommends a set of mitigations rather than claiming any single preventative control, and Microsoft explicitly calls for defence in depth, combining probabilistic and deterministic measures. Both converge on layers — architecture, checks at the point of action, and governance above them — rather than one guardrail carrying everything.
No regulation I am aware of requires any of this specifically. But if a dispute ever turns on whether a system was built with reasonable care, published guidance describing a known attack and the expected response is exactly the sort of thing that gets cited.
Written in a personal capacity, from public sources. It is not legal advice and does not create any professional relationship. Where a specification, a piece of research or a set of guidance is named, it is named so a reader can go and check it. Nothing here is a judgement about any company's conduct or compliance.