The Confirmation Button is a Lie

5 min read
By Zach Herbet, Co-Founder & CEO of Foundation
Share
AI agents are now clicking, submitting, and approving actions on your behalf. The security model built around user confirmation sounds reasonable. But when the agent controls the environment, a confirm button is not a security boundary. It is a comfortable illusion.
The Confirmation Button is a Lie
Credits: DepositPhotos

For years, the loudest AI safety debate has been about what AI says. Content filters, jailbreaks, system prompt leaks, guardrails. But the most consequential developments in AI right now have almost nothing to do with outputs. Anthropic's computer use, OpenAI's Operator, GitHub's Copilot coding agent – these systems are beginning to act on behalf of users, clicking buttons, filling forms, reading emails, opening pull requests, and moving through websites designed for humans. The security model that has emerged around all of this sounds reasonable. It is not.

The implied model is broken

Most people assume the solution is obvious. Let the agent do the research, draft the email, queue the wire, then force the user to approve the final action. A button appears. The user clicks confirm. The system records consent.

At first, this sounds prudent. It seems to preserve human agency. It seems to draw a clear line between automation and authorization. The user is still in the loop. The machine still needs permission.

But where is that line actually drawn?

If the confirmation appears inside the same browser session the agent is using, it is part of the agent's world. If the agent can influence the page, the text, the form, the recipient, the amount, or the surrounding context, then the agent can influence what the human believes is being approved. OWASP lists prompt injection as the top risk for large language model applications. Anthropic warns about jailbreaks and malicious instructions. Once an agent has tools, these are not theoretical attacks. A confirmation button can slow an agent down and catch obvious mistakes. It cannot make prompt injection disappear. It cannot guarantee that the displayed action matches the executed action. It cannot protect the user from a compromised browser or a malicious page.

The confirm button is not a security boundary. It is a plausible deniability machine.

The new authorization model runs on separate authority

Bitcoin faced this problem years ago. If a laptop is compromised, the address or amount displayed by a software wallet may be false. This is why hardware wallets exist. They move the critical signing step to a separate device with its own display, its own input method, and its own security boundary. The signing ceremony matters. Bitcoin.org's wallet guidance warns users to secure wallets carefully precisely because transactions are irreversible. That irreversibility forged a better authorization model.

The same principle applies to AI agents. These systems are increasingly touching actions that are difficult or impossible to unwind – sending sensitive data, approving payments, modifying production code, granting account access. The question is not whether agents should act. They will act. The question is where authority lives.

If the agent can read the screen, click the mouse, fill out the form, submit the request, and shape the context around the approval, then the approval cannot be treated as independent. Involvement is not control. Consent is not authorization. Consent is a user saying yes. Authorization is a system proving that the right human approved the right action under conditions the attacker could not control. Those are not the same thing, and pretending otherwise creates security models that collapse the moment an attacker learns how agents work.

The skeptics' case: why it falls short

The obvious objection is that better UI solves this. If the confirmation screen is clearer, the user will catch discrepancies. If the approval flow is more prominent, humans will pay closer attention.

At first, this sounds reasonable. Better UX reduces mistakes. It gives product teams a clean answer to a hard problem. It lets everyone say the human is still in the loop.

But a user can approve a payment because the recipient name looked familiar while the underlying account number changed. A developer can approve a pull request because the diff looked harmless while the agent introduced a subtle dependency or permission change. The approval happened inside an environment full of ambiguity and attacker-controlled context. Better UI does not change that environment. It just makes the illusion more comfortable.

A second objection is that most agent actions are low-stakes. Not every task involves irreversible consequences. For routine automations, a confirm button is probably fine.

This is true, and it should drive a tiered authorization model, not an argument against better authorization. The goal is not to wrap every agent action in ceremony. The goal is to ensure that consequential actions, the ones that are hard to undo, are verified in a place the agent does not control.

Pay attention to the boundary

The AI agent story of the next few years will not be told in benchmark scores. It will be told in the quiet infrastructure underneath – the enterprise that loses credentials to a prompt injection attack, the payment that routes to the wrong account, the production deployment that goes out because a confirmation button got clicked at the wrong moment. Each is a small failure. Together, they form the case for why authorization design matters as much as model capability.

Anyone still arguing about whether agents need better confirm buttons is asking the wrong question. The right question is whether the authorization model you are building can survive contact with an attacker who understands how agents work. The builders who think clearly about where authority lives, and design systems where humans authorize across a boundary the agent does not control, will ship something genuinely safe. Those who keep polishing the confirm button will find out what it actually protects against.

Agents can prepare actions. Humans must authorize them across a boundary the agent does not control.

About the author

Written by Zach Herbret, Co-Founder & CEO of Foundation.

Related posts

© Copyright 2025 - Just AI News - All Rights Reserved
linkedin facebook pinterest youtube rss twitter instagram facebook-blank rss-blank linkedin-blank pinterest youtube twitter instagram