We got used to a scene surprisingly quickly that, until recently, would have seemed unlikely: opening a conversation with AI, typing a request, and waiting.

A piece of writing. An explanation. An idea. A snippet of code.

If the answer missed the mark, we tried again. More context, another example, a better instruction. Gradually, learning to use AI came to mean learning how to ask.

Even when the answer surprised us, we still set the pace. The AI delivered something and waited. We decided what happened next.

That pace is starting to change.

With agents, the system can take on part of the journey between a request and a result. It looks up information, uses tools, checks what happened, and decides how to continue. Anthropic describes this distinction as the ability to direct the process and choose tools during execution.

Instead of following every exchange, we can define an objective and come back later to see the work.

“I ask, AI answers” is beginning to share space with “I define an objective, AI works.”

This does not mean prompts will disappear. We will still need to explain what we want. But a single instruction can now start a sequence of actions, with decisions along the way that we never spelled out individually.

The discussion around always-on agents, designed to operate continuously, makes this shift more visible. NVIDIA already describes systems built for this kind of persistent operation. What interests me most is the thought of work continuing while our attention is elsewhere.

What do we want to find when we return?

Consider a seemingly simple request: “organize my schedule for next week.”

AI could identify scheduling conflicts and suggest different times. It could also reschedule appointments, decline invitations, or send messages on my behalf.

All of those actions fit, in some way, within the word “organize.” I would not necessarily authorize all of them.

While we are talking, we can clarify that difference in the next message. When the system is already taking action, an invitation might be canceled before we realize we had not agreed on the boundaries.

I am beginning to find “what am I willing to delegate to AI?” more interesting than “what can I ask it?”

The first question makes me think about consequences.

In engineering, this becomes very concrete. I might want an agent to investigate a failure, read the code, and prepare a fix. Deploying that fix to a system customers use is a separate decision.

The objective may be the same: solve the problem. The authorization needs to differ at each stage.

For software developers, there is plenty of work in that transition. We need to verify what was done, understand which tests support a change, and prevent the access needed to investigate from turning into permission to change anything.

A result backed by evidence is far more useful than a message saying “fixed.”

This also puts a limit on the promise of simply leaving AI to work. In its experiments with long-running agents, Anthropic reports difficulties such as losing continuity between sessions and declaring work complete before properly checking it. Keeping a system running does not guarantee that it is moving in the right direction.

From an engineering leadership perspective, this discussion feels familiar.

When we delegate work to someone, we need to share enough context for them to make decisions without consulting us every minute. The expected outcome matters, along with constraints, commitments to other teams, and questions that remain open.

With AI, we need to make some of that context explicit and build technical boundaries around its actions. A phrase like “use good judgment” is no substitute for a clearly defined permission.

Perhaps one of the less visible challenges of adopting agents is discovering how much of our work depends on things that were never written down.

An exception everyone knows about. A customer who needs a conversation before any change. A priority that shifted in a meeting but has not yet reached the documentation.

Trying to delegate exposes those gaps.

For a company, buying access to agents may turn out to be the easy part. The harder work is deciding who can authorize what, which information can be used, and who follows an execution that crosses several systems.

That is where governance stops feeling abstract. It shows up in a refund decision, access to an internal document, or the ability to send a message to a customer.

There is also a difference between making a mistake and being manipulated into making one. An agent reading pages, documents, and messages can encounter malicious instructions mixed into the content it is supposed to analyze. Anthropic discusses this risk, known as prompt injection, and stresses that protection depends on several layers, including tools, permissions, and the environment.

The trust I care about here needs to be specific. Trusting a system to look up information is different from trusting it to share that information. Trusting it to prepare an action is different from authorizing it to act.

If I need to approve every move, I am still directing almost all the work. If I let everything happen without visibility, I may discover a problem too late.

The challenge is choosing the moments when my attention makes a difference.

I want to be brought in when there is a consequential decision, an ambiguity only I can resolve, or an outcome I have not yet authorized. To make that possible, I need to recognize those moments before handing over the task.

In “This photo was never taken,” I reflected on how making images easier to create increases the importance of verifying their origin. I see a similar concern here: as it becomes easier to set actions in motion, the need to understand who authorized them and how they were carried out grows.

There is something very appealing about leaving a conversation and finding work further along when I return. Research organized, a proposal compared with alternatives, a change ready for review.

I want to benefit from that without turning my day into an approval queue, and without having to accept an execution I cannot inspect.

We are teaching AI to do things on its own. Now we need to work out how far we want to let it go.

When I return and find the task complete, I want to be able to answer a simple question: did it do what I wanted, within what I authorized?