An assistant that drafts an email and a system that sends it to a customer do not have the same authority. Add the ability to search records, choose recipients, and update an account, and the design problem changes again. The label "AI agents" can conceal these differences unless the organization specifies what the system may decide and do.
A useful starting point is to separate output generation, workflow execution, and dynamic action selection. These are design choices, not inevitable stages that every organization should progress through. The best arrangement is the simplest one that meets the actual service requirements.
Distinguish three operating patterns
In an assistance pattern, a model proposes content or analysis for someone to use. A person decides whether and how to act. An example is a tool that drafts a troubleshooting explanation while the support specialist retains control of the customer interaction.
In a predefined workflow, software determines the sequence or permitted branches. A model may classify a request or generate an intermediate result, but the application controls the process. A request might be classified, checked against a rule, and routed to the appropriate queue.
In an agent pattern, the model has more discretion over the next step and its use of tools. Anthropic's engineering guidance uses the distinction between predefined workflows and model-directed processes to clarify these architectures. [1] Product labels alone are insufficient evidence of which pattern a system implements. Our research on AI agents and multi-agent systems explores this topic further.
Ask whether discretion solves a real problem
A fixed process can be appropriate when the steps and exceptions are well understood. Allowing a model to invent a route through that process may add cost and uncertainty without improving the result. Conversely, an investigation with several possible information sources may benefit from choosing the next step based on what has already been found.
Picture a facilities help desk. Routing a request to an approved queue can follow explicit rules. Investigating an unfamiliar equipment fault may require selecting among manuals and maintenance records. Approving replacement equipment introduces a separate financial authority.
Those tasks can coexist within one service without receiving the same level of autonomy. A useful design review asks which decisions genuinely require flexibility and which can remain fixed.
Understand the loop behind an agent
An agent typically receives a goal, selects a next step, requests a tool operation, observes the result, and decides whether to continue. The surrounding application supplies the tools, maintains state, and enforces limits. The model does not need direct unrestricted access to every underlying system.
Context management matters as the interaction grows. Anthropic describes approaches for maintaining useful information during extended tasks, including selecting relevant context and preserving state. [2] For implementation planning, this means asking what the agent needs to remember and how the application will recover after an interruption.
Set an explicit completion condition. "Resolve the issue" may be too vague for unattended execution. "Identify the relevant maintenance procedure and create a draft ticket with supporting references" is easier to evaluate. The boundary should be visible to both the user and the operator.
Authority must be enforced outside the request
An instruction asking an agent to behave carefully is not equivalent to restricting what its credentials permit. OWASP describes excessive agency in terms of excessive functionality, permissions, or autonomy. [3] These categories help separate what the system can technically do from what a user intended it to do.
For the facilities example, searching manuals might require no approval, creating a draft ticket could be allowed within defined fields, and committing expenditure could require a named approver. The backend should enforce those distinctions even if the generated request proposes something broader.
Apply limits to time, calls, and expenditure where relevant. Provide a stop mechanism and make interrupted or partially completed work visible. Match the details to the task; no checklist guarantees safety.
Evaluate actions and outcomes together
A convincing explanation is not proof that an agent completed its task. Inspect the resulting system state. Was the correct record changed? Was a duplicate created? Did the tool fail while the final message reported success?
Anthropic's agent-evaluation guidance discusses grading both outcomes and execution evidence. [4] In a local test, combine checks of final state with checks of required constraints. Repeated trials can reveal inconsistency that a single demonstration misses.
Also test refusal and escalation. A system should be rewarded for stopping when a mandatory precondition is absent. If evaluation scores only completed actions, developers may inadvertently optimize away the very safeguards the organization intended to preserve.
Choose a bounded place to begin
Start with a task whose outputs are reviewable and whose errors can be corrected. Define the allowed information sources, actions, and stopping conditions before increasing scope. Compare the result with a simpler workflow and include the cost of supervision.
Control points sit inside the action loop because oversight must influence execution, not merely inspect a final narrative. The practical question is not whether a product qualifies as an agent. It is whether the organization can explain, test, and limit the discretion it has delegated.
References
- Anthropic. Building Effective Agents. December 19, 2024.
- Anthropic. Effective Context Engineering for AI Agents. September 29, 2025.
- OWASP. LLM06:2025 Excessive Agency. 2025 edition.
- Anthropic. Demystifying Evals for AI Agents. January 9, 2026.
Sources checked October 2026. CorpExcellence.com articles are best-effort research and analysis, not professional advice.
Leave A Comment