Writing · Agentic AI
Applying the AWS AI Security Framework to Agentic Workflows
How we translate AWS guidance into practical controls for identity, data, tools, authorization, observability and recovery across an agentic workflow.
An agent can be well prompted and still be badly secured.
The prompt may tell it to protect confidential data, remain inside scope and ask before taking a consequential action. Those are useful instructions. They are not the controls that should carry the risk. If the same model can reach excessive data, inherit broad credentials or invoke a high-impact tool without independent authorization, the workflow is depending on probabilistic behavior to enforce a deterministic boundary.
That is the architectural change introduced by agentic systems. A model that produces text creates one class of risk. A workflow that retrieves private context, maintains state, selects tools and changes another system creates several more. Security must follow the entire path from the requesting identity to the resulting action.
I like the AWS AI Security Framework because it is already an adequate structure for this work. There is no need to wrap it in a new methodology. It asks teams to classify what the AI can do, apply controls at several layers and strengthen those controls as the workload moves from prototype to production. The engineering work begins when those categories are translated into decisions at each workflow boundary.
Start with what the system can do
AWS separates AI use cases into three cumulative categories: systems that answer, systems that connect to organizational data, and systems that act. That distinction is more useful than calling everything an assistant, copilot or agent.
The classification should follow the most consequential capability in the deployed path, not the most reassuring description in the interface.
- An assistant that drafts text from only the user’s prompt primarily answers.
- A read-only assistant that retrieves internal documents or account data also connects.
- A workflow that can send, update, approve, execute or control something acts.
The controls accumulate. Giving an answering system access to private knowledge does not replace its input and output controls; it adds identity, authorization, data-classification and retrieval concerns. Giving that connected system a tool does not replace either set; it adds action authorization, agent identity, blast-radius control, human approval and a recoverable operating path.
The classification closes off a common design error: treating a low-frequency action as if it were outside the security model. If a tool exists in the path, its consequence belongs in the threat model even when the product expects it to be used rarely.
Apply controls across all three layers
The framework groups defense in depth into infrastructure, identity and data, and the AI application. Governance and compliance span all three. For agentic workflows, each layer should be able to limit failure without assuming that the others will work perfectly.
| Layer | Question it must answer | Typical engineering concerns |
|---|---|---|
| Infrastructure | Where can the workflow run and communicate? | Isolation, network paths, workload boundaries, encryption, resilience and resource limits. |
| Identity and data | Who is requesting, what identity is acting, and which data may each access? | Authentication, least privilege, scoped credentials, tenant boundaries, secrets and audit history. |
| AI application | How are probabilistic inputs, outputs, plans and tool requests constrained? | Input handling, guardrails, grounding, output validation, behavior monitoring and evaluation. |
Application guardrails can filter or shape model behavior; they do not authorize a user to read a record. Network isolation constrains reach, while identity policy decides who may access what. Neither proves that a model response is grounded enough to show or that an action is safe in the product’s current state.
The layers overlap by design. An indirect prompt injection arriving through retrieved content might be detected by application controls, but its damage should also be limited by data scope, tool authorization, network reach and action policy. Defense in depth means a miss at one layer does not silently become complete authority at the next.
Use a workflow-boundary ledger
Architecture diagrams tend to show components: user, agent, model, knowledge base, tools and services. For security review, the arrows are often more important than the boxes.
We use the idea of a workflow-boundary ledger to make those arrows reviewable. The ledger does not replace a threat model. It ensures the threat model covers every place where identity, data, instruction, authority or state changes hands.
| Boundary | Record explicitly | Failure to test |
|---|---|---|
| Requester to workflow | Human or service identity, intent, tenant and session scope | An authenticated caller uses a capability outside the intended role. |
| Source to context | Data classification, provenance, access decision and freshness | Retrieved content crosses a tenant boundary or supplies hostile instructions. |
| Model to decision | Proposed result, confidence limits and validation requirements | Plausible output is treated as verified fact or executable intent. |
| Decision to tool | Action class, parameters, target, policy result and resource limits | A valid tool is called with excessive scope or manipulated arguments. |
| Approval to authority | Approver identity, exact action, expiry and conditions | A vague or stale approval authorizes a different action later. |
| Execution to state | Observed outcome, side effects, rollback path and retained memory | A request accepted by a tool is recorded as a real-world result. |
| Workflow to operator | Escalation trigger, evidence, actions taken and unresolved risk | A person receives a transcript without the information needed to intervene. |
For every boundary, I want concrete answers to six questions:
- Which principal is acting?
- Which data is crossing, and how was access decided?
- Which capability is being requested?
- Which independent control can allow or deny it?
- Which event proves what actually happened?
- How can the workflow stop, degrade or recover safely?
This is where a framework becomes delivery practice. Broad principles are converted into interfaces, policies, traces, tests and operating responsibilities.
Phase one: build the prototype on a real security foundation
AWS calls its first maturity phase foundational. A prototype does not need every advanced control, but it cannot depend on an architecture that must later be replaced to become securable.
For an agentic workflow, the initial foundation normally includes:
- a defined requester identity and a distinct workload or agent identity;
- scoped access to only the required models, data and tools;
- encryption and managed secrets rather than credentials inside prompts or code;
- an audit event for model access, data retrieval and tool invocation;
- explicit input and output handling appropriate to the use case;
- a narrow action catalogue with resource, time and rate limits;
- human approval outside the model for consequential actions;
- a clear stop and escalation path.
The implementation may use different AWS services depending on the workload. The durable rule is that model choice does not carry identity, authorization, evidence or recovery. A model can be replaced without redefining who may act or what the action boundary means.
The model may propose an action. It should not be the component that grants itself the authority to perform it.
Phase two: harden the complete operating path
Moving from prototype to production changes the standard of evidence. A successful demo shows that the path can work. Production security must show how the path behaves when input is hostile, context is misleading, permissions are stale, dependencies fail or the model chooses an unexpected but technically valid route.
This phase adds depth around the foundation:
- classify the data available at each retrieval and tool boundary;
- threat-model both direct prompts and instructions embedded in external content;
- test tenant separation, identity propagation and confused-deputy paths;
- validate tool arguments and action policy independently of generated text;
- capture end-to-end traces without turning sensitive content into uncontrolled logs;
- define security events, containment actions and incident ownership;
- evaluate expected behavior before model, prompt, tool or policy changes are released.
The security test set should include more than jailbreak prompts. Agentic failure can begin with a legitimate user, a legitimate tool and a legitimate request whose combination is not authorized. It can also emerge over several individually plausible steps. Evaluation must therefore cover trajectories: what the workflow accessed, proposed, attempted, executed and reported over time.
A hypothetical document-analysis agent illustrates the point. The user is allowed to read one project. A retrieved document contains an instruction to search another project and send the result through an available messaging tool. Content filtering may identify the instruction. But the stronger design also prevents the second retrieval, denies the tool call, records both decisions and presents a safe explanation. No single control is asked to be perfect.
Phase three: govern the changing portfolio
At scale, the problem is no longer one workflow. It is a changing portfolio of agents, models, prompts, knowledge sources, tools, policies, identities and dependencies.
Advanced security therefore needs inventory and continuous evidence:
- register deployed agents, tools and externally connected protocol servers;
- version prompts, policies, models, evaluations and tool contracts together;
- detect permission, data-source and capability drift;
- maintain behavior baselines appropriate to each workflow’s agency and autonomy;
- automate policy checks and configuration review where the rules are deterministic;
- rehearse containment, rollback and graceful degradation;
- retain enough traceability to reconstruct the authorization and action chain;
- review whether each workflow has earned any increase in autonomy.
AWS’s companion Agentic AI Security Scoping Matrix distinguishes agency—the capabilities a system is permitted to exercise—from autonomy—how independently it may exercise them. That distinction matters operationally. Human approval can reduce autonomy while the underlying toolset still gives the system broad agency. Conversely, a continuously running monitor may have high autonomy but only narrow, read-only agency.
Progress should not be defined as moving every system toward maximum autonomy. The correct scope is the least autonomy and agency that can deliver the intended outcome. Increasing either should be a governed decision supported by evaluation, observability, incident readiness and a demonstrated reason.
Human oversight must be an engineered capability
“Human in the loop” is not a sufficient control description. The human needs a trustworthy view of the exact action, target, evidence, consequence and recovery path. The approval must be bound to that action, expire appropriately and be enforced by a component outside the model.
Oversight also changes as autonomy increases. Lower-autonomy workflows need people to approve individual actions. Higher-autonomy workflows need stronger portfolio review, behavior monitoring, intervention paths and audit discipline. Removing a confirmation step does not remove human responsibility; it moves that responsibility into system design and operation.
For high-consequence actions, ambiguity should reduce authority. The safe response may be to request more evidence, narrow the action, fall back to read-only operation or escalate. Graceful degradation is often more useful than a binary choice between full operation and a complete outage.
Physical-world workflows need a safety boundary too
When an agent can influence a connected product, industrial system or other physical process, information security is only part of the consequence model. The workflow must also respect product state, operational interlocks, physical authorization and a safe local fallback.
Cloud authorization should not silently override a device-level safety boundary. A model’s confidence should not replace deterministic checks. A successful API response should not be reported as proof that the physical outcome occurred. The system needs independent observation and a defined response when cloud state, device state and real-world conditions disagree.
Agentic workflows can be appropriate in physical products, provided their authority is designed around consequence rather than technical reach.
Framework use and compliance are separate claims
The AWS framework makes security work cumulative across use cases, layers and delivery phases. A list of configured services still proves very little about the security of the workflow that connects them.
We use it as a design and review structure:
- Classify the workflow by what it can answer, access and change.
- Map controls across infrastructure, identity and data, and the AI application.
- Record every workflow boundary where trust, authority or state changes.
- Build foundational controls into the prototype architecture.
- Add production evidence through threat modeling, evaluation and incident readiness.
- Govern scope, drift and control effectiveness as the portfolio grows.
The goal is deliberately modest: apply an established framework to the engineering path that data, identity and authority actually follow.
The final review question is simple. If the model behaves unexpectedly at this point in the workflow, which independent control contains the consequence—and which evidence proves that it did?
If the answer is only “the prompt tells it not to,” the workflow is not ready.
Sources and further reading
- AWS Security Blog, “The AWS AI Security Framework: Securing AI with the right controls, at the right layers, at the right phases”, defines three cumulative use cases, three defense-in-depth layers and three phases of security maturity.
- AWS Security Blog, “The Agentic AI Security Scoping Matrix”, distinguishes agency from autonomy and maps security considerations across four agentic scopes.
- AWS Security Blog, “Four security principles for agentic AI systems”, describes secure-by-design principles for agent identities, tools, permissions and observability.
- AWS Security Blog, “Threat modeling your generative AI workload to evaluate security risk”, applies threat modeling to generative-AI workload scope, actors, dependencies and trust boundaries.
- AWS Prescriptive Guidance, AWS Security Reference Architecture — AI security, provides architectural guidance for protecting AI workloads deployed on AWS.