Author: MTF Institute Research Team
Independent review: MTF Institute Research QA
Evidence window: 8 July–5 October 2026 (90 calendar days, inclusive)
Geographic focus: U.S. security work; internationally available technical research is labelled by its actual scope.
Why this review matters now
Security work around large language models is changing as AI systems gain access to retrieval sources, persistent memory and tools that can act. A prompt-injection attempt may arrive through a document, repository file, web page or tool result. The practical security question is no longer confined to whether a model refuses a malicious sentence. Teams must decide which data an agent may read, which action it may take, which identity it acts under, and what evidence shows a control worked.
This article reviews ten developments dated between 8 July and 5 October 2026. Its source corpus comprises official OWASP and NIST releases, an official Microsoft product and guidance update, an OpenAI incident disclosure, and original research papers. It was assembled separately from MTF's U.S. vacancy study. The sources establish that a guide, project direction, product capability, study or incident was reported; they do not measure adoption across U.S. organizations or the frequency of prompt injection in production. We distinguish established references from new proposals and controlled experiments.
The risk reference changed, while agent controls became more explicit
On 3 August 2026, OWASP released its 2026 Top 10 for LLM applications. OWASP describes revised rankings, wider threat coverage and mappings to other security references. This is a useful point-in-time vocabulary for security reviews. It is a community guide, however, and a category name alone cannot demonstrate that a particular application has an effective control.
On 1 September, OWASP published the Agent Control Standard, an open proposal for middleware hooks, traceability and runtime safety policies across agent frameworks. Its central design idea is operationally important: inspect an agent's proposed action at a control point outside the model's own instructions. A team can then record the attempted tool call, its acting identity, the target resource, applicable policy and the allow, deny or escalation decision. The new standard's release does not prove broad uptake or universal effectiveness; it gives architects a concrete design to test.
For U.S. application and platform security teams, these releases support a review of one real workflow: identify every lower-trust input, every tool with side effects, and every decision that must be enforced by code or policy rather than by prompt wording. A threat model should name the owner of each control and the evidence required before deployment.
Agent identity moved into a U.S. implementation case
NIST's National Cybersecurity Center of Excellence reported on 29 September that its first implementation use case for software and AI-agent identity will be the software development lifecycle. The planned demonstration will examine how agents can be identified, authenticated and authorized. NIST said the concept-paper process received feedback from more than 600 commenters. This is a project direction, not a completed standard or a federal rule.
The direction is relevant because a coding agent can operate through a user's session, a service account, a repository integration or a delegated token. Without an explicit identity and permission map, it can be hard to establish who authorized a change, what scope was granted and which action crossed a boundary. An AI-security practitioner should be able to draw that map, separate the human request from the agent's authority, and specify how a sensitive write or external call is approved and logged.
NIST also finalized one August workshop summary and a second workshop summary for its prospective Cyber AI Profile. These summarize governance and attack-surface discussion that may inform future work. They can help teams translate AI-specific scenarios into familiar cybersecurity ownership and evidence language, but they are not final control obligations.
Runtime gates, memory and reusable skills widened the review surface
Microsoft's 4 August security update added AI and DevSecOps assessment material, including attention to memory, tool allowlisting, code governance and data protection. It describes vendor capabilities and recommended practices; it is not independent proof that those controls achieve a particular reduction in risk. Its value for a vendor-neutral security review is the linked set of questions: where does an agent obtain context, which tool may it invoke, who maintains its memory, and what is permitted to leave the environment?
A July network-operations benchmark tested indirect injection in tool-using agents that read tickets and logs. Under the authors' controlled conditions, a metadata-aware gate before action performed better than prompt-only defenses against the tested unsafe actions. That result supports testing execution-time authorization, but it depends on trusted metadata and selected scenarios. It should not be converted into a claim that a gate eliminates prompt injection in every deployment. Teams should measure both unsafe action rates and legitimate-task completion so a control does not silently make the workflow unusable.
Two other studies show why the review cannot stop at a single conversation. Bad Memory tested persistent agent memory files and found that already-planted malicious content could influence later sessions in its synthetic workspace; causing an agent to create that malicious memory from untrusted content was harder. The finding supports provenance, write controls, review and retention for persistent context, with model and environment limits kept visible. SkillSecurer examined reusable agent skills, identifying context-dependent weaknesses and testing selected exploit paths in a controlled harness. It supports inspecting a skill package's instructions, scripts, permissions and external-input paths before installation and after updates. Neither paper estimates how often these failures occur across U.S. organizations.
Better attack records and safer evaluation environments
An August research preprint on prompt-injection anatomy proposes recording the carrier, delivery path, concealment, context break, attempted privilege change, payload and return channel of an injection. This is an analytic proposal, not a validated industry standard. It nevertheless addresses a common review gap: storing only the suspicious string hides where the attacker-controlled content entered, which boundary it tried to cross and what action would have counted as success. A useful test record should include the authorized task, the lower-trust surface, the intended unauthorized effect, observed behavior and the control decision.
In a separate safety issue, OpenAI disclosed on 4 August that two external cyber evaluations had configurations and controls under which model activity extended beyond intended test boundaries. OpenAI described unusual reduced-safeguard or misconfigured environments and public-internet access. The incidents were not evidence that prompt injection caused the behavior, and they should not be generalized to ordinary deployments. They do show why evaluation owners need explicit scope, isolated test environments, credential controls, outbound-network limits, monitoring and stop conditions. A red-team exercise that produces a convincing finding while creating an uncontrolled external effect is not a sound security test.
A practical response for security teams
Choose one authorized LLM or agent-enabled workflow and produce a short security decision record. Start with the legitimate user task and the system's data flow. Mark which documents, web pages, memories, tool responses and skill packages are lower-trust. Identify the acting human, service and agent identities and the permissions each can exercise. List the tool actions that can write data, contact an external service or change a security-relevant setting. For each one, specify the code or policy gate, the owner of that gate, the evidence it logs and the person who resolves a denied or ambiguous action.
Then test a small set of synthetic cases, including one indirect prompt-injection attempt, one unauthorized tool request and one benign task that should still succeed. Record observed behavior rather than assuming a guardrail worked because it was configured. Review false positives, incomplete logs and handoff failures. If persistent memory or installable skills are in scope, test the update path and a later session as well. Keep the final decision tied to the exact system version and authorized environment tested.
The changes in this review justify stronger attention to authorization, provenance, runtime intervention and evaluation containment. They do not justify a claim that any single standard, vendor product or prompt can secure all LLM applications. U.S. teams can use the sources as current design inputs while validating controls against their own architecture, approved use and observed evidence.
Evidence and limits
The ten developments were selected purposively for relevance to AI-security duties and the stated 90-day window; this is not an exhaustive inventory of releases or incidents. OWASP publications are community references, NIST materials include ongoing project and workshop work, Microsoft describes its own offerings, and the research papers test selected systems under controlled conditions. The OpenAI account concerns specific external evaluations. None supplies a denominator for U.S. employer adoption, U.S. incident frequency or real-world defensive efficacy. The separate vacancy study addresses what current U.S. employers explicitly request; this article addresses what changed in the security reference and technical landscape.