Back

Blog

Insights

Living off the AI: Microsoft's Digital Defense Report 2026 on AI agent security

Thierry Messmer

Microsoft published its Digital Defense Report 2026 on 5 October: 103 pages from Microsoft Threat Intelligence, covering July 2025 to June 2026. Its AI chapter contains one of the clearest descriptions we have read of the problem we work on every day: AI agents running on developer machines and CI/CD runners, with access to code, secrets and production.

Microsoft gives that problem a name: living off the AI. Below, we quote what the report says, with page numbers, then map it to what EDAMAME does. Microsoft does not mention EDAMAME; the mapping is ours.

1. Microsoft named the problem: living off the AI

Attackers have long "lived off the land", using signed system binaries so that their activity looks like administration. The report describes the same move with AI tools:

“Trusted binaries do the work. Signed, allow-listed AI CLIs and runtimes become ‘living-off-the-AI’ tooling—LOLBins with reasoning, web access, and code execution attached, making preventative host controls difficult to implement without affecting users.” (p. 16)

The example it gives is concrete. The s1ngularity malware, spread through trojanized Nx npm packages in the Pesky Pika supply chain attack of August 2025, probed infected machines for AI assistants. “If it found Claude Code, Gemini CLI, or Amazon Q CLI, it ran them with permissive overrides to hunt secrets, credentials, and Secure Shell (SSH) keys. Ultimately, it leaked ~2,000 secrets and ~20,000 files across 225 victims by targeting these trusted, pre-approved developer tools.” (p. 15)

What the attackers abused was not a flaw in those agents but their legitimate capabilities and the access the developer had given them: AI is now “both a tool and a target” (p. 8).

2. Why allow-lists and known indicators fall short

The report names three properties that defeat traditional host controls:

  • The tool is trusted. It is signed and allow-listed, so blocking it means blocking the developer: hard to do “without affecting users” (p. 16).

  • The behaviour changes on every run. “Behavior is not pre-programmed. The model generates a fresh action sequence on each run, so indicators captured from one incident generalize poorly to the next.” (p. 16)

  • The packages are new. On the supply chain side, “Malicious packages also often lack known signatures” (p. 36).

The question is no longer whether a binary is bad, but whether what it does matches what it was asked to do.

3. The developer workstation joined the supply chain

Open-source supply chain compromise is one of the three biggest threats Microsoft expects over the next year (p. 8), and the report is plain about where it lands: “The developer environment is the new perimeter. Repositories, build systems, and pipelines now sit inside the real security boundary, and code and configuration flow straight through them to production.” (p. 20)

“Within the past year, major developer ecosystems saw attacks roughly every ten days, with a clear acceleration through Q2 2026.” Axios alone “reached more than 100 million weekly downloads while under attacker control” (p. 59). The report attributes the Axios compromise of March 2026 to a state-sponsored group (p. 43). We covered how that attack behaves on a developer machine in our axios npm RAT write-up.

Two lines tie this to AI agents. Malicious packages “often execute automatically during installation or CI/CD workflows. This does not require any user interaction” (p. 36). And Microsoft expects adversaries to “manipulate package ecosystems to influence automated selection” by the AI agents that write code (p. 36): the agent picks the dependency, the attacker shapes what it picks. The conclusion: “the developer workstation is becoming part of the supply chain attack surface” (p. 60).

4. Guardrails inside the agent share its weaknesses

Pages 25 and 26, drawn from Microsoft's Red Team work, are the most useful for anyone securing AI agents. On memory poisoning: “Human-in-the-Loop (HITL) controls alone cannot protect memory writes. [...] The underlying issue is that the control and the thing it guards share a failure mode.” (p. 25) The team adds: “The Red Team has also observed instances where HITL is triggered but ignored by the agent and HITL being bypassed by a policy stored in memory in the same attack chain.” (p. 25) Without better defences, it warns, “the next class of incidents will be ones where the user ‘approved’ the breach themselves” (p. 26).

The same logic applies to what agents load. Self-hosted platforms such as OpenClaw pull executable skills from public marketplaces, and Microsoft treats them “as untrusted code execution with persistent credentials, fit only for isolated hosts under non-privileged identities” (p. 15). On skills: “Public registries already ship malware disguised as utilities; agent-initiated installs effectively run third-party code under the agent's identity.” (p. 15)

A control that lives inside the agent cannot be the only record of what the agent did: when the agent is manipulated, its approval prompts, its memory and its own account of its actions are manipulated with it.

5. What Microsoft says detection needs

Because agent misuse runs through “valid credentials and sanctioned agent workflows”, it “can closely resemble legitimate use”, so “Detection therefore depends less on identifying anomalous tooling than on establishing what normal looks like for each agent: which identity it uses, which tools it calls, and at what frequency, so that deviation becomes visible.” (p. 17)

It goes further: “Organizations need systems and controls that enforce policy in real time, keeping every action aligned to the agent's objective, its intent, and the organization's boundaries.” (p. 23) Its table of defences lists “intent validation” and “runtime gating” among the controls (p. 22).

And on the host: “Useful signal moves from process artifacts to patterns: unexpected egress to inference endpoints, AI CLIs run with permissive flags, anomalous shell history prompts, or assistant-driven file scans. This requires defenders to ensure they have a wider set of telemetry sources and develop new detections for these sources.” (p. 16)

A per-agent baseline, actions checked against intent, new host telemetry: that is watching the agent from outside the agent.

6. Governance that does not depend on the agent's vendor

Adoption makes this urgent. “Microsoft has observed 88% of enterprises are already experimenting with agents, and 82% of leaders plan broader rollouts within the next 12 to 18 months.” (p. 22) And: “Without centralized visibility, agent sprawl creates even greater risks than shadow IT.” (p. 23) Microsoft's first priority for leadership now includes agents: “Report identity exposure, patch latency, agent permissions, critical dependencies, dwell time, and recovery readiness.” (p. 7)

For teams running agents from several vendors, two recommendations stand out:

  • “Consistent governance policies need to apply regardless of each agent's origin. Using different rules for different platforms creates audit and visibility gaps that adversaries can exploit.” (p. 24)

  • “Treat revocation latency as a measured outcome. When an agent or its owner is compromised, the time to revoke across identity, posture, and threat protection controls is the operational metric that determines incident impact.” (p. 24)

In most engineering teams those agents are Claude Code, Codex, Cursor or OpenClaw, on macOS and Linux laptops and CI runners, and infostealers now “also target macOS and developer environments” (p. 52).

7. How EDAMAME maps to it, and what it does not do

EDAMAME is the AI agent trust layer: it proves the security posture and the runtime behavior of every AI agent and the machine it runs on, then binds access to code, secrets, production and critical company resources to that proof. The mapping below is ours, not Microsoft's.

  • Outside the agent, at the host. EDAMAME Security on laptops and EDAMAME Posture on CI/CD runners and any server running AI agents, cloud or self-hosted, observe Cursor, Claude Code, Claude Desktop, Codex, OpenClaw and Hermes from the operating system, with no plugin in the agent and no kernel driver of its own, so the record does not share the agent's failure mode (p. 25).

  • Intent divergence. EDAMAME compares the task the agent declared with what the host observed (process lineage, files opened, network connections): a shipped form of keeping actions “aligned to the agent's objective, its intent” (p. 23). More on the agents page.

  • Eleven deterministic attack-pattern checks. Credential harvest, token exfiltration, sensitive-material egress, skill supply chain, package install lifecycle, agent control tampering, process memory scrape, cloud metadata egress, sandbox exploitation, file-system tampering and denylist bypass. They grade the technique, not the indicator. In our September 2026 review, 11 of 12 publicly documented incidents were caught with no update (XZ Utils only partially, at exploitation), including two cases of what Microsoft calls “poisoned AI-assistant configurations” (p. 60), flagged as a write to the agent's configuration.

  • Access that follows the proof. EDAMAME Hub gates access on that evidence through Entra ID, GitHub, GitLab, Google, Netskope, Tailscale, NetBird and FortiGate. When a device stops passing, access is withdrawn through the Zero Trust controls the company already runs: revocation “across identity, posture, and threat protection controls” (p. 24).

  • AI governance across vendors. AI governance in EDAMAME Hub (Enterprise plan) discovers every AI agent and MCP server across the fleet, audits them against the same checks, and names the approved ones in allowlists, whichever vendor built the agent.

  • Not a sandbox. Microsoft's answer for self-hosted agents is isolation (p. 15). EDAMAME does not replace it: it detects an installed harness, reports agents still running unconfined next to it, and checks the security posture of the host itself.


EDAMAME Security, AI tab, Divergence monitor: agent activity mapped against the declared task, with aligned, outside and forbidden actions and a live divergence snapshot

Microsoft's own answers are identity, isolation, managed vaults with short-lived credentials, and correlation across its platform. EDAMAME complements all four and replaces none: it adds proof about the machine and the agent's runtime behavior to the access controls already in place.

What the report doesn't say. Microsoft does not mention or endorse EDAMAME or any other agent-security vendor; the mapping above is ours. The attack mix it reports for AI workloads (malicious link injection 52%, tool and agent abuse 12%, identity and credential abuse 10%, p. 23) comes from Azure AI workloads, not developer machines: it is not an endpoint statistic. The tools s1ngularity abused did what they were built to do; the report describes abuse of trusted tools, not a flaw in them. Autonomous attacks are a direction of travel, not today's norm: “In most observed real-world operations today, complex end-to-end intrusions still retain meaningful human direction” (p. 17). And the data stops in June 2026.

What to do with this

  • Run an agentic posture audit on your own machine. Download EDAMAME Security, free, and open the Agents tab: every AI agent on the machine, the MCP servers it uses and what each one can reach.

  • Read the white paper on intent divergence and the attack-pattern checks: EDAMAME white paper.

  • Read the report: the AI chapter (pages 9 to 30) and the supply chain pages (36, 59 and 60) are worth the time. Microsoft Digital Defense Report 2026.

Source: Microsoft, “2026 Microsoft Digital Defense Report”, version 1.1, October 2026 (reporting period July 2025 to June 2026). Quotes are verbatim, with the report's printed page numbers.

Thierry Messmer

Share this post