Back
Blog
Insights
Endpoint Security for AI Agents: Fast Signals, Slow Judgement

Minh Anh Day
On 5 October, Zscaler announced a limited preview that puts Anthropic's Claude Mythos 5.1 inside Zscaler Endpoint AI Security, powering a capability they call Exploit Paths: correlating endpoint AI activity into an attack chain, with confidence levels, and showing defenders where to interrupt the sequence.
I build the detection engine at EDAMAME, and I think this is good news — worth saying plainly before I say anything else. When the largest zero-trust vendor in the market puts a frontier model on the endpoint to reason about what AI agents are doing, it settles an argument we have been having for two years. The endpoint is where agent risk lands. Not the model provider, not the gateway, not the prompt. The machine where the agent holds a developer's credentials and reaches a developer's filesystem.
So the category question is closed. What is open is the engineering question underneath it, and that is the one worth a post: what has to be true on the endpoint before reasoning over agent behaviour produces something you can act on, quickly, at a cost that survives contact with a real fleet?
Everything below is about our own system and our own measurements. I am not making claims about anyone else's.
Three things the announcement gets right
The first is that the unit of analysis is the chain, not the event. An agent with broad permissions calls a third-party tool, then reads a sensitive file, then opens a socket. Each step is unremarkable. The sequence is the attack. Any system that grades events one at a time will be quiet through the whole thing — and we built our own scoring model around exactly that observation.
The second is that AI activity on the endpoint now needs its own discovery. Which assistants, which agents, which extensions, which local models, which connected services. You cannot reason about what you have not inventoried, and most organisations don't actually know what is installed on their developers' machines this quarter.
The third is the framing of inferred versus confirmed steps. An attack chain assembled after the fact will be a mix of both, and a tool that presents them as equally solid will burn an analyst's trust on the second investigation.
Where a model-first reading of the endpoint meets physics
We shipped a language-model adjudicator in our own detection pipeline, and we instrumented what it costs. The numbers below are from our internal test benches and our own continuous integration, not from customer production, and each names its run. They are the reason our architecture looks the way it does.
Latency. Across our last eight release-gate runs, an engine tick that consulted the model had a median duration of 17.0 seconds, with a 90th percentile of 22.2 seconds. A tick with nothing to adjudicate finished in under a second. The deterministic evaluation on its own — eleven checks over measured telemetry, severity grading, evidence scoring — takes tens of milliseconds. The model is therefore not one component among several in the latency budget. It is roughly three orders of magnitude more expensive in time than everything it sits on top of.
That is tolerable when the output is an explanation a human will read in the morning. It stops being tolerable the moment the output is supposed to interrupt something.
Cost, and what it buys. One week of our own CI traffic, in the seven days to 15 September, cost 10.85 million tokens of adjudication. That is a development pipeline, not a fleet of customer laptops. Multiply it by a real population of developer machines, each ticking continuously, and the arithmetic gets uncomfortable quickly, no matter where token costs settle.
Dependency. Over the same seven days, 0.19% of adjudication calls exceeded our client timeout and 0.17% failed at the provider. Each one is a tick where the detector had evidence and no verdict. Worse, in the preceding fortnight we recorded 23,436 quota refusals: our own company account, shared by every CI runner and dogfood host, had spent its lifetime token cap, and every device on it was refused for about two hours — our release gate among them. A model quota is a single point of failure across every device that shares it, and the reset cadence, not the cap, decides how long those devices stay blind.
None of that is an argument against frontier models. It is an argument about where in a detection pipeline they belong, and what has to be able to run without them.
Jev made the same point in public, in September
In September, TypeSafe released Jev, the first of what they call System One models — the name borrowed from Kahneman's distinction between fast, automatic judgement and slow, deliberate reasoning. Jev does not write sentences. It takes a state and a question and returns a typed answer with a calibrated probability, in well under a second. Their own documentation is explicit about the boundary: if a question needs extended reasoning or weighs several independent factors at once, decompose it rather than asking Jev to deliberate.
The pattern that fell out of it — accept the answers the fast layer is confident about, escalate only the rest — is not new in principle. What Jev made concrete is the economics. Once the routing decision itself costs a fraction of a cent and returns in milliseconds, the case for sending every event to a frontier model collapses on its own merits. You are paying deliberation prices for reflexes.
For security work the lesson is sharper than for most domains, because our question volume is enormous and our latency budget is small. A developer workstation running a coding agent generates a continuous stream of process executions, file opens, and network sessions. Most of them are nothing. Deciding that is reflex work. Explaining the handful that are something, in language a human can act on, is deliberation work. Those are different instruments, and the error is using one for both.
Brute force over weak signals is the expensive mistake
There is a failure mode I expect to see a lot of over the next year, and it is not really about models at all.
If what you feed a frontier model is a process name, a log line and a destination, then no amount of reasoning on top will recover what the sensor never captured. The model will produce a fluent, confident narrative built on three weak signals, and it will do so at frontier prices, thousands of times a day. Reasoning does not manufacture evidence. It can only weigh what the measurement layer actually gave it.
Which is why the hard, unglamorous work sits below the model, in the sensors. On our side that means kernel-level process execution with lineage, network sessions attributed to the process that opened them, file-integrity events that name the writing process, and the live process tree — captured through the platform's own mechanisms on macOS, Linux and Windows. The reason a split-process relay attack is catchable at all — one process holding credential files open and streaming them to localhost, a sibling reading that socket and exfiltrating externally, neither half suspicious alone — is that the measurement layer sees the lineage. It took us three attempts to close that one, and none of the attempts were model problems.
Models aren't magic. They can't make good decisions on weak information, and they can't make any decisions instantly.
What we built: deterministic first, model last, model optional
EDAMAME keeps two planes separate and never lets them merge.
The reasoning plane is the agent's own account of its work, recovered from the transcripts it writes to disk. It is rich, high-level, and directly attacker-influenced: an agent under a prompt injection will describe its actions in whatever terms the injection chose. We never trust it alone. The system plane is what the host measured. The agent cannot edit it. It is narrow, low-level, and knows nothing about intent.
Neither is sufficient. Their disagreement is the signal — that is what we call intent divergence, and it is the measurement at the centre of our research with the LISTIC laboratory.
Two engines consume those planes. The divergence engine correlates declared intent against observed telemetry, and treats a declaration that matches a known-bad destination as an aggravating factor rather than an excuse. The attack-pattern detector ignores intent entirely and runs eleven deterministic checks against measured telemetry alone — credential harvest, token exfiltration, sensitive-material egress, package install lifecycle, process memory scrape, cloud metadata egress, agent control tampering and the rest — so that an agent which declares nothing is not thereby invisible.
Both feed a scoring model that grades each finding on the strength of its evidence rather than on a severity fixed per check. A process whose only suspicious property is that it lives in a temporary directory is the shape of both a dropper and an ordinary installer; emitting that at HIGH forces an allowlist that never finishes, and every entry in it is a hole.
Then, and only then, the frontier model. Its authority is deliberately bounded: it never adds a finding and never raises a severity. It may keep, demote or suppress, under tier rules that give it no authority at all over corroborated critical findings. Two things follow by construction. Recall is set by the deterministic layer, so the model cannot be the reason we detect something. And its entire contribution lies on the false-positive side, plus the prose of the alert.
Because of that shape, the model is optional and the choice is the customer's. Each engine carries an adjudication mode: consult the model and withhold if it does not answer; consult it but publish the deterministic result when it is unavailable; or never consult it at all — no provider, no key, no quota on the path. Customers run our managed model, their own Anthropic or OpenAI key, or a local Ollama instance. Agent transcripts stay on the device.
What the model actually does, measured
We thought this was worth measuring rather than asserting, and the answer is not the flattering one a vendor would write.
Across 4,710 per-finding verdicts in the seven days to 15 September, the adjudicator returned keep on 97.6% of what the deterministic layer produced. It suppressed 2.1% and demoted 0.3%. Only 59 of 1,349 adjudications quieted anything at all. That sample is our continuous integration, which generates a narrow and repetitive candidate set — it is emphatically not a customer fleet, and we are explicit internally that the number sets a floor rather than an estimate for what a working machine looks like.
But the direction is the point. On our release gate of 23 attack-pattern scenarios plus a divergence scenario, run against real triggers on four platforms, the model's contribution to true positives was nil by construction: had the deterministic result been published directly, the same alerts would have been raised. On the last main-branch gate before this was written, run 34902330391 of 14 September, all 95 cells passed on all four platforms, and every platform also sat through a ten-minute idle window producing zero findings at any severity. That second constraint is the one that keeps the detector honest — a check that cannot stay quiet on an unremarkable host does not ship, whatever it catches.
The model earns its place on the other side of the ledger: it quiets the long tail of odd-but-benign behaviour on real machines, and it writes the explanation a human reads. Those are real and they are worth a quarter of a minute. They are not detection.
The autonomous endpoint is coming, and it needs a foundation
I agree with where the industry is pointing. Within a reasonable horizon we will have security agents on the endpoint that respond to threats on their own — isolating a process, revoking a token, cutting a session, without waiting for a human. That is a better world than the one where an alert queues behind forty others.
It will not arrive on top of a brute-forced language model over weak signals. An autonomous responder needs four things that a reasoning layer cannot supply by itself:
Signals fast enough to act on. A response decision taken seventeen seconds after the evidence is a response to history. The reflex layer has to be deterministic and local.
Provenance for every action taken. When an automated response cuts someone's access, the evidence trail has to say exactly what was measured, by which sensor, and which rule fired. "The model judged it risky" is not an audit record.
An observer the observed cannot silence. This is the load-bearing property of the whole design: we watch the agent from outside it, at the operating-system boundary, with no SDK and no plugin inside the agent. An agent that can disable its own monitor is not monitored, and the agents' own security documentation says the same thing.
Bounded authority. Automatic response is a privilege, and it needs the same kind of guard rails we put on our adjudicator: things it may never clear, severities it may never lower, actions that always require corroboration.
And it needs one more thing that is easy to skip in the excitement. The new attack surface is additive. AI agents are now a first-class part of the endpoint attack surface, but they sit on top of the old one, which did not go anywhere: operating-system hardening, network exposure, configuration drift, credentials lying in the open, the machine itself. An agent running on a compromised laptop is a compromised agent, and no amount of behavioural reasoning about the agent recovers that. This is why we have always paired the two halves — posture (is the machine hardened and compliant?) and behaviour (is the agent doing what it declared, and staying clear of attack patterns?). Trust is both, or it is neither.
That pairing is also what makes the detection generalise. In a September review of twelve publicly documented AI-agent and supply-chain incidents, our deterministic checks caught eleven with no update — because an attack that harvests credentials or exfiltrates a token has a measurable shape on the host, whatever its name turns out to be in the write-up.
End to end, inside the zero-trust stack you already run
The last piece is that none of this should require replacing anything.
Since July, EDAMAME has been integrated with Netskope. While a device passes your EDAMAME policy we apply a device tag in your Netskope tenant; when it stops passing, the tag comes off, and your own Real-time Policy — which you wrote, and which we never touch — decides what that tag grants. EDAMAME labels; Netskope enforces. Re-tagging is automatic once the device recovers.
What changed is the definition of "passes". It now covers the AI agents on the machine: observed, fenced by a harness, still inside that fence, and behaving as declared. When an agent starts behaving like an attack — credentials touched, tokens moving, intent diverging from activity — the tag comes off and access closes. Nobody opens a console.
That is endpoint detection and response for AI agents, end to end, assembled out of two planes that each do what they are good at. Netskope owns the network, cloud and AI-traffic plane and enforces there. EDAMAME owns the host and agent plane, where the attribution actually exists — which agent opened that socket, under which declared task, from which parent process, having just read which file. The network sees a destination and a volume. Only the host sees the sentence.
Two details matter to the people who have to deploy it. The device does not need to be managed: EDAMAME is user-installed with no MDM, which is what makes contractor laptops, BYOD and personal developer machines reachable at all. And the same observation runs headless on CI/CD runners and on servers hosting agents, which is where a growing share of agent execution now happens, and where nobody has an MDM to begin with.
Where this leaves us
A frontier model on the endpoint is a good idea, and the vendors moving there are moving in the right direction. Our argument is narrower and, I think, boring in the way infrastructure arguments usually are: the reasoning is the last ten percent, and it is only worth what the measurement underneath it is worth. Fast, deterministic, local signals with real provenance; a frontier model for the final judgement and the explanation; hard guard rails so the judgement can never quietly clear the evidence. Then automation on top of that, once the foundation holds.
We publish our blind spots too — a patient attacker exfiltrating a single credential slowly over ordinary HTTPS to a routine destination stays below our thresholds by construction, and that sits in the limitations chapter of our own documentation rather than buried in a footnote. A security system's blind spots are the part of its documentation most worth reading, including ours.
If you are evaluating how to secure AI agents on your endpoints and you want to see this on your own machines, the product page is here: https://www.edamame.tech/agents
And if you run Netskope, the write-up of how the two planes connect, including setup, is here: https://www.edamame.tech/blog-full/netskope-conditional-access-ai-agent-posture

Minh Anh Day
Share this post



