Work / Argus
Seeing what AI agents actually do
In-house R&D: a runtime observability layer that shows what an AI agent really did on a machine, not just what it reported.
Status: Research prototype
Why we built it
Teams are giving AI agents real access to files, networks and APIs. Logs show what an agent says it did. Nobody sees what it actually did.
Stack
- Cilium Tetragon (eBPF)
- Kubernetes
- Node.js
- TypeScript
- PostgreSQL
- Redis
- Next.js
- Python SDK
What we did
- Watches the system at kernel level with eBPF, below anything an agent can skip or fake.
- Links each agent action to the real system calls and network traffic underneath it.
- Flags behaviour the agent never declared.
- Streams sessions live to a dashboard as they run.
- Keeps the correlation engine honest with an evaluation harness built on real captured data.
Architecture
Cilium Tetragon captures file and network activity at the kernel. An ingestion service writes every event to Postgres and publishes it over Redis, and a WebSocket API streams it to a live dashboard. A small Python SDK marks where each agent session and action begins and ends.
How correlation works
A multi-signal engine scores every kernel event against the actions the agent declared, using timing, process identity, network destination, file path and the function it expected to run. Activity that no declared action explains is surfaced on its own. An evaluation harness built on real captured traffic runs in CI, so changes to the engine are measured, not guessed.
Security model
Argus runs beside the agent, not inside it, and records at the kernel, so an agent working with ordinary permissions has no way to switch it off or edit what it saw. That assumes the agent cannot gain root on the host or tamper with the kernel; an attacker who can do that is outside what Argus is built to catch. It filters its own components out of the event stream so it never observes itself into a feedback loop. Argus is a research prototype, not a product we sell.
Outcome
The same principle goes into every AI system we ship: if you cannot observe it, you cannot trust it in production. Argus is where we push that idea to its limit.

