Agent security and governance, dev through production

What you declared, and what is actually happening.

Everything we do is that one comparison, at three levels. The agents you wrote down against the ones we observe. The authority an agent was granted against the behaviour it shows. The changes somebody declared against the ones that merely happened.Whatever sits in the gap is the finding, and it is a finding in dev, at the gate before you ship, and in production.

Read-only to beginNo agent redeployDev, staging and productionAgents, tools, MCP servers and skills
Level one · the estate

The roster you wrote down, against the estate you actually run.

You maintain a list of the agents you know about. Agenor builds a second list from behaviour alone, out of telemetry you are already emitting, then reconciles them. An asset takes its kind from your declared roster, and an undeclared one is reported as a shadow rather than quietly typed. Every finding carries its app and its environment, so a shadow in dev reads differently from the same shadow in production, and neither is hidden from you. It is not only agents: tools, MCP servers, skills, identities and permissions are inventoried the same way, because that is where an agent's real reach comes from. Sources are first-class too, and a collector with no owner reads owner null, since an unowned source is a finding rather than a blank to fill with a default.

What a trace-based inventory cannot seeWe find agents from the telemetry they emit. An agent that emits nothing is invisible to us, so this is a floor on your estate and never a certified ceiling. Where a framework or a gateway can be made to emit, we will tell you which one and what it costs; where it cannot, that gap is named in the report rather than rounded away.
ESTATE · RECONCILED5 declared7 undeclared
billing-opsprodstripe · ledger-dbdeclared
support-triageprodzendesk · customer-db(read)declared
invoice-botprodnetsuite · s3://invoicesdeclared
deploy-helperstggithub · k8s-stagingSHADOW
crm-writerprodsalesforce(write)declared
spend-guardprodstripe(read) · slackdeclared
eval-harness-3devopenai · s3://eval-outSHADOW
mcp-filesystemdevlocal-fs · repo checkoutSHADOW
notebook-runnerdevwarehouse(read) · pypiSHADOW
media-buyerstggoogle-ads · stripeSHADOW
intern-copilotdevgithub · jira · customer-db(read)SHADOW
prod-debug-agentprodk8s-prod · logs · secrets-managerSHADOW
7 of 12 observed in telemetry appear on no declared roster · specimen data
One platform, six rungs

Start wherever your problem actually is.

Most teams come in at rung one or two and stop there for a while, which is fine. Each rung is useful on its own and each is built on the same ingest, so moving up is a setting rather than a migration. The three product lanes are named beside each rung.

01 / SEE

Inventory and shadows

for: platform and security leads with no register

Every agent, tool and MCP server observed in your traces, what each one reached, and which of them never appeared on your roster. Read-only.

lane Detect
in OTLP GenAI spans you already emit
out reconciled estate, shadows named
02 / QUALIFY

Is this ready to go to production?

for: the sign-off before an agent ships

A scoped assessment against a candidate agent in dev or staging: detection over blast radius, crown-jewel reachability and attack paths, plus probes fired at it. You get an answer for security and for governance, and the unknowns are listed as unknowns rather than averaged into a number.

lane Detect
answers a range, floor and ceiling, not a score
grades partial while anything is unknown
03 / WATCH

Continuous posture and drift

for: anyone whose estate changes weekly

The same detection on a live stream, with incidents that dedupe, recur and carry ownership. Claim, assign, triage, escalate, suppress, accept risk, and push to Jira, ServiceNow, PagerDuty, Slack or a webhook.

lane Detect
suppress requires an expiry, then reopens itself
per source fresh, idle, stopped, expired-credential
04 / GOVERN

Declared change, against change that just happened

for: governance, and anyone who owns the approval

Operators declare changes to a prompt, model, tool, MCP server, skill, identity, permission or configuration, with before and after versions and a stored diff. Agenor derives the other half from captured history. Both land on one timeline, which is what lets a declared change be told from one that merely happened. Approve or reject with a second pair of eyes, and drift is never treated as approval.

lane Detect
keeps digests and the diff, never your material
receipt immutable, hash-chained
05 / PROVE

Evidence and framework mapping

for: governance, GRC, and the audit you have coming

Evidence packs carrying a content hash, a versioned signer key, a validity window and a revocation URL, crosswalked to the frameworks you report against. A pack whose underlying claim is unsupported refuses to generate rather than attesting less.

lane Prove
verify offline, against a published keyring
maps 12 frameworks, crosswalked
06 / PROTECT

Decide before it happens

for: agents that spend, write or touch customer data

In the action path: signed authority envelopes naming exact targets, versioned signed policy, a human queue behind every deferral, and kill or quarantine containment. Staged shadow, then monitor, then enforce, so nothing blocks on day one.

lane Protect
budget 250ms, a deadline and not a checkpoint
unavailable per-action-class failure mode, receipted
There is no readiness score, deliberatelyYou will not get "this agent is 78% production ready" from us, because that number is made by dropping the requirements that could not be evidenced, so it inflates in exact proportion to how much is unknown. What you get is a range: a floor of what is actually evidenced, a ceiling of what would be evidenced if every unknown resolved in your favour, and a verdict that reads partial for as long as those two differ. Unassessed and not-applicable requirements stay in the denominator. It reads measured only when nothing is unknown, which is the one case where a single number would have meant anything.
Or have us run it against your agents

Once, or continuously. Pick per service.

Every rung above is software you can run yourself. These two are the same capabilities delivered as work we do for you, on your estate. Each comes two ways, and the difference is not the size of the invoice: a one-off answers "where are we right now", and continuous answers "what changed since, and who changed it". Most teams buy one-off first, because it needs no platform decision, and move to continuous once the first report is already out of date. The compliance corpus is sold separately and is further down.

Service 01

Assessment

We point Agenor at your agents and come back with your estate, the shadows in it, what each agent can reach, and findings ranked by blast radius. Read-only throughout.

ONE-OFFA scoped engagement with a date. A report plus a signed pack you keep and can verify without us. If it is all you ever buy, it was still worth the afternoon.
CONTINUOUSThe same detection on a live stream, with incidents that dedupe, recur and carry an owner. Banded by agents in scope.
Service 02

Red team

A battery fired at a candidate agent: prompt injection reaching a real action, tool poisoning, authority escalation, exfiltration paths. Every probe and its outcome recorded so the run can be repeated.

ONE-OFFFired at one agent before it ships, at the gate. Non-production by default, production only on your written go. You get the exact probe list for your scope before we run it.
CONTINUOUSThe battery reruns on a schedule and on change, so a new tool or a widened scope is probed rather than assumed. A pass in March says nothing about the agent in June.
What we will not sell youA number. Not a readiness percentage, not a security score, not a maturity grade out of five. Every engagement above returns counts, a range where something is unknown, and a list of what stayed unknown and why. If a single figure is what you need for a board slide, we are the wrong people, and that is worth finding out in the first conversation rather than the last.
Three lanes, counted separately

What fires on your telemetry, what we sweep, and what we provoke.

Vendors tend to add these together into one impressive number. They are different claims carrying different weight, and internally they are different evidence classes that cannot resolve each other's findings, so we keep them apart here too.

Passive · your telemetry

Detector families reading traces you already emit, with nothing fired at your agents. These run on live customer telemetry today.

blast radiuscrown-jewel reachabilityattack pathsinjection reaching an actionauthority escalationexfiltration pathsmissing human approval

Findings arrive as the security condition with its MITRE technique. The full detector catalogue goes to customers under NDA rather than onto this page.

Manifest · your connected tools

A signed sweep over the tool manifests your agents are actually wired to, re-scanning the latest signature per agent and per MCP server on every pass.

MCP serverstool definitionstool poisoningdefinition driftper-server identity

A finding is held per server, so a clean scan of one never clears a problem on another. Skills, plugins, identities and permissions are tracked as declared and observed change, which is a weaker claim than scanning and is why it is worded differently.

Active · probes we fire

A red-team battery run against a candidate agent, in dev or staging, at the gate before it ships.

prompt injection to actiontool poisoningauthority escalationexfiltration

We keep three separate counts internally: classes named, probes written, and probes wired into a battery that actually runs. They are not the same number and we will give you all three for your scope, in writing, before you buy anything. They are not on this page because the catalogue is the product.

Why these are three columns and not one numberPassive detection and active probing get added together in this market, and the sum flatters whichever half is thinner. Ours are counted separately, and for your scope we will tell you which lane each finding came from, what ran, and what did not run. The catalogue behind the third column is the part we do not publish, for the same reason you would not publish yours.
Why rung three is not rung two repeated

What it was allowed to do, against what it did.

A one-off assessment tells you the state of the estate on the day someone looked. The finding that keeps mattering is the gap between the authority an agent was granted and the behaviour it actually shows, because that gap opens quietly, between audits, as tools get added and scopes get widened. The distance from the diagonal is the finding. Drift shows up as geometry before it shows up as an incident.

Detection tells you an action happened. A record of granted authority tells you whether it was ever supposed to. Only the second survives the question"was that agent allowed to do that?", which is the question an auditor asks and the one a screenshot of a dashboard cannot answer.

MANDATE DIVERGENCE · 9 SPECIMEN AGENTSinside mandatedriftingoutside mandate
Hover a point to read the divergence. Specimen agents, no customer data.
Roots collapsed

Every agent now sits on the line. One key signed both the permission and the report of what happened, so the report can only agree with the permission. This is what a self-attested chart looks like, and it is why the system refuses to start in this configuration.

How it reaches your stack

Hooks and gateways, not another platform to migrate to.

Agenor attaches to what you already run, and how much work that is depends entirely on where your agents live. You will not find a wall of logos here. A logo implies a tested integration, and we only claim that for platforms we have actually built a collector for. Everything else reaches us through the open telemetry convention, which is genuinely most things and is not the same promise. We will tell you which case you are in before you buy, rather than letting you find out on a call.

Collector we built and test

Including LangGraph and Claude Code: a packaged collector with its own test suite, shipped and exercised by us on every build. Ask about yours and we will tell you straight whether it is on this list.

Any OTLP GenAI emitter

Point an OTLP endpoint at us and agent spans arrive unchanged. This works by the open convention rather than by an integration we wrote, which means it usually works and we have not tested your platform specifically. We will say which case you are in.

In the action path

An MCP gateway that needs no change to the agent. This is the layer that offers a decision, where the others offer a read, and it is also the one that can be routed around.

When the gateway is bypassed

An action that skipped the gateway wrote no receipt, so the enforcer's own log is the one place a bypass can never appear. We detect it anyway, from evidence the enforcer does not control and cannot erase. How that works is something we will walk you through directly.

A gateway you already run

If you have an enforcing control in front of your agents, we can take a read-only tap on it and attest what it decided, rather than asking you to replace it. Telemetry tells us what an agent did; this tells us what a control decided about it. We operate none of it, which is the point.

Your payloads stop at the boundary

Decision records from a gateway carry raw material, literal SQL and HTTP bodies included. That is dropped at ingest, before an event exists on our side, rather than filtered later by a reader. A filter downstream leaves the thing upstream still writing it.

Why this is harder than it soundsThe OpenTelemetry GenAI convention is still in development: it moved to its own repository, which has published no releases, and agent-shaped spans are the thinnest part of it. Frameworks emit different vintages, and some emit a rival convention entirely. Normalising across that is most of the work, so treat "it just reads your traces" as a claim to test rather than a feature to tick.
The fourth question, once the first three are answered

Why you should believe the output.

Agenor both enforces and attests, which should void the attestation. It does not, because the two planes hold different keys in different trust roots, and the system refuses to startif anyone collapses them into one. That is also why our output survives leaving us: a pack verifies offline against a published keyring, so your auditor checks it without taking our word for anything, and without us in the room.

Attestation plane

pipeline-verdict-p1

Signs evidence packs. Its verifier is the one your auditor or regulator runs.

SEPARATE
TRUST ROOTS

Enforcement plane

protect-receipt-p1.v2

Signs authority envelopes and decision receipts. Its own custody directory and pin.

Evidence that cannot be quietly edited

Findings, case history, changes and decisions are append-only and hash-chained. The roles are split so the component being recorded is not the component that can rewrite the record.

append-onlyhash-chained receiptsserver-derived digestsseparated roles

The practical consequence: a correction appends a superseding record rather than editing the original, so the history of a finding survives being wrong about it. Nobody at Agenor can quietly revise what was recorded, us included.

In the action path

The Protect lane answers "may this happen?" before it happens, at an MCP gateway that needs no change to the agent. Allow, deny or defer to a human, inside a 250ms deadline.

signed authority envelopesversioned signed policyhuman approval queuekill or quarantineshadow → monitor → enforce

Every policy version carries its enforcement rung inside its own signature, so a version cannot be moved between rungs without the content hash moving. New versions default to shadow, which means nothing you publish starts by blocking traffic.

A separate product · available without the platform

AI agent compliance and governance, as a data product.

This one is not about your estate. It is a corpus, and it is built for agents specifically: the Model Context Protocol specification, agent authority and tool permissions, autonomy and human oversight, alongside the AI frameworks your auditor already names. Requirements extracted one by one with a source anchor each, dispositioned, and crosswalked to the detections that can evidence them. Delivered as a versioned, content-hashed pack and kept current as the frameworks move. You can license it on its own, with no assessment, no telemetry and no platform decision. It is not self-serve yet: scope and terms are agreed per buyer, so the next step is a conversation rather than a checkout.

compliance mappinggovernance controlsframework crosswalkcontrol dispositionsversion driftaudit evidence
Buyer 01

Auditors

Working material for an engagement. Requirement-level detail with a source anchor on every row, so a finding traces back to the clause it came from rather than to our summary of it, and a disposition saying whether evidence for it can be produced from a system at all.

WHY IT HELPSThe awkward rows are still in the set. Requirements that no telemetry can ever evidence are marked as such rather than dropped, which is the difference between a mapping you can stand behind and one that looks complete.
Buyer 02

GRC and governance teams

Which agent governance obligations actually apply to you, split by what a machine can evidence and what a human has to attest, so the work can be assigned. Then kept current: a standard revises quietly and your mapping is wrong from that morning, with nothing to tell you.

WHY IT HELPSFramework versions are watched for drift, so a revision arrives as a diff against the mapping you already hold rather than as a surprise during fieldwork.
Buyer 03

Product and security vendors

A content and intelligence feed you license for your own product. Map your coverage to the frameworks your customers ask about without building and maintaining the extraction pipeline yourself, which is a standing cost rather than a project.

WHY IT HELPSVersioned and content-hashed, so what you shipped against is identifiable later. We are not competing with you for your customer; this is the part of us you can buy and resell inside your own surface.
951
Requirements
across 13 frameworks
43
Pins watched
for silent drift
835
Automatable
evidenced from telemetry
210
Human only
835 + 210 = 1045 controls, from 951 requirements
18
Of those, not feasible
inside the 835, not a third bucket

Crosswalked to one another

MITRE ATLASOWASP LLMOWASP MCPOWASP AISVSEU AI ActNIST AI RMFNIST AI 600-1NIST 800-53NIST 800-218AAIUC-1ATF ConformanceATF Crosswalks

Twelve carry crosswalk records, proposing a canonical reference so one control can answer several frameworks at once. The thirteenth, the MCP specification, is extracted but not yet crosswalked. A crosswalk is a proposed equivalence with a confidence and a reason attached, not a detection claim.

Agent-specific requirements extracted

MCP Specification 506MITRE ATLAS 224AIUC-1 53NIST AI 600-1 49ATF Conformance 25

The five deepest of the thirteen, 857 of the 951, extracted requirement by requirement with a source anchor each. The twelve beside it is the mapping surface; these five are how deep the extraction has gone. They are different numbers because they measure different things.

The bound, stated up frontThese counts describe our corpus, not a promise about what will be evidenced in your environment on day one. What a mapping is worth depends on what your estate actually emits, and we would rather you read that here than discover it during an audit.

Scope it yourself. We run the first one with you.

You can scope the engagement yourself in a few minutes, including for an agent you can only reach over the network and cannot instrument, and the credential we ask for is scoped to that one job. Getting from there to a first finding is still a run we do with you rather than one you kick off alone; the hands-off path is being built and is not finished.

Access is approved by us rather than open signup, which is deliberate while the estate is small, and we would rather say that than dress a waiting list as a product tour. Prefer it done for you? That is the assessment above. Ongoing work is banded by agents in scope andnever billed per finding, because a grader paid by the finding is not a grader.