Simam AI Lab Our applied AI research division is now open. Visit the lab
Simam AI Lab product · invite-only beta

A control point between AI agents and the tools they call.

AI agents now run shell commands, push code, query databases and call APIs on their own behalf. Coding assistants do it continuously; MCP-connected copilots and autonomous workflows do it without anyone watching. In most organisations there is no point in that chain where a dangerous action can be stopped, no approval step, and no record afterwards of what was attempted.

The Agent Harness is that missing point. Every tool call an agent makes passes through a policy pipeline before it runs, and every decision — allowed, held, or blocked — is written to a tamper-evident ledger with a downloadable receipt. It is the control and audit layer that endpoint detection gave security teams for laptops, applied to agents.

It is a working product in invite-only beta, built by Simam AI Lab. This page describes what it does, the decisions behind it, and what it does not do.

Built — September to October 2026  ·  Pipeline — six stages, fail-closed  ·  Threat mapping — OWASP Top 10 for LLM Apps, OWASP Agentic AI, MITRE ATLAS  ·  Tests — 980 across six suites, counted 6 October 2026  ·  Governed operations — 19, across three adapters

The pipeline

Every tool call is checked before it runs, not after.

Six stages, in order. A call has to survive all of them. If any stage cannot reach a decision the call is refused rather than allowed, because a governance layer that fails open is decoration.

Kill switch

One control that stops every agent action across the estate. The thing you reach for when something is going wrong and you do not yet know what.

Per-agent allow-list

Which tools this particular agent may use at all. An agent that never needs a shell does not get one, so a prompt injection cannot talk it into using one.

Policy rules

Organisation rules evaluated against the specific call — the command, its arguments, the target, the context it arrived in.

Governed operations

Nineteen named dangerous actions, matched by pattern across three adapters: seven on GitHub (repository creation, merging a pull request, managing secrets, writing workflows, changing branch protection, force-pushing, making a repository public), five on SQL (DROP, TRUNCATE, GRANT, and DELETE or UPDATE with no WHERE clause) and seven on the shell (rm -rf on a root or home directory, disk wipe, piping a download into a shell, reading credentials or private keys, pushing to main, force-pushing, and tampering with the harness’s own hook or config). Each is tagged with its OWASP Top 10 for LLM Applications entry, its OWASP Agentic AI threat and its MITRE ATLAS technique, so a finding arrives already mapped to the framework a security team reports against. Five of the shell operations are denied outright; the two pushes are held for a person. A twentieth rule is a guard rather than an operation: any shell call over 64 KB is held rather than partially inspected, so a long command cannot be used to push its dangerous half out of view.

Budget limits

Ceilings on how much an agent may spend or consume, so a loop is expensive for a minute rather than for a weekend.

Human approval

What is left: the call is held, a person sees exactly what was attempted, and nothing happens until they decide.

The brief

The problem is not that agents are dangerous. It is that nobody is watching.

Three ways of stating the same gap, because which one lands depends on who is reading.

There is no control point

An agent calling a tool goes straight from model output to execution. There is no place to stand between the two, which means there is nowhere to put a rule even if you had written one.

There is no approval step

Teams respond by granting agents either too much access or too little. Too much is the incident; too little is the agent that cannot do its job and gets switched off. The missing middle is "ask first for these specific things".

There is no record

After an incident the question is always what the agent actually did. Chat transcripts are not an audit trail: they show what the model said, not which calls reached a system, in what order, and which were refused.

Architecture

Three decisions shaped everything else.

These are the choices that are expensive to reverse, so they are the ones worth describing.

Fail closed, everywhere

If the gateway is unreachable, a rule cannot be evaluated or a dependency is down, the call does not run. This is the decision that costs the most in day-to-day friction and it is not negotiable: a control that disappears under load is worse than none, because people plan around it being there.

The ledger is evidence, not logging

Every decision is written to a tamper-evident Evidence Ledger with a downloadable receipt. Logs are for debugging and nobody trusts them in an argument. The ledger is built so that "this call was blocked at this time under this rule" survives being questioned months later by someone who was not there.

Governance has three modes, not two

Off, monitor, enforce. Monitor is the one that matters: it records what would have been blocked without blocking it, so a team can see the real shape of their agent traffic before switching enforcement on. Every tool of this kind that goes straight to enforce gets turned off in week one.

What is in it

The parts you actually touch.

The pipeline is the argument; these are the surfaces around it.

MCP proxy

Governs tools reached over the Model Context Protocol, sitting between the copilot and the servers it calls.

Posture scanning

Looks at how agents are configured and reports where the exposure is, before an incident rather than after.

Approvals queue

Where held calls wait. Shows what was attempted, by which agent, with the full arguments, and takes a decision.

Evidence Ledger and reports

The decision record, with receipts that can be downloaded and handed to someone who needs to verify a claim.

Role-based access

Who may change a policy, who may approve a held call, and who may only look. These are different jobs and the product treats them as such.

Invite-only beta

Anyone may create an account; a new account reaches an awaiting-approval screen with no access to data until an administrator approves it.

Next step, in progress

Governing a coding agent, live.

The current work is putting Claude Code itself behind the harness. Every Bash, Write and Read the coding agent attempts is checked before it runs; dangerous ones are blocked outright or held for approval.

Why this case first

A coding assistant is the agent with the most dangerous tool access in most organisations today, and the one people are most reluctant to restrict because it is useful. If governance can be made tolerable here, the easier cases follow.

What it proves

That the overhead is survivable. A control that adds a visible pause to every file read will be removed by the first engineer it annoys, so the test is whether the agent stays usable.

The honest risk

Rules are pattern-based. A coding agent is extremely good at producing commands that do not look like the pattern, which is the limitation described below rather than one discovered later.

Problems worth writing down

Six things that went wrong, and what changed.

Written down because a build account with no failures in it is not an account. Three of these were caught by review of the security product itself, before release.

Silent disk corruption that every test passed through

An unknown process overwrote around 140 files on the build drive with the word SKIP repeated — source files, installed dependencies, and one object inside the git database. It preserved every file’s original size and timestamp, so git reported a clean tree and Python kept running from stale bytecode caches. The full test suite stayed green against destroyed source. It was found by reading file contents rather than trusting any tool. Recovery: restore from git, rebuild the damaged git object from its two parents — the hash matched — and reinstall dependencies. What changed: work is pushed off-machine promptly, tests run with bytecode caching disabled, and the build refuses to ship any file beginning with the filler pattern.

A redirect that pointed at itself

Moving the console from / to /app required /app to forward to /app/. Firebase Hosting ignores trailing slashes when matching, so the redirect matched its own target and the console became unreachable. What changed: a rewrite instead of a redirect, and a configuration test that fails if a redirect for /app ever reappears.

Login that succeeded and then failed

On the first hosted deploy, sign-in returned 200 and every subsequent request returned 401. Firebase Hosting forwards only one cookie name, __session, to the backend. What changed: the session cookie was renamed, and smoke tests now perform a full login-to-profile round trip rather than asserting on the login status code — which had been passing throughout.

A live encryption key in a planning document

A real encryption key was pasted into a design document and committed. What changed: the key was retired, a replacement generated in Secret Manager, the document scrubbed from git history, and a fresh database stood up under the new key. Rotating is the easy half; assuming the old key is public from the moment it is committed is the half people skip.

Rules written for one agent fired on another

The final whole-branch review found that SQL and GitHub rules written for MCP tools were also matching Claude Code’s own calls. A commit message reading “Update README and set version” matched the pattern for UPDATE … SET with no WHERE clause, so an ordinary commit would have been held for approval in the middle of a demonstration. What changed: the older unscoped rules no longer apply to Claude Code tools, which have their own scoped shell rules instead. The general lesson is that a rule inherits the assumptions of the agent it was written for, and those assumptions do not survive being pointed at a different one.

Review found security bugs in the security product

Independent review of each task caught three: a shell-rule regular expression that took 26 seconds on a crafted thousand-word command, which is a denial-of-service route into the thing meant to protect you — rewritten with bounded matching, now around 0.05 seconds on 100 KB; a hook that followed HTTP redirects and would have sent the agent’s key to another host, now refused outright with the call blocked; and a 2 KB command preview that could be padded so a dangerous tail fell outside it, so commands are now always inspected in full.

Limits

What it does not do.

Stated plainly, because a governance product that oversells its coverage is actively harmful — people stop looking at the things it does not watch.

It does not sandbox the machine

It decides whether a tool call may run. It does not contain what that call does once it is allowed. It is a decision point, not an isolation boundary.

Rules are pattern-based

Obfuscated or run-time-assembled commands can slip past. This is inherent to matching on patterns and is the reason AI-assisted risk scoring is the next piece of work rather than a nice-to-have.

Implicit context is not always caught

A bare git push while already on main is not held, because the branch is implicit in the shell state rather than present in the command.

Two paths, governed separately

MCP tools are governed by the MCP proxy; Claude Code goes through its own hook. They are not one mechanism and should not be assumed to behave identically.

One shared workspace

The hosted beta is a single workspace. Per-company tenancy is not built yet, which is why the beta is invite-only and approved by hand.

No risk scoring yet

Decisions are rules and patterns today. There is no model judging how dangerous a given call looks.

Roadmap

What is next.

In the order it is being built.

AI-assisted risk scoring

Judging how dangerous a call looks, to catch what patterns miss.

Plain-English policy authoring

Writing a rule without writing a rule, so the person who owns the risk can write the policy.

An incident explainer

Turning a sequence of ledger entries into an account a human can read.

Per-company workspaces

Tenancy, which is what the beta is gated on.

A public demonstration

A video of Claude Code being governed live, so the claim can be checked rather than taken on trust.

Case study decision record

The commercial case, in one view.

A concise record of what the project set out to prove, the evidence available today, and the next responsible investment step.

Business challenge

Organisations are adopting AI agents that execute real actions — shell commands, code pushes, database queries, API calls — faster than they are adopting any way to control or evidence them. Security teams are being asked to sign off on systems they cannot see into.

Why the project mattered

It answers the objection that Simam Digital builds interfaces rather than AI systems. A governance gateway is an infrastructure product with a threat model, a fail-closed posture and an evidence standard — and the engineering account below includes security bugs found in it by review, which is the part a capability deck cannot fake.

What Simam Digital designed and built

A governance gateway for AI agents: a six-stage policy pipeline, governed operations mapped to OWASP Top 10 for LLM Applications, OWASP Agentic AI and MITRE ATLAS, a tamper-evident Evidence Ledger with downloadable receipts, an MCP proxy, posture scanning, an approvals queue, reporting and role-based access.

Important product decisions

Fail closed rather than fail open. Evidence rather than logging. Three governance modes rather than two, so monitor mode can show a team the real shape of their agent traffic before enforcement is switched on.

Measured evidence

980 automated tests across six suites as of 6 October 2026, all passing: 767 on the Python gateway, 103 on the Claude Code hook, 76 on the React console with 6 more on its build scripts, 20 on the landing page and 8 on hosting configuration. Independent per-task review found and closed three security defects before release, including a regular expression that took 26 seconds on a crafted input and now takes around 0.05 seconds on 100 KB.

Credible outcome

A working invite-only beta that can be opened and inspected, with a public landing page demonstrating an agent attack being stopped. The current work governs Claude Code itself — every Bash, Write and Read checked before it runs.

Recommended next engagement

An Idea Validation Sprint at £1,500 to map where agents already hold tool access in your estate and what a policy would need to say. Prototype work on a governed integration starts from £5,000.