AI agent action assurance

Before your agent acts, prove it knows where to stop.

A bounded rehearsal for agents that send, change, pay, delete, or publish. We test whether approval, permissions, retry protection, audit evidence, shutdown, and recovery work when the workflow is stressed.

No production credentials. Staging or a sanitised workflow first. Written delivery available.
Illustrative control replay

Refund agent · staging

AA-001
01

Request interpreted

Observed
02

Human approval

Passed
03

Arguments changed after approval

Failed
04

Action trace preserved

Partial
01No credentials by email
02Staging before production
03Replay evidence, not opinion
04No certification claims
Free no-data self-check

Map what the agent can do—then test what can stop it.

No workflow file, credentials, prompt, or customer data leaves your browser.

1. Which actions can this agent perform?
2. How are those actions controlled?
01

Must a person independently approve every high-impact action?

02

Are the exact approved arguments bound to the action that executes?

03

Does the agent use a restricted service identity rather than a broad user account?

04

Can retries occur without sending, charging, deleting, or publishing twice?

05

Can you reconstruct the input, decision, tool arguments, approver, and outcome?

06

Is there a tested emergency stop and a defined recovery path?

07

Is external content prevented from silently changing the agent's authority?

The action boundary rehearsal

Test the moment an answer becomes an external action.

Generic model evaluation asks whether an answer is good. This rehearsal asks whether the agent can be induced to perform an unauthorised, duplicated, altered, or unrecoverable action.

01

Authority boundary

Attempt an action outside the written purpose and permitted tool set.

02

Approval binding

Check that the reviewed action and arguments are exactly what executes.

03

Least privilege

Verify that the agent identity cannot reach unrelated records, tools, or secrets.

04

Duplicate action

Replay timeout and retry paths without double-send, double-charge, or double-delete.

05

Instruction isolation

Introduce untrusted content that tries to alter goals, permissions, or approval rules.

06

Stop and rollback

Exercise the emergency stop, failure owner, recovery path, and preserved evidence.

Two professionals reviewing an AI workflow at a human approval point before an external action
Human review pointConcept illustration
Human judgement stays in the loop

The approval must travel with the action.

A click is not enough. We check whether the approved destination, amount, record, and action remain bound to what actually executes—and whether the team can stop and recover when they do not.

01

Bind the reviewed arguments to execution.

02

Prevent retries from repeating the action.

03

Preserve a stop owner and usable evidence.

A deliberately narrow buyer

Built first for AI automation agencies.

Agencies already deploy agent workflows for clients. A reusable, client-facing evidence pack helps them explain what the agent can do, which actions require approval, what happens on failure, and what was actually tested.

Agency valueClose with evidence

Add a concrete action-control workstream to a client deployment.

Delivery valueCatch expensive edges

Find approval bypass, retry duplication, privilege, and rollback gaps before handover.

Scale valueRepeat the proof

Move from one workflow rehearsal toward a white-label partner programme.

Inspect the full synthetic evidence pack ↗
Founding technical pilot

One workflow. Six replay families. Evidence your client can read.

USD$650fixed founding pilot
01

Written action boundary

Purpose, tools, permissions, high-impact actions, approvers, exclusions, and stop authority.

02

Replay evidence

Pass, partial, or fail results tied to the exact test, observed trace, and limitation.

03

Remediation plan

Prioritised fixes plus one written clarification round and one bounded re-check.

5 business days after usable staging evidence · no production access by default · not a penetration test or certification

Request written pilot scope ↗
For builders who let agents act

Do not sell “safe AI.” Show the exact boundary you tested.

Start with the self-check. If the action deserves evidence, request a scope without sharing credentials or client data.

Run the action self-check ↗