The question arrives ad hoc
There is no standard form for agent behavior. One technical reviewer asks it live on a call, and “it won't do that” is not an accepted answer.
For teams selling AI agents to enterprise buyers
Enterprise reviewers ask what your agent does, not what your policies say. We keep the question bank, we run your real agent against it in a sandbox, and every failure lands on your desk. Privately.
Refunds, cancellations, order edits — whatever it can touch is what gets asked about.
01 We bring the questions
02 Real tool calls, real final state
03 Failures stay private
04 Reruns on every release
The gap
Access controls, encryption, change management — audited. Whether your agent can be talked into a second refund when your payments API times out — never tested. That is the question a reviewer asks, and it is not on any questionnaire.
There is no standard form for agent behavior. One technical reviewer asks it live on a call, and “it won't do that” is not an accepted answer.
Your evals cover what your team thought to test. The reviewer asks what they thought to ask. The overlap gets discovered mid-deal.
Weeks of prose answers into a spreadsheet — per buyer, per version — while the contract waits.
The dry run
A fixed-scope engagement that finds the failures before a buyer does — across the layers an agent is actually made of: context, tools, authority, harness. No buyer involvement, no publication. The results are yours and nobody else's.
A paid engagement. The fee is fixed in writing before anything starts, and it is the same whatever we find — we are not paid for good news. Three founding slots, credited toward whatever comes next. The runner is yours to keep; the question bank stays with us and grows with every engagement. Results go to you and to nobody else.
Book a scoping callThe part that stays on
A model swap can quietly undo a behavior you fixed in March. A dry run is a snapshot; your agent is a moving target. The subscription keeps the two honest with each other.
A failed critical scenario blocks the deploy like any other test — and each finding ships with an ordered remediation your agent can work from.
RUNS IN YOUR PIPELINENew attack classes and buyer-derived scenarios are added continuously — and run against your agent first.
NEW SCENARIOS EVERY MONTHA version-by-version log of behavior across releases. Evidence you can show a buyer, and cannot recreate later.
EVERY VERSION, KEPTFree and open source
Agent-native from the first step: download the skill file for your harness and drop it in. Your agent reads itself and writes down what it can touch — the tools it calls, what it can write to, the actions that move money, the customer data it can reach. It comes back as one standard declared profile. Ten minutes, your machine, no account.
A declared profile is not evidence. It is your agent's own account of itself, and your agent will miss things — the tool it forgot, the path it never mentions, the action it does not think of as spending money. Finding what it did not mention is the job. That is the dry run: cases you did not pick, actions executed in a sandbox we control, and findings that fail closed.
The deliverable
Three synthetic specimens — each says so in its own metadata — covering every conclusion the engine can reach. One fails on severity, one passes and still hits the evidence ceiling, one refuses to conclude at all. The fourth outcome cannot be produced, and that row is the point.
Every finding traces to a preserved trace, and the whole file recomputes from raw evidence. Yours arrives bound to one frozen version, under a digest that changes if a single character does, in three forms — a page you can read, Markdown, and the raw evidence. Read it before you buy anything.
Get yours privatelyMethod
Not a penetration test. A red team asks whether your agent can be attacked. We ask whether it follows the policy you sold when the tools fail — and whether the money actually moved. The full method and our commercial independence rules are published: how conclusions are reached · what we will not be paid for.
Some buyers will want evidence directly. That is a separate, later engagement — a buyer-visible Factfile under the published method, at the same fee whether the conclusion helps you or not. Behavioral review is where the market is heading: 26 of the 51 requirements AIUC-1 publishes test behavior rather than paperwork (Safety 12, Security 10, Reliability 4). The fastest route through any such review is having already failed it in private.
Talk it through with usStart here
Refunds, cancellations, order edits — whatever your agent is allowed to touch. We reply with the questions an enterprise reviewer will ask about each one, and which we would test first. Twenty minutes to walk through it, no deck. If a dry run is not worth it for you yet, we will say so.
What a dry run needs, so nothing is a surprise later