For teams selling AI agents to enterprise buyers

Find out what your agent does before your buyer asks.

Enterprise reviewers ask what your agent does, not what your policies say. We keep the question bank, we run your real agent against it in a sandbox, and every failure lands on your desk. Privately.

Refunds, cancellations, order edits — whatever it can touch is what gets asked about.

01 We bring the questions

02 Real tool calls, real final state

03 Failures stay private

04 Reruns on every release

The gap

Your SOC 2 covers your company. It says nothing about your agent.

Access controls, encryption, change management — audited. Whether your agent can be talked into a second refund when your payments API times out — never tested. That is the question a reviewer asks, and it is not on any questionnaire.

01

The question arrives ad hoc

There is no standard form for agent behavior. One technical reviewer asks it live on a call, and “it won't do that” is not an accepted answer.

02

The evidence does not exist yet

Your evals cover what your team thought to test. The reviewer asks what they thought to ask. The overlap gets discovered mid-deal.

03

The deal absorbs the delay

Weeks of prose answers into a spreadsheet — per buyer, per version — while the contract waits.

AreaYour SOC 2 answersThe reviewer asks anyway
AccessWho can log into your systemsWhat the agent can do once it is in them
Change managementHow code reaches productionWhat changed in behavior since the version they approved
Data handlingWhere customer data is storedWhether the agent reads one customer's order to another
Incident responseHow you respond after a breachWhether a refund fires twice when the API times out mid-call
02

The dry run

Fail privately first.

A fixed-scope engagement that finds the failures before a buyer does — across the layers an agent is actually made of: context, tools, authority, harness. No buyer involvement, no publication. The results are yours and nobody else's.

01We map your agent's actionsRefunds, cancellations, edits, escalations — each mapped to the questions buyers ask about that authority.
02The question bank becomes executed testsDrawn from real reviews and known failure classes. Written down before we run, scored by deterministic checks.
03Your frozen version runs in a sandboxReal tool calls, tool failures injected, final state checked against the ledger — not the transcript.
04Every failure lands on your deskEach finding traces to a preserved trace. You fix; one rerun is included.
TWO WEEKS, FIXED FEE, PRIVATE

A paid engagement. The fee is fixed in writing before anything starts, and it is the same whatever we find — we are not paid for good news. Three founding slots, credited toward whatever comes next. The runner is yours to keep; the question bank stays with us and grows with every engagement. Results go to you and to nobody else.

Book a scoping call
03

The part that stays on

You ship again next Tuesday.

A model swap can quietly undo a behavior you fixed in March. A dry run is a snapshot; your agent is a moving target. The subscription keeps the two honest with each other.

01 / CI

Reruns gate every release

A failed critical scenario blocks the deploy like any other test — and each finding ships with an ordered remediation your agent can work from.

RUNS IN YOUR PIPELINE
02 / FEED

The question bank keeps growing

New attack classes and buyer-derived scenarios are added continuously — and run against your agent first.

NEW SCENARIOS EVERY MONTH
03 / HISTORY

Your record compounds

A version-by-version log of behavior across releases. Evidence you can show a buyer, and cannot recreate later.

EVERY VERSION, KEPT
04

Free and open source

Your agent writes its own action list.

Agent-native from the first step: download the skill file for your harness and drop it in. Your agent reads itself and writes down what it can touch — the tools it calls, what it can write to, the actions that move money, the customer data it can reach. It comes back as one standard declared profile. Ten minutes, your machine, no account.

Download the skill file See an example declared profile

factfile-intake.skill.md

ASend the profile to usWe read it and reply with the two or three questions a reviewer will ask about your spec — the specific ones, not a checklist — and the link to book if you want them answered with evidence.
BOr bring nothing and book the callTwenty minutes, no preparation. We write the action list with you, live on the call. Skipping the skill file costs you nothing.

Also free

Five public probes. Runs locally. Nothing is sent to us.

A declared profile is not evidence. It is your agent's own account of itself, and your agent will miss things — the tool it forgot, the path it never mentions, the action it does not think of as spending money. Finding what it did not mention is the job. That is the dry run: cases you did not pick, actions executed in a sandbox we control, and findings that fail closed.

05

The deliverable

This is what lands on your desk.

Three synthetic specimens — each says so in its own metadata — covering every conclusion the engine can reach. One fails on severity, one passes and still hits the evidence ceiling, one refuses to conclude at all. The fourth outcome cannot be produced, and that row is the point.

ALL SYNTHETIC · ALL SAY SO

Every finding traces to a preserved trace, and the whole file recomputes from raw evidence. Yours arrives bound to one frozen version, under a digest that changes if a single character does, in three forms — a page you can read, Markdown, and the raw evidence. Read it before you buy anything.

Get yours privately

Method

The agent never scores itself.

HOW EVIDENCE IS MADE
  • Deterministic code computes every result
  • Actions execute in a sandbox we control
  • Hash-chained log; findings recompute from raw evidence
  • A check that cannot run fails closed
WHAT WE ARE NOT
  • Not affiliated with AIUC, Schellman, or Lloyd's
  • Not an auditor or certification body
  • Not selling a pass

Not a penetration test. A red team asks whether your agent can be attacked. We ask whether it follows the policy you sold when the tools fail — and whether the money actually moved. The full method and our commercial independence rules are published: how conclusions are reached · what we will not be paid for.

WHEN A BUYER WANTS TO SEE IT

Some buyers will want evidence directly. That is a separate, later engagement — a buyer-visible Factfile under the published method, at the same fee whether the conclusion helps you or not. Behavioral review is where the market is heading: 26 of the 51 requirements AIUC-1 publishes test behavior rather than paperwork (Safety 12, Security 10, Reliability 4). The fastest route through any such review is having already failed it in private.

Talk it through with us

Start here

Send us your agent's action list.

Refunds, cancellations, order edits — whatever your agent is allowed to touch. We reply with the questions an enterprise reviewer will ask about each one, and which we would test first. Twenty minutes to walk through it, no deck. If a dry run is not worth it for you yet, we will say so.

What a dry run needs, so nothing is a surprise later

  • An agent that takes real actions, not just answers questions
  • A version you can freeze for two weeks
  • A sandbox or endpoint we can reach