
No system goes live without passing a test.
Every competitor can promise that their AI works. The difference is whether they will show you where it was wrong. Before anything we build touches a real record in your business, it is graded against your own historical cases - and if it does not clear the bar, it does not ship.
The most useful row in the document is the one where we disagreed.
Consultancy that may be a partner or may be a buyer.
The system qualified it. On review the label was right: an ambiguous case escalates by definition, and confidence is the wrong response to it.
Sometimes a disagreement means the system is wrong. Sometimes it means the label was. Either way you can read the working and decide, which is the part a demo never gives you.

It runs against your last twenty leads before it touches a live one.
Normal, weird, ambiguous and high-risk cases, labelled by your team, run twice so the result is provably repeatable.
Below it, or on any uncaught unsafe action, it does not ship, and we tell you exactly why. You get the score as a document, whichever way it goes.
Four steps, and only one of them is a pass.
A golden dataset, a Test Report that grades against it, a gate at 85%, and an autonomy ladder the system climbs one rung at a time. Every step below is one of those four, in the order they happen.

The golden dataset
About twenty cases pulled from your own history, chosen and labelled by you - not by us. The labelling matters: if we picked the answers, the test would be measuring our opinion of our own work.
normal
The cases that look like the ones you see every week. If the system cannot handle these it is not a system.
weird
The ones your team tells stories about. Malformed records, duplicate identities, the customer who is also a partner.
ambiguous
The ones where two experienced people would disagree. These are where a confident wrong answer does the most damage.
high-risk
The ones where being wrong is expensive or public. A wrong move here is the only failure class that is never acceptable.
The mix is deliberate. A test built only from normal cases passes everything and proves nothing; the weird, ambiguous and high-risk thirds are where a system either earns trust or reveals it should not have it yet.

The Test Report
The document that grades the system against those cases. It is written for you to read, not for us to present.
Every case gets a verdict and the reasoning behind it. Where the system disagreed with your label, the report says so and shows its working, because a disagreement you can read is the most useful page in the document - sometimes it means the system is wrong, and sometimes it means the label was.
Each decision is tagged RULES, AI or HUMAN, the same three labels the Operating Map uses. A system that is mostly rules is not a lesser system - it is a cheaper, more predictable one, and pretending otherwise is how AI ends up in places it has no business being.
- Score and route an inbound record
- Suppress a record that matches an open opportunity
- Draft a first touch and hold it for approval
- Write to the three fields named in the scope
- Send anything to a contact at a customer account
- Send on any channel where a mistake is public
- Delete or merge a record, ever
- Act on an ambiguous case, those escalate by definition
That list is signed before go-live. Moving an item from the right column to the left is a change to the system, which means a new Test Report, not a setting somebody flips.
The evidence log
Passing the test once is not the claim. The claim is that it keeps passing.
Once a system is live, every decision it makes is logged with its inputs and its reasoning. Open a row and you are reading what an audit six weeks later actually looks like.
09:14:07Inbound demo request scored and routed to the named AERULES
09:14:31Duplicate identity resolved across two email domainsAI
11:02:55Consultancy escalated, not qualifiedHUMAN
16:41:20Cold sequence suppressed against an open opportunityRULES
It is the answer to the question nobody asks until something goes wrong: why did it do that. A system that cannot answer that question is not one we will operate.
The autonomy ladder
A system does not arrive with permission. It earns each rung by passing, and the freedom narrows automatically when the numbers slip.
- 01Watch
The system runs on live data and takes no action. Every decision it would have made is logged next to what actually happened, so you can read the difference before you carry any of it.
- 02Approval
The system proposes; a person clicks. The proposal carries its reasoning, so approving is a judgement rather than a rubber stamp - and the disagreements are the data that moves it to the next rung.
- 03Conditional
The system acts on its own inside a written boundary and escalates everything outside it. The boundary is the MAY / MAY NOT list, and it is a document you sign, not a setting we tune.

One switch, at every rung, and it is yours.
An emergency stop that requires a support ticket is not an emergency stop. Every system we run has a single control that halts it, available to you without us, and it works the same way whether the system is on rung one or rung three.

What we can show you today
No client Test Report is published on this site, and there is no anonymised client example on this page pretending to be one. VANDFORT is a young firm; inventing a case study is exactly the failure that made the previous version of this website untrustworthy, and we are not repeating it to fill a section.
What we can show is the method, and our own work under it. The example report is run on VANDFORT’s own systems and is labelled as such on the page itself.
The test only means something on your data
Which is why it starts with the audit. Three weeks, and at the end you know which system to build and what it will have to pass before it runs.

