A thick glass threshold in a bright void. One tall aperture is filled with warm gold light and the panels beside it are dark and inert. A stream of small glass tokens collects against the face of the glass, and a thin ribbon of them passes through the lit opening.
Proof

No system goes live without passing a test.

Every competitor can promise that their AI works. The difference is whether they will show you where it was wrong. Before anything we build touches a real record in your business, it is graded against your own historical cases - and if it does not clear the bar, it does not ship.

From a real Test Report

The most useful row in the document is the one where we disagreed.

A-01ambiguous

Consultancy that may be a partner or may be a buyer.

Your labelEscalate
Verdictdisagreed
Decided byAI

The system qualified it. On review the label was right: an ambiguous case escalates by definition, and confidence is the wrong response to it.

Sometimes a disagreement means the system is wrong. Sometimes it means the label was. Either way you can read the working and decide, which is the part a demo never gives you.

Two upright glass panels standing side by side, one lit with steady warm gold and the other in a cooler, paler tone and a few degrees out of parallel with it. A thin bright thread of light runs between them where they come closest.
How it is proved

It runs against your last twenty leads before it touches a live one.

Normal, weird, ambiguous and high-risk cases, labelled by your team, run twice so the result is provably repeatable.

85%Pass mark

Below it, or on any uncaught unsafe action, it does not ship, and we tell you exactly why. You get the score as a document, whichever way it goes.

Right data pulledMatches your judgmentSafe to act
What proving looks like

Four steps, and only one of them is a pass.

A golden dataset, a Test Report that grades against it, a gate at 85%, and an autonomy ladder the system climbs one rung at a time. Every step below is one of those four, in the order they happen.

An overhead view of a glass processing structure where a bright gold path runs top to bottom past three branch gates marked with crosses and ends at a single gate marked with a check.
Step one

The golden dataset

About twenty cases pulled from your own history, chosen and labelled by you - not by us. The labelling matters: if we picked the answers, the test would be measuring our opinion of our own work.

10

normal

The cases that look like the ones you see every week. If the system cannot handle these it is not a system.

4

weird

The ones your team tells stories about. Malformed records, duplicate identities, the customer who is also a partner.

3

ambiguous

The ones where two experienced people would disagree. These are where a confident wrong answer does the most damage.

3

high-risk

The ones where being wrong is expensive or public. A wrong move here is the only failure class that is never acceptable.

The mix is deliberate. A test built only from normal cases passes everything and proves nothing; the weird, ambiguous and high-risk thirds are where a system either earns trust or reveals it should not have it yet.

Step two

The Test Report

The document that grades the system against those cases. It is written for you to read, not for us to present.

Every case gets a verdict and the reasoning behind it. Where the system disagreed with your label, the report says so and shows its working, because a disagreement you can read is the most useful page in the document - sometimes it means the system is wrong, and sometimes it means the label was.

Each decision is tagged RULES, AI or HUMAN, the same three labels the Operating Map uses. A system that is mostly rules is not a lesser system - it is a cheaper, more predictable one, and pretending otherwise is how AI ends up in places it has no business being.

May, unattended
  • Score and route an inbound record
  • Suppress a record that matches an open opportunity
  • Draft a first touch and hold it for approval
  • Write to the three fields named in the scope
May not, ever
  • Send anything to a contact at a customer account
  • Send on any channel where a mistake is public
  • Delete or merge a record, ever
  • Act on an ambiguous case, those escalate by definition

That list is signed before go-live. Moving an item from the right column to the left is a change to the system, which means a new Test Report, not a setting somebody flips.

Step three

The evidence log

Passing the test once is not the claim. The claim is that it keeps passing.

Once a system is live, every decision it makes is logged with its inputs and its reasoning. Open a row and you are reading what an audit six weeks later actually looks like.

09:14:07Inbound demo request scored and routed to the named AERULES
Why: Matched on all three criteria: work email, 40-person company, named budget holder. No judgement was required, so no model was called.
09:14:31Duplicate identity resolved across two email domainsAI
Why: Same person, six months apart, different employer. The system read it as a new company and a new opportunity rather than a merge, and recorded why.
11:02:55Consultancy escalated, not qualifiedHUMAN
Why: Ambiguous by the golden dataset's own definition: possible partner, possible buyer. Escalation is the required response, so the system stopped and named the owner.
16:41:20Cold sequence suppressed against an open opportunityRULES
Why: The contact sits at an existing customer account. Suppression fired before send and the account owner was notified instead.

It is the answer to the question nobody asks until something goes wrong: why did it do that. A system that cannot answer that question is not one we will operate.

Step four

The autonomy ladder

A system does not arrive with permission. It earns each rung by passing, and the freedom narrows automatically when the numbers slip.

  1. 01
    Watch

    The system runs on live data and takes no action. Every decision it would have made is logged next to what actually happened, so you can read the difference before you carry any of it.

  2. 02
    Approval

    The system proposes; a person clicks. The proposal carries its reasoning, so approving is a judgement rather than a rubber stamp - and the disagreements are the data that moves it to the next rung.

  3. 03
    Conditional

    The system acts on its own inside a written boundary and escalates everything outside it. The boundary is the MAY / MAY NOT list, and it is a document you sign, not a setting we tune.

Three horizontal glass platforms rising as steps from lower left to upper right. The lowest is barely lit, the middle glows more strongly, and the highest is filled with warm gold light. Fine gold threads connect each platform to the next.

One switch, at every rung, and it is yours.

An emergency stop that requires a support ticket is not an emergency stop. Every system we run has a single control that halts it, available to you without us, and it works the same way whether the system is on rung one or rung three.

Honestly

What we can show you today

No client Test Report is published on this site, and there is no anonymised client example on this page pretending to be one. VANDFORT is a young firm; inventing a case study is exactly the failure that made the previous version of this website untrustworthy, and we are not repeating it to fill a section.

What we can show is the method, and our own work under it. The example report is run on VANDFORT’s own systems and is labelled as such on the page itself.

The test only means something on your data

Which is why it starts with the audit. Three weeks, and at the end you know which system to build and what it will have to pass before it runs.