Decision evidence for consequential AI
AI can make a million decisions in a second.
Notary accounts for every one.
Notary checks every reason your AI gives against the evidence it had, and keeps the result under your own reference.
Facts, never verdicts.
None of these reasons was accounted for. That's why nobody caught the bad one.
Four industries. In each one, three of the reasons would have held up and one would not, and nothing in the loop could tell them apart.
The good reasons were never checked either. That is the whole problem.
It runs both ways. Granted on a basis that doesn't exist, money leaves quietly. Refused on one, it comes back as an appeal.
You already owe this document. We produce it.
These documents were always required, and someone specific still turns up asking for each one. What changed is that the person who used to write them is gone.
What has to exist
A file detailed enough to reconstruct the decision: the notes, the working, and the dates they happened.
Pulled at random in a market conduct exam, and by every bad-faith suit that follows.
What your AI produces today
A decision log and a generated summary, with nothing tying either back to the file.
What Notary supplies
Every reason in the determination, resolved to the exact evidence in the file, or marked as resting on nothing.
What has to exist
The determination, the specific criteria applied to it, and the name of the clinician who reviewed it.
Produced on internal appeal, then again for an independent external reviewer.
What your AI produces today
A determination and a rationale the model wrote itself.
What Notary supplies
Each criterion the determination relied on, resolved to the policy text it came from, and a record of who reviewed it.
What has to exist
A statement of the reasons actually considered, not the ones that sound right afterwards.
The applicant can demand it by right, and examiners sample it every cycle.
What your AI produces today
A score, and a reason code mapped to it after the fact.
What Notary supplies
The basis behind each stated reason, resolved to the file it came from, so the reason given is the reason that operated.
What has to exist
The disposition, the narrative behind it, and whose decision it actually was.
The first thing a BSA examiner asks for, often years after the alert closed.
What your AI produces today
A closure note the agent wrote about its own work.
What Notary supplies
Every assertion in that narrative, checked against the case file, with anything unsupported named rather than buried.
What has to exist
The ticket and the basis for the recommendation, created at the time rather than assembled later.
Requested in customer arbitration, and in routine examination sweeps.
What your AI produces today
The recommendation, and the prompt that produced it.
What Notary supplies
The stated basis, resolved to the client record and product terms it drew on, sealed at the moment of the recommendation.
What has to exist
Why this one and not that one, in terms a person could argue with.
Needed for the annual bias audit, and by anyone who files a charge.
What your AI produces today
A rank, and a generated justification.
What Notary supplies
Each stated reason for the outcome, resolved to the application material it came from, or marked as resting on nothing.
A record layer that sits behind your agent.
When your agent decides something, Notary takes what it decided, the reasons it gave, and the evidence it had. Every reason gets checked against that evidence. The result is kept under the reference your customer would quote, and stays openable for as long as the decision can be questioned.
Grounded Decision Records for Consequential AI
The record format, the states a claim can resolve to, and the rule that a model may propose while only verification decides. Published, so anyone can check our working without us.
Notary AI Research · August 2026 · PDF
Capture
One call after the decision, off your request path.
- Takes the reference, the reasons given, and the evidence the agent held
- Runs asynchronously, so a failure on our side never fails your decision
- Nothing is sampled. Every decision is taken
Verify
Every reason checked on its own, against what the agent was actually given.
- The reply is split into individual claims before anything is checked
- A model may propose the source; only resolution decides
- Anything resolving to nothing is named, not quietly dropped
The document
The claim file, determination record or credit memo, produced for one decision.
- Opens years later at the number a complaint arrives with
- Each reason shows what it resolved to, or the space where a source should be
- Sealed at the decision. Corrections are added, never written over
The ledger
All of them together, showing which reasons keep failing.
- Ranks the reasons that rest on nothing across every decision
- Points at the cause: a missing source, a template, a change you shipped
- Gives you the share of decisions you can actually account for
The unit
One decision.Not a trace, a session, or a run. One judgement about one person.
The key
Your own reference.The claim, case, or application number a complaint actually arrives with.
The output
The document you already owe.A claim file, a determination record, a credit memo. Notary supplies the substantiated basis inside it.
And what it never does.
Never sits in your request path. Never samples. Never tells you whether your decision was right.
One call after your agent decides.
You hand us the reference, the reasons, and the evidence the agent had. Everything after that happens off your critical path, so a failure on our side never fails your decision.
// after your agent decides, never before await notary.record({ ref: "claim CL-40921", reasons: agent.decision.reasons, evidence: agent.context, });
That's everything we need. This is what comes back.
The claim file, issued.
One of the six documents above, produced for one decision. Six months later someone types the reference and it opens exactly as it was sealed, every reason resolved to what it came from or marked as resting on nothing.
We didn't build your agent and we don't sell you the decision, so we have nothing to gain from a favourable answer. That is why it counts for anything to the person reading it.
- Standing is computed, never granted. No API, administrator, or support ticket can set it. If a reason doesn't resolve, it stays unsupported.
- Your content is stored apart from the account. Erase it whenever you like and the record survives, saying what can no longer be checked.
- We don't train on it, and every subprocessor is named publicly.
One bad reason is never one decision.
One record answers one case. Every record together answers a different question: which of your agent's reasons keep resting on nothing, and what put them there.
| The reason | Given | Unsupported | Why |
|---|---|---|---|
| "exempt from the deductible" | 1,204 | 1,18798.6% of them | The agent is never shown the deductible schedule |
| "policy in force" | 9,880 | 310.3% | Healthy |
| "covered peril" | 7,412 | 540.7% | Healthy |
claims where the agent cited a deductible exemption your policy never contained. Notary reports the count and the exposure per claim. What that adds up to, and what to do about it, is yours to decide.
Works with what you already run.
Notary doesn't replace your observability, your evaluations, or your governance platform. It reads what they produce and turns it into the one thing none of them makes: an account of a single decision that survives being asked about.
One decision at a time
Decisions taken
4,102
Accounted for
4,071
Resting on nothing
31
Traces tell you the agent ran. Evaluations tell you it performed on the cases you wrote. Governance tells you a policy exists. All three are inputs here, and none of them answers what a single decision rested on.
Three uses for one record.
Get it live
Risk and legal keep asking for the same thing. Give it to them and the agent stops living in pilot.
Answer when someone asks
A customer disputes it, an examiner requests it, a lawyer subpoenas it. Open the decision and show what it rested on.
Watch the pattern
The ledger ranks the reasons that keep failing and points at the cause, so a pattern surfaces as a line in a report rather than a surprise at year end.
We're taking a handful of design partners.
What you give
- One kind of decision. Not your whole stack.
- What your agent was given when it made them.
- An hour a week with someone who knows the workflow.
What you get
- The number you'll eventually be asked for: what share of your AI's decisions you can actually account for. Almost nobody has seen this figure for their own system.
- The reasons that don't hold up, named and ranked, with the cause attached.
- Records you can put in front of a review, instead of a questionnaire.
- A say in the format while it's still being set. It will be set by whoever is in the room.
We run your own decisions first and show you the number. If it comes back clean, we'll tell you.
Run it on last month's decisions.
We'll check every reason your AI gave and tell you how much of it holds up.