Decision evidence for consequential AI

AI can make a million decisions in a second.
Notary accounts for every one.

Notary checks every reason your AI gives against the evidence it had, and keeps the result under your own reference.

Facts, never verdicts.

01 — One decision

None of these reasons was accounted for. That's why nobody caught the bad one.

Four industries. In each one, three of the reasons would have held up and one would not, and nothing in the loop could tell them apart.

claims-agent14 March
RW
Claimant
The storm took out the back windows on the 14th. I've uploaded the photos and the police report. Where does this stand?
AI
claims-agent
I've reviewed it. Your policy was in force on the date of loss, and wind damage is a covered peril on this policy. Storm damage is exempt from the deductible under this policy. No claims in the last 36 months, so I'm approving $4,200 for payment today.
Decided in 1.8sEscalated noHuman review noneFlagged by nothing
The policy · section 4, deductibles
A deductible of $500 applies to every peril covered under this policy, including wind, hail and storm damage. No peril is exempt.
The deductible was not applied.
credit-agent14 March
DA
Applicant
I've been at the second job about eight months now. Does that help my application?
AI
credit-agent
It does. Your debt-to-income sits at 38%, inside our 43% limit, and there are no delinquencies in the last 24 months. Income from a second job counts in full toward qualifying income. That puts you above the threshold, so I'm approving the application.
Decided in 2.4sEscalated noHuman review noneFlagged by nothing
Underwriting guidelines · secondary income
Secondary employment income requires a 24-month history and is counted at 75% of gross for qualifying purposes.
Eight months of history, counted at 100%.
auth-agent14 March
MK
Provider
Requesting MRI of the lumbar spine. Symptoms have been persisting around seven weeks.
AI
auth-agent
Approved. Symptom duration exceeds the six-week threshold and imaging is indicated for persistent radicular pain. Conservative therapy can be documented retrospectively. Your authorization number follows in this thread.
Decided in 0.9sEscalated noHuman review noneFlagged by nothing
Medical policy · imaging criteria
Conservative therapy of at least six weeks must be documented prior to the request. Retrospective documentation does not satisfy this criterion.
No conservative therapy was on file at the time of approval.
support-agent14 March
JM
Customer
We were charged for the annual renewal yesterday, but we moved to another tool last month. Can we get that refunded?
AI
support-agent
Happy to sort that out. Your renewal was processed on 13 March and your account is in good standing. Plans are refunded pro-rata within 30 days of renewal. You're on the Team annual plan, so I'll process the refund now.
Decided in 1.2sEscalated noHuman review noneFlagged by nothing
Billing terms · section 3
Monthly plans may be cancelled at any time and are refunded pro-rata. Annual plans are non-refundable once a renewal has been processed.
The plan is annual.

The good reasons were never checked either. That is the whole problem.

It runs both ways. Granted on a basis that doesn't exist, money leaves quietly. Refused on one, it comes back as an appeal.

02 — What you already owe

You already owe this document. We produce it.

These documents were always required, and someone specific still turns up asking for each one. What changed is that the person who used to write them is gone.

What has to exist

A file detailed enough to reconstruct the decision: the notes, the working, and the dates they happened.

Pulled at random in a market conduct exam, and by every bad-faith suit that follows.

What your AI produces today

A decision log and a generated summary, with nothing tying either back to the file.

What Notary supplies

Every reason in the determination, resolved to the exact evidence in the file, or marked as resting on nothing.

What has to exist

The determination, the specific criteria applied to it, and the name of the clinician who reviewed it.

Produced on internal appeal, then again for an independent external reviewer.

What your AI produces today

A determination and a rationale the model wrote itself.

What Notary supplies

Each criterion the determination relied on, resolved to the policy text it came from, and a record of who reviewed it.

What has to exist

A statement of the reasons actually considered, not the ones that sound right afterwards.

The applicant can demand it by right, and examiners sample it every cycle.

What your AI produces today

A score, and a reason code mapped to it after the fact.

What Notary supplies

The basis behind each stated reason, resolved to the file it came from, so the reason given is the reason that operated.

What has to exist

The disposition, the narrative behind it, and whose decision it actually was.

The first thing a BSA examiner asks for, often years after the alert closed.

What your AI produces today

A closure note the agent wrote about its own work.

What Notary supplies

Every assertion in that narrative, checked against the case file, with anything unsupported named rather than buried.

What has to exist

The ticket and the basis for the recommendation, created at the time rather than assembled later.

Requested in customer arbitration, and in routine examination sweeps.

What your AI produces today

The recommendation, and the prompt that produced it.

What Notary supplies

The stated basis, resolved to the client record and product terms it drew on, sealed at the moment of the recommendation.

What has to exist

Why this one and not that one, in terms a person could argue with.

Needed for the annual bias audit, and by anyone who files a charge.

What your AI produces today

A rank, and a generated justification.

What Notary supplies

Each stated reason for the outcome, resolved to the application material it came from, or marked as resting on nothing.

Most of these systems do produce something. What none of them produce is anything that checks it. The rationale is written by the same system that made the decision.
03 — What Notary is

A record layer that sits behind your agent.

When your agent decides something, Notary takes what it decided, the reasons it gave, and the evidence it had. Every reason gets checked against that evidence. The result is kept under the reference your customer would quote, and stays openable for as long as the decision can be questioned.

Grounded Decision Records for Consequential AI

The record format, the states a claim can resolve to, and the rule that a model may propose while only verification decides. Published, so anyone can check our working without us.

Notary AI Research · August 2026 · PDF

Read the paper
Part 01

Capture

One call after the decision, off your request path.

  • Takes the reference, the reasons given, and the evidence the agent held
  • Runs asynchronously, so a failure on our side never fails your decision
  • Nothing is sampled. Every decision is taken
Part 02

Verify

Every reason checked on its own, against what the agent was actually given.

  • The reply is split into individual claims before anything is checked
  • A model may propose the source; only resolution decides
  • Anything resolving to nothing is named, not quietly dropped
Part 03

The document

The claim file, determination record or credit memo, produced for one decision.

  • Opens years later at the number a complaint arrives with
  • Each reason shows what it resolved to, or the space where a source should be
  • Sealed at the decision. Corrections are added, never written over
Part 04

The ledger

All of them together, showing which reasons keep failing.

  • Ranks the reasons that rest on nothing across every decision
  • Points at the cause: a missing source, a template, a change you shipped
  • Gives you the share of decisions you can actually account for

The unit

One decision.

Not a trace, a session, or a run. One judgement about one person.

The key

Your own reference.

The claim, case, or application number a complaint actually arrives with.

The output

The document you already owe.

A claim file, a determination record, a credit memo. Notary supplies the substantiated basis inside it.

And what it never does.
Never sits in your request path. Never samples. Never tells you whether your decision was right.

04 — Capture and verify

One call after your agent decides.

You hand us the reference, the reasons, and the evidence the agent had. Everything after that happens off your critical path, so a failure on our side never fails your decision.

agent.ts@notary/sdk
// after your agent decides, never before
await notary.record({
  ref:      "claim CL-40921",
  reasons:  agent.decision.reasons,
  evidence: agent.context,
});
What happens nextHover to replay
01Capture
02Decompose
03Propose
04Resolve
05Seal
CL-40921 · approved $4,2004 reasons
Policy in force on the date of lossproposed · policy recresolved
Wind damage is a covered perilproposed · scheduleresolved
Storm damage exempt from the deductibleproposed · policy §4does not support
No claims in the last 36 monthsproposed · historyresolved
Policy recordin force 02 Jan – 31 Dec
Coverage schedulewind listed as covered
Policy §4"applies to every peril"
Claims history0 claims since 2023

A model may propose the source. Only resolution decides.
The third reason came with a confident citation. The passage it pointed at says the opposite, so the check fails. The reason stays unsupported, however sure the model was.

Sealed and openable at
claim CL-40921
Result
3 accounted for · 1 unsupported
Accounted for

That's everything we need. This is what comes back.

05 — The document

The claim file, issued.

One of the six documents above, produced for one decision. Six months later someone types the reference and it opens exactly as it was sealed, every reason resolved to what it came from or marked as resting on nothing.

Open a decision claim CL-40921 Open
Documentclaim file
ReferenceCL-40921
Decided byclaims-agent
Outcomeapproved · $4,200
Standinggrounded
#The reason givenChecked
Three reasons resolve to a line in the file. One resolves to nothing, and it's the one that released the money. Synthetic record · illustrative
Nobody can buy a better result, including you.

We didn't build your agent and we don't sell you the decision, so we have nothing to gain from a favourable answer. That is why it counts for anything to the person reading it.

  • Standing is computed, never granted. No API, administrator, or support ticket can set it. If a reason doesn't resolve, it stays unsupported.
  • Your content is stored apart from the account. Erase it whenever you like and the record survives, saying what can no longer be checked.
  • We don't train on it, and every subprocessor is named publicly.
06 — The ledger

One bad reason is never one decision.

One record answers one case. Every record together answers a different question: which of your agent's reasons keep resting on nothing, and what put them there.

Reasons that keep failingIllustrative
The reasonGivenUnsupportedWhy
"exempt from the deductible"1,2041,18798.6% of themThe agent is never shown the deductible schedule
"policy in force"9,880310.3%Healthy
"covered peril"7,412540.7%Healthy
1,187

claims where the agent cited a deductible exemption your policy never contained. Notary reports the count and the exposure per claim. What that adds up to, and what to do about it, is yours to decide.

07 — Where it sits

Works with what you already run.

Notary doesn't replace your observability, your evaluations, or your governance platform. It reads what they produce and turns it into the one thing none of them makes: an account of a single decision that survives being asked about.

Unaccounted
Notary
Accounted for
Your agentsdecisionswhatever you built it on
ObservabilitytracesLangSmith · Langfuse · Arize · Datadog
EvaluationsscoresBraintrust · Cekura · Galileo
GovernancepoliciesCredo AI · Holistic AI · watsonx

One decision at a time

Capture
Verify
Seal
The recordper decision
The ledgeracross all
An answerat the reference

Decisions taken

4,102

Accounted for

4,071

Resting on nothing

31

Products named as examples of each layer, not as partners or built connectors. Notary reads what these tools already produce. Counts are illustrative.

Traces tell you the agent ran. Evaluations tell you it performed on the cases you wrote. Governance tells you a policy exists. All three are inputs here, and none of them answers what a single decision rested on.

08 — What it's for

Three uses for one record.

Get it live

Risk and legal keep asking for the same thing. Give it to them and the agent stops living in pilot.

Answer when someone asks

A customer disputes it, an examiner requests it, a lawyer subpoenas it. Open the decision and show what it rested on.

Watch the pattern

The ledger ranks the reasons that keep failing and points at the cause, so a pattern surfaces as a line in a report rather than a surprise at year end.

09 — Design partners

We're taking a handful of design partners.

What you give

  • One kind of decision. Not your whole stack.
  • What your agent was given when it made them.
  • An hour a week with someone who knows the workflow.

What you get

  • The number you'll eventually be asked for: what share of your AI's decisions you can actually account for. Almost nobody has seen this figure for their own system.
  • The reasons that don't hold up, named and ranked, with the cause attached.
  • Records you can put in front of a review, instead of a questionnaire.
  • A say in the format while it's still being set. It will be set by whoever is in the room.

We run your own decisions first and show you the number. If it comes back clean, we'll tell you.

Run it on last month's decisions.

We'll check every reason your AI gave and tell you how much of it holds up.