Hetu / Industries / Financial services

AI decision verification for lending, underwriting & collections

You have forty agent pilots and nothing in production that touches a real decision.

Hetu helps banks, NBFCs, and insurers put AI agents on real credit decisions — with deterministic checks, causal attribution, and confidence labels model risk committees can accept. For heads of AI, CROs, and model risk teams where the blocker on autonomy was never capability. It was accountability.

The model risk committee

The demo was excellent. The agent read the portfolio, found the delinquency spike, and wrote a paragraph explaining it that was clear, specific, and confident.

It was also wrong. And there was no way to tell that from the output — the wrong answer and the right answer are rendered in the same font, with the same certainty, at the same speed.

So the committee asked the only question that matters: when it is wrong, how will we know, and what will it have done by then? You did not have an answer. The pilot is still a pilot.

Why this one is different

In lending, a confidently wrong answer isn't an embarrassment. It's a provision.

Silent
A hallucinated root cause produces a mistargeted intervention. It fails quietly, over a quarter, and looks like bad luck.
Written
Your supervisor will ask for the reasoning behind an automated action. A prompt log is not reasoning. It's a transcript.
Compounding
An agent that misreads sourcing quality as a collections problem keeps buying the bad cohort. The error funds itself.

The industry response has been to keep a human in the loop on everything, which means you have not built an agent — you have built an expensive suggestion box, and you are paying for both the model and the reviewer.

What Hetu does about it

The model is never allowed to compute. It only narrates what was proved.

"How do we know when it's wrong?"

It tells you before it's wrong

Every conclusion carries a label: confirmed, high, medium — two candidates, or hypothesis — requires your input. The label is computed from explained variance and interval overlap, not from the model's self-report.

"What if the evidence is thin?"

It refuses

Below the sample-size floor, the conclusion is withheld and the thin metric is named. A regulatory or macro event overlapping the window suppresses attribution entirely rather than assigning it to the nearest plausible cause.

"Can it hallucinate a number?"

Structurally, no

Figures come from a deterministic rule engine or a fitted causal model. The narration call is separate, receives only the verified object, and is rejected if its output contains a number or a causal claim not in the payload.

"What can it do unsupervised?"

Only what it has proved

The autonomy envelope starts closed. Decision types move inside it as measured outcomes demonstrate accuracy, and fall back out when they don't. Irreversible, fat-tailed actions never auto-execute at any confidence level.

"What do we show the regulator?"

One immutable file

Every decision generated, approved, deferred, executed and measured — who, when, what, why, on what evidence, and what happened. Refusals are logged as first-class records, not as gaps.

"Do we have to rebuild the graph?"

Your experts seed it

The causal graph and the WHY traversals are configuration, hand-seeded by the people who already know your failure patterns. Adding one is an insert, not a deployment — and it stays auditable in a table your risk team can read.

Running today

An NBFC origination pipeline, from field sourcing to 90 days past due.

Not a benchmark. A production deployment where the target variable is deliberately the hard one: not a disbursed loan, but a loan still performing at 90 DPD. Anything less lets a system claim success for originating bad credit quickly.

  • The graph extends past the sale. Sourcing → underwriting → disbursement → repayment. That's the only way "fewer loans disbursing" and "more disbursed loans going bad" can be told apart — they look identical in a throughput chart and demand opposite interventions.
  • Confounders are declared, not discovered. Policy changes and macro events sit in a maintained log. When one overlaps an anomaly window, the system says so instead of attributing around it.
  • Accuracy is validated on a labelled holdout, stratified by confidence label, against root causes a human analyst determined by file review.
TriggerOutcomeLabel
Performing rate down 14%, one zoneSourcing quality at two field officers, 60 days priorConfirmed
Throughput down while applications riseEarly delinquency dominant · 44%, CI 37–51High
Back-book DPD rising, policy change same windowAttribution suppressed. Two hypotheses, each with a falsification test.Refused

The third row is the one to take to your committee. A system that will say "I cannot attribute this, and here is what would settle it" is the only kind that can be trusted to act on the rows where it does not say that.

Deployment

It runs where your data already is.

In-VPCOn-premNo data used for trainingISO 27001SOC 2Immutable audit logModel-agnostic narration layerConfig-as-table causal graphCircuit breakersHuman gatesRegime-change detector

The narration model is swappable and the reasoning does not depend on it, because it never did any reasoning. A model upgrade improves the sentences and changes none of the maths.

FAQ

Financial services — common questions

How does Hetu help with model risk and AI governance?

Every decision carries a confidence label, a causal or deterministic rationale, and an immutable audit record — the artifacts a model risk committee asks for before an agent touches a real credit action.

Can this run in-VPC for banks and NBFCs?

Yes. Deployments support in-VPC and on-prem patterns. Your data is not used to train foundation models, and the reasoning layer does not depend on the narration model.

What is a Vertical Verification Pack for lending?

A pre-seeded causal graph and gate set for one regulated decision type — for example origination or collections — so you start from domain failure patterns instead of an empty prompt.

How do we start without a full platform rollout?

Most teams begin with one currently-manual decision via Consulting, then convert to the platform once the verification layer earns trust in shadow.

Explore

Where to go next

Design partners

Bring the decision your committee refused to let an agent make.

We start with one: a decision your team already makes manually, where you can articulate the failure patterns and where being wrong is expensive. We encode the constraint framework, seed the graph with your experts, and show you the first refusal — which is, reliably, the moment the room understands what this is.

Request access