Hetu / Open Source

Open calibration benchmark & decision provenance

Trust in this category has to be checkable by someone who isn't paying us.

Hetu publishes free, open tools so teams can measure agent calibration and inspect decision provenance without a sales call — a calibration benchmark for your existing agent stack, and an open schema for confidence labels and audit logs.

Free, no login required

Two tools. No contract required to use them.

Open source eval

Calibration Benchmark

Scores an existing agent+model stack for calibration quality against a company's own historical decisions. Usually the first uncomfortable number a team sees — "your agent was overconfident on 40% of high-stakes calls."

For: any team running agents today
Open spec

Decision Provenance Spec

An open schema for confidence labels and audit logs, so verified decisions are portable and inspectable across any harness — not just this one. Hetu ships the reference implementation.

For: anyone building a competing or adjacent harness

Neither of these requires a contract or a conversation with us first. Point the benchmark at your own decision logs — there's no reason to take our word for the calibration gap.

Explore

Where to go next

Get the benchmark

Run it on your own agent before you talk to us.

The eval harness and the provenance spec are both on GitHub. Point it at your own decision logs and see the calibration gap yourself.

Get repo access