Services / Open source
Trust in this category has to be checkable by someone who isn't paying us.
Two things we give away, free, no sales call required. This is how a team finds out it needs the platform — not how we make money on it directly.
Free, no login required
An eval you can run against your own logs, and a spec anyone can implement.
Calibration Benchmark
Open source eval
Scores an existing agent and model stack for calibration quality against a company's own historical decisions. Usually the first uncomfortable number a team sees — your agent was overconfident on 40% of high-stakes calls.
github.com/hetu-labs/calibration-benchmark
Decision Provenance Spec
Open spec
An open schema for confidence labels and audit logs, so verified decisions are portable and inspectable across any harness — not just this one. Hetu ships the reference implementation.
github.com/hetu-labs/decision-provenance-spec
Neither requires a contract or a conversation with us first. Point the benchmark at your own decision logs — there's no reason to take our word for the calibration gap.
Then
Run it on your own agent before you talk to us.
Point the harness at your own decision logs and see the calibration gap yourself. If the number is uncomfortable, that is the conversation.