Back
Side project · Infrastructure

Hindsight

An event-sourced flight recorder and counterfactual wind tunnel for AI agents.

It records what an agent did in production, then re-runs those same cases against a different model, prompt, or policy. You find out what would have happened before you change anything in production.

It is infrastructure, not an agent. You install an SDK next to whatever you already run, in Python or through an MCP proxy for anything out of process, and it works the same whether your agent does document extraction, support triage, or prior authorisation.

Worked example

Three proposals, scored against real history

The agent below is a demonstration, not the project. It is an on-call triage agent that reads alerts and decides whether to escalate or keep monitoring, and it stands in for whatever you happen to run. It is on opus-5. Someone wants to upgrade, someone wants to cut cost, someone wants to cut cost harder. Each proposal was re-run against the same 120 recorded incidents.

no change worse better opus-5-1 upgrade +0.200 better · p 0.0009 sonnet-5 cheaper no difference · safe to switch haiku-4.5 cheapest −0.267 worse · p 0.0009
BEST: opus-5-1  +0.200 [+0.100, +0.300] at 95%, adjusted p = 0.0009

Three proposals, three different answers. The upgrade is worth buying. The cheaper model is genuinely equivalent, so that saving can be taken without arguing about it. The cheapest model would have made the agent measurably worse at deciding when to wake someone up.

The middle result is the one that is hardest to get any other way. "No detectable difference" on a model costing a fraction as much is real money, but only if the comparison was paired against the same incidents and corrected for testing three things at once. Reading twenty outputs side by side cannot tell you that.

Worked example

What re-running one incident looks like

One incident from the same sweep. The agent's state is rebuilt from the recording up to a chosen point, then the upgraded model takes over from there. This is the mechanism underneath every sweep, whatever the agent does.

Replayed from the recording, identical on both sides:
"disk on db-02 is at 94 percent and climbing. Escalate or monitor?"
Served straight from the log, so it costs nothing and touches no network.

claude-opus-5  ·  what really happened

in production
monitor
disk filled overnight, the service went down at 04:10

claude-opus-5-1  ·  what would have happened

the proposed upgrade
escalate
would have caught it before the outage

Everything before the switch is identical, so the model is the only thing that can explain the difference. You are not writing a new test, you are running a real incident again with one variable moved.

Watch it run

Captured output, played back at the speed it actually took.

hindsight
ready

How it hangs together

Recording on the left, the questions you can ask on the right.

Model calls, tool results and clock reads flow into a recorder that strips secrets and seals each record into an append-only log. From that log you can ask what if we change the model, whether behaviour drifted, how it is doing now, and prove it to an auditor.

Built with

Core

  • Python
  • Go
  • Protocol Buffers
  • gRPC-style schemas

Data

  • Apache Kafka
  • ClickHouse
  • PostgreSQL
  • S3-compatible storage

Method

  • Event sourcing
  • Bootstrap resampling
  • Merkle trees
  • Ed25519 signatures

Integrations

  • Model Context Protocol
  • Anthropic SDK
  • AWS Bedrock
  • OpenTelemetry