Human Benchmark

Jul 2026 – Present·Live demo

A harness for measuring whether an agent actually reproduces a specific person's judgment, rather than producing generically competent output.

Grew out of needing a way to tell whether Cortex was working. Generic evals do not answer that question.