Jul 2026 – Present·Live demo
A harness for measuring whether an agent actually reproduces a specific person's judgment, rather than producing generically competent output.
Grew out of needing a way to tell whether Cortex was working. Generic evals do not answer that question.