Evaluation & Feedback
Ratings, corrections, side-by-side output comparison and eval results.
Components
5 components availableThumbs up and down that ask the one follow-up question worth asking — with a state for feedback that has been sent and cannot be taken back.
Inline CorrectionFix the answer where it is wrong, in place — the edit becomes training data, so the component keeps the original alongside it rather than overwriting it.
Output ComparisonTwo answers, side by side, with the model names hidden until a winner is picked — because a visible label is the thing being measured.
Eval ResultsPass rates per test case across two runs, with the regressions surfaced first and a sample size honest enough to say when a delta means nothing.
Edit Diff ViewA streamed AI edit as hunks you accept or reject one by one — id-keyed decisions, word-level marks, and buttons that stay disabled while a hunk is still arriving.