LLM Eval Compare

Baseline:

This is a post-unblind review view. Each column header shows its role, policy, model id, and original blinded candidate label.

The left column is the baseline unless the report was generated with fixed --left/--right policies. The right column is the computed winner, challenger, or fixed right policy for that case.