Judgment models for agents. One forward pass, calibrated probabilities over the options you define. No generated text, nothing to parse.
Open weights and a public recipe: which objective yields calibrated probabilities, base or instruct start, dense or MoE.
A coding-agent harness that filters every tool output through a judgment model. Filter on vs. off, measured on cost, tokens and pass rate.
Same questions for every model: accuracy, agreement when options are reordered, NLL, Brier score, calibration error.