I outlined three calibration levels for the AnyJev library, each trading label effort for confidence and latency.
L0 requires no labels. It rotates the answer options to reduce position bias, but its probabilities remain uncalibrated.
L1 uses 100-500 labeled examples per question to calibrate those probabilities through temperature scaling. It improves confidence estimates without changing the answer ranking.
L2 uses 100-300 labels to fit a closed-form head on an intermediate hidden state. It leaves model weights untouched and can stop inference before the full forward pass completes.