Researchers make TypedBench, a test for typed decision models
The hosted model follows its policy, but it is sensitive to wording and underconfident.
Claimed, not confirmed
Researchers made TypedBench, a benchmark for typed decision models. These models give calibrated likelihoods for typed answers, and software acts on them with thresholds and rules. The researchers did tests on a hosted model, an open encoder and open decoders from 0.8B to 9B parameters. When errors have different costs, the hosted model can give likelihoods that are less good than its top answer. They say that a model must have a test for policy, wording, likelihood quality and decision results.
Sources
Posted