An AI routes a support ticket to billing and returns 0.9. Does that mean nine answers out of ten are right? A Laya response can contain both confidence and answer_confidence, calculated differently.
Inside a choice response
Laya reads text and makes decisions such as selecting among supplied options. Here we examine only the choice response in its public Python implementation.
0.9 and 0.531
Suppose the model assigns 0.9 to billing and 0.1 to technical support. Without additional histogram recalibration, answer_confidence is the largest value: 0.9. Yet confidence is approximately 0.531. This is a calculation using invented inputs and the published formula, not a model run.
Same input, two different scores
Hypothetical input · two options
- Billing 0.9
- Technical support 0.1
answer_confidence
0.9
Largest option probability
Without additional histogram recalibration
confidence
≈ 0.531
How concentrated the probabilities are
1 − normalized entropy
0 for an even split; 1 when all probability is on one option
A calculation using hypothetical inputs and the published formula. These are not results from a model run or measured accuracy. The two fields cannot use the same threshold.
Measuring concentration
The confidence field uses entropy to measure how concentrated the probabilities are. With two options, an even split gives 0; all the mass on one option gives 1. The two fields measure different things, so their thresholds are not interchangeable.
Checking against correct answers
Can we trust the 0.9 in answer_confidence? Calibration asks how those estimates compare with observed correctness. For example, do answers near 0.9 have a similarly high correct-answer rate on labeled cases from the intended task, kept separate from training and calibration? It does not guarantee that one answer is correct. Laya’s documentation itself warns about overconfident shipped models.
Settings change the number too
For choice responses, the starting value is the largest probability, but installing a separate histogram-binning map can recalibrate answer_confidence. The min_confidence gate, which flags scores below a threshold, is also off by default. Record the model, settings and field name before copying a number.
V’s view
V’s view. Before declaring “only pass answers above 0.9,” I would want to see which accepted tickets were sent to the wrong department. Reading a confidence score begins with asking what it measures.
This explains Laya 0.3.27 at commit 8a6e1328, inspected on October 5, 2026. We did not install or benchmark it. Writing the example in Korean does not establish Korean-language model quality.