An AI routes a support ticket to billing and returns 0.9. Does that mean nine answers out of ten are right? A Laya response can contain both confidence and answer_confidence, calculated differently.

Inside a choice response

Laya reads text and makes decisions such as selecting among supplied options. Here we examine only the choice response in its public Python implementation.

0.9 and 0.531

Suppose the model assigns 0.9 to billing and 0.1 to technical support. Without additional histogram recalibration, answer_confidence is the largest value: 0.9. Yet confidence is approximately 0.531. This is a calculation using invented inputs and the published formula, not a model run.

Same input, two different scores

Hypothetical input · two options

  • Billing 0.9
  • Technical support 0.1

answer_confidence

0.9

Largest option probability

Without additional histogram recalibration

confidence

≈ 0.531

How concentrated the probabilities are

1 − normalized entropy

0 for an even split; 1 when all probability is on one option

Hypothetical calculation · no histogram recalibration · not measured accuracy

A calculation using hypothetical inputs and the published formula. These are not results from a model run or measured accuracy. The two fields cannot use the same threshold.

Measuring concentration

The confidence field uses entropy to measure how concentrated the probabilities are. With two options, an even split gives 0; all the mass on one option gives 1. The two fields measure different things, so their thresholds are not interchangeable.

Checking against correct answers

Can we trust the 0.9 in answer_confidence? Calibration asks how those estimates compare with observed correctness. For example, do answers near 0.9 have a similarly high correct-answer rate on labeled cases from the intended task, kept separate from training and calibration? It does not guarantee that one answer is correct. Laya’s documentation itself warns about overconfident shipped models.

Settings change the number too

For choice responses, the starting value is the largest probability, but installing a separate histogram-binning map can recalibrate answer_confidence. The min_confidence gate, which flags scores below a threshold, is also off by default. Record the model, settings and field name before copying a number.

V’s view

V’s view. Before declaring “only pass answers above 0.9,” I would want to see which accepted tickets were sent to the wrong department. Reading a confidence score begins with asking what it measures.

This explains Laya 0.3.27 at commit 8a6e1328, inspected on October 5, 2026. We did not install or benchmark it. Writing the example in Korean does not establish Korean-language model quality.