Suppose three AI judges cast two votes for answer A and one for B. Copy each recorded vote ten times: now A has 20 votes and B has 10. The decision is unchanged, but the count looks reassuring. What did this invented exercise actually teach us?

Ten times the records, the same observations

Take four hypothetical questions. Their three votes are A–A–B, A–B–B, A–A–A and B–B–B. We have supplied no correct answers. This counts support for A; it is not an experiment measuring AI accuracy.

Copy every vote ten times within its original question. Twelve records become 120. The four shares supporting A remain 2/3, 1/3, 1 and 0; majority decisions remain A–B–A–B. Overall support for A stays at 50%. Copying only one judge’s vote would change its weight and is a different exercise.

In an invented question, two A votes and one B vote become 20 A votes and 10 B votes. Support remains two-thirds. Across all four questions, the misapplied independent-vote standard error falls to 0.316 of its original size.
V’s hypothetical copy calculation. The upper cards show only question one. The lower ratio uses all four questions, with every vote copied equally. These are not fresh model calls or measured performance.

Why does the calculation become more certain?

A standard error describes uncertainty in an estimated mean. Deliberately pretend every vote is independent and apply √[p(1−p)/N], with p the share supporting A and N the vote count. Here it falls from about 0.144 to 0.0456. Copying leaves it at about 32% of its former size, exactly 1/√10. These values concern the support share, not the probability of a correct answer.

The calculation has counted copies as fresh observations. Analyze the four question-level means instead, and the inputs have not changed. Resampling the same question groups in the same way also gives unchanged results. This is an invariance of the copying exercise. Four invented questions do not validate an uncertainty estimate for the real world.

Download the 12-to-120-vote Python demonstration. It needs no extra packages and checks the records, support shares, majority choices and unchanged question-level calculations. Its independent-vote standard error is deliberately misapplied for comparison.

How closely do real AI judges resemble copies?

Calling a real model again can produce a different answer. Literal copying cannot tell us how much actual models’ errors overlap. It can expose the assumption behind treating a larger vote count as proportionally stronger evidence.

An ICML 2025 study investigated shared errors across language models. It gives a reason to examine error overlap, but cannot measure the accuracy of the particular panel you use today.

Different mistakes need not mean better decisions

Table 1 of a preprint posted September 18, 2026 supplies a useful countercase. On the same 400 RewardBench items, GPT, Claude and Grok together scored 93.3% majority accuracy; an eleven-judge mixed panel scored 91.5%. Lower error correlation in the mixed panel did not produce higher accuracy. This descriptive difference alone does not establish statistical superiority.

Before adding judges, V would ask which mistakes the next vote might catch. Diluting stronger judges with weaker ones is another risk. Avoiding duplicate evidence does not imply abandoning AI panels.

What would change V’s mind?

A panel that makes better decisions has a case for adoption. I would fix its membership and voting rule first, then compare it with the strongest single judge on the same questions, kept separate from tuning. I would examine final accuracy, unanimous mistakes and the share of questions left unanswered together. Answering less can make a system look more accurate.

If the panel repeatedly improves those decisions at an acceptable cost, V would choose it. If it adds names and speaking styles while preserving the same mistakes, I would spend the effort on checking original evidence or running executable tests. Tests also miss errors when their specifications or coverage are incomplete.

The only computation run for this article uses invented votes. We have not reproduced the paper’s model judgments or predicted any current panel’s performance. Real questions may themselves share dependencies, requiring a different choice of independent unit.

Reading an AI’s confidence score continues with the gap between one model’s stated confidence and measured correctness.