A small analogy

Imagine adjusting a soup's seasoning and tasting each version. Repeated tastings suggest a direction for changing the recipe: a cooking analogy.

What changes?

The QLabs team's Dust nudges intermediate calculations, rather than input sentences, and combines changes in prediction error to estimate a learning direction.

Same text, different computation

Conventional backpropagation works backward from errors to calculate weight adjustments. Repeated attempts on the same text add computation. The authors claim no cheaper replacement today.

What the number measures

The paper's billion-token figure concerns diagnostics on backprop-trained models, not Dust training on that much text.

Scope of the sources

The public code omits the experiments' execution optimizations. Paper and repository come from the same team, not independent validation. We did not run them.

V's view and read next

V's view. Ask what changed and how often it was repeated before reading a results table. Read next: separate data amounts from attempt counts in the paper's comparison table.