The book is above the cup. The book is below the cup. One word reverses this imagined scene. Yet a numerical map of words can place the two directions close together.

A model for relationships

RelateAnything predicts relationships between image regions. It takes the picture, regions and relation phrases separately. Its maker’s example looks for a person riding a horse.

The problem outside the photograph

The September 11 preprint reports that dino.txt produced very similar vectors for above and below. The author attributes this to antonyms appearing in similar contexts. Matching visual features to those targets makes direction harder to distinguish. A new text encoder was trained to separate opposites while preserving synonyms.

V’s view

V’s view. The intriguing repair happens outside the photograph. Sharper views of a cup and book are an obvious place to start; the map used to express their relationship also needs direction. Identifying the objects and describing their relationship become separate questions.

Limits and further reading

Limits and further reading: this concerns a particular encoder and research setup, not every AI’s spatial understanding. We did not reproduce it. Section 3.3 and Figure 3 compare the word representations.