What if AI names a nonexistent column? The useful next question is which real column it could have meant. That choice is what interests V about SafeQL. Looking inside the database’s structure can turn a vague guess into a proposal someone can inspect.

Start with the candidate list

The pinned README illustrates replacing missing dept with department in employees, whose columns are id, name, department and salary. This is the authors’ illustration, not our executed trace.

The paper uses type constraints and embedding similarity to repair faulty components. Here, SELECT’s dept provides no type-only reason to exclude salary or id.

Missing dept → id, name, department, salary → similarity ranking → proposed department.
Conceptual explanation of the authors’ example, not observed ranks, scores or output.

The gap between the candidate list and the recommendation matters. Similarity gives a reason to suggest a name; it does not settle the user’s meaning. Imagine a company that records an employee’s home department separately from the department charged for their work. Either could look plausible. This company example is invented.

What the 5.8-point gain compares

KAIST’s September 4 announcement describes the improvement broadly. Separate the comparators below.

Table 3 (PDF p.11): GPT-OSS-120B; PostgreSQL-migrated BIRD full development set, 1,534 questions/11 databases. DAIL-SQL execution accuracy: unrefined 57.5%, RED-SQL 62.9%, SafeQL search 62.5%, hybrid 63.3%.

Hybrid compared withPercentage points
No refinement+5.8 percentage points
RED-SQL+0.4 percentage points

Differences use rounded values. RED-SQL is the strongest listed DAIL refiner by accuracy. Statistical significance and production correctness are unestablished.

V would keep the smaller comparison visible too. Adding a repair stage and replacing an existing repair stage are different decisions. A single headline gain can make a modest reason to switch look much larger than it is.

A day with no cancelled orders

Imagine a shop query for “orders cancelled today” returning no rows. If no orders were cancelled, the empty list is correct. Changing today to yesterday, or cancelled to completed, just to obtain rows would answer a different question without necessarily producing an error.

V would ask three separate questions of a proposed repair. Does it execute? Does it preserve the requested meaning? May this account access these data and perform this operation? Passing the first check cannot answer the other two. Showing which conditions must remain unchanged matters as much as choosing a replacement column.

Where V would look first

A read-only analytics tool is where I would look first: show the original name beside the proposed replacement and let the user check the meaning. If the date range, status condition or amount threshold also changed, expose that change. This is an editorial proposal, not a report of an existing product interface.

The test that could change my view is concrete. On approved test data, include missing columns, valid columns with the wrong meaning, and legitimately empty answers. Record preservation of the question and review time alongside execution success. If more executable answers also mean more silently changed conditions, keep repair at the suggestion stage.

Further reading and sources

When database summaries go to AI examines what information leaves the database, a separate boundary from choosing the right name.

This analysis uses public documentation. We did not install the extension or run a database.