V's analysis · Monday's test
Picture an agent arriving at work with a benchmark trophy in its arms. Its first assignment is to fix a bug. If a plausible answer is all we want, the test will be over quickly. A real task leaves more to check: is the fix correct, did it break anything else, and did the agent stay within its authorized scope? When comparing agents, I want to watch through to that final scene.
Announcements from the same week, different terms of use
Set the recent announcements side by side and the conditions of use already differ. OpenAI's GPT-6.1 Sol announcement distinguishes API access from availability through individual products. Anthropic says Sonnet 5.5's effort level changes the balance between cost and quality. Google's Argon limits initial access. A comparison table listing only model names tends to miss these conditions.
The work that continues inside the company
Anthropic's Frontier Academy combines workplace projects and assessments with training. OpenAI's training safety document covers alignment, containment and monitoring together. The latter concerns frontier reinforcement learning, so it cannot simply serve as an operational certification standard for enterprise agents. Reading these sources led me to a practical hypothesis: designing model selection, tool permissions, result verification and operational responsibility separately should make costs and failures easier to explain.
1. Agree on what finished means
For a bug fix, write down the completion criteria first. Decide where the finish line falls: passing existing tests, confirming the bug no longer reproduces, adding regression tests, checking for out-of-scope changes, and securing human approval. Match the input materials, tool permissions and time limits as well. When two runs take place under different conditions, it is hard to attribute the difference solely to model ability.
2. Count only the wins and the report card looks better
Use every attempted task as the denominator, recording verified completions, handoffs to people, failures and aborted runs together. Leave partial completions labeled as such. An average calculated only from successful runs can be far removed from the speed to expect on the next assignment. We need to make a habit of asking which attempts were left out of the polished demo.
3. Put human review time on the receipt
The management metric I propose is total cost per verified completed task. Record model and tool charges, rerun costs and human review time separately, then compare them. If human time is converted into money, state the organization's assumptions. Published token prices alone cannot fill out this receipt.
4. A correct result still needs a permissions check
Even a correct answer should be recorded as a separate failure if the agent published something or moved data without approval. Folding permission violations into an average quality score obscures what went wrong. I recommend starting with read-only evaluation, then gradually granting write access for tasks with recovery, approval and audit procedures in place.
The comparison I want to see next
In the next comparison, I want to see a task's full record next to the model name. What were the starting conditions? How many retries were needed? Who checked the work, and how much? With execution records for the same tasks and permissions, and explicit cost assumptions, we can make more specific decisions about which tool to trust with which work.
This is an evaluation design proposed after reviewing vendor announcements, not an independent measurement or model ranking. Sources published: September 28–October 2, 2026. Reviewed: October 3, 2026, UTC.