Results
Compare agents on the four task subsets used in the paper.
CUA-World · 26 tasks
2 results · 26 tasks per result
Filter models and reasoning effort
Results
All included results, ordered by average performance.
| # | Model | Effort | Performance | Time / task | Cost / task | |
|---|---|---|---|---|---|---|
| 1 |
GPT-6 Astrasvc_139_c67de8
Cost / task:
$15.0610
Logged estimate
Source detailsRun: svc_139_c67de8 Closed weights · Released 2026-09-03Cost is short-context API-equivalent usage, not a subscription payment; full tool/response trace unavailable in CLI export. Astra_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges {"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":26,"codex_model_response_count_and_full_tool_calls_unavailable":26} |
xhigh | 92.98% | 1515.4s | $15.0610 Logged estimate | |
| 2 |
GPT-6 Astrasvc_137_284111
Cost / task:
$9.2869
Logged estimate
Source detailsRun: svc_137_284111 Closed weights · Released 2026-09-03Cost is short-context API-equivalent usage, not a subscription payment; full tool/response trace unavailable in CLI export. Astra_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges {"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":26,"codex_model_response_count_and_full_tool_calls_unavailable":26} |
low | 90.29% | 806.8s | $9.2869 Logged estimate |
Task performance
Select a measure to compare. Tap a point for its model, score, and timing.
Scatter plot comparing model performance with the selected horizontal metric.
Performance, time, and cost
Drag to rotate; select a point to inspect it. Only results with known costs appear here.
Rotatable three-dimensional scatter plot comparing cost, time, and performance.
All included results for this dataset are shown together. Execution settings can differ; source details are preserved with each run.
Data sources & metric definitions
Latest source inventory: 2026-09-20T20:08:33Z. Download all 146 source rows · Download displayed costs & assumptions · View dataset catalog
Times use the task clock, excluding setup and final verification. Costs are recorded usage or labeled estimates, not invoices. API estimates identify their pricing and cache assumptions in each result's source details. Where shown, the cache range runs from all input cached to none. Standard-price proxies do not claim the archived tier or version price. Partial-coverage estimates show their task counts. Additional pricing sources & assumptions are preserved alongside the original CSV references. Runs without token usage show unit prices but stay out of per-task cost plots. The full-OSWorld human baseline is not transferred to these dataset views.
0 retry-attempt or overlapping summary rows for this dataset remain in the CSV, not counted as independent evaluations. Excluded evaluations are listed below the results.
Model labels use publisher-sourced model metadata and explicit aliases, verified 2026-09-19. Open means downloadable weights, not necessarily an open-source license. Release dates mean first public availability of the identified model version. Unversioned archives with an unverified release date are labeled explicitly and omitted only from the release-date curve.
Steps count action batches sent to the environment: one step call counts once, regardless of the number of actions in its batch. Tool calls count individual recorded environment actions within those batches, including waits; a batch of five actions counts as five, not one. This is an environment-action count, not a count of model API tool-call wrappers. Separate observation and task-completion requests are not included. Observations and model responses are separate counts; steps are not observation count minus one. Output tokens per model response divide total generated tokens (including reasoning) by recorded responses, only when both cover every task and the usage/response counts agree.