Skip to content
Published evaluations

Results

Compare agents on the four task subsets used in the paper.

Dataset

CUA-World · 26 tasks

2 results · 26 tasks per result

2of 2 results
Filter models and reasoning effort
All models

Results

All included results, ordered by average performance.

#ModelEffortPerformanceTime / taskCost / task
1 GPT-6 Astrasvc_139_c67de8 Cost / task: $15.0610 Logged estimate
Source details

Run: svc_139_c67de8

Cost is short-context API-equivalent usage, not a subscription payment; full tool/response trace unavailable in CLI export.

Astra_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

{"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":26,"codex_model_response_count_and_full_tool_calls_unavailable":26}

xhigh 92.98% 1515.4s $15.0610 Logged estimate
2 GPT-6 Astrasvc_137_284111 Cost / task: $9.2869 Logged estimate
Source details

Run: svc_137_284111

Cost is short-context API-equivalent usage, not a subscription payment; full tool/response trace unavailable in CLI export.

Astra_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

{"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":26,"codex_model_response_count_and_full_tool_calls_unavailable":26}

low 90.29% 806.8s $9.2869 Logged estimate

All included results for this dataset are shown together. Execution settings can differ; source details are preserved with each run.

Data sources & metric definitions

Latest source inventory: 2026-09-20T20:08:33Z. Download all 146 source rows · Download displayed costs & assumptions · View dataset catalog

Times use the task clock, excluding setup and final verification. Costs are recorded usage or labeled estimates, not invoices. API estimates identify their pricing and cache assumptions in each result's source details. Where shown, the cache range runs from all input cached to none. Standard-price proxies do not claim the archived tier or version price. Partial-coverage estimates show their task counts. Additional pricing sources & assumptions are preserved alongside the original CSV references. Runs without token usage show unit prices but stay out of per-task cost plots. The full-OSWorld human baseline is not transferred to these dataset views.

0 retry-attempt or overlapping summary rows for this dataset remain in the CSV, not counted as independent evaluations. Excluded evaluations are listed below the results.

Model labels use publisher-sourced model metadata and explicit aliases, verified 2026-09-19. Open means downloadable weights, not necessarily an open-source license. Release dates mean first public availability of the identified model version. Unversioned archives with an unverified release date are labeled explicitly and omitted only from the release-date curve.

Steps count action batches sent to the environment: one step call counts once, regardless of the number of actions in its batch. Tool calls count individual recorded environment actions within those batches, including waits; a batch of five actions counts as five, not one. This is an environment-action count, not a count of model API tool-call wrappers. Separate observation and task-completion requests are not included. Observations and model responses are separate counts; steps are not observation count minus one. Output tokens per model response divide total generated tokens (including reasoning) by recorded responses, only when both cover every task and the usage/response counts agree.