Skip to content
Published evaluations

Results

Compare agents on the four task subsets used in the paper.

Dataset

OSWorld 2.0 · 52 tasks

21 results · 52 tasks per result

21of 21 results
Filter models and reasoning effort
All models

Results

All included results, ordered by average performance.

#ModelEffortPerformanceTime / taskCost / task
1 GPT-6 Astrasvc_119_69f289 Cost / task: $8.1746 Logged estimate
Source details

Run: svc_119_69f289

Cost is short-context API-equivalent usage, not a subscription payment; full tool/response trace unavailable in CLI export.

Astra_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

{"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":52,"codex_model_response_count_and_full_tool_calls_unavailable":52}

xhigh 76.90% 1237.3s $8.1746 Logged estimate
2 GPT-6 Astrasvc_73_9d68f5 Cost / task: $7.3909 Logged estimate
Source details

Run: svc_73_9d68f5

Cost is short-context API-equivalent usage, not a subscription payment; full tool/response trace unavailable in CLI export.

Astra_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

{"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":52,"codex_model_response_count_and_full_tool_calls_unavailable":52}

high 75.02% 914.0s $7.3909 Logged estimate
3 GPT-6 Astrasvc_123_d9dc75 Cost / task: $6.3363 Logged estimate
Source details

Run: svc_123_d9dc75

Cost is short-context API-equivalent usage, not a subscription payment; full tool/response trace unavailable in CLI export.

Astra_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

{"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":52,"codex_model_response_count_and_full_tool_calls_unavailable":52}

low 68.17% 904.2s $6.3363 Logged estimate
4 GPT-6 Astrasvc_74_19ec78 Cost / task: $6.5192 Logged estimate
Source details

Run: svc_74_19ec78

Cost is short-context API-equivalent usage, not a subscription payment; full tool/response trace unavailable in CLI export.

Astra_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

{"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":52,"codex_model_response_count_and_full_tool_calls_unavailable":52}

medium 67.18% 828.1s $6.5192 Logged estimate
5 Gemini 3.8 Flashsvc_55_ffdb2e Cost / task: $5.0195 Logged estimate
Source details

Run: svc_55_ffdb2e

recorded_template_price_estimate. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

high 60.18% 3058.8s $5.0195 Logged estimate
6 Gemini 3.8 Flashsvc_54_a54e88 Cost / task: $2.9148 Logged estimate
Source details

Run: svc_54_a54e88

recorded_template_price_estimate. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

medium 59.93% 2051.8s $2.9148 Logged estimate
7 Claude Opus 5svc_57_addd34 Cost / task: $10.0063 Logged estimate
Source details

Run: svc_57_addd34

Claude_standard_global_list_price_estimate_from_logged_usage. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

low 52.06% 1385.7s $10.0063 Logged estimate
8 Claude Sonnet 5svc_58_909a13 Cost / task: $8.8630 Logged estimate
Source details

Run: svc_58_909a13

Claude_standard_global_list_price_estimate_from_logged_usage. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

low 39.08% 1993.6s $8.8630 Logged estimate
9 Gemini 3.8 Flashsvc_56_08ebc5 Cost / task: $1.8752 Logged estimate
Source details

Run: svc_56_08ebc5

recorded_template_price_estimate. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

low 35.29% 1386.7s $1.8752 Logged estimate
10 Meta Muse Spark 1.3svc_80_6fea36 Cost / task: $0.3903 API estimate · no cache Cache range: $0.0110–$0.3903
Source details

Run: svc_80_6fea36

API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges

minimal 33.97% 1516.0s $0.3903 API estimate · no cache Cache range: $0.0110–$0.3903
11 Meta Muse Spark 1.3svc_81_e35f89 Cost / task: $0.5029 API estimate · no cache Cache range: $0.0150–$0.5029
Source details

Run: svc_81_e35f89

API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges

low 33.68% 2280.8s $0.5029 API estimate · no cache Cache range: $0.0150–$0.5029
12 Meta Muse Spark 1.3svc_82_a48586 Cost / task: $0.5883 API estimate · no cache Cache range: $0.0198–$0.5883
Source details

Run: svc_82_a48586

Resumed after earlier interruptions; final archived event is run_done at 2026-09-09T19:09:11.302490+00:00, with all 52 planned tasks scored and timed. Earlier run_failed events are retained as history, not the final outcome.

API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges

medium 32.31% 2292.1s $0.5883 API estimate · no cache Cache range: $0.0198–$0.5883
13 Meta Muse Spark 1.3svc_84_e2ef21 Cost / task: $0.6399 API estimate · no cache Cache range: $0.0231–$0.6399
Source details

Run: svc_84_e2ef21

Resumed after earlier interruptions; final archived event is run_done at 2026-09-10T01:35:24.715393+00:00, with all 52 planned tasks scored and timed. Earlier run_failed events are retained as history, not the final outcome.

API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges

xhigh 32.16% 2261.0s $0.6399 API estimate · no cache Cache range: $0.0231–$0.6399
14 Meta Muse Spark 1.3svc_83_03ee99 Cost / task: $0.5985 API estimate · no cache Cache range: $0.0206–$0.5985
Source details

Run: svc_83_03ee99

Resumed after earlier interruptions; final archived event is run_done at 2026-09-09T21:41:29.284094+00:00, with all 52 planned tasks scored and timed. Earlier run_failed events are retained as history, not the final outcome.

API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges

high 31.46% 2179.3s $0.5985 API estimate · no cache Cache range: $0.0206–$0.5985
15 Kimi K3svc_147_d5291e Cost / task: $5.6247 Recorded usage
Source details

Run: svc_147_d5291e

Completed September 19 resume of evaluation 147; replaces its historical billing-affected snapshot for comparison. Only the 15 authorized billing-interrupted task/seed pairs were rerun; 37 original results unchanged. All 52 final tasks complete, with recorded cost and no HTTP 402 errors. Original frozen plan and submission unchanged. Prior interrupted attempts retained separately and overlap historical archives; do not sum snapshots as independent evaluations. Completed max-effort resume; the earlier billing-interrupted snapshot remains in the downloadable CSV. One final task reports more reasoning tokens than total output tokens; raw provider counters are retained and the impossible derived non-reasoning count is left unknown.

recorded_provider_usage_cost. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

{"Kimi_finish_rejected_response_usage_included_but_tool_bodies_unlogged":10,"Kimi_noncompletion_reply_usage_included_body_unlogged":1,"thinking_exceeds_generated":1}

max 28.51% 4825.1s $5.6247 Recorded usage
16 Kimi K3svc_145_4091f6 Cost / task: $4.5173 Recorded usage
Source details

Run: svc_145_4091f6

Completed September 19 resume of evaluation 145; replaces its historical billing-affected snapshot for comparison. Only the 2 authorized billing-interrupted task/seed pairs were rerun; 50 original results unchanged. All 52 final tasks complete, with recorded cost and no HTTP 402 errors. Original frozen plan and submission unchanged. Prior interrupted attempts retained separately and overlap historical archives; do not sum snapshots as independent evaluations. Completed high-effort resume; the earlier billing-interrupted snapshot remains in the downloadable CSV.

recorded_provider_usage_cost. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

{"Kimi_finish_rejected_response_usage_included_but_tool_bodies_unlogged":5}

high 22.60% 3733.9s $4.5173 Recorded usage
17 Kimi K3svc_146_518434 Cost / task: $2.0880 Recorded usage
Source details

Run: svc_146_518434

Resumed after earlier interruptions; final archived event is run_done at 2026-09-13T12:11:51.227606+00:00, with all 52 planned tasks scored and timed. Earlier run_failed events are retained as history, not the final outcome.

recorded_provider_usage_cost. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

{"Kimi_finish_rejected_response_usage_included_but_tool_bodies_unlogged":1}

low 19.19% 1860.9s $2.0880 Recorded usage
18 MiniMax M3svc_78_35a939 Cost / task: Usage not logged API $0.30 in / $1.20 out per 1M tokens · standard · ≤512K context
Source details

Run: svc_78_35a939

Standard API rates through 512K context; above 512K, input/output/cache rates double. Priority pricing differs. Archived responses do not retain token usage, so unit prices cannot yield a per-task total.

{"output_truncated_at_400_characters_and_usage_not_logged":52}

thinking_on 4.10% 1673.2s Usage not logged API $0.30 in / $1.20 out per 1M tokens · standard · ≤512K context
19 MiniMax M3svc_77_abf121 Cost / task: Usage not logged API $0.30 in / $1.20 out per 1M tokens · standard · ≤512K context
Source details

Run: svc_77_abf121

Standard API rates through 512K context; above 512K, input/output/cache rates double. Priority pricing differs. Archived responses do not retain token usage, so unit prices cannot yield a per-task total.

{"output_truncated_at_400_characters_and_usage_not_logged":52}

thinking_off 3.24% 1446.9s Usage not logged API $0.30 in / $1.20 out per 1M tokens · standard · ≤512K context
20 GLM-5V Turbosvc_88_3a6bc6 Cost / task: $0.7395 API estimate · no cache Cache range: $0.2393–$0.7395
Source details

Run: svc_88_3a6bc6

Publisher-designated corrected thinking-off/Pillow-fixed rerun; Pillow 11.3.0 installation is present in init.log. Recorded dollar cost and reasoning/cache token breakdown are unavailable, not zero. Separate reference-cost bounds use the cited OpenRouter rate snapshot and assume all versus no input cached; they are not billing.

API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges

thinking_off 1.53% 1979.4s $0.7395 API estimate · no cache Cache range: $0.2393–$0.7395
21 GLM-5V Turbosvc_79_c47638 Cost / task: $0.7250 API estimate · no cache Cache range: $0.2358–$0.7250
Source details

Run: svc_79_c47638

API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges

thinking_on 1.45% 2641.9s $0.7250 API estimate · no cache Cache range: $0.2358–$0.7250

All included results for this dataset are shown together. Execution settings can differ; source details are preserved with each run.

Data sources & metric definitions

Latest source inventory: 2026-09-20T20:08:33Z. Download all 146 source rows · Download displayed costs & assumptions · View dataset catalog

Times use the task clock, excluding setup and final verification. Costs are recorded usage or labeled estimates, not invoices. API estimates identify their pricing and cache assumptions in each result's source details. Where shown, the cache range runs from all input cached to none. Standard-price proxies do not claim the archived tier or version price. Partial-coverage estimates show their task counts. Additional pricing sources & assumptions are preserved alongside the original CSV references. Runs without token usage show unit prices but stay out of per-task cost plots. The full-OSWorld human baseline is not transferred to these dataset views.

14 retry-attempt or overlapping summary rows for this dataset remain in the CSV, not counted as independent evaluations. Excluded evaluations are listed below the results.

Model labels use publisher-sourced model metadata and explicit aliases, verified 2026-09-19. Open means downloadable weights, not necessarily an open-source license. Release dates mean first public availability of the identified model version. Unversioned archives with an unverified release date are labeled explicitly and omitted only from the release-date curve.

Steps count action batches sent to the environment: one step call counts once, regardless of the number of actions in its batch. Tool calls count individual recorded environment actions within those batches, including waits; a batch of five actions counts as five, not one. This is an environment-action count, not a count of model API tool-call wrappers. Separate observation and task-completion requests are not included. Observations and model responses are separate counts; steps are not observation count minus one. Output tokens per model response divide total generated tokens (including reasoning) by recorded responses, only when both cover every task and the usage/response counts agree.