Results
Compare agents on the four task subsets used in the paper.
OSWorld 2.0 · 52 tasks
21 results · 52 tasks per result
Filter models and reasoning effort
Results
All included results, ordered by average performance.
| # | Model | Effort | Performance | Time / task | Cost / task | |
|---|---|---|---|---|---|---|
| 1 |
GPT-6 Astrasvc_119_69f289
Cost / task:
$8.1746
Logged estimate
Source detailsRun: svc_119_69f289 Closed weights · Released 2026-09-03Cost is short-context API-equivalent usage, not a subscription payment; full tool/response trace unavailable in CLI export. Astra_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges {"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":52,"codex_model_response_count_and_full_tool_calls_unavailable":52} |
xhigh | 76.90% | 1237.3s | $8.1746 Logged estimate | |
| 2 |
GPT-6 Astrasvc_73_9d68f5
Cost / task:
$7.3909
Logged estimate
Source detailsRun: svc_73_9d68f5 Closed weights · Released 2026-09-03Cost is short-context API-equivalent usage, not a subscription payment; full tool/response trace unavailable in CLI export. Astra_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges {"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":52,"codex_model_response_count_and_full_tool_calls_unavailable":52} |
high | 75.02% | 914.0s | $7.3909 Logged estimate | |
| 3 |
GPT-6 Astrasvc_123_d9dc75
Cost / task:
$6.3363
Logged estimate
Source detailsRun: svc_123_d9dc75 Closed weights · Released 2026-09-03Cost is short-context API-equivalent usage, not a subscription payment; full tool/response trace unavailable in CLI export. Astra_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges {"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":52,"codex_model_response_count_and_full_tool_calls_unavailable":52} |
low | 68.17% | 904.2s | $6.3363 Logged estimate | |
| 4 |
GPT-6 Astrasvc_74_19ec78
Cost / task:
$6.5192
Logged estimate
Source detailsRun: svc_74_19ec78 Closed weights · Released 2026-09-03Cost is short-context API-equivalent usage, not a subscription payment; full tool/response trace unavailable in CLI export. Astra_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges {"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":52,"codex_model_response_count_and_full_tool_calls_unavailable":52} |
medium | 67.18% | 828.1s | $6.5192 Logged estimate | |
| 5 |
Gemini 3.8 Flashsvc_55_ffdb2e
Cost / task:
$5.0195
Logged estimate
Source detailsRun: svc_55_ffdb2e Closed weights · Released 2026-09-02recorded_template_price_estimate. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
high | 60.18% | 3058.8s | $5.0195 Logged estimate | |
| 6 |
Gemini 3.8 Flashsvc_54_a54e88
Cost / task:
$2.9148
Logged estimate
Source detailsRun: svc_54_a54e88 Closed weights · Released 2026-09-02recorded_template_price_estimate. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
medium | 59.93% | 2051.8s | $2.9148 Logged estimate | |
| 7 |
Claude Opus 5svc_57_addd34
Cost / task:
$10.0063
Logged estimate
Source detailsRun: svc_57_addd34 Closed weights · Released 2026-07-24Claude_standard_global_list_price_estimate_from_logged_usage. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
low | 52.06% | 1385.7s | $10.0063 Logged estimate | |
| 8 |
Claude Sonnet 5svc_58_909a13
Cost / task:
$8.8630
Logged estimate
Source detailsRun: svc_58_909a13 Closed weights · Released 2026-06-30Claude_standard_global_list_price_estimate_from_logged_usage. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
low | 39.08% | 1993.6s | $8.8630 Logged estimate | |
| 9 |
Gemini 3.8 Flashsvc_56_08ebc5
Cost / task:
$1.8752
Logged estimate
Source detailsRun: svc_56_08ebc5 Closed weights · Released 2026-09-02recorded_template_price_estimate. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
low | 35.29% | 1386.7s | $1.8752 Logged estimate | |
| 10 |
Meta Muse Spark 1.3svc_80_6fea36
Cost / task:
$0.3903
API estimate · no cache
Cache range: $0.0110–$0.3903
Source detailsRun: svc_80_6fea36 Closed weights · Released 2026-09-02API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges |
minimal | 33.97% | 1516.0s | $0.3903 API estimate · no cache Cache range: $0.0110–$0.3903 | |
| 11 |
Meta Muse Spark 1.3svc_81_e35f89
Cost / task:
$0.5029
API estimate · no cache
Cache range: $0.0150–$0.5029
Source detailsRun: svc_81_e35f89 Closed weights · Released 2026-09-02API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges |
low | 33.68% | 2280.8s | $0.5029 API estimate · no cache Cache range: $0.0150–$0.5029 | |
| 12 |
Meta Muse Spark 1.3svc_82_a48586
Cost / task:
$0.5883
API estimate · no cache
Cache range: $0.0198–$0.5883
Source detailsRun: svc_82_a48586 Closed weights · Released 2026-09-02Resumed after earlier interruptions; final archived event is run_done at 2026-09-09T19:09:11.302490+00:00, with all 52 planned tasks scored and timed. Earlier run_failed events are retained as history, not the final outcome. API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges |
medium | 32.31% | 2292.1s | $0.5883 API estimate · no cache Cache range: $0.0198–$0.5883 | |
| 13 |
Meta Muse Spark 1.3svc_84_e2ef21
Cost / task:
$0.6399
API estimate · no cache
Cache range: $0.0231–$0.6399
Source detailsRun: svc_84_e2ef21 Closed weights · Released 2026-09-02Resumed after earlier interruptions; final archived event is run_done at 2026-09-10T01:35:24.715393+00:00, with all 52 planned tasks scored and timed. Earlier run_failed events are retained as history, not the final outcome. API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges |
xhigh | 32.16% | 2261.0s | $0.6399 API estimate · no cache Cache range: $0.0231–$0.6399 | |
| 14 |
Meta Muse Spark 1.3svc_83_03ee99
Cost / task:
$0.5985
API estimate · no cache
Cache range: $0.0206–$0.5985
Source detailsRun: svc_83_03ee99 Closed weights · Released 2026-09-02Resumed after earlier interruptions; final archived event is run_done at 2026-09-09T21:41:29.284094+00:00, with all 52 planned tasks scored and timed. Earlier run_failed events are retained as history, not the final outcome. API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges |
high | 31.46% | 2179.3s | $0.5985 API estimate · no cache Cache range: $0.0206–$0.5985 | |
| 15 |
Kimi K3svc_147_d5291e
Cost / task:
$5.6247
Recorded usage
Source detailsRun: svc_147_d5291e Open weights · Released 2026-07-16Completed September 19 resume of evaluation 147; replaces its historical billing-affected snapshot for comparison. Only the 15 authorized billing-interrupted task/seed pairs were rerun; 37 original results unchanged. All 52 final tasks complete, with recorded cost and no HTTP 402 errors. Original frozen plan and submission unchanged. Prior interrupted attempts retained separately and overlap historical archives; do not sum snapshots as independent evaluations. Completed max-effort resume; the earlier billing-interrupted snapshot remains in the downloadable CSV. One final task reports more reasoning tokens than total output tokens; raw provider counters are retained and the impossible derived non-reasoning count is left unknown. recorded_provider_usage_cost. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges {"Kimi_finish_rejected_response_usage_included_but_tool_bodies_unlogged":10,"Kimi_noncompletion_reply_usage_included_body_unlogged":1,"thinking_exceeds_generated":1} |
max | 28.51% | 4825.1s | $5.6247 Recorded usage | |
| 16 |
Kimi K3svc_145_4091f6
Cost / task:
$4.5173
Recorded usage
Source detailsRun: svc_145_4091f6 Open weights · Released 2026-07-16Completed September 19 resume of evaluation 145; replaces its historical billing-affected snapshot for comparison. Only the 2 authorized billing-interrupted task/seed pairs were rerun; 50 original results unchanged. All 52 final tasks complete, with recorded cost and no HTTP 402 errors. Original frozen plan and submission unchanged. Prior interrupted attempts retained separately and overlap historical archives; do not sum snapshots as independent evaluations. Completed high-effort resume; the earlier billing-interrupted snapshot remains in the downloadable CSV. recorded_provider_usage_cost. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges {"Kimi_finish_rejected_response_usage_included_but_tool_bodies_unlogged":5} |
high | 22.60% | 3733.9s | $4.5173 Recorded usage | |
| 17 |
Kimi K3svc_146_518434
Cost / task:
$2.0880
Recorded usage
Source detailsRun: svc_146_518434 Open weights · Released 2026-07-16Resumed after earlier interruptions; final archived event is run_done at 2026-09-13T12:11:51.227606+00:00, with all 52 planned tasks scored and timed. Earlier run_failed events are retained as history, not the final outcome. recorded_provider_usage_cost. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges {"Kimi_finish_rejected_response_usage_included_but_tool_bodies_unlogged":1} |
low | 19.19% | 1860.9s | $2.0880 Recorded usage | |
| 18 |
MiniMax M3svc_78_35a939
Cost / task:
Usage not logged
API $0.30 in / $1.20 out
per 1M tokens · standard · ≤512K context
Source detailsRun: svc_78_35a939 Open weights · Released 2026-06-01Standard API rates through 512K context; above 512K, input/output/cache rates double. Priority pricing differs. Archived responses do not retain token usage, so unit prices cannot yield a per-task total. {"output_truncated_at_400_characters_and_usage_not_logged":52} |
thinking_on | 4.10% | 1673.2s | Usage not logged API $0.30 in / $1.20 out per 1M tokens · standard · ≤512K context | |
| 19 |
MiniMax M3svc_77_abf121
Cost / task:
Usage not logged
API $0.30 in / $1.20 out
per 1M tokens · standard · ≤512K context
Source detailsRun: svc_77_abf121 Open weights · Released 2026-06-01Standard API rates through 512K context; above 512K, input/output/cache rates double. Priority pricing differs. Archived responses do not retain token usage, so unit prices cannot yield a per-task total. {"output_truncated_at_400_characters_and_usage_not_logged":52} |
thinking_off | 3.24% | 1446.9s | Usage not logged API $0.30 in / $1.20 out per 1M tokens · standard · ≤512K context | |
| 20 |
GLM-5V Turbosvc_88_3a6bc6
Cost / task:
$0.7395
API estimate · no cache
Cache range: $0.2393–$0.7395
Source detailsRun: svc_88_3a6bc6 Closed weights · Released 2026-04-01Publisher-designated corrected thinking-off/Pillow-fixed rerun; Pillow 11.3.0 installation is present in init.log. Recorded dollar cost and reasoning/cache token breakdown are unavailable, not zero. Separate reference-cost bounds use the cited OpenRouter rate snapshot and assume all versus no input cached; they are not billing. API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges |
thinking_off | 1.53% | 1979.4s | $0.7395 API estimate · no cache Cache range: $0.2393–$0.7395 | |
| 21 |
GLM-5V Turbosvc_79_c47638
Cost / task:
$0.7250
API estimate · no cache
Cache range: $0.2358–$0.7250
Source detailsRun: svc_79_c47638 Closed weights · Released 2026-04-01API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges |
thinking_on | 1.45% | 2641.9s | $0.7250 API estimate · no cache Cache range: $0.2358–$0.7250 |
Task performance
Select a measure to compare. Tap a point for its model, score, and timing.
Scatter plot comparing model performance with the selected horizontal metric.
Performance, time, and cost
Drag to rotate; select a point to inspect it. Only results with known costs appear here.
Rotatable three-dimensional scatter plot comparing cost, time, and performance.
All included results for this dataset are shown together. Execution settings can differ; source details are preserved with each run.
Data sources & metric definitions
Latest source inventory: 2026-09-20T20:08:33Z. Download all 146 source rows · Download displayed costs & assumptions · View dataset catalog
Times use the task clock, excluding setup and final verification. Costs are recorded usage or labeled estimates, not invoices. API estimates identify their pricing and cache assumptions in each result's source details. Where shown, the cache range runs from all input cached to none. Standard-price proxies do not claim the archived tier or version price. Partial-coverage estimates show their task counts. Additional pricing sources & assumptions are preserved alongside the original CSV references. Runs without token usage show unit prices but stay out of per-task cost plots. The full-OSWorld human baseline is not transferred to these dataset views.
14 retry-attempt or overlapping summary rows for this dataset remain in the CSV, not counted as independent evaluations. Excluded evaluations are listed below the results.
Model labels use publisher-sourced model metadata and explicit aliases, verified 2026-09-19. Open means downloadable weights, not necessarily an open-source license. Release dates mean first public availability of the identified model version. Unversioned archives with an unverified release date are labeled explicitly and omitted only from the release-date curve.
Steps count action batches sent to the environment: one step call counts once, regardless of the number of actions in its batch. Tool calls count individual recorded environment actions within those batches, including waits; a batch of five actions counts as five, not one. This is an environment-action count, not a count of model API tool-call wrappers. Separate observation and task-completion requests are not included. Observations and model responses are separate counts; steps are not observation count minus one. Output tokens per model response divide total generated tokens (including reasoning) by recorded responses, only when both cover every task and the usage/response counts agree.