Results
Compare agents on the four task subsets used in the paper.
OSWorld · 50 tasks
58 results · 50 tasks per result · 8 excluded evaluations
Filter models and reasoning effort
Results
All included results, ordered by average performance.
| # | Model | Effort | Performance | Time / task | Cost / task | |
|---|---|---|---|---|---|---|
| 1 |
GPT-6 Astra · Normal I/Osvc_121_b4bca3
Cost / task:
$0.7081
Logged estimate
Source detailsRun: svc_121_b4bca3 Closed weights · Released 2026-09-03Cost is short-context API-equivalent usage, not a subscription payment; full tool/response trace unavailable in CLI export. Astra_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges {"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":50,"codex_model_response_count_and_full_tool_calls_unavailable":50} |
xhigh | 91.62% | 126.8s | $0.7081 Logged estimate | |
| 2 |
Gemini 3.8 Flashsvc_59_708d1b
Cost / task:
$0.1080
Logged estimate
Source detailsRun: svc_59_708d1b Closed weights · Released 2026-09-02New no-preload Energy50 evaluation; preserve its distinct track and measurement contract when comparing earlier runs. Cost is the logged template estimate from final cumulative snapshots, not an invoice. recorded_template_price_estimate. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
low | 91.62% | 127.1s | $0.1080 Logged estimate | |
| 3 |
Claude Opus 5svc_31_878995
Cost / task:
$0.6611
Logged estimate
Source detailsRun: svc_31_878995 Closed weights · Released 2026-07-24Claude_standard_global_list_price_estimate_from_logged_usage. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
high | 91.62% | 136.0s | $0.6611 Logged estimate | |
| 4 |
GPT-6 Astra · Single actionsvc_157_c2dbaf
Cost / task:
$1.3637
Logged estimate
Source detailsRun: svc_157_c2dbaf Closed weights · Released 2026-09-03Completed local evaluation 157; all 50 tasks. Single-action mode enforces one GUI action followed by a fresh screenshot before another action. Same model, effort and configured limits as the normal-I/O xhigh baseline; owner-approved current runtime and Codex subscription. Cost is the same historical standard API-equivalent estimate from logged usage, not subscription spending. The source link contains derived task metrics and source hashes; the trajectory archive has not been published to Hugging Face. One GUI action per request, with a fresh screenshot required before the next action. Astra_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges {"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":50,"codex_model_response_count_and_full_tool_calls_unavailable":50} |
xhigh | 91.62% | 182.7s | $1.3637 Logged estimate | |
| 5 |
Gemini 3.8 Flashsvc_61_1ec4be
Cost / task:
$0.2209
Logged estimate
Source detailsRun: svc_61_1ec4be Closed weights · Released 2026-09-02New no-preload Energy50 evaluation; preserve its distinct track and measurement contract when comparing earlier runs. Cost is the logged template estimate from final cumulative snapshots, not an invoice. recorded_template_price_estimate. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
medium | 91.62% | 250.1s | $0.2209 Logged estimate | |
| 6 |
GPT-6 Astra · Normal I/Osvc_151_181620
Cost / task:
$0.6001
Logged estimate
Source detailsRun: svc_151_181620 Closed weights · Released 2026-09-03Added 2026-09-14 at owner request. Separate no-preload track/measurement contract; not a same-track ranking against older preload runs. Cost uses the prior September11 short-context API-equivalent pricing snapshot, not subscription billing. Full CLI model-response/tool trace remains unavailable. Final-task metrics exclude retained non-final attempts. Astra_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges {"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":50,"codex_model_response_count_and_full_tool_calls_unavailable":50} |
medium | 89.62% | 90.2s | $0.6001 Logged estimate | |
| 7 |
GPT-6 Astra · Normal I/Osvc_152_8d02f4
Cost / task:
$0.6364
Logged estimate
Source detailsRun: svc_152_8d02f4 Closed weights · Released 2026-09-03Added 2026-09-14 at owner request. Separate no-preload track/measurement contract; not a same-track ranking against older preload runs. Cost uses the prior September11 short-context API-equivalent pricing snapshot, not subscription billing. Full CLI model-response/tool trace remains unavailable. Final-task metrics exclude retained non-final attempts. One retained infrastructure attempt is represented separately; final-task API-equivalent total=31.818674 USD, extra attempt=0.263074 USD, all recorded sessions=32.081748 USD. Astra_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges {"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":50,"codex_model_response_count_and_full_tool_calls_unavailable":50} |
high | 89.62% | 100.8s | $0.6364 Logged estimate | |
| 8 |
GPT-6 Astra · Fast I/Osvc_127_422895
Cost / task:
$0.7367
Logged estimate
Source detailsRun: svc_127_422895 Closed weights · Released 2026-09-03Cost is short-context API-equivalent usage, not a subscription payment; full tool/response trace unavailable in CLI export. Astra_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges {"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":50,"codex_model_response_count_and_full_tool_calls_unavailable":50} |
low | 89.62% | 105.2s | $0.7367 Logged estimate | |
| 9 |
Claude Sonnet 5svc_24_ddcb3b
Cost / task:
$0.3086
Logged estimate
Source detailsRun: svc_24_ddcb3b Closed weights · Released 2026-06-30Claude_standard_global_list_price_estimate_from_logged_usage. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
high | 89.62% | 140.8s | $0.3086 Logged estimate | |
| 10 |
Gemini 3.7 Flashsvc_85_948d7f
Cost / task:
$0.9437
Logged estimate
Source detailsRun: svc_85_948d7f Closed weights · Released 2026-08-13sum_of_rounded_per_response_template_estimates. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
medium | 89.62% | 198.9s | $0.9437 Logged estimate | |
| 11 |
Gemini 3.7 Flashsvc_86_9e7a56
Cost / task:
$1.4121
Logged estimate
Source detailsRun: svc_86_9e7a56 Closed weights · Released 2026-08-13sum_of_rounded_per_response_template_estimates. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
high | 89.62% | 254.3s | $1.4121 Logged estimate | |
| 12 |
Kimi K3 · Batched tool callssvc_141_f28fed
Cost / task:
$0.5034
Recorded usage
Source detailsRun: svc_141_f28fed Open weights · Released 2026-07-16recorded_provider_usage_cost. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
max | 89.62% | 443.7s | $0.5034 Recorded usage | |
| 13 |
Claude Opus 5svc_32_909005
Cost / task:
$0.3896
Logged estimate
Source detailsRun: svc_32_909005 Closed weights · Released 2026-07-24Claude_standard_global_list_price_estimate_from_logged_usage. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges {"Claude_cost_assumes_standard_global_for_responses_with_missing_tier_geo":1} |
low | 87.62% | 85.9s | $0.3896 Logged estimate | |
| 14 |
GPT-6 Astra · Normal I/Osvc_122_93c9c5
Cost / task:
$0.5665
Logged estimate
Source detailsRun: svc_122_93c9c5 Closed weights · Released 2026-09-03Cost is short-context API-equivalent usage, not a subscription payment; full tool/response trace unavailable in CLI export. Astra_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges {"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":50,"codex_model_response_count_and_full_tool_calls_unavailable":50} |
low | 87.62% | 86.1s | $0.5665 Logged estimate | |
| 15 |
Gemini 3.7 Flashsvc_87_735f9a
Cost / task:
$0.9264
Logged estimate
Source detailsRun: svc_87_735f9a Closed weights · Released 2026-08-13sum_of_rounded_per_response_template_estimates. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
low | 87.62% | 207.3s | $0.9264 Logged estimate | |
| 16 |
Gemini 3.8 Flashsvc_60_8db123
Cost / task:
$0.3475
Logged estimate
Source detailsRun: svc_60_8db123 Closed weights · Released 2026-09-02New no-preload Energy50 evaluation; preserve its distinct track and measurement contract when comparing earlier runs. Cost is the logged template estimate from final cumulative snapshots, not an invoice. recorded_template_price_estimate. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
high | 87.62% | 347.7s | $0.3475 Logged estimate | |
| 17 |
Kimi K3 · Single tool callsvc_56_9e002a
Cost / task:
$0.7158
Recorded usage
Source detailsRun: svc_56_9e002a Open weights · Released 2026-07-16Published Energy50 max-reasoning subset of Kimi evaluation56; not a separate execution. One model tool call per response; the code in that call can contain multiple GUI actions. recorded_provider_usage_cost. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges {"Kimi_finish_rejected_response_usage_included_but_tool_bodies_unlogged":1,"Kimi_noncompletion_replies_have_no_logged_usage_or_body":2} |
max | 85.62% | 606.3s | $0.7158 Recorded usage | |
| 18 |
GPT-5.6 Luna · Direct APIsvc_20_9456b8
Cost / task:
$0.0256
Logged estimate
Source detailsRun: svc_20_9456b8 Closed weights · Released 2026-07-09sum_of_rounded_per_response_template_estimates. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
medium | 81.62% | 100.6s | $0.0256 Logged estimate | |
| 19 |
GPT-5.6 Solsvc_23_753ef9
Cost / task:
$0.5267
Logged estimate
Source detailsRun: svc_23_753ef9 Closed weights · Released 2026-07-09sum_of_rounded_per_response_template_estimates. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
xhigh | 81.62% | 108.3s | $0.5267 Logged estimate | |
| 20 |
Kimi K3 · Single tool callsvc_90_357e81
Cost / task:
$0.4141
Recorded usage
Source detailsRun: svc_90_357e81 Open weights · Released 2026-07-16recorded_provider_usage_cost. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
high | 81.62% | 305.1s | $0.4141 Recorded usage | |
| 21 |
GPT-5.6 Solsvc_22_3a244d
Cost / task:
$0.5394
Logged estimate
Source detailsRun: svc_22_3a244d Closed weights · Released 2026-07-09sum_of_rounded_per_response_template_estimates. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
medium | 79.62% | 121.3s | $0.5394 Logged estimate | |
| 22 |
Yutori n2svc_13_01885e
Cost / task:
$0.1036
API estimate · logged usage
Source detailsRun: svc_13_01885e Closed weights · Released 2026-08-26Published HF PR23, merged before inventory. All 50 final tasks complete; original no-preload contract and GUI-only tool configuration retained. Full response usage recovered from agent.stdout despite the published README marking usage unavailable; duplicated SDK step usage is not counted again. Cost is an API-rate estimate using logged input/output and billed cached input at publisher rates verified September 19, not a recorded charge. Reasoning-token breakdown is incomplete and remains unknown. published_API_rate_estimate_from_logged_usage_and_billed_cache_not_recorded_charge. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges {"reasoning_token_breakdown_not_logged_for_all_responses":50} |
xhigh | 79.62% | 225.0s | $0.1036 API estimate · logged usage | |
| 23 |
GPT-5.6 Luna · Codexsvc_155_7fbf62
Cost / task:
$0.0759
Logged estimate
Source detailsRun: svc_155_7fbf62 Closed weights · Released 2026-07-09Added 2026-09-16 at owner request. GPT-5.6 Luna via Codex, not the direct-API gpt54 template. All 50 original tasks included, including failures. No rerun or regrading. No-preload track/measurement contract; not a same-track ranking against older preload runs. Costs are standard short-context API-equivalent estimates using official 2026-09-16 Luna rates, not subscription charges. All 50 session token summaries complete; cached input is a subset of input, reasoning is a subset of output, recorded cache-write counters are zero. Full CLI model-response/tool trace remains unavailable; observed CLI tool items and harness steps are separate metrics. No additional retained attempts found. Submitted agent.py SHA-256=306fe8c02f9c85a5ff90be95d8efef897a3555ab37fe4f46e84dad68dad0058a. Full per-task normalized evidence and usage breakdowns are linked in companion_cost_usage_url. Luna_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges; assumes standard short-context rates; per-request long-context/fast-mode adjustments and tool charges cannot be reconstructed from session totals; output already includes reasoning {"Luna_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":50,"codex_model_response_count_and_full_tool_calls_unavailable":50} |
high | 79.62% | 324.8s | $0.0759 Logged estimate | |
| 24 |
Kimi K3 · Batched tool callssvc_144_1f26e4
Cost / task:
$0.3823
Recorded usage
Source detailsRun: svc_144_1f26e4 Open weights · Released 2026-07-16recorded_provider_usage_cost. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
high | 79.62% | 358.1s | $0.3823 Recorded usage | |
| 25 |
Meta Muse Spark 1.1svc_69_41f88a
Cost / task:
$1.6849
Standard-price proxy · no cache
Cache range: $0.2422–$1.6849
Source detailsRun: svc_69_41f88a Closed weights · Released 2026-07-09Standard API price proxy, not the archived contributor-tier charge; historical contributor pricing is unverified. Recorded input/output tokens; no input caching assumed. Rates verified 2026-09-16; excludes other charges. |
xhigh | 79.62% | 417.2s | $1.6849 Standard-price proxy · no cache Cache range: $0.2422–$1.6849 | |
| 26 |
GPT-5.6 Luna · Codexsvc_156_2bc119
Cost / task:
$0.0631
Logged estimate
Source detailsRun: svc_156_2bc119 Closed weights · Released 2026-07-09Added 2026-09-16 at owner request. GPT-5.6 Luna via Codex, not the direct-API gpt54 template. All 50 original tasks included, including failures. No rerun or regrading. No-preload track/measurement contract; not a same-track ranking against older preload runs. Costs are standard short-context API-equivalent estimates using official 2026-09-16 Luna rates, not subscription charges. All 50 session token summaries complete; cached input is a subset of input, reasoning is a subset of output, recorded cache-write counters are zero. Full CLI model-response/tool trace remains unavailable; observed CLI tool items and harness steps are separate metrics. No additional retained attempts found. Submitted agent.py SHA-256=306fe8c02f9c85a5ff90be95d8efef897a3555ab37fe4f46e84dad68dad0058a. Full per-task normalized evidence and usage breakdowns are linked in companion_cost_usage_url. Luna_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges; assumes standard short-context rates; per-request long-context/fast-mode adjustments and tool charges cannot be reconstructed from session totals; output already includes reasoning {"Luna_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":50,"codex_model_response_count_and_full_tool_calls_unavailable":50} |
xhigh | 77.62% | 297.2s | $0.0631 Logged estimate | |
| 27 |
Yutori n2svc_11_0dddb1
Cost / task:
$0.1569
API estimate · logged usage
Source detailsRun: svc_11_0dddb1 Closed weights · Released 2026-08-26Published HF PR23, merged before inventory. All 50 final tasks complete; original no-preload contract and GUI-only tool configuration retained. Full response usage recovered from agent.stdout despite the published README marking usage unavailable; duplicated SDK step usage is not counted again. Cost is an API-rate estimate using logged input/output and billed cached input at publisher rates verified September 19, not a recorded charge. Reasoning-token breakdown is incomplete and remains unknown. published_API_rate_estimate_from_logged_usage_and_billed_cache_not_recorded_charge. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges {"reasoning_token_breakdown_not_logged_for_all_responses":50} |
low | 77.62% | 331.0s | $0.1569 API estimate · logged usage | |
| 28 |
Meta Muse Spark 1.3svc_68 + rerun3_xhigh13
Cost / task:
$0.1607
API estimate · no cache
Cache range: $0.0055–$0.1607
Source detailsRun: svc_68 + rerun3_xhigh13 Closed weights · Released 2026-09-02Published STITCH-musespark13-xhigh-68.txt: 47 original tasks plus three rerun replacements; overlaps both components and is not an independent run. Published 664.79 seconds/task is mean(agent_wall_sec + env_boot_sec)=664.78846; timed task clock=612.07638625182 seconds/task. Published as one combined 50-task result. Original tasks ran at concurrency 8; the three recovery tasks ran at concurrency 3. Both source archives are retained below. API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges svc_68_bca75f · rerun3_xhigh13 {"result_json_absent_metrics_reduced_from_original_runlog_events":47} |
xhigh | 77.62% | 612.1s | $0.1607 API estimate · no cache Cache range: $0.0055–$0.1607 | |
| 29 |
Meta Muse Spark 1.3svc_65_46a645
Cost / task:
$0.1258
API estimate · no cache
Cache range: $0.0043–$0.1258
Source detailsRun: svc_65_46a645 Closed weights · Released 2026-09-02API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges |
medium | 77.54% | 397.6s | $0.1258 API estimate · no cache Cache range: $0.0043–$0.1258 | |
| 30 |
Claude Sonnet 5svc_25_adbcee
Cost / task:
$0.2941
Logged estimate
Source detailsRun: svc_25_adbcee Closed weights · Released 2026-06-30Claude_standard_global_list_price_estimate_from_logged_usage. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
low | 75.81% | 129.4s | $0.2941 Logged estimate | |
| 31 |
GPT-5.6 Luna · Codexsvc_153_795ea7
Cost / task:
$0.0254
Logged estimate
Source detailsRun: svc_153_795ea7 Closed weights · Released 2026-07-09Added 2026-09-16 at owner request. GPT-5.6 Luna via Codex, not the direct-API gpt54 template. All 50 original tasks included, including failures. No rerun or regrading. No-preload track/measurement contract; not a same-track ranking against older preload runs. Costs are standard short-context API-equivalent estimates using official 2026-09-16 Luna rates, not subscription charges. All 50 session token summaries complete; cached input is a subset of input, reasoning is a subset of output, recorded cache-write counters are zero. Full CLI model-response/tool trace remains unavailable; observed CLI tool items and harness steps are separate metrics. No additional retained attempts found. Submitted agent.py SHA-256=306fe8c02f9c85a5ff90be95d8efef897a3555ab37fe4f46e84dad68dad0058a. Full per-task normalized evidence and usage breakdowns are linked in companion_cost_usage_url. Luna_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges; assumes standard short-context rates; per-request long-context/fast-mode adjustments and tool charges cannot be reconstructed from session totals; output already includes reasoning {"Luna_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":50,"codex_model_response_count_and_full_tool_calls_unavailable":50} |
low | 75.62% | 144.8s | $0.0254 Logged estimate | |
| 32 |
MiniMax M3svc_51_121311
Cost / task:
Usage not logged
API $0.30 in / $1.20 out
per 1M tokens · standard · ≤512K context
Source detailsRun: svc_51_121311 Open weights · Released 2026-06-01Standard API rates through 512K context; above 512K, input/output/cache rates double. Priority pricing differs. Archived responses do not retain token usage, so unit prices cannot yield a per-task total. {"output_truncated_at_400_characters_and_usage_not_logged":50} |
thinking_off | 75.62% | 253.8s | Usage not logged API $0.30 in / $1.20 out per 1M tokens · standard · ≤512K context | |
| 33 |
Yutori n2svc_10_cd0ed2
Cost / task:
$0.0811
API estimate · logged usage
Source detailsRun: svc_10_cd0ed2 Closed weights · Released 2026-08-26Published HF PR23, merged before inventory. All 50 final tasks complete; original no-preload contract and GUI-only tool configuration retained. Full response usage recovered from agent.stdout despite the published README marking usage unavailable; duplicated SDK step usage is not counted again. Cost is an API-rate estimate using logged input/output and billed cached input at publisher rates verified September 19, not a recorded charge. Reasoning-token breakdown is incomplete and remains unknown. published_API_rate_estimate_from_logged_usage_and_billed_cache_not_recorded_charge. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges {"reasoning_token_breakdown_not_logged_for_all_responses":50} |
none | 75.62% | 298.5s | $0.0811 API estimate · logged usage | |
| 34 |
GPT-5.6 Luna · Direct APIsvc_21_00be0a
Cost / task:
$0.0386
Logged estimate
Source detailsRun: svc_21_00be0a Closed weights · Released 2026-07-09sum_of_rounded_per_response_template_estimates. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
xhigh | 73.62% | 135.3s | $0.0386 Logged estimate | |
| 35 |
Kimi K3 · Batched tool callssvc_143_e8d33d
Cost / task:
$0.1740
Recorded usage
Source detailsRun: svc_143_e8d33d Open weights · Released 2026-07-16recorded_provider_usage_cost. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
low | 73.62% | 149.5s | $0.1740 Recorded usage | |
| 36 |
Meta Muse Spark 1.3svc_70_8f6cda
Cost / task:
$0.1479
API estimate · no cache
Cache range: $0.0050–$0.1479
Source detailsRun: svc_70_8f6cda Closed weights · Released 2026-09-02API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges |
high | 73.62% | 506.9s | $0.1479 API estimate · no cache Cache range: $0.0050–$0.1479 | |
| 37 |
Kimi K3 · Single tool callsvc_89_b4886a
Cost / task:
$0.2802
Recorded usage
Source detailsRun: svc_89_b4886a Open weights · Released 2026-07-16recorded_provider_usage_cost. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
low | 71.62% | 176.6s | $0.2802 Recorded usage | |
| 38 |
MiniMax M3svc_52_554434
Cost / task:
Usage not logged
API $0.30 in / $1.20 out
per 1M tokens · standard · ≤512K context
Source detailsRun: svc_52_554434 Open weights · Released 2026-06-01Standard API rates through 512K context; above 512K, input/output/cache rates double. Priority pricing differs. Archived responses do not retain token usage, so unit prices cannot yield a per-task total. {"output_truncated_at_400_characters_and_usage_not_logged":50} |
thinking_on | 69.62% | 345.1s | Usage not logged API $0.30 in / $1.20 out per 1M tokens · standard · ≤512K context | |
| 39 |
Meta Muse Spark 1.1svc_59_d8a6d4
Cost / task:
$1.3785
Standard-price proxy · no cache
Cache range: $0.1929–$1.3785
Source detailsRun: svc_59_d8a6d4 Closed weights · Released 2026-07-09Standard API price proxy, not the archived contributor-tier charge; historical contributor pricing is unverified. Recorded input/output tokens; no input caching assumed. Rates verified 2026-09-16; excludes other charges. |
medium | 69.62% | 359.8s | $1.3785 Standard-price proxy · no cache Cache range: $0.1929–$1.3785 | |
| 40 |
Meta Muse Spark 1.1svc_58_38520a
Cost / task:
$1.7059
Standard-price proxy · no cache
Cache range: $0.2426–$1.7059
Source detailsRun: svc_58_38520a Closed weights · Released 2026-07-09Standard API price proxy, not the archived contributor-tier charge; historical contributor pricing is unverified. Recorded input/output tokens; no input caching assumed. Rates verified 2026-09-16; excludes other charges. |
high | 69.62% | 368.5s | $1.7059 Standard-price proxy · no cache Cache range: $0.2426–$1.7059 | |
| 41 |
Meta Muse Spark 1.1svc_60_72500b
Cost / task:
$1.3116
Standard-price proxy · no cache
Cache range: $0.1818–$1.3116
Source detailsRun: svc_60_72500b Closed weights · Released 2026-07-09Standard API price proxy, not the archived contributor-tier charge; historical contributor pricing is unverified. Recorded input/output tokens; no input caching assumed. Rates verified 2026-09-16; excludes other charges. |
low | 69.62% | 370.4s | $1.3116 Standard-price proxy · no cache Cache range: $0.1818–$1.3116 | |
| 42 |
Meta Muse Spark 1.2svc_39_025394
Cost / task:
$0.2445
API estimate · no cache
Cache range: $0.0083–$0.2445
Source detailsRun: svc_39_025394 Closed weights · Released 2026-08-05API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges |
xhigh | 69.62% | 720.9s | $0.2445 API estimate · no cache Cache range: $0.0083–$0.2445 | |
| 43 |
GPT-5.6 Luna · Codexsvc_154_323b97
Cost / task:
$0.0353
Logged estimate
Source detailsRun: svc_154_323b97 Closed weights · Released 2026-07-09Added 2026-09-16 at owner request. GPT-5.6 Luna via Codex, not the direct-API gpt54 template. All 50 original tasks included, including failures. No rerun or regrading. No-preload track/measurement contract; not a same-track ranking against older preload runs. Costs are standard short-context API-equivalent estimates using official 2026-09-16 Luna rates, not subscription charges. All 50 session token summaries complete; cached input is a subset of input, reasoning is a subset of output, recorded cache-write counters are zero. Full CLI model-response/tool trace remains unavailable; observed CLI tool items and harness steps are separate metrics. No additional retained attempts found. Submitted agent.py SHA-256=306fe8c02f9c85a5ff90be95d8efef897a3555ab37fe4f46e84dad68dad0058a. One original agent_error on osworld_libreoffice_writer_adf5e2c3-64c7-4644-b7b6-d2f0167927e7/seed_1031319744: the environment HTTPS endpoint refused the /done connection and the gateway failed. Its zero score, timing and recorded usage remain included; not a clean failure-free comparison. Full per-task normalized evidence and usage breakdowns are linked in companion_cost_usage_url. Luna_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges; assumes standard short-context rates; per-request long-context/fast-mode adjustments and tool charges cannot be reconstructed from session totals; output already includes reasoning {"Luna_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":50,"codex_model_response_count_and_full_tool_calls_unavailable":50} |
medium | 67.62% | 188.8s | $0.0353 Logged estimate | |
| 44 |
Yutori n2svc_12_026ed7
Cost / task:
$0.1100
API estimate · logged usage
Source detailsRun: svc_12_026ed7 Closed weights · Released 2026-08-26Published HF PR23, merged before inventory. All 50 final tasks complete; original no-preload contract and GUI-only tool configuration retained. Full response usage recovered from agent.stdout despite the published README marking usage unavailable; duplicated SDK step usage is not counted again. Cost is an API-rate estimate using logged input/output and billed cached input at publisher rates verified September 19, not a recorded charge. Reasoning-token breakdown is incomplete and remains unknown. published_API_rate_estimate_from_logged_usage_and_billed_cache_not_recorded_charge. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges {"reasoning_token_breakdown_not_logged_for_all_responses":50} |
medium | 67.62% | 265.4s | $0.1100 API estimate · logged usage | |
| 45 |
Meta Muse Spark 1.1svc_61_96ce3f
Cost / task:
$1.1840
Standard-price proxy · no cache
Cache range: $0.1644–$1.1840
Source detailsRun: svc_61_96ce3f Closed weights · Released 2026-07-09Standard API price proxy, not the archived contributor-tier charge; historical contributor pricing is unverified. Recorded input/output tokens; no input caching assumed. Rates verified 2026-09-16; excludes other charges. |
minimal | 67.62% | 318.5s | $1.1840 Standard-price proxy · no cache Cache range: $0.1644–$1.1840 | |
| 46 |
Meta Muse Spark 1.3svc_66_660b4a
Cost / task:
$0.1071
API estimate · no cache
Cache range: $0.0035–$0.1071
Source detailsRun: svc_66_660b4a Closed weights · Released 2026-09-02API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges |
low | 67.62% | 361.0s | $0.1071 API estimate · no cache Cache range: $0.0035–$0.1071 | |
| 47 |
GPT-5.6 Solsvc_62_9b9e92
Cost / task:
$0.3191
Logged estimate
Source detailsRun: svc_62_9b9e92 Closed weights · Released 2026-07-09Added from JY's HF PR21 at immutable revision 5d04d3b4b91dceb4c90e8592ce3235c1965c9db2; PR open at inventory. All 50 final tasks completed. No-preload track and original run contract retained. Cost is the sum of rounded per-response template estimates from final-task logs, not an invoice. Published-rate reconstruction independently agrees within the saved rounding bound. One prior infrastructure attempt is retained separately; the publisher's $16.43 total includes it, while this final-task row excludes its cost and usage. sum_of_rounded_per_response_template_estimates. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
low | 65.81% | 100.1s | $0.3191 Logged estimate | |
| 48 |
Meta Muse Spark 1.2svc_48_4e20fb
Cost / task:
$0.1054
API estimate · no cache
Cache range: $0.0030–$0.1054
Source detailsRun: svc_48_4e20fb Closed weights · Released 2026-08-05API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges |
low | 63.62% | 363.6s | $0.1054 API estimate · no cache Cache range: $0.0030–$0.1054 | |
| 49 |
Meta Muse Spark 1.2svc_41_6ef6f1
Cost / task:
$0.1365
API estimate · no cache
Cache range: $0.0043–$0.1365
Source detailsRun: svc_41_6ef6f1 Closed weights · Released 2026-08-05API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges |
medium | 63.62% | 466.2s | $0.1365 API estimate · no cache Cache range: $0.0043–$0.1365 | |
| 50 |
GPT-5.6 Luna · Direct APIsvc_63_ccb407
Cost / task:
$0.0107
Logged estimate
Source detailsRun: svc_63_ccb407 Closed weights · Released 2026-07-09Added from JY's HF PR21 at immutable revision 5d04d3b4b91dceb4c90e8592ce3235c1965c9db2; PR open at inventory. All 50 final tasks completed. No-preload track and original run contract retained. Cost is the sum of rounded per-response template estimates from final-task logs, not an invoice. Published-rate reconstruction independently agrees within the saved rounding bound. sum_of_rounded_per_response_template_estimates. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
low | 61.81% | 73.6s | $0.0107 Logged estimate | |
| 51 |
Meta Muse Spark 1.3svc_67_b7d034
Cost / task:
$0.0789
API estimate · no cache
Cache range: $0.0023–$0.0789
Source detailsRun: svc_67_b7d034 Closed weights · Released 2026-09-02API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges |
minimal | 61.62% | 264.4s | $0.0789 API estimate · no cache Cache range: $0.0023–$0.0789 | |
| 52 |
Meta Muse Spark 1.2svc_40_c6a678
Cost / task:
$0.1736
API estimate · no cache
Cache range: $0.0053–$0.1736
Source detailsRun: svc_40_c6a678 Closed weights · Released 2026-08-05API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges |
high | 61.62% | 669.8s | $0.1736 API estimate · no cache Cache range: $0.0053–$0.1736 | |
| 53 |
Gemini 3 Flash Previewsvc_81_c8b73b
Cost / task:
$0.8091
Logged estimate
Source detailsRun: svc_81_c8b73b Closed weights · Released 2025-12-17sum_of_rounded_per_response_template_estimates. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
high | 59.62% | 337.2s | $0.8091 Logged estimate | |
| 54 |
GLM-5V Turbosvc_87_ed8093
Cost / task:
$0.2203
API estimate · no cache
Cache range: $0.0701–$0.2203
Source detailsRun: svc_87_ed8093 Closed weights · Released 2026-04-01Publisher-designated corrected thinking-off/Pillow-fixed rerun; Pillow 11.3.0 installation is present in init.log. Recorded dollar cost and reasoning/cache token breakdown are unavailable, not zero. Separate reference-cost bounds use the cited OpenRouter rate snapshot and assume all versus no input cached; they are not billing. API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges |
thinking_off | 57.81% | 752.4s | $0.2203 API estimate · no cache Cache range: $0.0701–$0.2203 | |
| 55 |
Gemini 3 Flash Previewsvc_83_d83faa
Cost / task:
$0.5923
Logged estimate
Source detailsRun: svc_83_d83faa Closed weights · Released 2025-12-17sum_of_rounded_per_response_template_estimates. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
medium | 57.62% | 268.0s | $0.5923 Logged estimate | |
| 56 |
GLM-5V Turbosvc_49_c74b13
Cost / task:
$0.2350
API estimate · no cache
Cache range: $0.0744–$0.2350
Source detailsRun: svc_49_c74b13 Closed weights · Released 2026-04-01API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges |
thinking_on | 53.81% | 359.2s | $0.2350 API estimate · no cache Cache range: $0.0744–$0.2350 | |
| 57 |
Meta Muse Spark 1.2svc_47_1eed89
Cost / task:
$0.0697
API estimate · no cache
Cache range: $0.0019–$0.0697
Source detailsRun: svc_47_1eed89 Closed weights · Released 2026-08-05API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges |
minimal | 53.62% | 248.8s | $0.0697 API estimate · no cache Cache range: $0.0019–$0.0697 | |
| 58 |
Gemini 3 Flash Previewsvc_88_da1388
Cost / task:
$1.4006
Logged estimate
Source detailsRun: svc_88_da1388 Closed weights · Released 2025-12-17sum_of_rounded_per_response_template_estimates. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges |
low | 33.62% | 492.2s | $1.4006 Logged estimate |
Task performance
Select a measure to compare. Tap a point for its model, score, and timing.
Scatter plot comparing model performance with the selected horizontal metric.
Performance, time, and cost
Drag to rotate; select a point to inspect it. Only results with known costs appear here.
Rotatable three-dimensional scatter plot comparing cost, time, and performance.
All included results for this dataset are shown together. Execution settings can differ; source details are preserved with each run.
Data sources & metric definitions
Latest source inventory: 2026-09-20T20:08:33Z. Download all 146 source rows · Download displayed costs & assumptions · View dataset catalog
Times use the task clock, excluding setup and final verification. Costs are recorded usage or labeled estimates, not invoices. API estimates identify their pricing and cache assumptions in each result's source details. Where shown, the cache range runs from all input cached to none. Standard-price proxies do not claim the archived tier or version price. Partial-coverage estimates show their task counts. Additional pricing sources & assumptions are preserved alongside the original CSV references. Runs without token usage show unit prices but stay out of per-task cost plots. The full-OSWorld human baseline is not transferred to these dataset views.
19 retry-attempt or overlapping summary rows for this dataset remain in the CSV, not counted as independent evaluations. Excluded evaluations are listed below the results.
Model labels use publisher-sourced model metadata and explicit aliases, verified 2026-09-19. Open means downloadable weights, not necessarily an open-source license. Release dates mean first public availability of the identified model version. Unversioned archives with an unverified release date are labeled explicitly and omitted only from the release-date curve.
Steps count action batches sent to the environment: one step call counts once, regardless of the number of actions in its batch. Tool calls count individual recorded environment actions within those batches, including waits; a batch of five actions counts as five, not one. This is an environment-action count, not a count of model API tool-call wrappers. Separate observation and task-completion requests are not included. Observations and model responses are separate counts; steps are not observation count minus one. Output tokens per model response divide total generated tokens (including reasoning) by recorded responses, only when both cover every task and the usage/response counts agree.
Excluded evaluations (8)
Excluded evaluations 8
These runs remain available for inspection but are not included in the results or plots above.
- GLM-5V Turbo · thinking_onsvc_50_592ed5Closed weights · Released 2026-04-01superseded configuration label
- Gemini 3.8 Flash · mediumsvc_97_82e253Closed weights · Released 2026-09-02Older reference (2026-09-03); retaining newer medium run svc_61_1ec4be (2026-09-13, no preload). Not a failed run.
- Meta Muse Spark 1.1 · highsvc_29_8c9590Closed weights · Released 2026-07-09Rate-limited playground contrast; contributor-tier high run svc_58_38520a retained.
- Meta Muse Spark 1.1 · lowsvc_31_c00240Closed weights · Released 2026-07-09Rate-limited playground contrast; contributor-tier low run svc_60_72500b retained.
- Meta Muse Spark 1.1 · mediumsvc_30_d46514Closed weights · Released 2026-07-09Rate-limited playground contrast; contributor-tier medium run svc_59_d8a6d4 retained.
- Meta Muse Spark 1.1 · minimalsvc_34_8c4108Closed weights · Released 2026-07-09Rate-limited playground contrast; contributor-tier minimal run svc_61_96ce3f retained.
- Meta Muse Spark 1.1 · xhighsvc_53_1cade5Closed weights · Released 2026-07-09billing-affected
- Meta Muse Spark 1.1 · xhighsvc_32_109178Closed weights · Released 2026-07-09Rate-limited playground contrast; contributor-tier xhigh run svc_69_41f88a retained.