Skip to content
Published evaluations

Results

Compare agents on the four task subsets used in the paper.

Dataset

OSWorld · 50 tasks

58 results · 50 tasks per result · 8 excluded evaluations

58of 58 results
Filter models and reasoning effort
All models

Results

All included results, ordered by average performance.

#ModelEffortPerformanceTime / taskCost / task
1 GPT-6 Astra · Normal I/Osvc_121_b4bca3 Cost / task: $0.7081 Logged estimate
Source details

Run: svc_121_b4bca3

Cost is short-context API-equivalent usage, not a subscription payment; full tool/response trace unavailable in CLI export.

Astra_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

{"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":50,"codex_model_response_count_and_full_tool_calls_unavailable":50}

xhigh 91.62% 126.8s $0.7081 Logged estimate
2 Gemini 3.8 Flashsvc_59_708d1b Cost / task: $0.1080 Logged estimate
Source details

Run: svc_59_708d1b

New no-preload Energy50 evaluation; preserve its distinct track and measurement contract when comparing earlier runs. Cost is the logged template estimate from final cumulative snapshots, not an invoice.

recorded_template_price_estimate. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

low 91.62% 127.1s $0.1080 Logged estimate
3 Claude Opus 5svc_31_878995 Cost / task: $0.6611 Logged estimate
Source details

Run: svc_31_878995

Claude_standard_global_list_price_estimate_from_logged_usage. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

high 91.62% 136.0s $0.6611 Logged estimate
4 GPT-6 Astra · Single actionsvc_157_c2dbaf Cost / task: $1.3637 Logged estimate
Source details

Run: svc_157_c2dbaf

Completed local evaluation 157; all 50 tasks. Single-action mode enforces one GUI action followed by a fresh screenshot before another action. Same model, effort and configured limits as the normal-I/O xhigh baseline; owner-approved current runtime and Codex subscription. Cost is the same historical standard API-equivalent estimate from logged usage, not subscription spending. The source link contains derived task metrics and source hashes; the trajectory archive has not been published to Hugging Face. One GUI action per request, with a fresh screenshot required before the next action.

Astra_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

{"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":50,"codex_model_response_count_and_full_tool_calls_unavailable":50}

xhigh 91.62% 182.7s $1.3637 Logged estimate
5 Gemini 3.8 Flashsvc_61_1ec4be Cost / task: $0.2209 Logged estimate
Source details

Run: svc_61_1ec4be

New no-preload Energy50 evaluation; preserve its distinct track and measurement contract when comparing earlier runs. Cost is the logged template estimate from final cumulative snapshots, not an invoice.

recorded_template_price_estimate. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

medium 91.62% 250.1s $0.2209 Logged estimate
6 GPT-6 Astra · Normal I/Osvc_151_181620 Cost / task: $0.6001 Logged estimate
Source details

Run: svc_151_181620

Added 2026-09-14 at owner request. Separate no-preload track/measurement contract; not a same-track ranking against older preload runs. Cost uses the prior September11 short-context API-equivalent pricing snapshot, not subscription billing. Full CLI model-response/tool trace remains unavailable. Final-task metrics exclude retained non-final attempts.

Astra_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

{"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":50,"codex_model_response_count_and_full_tool_calls_unavailable":50}

medium 89.62% 90.2s $0.6001 Logged estimate
7 GPT-6 Astra · Normal I/Osvc_152_8d02f4 Cost / task: $0.6364 Logged estimate
Source details

Run: svc_152_8d02f4

Added 2026-09-14 at owner request. Separate no-preload track/measurement contract; not a same-track ranking against older preload runs. Cost uses the prior September11 short-context API-equivalent pricing snapshot, not subscription billing. Full CLI model-response/tool trace remains unavailable. Final-task metrics exclude retained non-final attempts. One retained infrastructure attempt is represented separately; final-task API-equivalent total=31.818674 USD, extra attempt=0.263074 USD, all recorded sessions=32.081748 USD.

Astra_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

{"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":50,"codex_model_response_count_and_full_tool_calls_unavailable":50}

high 89.62% 100.8s $0.6364 Logged estimate
8 GPT-6 Astra · Fast I/Osvc_127_422895 Cost / task: $0.7367 Logged estimate
Source details

Run: svc_127_422895

Cost is short-context API-equivalent usage, not a subscription payment; full tool/response trace unavailable in CLI export.

Astra_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

{"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":50,"codex_model_response_count_and_full_tool_calls_unavailable":50}

low 89.62% 105.2s $0.7367 Logged estimate
9 Claude Sonnet 5svc_24_ddcb3b Cost / task: $0.3086 Logged estimate
Source details

Run: svc_24_ddcb3b

Claude_standard_global_list_price_estimate_from_logged_usage. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

high 89.62% 140.8s $0.3086 Logged estimate
10 Gemini 3.7 Flashsvc_85_948d7f Cost / task: $0.9437 Logged estimate
Source details

Run: svc_85_948d7f

sum_of_rounded_per_response_template_estimates. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

medium 89.62% 198.9s $0.9437 Logged estimate
11 Gemini 3.7 Flashsvc_86_9e7a56 Cost / task: $1.4121 Logged estimate
Source details

Run: svc_86_9e7a56

sum_of_rounded_per_response_template_estimates. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

high 89.62% 254.3s $1.4121 Logged estimate
12 Kimi K3 · Batched tool callssvc_141_f28fed Cost / task: $0.5034 Recorded usage
Source details

Run: svc_141_f28fed

recorded_provider_usage_cost. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

max 89.62% 443.7s $0.5034 Recorded usage
13 Claude Opus 5svc_32_909005 Cost / task: $0.3896 Logged estimate
Source details

Run: svc_32_909005

Claude_standard_global_list_price_estimate_from_logged_usage. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

{"Claude_cost_assumes_standard_global_for_responses_with_missing_tier_geo":1}

low 87.62% 85.9s $0.3896 Logged estimate
14 GPT-6 Astra · Normal I/Osvc_122_93c9c5 Cost / task: $0.5665 Logged estimate
Source details

Run: svc_122_93c9c5

Cost is short-context API-equivalent usage, not a subscription payment; full tool/response trace unavailable in CLI export.

Astra_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

{"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":50,"codex_model_response_count_and_full_tool_calls_unavailable":50}

low 87.62% 86.1s $0.5665 Logged estimate
15 Gemini 3.7 Flashsvc_87_735f9a Cost / task: $0.9264 Logged estimate
Source details

Run: svc_87_735f9a

sum_of_rounded_per_response_template_estimates. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

low 87.62% 207.3s $0.9264 Logged estimate
16 Gemini 3.8 Flashsvc_60_8db123 Cost / task: $0.3475 Logged estimate
Source details

Run: svc_60_8db123

New no-preload Energy50 evaluation; preserve its distinct track and measurement contract when comparing earlier runs. Cost is the logged template estimate from final cumulative snapshots, not an invoice.

recorded_template_price_estimate. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

high 87.62% 347.7s $0.3475 Logged estimate
17 Kimi K3 · Single tool callsvc_56_9e002a Cost / task: $0.7158 Recorded usage
Source details

Run: svc_56_9e002a

Published Energy50 max-reasoning subset of Kimi evaluation56; not a separate execution. One model tool call per response; the code in that call can contain multiple GUI actions.

recorded_provider_usage_cost. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

{"Kimi_finish_rejected_response_usage_included_but_tool_bodies_unlogged":1,"Kimi_noncompletion_replies_have_no_logged_usage_or_body":2}

max 85.62% 606.3s $0.7158 Recorded usage
18 GPT-5.6 Luna · Direct APIsvc_20_9456b8 Cost / task: $0.0256 Logged estimate
Source details

Run: svc_20_9456b8

sum_of_rounded_per_response_template_estimates. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

medium 81.62% 100.6s $0.0256 Logged estimate
19 GPT-5.6 Solsvc_23_753ef9 Cost / task: $0.5267 Logged estimate
Source details

Run: svc_23_753ef9

sum_of_rounded_per_response_template_estimates. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

xhigh 81.62% 108.3s $0.5267 Logged estimate
20 Kimi K3 · Single tool callsvc_90_357e81 Cost / task: $0.4141 Recorded usage
Source details

Run: svc_90_357e81

recorded_provider_usage_cost. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

high 81.62% 305.1s $0.4141 Recorded usage
21 GPT-5.6 Solsvc_22_3a244d Cost / task: $0.5394 Logged estimate
Source details

Run: svc_22_3a244d

sum_of_rounded_per_response_template_estimates. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

medium 79.62% 121.3s $0.5394 Logged estimate
22 Yutori n2svc_13_01885e Cost / task: $0.1036 API estimate · logged usage
Source details

Run: svc_13_01885e

Published HF PR23, merged before inventory. All 50 final tasks complete; original no-preload contract and GUI-only tool configuration retained. Full response usage recovered from agent.stdout despite the published README marking usage unavailable; duplicated SDK step usage is not counted again. Cost is an API-rate estimate using logged input/output and billed cached input at publisher rates verified September 19, not a recorded charge. Reasoning-token breakdown is incomplete and remains unknown.

published_API_rate_estimate_from_logged_usage_and_billed_cache_not_recorded_charge. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

{"reasoning_token_breakdown_not_logged_for_all_responses":50}

xhigh 79.62% 225.0s $0.1036 API estimate · logged usage
23 GPT-5.6 Luna · Codexsvc_155_7fbf62 Cost / task: $0.0759 Logged estimate
Source details

Run: svc_155_7fbf62

Added 2026-09-16 at owner request. GPT-5.6 Luna via Codex, not the direct-API gpt54 template. All 50 original tasks included, including failures. No rerun or regrading. No-preload track/measurement contract; not a same-track ranking against older preload runs. Costs are standard short-context API-equivalent estimates using official 2026-09-16 Luna rates, not subscription charges. All 50 session token summaries complete; cached input is a subset of input, reasoning is a subset of output, recorded cache-write counters are zero. Full CLI model-response/tool trace remains unavailable; observed CLI tool items and harness steps are separate metrics. No additional retained attempts found. Submitted agent.py SHA-256=306fe8c02f9c85a5ff90be95d8efef897a3555ab37fe4f46e84dad68dad0058a. Full per-task normalized evidence and usage breakdowns are linked in companion_cost_usage_url.

Luna_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges; assumes standard short-context rates; per-request long-context/fast-mode adjustments and tool charges cannot be reconstructed from session totals; output already includes reasoning

{"Luna_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":50,"codex_model_response_count_and_full_tool_calls_unavailable":50}

high 79.62% 324.8s $0.0759 Logged estimate
24 Kimi K3 · Batched tool callssvc_144_1f26e4 Cost / task: $0.3823 Recorded usage
Source details

Run: svc_144_1f26e4

recorded_provider_usage_cost. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

high 79.62% 358.1s $0.3823 Recorded usage
25 Meta Muse Spark 1.1svc_69_41f88a Cost / task: $1.6849 Standard-price proxy · no cache Cache range: $0.2422–$1.6849
Source details

Run: svc_69_41f88a

Standard API price proxy, not the archived contributor-tier charge; historical contributor pricing is unverified. Recorded input/output tokens; no input caching assumed. Rates verified 2026-09-16; excludes other charges.

xhigh 79.62% 417.2s $1.6849 Standard-price proxy · no cache Cache range: $0.2422–$1.6849
26 GPT-5.6 Luna · Codexsvc_156_2bc119 Cost / task: $0.0631 Logged estimate
Source details

Run: svc_156_2bc119

Added 2026-09-16 at owner request. GPT-5.6 Luna via Codex, not the direct-API gpt54 template. All 50 original tasks included, including failures. No rerun or regrading. No-preload track/measurement contract; not a same-track ranking against older preload runs. Costs are standard short-context API-equivalent estimates using official 2026-09-16 Luna rates, not subscription charges. All 50 session token summaries complete; cached input is a subset of input, reasoning is a subset of output, recorded cache-write counters are zero. Full CLI model-response/tool trace remains unavailable; observed CLI tool items and harness steps are separate metrics. No additional retained attempts found. Submitted agent.py SHA-256=306fe8c02f9c85a5ff90be95d8efef897a3555ab37fe4f46e84dad68dad0058a. Full per-task normalized evidence and usage breakdowns are linked in companion_cost_usage_url.

Luna_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges; assumes standard short-context rates; per-request long-context/fast-mode adjustments and tool charges cannot be reconstructed from session totals; output already includes reasoning

{"Luna_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":50,"codex_model_response_count_and_full_tool_calls_unavailable":50}

xhigh 77.62% 297.2s $0.0631 Logged estimate
27 Yutori n2svc_11_0dddb1 Cost / task: $0.1569 API estimate · logged usage
Source details

Run: svc_11_0dddb1

Published HF PR23, merged before inventory. All 50 final tasks complete; original no-preload contract and GUI-only tool configuration retained. Full response usage recovered from agent.stdout despite the published README marking usage unavailable; duplicated SDK step usage is not counted again. Cost is an API-rate estimate using logged input/output and billed cached input at publisher rates verified September 19, not a recorded charge. Reasoning-token breakdown is incomplete and remains unknown.

published_API_rate_estimate_from_logged_usage_and_billed_cache_not_recorded_charge. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

{"reasoning_token_breakdown_not_logged_for_all_responses":50}

low 77.62% 331.0s $0.1569 API estimate · logged usage
28 Meta Muse Spark 1.3svc_68 + rerun3_xhigh13 Cost / task: $0.1607 API estimate · no cache Cache range: $0.0055–$0.1607
Source details

Run: svc_68 + rerun3_xhigh13

Published STITCH-musespark13-xhigh-68.txt: 47 original tasks plus three rerun replacements; overlaps both components and is not an independent run. Published 664.79 seconds/task is mean(agent_wall_sec + env_boot_sec)=664.78846; timed task clock=612.07638625182 seconds/task. Published as one combined 50-task result. Original tasks ran at concurrency 8; the three recovery tasks ran at concurrency 3. Both source archives are retained below.

API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges

svc_68_bca75f · rerun3_xhigh13

{"result_json_absent_metrics_reduced_from_original_runlog_events":47}

xhigh 77.62% 612.1s $0.1607 API estimate · no cache Cache range: $0.0055–$0.1607
29 Meta Muse Spark 1.3svc_65_46a645 Cost / task: $0.1258 API estimate · no cache Cache range: $0.0043–$0.1258
Source details

Run: svc_65_46a645

API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges

medium 77.54% 397.6s $0.1258 API estimate · no cache Cache range: $0.0043–$0.1258
30 Claude Sonnet 5svc_25_adbcee Cost / task: $0.2941 Logged estimate
Source details

Run: svc_25_adbcee

Claude_standard_global_list_price_estimate_from_logged_usage. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

low 75.81% 129.4s $0.2941 Logged estimate
31 GPT-5.6 Luna · Codexsvc_153_795ea7 Cost / task: $0.0254 Logged estimate
Source details

Run: svc_153_795ea7

Added 2026-09-16 at owner request. GPT-5.6 Luna via Codex, not the direct-API gpt54 template. All 50 original tasks included, including failures. No rerun or regrading. No-preload track/measurement contract; not a same-track ranking against older preload runs. Costs are standard short-context API-equivalent estimates using official 2026-09-16 Luna rates, not subscription charges. All 50 session token summaries complete; cached input is a subset of input, reasoning is a subset of output, recorded cache-write counters are zero. Full CLI model-response/tool trace remains unavailable; observed CLI tool items and harness steps are separate metrics. No additional retained attempts found. Submitted agent.py SHA-256=306fe8c02f9c85a5ff90be95d8efef897a3555ab37fe4f46e84dad68dad0058a. Full per-task normalized evidence and usage breakdowns are linked in companion_cost_usage_url.

Luna_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges; assumes standard short-context rates; per-request long-context/fast-mode adjustments and tool charges cannot be reconstructed from session totals; output already includes reasoning

{"Luna_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":50,"codex_model_response_count_and_full_tool_calls_unavailable":50}

low 75.62% 144.8s $0.0254 Logged estimate
32 MiniMax M3svc_51_121311 Cost / task: Usage not logged API $0.30 in / $1.20 out per 1M tokens · standard · ≤512K context
Source details

Run: svc_51_121311

Standard API rates through 512K context; above 512K, input/output/cache rates double. Priority pricing differs. Archived responses do not retain token usage, so unit prices cannot yield a per-task total.

{"output_truncated_at_400_characters_and_usage_not_logged":50}

thinking_off 75.62% 253.8s Usage not logged API $0.30 in / $1.20 out per 1M tokens · standard · ≤512K context
33 Yutori n2svc_10_cd0ed2 Cost / task: $0.0811 API estimate · logged usage
Source details

Run: svc_10_cd0ed2

Published HF PR23, merged before inventory. All 50 final tasks complete; original no-preload contract and GUI-only tool configuration retained. Full response usage recovered from agent.stdout despite the published README marking usage unavailable; duplicated SDK step usage is not counted again. Cost is an API-rate estimate using logged input/output and billed cached input at publisher rates verified September 19, not a recorded charge. Reasoning-token breakdown is incomplete and remains unknown.

published_API_rate_estimate_from_logged_usage_and_billed_cache_not_recorded_charge. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

{"reasoning_token_breakdown_not_logged_for_all_responses":50}

none 75.62% 298.5s $0.0811 API estimate · logged usage
34 GPT-5.6 Luna · Direct APIsvc_21_00be0a Cost / task: $0.0386 Logged estimate
Source details

Run: svc_21_00be0a

sum_of_rounded_per_response_template_estimates. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

xhigh 73.62% 135.3s $0.0386 Logged estimate
35 Kimi K3 · Batched tool callssvc_143_e8d33d Cost / task: $0.1740 Recorded usage
Source details

Run: svc_143_e8d33d

recorded_provider_usage_cost. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

low 73.62% 149.5s $0.1740 Recorded usage
36 Meta Muse Spark 1.3svc_70_8f6cda Cost / task: $0.1479 API estimate · no cache Cache range: $0.0050–$0.1479
Source details

Run: svc_70_8f6cda

API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges

high 73.62% 506.9s $0.1479 API estimate · no cache Cache range: $0.0050–$0.1479
37 Kimi K3 · Single tool callsvc_89_b4886a Cost / task: $0.2802 Recorded usage
Source details

Run: svc_89_b4886a

recorded_provider_usage_cost. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

low 71.62% 176.6s $0.2802 Recorded usage
38 MiniMax M3svc_52_554434 Cost / task: Usage not logged API $0.30 in / $1.20 out per 1M tokens · standard · ≤512K context
Source details

Run: svc_52_554434

Standard API rates through 512K context; above 512K, input/output/cache rates double. Priority pricing differs. Archived responses do not retain token usage, so unit prices cannot yield a per-task total.

{"output_truncated_at_400_characters_and_usage_not_logged":50}

thinking_on 69.62% 345.1s Usage not logged API $0.30 in / $1.20 out per 1M tokens · standard · ≤512K context
39 Meta Muse Spark 1.1svc_59_d8a6d4 Cost / task: $1.3785 Standard-price proxy · no cache Cache range: $0.1929–$1.3785
Source details

Run: svc_59_d8a6d4

Standard API price proxy, not the archived contributor-tier charge; historical contributor pricing is unverified. Recorded input/output tokens; no input caching assumed. Rates verified 2026-09-16; excludes other charges.

medium 69.62% 359.8s $1.3785 Standard-price proxy · no cache Cache range: $0.1929–$1.3785
40 Meta Muse Spark 1.1svc_58_38520a Cost / task: $1.7059 Standard-price proxy · no cache Cache range: $0.2426–$1.7059
Source details

Run: svc_58_38520a

Standard API price proxy, not the archived contributor-tier charge; historical contributor pricing is unverified. Recorded input/output tokens; no input caching assumed. Rates verified 2026-09-16; excludes other charges.

high 69.62% 368.5s $1.7059 Standard-price proxy · no cache Cache range: $0.2426–$1.7059
41 Meta Muse Spark 1.1svc_60_72500b Cost / task: $1.3116 Standard-price proxy · no cache Cache range: $0.1818–$1.3116
Source details

Run: svc_60_72500b

Standard API price proxy, not the archived contributor-tier charge; historical contributor pricing is unverified. Recorded input/output tokens; no input caching assumed. Rates verified 2026-09-16; excludes other charges.

low 69.62% 370.4s $1.3116 Standard-price proxy · no cache Cache range: $0.1818–$1.3116
42 Meta Muse Spark 1.2svc_39_025394 Cost / task: $0.2445 API estimate · no cache Cache range: $0.0083–$0.2445
Source details

Run: svc_39_025394

API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges

xhigh 69.62% 720.9s $0.2445 API estimate · no cache Cache range: $0.0083–$0.2445
43 GPT-5.6 Luna · Codexsvc_154_323b97 Cost / task: $0.0353 Logged estimate
Source details

Run: svc_154_323b97

Added 2026-09-16 at owner request. GPT-5.6 Luna via Codex, not the direct-API gpt54 template. All 50 original tasks included, including failures. No rerun or regrading. No-preload track/measurement contract; not a same-track ranking against older preload runs. Costs are standard short-context API-equivalent estimates using official 2026-09-16 Luna rates, not subscription charges. All 50 session token summaries complete; cached input is a subset of input, reasoning is a subset of output, recorded cache-write counters are zero. Full CLI model-response/tool trace remains unavailable; observed CLI tool items and harness steps are separate metrics. No additional retained attempts found. Submitted agent.py SHA-256=306fe8c02f9c85a5ff90be95d8efef897a3555ab37fe4f46e84dad68dad0058a. One original agent_error on osworld_libreoffice_writer_adf5e2c3-64c7-4644-b7b6-d2f0167927e7/seed_1031319744: the environment HTTPS endpoint refused the /done connection and the gateway failed. Its zero score, timing and recorded usage remain included; not a clean failure-free comparison. Full per-task normalized evidence and usage breakdowns are linked in companion_cost_usage_url.

Luna_standard_short_context_API_equivalent_not_subscription_bill. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges; assumes standard short-context rates; per-request long-context/fast-mode adjustments and tool charges cannot be reconstructed from session totals; output already includes reasoning

{"Luna_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":50,"codex_model_response_count_and_full_tool_calls_unavailable":50}

medium 67.62% 188.8s $0.0353 Logged estimate
44 Yutori n2svc_12_026ed7 Cost / task: $0.1100 API estimate · logged usage
Source details

Run: svc_12_026ed7

Published HF PR23, merged before inventory. All 50 final tasks complete; original no-preload contract and GUI-only tool configuration retained. Full response usage recovered from agent.stdout despite the published README marking usage unavailable; duplicated SDK step usage is not counted again. Cost is an API-rate estimate using logged input/output and billed cached input at publisher rates verified September 19, not a recorded charge. Reasoning-token breakdown is incomplete and remains unknown.

published_API_rate_estimate_from_logged_usage_and_billed_cache_not_recorded_charge. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

{"reasoning_token_breakdown_not_logged_for_all_responses":50}

medium 67.62% 265.4s $0.1100 API estimate · logged usage
45 Meta Muse Spark 1.1svc_61_96ce3f Cost / task: $1.1840 Standard-price proxy · no cache Cache range: $0.1644–$1.1840
Source details

Run: svc_61_96ce3f

Standard API price proxy, not the archived contributor-tier charge; historical contributor pricing is unverified. Recorded input/output tokens; no input caching assumed. Rates verified 2026-09-16; excludes other charges.

minimal 67.62% 318.5s $1.1840 Standard-price proxy · no cache Cache range: $0.1644–$1.1840
46 Meta Muse Spark 1.3svc_66_660b4a Cost / task: $0.1071 API estimate · no cache Cache range: $0.0035–$0.1071
Source details

Run: svc_66_660b4a

API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges

low 67.62% 361.0s $0.1071 API estimate · no cache Cache range: $0.0035–$0.1071
47 GPT-5.6 Solsvc_62_9b9e92 Cost / task: $0.3191 Logged estimate
Source details

Run: svc_62_9b9e92

Added from JY's HF PR21 at immutable revision 5d04d3b4b91dceb4c90e8592ce3235c1965c9db2; PR open at inventory. All 50 final tasks completed. No-preload track and original run contract retained. Cost is the sum of rounded per-response template estimates from final-task logs, not an invoice. Published-rate reconstruction independently agrees within the saved rounding bound. One prior infrastructure attempt is retained separately; the publisher's $16.43 total includes it, while this final-task row excludes its cost and usage.

sum_of_rounded_per_response_template_estimates. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

low 65.81% 100.1s $0.3191 Logged estimate
48 Meta Muse Spark 1.2svc_48_4e20fb Cost / task: $0.1054 API estimate · no cache Cache range: $0.0030–$0.1054
Source details

Run: svc_48_4e20fb

API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges

low 63.62% 363.6s $0.1054 API estimate · no cache Cache range: $0.0030–$0.1054
49 Meta Muse Spark 1.2svc_41_6ef6f1 Cost / task: $0.1365 API estimate · no cache Cache range: $0.0043–$0.1365
Source details

Run: svc_41_6ef6f1

API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges

medium 63.62% 466.2s $0.1365 API estimate · no cache Cache range: $0.0043–$0.1365
50 GPT-5.6 Luna · Direct APIsvc_63_ccb407 Cost / task: $0.0107 Logged estimate
Source details

Run: svc_63_ccb407

Added from JY's HF PR21 at immutable revision 5d04d3b4b91dceb4c90e8592ce3235c1965c9db2; PR open at inventory. All 50 final tasks completed. No-preload track and original run contract retained. Cost is the sum of rounded per-response template estimates from final-task logs, not an invoice. Published-rate reconstruction independently agrees within the saved rounding bound.

sum_of_rounded_per_response_template_estimates. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

low 61.81% 73.6s $0.0107 Logged estimate
51 Meta Muse Spark 1.3svc_67_b7d034 Cost / task: $0.0789 API estimate · no cache Cache range: $0.0023–$0.0789
Source details

Run: svc_67_b7d034

API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges

minimal 61.62% 264.4s $0.0789 API estimate · no cache Cache range: $0.0023–$0.0789
52 Meta Muse Spark 1.2svc_40_c6a678 Cost / task: $0.1736 API estimate · no cache Cache range: $0.0053–$0.1736
Source details

Run: svc_40_c6a678

API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges

high 61.62% 669.8s $0.1736 API estimate · no cache Cache range: $0.0053–$0.1736
53 Gemini 3 Flash Previewsvc_81_c8b73b Cost / task: $0.8091 Logged estimate
Source details

Run: svc_81_c8b73b

sum_of_rounded_per_response_template_estimates. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

high 59.62% 337.2s $0.8091 Logged estimate
54 GLM-5V Turbosvc_87_ed8093 Cost / task: $0.2203 API estimate · no cache Cache range: $0.0701–$0.2203
Source details

Run: svc_87_ed8093

Publisher-designated corrected thinking-off/Pillow-fixed rerun; Pillow 11.3.0 installation is present in init.log. Recorded dollar cost and reasoning/cache token breakdown are unavailable, not zero. Separate reference-cost bounds use the cited OpenRouter rate snapshot and assume all versus no input cached; they are not billing.

API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges

thinking_off 57.81% 752.4s $0.2203 API estimate · no cache Cache range: $0.0701–$0.2203
55 Gemini 3 Flash Previewsvc_83_d83faa Cost / task: $0.5923 Logged estimate
Source details

Run: svc_83_d83faa

sum_of_rounded_per_response_template_estimates. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

medium 57.62% 268.0s $0.5923 Logged estimate
56 GLM-5V Turbosvc_49_c74b13 Cost / task: $0.2350 API estimate · no cache Cache range: $0.0744–$0.2350
Source details

Run: svc_49_c74b13

API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges

thinking_on 53.81% 359.2s $0.2350 API estimate · no cache Cache range: $0.0744–$0.2350
57 Meta Muse Spark 1.2svc_47_1eed89 Cost / task: $0.0697 API estimate · no cache Cache range: $0.0019–$0.0697
Source details

Run: svc_47_1eed89

API-equivalent estimate using recorded tokens and no input caching; not historical billing. Current OpenRouter public API-equivalent token price interval, not historical billing; lower assumes all input cached, upper none; excludes other charges

minimal 53.62% 248.8s $0.0697 API estimate · no cache Cache range: $0.0019–$0.0697
58 Gemini 3 Flash Previewsvc_88_da1388 Cost / task: $1.4006 Logged estimate
Source details

Run: svc_88_da1388

sum_of_rounded_per_response_template_estimates. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

low 33.62% 492.2s $1.4006 Logged estimate

All included results for this dataset are shown together. Execution settings can differ; source details are preserved with each run.

Data sources & metric definitions

Latest source inventory: 2026-09-20T20:08:33Z. Download all 146 source rows · Download displayed costs & assumptions · View dataset catalog

Times use the task clock, excluding setup and final verification. Costs are recorded usage or labeled estimates, not invoices. API estimates identify their pricing and cache assumptions in each result's source details. Where shown, the cache range runs from all input cached to none. Standard-price proxies do not claim the archived tier or version price. Partial-coverage estimates show their task counts. Additional pricing sources & assumptions are preserved alongside the original CSV references. Runs without token usage show unit prices but stay out of per-task cost plots. The full-OSWorld human baseline is not transferred to these dataset views.

19 retry-attempt or overlapping summary rows for this dataset remain in the CSV, not counted as independent evaluations. Excluded evaluations are listed below the results.

Model labels use publisher-sourced model metadata and explicit aliases, verified 2026-09-19. Open means downloadable weights, not necessarily an open-source license. Release dates mean first public availability of the identified model version. Unversioned archives with an unverified release date are labeled explicitly and omitted only from the release-date curve.

Steps count action batches sent to the environment: one step call counts once, regardless of the number of actions in its batch. Tool calls count individual recorded environment actions within those batches, including waits; a batch of five actions counts as five, not one. This is an environment-action count, not a count of model API tool-call wrappers. Separate observation and task-completion requests are not included. Observations and model responses are separate counts; steps are not observation count minus one. Output tokens per model response divide total generated tokens (including reasoning) by recorded responses, only when both cover every task and the usage/response counts agree.

Excluded evaluations (8)
Source coverage

Excluded evaluations 8

These runs remain available for inspection but are not included in the results or plots above.

  • GLM-5V Turbo · thinking_onsvc_50_592ed5
    superseded configuration label
  • Gemini 3.8 Flash · mediumsvc_97_82e253
    Older reference (2026-09-03); retaining newer medium run svc_61_1ec4be (2026-09-13, no preload). Not a failed run.
  • Meta Muse Spark 1.1 · highsvc_29_8c9590
    Rate-limited playground contrast; contributor-tier high run svc_58_38520a retained.
  • Meta Muse Spark 1.1 · lowsvc_31_c00240
    Rate-limited playground contrast; contributor-tier low run svc_60_72500b retained.
  • Meta Muse Spark 1.1 · mediumsvc_30_d46514
    Rate-limited playground contrast; contributor-tier medium run svc_59_d8a6d4 retained.
  • Meta Muse Spark 1.1 · minimalsvc_34_8c4108
    Rate-limited playground contrast; contributor-tier minimal run svc_61_96ce3f retained.
  • Meta Muse Spark 1.1 · xhighsvc_53_1cade5
    billing-affected
  • Meta Muse Spark 1.1 · xhighsvc_32_109178
    Rate-limited playground contrast; contributor-tier xhigh run svc_69_41f88a retained.