Results
Compare agents on the four task subsets used in the paper.
MyPCBench · 38 tasks
4 results · 38 tasks per result
Filter models and reasoning effort
Results
All included results, ordered by average performance.
| # | Model | Effort | Performance | Time / task | Cost / task | |
|---|---|---|---|---|---|---|
| 1 |
GPT-6 Astra · Codexsvc_96_fee517
Cost / task:
$4.8099
API estimate · logged usage
Source detailsRun: svc_96_fee517 Closed weights · Released 2026-09-03Published HF PR25, merged before inventory. GPT-6 Astra via Codex; 100 CUA batches, 7200-second timeout, no-preload track, 38 original partial-credit scores retained. Cost is an API-equivalent estimate from complete logged usage and cache counts at the published September 12 reference rates, not Codex-plan spending or a current-price quote. 0/38 judge bundles lacked supplied final-response text. published_API_rate_estimate_from_logged_usage_and_cache_not_recorded_charge. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges {"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":38,"codex_model_response_count_and_full_tool_calls_unavailable":38} |
xhigh | 93.55% | 619.2s | $4.8099 API estimate · logged usage | |
| 2 |
GPT-6 Astra · Codexsvc_93_cbb0ad
Cost / task:
$4.5741
API estimate · logged usage
Source detailsRun: svc_93_cbb0ad Closed weights · Released 2026-09-03Published HF PR25, merged before inventory. GPT-6 Astra via Codex; 100 CUA batches, 7200-second timeout, no-preload track, 38 original partial-credit scores retained. Cost is an API-equivalent estimate from complete logged usage and cache counts at the published September 12 reference rates, not Codex-plan spending or a current-price quote. 2/38 judge bundles lacked supplied final-response text. Resumed after six preparation failures before agents started; 32 completed records retained and six incomplete instances reopened. No extra timed trajectories. 1 agent(s) exited 1 after the gateway finalized at the 100-batch limit; recorded verdicts unchanged. published_API_rate_estimate_from_logged_usage_and_cache_not_recorded_charge. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges {"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":38,"agent_exit_1_after_gateway_finalized_at_100_batches":1,"codex_model_response_count_and_full_tool_calls_unavailable":38,"judge_did_not_receive_final_response_text":2} |
low | 89.87% | 514.3s | $4.5741 API estimate · logged usage | |
| 3 |
GPT-6 Astra · Codexsvc_95_5ae456
Cost / task:
$4.8124
API estimate · logged usage
Source detailsRun: svc_95_5ae456 Closed weights · Released 2026-09-03Published HF PR25, merged before inventory. GPT-6 Astra via Codex; 100 CUA batches, 7200-second timeout, no-preload track, 38 original partial-credit scores retained. Cost is an API-equivalent estimate from complete logged usage and cache counts at the published September 12 reference rates, not Codex-plan spending or a current-price quote. 1/38 judge bundles lacked supplied final-response text. 1 agent(s) exited 1 after the gateway finalized at the 100-batch limit; recorded verdicts unchanged. published_API_rate_estimate_from_logged_usage_and_cache_not_recorded_charge. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges {"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":38,"agent_exit_1_after_gateway_finalized_at_100_batches":1,"codex_model_response_count_and_full_tool_calls_unavailable":38,"judge_did_not_receive_final_response_text":1} |
high | 88.63% | 568.3s | $4.8124 API estimate · logged usage | |
| 4 |
GPT-6 Astra · Codexsvc_94_2d132c
Cost / task:
$4.3321
API estimate · logged usage
Source detailsRun: svc_94_2d132c Closed weights · Released 2026-09-03Published HF PR25, merged before inventory. GPT-6 Astra via Codex; 100 CUA batches, 7200-second timeout, no-preload track, 38 original partial-credit scores retained. Cost is an API-equivalent estimate from complete logged usage and cache counts at the published September 12 reference rates, not Codex-plan spending or a current-price quote. 3/38 judge bundles lacked supplied final-response text. 2 agent(s) exited 1 after the gateway finalized at the 100-batch limit; recorded verdicts unchanged. One large paste action was omitted from the judge bundle for preference_inference_f025; full action remains in runlog, and the original score of 43 is retained. published_API_rate_estimate_from_logged_usage_and_cache_not_recorded_charge. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges {"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":38,"agent_exit_1_after_gateway_finalized_at_100_batches":2,"codex_model_response_count_and_full_tool_calls_unavailable":38,"judge_did_not_receive_final_response_text":3,"publisher_reports_one_large_action_omitted_from_judge_bundle_full_runlog_preserved":1} |
medium | 88.34% | 495.4s | $4.3321 API estimate · logged usage |
Task performance
Select a measure to compare. Tap a point for its model, score, and timing.
Scatter plot comparing model performance with the selected horizontal metric.
Performance, time, and cost
Drag to rotate; select a point to inspect it. Only results with known costs appear here.
Rotatable three-dimensional scatter plot comparing cost, time, and performance.
All included results for this dataset are shown together. Execution settings can differ; source details are preserved with each run.
Data sources & metric definitions
Latest source inventory: 2026-09-20T20:08:33Z. Download all 146 source rows · Download displayed costs & assumptions · View dataset catalog
Times use the task clock, excluding setup and final verification. Costs are recorded usage or labeled estimates, not invoices. API estimates identify their pricing and cache assumptions in each result's source details. Where shown, the cache range runs from all input cached to none. Standard-price proxies do not claim the archived tier or version price. Partial-coverage estimates show their task counts. Additional pricing sources & assumptions are preserved alongside the original CSV references. Runs without token usage show unit prices but stay out of per-task cost plots. The full-OSWorld human baseline is not transferred to these dataset views.
0 retry-attempt or overlapping summary rows for this dataset remain in the CSV, not counted as independent evaluations. Excluded evaluations are listed below the results.
Model labels use publisher-sourced model metadata and explicit aliases, verified 2026-09-19. Open means downloadable weights, not necessarily an open-source license. Release dates mean first public availability of the identified model version. Unversioned archives with an unverified release date are labeled explicitly and omitted only from the release-date curve.
Steps count action batches sent to the environment: one step call counts once, regardless of the number of actions in its batch. Tool calls count individual recorded environment actions within those batches, including waits; a batch of five actions counts as five, not one. This is an environment-action count, not a count of model API tool-call wrappers. Separate observation and task-completion requests are not included. Observations and model responses are separate counts; steps are not observation count minus one. Output tokens per model response divide total generated tokens (including reasoning) by recorded responses, only when both cover every task and the usage/response counts agree.