Skip to content
Published evaluations

Results

Compare agents on the four task subsets used in the paper.

Dataset

MyPCBench · 38 tasks

4 results · 38 tasks per result

4of 4 results
Filter models and reasoning effort
All models

Results

All included results, ordered by average performance.

#ModelEffortPerformanceTime / taskCost / task
1 GPT-6 Astra · Codexsvc_96_fee517 Cost / task: $4.8099 API estimate · logged usage
Source details

Run: svc_96_fee517

Published HF PR25, merged before inventory. GPT-6 Astra via Codex; 100 CUA batches, 7200-second timeout, no-preload track, 38 original partial-credit scores retained. Cost is an API-equivalent estimate from complete logged usage and cache counts at the published September 12 reference rates, not Codex-plan spending or a current-price quote. 0/38 judge bundles lacked supplied final-response text.

published_API_rate_estimate_from_logged_usage_and_cache_not_recorded_charge. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

{"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":38,"codex_model_response_count_and_full_tool_calls_unavailable":38}

xhigh 93.55% 619.2s $4.8099 API estimate · logged usage
2 GPT-6 Astra · Codexsvc_93_cbb0ad Cost / task: $4.5741 API estimate · logged usage
Source details

Run: svc_93_cbb0ad

Published HF PR25, merged before inventory. GPT-6 Astra via Codex; 100 CUA batches, 7200-second timeout, no-preload track, 38 original partial-credit scores retained. Cost is an API-equivalent estimate from complete logged usage and cache counts at the published September 12 reference rates, not Codex-plan spending or a current-price quote. 2/38 judge bundles lacked supplied final-response text. Resumed after six preparation failures before agents started; 32 completed records retained and six incomplete instances reopened. No extra timed trajectories. 1 agent(s) exited 1 after the gateway finalized at the 100-batch limit; recorded verdicts unchanged.

published_API_rate_estimate_from_logged_usage_and_cache_not_recorded_charge. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

{"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":38,"agent_exit_1_after_gateway_finalized_at_100_batches":1,"codex_model_response_count_and_full_tool_calls_unavailable":38,"judge_did_not_receive_final_response_text":2}

low 89.87% 514.3s $4.5741 API estimate · logged usage
3 GPT-6 Astra · Codexsvc_95_5ae456 Cost / task: $4.8124 API estimate · logged usage
Source details

Run: svc_95_5ae456

Published HF PR25, merged before inventory. GPT-6 Astra via Codex; 100 CUA batches, 7200-second timeout, no-preload track, 38 original partial-credit scores retained. Cost is an API-equivalent estimate from complete logged usage and cache counts at the published September 12 reference rates, not Codex-plan spending or a current-price quote. 1/38 judge bundles lacked supplied final-response text. 1 agent(s) exited 1 after the gateway finalized at the 100-batch limit; recorded verdicts unchanged.

published_API_rate_estimate_from_logged_usage_and_cache_not_recorded_charge. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

{"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":38,"agent_exit_1_after_gateway_finalized_at_100_batches":1,"codex_model_response_count_and_full_tool_calls_unavailable":38,"judge_did_not_receive_final_response_text":1}

high 88.63% 568.3s $4.8124 API estimate · logged usage
4 GPT-6 Astra · Codexsvc_94_2d132c Cost / task: $4.3321 API estimate · logged usage
Source details

Run: svc_94_2d132c

Published HF PR25, merged before inventory. GPT-6 Astra via Codex; 100 CUA batches, 7200-second timeout, no-preload track, 38 original partial-credit scores retained. Cost is an API-equivalent estimate from complete logged usage and cache counts at the published September 12 reference rates, not Codex-plan spending or a current-price quote. 3/38 judge bundles lacked supplied final-response text. 2 agent(s) exited 1 after the gateway finalized at the 100-batch limit; recorded verdicts unchanged. One large paste action was omitted from the judge bundle for preference_inference_f025; full action remains in runlog, and the original score of 43 is retained.

published_API_rate_estimate_from_logged_usage_and_cache_not_recorded_charge. not an invoice; excludes unlogged requests, startup credential probes, infrastructure and separate grading charges

{"Astra_cost_excludes_per_request_long_context_fast_mode_and_tool_adjustments":38,"agent_exit_1_after_gateway_finalized_at_100_batches":2,"codex_model_response_count_and_full_tool_calls_unavailable":38,"judge_did_not_receive_final_response_text":3,"publisher_reports_one_large_action_omitted_from_judge_bundle_full_runlog_preserved":1}

medium 88.34% 495.4s $4.3321 API estimate · logged usage

All included results for this dataset are shown together. Execution settings can differ; source details are preserved with each run.

Data sources & metric definitions

Latest source inventory: 2026-09-20T20:08:33Z. Download all 146 source rows · Download displayed costs & assumptions · View dataset catalog

Times use the task clock, excluding setup and final verification. Costs are recorded usage or labeled estimates, not invoices. API estimates identify their pricing and cache assumptions in each result's source details. Where shown, the cache range runs from all input cached to none. Standard-price proxies do not claim the archived tier or version price. Partial-coverage estimates show their task counts. Additional pricing sources & assumptions are preserved alongside the original CSV references. Runs without token usage show unit prices but stay out of per-task cost plots. The full-OSWorld human baseline is not transferred to these dataset views.

0 retry-attempt or overlapping summary rows for this dataset remain in the CSV, not counted as independent evaluations. Excluded evaluations are listed below the results.

Model labels use publisher-sourced model metadata and explicit aliases, verified 2026-09-19. Open means downloadable weights, not necessarily an open-source license. Release dates mean first public availability of the identified model version. Unversioned archives with an unverified release date are labeled explicitly and omitted only from the release-date curve.

Steps count action batches sent to the environment: one step call counts once, regardless of the number of actions in its batch. Tool calls count individual recorded environment actions within those batches, including waits; a batch of five actions counts as five, not one. This is an environment-action count, not a count of model API tool-call wrappers. Separate observation and task-completion requests are not included. Observations and model responses are separate counts; steps are not observation count minus one. Output tokens per model response divide total generated tokens (including reasoning) by recorded responses, only when both cover every task and the usage/response counts agree.