Run an evaluation
Evaluate a template or your own agent on benchmark tasks. Runs use your Modal account and model API keys.
Checking evaluation service…
Run from your own machine
The command-line toolkit is available now. It starts the agent and desktop environments on Modal and saves the results to your machine.
- Install CUA-Speedrun and connect your Modal account.
- Choose an agent template and add its model API key.
- Run a benchmark and inspect the saved scores and trajectories.