We built a computer use agent once and got it benchmarked at the 2nd position in the biggest benchmark for this field but in the process, we realized one thing, how outdated and inefficient these benchmarks are and how they don't truly capture how these tasks went. So we thought a community led ...
Source: [Hacker News](https://coarena.ai)