代码库

A benchmark for evaluating AI agents on realistic business workflows
Python
benchmarksevalsllmprimeintellect