Pinned
There is a fundamental problem with most AI benchmarks:
They evaluate outputs, while production systems depend on actions.
That’s precisely the gap @Accio_official’s newly open-sourced CommerceAgentBench aims to close.
Take one of its procurement tasks.
The agent receives




