If your agent needs 85 MCP turns to answer a SQL-shaped question, the problem may be the interface you gave it.
Retrieval is great when a coding agent needs to inspect a few traces. It gets clumsy when the task is really about counting, filtering, joining, or aggregating across
Better models don’t fix every agent failure.
As models get more capable, the engineering around them matters even more: context, prompts, evals, feedback loops, and the way you measure the agent’s behavior.
We spoke with @stuart__sy from @OpenAI about what developers should
A prompt-level experiment tests one part of an agent.
An agent experiment runs your test cases through the deployed workflow, including routing, tools, and multi-step execution.
Join us on Sep 3 to see how to run that workflow in Arize AX, attach evals, and compare changes
The final answer is only one part of an agent eval.
This Wednesday in SF, our cofounder Aparna Dhinakaran will talk through what “good” actually means for agent evals: sessions, traces, failure modes that look like success.
This will be hosted by Engineering AI Reading Group.