Model Benchmarks
Performance of coding agents creating and changing eve projects, measured by deterministic checks against the files, commands, and simulated external interactions each run produces1.
Last published: August 31, 2026eve revision: 450681ad
Current results
| Claude Code | 110.3s | $4.37 | 100% | 100% | |
| Claude Code | 128.1s | $5.60 | 100% | 100% | |
| Claude Code | 68.8s | $1.20 | 100% | 100% | |
| Codex | 122.7s | $0.23 | 100% | 100% | |
| Codex | 110.4s | $0.15 | 100% | 100% | |
| OpenCode | 77.4s | $0.14 | 86% | 100% | |
| OpenCode | 117.7s | $0.11 | 86% | 100% | |
| OpenCode | 124.5s | $0.06 | 100% | 81% | |
| OpenCode | 95.5s | $0.21 | 57% | 71% |
Retired Models
| OpenCode | 233.1s | $0.07 | 19% | 38% |
1 Each model completes the same deterministic eve authoring tasks three times with and without the guidance files generated by eve init; success rates include every scheduled run, while duration, tokens, and tool calls average valid runs. Cost estimates apply provider public list prices to recorded token usage.