Reproducible on-device LLM benchmarks — which models actually run on an iPhone, how smart they stay, and how fast they decode. Measurement, not voting.
- Leaderboard — Pareto chart + full table
- Methodology — what is measured, proof strength, CIs
- Results dataset — CC-BY-4.0, on Hugging Face
- Submit a model / object to a score
Maintained by Daisuke Majima (MLBoy) — who also ports the
Core AI model zoo most aimodel rows come from
(disclosed; every gate and recipe is public) and wrote the textbook
The Art of Core AI (JA).