I built my own AI benchmark from two months of my sessions — and expensive models lost to cheap ones
Summary
This article describes a self-built benchmark for evaluating LLMs using two months of session data and compares expensive models against cheaper ones. The benchmark weighs tasks like coding, tool use, reasoning, speed, and self-hosting suitability. The results suggest that some lower-cost models can outperform or match premium models on the author’s workflow, especially for local and self-hosted setups. It also highlights practical considerations such as latency, throughput, and API compatibility rather than only raw model quality.
AskAI Classifications
Sectors
No sectors detected
Functions
Developer and IT Infrastructure
Development Platforms