I built my own AI benchmark from two months of my sessions — and expensive models lost to cheap ones

General News

Summary

This article describes a self-built benchmark for evaluating LLMs using two months of session data and compares expensive models against cheaper ones. The benchmark weighs tasks like coding, tool use, reasoning, speed, and self-hosting suitability. The results suggest that some lower-cost models can outperform or match premium models on the author’s workflow, especially for local and self-hosted setups. It also highlights practical considerations such as latency, throughput, and API compatibility rather than only raw model quality.

AskAI Classifications

Sectors
No sectors detected
Functions
Developer and IT Infrastructure Development Platforms

Linked Companies

Google LLC
$100M to $250M
OpenAI
$25M to $50M
Anthropic
$10M to $25M