Local LLM Lab.
2 posts in this series. All writing →
Cost-Aware Model Selection: When Local Beats Frontier
The most capable model is not always the correct choice. A routing policy that defaults bounded tasks to a local engine, with a CLI you can run yourself, and two real bugs found while proving the policy worked as documented.
LLM-as-Judge, and How Not to Fool Yourself
A model judging another model's output sounds confident by default. Confident is not the same as calibrated. The measurement methodology that tells the difference, and a bug the methodology caught in its own tool.