Public benchmarks · run on demand

Neutral, receipted benchmarks, re-provable on demand.

RedCrown runs the same pinned tasks through each model in the set under identical conditions, scores them on a real labeled set, and names the cheapest model that clears the quality bar. Every number carries a receipt, and the answer is re-checked as prices and models change. No vendor's thumb on the scale.

Each benchmark runs on the published redcrown CLI, prices each model at its published list rate, counts and shows errors, and links a public proof you can check item by item for every published run. Run benchmarks anywhere; prove them here.