Real ranked decision records on illustrative synthetic data, one per modality: transcription, summarization, extraction, vision, and reasoning. Open any of them and check the evidence item by item, the way a client or an auditor would.
All examples use small, clearly-synthetic data. No real client or PHI data.Whisper large-v3-turbo wins, 49% cheaper at equal medical accuracy
View proofQwen3 235B wins, 99% cheaper; the flagship model was both priciest and lowest quality
View proofCheapest-per-correct wins; the most accurate model is not the economic winner (96% cheaper)
View proof