Benchmarks
We test frontier models against human-verified data drawn from real conversations and real work, the conditions models actually meet in production, not curated test sets.
bench-1Read report →
Speech-to-Text
Fifteen frontier transcription models against real multilingual conversation. Best-in-class still misses every second word of dialectal Arabic.
bench-2Read report →
Real-SWE
Real-SWE, our software-engineering benchmark on private, real world production codebases that we license from real companies.
bench-3
In progress.