AI model benchmarks

Each bar is a score with its 95% interval. Orange: ahead alone. Green: tied for the lead. Gray: the rest. Hatched with *: trained on the test data, so not ranked. Average: the unweighted mean of a model's task scores; tasks use different metrics, so read it as a summary, not a ranking.