HomeReleasesSnorkel AI Backs Open-Source Benchmarks to Stress-Test Front
Releases

Snorkel AI Backs Open-Source Benchmarks to Stress-Test Frontier Models

As AI systems evolve faster than traditional evaluation methods, Snorkel AI has unveiled the inaugural group of projects supported by its $3 million Open Benchmarks Grants program. The initiative aims to provide researchers with the tools needed to rigorously measure how models perform on complex, real-world professional tasks.

Snorkel AI Backs Open-Source Benchmarks to Stress-Test Frontier Models

The grants provide selected teams with a mix of direct funding, engineering collaboration, and access to platform resources. Among the projects receiving support are Agents' Last Exam, a collaboration with UC Berkeley RDI that tasks agents with complex professional workflows across 55 sub-industries, and OSWorld 2.0, which evaluates computer-use agents in self-hosted web and desktop environments. Other funded efforts include the Continual Learning Bench and the SlopCode Bench, designed to monitor how code quality fluctuates when agents repeatedly modify their own work.

Fred Sala, a steering committee member and assistant professor at the University of Wisconsin–Madison, noted that these projects address the field's most difficult evaluation hurdles. Beyond the grant recipients, Snorkel AI also recently collaborated with Princeton University and the University of Wisconsin–Madison to develop Senior SWE-Bench, a framework focused on high-level software engineering tasks. Supported by partners including Hugging Face and Together AI, the grant program continues to accept applications on a rolling basis, seeking to standardize how the industry defines and measures machine intelligence.

Comments (0)

Leave a comment

No comments yet. Be the first!