Snorkel AI
AI data development platform for enterprise model fine-tuning, evaluation, and curation.
Snorkel AI has become an AI evaluation research publisher, not just a data-labeling platform.
◆Recent moves
- 12d ago
OSWorld 2.0: Frontier Agents Complete Only 1 in 5 Long-Horizon Computer-Use Tasks
OSWorld 2.0 introduces 108 long-horizon computer-use tasks where even the best agents complete only 20.6% — a direct challenge to optimistic agent capability claims. For Snorkel, publishing this research positions them as a credible evaluator in the agentic AI market, which feeds demand for their training-data tooling.
View source ↗ - 14d ago
Fable 5.1 on Frontier Coding Tasks: Efficient Successes, Distinct Failure Modes
Snorkel evaluates Fable 5.1 against Opus 5 on proprietary Terminal-Bench+ tasks — finding Fable competitive on most categories but weaker on terminal-heavy and build/dependency work. Publishing head-to-head model evals with proprietary benchmarks is itself a product signal: it demonstrates both benchmark depth and Snorkel's model-agnostic evaluation position.
View source ↗ - 18d ago
Terminal-Bench 4.0: Why Continuous Benchmarks Require Continuous QA
Terminal-Bench 4.0 addresses benchmark saturation by building a continuous QA process — tasks are actively maintained and updated as models improve, preventing the rapid obsolescence that plagues static benchmarks. This is a meaningful infrastructure move for evaluation reliability, consistent with Snorkel's trajectory toward productized evaluation.
View source ↗ - 22d ago
Why Frontier Agents Fail Real Engineering Work: Two Terminal-Bench 3.0 Task Deep Dives
Deep dives into two Terminal-Bench 3.0 task failures expose where frontier agents break down on authentic engineering work — a concrete demonstration of the gap between benchmark marketing and real-world agent reliability that underpins Snorkel's evaluation thesis.
View source ↗ - 26d ago
Continual Learning Bench: measuring whether AI systems actually improve with experience
A research presentation on continual learning evaluation — framed as a third-party contribution (Nicholas Roberts) rather than a Snorkel product release. Relevant to the evaluation positioning but not a Snorkel product change.
View source ↗ - 28d ago
Train-to-Test (T²) Scaling Laws: Why Reasoning Models Should Be Overtrained
A presentation of T² scaling laws showing reasoning models should be overtrained beyond Chinchilla-optimal ratios — academic research content, not a Snorkel product feature or capability expansion.
View source ↗