← Back to all sparks
S

Snorkel AI

AI-ASSISTANTS
Velocity5.0

AI data development platform for enterprise model fine-tuning, evaluation, and curation.

Snorkel AI has become an AI evaluation research publisher, not just a data-labeling platform.

ai-evaluationbenchmarkingagent-trainingdata-infrastructurellm-evals
Current state
Snorkel AI's public changelog is entirely research blog posts covering AI agent benchmarks — OSWorld 2.0, Terminal-Bench 3.0 and 4.0, T² scaling laws, and continual learning evaluation. These are not product release notes but research contributions Snorkel is publishing to establish credibility in the AI evaluation and training space. The company appears to be repositioning from data-labeling infrastructure toward AI evaluation and training-data intelligence.
Where it's heading
The consistent theme is that frontier AI agents fail at real-world tasks at far higher rates than benchmarks imply — OSWorld 2.0 shows 20.6% completion on long-horizon computer-use tasks, Terminal-Bench 3.0 has Claude Opus 5 at 43.5%. Snorkel is building a position as the entity that measures this gap and, by extension, sells the training data and tooling to close it. Terminal-Bench becoming a 'continuous benchmark' suggests a product motion, not just research.
Prediction
Expect Snorkel to productize Terminal-Bench and OSWorld-class evaluations as a paid eval-as-a-service offering, targeting enterprise AI teams that need to benchmark agents against real workflows before deployment.

Recent moves

  1. 12d ago

    OSWorld 2.0: Frontier Agents Complete Only 1 in 5 Long-Horizon Computer-Use Tasks

    OSWorld 2.0 introduces 108 long-horizon computer-use tasks where even the best agents complete only 20.6% — a direct challenge to optimistic agent capability claims. For Snorkel, publishing this research positions them as a credible evaluator in the agentic AI market, which feeds demand for their training-data tooling.

    View source ↗
  2. 14d ago

    Fable 5.1 on Frontier Coding Tasks: Efficient Successes, Distinct Failure Modes

    Snorkel evaluates Fable 5.1 against Opus 5 on proprietary Terminal-Bench+ tasks — finding Fable competitive on most categories but weaker on terminal-heavy and build/dependency work. Publishing head-to-head model evals with proprietary benchmarks is itself a product signal: it demonstrates both benchmark depth and Snorkel's model-agnostic evaluation position.

    View source ↗
  3. 18d ago

    Terminal-Bench 4.0: Why Continuous Benchmarks Require Continuous QA

    Terminal-Bench 4.0 addresses benchmark saturation by building a continuous QA process — tasks are actively maintained and updated as models improve, preventing the rapid obsolescence that plagues static benchmarks. This is a meaningful infrastructure move for evaluation reliability, consistent with Snorkel's trajectory toward productized evaluation.

    View source ↗
  4. 22d ago

    Why Frontier Agents Fail Real Engineering Work: Two Terminal-Bench 3.0 Task Deep Dives

    Deep dives into two Terminal-Bench 3.0 task failures expose where frontier agents break down on authentic engineering work — a concrete demonstration of the gap between benchmark marketing and real-world agent reliability that underpins Snorkel's evaluation thesis.

    View source ↗
  5. 26d ago

    Continual Learning Bench: measuring whether AI systems actually improve with experience

    A research presentation on continual learning evaluation — framed as a third-party contribution (Nicholas Roberts) rather than a Snorkel product release. Relevant to the evaluation positioning but not a Snorkel product change.

    View source ↗
  6. 28d ago

    Train-to-Test (T²) Scaling Laws: Why Reasoning Models Should Be Overtrained

    A presentation of T² scaling laws showing reasoning models should be overtrained beyond Chinchilla-optimal ratios — academic research content, not a Snorkel product feature or capability expansion.

    View source ↗