← Back to all sparks
C

Comet

AI-ASSISTANTS
Velocity5.0

ML experiment tracking and LLM observability platform, including Opik for evaluating LLM apps.

Comet's feed has gone all-in on category education while Opik ships out of frame.

llm-observabilityopikagent-tracingcost-intelligencedeveloper-educationseo
Current state
Every entry in this window is educational or category content: what AI observability is, how to select a model per agentic task, which observability platforms rank in 2026, a from-scratch treatment of diffusion language models, and a build-log for a self-grading F1 radio RAG pipeline. The real Opik engineering, Agent Diagnostics for cross-trace analysis and the cost work behind it, has now fallen out of the recent window entirely. The old experiment-tracking identity is not visible at all.
Where it's heading
Comet has completed its pivot from ML experiment tracking to LLM and agent observability, and the writing is aimed at defining the category rather than documenting releases. Cost is the recurring hook across the teaching posts, with model selection framed as a billing problem, which suggests the wedge is budget pressure rather than debugging alone. The ratio has tipped far enough that buyer education is now running well ahead of anything shipped in view.
Prediction
On this evidence the feed will keep publishing category guides faster than product notes, so what Opik actually ships next is not readable here. The consistent framing of model choice as a cost decision is the one thread pointing at where the product work is going.

Recent moves

  1. 19d ago

    Diffusion Language Models, From Scratch to Production

    A technical explainer tracing how the term language model narrowed after GPT-3.5, then walking diffusion-based alternatives through to production. Well outside Opik's surface and typical of a feed now dominated by teaching content.

    View source ↗
  2. 27d ago

    What is AI Observability? A Complete Guide to Debugging and Monitoring Modern AI Systems at Scale

    A category-defining guide opening on green infrastructure dashboards next to users asking why the agent did that. It is the clearest statement of the market Comet wants to own, with no product news in it.

    View source ↗
  3. 29d ago

    LLM Model Selection: How to Pick the Right Model for Every Agentic Task

    Model selection framed as a billing problem, starting from a team that defaulted every tool call to the flagship model and never revisited it. Education, but it names the cost angle the product work keeps returning to.

    View source ↗
  4. 29d ago

    Best LLM Observability Tools of 2026: Top Platforms & Features

    A ranked roundup of LLM observability platforms for 2026. Search-driven comparison content of the kind a vendor publishes to sit at the top of its own category page.

    View source ↗
  5. 1mo ago

    I Built a RAG Pipeline for F1 Team Radio, Then Made It Grade Itself

    A build log for a RAG pipeline over Formula 1 team radio that grades its own output, presented as five commands. A demonstration of the evaluation workflow rather than a change to it.

    View source ↗
  6. 1mo ago

    One Prompt, 24 Versions: How Digibee Builds Prompts with Opik to Power Their AI-Native Integration Platform

    A customer story about Digibee iterating one prompt across 24 versions in Opik. Proof material showing the prompt-versioning workflow in use, with nothing newly shipped.

    View source ↗