← Back to all sparks
Perplexity logo

Perplexity

AI-ASSISTANTS
Velocity8.8

Conversational answer engine combining LLMs with live web search.

Perplexity is selling access to other people's models, and now repricing them weekly.

gateway-apimodel-routingagent-apimcppricing-tiers
Current state
The Gateway API put Anthropic, OpenAI, Google, xAI, and Perplexity models behind one endpoint reachable with an existing Perplexity key, and the MCP server became a remote service hosted by Perplexity with no local installation. Since then the traffic has been commercial rather than structural: GPT-5.6 price cuts, a Sol Fast mode, and successive preset re-pointings — low and fast both now run openai/gpt-5.6-luna, with the fast preset carrying priority processing at twice standard token prices. Frozen configurations have to be updated by hand each time.
Where it's heading
Perplexity is behaving like an infrastructure vendor rather than an answer engine: the differentiator is the credential and the routing, not the model. The preset churn is the visible cost of that position — when the models underneath are someone else's, keeping a named tier meaningful means re-pointing it whenever the market moves, and passing the migration work to customers who pinned a configuration. Inline citations across the search-backed presets remain the one capability that is distinctly Perplexity's own.
Prediction
Expect the preset re-pointings to keep arriving at this cadence and the priority-processing tier to spread beyond the fast preset, since a 2x price band is easier to extend than to justify on one preset alone.

Recent moves

  1. 1mo ago

    Low preset updated

    The low preset re-points to openai/gpt-5.6-luna with minimal reasoning and a 32,768-token output cap. Another instance of the recurring pattern where a named tier changes underneath and anyone on a frozen configuration has to reconcile it manually.

  2. 1mo ago

    GPT-5.6 price cuts and Sol Fast mode

    Price cuts across the GPT-5.6 family plus a Sol Fast mode. Pricing moves are the main lever available to a vendor reselling models it does not train, and passing cuts through quickly is what keeps the gateway worth routing through.

  3. 1mo ago

    Remote MCP Server

    The MCP server becomes a Perplexity-hosted remote endpoint reachable over Streamable HTTP with an API key as bearer token, removing local installation. It is the same consolidation as the Gateway — one credential, one hosted surface — applied to the tool-calling side.

  4. 1mo ago

    New: Gateway API

    ⚡ SPARK

    The release that set the current direction: Perplexity became a routing layer for the whole frontier-model market rather than a seller of its own answers. Every preset re-pointing and price adjustment since is maintenance on that position.

  5. 1mo ago

    Agent API: New Models

    The Agent API added several models including anthropic/claude-opus-5, all at direct first-party token pricing. Catalog breadth is the product here, and first-party pricing is what makes routing through Perplexity a neutral choice rather than a markup.

  6. 1mo ago

    Inline citations for research presets

    Search-backed presets now carry inline numbered citations, with low, medium, and high citing claims drawn from tool results. This is the capability Perplexity owns rather than resells, and it is what distinguishes its agent presets from a bare model call.