← Back to all sparks
A

AWS Machine Learning

AI-ASSISTANTS
Velocity10.0

Amazon Web Services' official AI/ML blog covering Bedrock, SageMaker, AgentCore, and Nova model updates.

AWS is widening where its models run and what they cost, not what they can do.

agent-infrastructurebedrockdata-residencyinference-costpartner-integrationsobservability
Current state
The feed keeps its high volume and its roughly even split between launch posts and tutorials. This window is lighter on new agent capability than recent ones: the platform news is OpenAI's GPT-5.6 Terra and Luna becoming available for in-country inference in India, and Deepgram pushing billing, usage and per-GPU metrics out of its own container into customer CloudWatch accounts on SageMaker. Around them sit two how-to posts, an MCP-connected agent harness joining Amazon Quick to fal, and an NVIDIA MPS configuration that cuts ASR GPU cost by 75%. The framework-agnostic AgentCore Evaluations contract from the day before remains the most consequential recent launch.
Where it's heading
The agent-operations buildout described in previous windows is still the spine, but the newest work is about reach and unit economics rather than new capability. Geographic expansion has become a routine cadence: cross-Region inference for GPT-5.6 landed a week ago, India in-country inference follows it, and single-Region Claude Code preceded both, which reads as data residency becoming something AWS expects to tick off per model and per jurisdiction. The partner posts point the same way, since the Deepgram and NVIDIA material is about making someone else's model cheaper or more legible to run on AWS infrastructure rather than about AWS shipping a model.
Prediction
Expect the residency cadence to continue onto the next regulated market rather than the next model, with the cost-per-GPU material continuing to run alongside it. On the evidence of these entries AWS is competing on where and how cheaply a model runs more than on which models it carries.

Recent moves

  1. 19d ago

    Build agentic creative workflows with Amazon Quick and fal

    A build-along rather than a launch: a reusable agent harness wiring Amazon Quick to fal over MCP, demonstrated on an eight-panel storyboard and a music-video prototype. It points the MCP surface AWS has been assembling at a third-party generative provider, but nothing ships here.

    View source ↗
  2. 19d ago

    Introducing OpenAI models on Amazon Bedrock for in-country inferencing in India

    The latest entry in the GPT-5.6-on-Bedrock line and the second on geography in a week, following cross-Region inference on August 20. Terra and Luna now serve inference inside India for customers with local data-processing requirements, which is reach for an already-shipped model rather than new capability.

    View source ↗
  3. 19d ago

    Deepgram deepens Amazon SageMaker AI observability with Enhanced Metrics

    Deepgram, not AWS, closes the gap here: billing, usage and per-GPU metrics leave the vendor container and land in the customer's own CloudWatch account. It fits the platform's push to make self-hosted third-party models legible to the same tooling that watches native ones.

    View source ↗
  4. 19d ago

    Reduce ASR inference costs by 75% with NVIDIA MPS on Amazon EC2

    A configuration walkthrough with real numbers behind it, 75% less GPU infrastructure while holding 92.1 requests per second per GPU, but a how-to rather than a release. Its presence next to the Deepgram post marks how much of this window is about the cost of serving models instead of the models themselves.

    View source ↗
  5. 20d ago

    Evaluate any agent framework with Amazon Bedrock AgentCore Evaluations

    The most substantial launch in this window and the one the current trajectory turns on: evaluation decoupled from the framework, scoring any agent that emits OpenTelemetry. It widens AgentCore's reach across LangGraph, LlamaIndex, the OpenAI Agents SDK, Google ADK, the Claude Agent SDK and Strands rather than adding an operational surface that did not exist before.

    View source ↗
  6. 20d ago

    How GoDaddy transformed its analytics with Amazon Quick

    A two-year business intelligence migration case study carrying customer-side numbers. Standard proof material for Amazon Quick, with no product change behind it.

    View source ↗