Amazon SageMaker HyperPod Introduces Model Caching to Cut Cold Starts
By slashing inference cold‑start latency, enterprises can deploy larger models with tighter scaling, improving user experience and reducing operational costs in AI‑driven services.
-
Amazon SageMaker Adds Prefix-Aware Routing to Cut LLM Latency
Amazon SageMaker Inference now supports prefix‑aware routing, a strategy that sends requests with identical prompt beginnings to the same instance. By keeping the key‑value cache warm, the feature…
1 source primary source -
Amazon SageMaker HyperPod Introduces Model Caching to Cut Cold Starts
Amazon Web Services has added model caching to its SageMaker HyperPod inference platform, allowing large language models to start serving traffic in seconds instead of minutes. The new feature…
1 source primary source -
Alibaba Releases Open‑Weight Qwen3.8‑2.4T‑A95B; Deploys on SageMaker HyperPod
Alibaba’s Qwen team unveiled the Qwen3.8‑2.4T‑A95B model on August 12, 2026, marking the first open‑weight release of a Qwen‑Max‑class model. The 2.4 trillion‑parameter network activates 95 billion…
1 source primary source -
Heurist Finance Builds AI‑Native Investment Workbench on Amazon Bedrock AgentCore
Heurist Finance, a fintech startup, leveraged Amazon Bedrock AgentCore to create an AI‑native investment workbench that consolidates market data, filings, news, portfolio construction, and scenario…
1 source primary source -
Simplify TorchServe Workloads with Ray Serve DLC
AWS has introduced the Ray Serve Deep Learning Containers (DLC) to simplify and support TorchServe workloads. TorchServe, no longer actively maintained, requires teams to manage the entire…
1 source primary source -
AWS details cross-account MLflow and SageMaker model governance topologies
AWS has published a technical guide outlining two cross-account architectures for governing machine learning models using MLflow and Amazon SageMaker AI. The post extends previous single-account…
1 source primary source