DigestAI news desk
Company

Amazon SageMaker

2 stories mentioning Amazon SageMaker, newest first, each with its sources and discussion. Follow to see new ones on your front page.

  1. Amazon SageMaker Adds Prefix-Aware Routing to Cut LLM Latency

    Amazon SageMaker Inference now supports prefix‑aware routing, a strategy that sends requests with identical prompt beginnings to the same instance. By keeping the key‑value cache warm, the feature cuts the…

    1 source primary source
    AWS Machine Learning Blog
  2. Amazon SageMaker HyperPod Introduces Model Caching to Cut Cold Starts

    Amazon Web Services has added model caching to its SageMaker HyperPod inference platform, allowing large language models to start serving traffic in seconds instead of minutes. The new feature pre‑loads both the…

    1 source primary source
    AWS Machine Learning Blog