AWS
8 stories mentioning AWS, newest first, each with its sources and discussion. Follow to see new ones on your front page.
-
Comparing OpenAI Models on AWS for Workload Efficiency
Organizations often compare models based solely on the price per million tokens, but this overlooks critical factors like how often a model is correct and the number of tokens needed. This post uses an open-source…
1 source primary sourceAWS Machine Learning Blog -
Arm Launches Total Design for Physical AI & Robotics
Arm has unveiled Arm Total Design for Physical AI alongside a new robotics framework to unify standards across various physical industries. This initiative aims to address fragmentation in engineering practices,…
1 sourceAI News -
Simplify TorchServe Workloads with Ray Serve DLC
AWS has introduced the Ray Serve Deep Learning Containers (DLC) to simplify and support TorchServe workloads. TorchServe, no longer actively maintained, requires teams to manage the entire dependency chain, including…
1 source primary sourceAWS Machine Learning Blog -
AWS benchmarks G7 Blackwell GPUs for LLM inference, showing up to 5.6x cost savings
AWS has published detailed benchmarks comparing its new G7 instances, powered by NVIDIA Blackwell GPUs, against previous-generation G5 and G6 instances for deploying large language models on SageMaker AI. The study…
1 source primary sourceAWS Machine Learning Blog -
AWS details cross-account MLflow and SageMaker model governance topologies
AWS has published a technical guide outlining two cross-account architectures for governing machine learning models using MLflow and Amazon SageMaker AI. The post extends previous single-account workflows to address…
1 source primary sourceAWS Machine Learning Blog -
AWS introduces Agent Evaluation Metric for multi-turn AI conversations
AWS has introduced the Agent Evaluation Metric (AEM), a new framework designed to address the limitations of holistic scoring in multi-turn agentic workflows. Traditional evaluation methods often treat agent quality as…
1 source primary sourceAWS Machine Learning Blog -
Amazon SageMaker Adds Prefix-Aware Routing to Cut LLM Latency
Amazon SageMaker Inference now supports prefix‑aware routing, a strategy that sends requests with identical prompt beginnings to the same instance. By keeping the key‑value cache warm, the feature cuts the…
1 source primary sourceAWS Machine Learning Blog -
Model-Agnostic PII Detector for LLMs
A new model-agnostic detector for personally identifiable information (PII) has been released, designed to run on any large language model (LLM) managed on Amazon Bedrock. The detector, evaluated on five public PII…
1 source primary sourceAWS Machine Learning Blog