Tag: AI inference

Liquid Cooling at 45°C: The New Thermal Math of AI Factories
Data Centers

Liquid Cooling at 45°C: The New Thermal Math of AI Factories

AI data center cooling has moved from a facilities concern to a hard limit on usable GPU capacity. NVIDIA says its Rubin…

Intel Crescent Island: The Case for an Inference-First Data Center GPU
GPUs & Hardware

Intel Crescent Island: The Case for an Inference-First Data Center GPU

Inference data center GPU design is moving from a broad promise to a clear product decision. Intel has positioned Crescent Island as…

Qualcomm Dragonfly AI300: Why Memory Bandwidth Is an Inference Product
GPUs & Hardware

Qualcomm Dragonfly AI300: Why Memory Bandwidth Is an Inference Product

AI accelerator memory bandwidth has become a product decision, not a footnote beneath peak TOPS. Qualcomm positions its Dragonfly AI300 as a…

AMD MI355X and MLPerf 6.0: How to Read an Inference Benchmark
Benchmarks

AMD MI355X and MLPerf 6.0: How to Read an Inference Benchmark

Inference benchmark results can help size an AI serving fleet, but only if the numbers are read in context. AMD’s MLPerf Inference…

AI Networking in 2026: Why the GPU Is Not the Whole Inference Story
AI Infrastructure

AI Networking in 2026: Why the GPU Is Not the Whole Inference Story

AI networking now shapes inference quality of service as directly as GPU selection. A production agent does not send one prompt to…

AI Context Caching: The New Constraint in Inference Infrastructure
Data Centers

AI Context Caching: The New Constraint in Inference Infrastructure

AI context caching has become a capacity problem, not a minor serving tweak. A production agent may keep a long system prompt,…

AI Inference Hardware in 2026: The Real Bottleneck Is the System
GPUs & Hardware

AI Inference Hardware in 2026: The Real Bottleneck Is the System

AI inference hardware is becoming a system-design problem, not a race to buy the accelerator with the highest quoted throughput. Modern serving…