¿Qué hay de nuevo en la infraestructura y la orquestación de IA este mes?
ARCHITECT ANÁLISIS DESTACADO

¿Qué hay de nuevo en la infraestructura y la orquestación de IA este mes?

POR

Alex Barrett

FUENTE

Cloud Blog

DATE

READ

3 min de lectura

El ecosistema de IA de Google abarca modelos, herramientas e infraestructura, liberando actualizaciones mensualmente. Las últimas liberaciones incluyen Managed Lustre (versión final, hasta 8 PB), VMs C4N (400 Gbps, 25 …

At Google, AI is a comprehensive endeavor. We offer leading AI models like Gemini and Nano Banana. We integrate AI into tools like Gmail, BigQuery, AlloyDB, Google Cloud Code, and Google Cloud Assist. We provide software frameworks like Gemini Enterprise Agent Platform, JAX, and MaxTest to help build with AI. We also co-design the infrastructure platform, including compute, accelerators like TPUs and GPUs, optimized networks and storage, and orchestration software like GKE and Cluster Director. This allows us to power AI Hypercomputer, enabling industry-wide transformation to an AI and agentic future. This is critical in today’s agentic era, where AI is evolving from answering questions to reasoning and taking action. Companies that want to lead in this next phase of AI need computing infrastructure designed and optimized for these new requirements, so they can innovate faster, deliver compelling user and customer experiences, and optimize for cost and energy efficiency at massive scale. To support this, we are making AI infrastructure and orchestration news at a furious pace. In this blog, we provide a monthly snapshot of the recent launches and milestones that you need to know about, in-depth analyses of architecture and performance tuning, and discussions of specialized use cases, always with pointers to where you can learn more. Keep an eye out for updates to this blog every month. Google Cloud Managed Lustre is now GA, and available in four distinct performance tiers that deliver throughput ranging from 125 MB/s, 250 MB/s, 500 MB/s, to 1000 MB/s per TiB of capacity — with the ability to scale up to 8 PB of storage capacity. The Managed Lustre solution is powered by DDN’s EXAScaler, combining DDN’s decades of leadership in high-performance storage with Google Cloud’s expertise in cloud infrastructure. C4N network and storage optimized VMs are now GA. C4N is our first network- and block-storage-optimized VM series built to eliminate data-transfer bottlenecks. Powered by 5th Gen Intel Xeon Scalable processors and built on Google’s Titanium offloading hardware, it achieves 400 Gbps network bandwidth, 95 million packets per second (MPPS), and up to 25 GiB/s of block storage throughput when paired with Hyperdisk Extreme. GKE Dataplane V2 up to 15K Nodes with Network Policies (GA). This capability enables standard GKE clusters to scale up to 15,000 nodes while maintaining full active Network Policy enforcement, supporting the massive infrastructure needs of large enterprise and AI/ML customers. Co-operative time-slicing in llm-d. If you’re running reinforcement learning (RL) workloads, you can now interleave independent RL jobs onto shared physical hardware, increasing aggregate accelerator duty cycles from a ~40% baseline up to 70% without impacting model convergence or accuracy. Looking to secure your AI supply chain on GKE, deploy AI workloads safely, and cut down on shadow AI? We open-sourced k8s-aibom, a lightweight, unprivileged Kubernetes controller that continuously monitors container clusters to automatically detect running AI runtimes (like vLLM and Triton) and generate standard CycloneDX Machine Learning Bill of Materials (ML-BOMs). C4N network and storage optimized VMs are now GA. GKE Agent Sandbox is now generally available. Agent Substrate is a new open-source project aimed at continuing to push the limits of agentic infrastructure density. Google AI Edge Portal, a solution for testing and benchmarking on-device machine learning (ML) at scale, now supports benchmarking and debugging on-device LLMs. Cloud Storage Rapid, a new family of high-performance storage offerings for AI workloads.