Category index

Architect

325 articles

325 ARTICLES

What Google Cloud announced in AI this month
ARCHITECT

What Google Cloud announced in AI this month

Editor’s note: Want to keep up with the latest from Google Cloud? Check back here for a monthly recap of our latest updates, announcements, resources, events, learning opportunities, and more. Since launching the Gemini Enterprise Agent Platform a few months ago, we’ve watched businesses move from basic experiments to serious, production-grade builds. We want to make it even easier — and more secure — for you to scale those systems. Along with a batch of new platform updates, this month we’ve put together 13 practical demos and 20 diagnostic questions to help your engineering teams align on a strong architectural blueprint. Let’s dive in! Top announcements What’s new in Gemini Enterprise Agent Platform: In this helpful recap, we announced some of our most popular capabilities are available for everyone, from Agent Runtime to Agent Identity. Now in preview: Find and fix software vulnerabilities with CodeMender: As adversarial AI threats accelerate attacks on code, security teams must counter them with machine-speed defenses that can automate code remediation and fight AI with AI. You can learn more about CodeMender and review the documentation here. Solve harder problems with AlphaEvolve, now available to everyone on Google Cloud: AlphaEvolve is a code optimization and discovery agent built on top of Gemini that helps solve the hardest algorithmic problems and achieve breakthroughs for your business and research. Thought leadership (editor’s pick): If automation requires delegation, then delegation requires trust. But letting an AI agent run on its own is a big leap for any business. While the productivity gains are clear, the fear of losing control is very real. This month, we sat down with our experts to discuss how leaders can navigate this shift by focusing on transparency, predictability, and setting clear boundaries for how agents handle weird data exceptions. Here’s our picks for the month to learn more. How leaders can scale AI by trading control for trust (Q&A): We sat down with Michael Gerstenhaber, VP of Product Management for Gemini Enterprise, to discuss why the future of AI is about defining safe boundaries. What makes an AI agent trustworthy: Context is fast becoming one of the most valuable assets a company owns. Prajakta Damle, Senior Director, Product Management, shares what it takes to get trustworthy AI right. News you can use: What if you’re looking for a steer on your basic foundation, inspiration for recipes, or some inspiration? Take a look at some of our favorite how-to guides from July: Automate your agent development lifecycle using any coding agent: Stuck prototyping? With Agents CLI skills, you can go through the different phases of the entire agent lifecycle without ever leaving your coding agent. Why AI apps fail in production (And how Google solved it): Only 5% of AI prototypes make it to production, and the other 95% fall into the validation abyss. How can you move confidently into production? 13 hands-on demos to build on Gemini Enterprise Agent Platform: Not sure where to start with Agent Platform? Here’s 13 ways you can stir up some creativity. Stay tuned for monthly updates on Google Cloud’s AI announcements, news, and best practices. For a deeper dive into the latest from Google Cloud customers, read our monthly recap, Cool stuff customers built. aside_block <ListValue: [StructValue([(’title’, ‘$300 in free credit to try Google Cloud AI and ML’), (‘body’, <wagtail.rich_text.RichText object at 0x7fa074124250>), (‘btn_text’, ‘Start building for free’), (‘href’, ‘http://console.cloud.google.com/freetrial?redirectPath=/vertex-ai/’), (‘image’, None)])]> June Our main focus in June was helping your teams build, scale, and secure AI. Today, we’re sharing a fresh roundup of updates designed to help you run smarter, more secure applications while keeping everything under your control. We even shared a cool virtual shopping demo at Cannes to show how retailers can make product discovery more exciting. Let’s dive in! Top announcements Introducing the Open Knowledge Format: We introduced the Open Knowledge Format (OKF), an open specification that formalizes the LLM-wiki pattern into a portable, interoperable format. This is a vendor-neutral, agent- and human-friendly standard for representing the metadata, context, and curated knowledge that modern AI systems need. Collaboration with Apple on its expanded Private Cloud Compute (PCC) systems: Our collaboration with Apple is built on a foundation of deep commitment to privacy that leverages Google Cloud’s security and privacy technologies. At the heart of this collaboration is our Confidential Computing portfolio and our Titanium security architecture. Claude Fable 5: Available on Google Cloud: Claude Fable 5, Anthropic’s latest frontier model, is now generally available on Google Cloud. This launch is the latest proof point of our ongoing commitment to bring the industry’s latest models straight to our Agent Platform. Cloud Atelier: How Gemini Enterprise is helping restyle the retail playbook: This year at Cannes, we showcased Cloud Atelier — a destination-based, virtual shopping experience that highlights how retail brands can turn this classic dilemma into an exciting moment of product discovery. Thought leadership (editor’s pick): How Google Cloud Security uses AI internally: To counter machine-speed, AI-driven threats, we’ve worked hard to transition Google Cloud’s security posture to an autonomous, proactive model. By embedding specialized AI agents directly into our software development lifecycle (SDLC), we’ve created automated guardrails that protect code at a scale and speed unreachable by human teams — and we’re taking steps to make those same guardrails widely available. The 4 lessons that guided AI Threat Defense: We introduced Chris Betz as the new CISO of Google Cloud. For his first Cloud CISO Perspectives, Chris shares four key lessons we learned about using AI to the defender’s advantage while building AI Threat Defense. News you can use: 5 lessons from red teaming AI applications: To help you build AI securely, Mandiant has developed a proactive, risk-based approach centered on the Good AI Assessment (GAIA) Top 10, outlined in our new report, Secure Development of Generative AI Applications: A Proactive Approach. How to unlock true ROI in software development – a deep dive into the latest DORA research: To help you evaluate the costs and business benefits of AI, we recently shared the DORA: ROI of AI-assisted software development report. This research offers a practical approach to help your team work through early adoption challenges, align engineering plans, and drive business growth. Agent Factory Recap: 100X engineering with AI agents in Google Antigravity 2.0: In this episode of the Agent Factory, Shir Meir Lador, Head of AI Engineering, Google Cloud Developer Relations, sat down with Rody Davis, one of Google’s top agentic engineers. They dive into the massive shift from traditional IDEs to agent-first platforms, the reality of code reviews in an AI-driven world, and how to use “skills” to perform at a 100X level. Stay tuned for monthly updates on Google Cloud’s AI announcements, news, and best practices. For a deeper dive into the latest from Google Cloud customers, read our monthly recap, Cool stuff customers built. May We’ve had a busy month! Between announcing Gemini Spark and Gemini 3.5 at Google I/O – and unveiling Google AI Threat Defense, our latest AI-powered cybersecurity solution, we had a lot to share with Google Cloud customers. Keeping up with the latest news takes time, so we gathered the most important announcements, thought leadership, and technical guides in one place to help you quickly catch up. To learn more about our I/O announcements, here’s everything you need to know for Google Cloud customers, and top news for startups. Top announcements Introducing Google AI Threat Defense to help you outpace the adversary: Google Cloud is introducing a comprehensive AI-powered cybersecurity solution — Google AI Threat Defense — an always-on autonomous security platform. Learn more here. Gemini 3.5: Our latest family of models combines frontier intelligence with action – starting with Gemini 3.5 Flash. Gemini Omni: Our new model is a leap forward in world understanding, multimodality, and editing, letting you generate any output from any input, starting with video. Google Antigravity: Google Antigravity’s expanded capabilities and new integration with Agent Platform bring agentic development to your entire organization. Gemini Spark: For Gemini Enterprise and Workspace customers, Gemini Spark is your 24/7 personal AI agent that helps you work more efficiently by autonomously taking action on your behalf, under your direction. Google Workspace: Google Pics, our new image generation and editing tool, and new voice features in Gmail, Docs and Keep, help reimagine how you work. Managed Agents API on Agent Platform: Allows developers to build and run custom agents inside secure, Google-hosted environments that seamlessly integrate with Agent Platform. CodeMender: A powerful AI security agent provided through Agent Platform, CodeMender can help find and fix vulnerabilities in your code. Nano Banana 2 and Nano Banana Pro are generally available: Available today via Gemini Enterprise Agent Platform, organizations are already putting the models to work. Learn more here. Thought leadership (editor’s pick): Cloud CISO Perspectives: How Google + Wiz changes multicloud strategy for CISOs: Vinod D’Souza, director, Office of the CISO, shares highlights from his RSA Conference fireside chat with Anthony Belfiore, chief strategy officer, Wiz. While threat actors have seen gains from the adversarial misuse of AI, Google and Wiz are tackling these challenges head-on by combining Wiz’s deep cloud telemetry with Google’s world-class AI and quantum research to help CISOs and their organizations meet the needs of the agentic enterprise era. Read more here. News you can use: What Google I/O ‘26 means for developing agents on Google Cloud: Dig deep into how Gemini Enterprise Agent Platform and the new developer tools shared at I/O fit together, unpack the spectrum of choice for building, and share what we’d actually try first. Learn more here. Five must-have guides to move agents into production with Gemini Enterprise Agent Platform: Here is a look back at our five-part series covering the architecture patterns and best practices you need to move your agents into production. Learn more here. How to build an AI-ready security program for the public sector: From industrial control systems to decades-old municipal databases, here’s our CISO guidance to prep AI-ready security programs for the public sector. Learn more here. Stay tuned for monthly updates on Google Cloud’s AI announcements, news, and best practices. For a deeper dive into the latest from Google Cloud customers, read our monthly recap, Cool stuff customers built. April We hosted Google Cloud Next in Las Vegas on April 22, announcing incredible innovations from Gemini Enterprise Agent Platform to our eight-generation TPUs. We also expanded the Gemini Enterprise app in collaborative ways – now, with new features like Projects, you can work side-by-side with your agents and colleagues. If you missed the livestream, take a look at our Day 1 recap. It’s been incredible to see how customers have been applying AI in thousands of ways — so far, we’ve counted more than 1,300 examples. Top announcements 1. Gemini Enterprise Agent Platform: Our new, comprehensive platform to build, scale, govern, and optimize agents. Moving forward, all Vertex AI services and roadmap evolutions will be delivered exclusively through the Agent Platform, rather than as a standalone service, to power the next generation of agent development. The platform is designed around four core pillars — build, scale, govern, and optimize — that allow teams to collaborate seamlessly. Learn more about Agent Platform here. 2. Gemini Enterprise app has all the key components to let teams discover, create, share, and run AI agents in a single environment. At Next ‘26, we introduced several new capabilities in the Gemini Enterprise app: Agent Designer uses the same no-code agent designer experience of Agent Platform and lets employees build sophisticated schedule- and trigger-based agents using any enterprise connector. It gives you a virtual flowchart of your agent, allowing you to inspect, test, and approve workflows, ensuring total transparency for executing critical business processes. Long-running agents are designed to execute complex business processes. They can work autonomously in secure cloud sandboxes, giving agents the ability to orchestrate business logic, write code to build custom tools, and complete multi-step work like reconciliation activities or sales prospect sequencing — without needing constant prompting. Inbox in Gemini Enterprise provides a central location to monitor, guide, and help manage all of your agent activity, including your long-running agents. Notifications are intuitively categorized into actionable groups like “Needs your input,” “Errors,” and “Completed.” Projects create a dedicated space where the agent’s memory is confined to the files and conversations your team adds. By connecting it to data sources including Google Drive, NotebookLM, and Google Group Chats, the agent becomes an expert on a specific topic and can provide team members daily briefings or status updates without digging through months of documents. Skills create simple shortcuts using an “@” mention for repetitive tasks such as applying brand guidelines, formatting a report, and accessing specific data. Canvas gives our customers an interactive editor directly within Gemini Enterprise. It allows teams to easily create and edit Docs and Slides, and even export to Microsoft 365 files, within the same experience. Agent Gallery provides access to third-party agents from partners like Adobe, Atlassian, Lovable, and ServiceNow, and is adding more third-party connectors for Asana, Mailchimp, Workday, and more. These integrations enable your agents to retrieve data and execute tasks with your systems-of-record. 3. AI Hypercomputer: Designed specifically for demanding AI workloads, our AI Hypercomputer is an advanced, purpose-built architecture that unites performance-optimized hardware for compute, storage, networking, open software and machine learning frameworks — as well as flexible consumption models — into a single, integrated system. We are announcing innovations at every layer of the AI Hypercomputer: TPU 8t, optimized for training, uses breakthrough Inter-Chip Interconnect (ICI) technology to scale up to 9,600 TPUs and 2 PB of shared, high-bandwidth memory in a single superpod. It achieves 3x the processing power of Ironwood and delivers up to 2x more performance/Watt. TPU 8i, optimized for inference, uses our new Boardfly topology to directly connect 1,152 TPUs in a single pod. It features 3x more on-chip SRAM compared to previous versions to host larger KV caches entirely on-silicon and integrates a specialized Collectives Acceleration Engine. Taken together, TPU 8i delivers 80% better performance per dollar for inference than the prior generation, enabling millions of concurrent agents to run cost-effectively. 4. The Agentic Data Cloud: A new data architecture built for the speed and scale of agentic AI. The Agentic Data Cloud delivers an AI-native architecture, allowing agents to perceive, reason, and act on your behalf in real-time, including: Cross-Cloud Lakehouse, standardized on Apache Iceberg, is our Lakehouse that enables you to leave your data in AWS or Azure (coming later this year) while querying it instantly — without the friction of vendor lock-in or the cost of data movement Knowledge Catalog constructs a unified, dynamic context graph of your entire business enabling you to ground agents in all of your business data and semantics. With Smart Storage and the Object Context API, files in Google Cloud Storage are instantly tagged and enriched with metadata before an agent touches them. Then our Knowledge Engine uses Gemini to autonomously tag, define logic and instantly map complex relationships across your entire enterprise, providing the semantic definition your agents have been missing. 5. Protecting the agentic enterprise: Security built for the AI era. Our full-stack AI approach, from the chips to the models, gives you a competitive advantage with better integration and velocity to help protect customers. Not only can Google action insights from the world’s largest threat observatory and Mandiant frontline experts, but we also bring cutting-edge insights and breakthroughs from Google DeepMind, to help make your platforms more secure. Agentic defense: Three new agents in Google Security Operations can help hunt threats, engineer detections, and provide context on third parties. You can build your own security agents with remote Google Cloud model context protocol (MCP) server support for Google Security Operations, now generally available. You can also access the MCP server client directly from the Google Security Operations chat interface, available in preview. Protecting AI and cloud apps across any infrastructure with Wiz: Newly expanded AI coverage helps build secure agents across clouds and AI studios. New AI-Bill of Materials in development tools can help secure AI-generated code and mitigate the risk of shadow AI. Learn more. Securing agents and the agentic web: Model Armor can integrate with Agent Gateway, and new Agent Identities provide more layers of defense against shadow AI. Google Cloud Fraud Defense, the next evolution of reCAPTCHA, offers agent-specific capabilities that can help secure the agentic web as well as the entire user and customer journey. Trusted Cloud: We’re simplifying permissions with modern IAM, and advancing Google Cloud security with new capabilities in Security Command Center plus new innovations in data and network security. New partner-supported workflows for Google Security Operations: This new robust cohort of partner integrations includes partners developing their own agentic security operations centers (SOCs). You can catch up on all our security announcements from Next ‘26 here. News you can use Guide to prompting Gemini 3.1 Flash TTS (text-to-speech): The new TTS model introduces a high level of controllability by allowing you to steer the delivery using more than 200 audio tags. We’ll share how to get strong results from the model, whether you are building accessible gaming soundtracks, banking systems, or audiobooks. Learn more about the model here. Ultimate prompting guide for Lyria 3 models: Lyria 3, Google’s family of music-generation models, is designed to give you granular control over vocals, instrumentation, and arrangement. So we spent weeks testing against every musical genre and use case we could imagine. We put together this guide to share exactly what we learned and how you can get the best results. How to find the sweet spot between cost and performance: This guide will walk you through Google Cloud’s flexible gen AI infrastructure options, showing you how to find that sweet spot on the efficient frontier between cost and performance. We’ll start with the foundational pay-as-you-go (PayGo) models and then explore how to layer on more specialized options to build a robust and cost-effective gen AI strategy. Essential AI and cloud security now on by default: To support the next generation of AI innovators, we are offering on by default essential AI security and cloud security in Security Command Center Standard. Securing AI inference on GKE with Model Armor: Here’s how to secure AI inference on Google Kubernetes Engine with Model Armor and high-performance storage. Cloud CISO Perspectives: AI, security, and the workforce of the future: You can’t bring traditional security to an AI fight, so how do we defend against AI-powered attacks, boost defenders with AI, and secure AI use? Drop in on this RSA Conference fireside chat between Francis deSouza, Google Cloud COO and President, Security Products, and Nick Godfrey, senior director, Office of the CISO. Stay tuned for monthly updates on Google Cloud’s AI announcements, news, and best practices. For a deeper dive into the latest from Google Cloud customers, read our monthly recap, Cool stuff customers built. March March was a busy month for our AI teams. We launched Gemini Embedding 2, rolled out a highly cost-effective Veo 3.1 Lite model, and officially welcomed the Wiz team to Google Cloud to help redefine security in the AI era. Alongside these launches, we created comprehensive guides to help you get the most out of these models, from prompting formulas for Nano Banana 2, to practical advice for optimizing your TPU training. Here’s a quick look at the latest news and resources to help your team build what’s next. Top hits: Gemini Embedding 2: Our first natively multimodal embedding model: Gemini Embedding 2 is our first natively multimodal embedding model that maps text, images, video, audio and documents into a single embedding space, enabling multimodal retrieval and classification across different types of media — and it’s available now in public preview. Build with Veo 3.1 Lite, our most cost-effective video generation model: This model empowers developers to build high-volume video applications, at less than 50% of the cost of Veo 3.1 Fast, but with the same speed. This rounds out the Veo 3.1 model family, giving developers flexibility based on needs. For Cloud customers, it’s now available on Vertex AI. Here’s a fun bonus: Check out our ultimate prompting guide for Veo 3.1 to get started. Veo 3.1 Lite Welcoming Wiz to Google Cloud: Redefining security for the AI era: Google has completed its acquisition of Wiz, a leading cloud and AI security platform. The Wiz team will join Google Cloud, and we will retain the Wiz brand. With the addition of Wiz, we will provide customers with a comprehensive platform to secure their cloud and hybrid environments, as well as accelerate threat prevention, detection, and response. Gemini 3.1 Flash Live: Making audio AI more natural and reliable: We’ve improved 3.1 Flash Live’s overall quality, making it more reliable for developers and enterprises to build voice-first agents that can complete complex tasks at scale. On ComplexFuncBench Audio, a benchmark that captures multi-step function calling with various constraints, it leads with a score of 90.8% compared to our previous model. News you can use: The ultimate Nano Banana prompting guide: This is a must-read for anyone working with Nano Banana. We spent weeks testing Nano Banana 2 and Nano Banana Pro against every use case we could imagine to test its limits. We put together this guide to share exactly what we learned and how you can get the best results. Here’s an example formula: [Reference images] + [Relationship instruction] + [New scenario] A developer’s guide to training with Ironwood TPUs: In this guide, we hear from Lillian Yu, CPA, CA , Product Strategy and Operation, and Liat Berry, Product Manager, on five strategies within the JAX and MaxText ecosystems designed to help developers refine training efficiency and hit peak performance on Ironwood hardware. How to build production-ready AI agents with Google-managed MCP servers: In this guide, we anchor on a specific example. Cityscape is a demo agent built with Google’s Application Development Kit (ADK) that turns a simple text prompt — like “Generate a cityscape for Kyoto” — into a unique, AI-generated city image. Check out the guide to learn more. Stay tuned for monthly updates on Google Cloud’s AI announcements, news, and best practices. For a deeper dive into the latest from Google Cloud customers, read our monthly recap, Cool stuff customers built. February In February, we’re giving developers more reasoning power with Gemini 3.1 Pro and Claude 4.6, and faster creative scaling with Nano Banana 2. We’re also opening up new training programs and step-by-step guides to help you tackle the hardest parts of the AI lifecycle, from capacity planning to mounting defenses against AI-powered attacks. Here’s a rundown of our latest news, tools, and resources to help you build what’s next. Top hits Pro-level image generation gets faster and more accessible with Nano Banana 2: To build creative that stands out, you need models that naturally integrate into your workflows and scale with ease. Check out our blog to see how this comes to life (and how customers are putting the model to work). Introducing Gemini 3.1 Pro on Google Cloud: Gemini 3.1 Pro is a clear step forward in reasoning, designed to solve tougher problems, giving you the reasoning depth your business needs. Gemini 3.1 Pro is available starting today in preview in Vertex AI and Gemini Enterprise. Developers can access the model in preview via the Gemini API in Google AI Studio, Android Studio, Google Antigravity, and Gemini CLI. Announcing Claude Opus 4.6 and Claude Sonnet 4.6 on Vertex AI: Now generally available on Vertex AI, explore our sample notebook to get started and visit our documentation for comprehensive pricing and regional availability details. New AI threats report: Distillation, experimentation, and integration: John Hultquist, chief analyst, Google Threat Intelligence Group, details what security leaders should know from our newest AI threat report on experimentation, integration, and distillation attacks. News you can use A developer’s guide to production-ready AI agents: To help developers work through these challenges, we’ve published a collection of guides covering the full agent lifecycle. These resources first appeared during Kaggle’s 5 days of AI Agents Intensive, and they’ve proven so popular and useful, we wanted to make sure a wider audience had access, as well. Gemini Enterprise Agent Ready (GEAR) program now available: We opened the Gemini Enterprise Agent Ready (GEAR) learning program to everyone. As a new specialized pathway within the Google Developer Program, GEAR empowers developers and pros to build and deploy enterprise-grade agents with Google AI. Your guide to Provisioned Throughput (PT) on Vertex AI: Check out this deep-dive blog designed to show you the resources available to you today on Vertex AI, and how you can get started capacity planning. How AI can boost defenders, from defense in depth to the cyber kill chain (Q&A): We know that defenders are also developing powerful AI tools, but what’s still unknown is what it could mean for enterprise software ownership if companies have to constantly mount AI-directed defenses at AI-powered attacks? Stay tuned for monthly updates on Google Cloud’s AI announcements, news, and best practices. For a deeper dive into the latest from Google Cloud customers, read our monthly recap, Cool stuff customers built. Janurary We used to have to learn the language of computers. In 2026, they’re learning ours. We kicked off the year by exploring the future of agentic commerce, where AI agents navigate the web to find and buy products for us. Our leaders call this the “invisible shelf” — a world where commerce isn’t tied to a specific website. To make this reality scalable, we announced the Universal Commerce Protocol (UCP), a shared language that allows agents and retailers to understand each other. We brought that same fluency to our creative and technical tools: Updates to Veo 3.1 allow creators to use simple inputs — like reference images — to generate precise, mobile-ready video. Natural language queries: With Comments to SQL in BigQuery, we’re removing the language barrier to data. Engineers can now write queries by describing their intent in natural language, prioritizing the question over the code. Let’s dive in. Top hits 1. Gemini Enterprise for Customer Experience (CX): Specifically built for agentic retail, this platform transforms fragmented search, commerce and service touch points into one seamless journey — whether you need a shopping assistant, a support bot, agentic search or help with merchandising. 2. We announced Universal Commerce Protocol (UCP): A new open standard for agentic commerce that works across the entire shopping journey — from discovery and buying to post-purchase support. UCP establishes a common language for agents and systems to operate together across consumer surfaces, businesses and payment providers. So instead of requiring unique connections for every individual agent, UCP enables all agents to interact easily. UCP is built to work across verticals and is compatible with existing industry protocols like Agent2Agent (A2A), Agent Payments Protocol (AP2) and Model Context Protocol (MCP). 3. We updated Veo 3.1, including improvements to Ingredients to Video and Portrait mode: Veo is getting more expressive, with improvements that help you create more fun, creative, high-quality videos based on ingredient images, built directly for the mobile format. This includes: Improvements to Veo 3.1 Ingredients to Video, our capability that lets you create videos based on reference images. Native vertical outputs for Ingredients to Video (portrait mode) to power mobile-first, short-form video creation. State-of-the-art upscaling to 1080p and 4K resolution 1 for high-fidelity production workflows. These updates are launching in the Gemini app, YouTube, Flow, Google Vids, the Gemini API and Vertex AI. 4. Vibe querying with comments-to-SQL: Crafting complex SQL queries can be challenging. Often, engineers simply want to express their data needs in plain English directly within their SQL workflow. That’s why we’re introducing Comments to SQL in BigQuery. This feature makes writing queries using natural language – ‘vibe querying’ – a reality. Learn more in the blog. News you can use Mastering Gemini CLI: Your complete guide from installation to advanced use-cases: We’ve teamed up with DeepLearning.ai and are excited to announce a free course – Gemini CLI: Code & Create with an Open-Source Agent. This course isn’t just for developers; we dive into practical use cases for various tasks such as data analysis, content creation, and personalized learning. How Google SREs use Gemini CLI to solve real-world outages: In this article, we’ll delve into real scenarios that Google SREs are solving today using Gemini 3 (our latest foundation model) and Gemini CLI—the go-to tool for bringing agentic capabilities to the terminal. Getting started with Gemini 3: Deploy your first Gemini 3 app to Google Cloud Run: In this blog, we will show you how to vibe code your first app—which leverages the Gemini 3 Flash Preview model and deploy it as a publicly accessible URL on Google Cloud Run. Google AI Studio lets you go from idea to app quickly by using natural language to generate fully functional apps using the power of Gemini 3. Practical guidance: Building with the Secure AI Framework (SAIF) on Google Cloud: We know that security and data privacy are the top concern for executives when evaluating AI providers, and security is the top use case for AI agents in a majority of industries. To help you build AI boldly and responsibly, here’s our guide to developing AI with the Secure AI Framework (SAIF) on Google Cloud. The truths about AI hacking that every CISO needs to know (Q&A): How will AI boost threat actors? And what can chief information security officers do about it? Google’s Heather Adkins, vice-president, Security Engineering, explores how securing the enterprise is about to change. Stay tuned for monthly updates on Google Cloud’s AI announcements, news, and best practices. For a deeper dive into the latest from Google Cloud customers, read our monthly recap, Cool stuff customers built. Related Article What Google Cloud announced in AI this month - 2025 Learn about the latest announcements, innovations, and guides when it comes to Google Cloud AI. Read Article

24 MIN READ arrow_forward
What’s new in AI infrastructure and orchestration this month
ARCHITECT

What’s new in AI infrastructure and orchestration this month

At Google, AI is a soup-to-nuts endeavor. Obviously, we make leading AI models like Gemini and Nano Banana. We incorporate AI into the tools you use every day (think Gmail, BigQuery, AlloyDB, Google Cloud Code and Google Cloud Assist). We make software frameworks to help you build with AI, like Gemini Enterprise Agent Platform, JAX, or MaxTest. And we co-design the powerful infrastructure platform that runs underneath it all, including a broad range of standard compute, accelerators like TPUs and GPUs, optimized networks and storage, as well as orchestration software like GKE and Cluster Director. Then we package them all up into supercomputing platforms like AI Hypercomputer to power the industry-wide transformation to the AI and agentic future. This is critical in today’s agentic era, where AI is evolving from answering questions to reasoning and taking action. Companies that want to lead in this next phase of AI need computing infrastructure that’s designed and optimized for these new requirements, so they can innovate faster, deliver compelling user and customer experiences, and optimize for cost and energy efficiency — all at massive scale. To support this, we are making AI infrastructure and orchestration news at a furious pace. In this blog, we provide a monthly snapshot of the recent launches and milestones that you need to know about to keep up-to-date, deep dives on architecture and performance tuning, and discussions of specialized use cases, always with pointers to where you can learn more. Keep an eye out for updates to this blog every month. July 2026 Product, technology, and tools updates Product update: Google Cloud Managed Lustre is now GA, and available in four distinct performance tiers that deliver throughput ranging from 125 MB/s, 250 MB/s, 500 MB/s, to 1000 MB/s per TiB of capacity — with the ability to scale up to 8 PB of storage capacity. The Managed Lustre solution is powered by DDN’s EXAScaler, combining DDN’s decades of leadership in high-performance storage with Google Cloud’s expertise in cloud infrastructure. Product update: C4N network and storage optimized VMs are now GA. C4N is our first network- and block-storage-optimized VM series built to eliminate data-transfer bottlenecks. Powered by 5th Gen Intel Xeon Scalable processors and built on Google’s Titanium offloading hardware, it achieves 400 Gbps network bandwidth, 95 million packets per second (MPPS), and up to 25 GiB/s of block storage throughput when paired with Hyperdisk Extreme. New feature: GKE Dataplane V2 up to 15K Nodes with Network Policies (GA). This capability enables standard GKE clusters to scale up to 15,000 nodes while maintaining full active Network Policy enforcement, supporting the massive infrastructure needs of large enterprise and AI/ML customers. New feature: Co-operative time-slicing in llm-d. If you’re running reinforcement learning (RL) workloads, you can now interleave independent RL jobs onto shared physical hardware, increasing aggregate accelerator duty cycles from a ~40% baseline up to 70% without impacting model convergence or accuracy. New AI security tool: Looking to secure your AI supply chain on GKE, deploy AI workloads safely, and cut down on shadow AI? We open-sourced k8s-aibom, a lightweight, unprivileged Kubernetes controller that continuously monitors container clusters to automatically detect running AI runtimes (like vLLM and Triton) and generate standard CycloneDX Machine Learning Bill of Materials (ML-BOMs). Check out the k8s-aibom project and get involved. Practitioner guides and how-tos How-to guide: On July 27, Google announced Day 0 support for Moonshot AI’s Kimi K3 2.8-trillion-parameter open-weight model, the day weights were released. Whichever your preferred deployment path — via Model Garden, custom orchestration, or GKE with llm-d recipes — this guide offers detailed step-by-step instructions to help you evaluate and pilot Kimi K3 in Google Cloud. How-to guide: Google Kubernetes Engine (GKE) managed DRANET supports both GPUs and TPUs. There are several configurations to use this implementation, including standard cluster (where you have full control) and autopilot cluster (where Google does the heavy configs for you). Take a deeper dive in the hands-on lab, GKE Autopilot clusters with TPUs, GKE managed DRANET and Gemma 4. How-to guide: Learn to run Ray on TPUs, not GPUs. In Part 1 of this two-part series, we discuss TPU slices (hint: Ray thinks of them as just another accelerator on which to schedule), then walk through Ray’s various AI libraries (Part 2). How-to guide: Evaluate TPUs for sample workloads using a new microbenchmark suite that helps you accurately assess whether a device is achieving its theoretical performance specifications, and to identify specific performance gaps or architecture-specific bottlenecks. Dive in here. How-to guide: Scale your agents without killing your budget. Learn how GKE orchestration can help you safely pack more agents onto a fixed compute footprint with GKE Agent Sandbox and Pod snapshots. Whether your goal is performance or cost optimization, we teach you how to turn the right dials for optimal agent efficiency. Technical blueprint: Inside the optimization of Mistral 3 large inference on Ironwood. This blog outlines how one Google team optimized Mistral 3 large MoE model inference on Google’s Ironwood (TPU v7x), achieving a 1.5x performance gain. They did so with hybrid sharding, replacing linear VPU summations with tree reductions, optimizing GMM/MLA kernels, and adopting asynchronous scheduling. As a result, they boosted throughput by up to 48% while maintaining benchmark accuracy neutrality. Read the full blog here. Research, reports and deep-dives Report: Google was named a Leader in the inaugural GartnerⓇ Magic Quadrant™ for AI Infrastructure, positioned highest for ‘Ability to Execute’ and furthest for ‘Completeness of Vision’. Gartner called out Google’s proprietary scalable compute, integrated AI Hypercomputer architecture, and the scale of our AI compute capacity as key strengths. Download a copy here. Report: We recently surveyed more than 1,400 senior IT leaders for our State of AI Infrastructure report, and a resounding pattern emerged: The gap between AI ambition and infrastructure reality is widening. In fact, 83% of organizations say they require infrastructure upgrades to support production-grade agentic AI. Read the accompanying blog to understand how adapting your infrastructure to meet the demands that agentic applications place on your systems will help you move from pilot to production. June 2026 Product, technology and tool updates Product update: Protecting sensitive data used with AI is a critical part of advanced and secure cloud infrastructure. Confidential Computing cryptographically protects data in use in hardware-based Trusted Execution Environments (TEEs) with verifiable data integrity, and is now available on the accelerator-optimized G4 machine series, featuring NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs. Get started with Confidential G4 VMs and Confidential G4 GKE Nodes. Developer resource: The new TPU Developer Hub is the place to go for model builders, optimizers, and developers to learn to unlock the full performance of Google Cloud TPUs. Read more in this blog. New product: Scale your AI workloads with the new OpenTelemetry-Based TPU AI Telemetry Collector Agent. For the first time, you can route high-fidelity TPU hardware telemetry to Google Cloud Monitoring, Google Managed Prometheus, or your own self-hosted Grafana stack. Practitioner guides and how-tos How-to guide: Learn how to build high availability into an AI inference workload running on GKE Inference Gateway with TPUs, Cloud Storage FUSE and Dynamic Resource Allocation (DRA). This blog provides an overview, or you can get all the technical details in the hands-on codelab. How-to guide: Did you know you can connect your AI agents to unstructured data in Cloud Storage via Model Context Protocol (MCP)? In this blog, learn about why would want to do that from three customer examples, then how to do it, choosing either a fully managed service, or a self-managed local server for more customization and control. Research, reports and deep-dives Report: According to an independent benchmark report, GKE Inference Gateway outperforms the next leading managed Kubernetes service with 15.7% higher throughput, 92.8% shorter wait times, and 62.6% lower inter-token latency. This performance can be attributed to its use of prefix caching, which optimizes LLM performance by storing the KV cache (activation states) of long, repetitive prompt prefixes. Learn more in the blog. Architecture deep dive: A closer look at the cold start problem, this time for TPUs and GKE, and how the Run:ai Model Streamer can help change the dynamic. Customer and partner updates Customer win: Leveraging GKE, BigQuery, Cloud SQL, and Gemini Enterprise Agent Platform, Pager Health is eliminating operational fragmentation to deliver a simplified, personalized U.S. healthcare experience that transforms lives. Customer win: Trustpilot, the customer review platform, built a high-volume streaming pipeline using fine-tuned Gemma models with Dataflow and Gemini Enterprise Agent Platform running on cost-optimized A2 VMs using A100 GPUs, as well as optimized version of vLLM maintained by Gemini Enterprise Agent Platform. May 2026 Product, technology and tool updates Product update: GKE Agent Sandbox is now generally available. New open-source project: Agent Substrate is a new open-source project aimed at continuing to push the limits of agentic infrastructure density New feature: Google AI Edge Portal, a solution for testing and benchmarking on-device machine learning (ML) at scale, now supports benchmarking and debugging on-device LLMs. Read more here. Product deep dive: We went into depth about Cloud Storage Rapid, a new family of high-performance storage offerings for AI workloads. At launch, offerings include Rapid Bucket (formerly Rapid Storage), a high-performance zonal object storage offering, and Rapid Cache (formerly Anywhere Cache), which accelerates reads on-demand and colocates compute and data for workloads in existing buckets. Research, reports and deep dives Architecture deep dive: Google Global Infrastructure VP Bikash Koley and Engineering Fellow Arjun Singh provide a high-level overview of the challenges that AI workloads pose to network infrastructure, and discuss the deep enhancements we’ve made to our data center fabrics, WAN, and global networks to better support them. Architecture deep dive: We unveiled a new cluster-level reliability model for developing frontier AI models on TPUs, ditching instance-level reliability Customer and partner updates Customer win: Visual media provider Imgix serves more than 8 billion images and videos from AI Hypercomputer equipped with G4 VMs powered by NVIDIA RTX PRO 6000 Blackwell GPUs.

8 MIN READ arrow_forward
Dropbox Integrates MCP and Dash to Close the Gap Between Security Design and Code Review
ARCHITECT

Dropbox Integrates MCP and Dash to Close the Gap Between Security Design and Code Review

Dropbox has integrated Model Context Protocol (MCP) with its internal knowledge platform, Dash, to surface security design context during AI assisted code reviews. The system retrieves threat models and security requirements for pull requests, helping reviewers validate implementation against design intent. An InfoQ Q&A explores the architecture and key lessons learned. By Leela Kumili

1 MIN READ arrow_forward
Cloud CISO Perspectives: Why AI Threat Defense is the new boardroom baseline
ARCHITECT

Cloud CISO Perspectives: Why AI Threat Defense is the new boardroom baseline

Welcome to the second Cloud CISO Perspectives for July 2026. Today, Chris Betz, CISO, Google Cloud, and Alicja Cade, Senior Director, Office of the CISO, Google Cloud, explain what boards of directors need to know about AI security and how to prepare their organizations for security governance and business agility in the AI era. As with all Cloud CISO Perspectives, the contents of this newsletter are posted to the Google Cloud blog. If you’re reading this on the website and you’d like to receive the email version, you can subscribe here. aside_block <ListValue: [StructValue([(’title’, ‘Get vital board insights with Google Cloud’), (‘body’, <wagtail.rich_text.RichText object at 0x7f978cfcc430>), (‘btn_text’, ‘Visit the hub’), (‘href’, ‘https://cloud.google.com/solutions/security/board-of-directors?utm_source=cgc-site&utm_medium=et&utm_campaign=FY26-Q2-GLOBAL-GCP39634-email-dl-dgcsm-CISOP-NL-177159&utm_content=-&utm_term=-’), (‘image’, <GAEImage: GCAT-replacement-logo-A>)])]> Why AI Threat Defense is the new boardroom baseline By Chris Betz, CISO, and Alicja Cade, Senior Director, Office of the CISO, Google Cloud Chris Betz, CISO, Google Cloud Modern security governance has become a critical part of the foundation for business agility. Often treated as an operational cost center, security is increasingly recognized as a primary business enabler, a runway that empowers your organization to move fast, adopt cutting-edge generative AI, and capture new markets securely. In today’s environment, every major business initiative is an AI initiative, and every AI initiative requires a secure foundation. Ensuring your company is investing in the right technologies and using the right tools will be crucial in leading through the rapid AI transformation. Alicja Cade, Senior Director, Office of the CISO, Google Cloud To operate against AI speed threats, boards of directors should encourage their CISOs and business leaders to transform their strategic approach for speed, scope, and scale. We need to emphasize risk and vulnerability management with a defensive strategy that’s AI native, agentic, and open. By aligning defensive speeds with automated attack cycles, using deep internal business context, and integrating tools into unified platforms, AI-powered defense can help you confidently manage today’s threats at machine speed, and simultaneously greenlight aggressive innovation. Based on our learnings defending ourselves and our customers, Google developed AI Threat Defense (AITD) to help transition security from manual, reactive firefighting to an automated, continuous capability. While directors don’t need to manage the execution of these technologies, they have to provide the governance frameworks that encourage operational modernization. To help guide your organization’s leadership team in this transition, we recommend focusing on these five strategic, constructive areas of inquiry. For boards of directors, investing in these capabilities helps build the resilience required to drive business velocity. Key questions for CISOs, business, and tech leadership While directors don’t need to manage the execution of these technologies, they have to provide the governance frameworks that encourage operational modernization. To help guide your organization’s leadership team in this transition, we recommend focusing on these five strategic, constructive areas of inquiry. 1. Business enablement: When an enterprise transitions to automated threat defense, it is not just closing a security gap — it’s reclaiming engineering productivity and protecting operational continuity. Ask your team: How will modernization investments speed up our business to deliver value to our customers? What additional resources do we need (if any) to create this business value more quickly, and create a competitive advantage? Governance objective: Ensure that any decisions about investments align with business strategy. Speed up time to market on new features. Create competitive agility advantage for security and shareholders. Expected operational standard: Consolidate business process, speed up execution and time to market. 2. Remediation cycle: By integrating business logic and context into defensive platforms, AI can help filter out the background noise that has historically overwhelmed security operations, and also keep you on top of the complex threat landscape. Ask your team: How are we managing the organization’s risk in the era of fighting AI with AI? Governance objective: Expect a management plan with CISO input for balancing business operations, risk, and profitability with speed and reliability in an AI threat-driven world. Expected operational standard: Your organizational mean time to remediate (MTTR) exposures and other desired changes into production goes down and to the right. 3. System consolidation: Boards should look beyond standalone AI features and point products to address systemic risk and truly enable business speed. Ask your team: Are we moving toward a unified security platform, or maintaining a patchwork of point tools? Governance objective: Reduce visibility gaps and operational friction created by fragmented vendor environments. Expected operational standard: Consolidate scanning, risk prioritization, and code remediation into an integrated workflow. 4. Contextual prioritization: Your organization knows exactly how applications are interconnected, where critical data assets reside, who has access privileges, and which workflows drive actual business logic. That deep context becomes the defender’s advantage when you are using AI powered defenses, including those in AI Threat Defense. Ask your team: How are we using our deep business context to reduce security alert fatigue? Governance objective: Optimize engineering resources by ensuring teams are not consumed by false-positive alerts. Expected operational standard: Direct AI systems to prioritize vulnerabilities based on actual reachability and business context. 5. AI safety and policy: Every AI conversation is a security conversation. Securing AI infrastructure starts with directing teams toward approved architectures with proper governance. Ask your team: What frameworks do we have in place to secure our internal AI pipelines and monitor shadow AI? Governance objective: Protect intellectual property and maintain compliance as the enterprise adopts generative tools. Expected operational standard: Implement clear runtime visibility, data egress controls, and secure development standards for AI. Innovate with confidence In a highly automated digital environment, passive oversight is no longer practical. Your teams should be looking at how they are using AI to accelerate security and respond to AI-driven threats at AI speed. By steering the enterprise toward a platform-centered, context-driven security posture, boards can support long-term business resilience, protect asset value, and give the organization the confidence to innovate, scale, and lead in its next phase of growth safely. Consider technologies like AI Threat Defense as part of your defenses in this new world. For more insight, check out our Board of Directors hub here. aside_block <ListValue: [StructValue([(’title’, ‘Learn something new’), (‘body’, <wagtail.rich_text.RichText object at 0x7f978e256700>), (‘btn_text’, ‘Watch now’), (‘href’, ‘https://www.youtube.com/watch?v=CmGWIwgHR60’), (‘image’, <GAEImage: Cloud-CISO-Perspectives-logo-A>)])]> In case you missed it Here are the latest updates, products, services, and resources from our security teams so far this month: Now in preview: Find and fix software vulnerabilities with CodeMender: Our AI code security agent CodeMender can scan and fix software vulnerabilities, and is now available in preview through Agent Platform and AI Threat Defense. Read more. Cyber Snapshot Report: Enterprise resilience key to toolchain success: Check out curated frontline insights and blueprints to turn potential crises into manageable events in the newest Cyber Snapshot Report. Read more. Future-proofing data integrity: Quantum-safe digital signatures in Cloud KMS: TWe are extending the PQC digital signature algorithms suite available in Google Cloud Key Management System to include ML-DSA and SLH-DSA. Here’s why. Read more. Atlas, Wiz’s autonomous vulnerability-research agent, has been ranked #1 on CyberGym: See how Wiz built Atlas, an autonomous AI system for vulnerability research that validates every finding with a real, working exploit. Read more. Best Buy scales AI workloads and secures access with Workforce Identity Federation: As Best Buy expanded its use of Google Cloud for advanced analytics and AI, its technology teams faced two significant scaling challenges: Mitigating risk and managing administrative friction when syncing thousands of backend users from Microsoft Entra ID. Here’s how Workforce Identity Federation helped them solve both problems. Read more. The risk hiding behind exposed MCP servers: Learn how unauthenticated model context protocol (MCP) servers are opening doors to sensitive cloud data, IAM, and command execution. Read more. Agentless threat detection: Illuminating cloud blind spots: Learn how Agentless Workload Detection exposes hidden threats in virtual appliances and modern cloud networks. Read more. AlloyDB adds group authentication to secure enterprise scale and AI agents: We’re bringing identity-driven access control to your enterprise workloads through IAM group authentication for AlloyDB, now available in preview. Read more. Please visit the Google Cloud blog for more security stories published this month. aside_block <ListValue: [StructValue([(’title’, ‘Join the Google Cloud CISO Community’), (‘body’, <wagtail.rich_text.RichText object at 0x7f978e256550>), (‘btn_text’, ‘Learn more’), (‘href’, ‘https://rsvp.withgoogle.com/events/google-cloud-ciso-community-interest-form-2026?utm_source=cgc-blog&utm_medium=blog&utm_campaign=FY25-Q1-global-GCP30328-physicalevent-er-dgcsm-parent-CISO-community-2025&utm_content=cisop_&utm_term=-’), (‘image’, <GAEImage: GCAT-replacement-logo-A>)])]> Threat Intelligence news Updated cyber threat actor naming system: Google Threat Intelligence Group (GTIG) has begun rolling out a unified naming schema for tracking threat actors. This new naming taxonomy represents an effort to standardize tracking across platforms and public reporting. Read more. Demystifying AI exploits: A blueprint for AI-assisted vulnerability management: Concerned about how to safely integrate AI capabilities into vulnerability management workflows? Here’s actionable guidance from Mandiant Consulting on establishing operational guardrails for AI assisted vulnerability management, including detailed scenarios. Read more. GhostApproval: A trust boundary gap in AI coding assistants: Learn how Wiz uncovered a category-level blind spot in modern AI coding assistants, and why the human-in-the-loop safety model fails against this classic threat. Read more. The risk of exposed cloud functions and how to harden: Mandiant uses recent lessons from customer engagements to describe attack scenarios and provide actionable guidance on how to secure serverless environments. While this analysis focuses on hardening strategies for Google Cloud Run services and functions that must remain publicly accessible, these principles apply universally to any public serverless deployment. Read more. Please visit the Google Cloud blog for more threat intelligence stories published this month. Now hear this: Podcasts from Google Cloud Cloud Security Podcast: CISO tested, board approved: Noah Korba, vice-president, Digital Core, Cybersecurity, and Enterprise Architecture, General Mills, goes under the hood of Mills Collaborative Recovery, the company’s intensive, annual two-week drill that recovers 90% of their Google Cloud estate to test real-world cyber resilience. Listen here. Cloud Security Podcast: Creating trust at global scale with local AI: Shuman Ghosemajumder, CEO, Reken, traces the evolution of automated fraud, from Gmail’s early invite days to the origin of credential stuffing. Listen here. Defender’s Advantage: Shadow LLMs, agentic identities, and securely integrating AI: Join Muhammad Muneer, technical manager, Incident Response, Mandiant, as he unpacks the stark realities of enterprise AI adoption. Listen here. Behind the Binary: The challenges of reversing modern languages: Jae Young Kim from the Mandiant FLARE team discusses navigating how software has evolved, and what it actually takes to reverse engineer modern compiled languages like Go and Rust. Listen here. To have our Cloud CISO Perspectives post delivered twice a month to your inbox, sign up for our newsletter. We’ll be back in a few weeks with more security-related updates from Google Cloud.

9 MIN READ arrow_forward
Anthropic lost control of Claude in latest AI cyber blunder
ARCHITECT

Anthropic lost control of Claude in latest AI cyber blunder

Days after two OpenAI frontier AI models conducted their own real-world cyber attacks, Anthropic admits that three of its models went off the rails and hacked external organisations thanks to a “misunderstanding” with one of its technical partners

1 MIN READ arrow_forward