Category index

Architect

325 articles

325 ARTICLES

FEATURED REPORT

Building multi-Region resiliency for AWS CloudFormation custom resource deployment

AWS CloudFormation is the foundational tool of infrastructure-as-code for thousands of organizations running workloads on AWS. But as teams push the boundaries of what CloudFormation can do natively, custom resources have emerged as a powerful extension mechanism that unlocks a broad range of possibilities. Yet, when it comes to building resilient, multi-Region deployments with custom

BY Raman Pujani
MIN READ 1 MIN READ
EXPLORE north_east
Building multi-Region resiliency for AWS CloudFormation custom resource deployment
GitHub Increased Instant Navigation from 4% to 22% by Rethinking Client Side Architecture
ARCHITECT

GitHub Increased Instant Navigation from 4% to 22% by Rethinking Client Side Architecture

GitHub redesigned GitHub Issues navigation using a client-side architecture that combines caching, predictive prefetching, and service workers to reduce perceived latency. The approach uses IndexedDB, in-memory caching, and background synchronization to serve data faster. GitHub reported instant navigation improvements from 4% to 22%, with latency reductions across multiple navigation By Leela Kumili

1 MIN READ arrow_forward
From maintenance to innovation: Checkout's migration to Managed Service for Apache Airflow
ARCHITECT

From maintenance to innovation: Checkout's migration to Managed Service for Apache Airflow

Data engineering teams often face a “Day 2” operational reality after building a data platform: the ongoing work of maintaining the orchestrator itself. For the Data Platform team at Checkout.com, managing a self-hosted Apache Airflow environment on another hyperscaler was consuming time the team wanted to spend elsewhere as server management, patching, and incident response were pulling focus from building pipelines. By migrating to Managed Service for Apache Airflow (Gen 3), Google Cloud’s fully managed Airflow service, Checkout.com transformed its reliability and cost structure. Here’s how they built a more scalable, cost-efficient, and robust data foundation. The starting point: self-managed Airflow Before the migration, Checkout.com ran Airflow on self-managed infrastructure. While functional, maintaining the underlying resources required significant attention. Patching, upgrades, and server management created regular interruptions. Operational data from the past year illustrates some of the challenges the company was navigating: Reducing operational friction: In its self-managed environment, Checkout.com faced stability challenges, particularly during high-load periods. Complex dependency management: Upgrading packages and ensuring compatibility was a constant, manual struggle. With Managed Airflow (Gen 3), the company was able to simplify this by handling dependencies at the image level, ensuring seamless compatibility out-of-the-box during routine environment upgrades. DAG sync time: Syncing DAGs to the scheduler took approximately six minutes after deployment to S3, which affected iteration speed. Manual processes: Scaling required manual intervention, and onboarding new teams meant manually creating secrets and variables for dbt. The solution: Managed Service for Apache Airflow (Gen 3) Checkout.com’s team migrated to Managed Airflow to offload infrastructure responsibility and take advantage of Google Cloud’s managed scalability. The results were immediate and measurable across three areas: reliability, cost, and developer velocity. Dynamic scaling in action In the previous elastic container service setup, the team allocated the maximum number of workers required for peak loads. This meant paying for peak capacity around the clock, regardless of actual usage. Managed Airflow provides built-in dynamic scaling, eliminating the need for manual resource management. The environment automatically adjusts the number of workers based specifically on the workload demands. When tasks spike, the system scales up; when they drop, it scales down to save resources. Similarly, moving from fixed provisioning to dynamic scaling reduced monthly costs by an estimated 30%. Reliability and DAG isolation Achieving increased stability was a primary driver for Checkout.com’s migration since in the past, a single problematic DAG could affect its entire environment. Managed Airflow introduced a number of critical architecture improvements: DAG isolation: Each DAG runs in its own execution environment. If one DAG fails or consumes excessive resources, it doesn’t affect the entire environment. Managed operations: Google Cloud handles patching and upgrades during scheduled windows, removing the need for manual upgrade management. Improved visibility: Integration with Google Cloud Monitoring and Cloud Logging provides clear visibility into task execution. Engineers can now debug issues independently without escalating to the platform team. Faster developer workflows The migration also improved day-to-day workflows for Checkout.com’s data engineers. Faster deployments: Using Cloud Storage for DAGs enabled near-instant syncing. Simpler onboarding: Teams no longer needed platform support to create variables before onboarding. Modernizing dbt execution: One of the company’s most significant wins was changing how it runs dbt. Previously, its engineers had to manually install and manage complex virtual environments for every supported dbt version. By leveraging containerized dbt runs, Managed Airflow (Gen 3) eliminates dependency bottlenecks. This ensures complete dependency isolation, allowing teams to run any required dbt model with minimal setup and no manual infrastructure overhead. Environment updates: The company no longer needs to redeploy the entire Airflow environment to add new roles or update Python packages. AI-powered troubleshooting with Gemini Cloud Assist: In a self-managed environment, a failed task often triggered a frantic hunt through fragmented logs and metrics. With Managed Airflow, Checkout.com can initiate a Gemini investigation directly from its Airflow DAG UI in the Google Cloud console. Gemini doesn’t just provide generic error messages; it generates a scorecard that evaluates different hypotheses with both supporting and contradictory evidence, which can drastically reduce mean time to recovery. Conclusion For Checkout.com, the move to Managed Airflow (Gen 3) marked a strategic shift, one that freed its engineers to focus on delivering value. “With Managed Service for Apache Airflow, we’ve achieved significant improvements in efficiency, scalability, and reliability. Managed infrastructure, automated scaling, faster deployments, and isolated execution environments have transformed how we operate.” — Keisi Mancellari, Data Platform Engineer, Checkout.com With a stable, scalable, and cost-efficient platform in place, Checkout.com is now able to focus on the future of its data pipelines, confident that its orchestration layer is ready for whatever comes next. Learn more about Managed Service for Apache Airflow and how it can support your data platform. Special thanks to the following contributors to this post: Serge Bouschet and Keisi Mancellari

4 MIN READ arrow_forward
Automate custom PII detection at scale with Amazon Macie and Step Functions
ARCHITECT

Automate custom PII detection at scale with Amazon Macie and Step Functions

Organizations in regulated industries like financial services, insurance, healthcare, and government ingest large volumes of data containing personally identifiable information (PII). Your applications, claims processing systems, partner data feeds, and internal workflows produce files that may include names, addresses, Social Security numbers, and domain-specific identifiers such as policy numbers, member IDs, and medical record numbers.

1 MIN READ arrow_forward
Accelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis Mission
ARCHITECT

Accelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis Mission

Scientists today face challenges of extraordinary scale and complexity. From shaping and simulating the intricate dynamics of fusion plasma, to exploring the vast search space of new materials, to making sense of the exabytes of data pouring out of the world’s most advanced experimental facilities. The demands on modern research are unprecedented. Frontier AI can help address these challenges, while accelerating groundbreaking scientific discoveries. In December, we shared our commitment to the White House’s Genesis Mission — the national effort to harness AI and double the pace of American scientific discovery within a decade. Since then, Google DeepMind (GDM) announced an early access program that provides AI for science tools to all 17 Department of Energy (DOE) National Laboratories, and Google Public Sector shared how Gemini for Government could serve as an AI backbone for the DOE. Today, at the DOE Genesis Mission Summit 2026, we are expanding this by committing $40 million of AI tokens and cloud credits for researchers in support of the Genesis Mission. Frontier AI tools for scientific discovery Under this expanded commitment, we will first provide DOE’s Genesis Mission awardees in-kind access to GDM’s frontier AI for science portfolio, including: AlphaEvolve — a Gemini-powered coding and discovery agent, for designing advanced algorithms. AlphaFold 3 — a model for predicting the structure and interactions of proteins and other biomolecules. AlphaGenome — a tool for understanding how variation in DNA, including the non-coding genome, shapes biology and disease. WeatherNext — a state-of-the-art family of AI weather forecasting models for mapping weather conditions. AlphaEarth Foundations — a foundational AI model for mapping and understanding our planet in unprecedented detail. Second, we will provide Gemini for Government seats and tokens for one year to tens of thousands of users across the DOE National Laboratories’ operations, research, and management teams. This secure platform supports the full breadth of work from the research bench to the administration of specialized user facilities serving the entire scientific community, providing a single secure foundation that the DOE mission can depend on. AI for science tools in action across the laboratory ecosystem While we have a lot of work still to do, the practical impact of the Genesis Mission is already coming to life across the laboratory ecosystem. At Pacific Northwest National Laboratory (PNNL), senior scientist Dr. Henry Kvinge is using AlphaEvolve to map out massive mathematical systems that are far too complex for humans to explore by hand. The AI uncovers hidden connections automatically, fast-tracking discoveries that would normally take researchers years to find. “Modern math relies on abstraction, but combinatorics offers concrete models that make complex geometry and algebra easier to grasp. We’ve found that systems like AlphaEvolve are perfect for this search,” said Dr. Kvinge. “By leveraging the broad mathematical knowledge of LLMs, we can automate the exploration of countless angles. We’re still experimenting, but the discoveries are already shaping our future research.” At the National Laboratory of the Rockies (NLR), researchers are utilizing Gemini to fundamentally change how they interact with physical laboratory hardware. Dr. Steven R. Spurgeon, a senior materials data scientist at NLR, leads a pioneering program in autonomous materials discovery. “Our collaboration has allowed us to build an autonomous experimentation capability,” said Dr. Spurgeon. “By deploying Gemini in our instruments, we cut microscope calibration time from over 90 minutes to about 13 minutes (eight times faster) and reduced the manual steps needed to focus an image from as many as 50 down to two. That’s time and attention we’ve given back to the science itself, enabling genuinely autonomous workflows that observe, reason, and decide in real time. This has helped us explore parts of the material design space we simply could not have reached through manual operation alone.” Driving American innovation The Genesis Mission represents an opportunity to transform research and science across America. By providing access to advanced AI tools, we aim to help scientists accelerate breakthroughs across critical energy, security, and scientific challenges. To learn more about how these AI capabilities can support your research initiatives, join us at the upcoming Google Public Sector Summit in October.

4 MIN READ arrow_forward
Anthropic Details How It Contains Claude Across Web, Code, and Cowork
ARCHITECT

Anthropic Details How It Contains Claude Across Web, Code, and Cowork

Anthropic detailed the containment architectures it uses for Claude across its products. It argues that agent safety depends on placing deterministic limits on an agent’s filesystem, network, and execution environment rather than on permission prompts or safeguards. Most notably, it examines failures at trust boundaries and along permitted egress paths that led Anthropic to revise those designs. By Eran Stiller

1 MIN READ arrow_forward
I made a policy engine think it was in production
ARCHITECT

I made a policy engine think it was in production

Kyverno is a Kubernetes-native policy engine that validates, mutates, and generates resources before workloads reach your cluster, enforcing security and compliance rules as code, without requiring a separate policy language. Enterprise teams running Kyverno policies in…

1 MIN READ arrow_forward
AWS Billing Bug Shows Customers Trillion-Dollar Estimates While Its Own Cost Alarms Fail to Act
ARCHITECT

AWS Billing Bug Shows Customers Trillion-Dollar Estimates While Its Own Cost Alarms Fail to Act

A configuration change in AWS’s bill computation system showed customers estimated bills in the billions and trillions of dollars for over 24 hours. AWS’s own alarms detected the anomalies but failed to halt bill generation or page engineers; customer escalations alerted the company 4.5 hours later. Budget and cost anomaly alerts were disabled platform-wide during mitigation. By Steef-Jan Wiggers

1 MIN READ arrow_forward