Anthropic· Software Engineering - Infrastructure· San Francisco, CA | New York City, NY | Seattle, WA
Staff Engineer, Datacenter Server Lifecycle
Classified Tasks (10)
Automate 0%Augment 70%Human-Only 30%
Augment (7)
AI assists, human decides
Lead the build-out of automation to support datacenters containing tens of thousands of servers.
leadership
Own the end-to-end operational journey of every machine in the facility from initial provisioning through decommissioning.
operational
Define processes, tooling, and operational standards for running and retiring hardware at scale.
operational
Maintain automation and operational procedures for common lifecycle events such as hardware failures, firmware upgrades, and fleet rotations.
operational
Implement and enforce secure provisioning and end-of-life handling to ensure machine trust and attestation.
technical
Build and maintain tooling to track machine health, configuration, and operational status across the full datacenter fleet.
technical
Ensure machines operate with a verified chain of integrity from hardware upward by applying attestation and integrity verification practices.
technical
Human-Only (3)
Requires human judgment
Define and own the end-to-end server lifecycle strategy, including provisioning, deployment, operation, maintenance, refresh, and decommissioning.
leadership
Partner with Infrastructure Security to design and enforce trusted compute standards across the server lifecycle.
operational
Work closely with the Networking team to ensure end-to-end connectivity across all sites.
communication
Job description
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role Anthropic is expanding beyond cloud infrastructure, and this role sits at the heart of that effort. As a Staff Engineer on the Datacenter Server Lifecycle team, you will own the end-to-end operational journey of every machine in our facility — from initial provisioning and deployment, across its working life, through maintenance and refresh, and all the way to decommissioning. This is greenfield work: you will help define the processes, tooling, and operational standards that govern how we run and retire hardware at scale. A distinguishing aspect of this role is its deep intersection with security. The machines in our datacenter handle some of the most sensitive workloads in AI — training frontier models and serving millions of users interacting with Claude. Ensuring that every machine in the fleet is trusted, attested, and operating with a verified chain of integrity from the hardware up is a core part of the job, not an afterthought. You will partner closely with our Infrastructure Security team to define and enforce trusted compute standards across the lifecycle, from secure provisioning through end-of-life handling. Key responsibilities Lead the build-out of automation to support datacenters containing tens of thousands of servers. Define and own the end-to-end server lifecycle strategy — from provisioning and deployment through operation, maintenance, refresh, and decommissioning — and maintain automation and operational procedures for common lifecycle events (e.g., hardware failures, firmware upgrades, fleet rotations). Partner closely with Infrastructure Security to design and enforce trusted compute standards across the server lifecycle. Work closely with our Networking team to ensure end-to-end connectivity across all sites. Build and maintain tooling to track machine health, configuration, and operational status across the full datacenter fleet. Minimum qualifications Hands-on experience with server hardware, including rack deployment, cabling, troubleshooting, and understanding failure modes at scale. <li class="whi