at Google
Location
Sunnyvale, CA, USA
Compensation
$364k–$505k USD
Type
full time
Posted
2 weeks ago
Market range · company + function + seniority
p25 · target · p75 · n=183
Posted $505k · well above market
Tailor your résumé to this role in 30 seconds.
Free account · ATS keyword check · per-job bullet rewrite by Claude.
As a Principal/Distinguished Engineer for AI Capacity Delivery, you will bridge the gap between low-level hardware design and intelligent software automation. You will architect the intelligent software control plane that manages, validates, and initializes Google’s next-generation global AI fleet. Your core mission is to drastically accelerate the onboarding and validation of Google’s next-generation AI fleet, including TPUs and GPUs. You will deliver and operate highly available, cost optimized data center infrastructure at speed and scale.
You will leverage modern machine learning platforms to build adaptive, self-training systems that analyze physical fleet behavior, predict anomalies, and automatically configure and validate hardware at scale. This person must develop the necessary hardware qualification tests that ensure that our fleet is reliable, properly configured and healthy enough to perform its mission. You will leverage advanced ML platforms to develop self-training systems that analyze physical hardware behavior and guide engineering teams to write highly optimized software for testing.
The AI and Infrastructure team is redefining what’s possible. We empower Google customers with breakthrough capabilities and insights by delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google users worldwide.
We're the driving force behind Google's groundbreaking innovations, empowering the development of our cutting-edge AI models, delivering unparalleled computing power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more.
Individual pay is determined by factors including job-related skills, experience, and relevant education or training.More open roles at Google
Hiring velocity, headcount trend, and every open posting on one page.
Open postings ranked by description similarity — useful if this role isn't quite right.