at Apple
Location
New York City, United States of America
Compensation
$150k–$225k USD
Type
full time
Posted
2 days ago
Market range · company + function + seniority
p25 · target · p75 · n=800
Posted $225k · well below market
Tailor your résumé to this role in 30 seconds.
Free account · ATS keyword check · per-job bullet rewrite by Claude.
As a site reliability engineer in Apple Ads focused on machine learning, you will own the health, performance, and scalability of large scale infrastructure powering ML training, inference, serving workloads and associated platform tooling. Your focus will be on building automation that eliminates manual processes, improves platform resilience, and enables teams to move faster with confidence.
This is not a DevOps-only or CI/CD-focused role. We are looking for engineers who build platform solutions, not just configure pipelines.
Build and operate distributed systems using AWS managed services such as EKS, ElasticCache and ML technologies like Ray over Kubernetes and NVIDIA Triton Inference Server.
Develop internal tooling and automation frameworks to improve infrastructure reliability, cost-efficiency, and operational visibility.
Collaborate with engineering teams to define infrastructure architecture, troubleshoot complex issues, and drive production excellence.
Design and manage Infrastructure as Code with Terraform, ensuring repeatable, secure, and scalable deployments.
Lead or participate in incident response, postmortems, and continuous improvement cycles to reduce future risk.
3+ years of experience in internet-facing backend production systems, SRE or ML Operations focused roles on large scale distributed cloud infrastructure
Proven expertise with AWS-managed infrastructure
Familiarity with ML lifecycle and associated technologies such as NVIDIA Triton, AnyScale Ray, Apache Airflow etc.
Strong programming skills in at least one of: Python, Java, Rust, Go or similar languages
Hands-on experience with Linux systems and deep knowledge of its internals.
Demonstrated experience with Infrastructure as Code, especially Terraform.
Strong foundation in SRE concepts: Monitoring, alerting, observability, Incident response and root cause analysis, Error budgets, SLAs/SLOs, and system reliability
Built tools or services that automate platform operations, reduce toil, or improve cost efficiency.
Experience managing Kubernetes clusters at scale in production environments.
Hands-on experience troubleshooting distributed systems under real-world load.
Clear communication skills and comfort collaborating across engineering, infrastructure, and product teams.
AWS certifications or broad experience across multiple AWS services is a plus.
Understanding of modern GPU hardware architectures (such as NVIDIA H100, B200, or GB200, AWS Inferentia ), associated drivers
Understanding of high-performance fabrics and network architecture, power, and thermal limits
At Apple, we focus deeply on our customers’ experience. Apple Ads brings this same approach to advertising, helping people find exactly what they’re looking for and helping advertisers grow their businesses.
Our technology powers ads and sponsorships across Apple Services, including the App Store, Apple News, and MLS Season Pass. Everything we do is designed for trust, connection, and impact: We respect user privacy, integrate advertising thoughtfully into the experience, and deliver value for advertisers of all sizes—from small app developers to big, global brands. Because when advertising is done right, it benefits everyone.
The Site Reliability Engineering team within Apple Ads ensures the reliability, performance, and availability of ML Platform and Services at scale. The team partners closely with Ads engineering, data science and ML platform teams to enable product delivery through design, configuration, and automation of machine learning infrastructure powering Apple Ads applications.
We are looking for a ML Platform Infrastructure Engineer to help build and evolve the next generation of Apple Ads machine learning platform — enabling fast, reliable, and scalable operations across AWS-based environments supporting transactional and analytical workloads.
Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant
At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.
Learn about accessibility in Apple’s workplace
Learn about reasonable accommodations for job applicants
Apple accepts applications to this posting on an ongoing basis.
More open roles at Apple
Hiring velocity, headcount trend, and every open posting on one page.
Open postings ranked by description similarity — useful if this role isn't quite right.