at Apple
Location
Sunnyvale, United States of America
Compensation
$185k–$278k USD
Type
full time
Posted
2 weeks ago
Market range · company + function + seniority
p25 · target · p75 · n=800
Posted $278k · in the market band
Posting health
Aging · 70Tailor your résumé to this role in 30 seconds.
Free account · ATS keyword check · per-job bullet rewrite by Claude.
We're seeking an experienced software engineer to join our Platform Reliability Engineering team and drive the design, operation, and optimization of large-scale distributed systems that power our GenAI, ML, and big data platforms. You'll leverage cutting-edge open source technologies in hybrid cloud environments to build resilient infrastructure that enables seamless inference, data processing, and machine learning workloads at scale. In this role, you'll own mission-critical platform components, respond to production incidents, and collaborate across teams to shape the future of our data and AI infrastructure.
Design, build, and maintain scalable multi-tenant systems that support diverse workloads and technologies at enterprise scale
Own the full lifecycle of infrastructure and platform projects—from architectural design and implementation through deployment, monitoring, and optimization
Operate and optimize high-throughput, mission-critical services to ensure reliability, performance, and cost-efficiency
Participate in on-call rotations to respond to production incidents; diagnose root causes, implement rapid fixes, and drive post-incident improvements
Lead cross-functional collaboration with engineering teams to define requirements, validate designs, and deliver customer-impacting features and improvements
Proactively identify operational bottlenecks and systemic issues; implement preventive measures to reduce incident frequency and improve system resilience
Establish observability practices and continuously refine operational excellence standards across the platform
Bachelor's degree in Computer Science, Computer Engineering, or equivalent professional experience
Proficiency in at least one systems programming language (Python, Go, Java, or similar)
Strong expertise in distributed systems architecture, with deep knowledge of reliability, scalability, and containerization principles
Hands-on experience with cloud platforms and data processing infrastructure (Kubernetes, Spark, Flink, Ray, Trino, or equivalent technologies)
7+ years of experience in SRE, DevOps, or infrastructure engineering, with demonstrated expertise managing distributed systems at scale.
Proficiency in diagnosing and resolving complex production incidents and performance bottlenecks in large-scale distributed environments.
Familiarity with open source codebases; ability to read, understand, and explain complex system implementations
Strong understanding of system architecture and proven ability to collaborate effectively across engineering teams
Hands-on experience with big data technologies (Spark, Flink, Iceberg) and/or ML/AI platforms (Ray, MLflow, model serving infrastructure).
Strong foundational knowledge of Linux, databases, and security principles
Proactive mindset with demonstrated commitment to optimizing reliability and uptime for mission-critical services
Excellent written and verbal communication skills with ability to articulate technical concepts and strategies to both engineering teams and non-technical leadership
Demonstrated track record of designing and operating systems at scale
AI & Data Platforms (AiDP) is IS&T's engine for AI-powered innovation. The team brings together data, application development, and machine learning — including generative AI — along with data services and customer success functions, to help IS&T build solutions more efficiently and streamline the adoption and embedding of generative AI across Apple.
The Applied Machine Learning team in AI and Data Platform organization is building the foundation for Apple's enterprise-wide machine learning and data capabilities. Our Applied Machine Learning team designs, builds, and operates mission-critical platforms and services spanning ML, GenAI, inference, and big data—enabling teams across the company to harness AI and analytics at scale. We tackle complex technical challenges in reliability, performance, and scalability across a diverse ecosystem of open source and cutting-edge technologies, serving some of Apple's most demanding workloads.
Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant
At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.
Learn about accessibility in Apple’s workplace
Learn about reasonable accommodations for job applicants
Apple accepts applications to this posting on an ongoing basis.
More open roles at Apple
Hiring velocity, headcount trend, and every open posting on one page.
Open postings ranked by description similarity — useful if this role isn't quite right.