at Apple
Location
Cupertino, United States of America
Compensation
$185k–$325k USD
Type
full time
Posted
6 days ago
Market range · company + function + seniority
p25 · target · p75 · n=771
Posted $325k · above the band
Tailor your résumé to this role in 30 seconds.
Free account · ATS keyword check · per-job bullet rewrite by Claude.
In this role you'll contribute across several interconnected work streams spanning evaluation quality, reward/alignment signals, and data science. A core part of the job is bringing "evals first" thinking to the team — building tooling and harnesses grounded in real workflows, and designing for observability and reproducibility so quality can be measured clearly and issues caught early. This is a largely unexplored space with few established playbooks, so being self-driven is a must — you'll define your own path as much as execute one. Scope and priorities will evolve, and we're looking for someone who moves fluidly across these areas, bringing strong software engineering fundamentals with enough ML/LLM depth to build and ship AI-facing tooling
Supporting the evaluation of new Siri features and interaction modalities, working from ambiguous early requirements toward concrete, automated coverage
Turning product goals into measurable system behavior — instrumenting the product, building eval harnesses, and creating test datasets grounded in real user workflows
Building and improving reward models and alignment signals that measure whether Siri responses meet user needs
Designing and shipping evaluation tooling, pipelines, and architecture end-to-end — from data ingestion through scoring to monitoring in production — with an eye toward observability, logging, and reproducibility
Diagnosing failures across the stack, from environment provisioning through pipeline execution to scoring — enabling auto-diagnostics and driving durable fixes by partnering across engineering, infrastructure, and program teams to align on interfaces, priorities, and shared standards
Strong programming skills in one or more compiled languages (Swift, C++, or Objective-C)
Strong Python skills and solid computer science fundamentals, including data structures, algorithms, and clean, testable code
Ability to quickly learn and adapt to evolving technologies and tools, such as GenAI-assisted coding, new ML frameworks, and emerging LLM/agent tooling
Experience with backend/API development and production debugging
Excellent communication and cross-team collaboration skills, with experience working effectively within large, cross-functional organizations
M.S. or B.S. in Computer Science, Machine Learning, or a related field (or equivalent experience)
Experience evaluating ML, LLM, or agent-based systems, including familiarity with metrics, scoring methodology, trajectory and outcome analysis, and techniques like prompting, RAG, or LLM as judge
Understanding of reinforcement learning and the underlying techniques behind modern LLMs (e.g. transformer architectures, RLHF/RLAIF, fine-tuning, reward modeling ) and frameworks such as PyTorch or Hugging Face, as applied to evaluation and reward signal design
Familiarity with eval-driven development — defining success criteria and test cases from product goals and real user workflows rather than abstract benchmarks
Experience with data science methods applied to quality measurement — defining ground truth, measuring inter-rater agreement (e.g. Cohen's/Fleiss' kappa), and validating automated scorers using basic statistical techniques (e.g. correlation, confidence intervals, hypothesis testing)
Experience with MLOps, deployment, and test/eval environment management — containerization, CI/CD, model versioning, monitoring, cloud platforms (AWS, GCP, or similar), and staging or provisioning environments to produce repeatable, deterministic conditions
Comfort communicating and collaborating effectively across multicultural teams and time zones
As part of the Siri organization, you will build the systems and tooling that make evaluation a first-class part of how Siri is developed — not an after-the-fact check — spanning human evaluation, real user feedback, reward and alignment signals, and data science rigor across iOS, iPadOS, macOS, watchOS, and visionOS.
This is a rare opportunity to work at the intersection of software engineering and rigorous evaluation science — applying machine learning engineering techniques, from model evaluation to reward modeling, to build the infrastructure that keeps this quality signal trustworthy. What you build will directly shape the direction of one of the world's most widely used assistants.
Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant
At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.
Learn about accessibility in Apple’s workplace
Learn about reasonable accommodations for job applicants
Apple accepts applications to this posting on an ongoing basis.
More open roles at Apple
Hiring velocity, headcount trend, and every open posting on one page.
Open postings ranked by description similarity — useful if this role isn't quite right.