at Apple
Location
San Francisco Bay Area, United States of America
Compensation
$207k–$373k USD
Type
full time
Posted
3 weeks ago
Market range · company + function + seniority
p25 · target · p75 · n=800
Posted $373k · well above market
Posting health
Aging · 70Tailor your résumé to this role in 30 seconds.
Free account · ATS keyword check · per-job bullet rewrite by Claude.
In this role, you will manage Apple's relationship with internal crowd as well as a portfolio of external vendorpartners spanning evaluation, red-teaming, and data generation. You will own how this work gets resourced, scoped, and executed — from statements of work and workforce calibration to quality audits and ongoing partner performance — ensuring these partnerships can scale reliably as safety evaluation needs grow across products, features, and markets.You will also own the feedback loop and establish the processes that connect Red-Teaming, Evaluation, and Post-Ship Insights groups. Findings surfaced through red-teaming and evaluation need to inform what we monitor once features ship, and post-ship signals need to flow back into how we prioritize and design future red-teaming and evaluation coverage. You will build and operate the mechanisms — shared reporting, recurring syncs, and closed-loop tracking — that keep these three functional areas working from a single, current picture of risk rather than in isolation. Ensuring compliant and timely reporting across Apple Intelligence features is a key element of this role.You will shape the tools and infrastructure roadmap needed to power these functions. This includes maintaining clear, current documentation of processes, systems, and data flows across these functional areas, and identifying and prioritizing the tooling and infrastructure investments needed to support them as they scale — partnering with engineering and technical stakeholders to translate operational needs into build requirements.Lastly, you will also drive the operational scaling of our data infrastructure: how we manage, store, and govern the human-labeled goldsets that underpin safety benchmarking and LLM-judge validation, and how those pipelines extend to new languages, markets, and cultural contexts.
Bachelor's degree in a related field5+ years of experience in vendor/partner management, data operations, or program management, ideally supporting ML/AI evaluation, annotation, or data labeling programs
Demonstrated experience managing external vendor relationships and/or internal annotation operations partnerships at scale, including statements of work, budget, and performance management
Experience scaling human-labeled data pipelines, including dataset creation, quality calibration, storage, and governance
Experience standing up or scaling operations across multiple international markets, including working with multilingual and multicultural vendor workforces
Experience coordinating across functionally distinct but interdependent teams to close feedback loops and keep shared priorities aligned
Experience owning documentation and translating operational needs into tooling or infrastructure roadmaps, partnering with engineering to scope and prioritize build work
Strong ability to think strategically about operational tradeoffs while managing multiple concurrent initiatives in a fast-paced, evolving environment
Proven track record managing budgets and vendor relationships at scale
Proficiency with data tools (e.g., SQL, spreadsheets/BI tools) sufficient to track vendor performance, cost, and data quality metrics
Excellent written and verbal communication skills, with demonstrated ability to translate operational risk and status for both technical and non-technical stakeholders
Experience managing vendor or annotation partnerships specifically for AI safety evaluation, red-teaming, or Responsible AI programs
Familiarity with goldset design, labeling guidelines, or LLM-judge/evaluation workflows sufficient to partner effectively with technical teams
Experience building internationalization or localization operations for data labeling or evaluation programs
Familiarity with regulatory frameworks relevant to AI safety and content risk across international markets
Demonstrated success building operations or vendor programs from the ground up, including establishing new processes or partner relationships
Executive presence and experience presenting operational risk assessments and scaling plans to senior leadership
At Apple, we don't just build products — we build experiences fueled by world-class data. The Responsible AI team is looking for a senior operations program manager to own the partnerships, vendor relationships, and data operations behind our Safety Evaluation and Red Teaming programs, scaling the human-labeled data, workforce, and localizationinfrastructure that generative AI features across our Software ecosystem.
At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $207,400 and $372,700, and your base pay will depend on your skills, qualifications, experience, and location.Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant
At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.
Learn about accessibility in Apple’s workplace
Learn about reasonable accommodations for job applicants
Apple accepts applications to this posting on an ongoing basis.
More open roles at Apple
Hiring velocity, headcount trend, and every open posting on one page.
Open postings ranked by description similarity — useful if this role isn't quite right.