About the job
CADET (Customer and Analytics Driven Evals Team) is building a customer-grounded quality system for Copilot. Our mission is to rapidly identify the customer scenarios that matter most, represent them faithfully in evaluation and learning assets, run quality gates continuously, and turn every important failure into reusable product and model improvements. We bring together DSAT and other product signals, deep customer engagements to create representative eval sets. Operating in a fast-paced environment, we connect customer grounded quality issues with quality teams to advance Copilot quality and product innovation. We are looking for a Principal Applied Scientist to work directly with enterprise customers and Copilot teams, translating high-value workflows, expected outcomes, and recurring pain points into trusted evaluation and learning signals. You will set the scientific direction for customer-grounded quality: define what “good” means, assess whether eval portfolios represent real needs, diagnose model and agent failures, and convert evidence into reusable evals, RLEs, reward signals, and post-training priorities. The ideal candidate combines scientific depth with product judgment and can turn ambiguous customer problems into rigorous, scalable methods in partnership with applied researchers and ML engineers.
Responsibilities
Set the scientific strategy for customer-grounded quality across priority Copilot intents, defining what good means and the tradeoffs across various quality and safety attributes
Translate user research, enterprise customer feedback, DSAT, and production incidents into evaluation and post-training priorities, then lead cross-team creation of reusable evaluation, regression, RLE, and post-training assets for the highest-value workflows and failure patterns
Develop methods to assess evaluation-set representativeness, coverage, freshness, discrimination, grader reliability, and alignment with production outcomes, identifying material gaps, drift, and emerging loss patterns
Design behavior and task evaluations, rubrics, graders, and calibration methods that translate qualitative customer expectations into measurable release-over-release quality
Establish methods to attribute quality losses across grounding, retrieval, tools, orchestration, model reasoning, response generation, and evaluation, linking offline movement with online signals such as DSAT, task completion, retries, abandonment, and escalation
Set a high bar for scientific rigor, reproducibility, documentation, and interpretation of results while mentoring scientists and engineers and influencing evaluation and post-training strategy across organizational boundaries
Qualifications
Minimum
Bachelor's Degree in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 6+ years related experience (e.g., statistics, predictive analytics, research)
OR Master's Degree in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 4+ years related experience (e.g., statistics, predictive analytics, research)
OR Doctorate in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 3+ years related experience (e.g., statistics, predictive analytics, research)
OR equivalent experience
Preferred
Advanced degree in computer science, machine learning, statistics, applied mathematics, or a related quantitative field, or equivalent practical experience
Significant experience applying machine learning, natural language processing, information retrieval, reinforcement learning, experimentation, or evaluation methods to complex production systems
Experience designing evaluations, metrics, experiments, datasets, graders, or reward functions for AI or agentic systems
Hands-on ability to inspect model outputs, identify behavioral patterns, and translate qualitative judgments into testable hypotheses and measurable evaluation criteria
Solid understanding of statistical inference, sampling, measurement validity, bias, uncertainty, and experimental design
Solid written and verbal communication skills, with a demonstrated ability to drive results across organizational boundaries by aligning science, engineering, product, and platform teams around shared quality goals, explicit ownership boundaries, and measurable outcomes
Experience with large language models, copilots, agents, tool use, retrieval-augmented generation, or enterprise grounding
Experience running or partnering on RLHF, direct preference optimization, instruction tuning, fine-tuning, or human-preference data programs
Experience connecting offline metrics with online product behavior and customer outcomes
Experience working directly with enterprise customers or translating qualitative research and customer signals into scientific assets
Publication record, patents, or demonstrated industry impact in relevant applied research areas