About the job
As a Staff Research Scientist, you will independently lead a major workstream in agent learning and recursive self-improvement. You will turn systematic failures and successful trajectories into hypotheses, experiments, training signals, and deployable improvements to model weights and/or the executable harness around the model. This is a research role for someone who can move between scientific reasoning, training code, agent systems, and production constraints.
Responsibilities
Design and execute end-to-end research projects that improve long-horizon enterprise agents across planning, reasoning, memory, tool use, retrieval, computer use, multi-agent coordination, and verification.
Research model post-training methods such as continued pretraining, supervised fine-tuning (SFT), RL, DPO/GRPO, reward modeling, and distillation.
Research harness-level optimization across prompts and task framing, tool and schema design, skills, MCP-backed providers, subagents, context and memory management, agent-loop policy, and reliable verifiers.
Build improvement flywheels that mine trajectories and production-safe signals, identify recurring failure modes, generate or curate data, propose interventions, and measure generalization before promotion.
Create realistic, stateful training environments and benchmarks for enterprise workflows, with programmatic verifiers and calibrated human or model-based graders where deterministic grading is not possible.
Run rigorous ablations and scaling experiments; reason explicitly about variance, contamination, reward hacking, distribution shift, cross-model transfer, cost, and latency.
Qualifications
Minimum
10+ years of relevant AI/ML research or engineering experience, or equivalent research depth and impact; PhD or other advanced degree required.
Track record of setting technical direction and leading multiple ambiguous, high-impact research efforts across team boundaries.
Strong foundations in machine learning, deep learning, reinforcement learning, and experimentation, with hands-on experience training or adapting large language or multimodal models.
Advanced Python and PyTorch skills, including modifying training code, data pipelines, evaluators, or research infrastructure.
Practical depth in agentic AI, including tool use, planning, memory, retrieval, environments, or long-horizon execution.
Experience designing decision-useful evaluations using robust datasets, trajectory analysis, graders or verifiers, and error analysis.
Experience with distributed training, rollout, or inference and modern post-training or serving stacks.
Strong software engineering fundamentals and evidence of research impact through publications, shipped systems, patents, benchmarks, or open source.
Preferred
Experience in multimodal, document AI, computer vision, speech/audio, or multilingual modeling.
Experience with enterprise agents, stateful workflows, computer use, tool protocols, or simulation environments.
Experience with synthetic data, model-generated feedback, automated experimentation, or search-based optimization.
Experience operating distributed GPU and experiment infrastructure.