Research Engineer, Core ML

Together AI
San Francisco / San Francisco, San Francisco, California, United States2026-02-18

About the job

This is a research engineering role with direct production impact. You won’t be publishing ideas in isolation—you will translate new RL algorithms, scheduling methods, and inference optimizations into production-grade systems that power Together’s API. Success in this role means shipping measurable improvements in latency, throughput, cost, and model quality at scale. We are looking for researchers who enjoy owning systems end-to-end and turning frontier ideas into robust infrastructure.

Responsibilities

Advance inference efficiency end-to-end

Design and prototype algorithms, architectures, and scheduling strategies for low-latency, high-throughput inference.

Implement and maintain changes in high-performance inference engines (e.g., SGLang- or vLLM-style systems and Together’s inference stack), including kernel backends, speculative decoding (e.g., ATLAS), quantization, etc.

Profile and optimize performance across GPU, networking, and memory layers to improve latency, throughput, and cost.

Unify inference with RL / post-training

Design and operate RL and post-training pipelines (e.g., RLHF, RLAIF, GRPO, DPO-style methods, reward modeling) where 90+% of the cost is inference, jointly optimizing algorithms and systems.

Make RL and post-training workloads more efficient with inference-aware training loops—for example, async RL rollouts, speculative decoding, and other techniques that make large-scale rollout collection and evaluation cheaper.

Use these pipelines to train, evaluate, and iterate on frontier models on top of our inference stack.

Co-design algorithms and infrastructure so that objectives, rollout collection, and evaluation are tightly coupled to efficient inference, and quickly identify bottlenecks across the training engine, inference engine, data pipeline, and user-facing layers.

Run ablations and scale-up experiments to understand trade-offs between model quality, latency, throughput, and cost, and feed these insights back into model, RL, and system design.

Own critical systems at production scale

Profile, debug, and optimize inference and post-training services under real production workloads, taking research ideas all the way to stable, measurable improvements in deployed systems.

Drive roadmap items that require real engine modification—changing kernels, memory layouts, scheduling logic, and APIs as needed.

Establish metrics, benchmarks, and experimentation frameworks to validate improvements rigorously.

Provide technical leadership (Staff level)

Set technical direction for cross-team efforts at the intersection of inference, RL, and post-training.

Mentor other engineers and researchers on full-stack ML systems work and performance engineering.

Qualifications

Minimum

3+ years of experience working on ML systems, large-scale model training, inference, or adjacent areas (or equivalent experience via research / open source).

Advanced degree in Computer Science, EE, or a related field, or equivalent practical experience.

Demonstrated experience owning complex technical projects end-to-end.

Preferred

No preferred qualifications listed.