Senior ML Solutions Architect - Token Factory

Nebius
United States2025-08-05

About the job

This position sits within Nebius Token Factory, our serverless platform for running and customizing open-source LLMs in production. Token Factory allows for serverless inference and fine-tuning (LoRA, full FT, RFT) backed by in-house optimizations like custom speculative decoding, quantization, cache-aware routing and dedicated endpoints. Customers come to us to move from prototype to scaled production without the cost and complexity of building and tuning their own inference stack.

Responsibilities

Optimize LLM inference across various modalities to drive business value and support customer goals

Provide support in supervised and reinforcement learning fine-tuning to maximize model quality for the customers

Design and implement LLM-based solutions using Nebius Token Factory’s inference services

Build production-ready applications leveraging our serverless LLM APIs, including multimodal models (text, vision, audio) and domain-specific models

Provide technical expertise in prompt engineering, RAG architectures and model selection

Collaborate with product and engineering teams to surface customer feedback and shape the platform roadmap

Guide customers in scaling from POC to production with a focus on performance, reliability, and cost efficiency

Qualifications

Minimum

5+ years of experience in ML/AI systems, with at least 2 years focused on LLMs and generative AI

Deep knowledge of the LLM ecosystem, including model architectures and fine-tuning approaches

Hands-on experience with:

Running LLMs in production: deploying and operating inference workloads

LLM fine-tuning, including supervised fine-tuning (SFT/LoRA) and data preparation/curation; experience with RL-based fine-tuning is a strong plus

LLM evaluation: building task-specific benchmarks and offline/online eval pipelines, including LLM-as-a-judge setups

Inference frameworks and libraries (e.g., vLLM, SGLang, TensorRT-LLM, Transformers)

Deploying LLM-powered applications using APIs from OpenAI, Anthropic, or open-source models

Strong Python programming skills

Excellent communication skills, with the ability to clearly explain technical concepts to diverse audiences

Preferred

Experience with inference frameworks and libraries (e.g., vLLM, SGLang, TensorRT-LLM)

Work with multimodal AI models (e.g., vision-language, speech)

Proficiency with DevOps tools (Docker, Kubernetes)

Contributions to open-source ML/AI projects