About the job
This position sits within Nebius Token Factory, our serverless platform for running and customizing open-source LLMs in production. Token Factory allows for serverless inference and fine-tuning (LoRA, full FT, RFT) backed by in-house optimizations like custom speculative decoding, quantization, cache-aware routing and dedicated endpoints. Customers come to us to move from prototype to scaled production without the cost and complexity of building and tuning their own inference stack.
Responsibilities
Optimize LLM inference across various modalities to drive business value and support customer goals
Provide support in supervised and reinforcement learning fine-tuning to maximize model quality for the customers
Design and implement LLM-based solutions using Nebius Token Factory’s inference services
Build production-ready applications leveraging our serverless LLM APIs, including multimodal models (text, vision, audio) and domain-specific models
Provide technical expertise in prompt engineering, RAG architectures and model selection
Collaborate with product and engineering teams to surface customer feedback and shape the platform roadmap
Guide customers in scaling from POC to production with a focus on performance, reliability, and cost efficiency
Qualifications
Minimum
5+ years of experience in ML/AI systems, with at least 2 years focused on LLMs and generative AI
Deep knowledge of the LLM ecosystem, including model architectures and fine-tuning approaches
Hands-on experience with:
Running LLMs in production: deploying and operating inference workloads
LLM fine-tuning, including supervised fine-tuning (SFT/LoRA) and data preparation/curation; experience with RL-based fine-tuning is a strong plus
LLM evaluation: building task-specific benchmarks and offline/online eval pipelines, including LLM-as-a-judge setups
Inference frameworks and libraries (e.g., vLLM, SGLang, TensorRT-LLM, Transformers)
Deploying LLM-powered applications using APIs from OpenAI, Anthropic, or open-source models
Strong Python programming skills
Excellent communication skills, with the ability to clearly explain technical concepts to diverse audiences
Preferred
Experience with inference frameworks and libraries (e.g., vLLM, SGLang, TensorRT-LLM)
Work with multimodal AI models (e.g., vision-language, speech)
Proficiency with DevOps tools (Docker, Kubernetes)
Contributions to open-source ML/AI projects