AI Field Engineer - Microsoft Foundry

Fireworks AI
San Mateo, CA / New York, New York, New York, United States / San Mateo, San Mateo, California, United States2026-06-11

About the job

As an AI Field Engineer for Microsoft Foundry, you will be one of the technical owners of Fireworks' most strategic partnership. You’ll work closely with Microsoft's field teams, Azure-aligned ISVs, and the SIs that run enterprise AI transformation programs to make Fireworks the default inference and fine-tuning layer in every Azure AI architecture your partners touch. The role sits at the intersection of engineering, partner development, and customer delivery.

Responsibilities

Be the technical lead on co-sell motions with Microsoft — joint reference architectures, Azure Foundry integration patterns, and shared POCs for strategic accounts.

Build end-to-end POCs and MVPs alongside partner engineering teams, working inside their codebases, infrastructure, and constraints.

Run load tests and establish latency, throughput, and cost baselines against realistic customer traffic profiles, and tune deployments to hit those targets.

Deploy and validate new model families on inference frameworks (vLLM, SGLang), determining optimal shapes, quantization configs, and serving patterns across workloads.

Guide Microsoft’s customers on model selection, fine-tuning strategy (SFT, DPO, RFT), and evaluation methodology.

Own the feedback loop — surface partner-driven product gaps to Fireworks engineering, and translate the roadmap back into partner messaging.

Qualifications

Minimum

3+ years in a pre-sales, partner engineering, forward-deployed, or technical consulting role.

Demonstrated ability to build production software with customers, not just advise on it. You have shipped code running in someone else's production environment.

Strong Python skills. Comfortable reading, writing, and debugging production code. Familiarity with Kubernetes and infrastructure engineering.

Hands-on fluency with LLM inference: latency/throughput tradeoffs, batching strategies, quantization, structured outputs, function calling. You can explain why 50ms p99 matters to an enterprise CTO.

Real experience with fine-tuning — LoRA at minimum, RFT a strong plus. You understand when SFT is enough and when it isn't.

Deep familiarity with the Azure AI stack: Azure Foundry, Azure OpenAI Service, Azure ML, AKS, Entra/RBAC for AI workloads. You know where Fireworks fits and where it doesn't.

Exceptional communication: able to run a sharp discovery call, present to a VP, and debug a latency issue with an ML engineer in the same afternoon.

Preferred

5+ years in technical field or engineering roles where you've owned a technical relationship with a hyperscaler or major SI, not just supported one

Experience with inference serving frameworks (vLLM, SGLang, TensorRT-LLM) and tuning deployments for real workloads.

Prior role at a hyperscaler, AI-native cloud, or inference provider.

Experience with agentic frameworks (LangChain, LlamaIndex, or custom tool-use pipelines) — you understand how inference latency and reliability shapes agent behavior at scale.

Background in model evaluation — you understand why benchmark gaming is rampant and what rigorous evals actually look like.

You've written a technical blog post or reference architecture that people actually read.

Track record taking GenAI POCs from prototype to production-scale deployments.