About the job
We are hiring a Senior Staff Machine Learning Engineer to architect and lead the data processing, indexing, and search infrastructure behind Firefly Foundry's media intelligence — the systems that turn massive volumes of customer media (image, video, 3D, audio) and model-derived signals (embeddings, captions, entities, shot and scene structure, aesthetic, safety, and IP labels) into structured, low-latency, searchable intelligence, and that expose it as agentic search: retrieval designed to be driven by AI agents, not only by people.
Responsibilities
Design and build scalable data-processing pipelines that transform raw customer media and model-derived signals into structured, searchable intelligence.
Contribute to the technical vision and architecture for Firefly Foundry's media-intelligence data platform and search stack.
Architect the indexing and search infrastructure — hybrid lexical + vector (ANN) retrieval, multimodal and cross-modal search, ranking and reranking, faceting and rich metadata filtering.
Make search a first-class capability for agents — tool/function-call retrieval interfaces, multi-hop query planning, iterative retrieval, and grounded results with citations and provenance.
Own index lifecycle and freshness — incremental and streaming indexing, backfills and reprocessing, and schema and embedding-model versioning.
Engineer for enterprise from the ground up — per-tenant index isolation, data residency, and the access controls that let us honor customer IP contracts under audit.
Define and enforce retrieval quality gates — offline and online evaluation, regression detection, and drift monitoring.
Own the performance and cost envelope of the platform — query latency and throughput SLAs, ANN index tuning, GPU-accelerated enrichment at scale, and right-sizing storage, serving, and accelerator fleets.
Build the platform underneath it all — rapid pipeline and index deployment, observability, monitoring, and alerting across data and search systems.
Run these systems operationally at enterprise scale — on-call, incident response, and postmortems for availability, freshness, and latency regressions.
Lead technically across teams — set standards, drive build/buy and design decisions, mentor senior engineers, and represent Firefly Foundry's data and search architecture to leadership and partner orgs.
Qualifications
Minimum
10+ years in machine learning, data, or infrastructure engineering, including deep ownership of large-scale data processing and/or search & retrieval systems in production.
Deep expertise designing and operating search and retrieval infrastructure at scale — vector/ANN, lexical search, hybrid retrieval, ranking and reranking, and query understanding.
Strong data-engineering foundations — large-scale batch and streaming pipelines, data modeling, and storage systems.
Experience building retrieval for LLM and agentic systems — RAG, multimodal and cross-modal search, grounding and provenance, and retrieval evaluation.
Strong Python.
Track record building observability, monitoring, and alerting for data and search systems.
Experience with multi-tenant systems and data isolation in an enterprise or regulated context.
Fluency with containers and orchestration (Docker, Kubernetes), CI/CD, and a major cloud (AWS or Azure).
Comfort reasoning about retrieval quality and relevance across modalities (text, image, video, 3D, audio).
Proven technical leadership — mentoring senior engineers, driving cross-org design and build/buy decisions, and influencing roadmap and standards.
MS or PhD in Computer Science, Computer Engineering, or a related field — or equivalent practical experience building and operating large-scale data and search systems.
Preferred
A systems language (Go, Rust, or C++) a plus.
Hands-on familiarity with embedding models and the inference paths that produce them (PyTorch).