Senior Staff Deployment Automation Engineer

Crusoe
San Francisco, CA, USA / Sunnyvale, CA, USA / Bellevue, WA, USA2026-08-13OnSite

About the job

As a Senior Staff/Principal Deployment Automation Engineer for the Compute Team, you will be responsible for deployment and testing automation of large-scale, multi-node GPU clusters. You will own the CI/CD infrastructure, including both deployment and integration testing, for a rapidly scaling fleet of virtualized GPU and CPU hosts across our AI Cloud. Your role is critical in ensuring the stability of the low-level infrastructure and enabling teams across our Cloud Infrastructure organization to quickly and reliably release, test, and deploy their artifacts across our datacenters.

Responsibilities

- Deployment and Integration Testing Ownership: Completely own deployment and integration testing automation for all bare-metal, on-premise systems across Crusoe’s AI Cloud Stack.

- CI/CD Automation and Tooling: Build CI/CD platforms that enable developers to quickly test, iterate, and deploy critical, low-level systems and applications.

- Multi-Node Scaling Validation: Design and execute large-scale validation tests across multi-node virtualized clusters to ensure linear scaling and stability of GPU workloads.

- Configuration Management and Observability: Maintain and scale bare-metal Linux configurations using a mix of custom and off the shelf tooling such as Gitlab, Ansible, AWX, osquery, etc.

- Deployment Orchestration: Create control applications to coordinate canary deployments on live production systems, run Blue/Green testing, and perform automatic rollback where necessary.

- Cluster Orchestration: Develop and maintain automation frameworks in Python or Go to dynamically provision, configure, and stress-test multi-node virtualized environments.

- Create automated test suites leveraging tools like fio, stress-ng, and iperf to ensure performance and multi-tenant isolation of CPU and GPU hosts.

Qualifications

Minimum

- 12+ YOE demonstrated ability to competently and independently perform responsibilities plus Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related technical field.

- Experience building and deploying automated integration testing for an AI Cloud Environment, ranging from low-level Linux Systems up to Distributed Control Planes.

- Working knowledge of the modern infrastructure stack, including Kubernetes, Docker, Terraform, and Postgres.

- Intimate knowledge of CI/CD pipelines and Gitlab Tooling to enable stable infrastructure releases across multiple datacenters.

- Previous experience with at least 1-2 configuration management systems, including Ansible, Puppet, Chef, or SaltStack.

- Advanced proficiency in Python and/or Bash for automating complex cluster-wide test scenarios.

- Knowledge of Linux kernel internals, specifically PCIe topology, VFIO, and memory management (HugePages, IOMMU).

- Familiarity with NVIDIA (CUDA/NCCL) and/or AMD (ROCm/RCCL) stacks in a multi-node context.

- Strong understanding of RDMA, RoCE, and InfiniBand protocols and their implementation in virtualized systems.

Preferred

- Experience with MNNVL (Multi-Node NVLink) or specialized AI fabric architectures.

- Familiarity with hardware-level debugging tools and performance profilers (e.g., NVIDIA Nsight, AMD Omniperf).

- Knowledge of containerized orchestration for GPUs (e.g., Kubernetes with specialized device plugins).