Director of Engineering, Flex Compute

Crusoe
San Francisco, CA - US / Sunnyvale, CA - US2026-08-24OnSite

About the job

You'll operate at the seam between one of the industry's fastest-growing energy footprints and the software that controls it. This is a utility-facing, safety-critical control system, and a ground-up build. You’ll define the architecture, ship the first production system, and build the team from scratch.

Responsibilities

- Own the curtailment orchestration and decision layer: grid-signal ingestion, staged shedding, per-SKU power capping without ever silently dropping a paid workload

- Land a utility-validated pilot: fast ramp to setpoint, tight accuracy, high-fidelity telemetry, and the test harness that proves it before we commit to a utility

- Deliver dynamic power management for oversubscription — more GPUs per megawatt with no customer-visible impact

- Build fast, workload-aware GPU power estimation, validated against fleet telemetry: the engine behind shed forecasting, oversubscription admission, and power planning for new silicon

- Design for safety: authenticated signal ingress, bounded blast radius, fail-safe defaults, manual backstops

- Own graceful ride-through of power-loss and grid-stress events, integrated with on-site battery and generation backstops

- Partner deeply with Data Center Engineering (mechanical, electrical, controls) and Energy teams on interconnection commitments, curtailment program design, and BESS/generation integration

- Integrate with our cloud control plane across Kubernetes and Slurm fleets: one control plane, not two

- Make build-vs-leverage calls across vendor power-management stacks and grid integration layers

- Hire and lead the team from the ground up, keeping headcount sublinear to fleet growth through automation

Qualifications

Minimum

- 12+ years in software engineering, including 5+ leading engineering teams, ideally taking a system from 0→1 to production scale

- Deep experience with distributed control planes, orchestration, or fleet automation (e.g., Temporal, Kubernetes, Slurm)

- A track record with safety-critical or physically-actuating systems, where a misfire has a bounded, designed-for worst case

- Working knowledge of data-center power systems: utility interconnection, switchgear/UPS/BESS, rack/PDU distribution, power telemetry, and GPU power management

- Experience with energy markets or grid programs demand response, curtailable-load tariffs, ISO/RTO market signals (e.g., ERCOT, PJM, CAISO) or demonstrated speed to fluency in them

- Strong build-vs-buy judgment and comfort deciding with incomplete data

- A record of hiring senior engineers and running healthy operations for systems that must never fail silently

Preferred

- GPU cluster operations or AI cloud infrastructure experience

- Prior work in the energy sector: power generation, storage, utility software, or grid-scale controls

- Background in GPU power/performance modeling, DVFS, or turning research prototypes into production systems

- Experience with checkpoint/preemption for large training jobs or power-aware scheduling