Fiatlux: A Long-Horizon Benchmark for Humanoid Ladder Climbing and Light-Bulb Replacement

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the absence of benchmarks evaluating long-horizon tasks that combine vertical locomotion with dexterous manipulation of fragile objects. Built upon NVIDIA Isaac Lab, we introduce a simulation benchmark for humanoid robots performing lightbulb replacement, encompassing ladder climbing, fine-grained manipulation, and disposal. This project pioneers a long-horizon evaluation protocol integrating vertical mobility with fragile payload handling, providing both privileged and standard observation modes alongside teleoperation data. We establish baselines using RSL-RL PPO, the GR00T VLA model, and whole-body controllers. Furthermore, we define difficulty-weighted scoring and partial credit mechanisms to rigorously assess performance. By fully open-sourcing the codebase, evaluation protocols, and reference implementations, this benchmark offers a standardized metric for advancing humanoid robotics research.
📝 Abstract
Existing benchmarks evaluate tabletop manipulation, flat-floor household activity, or humanoid locomotion and manipulation as separate task groups; none scores vertical mobility and dexterous work on a fragile payload in one long-horizon episode. We present Fiatlux, a light-bulb replacement benchmark built on NVIDIA Isaac Lab. In one episode, a Unitree G1 humanoid positions a step ladder under a ceiling or wall fixture, climbs it, exchanges a spent bulb in a socket for a fresh one, and leaves the spent one in a disposal crate. We decompose the episode into twelve subtask environments scored on difficulty-weighted gates. The goal is a successful replacement, with the fresh bulb seated, the spent one disposed of, neither dropped, and a fragility bound not crossed. Runs that fall short can earn partial credit. Observations are split into a standard mode (signals a physical robot could sense or estimate) and a privileged mode (exact simulator state). We specify the evaluation protocol and provide reference baseline implementations spanning RSL-RL PPO, zero-shot NVIDIA GR00T N1.7 Vision-Language-Action (VLA) models, and whole-body controllers. Additionally, we provide the teleoperated recordings used to specify and check the success gates. The benchmark code and the teleoperated recordings are available at fiatlux-bench.github.io.
Problem

Research questions and friction points this paper is trying to address.

Humanoid Robot
Long-Horizon Benchmark
Ladder Climbing
Dexterous Manipulation
Fragile Payload
Innovation

Methods, ideas, or system contributions that make the work stand out.

Humanoid Robot
Long-Horizon Benchmark
Dexterous Manipulation
Vertical Mobility
Vision-Language-Action
🔎 Similar Papers