TurboSens is a turbofan engine simulator built to study the inner life of learned models on a substrate where the latent ground truth is exposed. Sensor streams are paired with the full health state of the engine, so an encoder’s representation can be probed directly against the state it should have recovered. The substrate is designed for representation learning, prognostics, and reinforcement learning, with deterministic replay for counterfactuals.
What TurboSens is
The substrate exposes a 10-dimensional health state (efficiency and flow capacity deviations across engine subsystems), a 5-dimensional reversible fouling layer (dust and deposits that mask wear and which maintenance can clear), and 7 noisy sensors observed across 12 or 16 operating contexts (flight phases and conditions). Maintenance actions form a small discrete vocabulary (idle, targeted replacements per subsystem, engine wash), and the simulator can be queried online under any action sequence or replayed deterministically to make counterfactuals reproducible.
Two configurations ship together: TurboSens1 (7 × 12 observations, 6 actions, no fouling) and TurboSens2 (7 × 16 observations, 7 actions, fouling & phantom state). Both come with paired ground truth, HDF5 datasets, schema docs, generation scripts, multi-seed tooling, and an evaluation pipeline.
Reward, or the Observation Stream? · ICML 2026 RLxF
Lucas Thil, Jesse Read, Rim Kaddah, Guillaume Doquet
Accepted at the RLxF: Reinforcement Learning from World Feedback workshop at ICML 2026, a workshop on training reinforcement learning systems with real-world signals (efficiency, safety, health, performance, economic outcomes) rather than human preference alone. TurboSens is one such world signal: noisy, delayed, multi-context, and paired with measurable ground-truth consequences of action.
The paper arbitrates a long-standing question about what should shape an agent’s internal model of its environment. One view holds that maximising reward in a loop of action and observation is sufficient: a good policy implies a good internal model as a byproduct. The opposing view holds that reward pushes representations toward a minimum policy-sufficient statistic and discards the rest, and that faithful world models must instead emerge from self-supervised prediction on the observation stream.
On TurboSens, where the true latent state is exposed, we find that the two objectives shape the encoder along orthogonal axes. PPO tracks the fouling, because the policy must react to it, and discards the underlying wear. Self-supervised pretraining (JEPA) does the opposite: it recovers the wear well but is invariant to fouling. Both build internal models of the engine, but of different parts of it.
We show that the two compose cleanly sequentially rather than jointly: pretrain the encoder once with JEPA, freeze it, and train PPO on top per task. The recipe matches end-to-end PPO on the default reward, improves wear recovery, and yields a reusable encoder. We close with implications for representation learning in safety-critical domains.
Code & data
The TurboSens generation scripts, datasets, and evaluation pipeline will be released alongside the workshop. The repository is under internal review and will be made public once that process completes.
More to come
A number of follow-up papers using TurboSens are currently under review at other venues: on prognostics, on inverse problems for component-level health estimation, and on benchmarks for self-supervised learning under realistic degradation and maintenance patterns. I’ll update this page as they land.