LASER: Latent Space Adjoint Matching for Support-Constrained Entropy-Regularized Offline RL

NeurIPS 2026

1Massachusetts Institute of Technology2MIT Lincoln Laboratory

Expressive offline RL with support constraints and explicit entropy regularization
using latent space adjoint matching.

LASER policies on four OGBench environments.

Reach-avoid comparison: dataset trajectories go around a wall; removing support constraints leads to out-of-distribution failures; removing entropy regularization gets agents stuck; LASER reaches the goal.
Staying in support is only part of the story. Entropy regularization helps avoid brittle policies and artifacts in the learned critic. LASER combines both ingredients in an expressive latent flow policy.

Overview

Offline reinforcement learning learns from a fixed dataset. A policy must improve on the behavior in that dataset while avoiding out-of-distribution actions. Constraining the policy to a learned latent support helps, but directly maximizing the critic can still cause mode collapse.

40task–dataset settings
in the main OGBench suite
34/40settings reaching ≥95% of
the best mean success rate
13/40settings with 100%
mean success rate

Method

LASER learns where to sample in latent space, then decodes those samples into actions. Its four components are trained jointly: a behavior-cloning decoder, a reference latent flow, a critic, and an entropy-regularized latent policy.

LASER architecture: Gaussian base noise is mapped into bounded latent space; a behavior-cloning decoder maps latents to actions, while the learned latent policy favors high-value actions.
The behavior-cloning decoder models dataset actions. The latent policy changes how the bounded latent space is sampled to improve the induced action policy.

1 Learn the decoder

Flow matching maps a uniform distribution on a bounded ball to dataset actions, conditioned on the state.

2 Learn the reference flow

A second flow maps Gaussian base noise to the uniform latent distribution used as the reference for adjoint matching.

3 Learn the critic

Temporal-difference learning estimates action values. Composing the critic with the decoder gives a value function in latent space.

4 Optimize the latent policy

Adjoint matching learns a value-tilted latent distribution, balancing higher values with entropy.

The target latent distribution, for a fixed state, critic, and decoder

π_L^⋆(·|s) ∝ q_L^unif(·) exp(Q_φ^{π_L}(s, ·) / α): the optimal latent policy is proportional to the uniform latent density times the exponential of the critic divided by temperature alpha.

The temperature α controls the balance between value and entropy. The uniform reference defines the permitted latent support.

In practice, LASER projects numerical latent outputs onto the bounded ball before decoding. Action-space support depends on how faithfully the learned decoder captures behavior support. Adjoint matching avoids BPTT through the latent flow.

Why an expressive flow policy?

A Gaussian latent policy is simple, but limits expressiveness. In our illustrative comparison, the flow policy achieves a stronger tradeoff between entropy and expected value.

For every operating point achieved by the Gaussian policy in this example, the flow policy can attain both higher entropy and higher expected Q.

Experiments

We evaluate LASER on OGBench, spanning locomotion and manipulation. Four environments × five tasks × two dataset qualities give 40 evaluation settings, each run with three training seeds.

AntMaze

Locomotion & navigation

Cube-single

Single-object manipulation

Cube-double

Two-object manipulation

Scene

Multi-object manipulation

We compare with ReFORM, DSRL, FQL, IFQL, QAM, and QAM-E on both clean and noisy datasets. LASER uses one set of method-specific hyperparameters across tasks.

Strong performance across dataset qualities

Performance profiles on clean and noisy datasets: LASER has the strongest aggregate profile compared with six baselines.
At a success-rate threshold τ, each curve shows the fraction of runs with success rate ≥ τ. Shading gives pointwise 95% bootstrap confidence intervals from 2,000 resamples of training seeds within each task. Higher is better.

LASER outperforms the evaluated baselines in aggregate, reaching at least 95% of the best mean success rate in 34 of 40 settings and a mean success rate of 100% in 13 settings.

What makes LASER work?

Ablations on cube-double-play-singletask-task2-v0 examine entropy strength, optimization, and the decoder prior.

Entropy strength

Training curves comparing inverse temperatures 3, 5, 10, 20, and 30.
Moderate inverse temperatures, 1/α ∈ [5, 20], produce comparable results in this ablation. The default is 1/α = 10 for all environments.

Adjoint matching

Training curves showing more stable and stronger performance with adjoint matching than BPTT.
Adjoint matching gives stronger, more stable results than BPTT, with 3.6× higher reported training throughput (123 vs. 34 iterations/s).

Changing the decoder prior

Training curves comparing uniform and mixture decoder priors with the same latent support.
A different decoder-prior density with the same support yields comparable final performance in this task.

Lines show mean success rates; shading spans the minimum and maximum across runs.

See the paper and appendix for proofs, implementation details, additional ablations, and the harder Puzzle-4×4 evaluation.

Abstract

While offline reinforcement learning (RL) enables policy optimization from static datasets without costly online interaction, it remains bottlenecked by the risk of executing out-of-distribution (OOD) actions. Recent approaches mitigate this by learning a behavior-cloning policy through flow matching and then performing RL within its constrained latent space. However, naïvely optimizing the latent policy can easily cause the policy to collapse into a brittle mode or exploit sharp artifacts of the learned critic. In this work, we find that entropy regularization is essential in latent-space RL for addressing these challenges. We introduce LASER, a novel offline RL algorithm that applies latent-space adjoint matching to achieve entropy-regularized latent-space RL with expressive flow policies while avoiding backpropagation through time. Through comprehensive experiments on 40 challenging OGBench tasks with varying dataset qualities, we show that LASER achieves state-of-the-art performance. Notably, LASER uses fixed method-specific hyperparameters across all tasks and outperforms the evaluated baselines, including those with task- and dataset-specific tuning, which highlights the robust applicability of LASER.

BibTeX

@inproceedings{zhang2026laser,
  title={{LASER}: Latent Space Adjoint Matching for Support-Constrained Entropy-Regularized Offline {RL}},
  author={Zhang, Songyuan and So, Oswin and Yu, Eric Yang and Cleaveland, Matthew and Crowley-Dolen, Peter and Fan, Chuchu},
  booktitle={Advances in Neural Information Processing Systems},
  year={2026}
}

Figure

Scroll to inspect the full figure. Press Esc or Close to return.