Planning-Aligned Pretraining of BEV Representations with Sparse Action-Conditioned Targets for End-to-End Autonomous Driving
PAVER aligns BEV pretraining with candidate ego motions: a single LiDAR sweep provides sparse targets for occupied and unobserved regions along each path. A 10K-parameter head learns to predict these targets without driving-task labels or dense scene reconstruction. Transferring only the BEV encoder improves end-to-end driving under shorter training schedules, while preserving the downstream architecture and camera-only inference.
Teaser
Predictions and Grad-CAM
Six surround cameras with predicted boxes and futures, over the Grad-CAM of the BEV features behind them. PAVER captures the objects and the surrounding scene more completely, and plans better for it.
Abstract
End-to-end autonomous driving relies on a shared bird’s-eye-view (BEV) representation for perception, prediction, and planning, yet the BEV encoder is typically optimized only through heterogeneous downstream tasks. This leaves no dedicated stage for organizing the representation around geometry relevant to candidate actions. We introduce PAVER, Planning-Aligned BEV Encoder Pretraining, which rasterizes a LiDAR sweep into sparse free, occupied, and unknown cells. Rule-based candidate ego motions are rolled out and queried laterally across the vehicle width; the fractions of queries falling on measured returns and on unobserved cells form sparse action targets.
To limit direct copying from the supervised regions, PAVER replaces action-corridor features with a shared learnable mask token and predicts these targets from masked camera-derived BEV features conditioned on the corresponding action state. The framework requires no driving-task annotations or learned teacher and introduces only 10K trainable auxiliary parameters. After pretraining, the target builder and prediction components are discarded and only the BEV encoder is transferred, leaving the downstream architecture and inference pipeline unchanged.
On nuScenes, PAVER lowers VAD-Tiny’s average collision rate from 0.51% to 0.19% and its planning L2 from 0.66 to 0.60, while improving motion ADE from 0.91 to 0.80 and map mAP from 0.42 to 0.44. Planning L2 improves on VAD-Base and GenAD as well, although the collision gain does not carry over to VAD-Base.
Results overview
Planning performance and pretraining head size
Arrows run from each baseline to the same architecture with PAVER. Hollow markers are reconstruction-based pretraining, as its authors report it.
Multi-task performance
Planning, motion prediction, detection, and mapping results on nuScenes for VAD-Tiny, VAD-Base, and GenAD, with and without PAVER pretraining.
| Epochs | Auxiliary tasks | Planning | Motion | Detection | Map | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | PT (epochs) | FT (epochs) | Pretraining | Fine-tuning | L2 (m) ↓ | Col. (%) ↓ | ADE (m) ↓ | FDE (m) ↓ | MR ↓ | mAP ↑ | NDS ↑ | mAP ↑ |
| VAD-Tiny† | 0 | 60 | None | Det., Map, Mot., Plan. | 0.66 | 0.51 | 0.91 | 1.25 | 0.13 | 0.23 | 0.34 | 0.42 |
| + Ours | 20 | 30 | Actions | Det., Map, Mot., Plan. | 0.60 | 0.19 | 0.80 | 1.09 | 0.12 | 0.28 | 0.40 | 0.44 |
| VAD-Base | 0 | 60 | None | Det., Map, Mot., Plan. | 0.74 | 0.31 | 0.76 | 1.03 | 0.11 | 0.29 | 0.42 | 0.50 |
| + Ours | 20 | 30 | Actions | Det., Map, Mot., Plan. | 0.56 | 0.40 | 0.69 | 0.90 | 0.09 | 0.33 | 0.45 | 0.50 |
| GenAD† | 0 | 60 | None | Det., Map, Mot., Plan. | 0.59 | 0.37 | 0.87 | 1.18 | 0.14 | 0.19 | 0.26 | 0.46 |
| + Ours | 20 | 30 | Actions | Det., Map, Mot., Plan. | 0.54 | 0.21 | 0.80 | 1.05 | 0.11 | 0.21 | 0.28 | 0.44 |
Multi-task nuScenes validation results for three architectures, each without and with PAVER pretraining.
A bold value is the better of the pair it is compared against. † Reproduced results . Actions denotes PAVER’s sparse action targets. PT and FT denote pretraining and fine-tuning epochs. L2 is the trajectory displacement error and Col. the box collision rate in percent, both averaged over 1, 2, and 3 seconds.
Interactive 3D visualization
Both models decode the same clip: 3D boxes with object futures, the vector map, and the planned trajectory at one instant.
Model comparison
Pretraining objectives
Supervision sources and auxiliary-head requirements of driving pretraining methods.
Comparison of pretraining objectives
Training efficiency
Estimated total training time, including pretraining, compared with scratch training.
Total training time
Measured end to end on four NVIDIA RTX 5090 GPUs. The PAVER rows carry their own 20 pretraining epochs and still finish sooner than the scratch baselines.
Method
Perception supervision needs dense annotation; reconstruction needs a dense 3D objective. PAVER needs neither.
nuScenes results
Planning, mapping, and pretraining comparisons on the nuScenes validation split.
Planning by horizon
Planning L2 error and collision rate at 1, 2, and 3 seconds for VAD-Tiny, VAD-Base, and GenAD, with and without PAVER pretraining.
| L2 (m) ↓ | Col. (%) ↓ | |||||
|---|---|---|---|---|---|---|
| Method | 1s | 2s | 3s | 1s | 2s | 3s |
| VAD-Tiny† | 0.35 | 0.63 | 1.01 | 0.38 | 0.45 | 0.71 |
| + Ours | 0.32 | 0.58 | 0.92 | 0.09 | 0.15 | 0.32 |
| VAD-Base | 0.41 | 0.71 | 1.10 | 0.20 | 0.29 | 0.43 |
| + Ours | 0.30 | 0.54 | 0.85 | 0.24 | 0.41 | 0.55 |
| GenAD† | 0.33 | 0.56 | 0.88 | 0.21 | 0.34 | 0.56 |
| + Ours | 0.28 | 0.51 | 0.82 | 0.09 | 0.18 | 0.35 |
Planning error and collision rate at each prediction horizon.
Comparison with reconstruction-based pretraining
Planning performance and temporary auxiliary-parameter counts for PAVER and reconstruction-based pretraining with UniPAD and MIM4D.
| L2 (m) ↓ | Col. (%) ↓ | |||||||
|---|---|---|---|---|---|---|---|---|
| Method | Target | Auxiliary parameters ↓ | 1s | 2s | 3s | 1s | 2s | 3s |
| VAD-Tiny† | — | 0.00M | 0.46 | 0.76 | 1.12 | 0.21 | 0.35 | 0.58 |
| + UniPAD‡ | RGB, Depth | 6.41M | 0.45 | 0.75 | 1.11 | 0.19 | 0.31 | 0.46 |
| + MIM4D‡ | RGB, Depth | 13.43M | 0.39 | 0.69 | 1.06 | 0.21 | 0.28 | 0.38 |
| + Ours | Sparse Actions | 0.01M | 0.32 | 0.58 | 0.92 | 0.09 | 0.15 | 0.32 |
| VAD-Base | — | 0.00M | 0.41 | 0.71 | 1.10 | 0.20 | 0.29 | 0.43 |
| + MIM4D‡ | RGB, Depth | 13.43M | 0.36 | 0.64 | 1.00 | 0.08 | 0.14 | 0.36 |
| + Ours | Sparse Actions | 0.01M | 0.30 | 0.54 | 0.85 | 0.24 | 0.41 | 0.55 |
Comparison against reconstruction-based driving pretraining at matched downstream budgets.
Auxiliary parameters are temporary and discarded after pretraining. RGB denotes camera-color reconstruction and Depth denotes LiDAR-projected sparse metric-depth reconstruction. † Reproduced results. ‡ Reported results.
Map prediction results
Average precision for dividers, pedestrian crossings, and road boundaries on nuScenes, with and without PAVER pretraining.
| AP ↑ | ||||
|---|---|---|---|---|
| Method | Divider | Ped. Crossing | Boundary | mAP ↑ |
| VAD-Tiny† | 0.48 | 0.32 | 0.46 | 0.42 |
| + Ours | 0.48 | 0.34 | 0.49 | 0.44 |
| VAD-Base | 0.53 | 0.44 | 0.53 | 0.50 |
| + Ours | 0.53 | 0.44 | 0.52 | 0.50 |
| GenAD† | 0.50 | 0.38 | 0.49 | 0.46 |
| + Ours | 0.47 | 0.35 | 0.48 | 0.44 |
Vector map average precision per class on nuScenes.
AP and mAP denote average precision per map class and its mean. † Reproduced results.
Pseudo-LiDAR targets
VAD-Tiny downstream performance with no BEV pretraining, pseudo-LiDAR targets from monocular depth, or targets from measured LiDAR.
| Planning | Motion | Det. | Map | ||||
|---|---|---|---|---|---|---|---|
| Target source | L2 (m) ↓ | Col. (%) ↓ | ADE (m) ↓ | FDE (m) ↓ | MR ↓ | NDS ↑ | mAP ↑ |
| None† | 0.662 | 0.513 | 0.905 | 1.250 | 0.135 | 0.338 | 0.419 |
| Pseudo-LiDAR | 0.563 | 0.560 | 0.831 | 1.101 | 0.121 | 0.386 | 0.439 |
| LiDAR | 0.603 | 0.187 | 0.803 | 1.086 | 0.121 | 0.397 | 0.439 |
Replacing the LiDAR target source with monocular pseudo-LiDAR.
Closed-loop evaluation
Bench2Drive Town05 Long: nine routes, one repetition, identical protocol.
Closed-loop metrics
Driving Score, Route Completion, and Infraction Score on Bench2Drive Town05 Long for UniAD-Tiny and GenAD, with and without PAVER pretraining.
| Method | DS ↑ | RC ↑ | IS ↑ |
|---|---|---|---|
| UniAD-Tiny | 48.45 | 60.96 | 0.85 |
| + Ours | 58.79 | 79.06 | 0.79 |
| GenAD‡ | 34.53 | none | none |
| + Ours | 49.07 | 58.95 | 0.90 |
Closed-loop driving on Town05 Long, 9 routes with one repetition.
DS, RC, and IS denote Driving Score, Route Completion, and Infraction Score. none marks a value the source does not report. ‡ Reported by GenAD.
Ablation studies
Effects of pretraining targets, feature masking, action conditioning, and training schedules.
Pretraining targets and head size
VAD-Tiny downstream performance across pretraining targets and auxiliary-head sizes, using 20 pretraining epochs for every pretrained variant.
| Pretraining supervision | Plan. | Mot. | Det. | Map | |||||
|---|---|---|---|---|---|---|---|---|---|
| Det. | Map | Occ. | Act. | Auxiliary parameters ↓ | L2 (m) ↓ | Col. (%) ↓ | ADE (m) ↓ | NDS ↑ | mAP ↑ |
Pretraining supervision against auxiliary-head budget.
All pretraining runs use 20 epochs. Occ. and Act. denote occupancy and PAVER’s action conditioning. No task-supervised row matches the collision rate and NDS of sparse action targets, at 1/300 of the auxiliary parameters.
Feature masking and action conditioning
VAD-Tiny downstream performance with action-corridor feature masking and action-state conditioning enabled separately or together.
| Planning | Motion | Det. | Map | |||||
|---|---|---|---|---|---|---|---|---|
| Mask | Action | L2 (m) ↓ | Col. (%) ↓ | ADE (m) ↓ | FDE (m) ↓ | MR ↓ | NDS ↑ | mAP ↑ |
Feature masking and action-state conditioning, separately and together.
Mask denotes the corridor mask; Action state denotes the 4-D state concatenated at readout.
Pretraining and fine-tuning schedules
Downstream performance and estimated total training time for VAD-Tiny and VAD-Base across pretraining and fine-tuning schedules.
| Training cost | Planning | Mot. | Det. | Map | ||||
|---|---|---|---|---|---|---|---|---|
| Method | PT (epochs) | FT (epochs) | Time | L2 (m) ↓ | Col. (%) ↓ | ADE (m) ↓ | NDS ↑ | mAP ↑ |
| VAD-Tiny† | 48‡ | 12 | 20.5h | 0.701 | 0.730 | 0.840 | 0.363 | 0.413 |
| VAD-Tiny† | 0 | 60 | 21.3h | 0.662 | 0.513 | 0.905 | 0.338 | 0.419 |
| + Ours | 20 | 10 | 6.5h | 0.637 | 0.300 | 0.851 | 0.357 | 0.361 |
| + Ours | 20 | 20 | 10.1h | 0.647 | 0.220 | 0.781 | 0.387 | 0.425 |
| + Ours | 20 | 30 | 13.6h | 0.603 | 0.187 | 0.803 | 0.397 | 0.439 |
| VAD-Base | 0 | 60 | 103h | 0.742 | 0.307 | 0.760 | 0.422 | 0.502 |
| + Ours | 20 | 10 | 31.8h | 0.646 | 0.530 | 0.772 | 0.371 | 0.403 |
| + Ours | 20 | 20 | 49.0h | 0.757 | 0.280 | 0.698 | 0.442 | 0.460 |
| + Ours | 20 | 30 | 66.1h | 0.562 | 0.401 | 0.685 | 0.446 | 0.498 |
Pretraining and fine-tuning epoch budgets and estimated total training time.
Total training time, including pretraining, is estimated from mean iteration times on 4 NVIDIA RTX 5090 GPUs. † Reproduced. ‡ Stage 1 pretrains detection, mapping, and motion prediction without planning.
Representation analysis
Analyses of BEV representations, pretraining targets, and downstream learning dynamics.
Downstream learning curves
Supplementary results
Full metrics
Complete planning, motion prediction, detection, and mapping results, followed by class-wise detection and map metrics.
Citation
BibTeX
@misc{song2026paver,
title = {Planning-Aligned Pretraining of BEV Representations with Sparse Action-Conditioned Targets for End-to-End Autonomous Driving},
author = {Jaeha Song and Soonmin Hwang},
year = {2026},
eprint = {2609.22868},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
doi = {10.48550/arXiv.2609.22868},
url = {https://arxiv.org/abs/2609.22868}
}