AutoE2E: End-to-End AI for Self-Driving
AutoE2E is a free and fully open-source End-to-End AI model from the Autoware Foundation. It enables autonomous driving across highways, arterial roads and city streets using cameras only, without reliance on HD maps. One network turns surround-view camera images and the vehicle's own motion history into the future driving trajectory.
AutoE2E outputs can be fused with physics-based sensors such as LiDAR and radar to power fully driverless robotaxi applications. The baseline camera-only model can enable L2++ automotive ADAS applications for point-to-point hands-free navigation.
This repository hosts the AutoE2E v1.0 release: two trajectory-planning checkpoints, the evaluation reports behind the published metrics, and the KITScenes validation shards used by the community benchmark.
| Checkpoint | Trained on | Use it for |
|---|---|---|
models/nuplan-epoch-5 |
nuPlan v1.1 train sensor logs | The baseline model. Zero-shot evaluation on new cities, and the starting point for fine-tuning on your own data |
models/kitscenes-epoch-5 |
nuPlan Epoch 5, fine-tuned on KITScenes-Multimodal train | Driving in Karlsruhe and the KITScenes camera rig |
Vision and goals
Autonomous driving software has been built from many hand-engineered modules and, increasingly, from closed end-to-end models that nobody outside a company can study or improve. AutoE2E takes the opposite path. Everything is developed in the open by the Autoware community: the model code, the data pipelines, the training platform, the evaluation scripts, and the checkpoints published here.
The goals of the project are:
- deliver a camera-first driving model that works on highways, arterial roads and city streets without HD maps;
- keep the full stack reproducible, so anyone can retrain, benchmark and fine-tune AutoE2E on their own vehicles and cities;
- provide a planning core that can be combined with LiDAR/radar safety layers for driverless robotaxis, and used camera-only for L2++ ADAS;
- benchmark openly against other end-to-end models on shared data and shared scripts (issue #210).
The work happens in the Autoware Robotaxi working group. New contributors are welcome; the onboarding guide explains how to join the weekly meetings.
Architecture
AutoE2E is designed as three cooperating models:
- a Reactive model at 10 Hz that plans the trajectory in real time from multi-view images, a map raster and ego-motion history;
- a World Action Model at 1 Hz that learns to predict future visual features (JEPA-style) from the encoded visual history;
- a Reasoning model at 1 Hz that detects edge-case scenarios and conditions the planner.
The v1.0 checkpoints contain the Reactive model, which is the part that drives. The World Action and Reasoning branches are disabled in these weights and remain under active development.
| Component | v1.0 checkpoint configuration |
|---|---|
| Cameras | 6 surround views at 512 × 512; the front camera is also encoded at 1024 × 1024 |
| Camera history | 8 frames at 0.5 s spacing (3.5 s), BEVFormer V2 temporal fusion |
| Camera BEV encoder | ResNet-50 + BEVFormer V2 (T8), 300 × 200 BEV grid covering 180 m × 120 m (60 m behind to 120 m ahead), frozen from the official BEVFormer V2 R50 T8 initialization |
| Map input (optional) | 14-channel semantic map raster plus a 2-channel route raster, 450 × 300 cells at 0.4 m. It can be left empty; the KITScenes Test results below use cameras only |
| Map/image fusion | Deformable cross-attention from map features into the camera BEV |
| Ego-motion history | 6.4 s at 10 Hz: speed, longitudinal acceleration, yaw rate, curvature |
| Trajectory planner | Deterministic GRU planner with deformable BEV feature lookup |
| Output | 64 steps × (longitudinal acceleration, curvature) at 10 Hz = 6.4 s, integrated into an ego-frame XY trajectory |
| Auxiliary heads | 8-class BEV segmentation and route reconstruction (training signals) |
| Parameters | 79.9 M |
Results
All values are ADE / FDE in metres for a single predicted trajectory; lower is better. "Lateral" and "longitudinal" are the mean absolute errors across the valid horizon.
KITScenes-Multimodal Test (official split, cameras only)
206 scenes, 23,690 samples. The model receives cameras and ego-motion only; map and route inputs are absent. Ground truth comes from the official KITScenes poses.txt. Neither checkpoint saw these scenes during training.
| Model | 1 s | 2 s | 3 s | 5 s | Lateral | Longitudinal |
|---|---|---|---|---|---|---|
| nuPlan Epoch 5 (zero-shot) | 0.160 / 0.313 | 0.398 / 0.973 | 0.792 / 2.162 | 2.112 / 6.175 | 1.058 | 1.550 |
| KITScenes Epoch 5 | 0.146 / 0.311 | 0.412 / 1.045 | 0.826 / 2.220 | 2.104 / 5.942 | 1.204 | 1.405 |
KITScenes-Multimodal Validation (cameras + HD map + route)
117 scenes, 11,035 samples. The route is an oracle reconstructed from the logged future trajectory, so these numbers measure route-conditioned driving rather than online route planning. Ground truth is the trajectory_xy_m field of the AutoE2E KITScenes v3.5 shards published under datasets/kitscenes-val. That field was produced by an older coordinate conversion; rescoring against the official poses.txt is tracked in issue #210.
| Model | 1 s | 2 s | 3 s | 5 s | Lateral | Longitudinal |
|---|---|---|---|---|---|---|
| nuPlan Epoch 5 (zero-shot) | 0.195 / 0.378 | 0.472 / 1.133 | 0.915 / 2.423 | 2.295 / 6.388 | 1.191 | 1.663 |
| KITScenes Epoch 5 | 0.147 / 0.285 | 0.373 / 0.924 | 0.745 / 2.020 | 1.941 / 5.564 | 0.984 | 1.402 |
KITScenes labels stop at 5 s, so the 6.4 s horizon cannot be scored on KITScenes.
Held-out validation during training (6.4 s)
| Model | Validation set | Samples | ADE @ 6.4 s | FDE @ 6.4 s |
|---|---|---|---|---|
| nuPlan Epoch 5 | nuPlan held-out training logs | 1,024 | 1.267 | 3.898 |
| KITScenes Epoch 5 | KITScenes train-dev scenes | 3,820 | 2.633 | 7.557 |
These sets were used to choose the epoch, so treat them as in-distribution reference values rather than benchmark scores.
The per-run reports, including sample counts and dataset digests, are under evaluations/overlay-replay-v1. They score the predictions stored for the dashboard. The GRU planner is deterministic, so these stored controls are the model's predictions for every sample. Comparisons with other end-to-end models (METEOR, VAD, UniAD, Drive-JEPA, Qwen-Drive, Alpamayo, SimForge) are coordinated in issue #210, where each protocol difference is documented.
Predictions
Purple is the logged ground-truth path and green is the AutoE2E prediction, drawn into the cameras and the bird's-eye view. Every frame of these scenes can be played back in the DataModelConsole dashboard.
KITScenes Epoch 5 entering a roundabout in Karlsruhe (validation, map and route available).
nuPlan Epoch 5, never trained on German roads, turning left at a night-time intersection (KITScenes Test, cameras only).
nuPlan Epoch 5 following a left curve on a forest road (KITScenes Test, cameras only).
Camera images: KITScenes-Multimodal, CC BY-NC 4.0.
Quick start
The checkpoints are PyTorch training checkpoints. They load with the AutoE2E source at the revision used to train them:
git clone https://github.com/autowarefoundation/auto_e2e.git
cd auto_e2e
git checkout 40a75cf8a26b60f6b094125e3924fe0bdb427b02 # experimental/reactive-bev-learning
pip install -r requirements.txt huggingface_hub
import json
import sys
import torch
from huggingface_hub import hf_hub_download
sys.path.insert(0, "Model")
from model_components.auto_e2e import AutoE2E
from training.reactive_multitask import ReactiveTrainingStage, reactive_model_kwargs
repo_id = "AutowareFoundation/auto_e2e"
release = json.load(open(hf_hub_download(repo_id, "config.json")))
entry = release["checkpoints"]["nuplan-epoch-5"] # or "kitscenes-epoch-5"
checkpoint_path = hf_hub_download(repo_id, entry["path"])
payload = torch.load(checkpoint_path, map_location="cpu", weights_only=True)
config = payload["config"]
model = AutoE2E(
backbone=config["backbone"],
embed_dim=config["embed_dim"],
is_pretrained=False,
**reactive_model_kwargs(
ReactiveTrainingStage(config["training_stage"]),
num_views=config["num_views"],
),
)
model.load_state_dict(payload["model_state_dict"], strict=True)
model.eval()
Each checkpoint also contains the optimizer, scheduler and training state, so it can resume or seed fine-tuning directly. The evaluation entry point for KITScenes is evaluate_reactive_kitscenes_checkpoint in Platform/pipelines/distributed_training.py.
Files
| Path | Contents |
|---|---|
config.json |
Release description: architecture, input/output contract, checkpoint paths and SHA-256 digests |
models/<id>/checkpoint.pt |
PyTorch checkpoint: model_state_dict, config, metrics, optimizer, scheduler, training state |
models/<id>/metadata.json |
Lineage, training stage, evaluation summary and MLflow registry identifiers |
evaluations/overlay-replay-v1/ |
Evaluation reports for the tables above |
datasets/kitscenes-val/ |
AutoE2E KITScenes validation shards (v3.5) used by the issue #210 benchmark |
release_manifest.json, SHA256SUMS |
Release identity and checksums of every release file |
Verify a download with sha256sum -c SHA256SUMS. This release is also available at the Git tag v1.0.
Training
| nuPlan Epoch 5 | KITScenes Epoch 5 | |
|---|---|---|
| Initialization | Official BEVFormer V2 R50 T8 weights | nuPlan Epoch 5 (ca8b43d7…) |
| Data | nuPlan v1.1 train sensor logs, 6 cameras (CAM_F0, CAM_L0, CAM_R0, CAM_L2, CAM_B0, CAM_R2) |
KITScenes-Multimodal train, 6 cameras, 10 % of scenes held out |
| Objective | Trajectory imitation and route reconstruction, camera BEV frozen | Same objective, camera BEV frozen |
| Hardware | 8 GPUs, bf16, global batch 16 | 8 GPUs, bf16, global batch 16 |
Intended use and limitations
AutoE2E v1.0 is a research model for open-loop evaluation, visualization, benchmarking and further training. It is not a certified driving system and must not control a vehicle on public roads.
- The results are open-loop: each prediction is compared with the logged human trajectory. Closed-loop behaviour is not measured here.
- The camera rigs of nuPlan and KITScenes differ. Other rigs need fine-tuning before the numbers above can be expected.
- The KITScenes Validation route input is an oracle derived from the future trajectory.
- Confidence intervals and per-scenario breakdowns are not yet published.
License
The AutoE2E source code is released under Apache-2.0. These checkpoints are distributed under the following notice instead, because they are derived from third-party weights and datasets:
- the camera encoder was initialized from the official BEVFormer V2 R50 T8 checkpoint, whose weight license is not asserted upstream and which was trained on nuScenes;
- the models were trained on nuPlan and KITScenes-Multimodal (CC BY-NC 4.0).
Users are responsible for complying with the terms of the upstream weights and datasets, which include non-commercial restrictions. No broader license is granted for the checkpoint files. The KITScenes validation shards under datasets/kitscenes-val remain subject to CC BY-NC 4.0 and the additional terms on the original dataset page.
Citation
@software{autoware_autoe2e_2026,
author = {{The Autoware Foundation}},
title = {AutoE2E: Open End-to-End AI for Self-Driving},
year = {2026},
version = {1.0},
url = {https://github.com/autowarefoundation/auto_e2e},
note = {Checkpoints: https://huggingface.co/AutowareFoundation/auto_e2e}
}
- Downloads last month
- 11