# DM0.5 for LeRobot

DM05, also released as DM0.5, is Dexmal's Vision-Language-Action model for open-world robot control. The LeRobot
adapter preserves the OpenDM model path while using standard LeRobot datasets, processors, training,
checkpointing, Hub loading, and evaluation.

For model details, see the [DM0.5 technical blog](https://www.dexmal.com/blog/dm0.5/index_en.html), the
[OpenDM repository](https://github.com/Dexmal/OpenDM), and the raw
[Dexmal/DM05](https://huggingface.co/Dexmal/DM05) release.

## Installation

```bash
pip install -e ".[training,dm05]"   # training
pip install -e ".[libero,dm05]"     # LIBERO evaluation on Linux
```

## Checkpoint

Use `lerobot/dm05_base` with `--policy.path`. It is a self-contained LeRobot conversion of the raw OpenDM
checkpoint and supports `DM05Policy.from_pretrained()`.

The base config records OpenDM's 14-dimensional state/action contract. DM05's core model supports up to 32
dimensions; fresh fine-tuning resolves the effective image, state, and action features from the target dataset.

## Dataset contract

DM05 accepts standard LeRobot datasets with:

- one or more image or video observations;
- `observation.state`;
- `action`;
- task descriptions.

Set `policy.image_keys` when camera order must be explicit. RoboTwin, ALOHA, and other sources do not need a
DM05-specific reader after conversion to the standard LeRobot schema. Keep the standard `observation.state` and
`action` keys; use `rename_map` only to align camera names.

## Training

The LIBERO recipe follows OpenDM:

```bash
lerobot-train \
  --dataset.repo_id=lerobot/libero \
  --rename_map='{"observation.images.image": "observation.images.front", "observation.images.image2": "observation.images.wrist"}' \
  --dataset.video_backend=pyav \
  --policy.path=lerobot/dm05_base \
  --policy.add_state=false \
  --policy.chunk_size=10 \
  --policy.n_action_steps=10 \
  --policy.repo_id=your_repo_id \
  --output_dir=outputs/train/dm05-libero \
  --steps=50000 \
  --batch_size=8 \
  --policy.device=cuda
```

### Key training parameters

| Parameter                     | LIBERO value                        | Meaning                                                              |
| ----------------------------- | ----------------------------------- | -------------------------------------------------------------------- |
| `dataset.repo_id`             | `lerobot/libero`                    | Standard LeRobot training dataset                                    |
| `policy.path`                 | `lerobot/dm05_base`                 | Self-contained base checkpoint                                       |
| `rename_map`                  | `image`/`image2` to `front`/`wrist` | Camera names used by the eval command and `lerobot/dm05_libero`      |
| `policy.add_state`            | `false`                             | Matches OpenDM's LIBERO prompt without state tokens                  |
| `policy.use_relative_actions` | `false` (checkpoint default)        | Learns stored actions unchanged; relative mode is an explicit opt-in |
| `policy.chunk_size`           | `10`                                | Number of actions predicted per chunk                                |
| `policy.n_action_steps`       | `10`                                | Number executed before the next model call; at most `chunk_size`     |
| `policy.repo_id`              | `your_repo_id`                      | Hub destination; use `policy.push_to_hub=false` for local-only runs  |

Environment control mode is configured separately from the policy action representation. Relative mode requires
matching state/action dimensions.

### Cameras

The camera set is part of the policy config, and the prompt always renders exactly those cameras. Fine-tuning from
`lerobot/dm05_base`, which declares none, takes the training dataset's cameras. A fine-tuned checkpoint keeps its
own camera names: map a dataset with different names onto them with `--rename_map`, for example
`--rename_map='{"observation.images.top": "observation.images.front"}'`.

### Normalization statistics

Training uses the dataset's state/action statistics in `meta/stats.json`. Action statistics must match the representation used for training. With `policy.use_relative_actions=true`, compute them after the relative-action transform, using the same chunk size and excluded dimensions. See the [relative-action guide](./action_representations#using-relative-actions-in-lerobot).

## Evaluation

The LIBERO checkpoint `lerobot/dm05_libero` was trained for 50,000 steps and evaluated on all 40
LIBERO tasks with 5 episodes per task:

```bash
MUJOCO_GL=egl lerobot-eval \
  --policy.path=lerobot/dm05_libero \
  --env.type=libero \
  --env.task=libero_spatial,libero_object,libero_goal,libero_10 \
  --env.camera_name_mapping='{"agentview_image":"front","robot0_eye_in_hand_image":"wrist"}' \
  --env.observation_height=256 \
  --env.observation_width=256 \
  --env.control_mode=relative \
  --eval.n_episodes=5 \
  --eval.batch_size=1 \
  --seed=7 \
  --policy.device=cuda
```

| Suite     |           Successes |
| --------- | ------------------: |
| Spatial   |               49/50 |
| Object    |               50/50 |
| Goal      |               50/50 |
| LIBERO-10 |               48/50 |
| **Total** | **197/200 (98.5%)** |

This is a 200-episode reproduction check, not the 2,000-episode LIBERO benchmark protocol.

## Checkpoint layout

Use the complete checkpoint directory, which contains policy weights, config, and processor state. A standalone
`model.safetensors` is not sufficient for training or inference.

## License

The LeRobot integration is Apache-2.0. Model weights follow the license attached to the corresponding DM05
model card.

