Kamil Dreczkowski*1, Pietro Vitiello*1, Vitalis Vosylius1, Edward Johns1
*These authors contributed equally to this work.
1Imperial College London
This repository contains the implementation of all methods evaluated in the paper "Learning a Thousand Tasks in a Day". We provide model architectures, training scripts, and deployment examples.
Paper published on Science Robotics: https://www.science.org/doi/10.1126/scirobotics.adv7594
Paper published on Arxiv: https://arxiv.org/abs/2511.10110
Project Website: https://www.robot-learning.uk/learning-1000-tasks
This codebase implements five methods for learning manipulation tasks from limited per-task demonstrations:
- MT3: Retrieval-based alignment + retrieval-based interaction (no training required)
- Ret-BC: Retrieval-based alignment + behavioral cloning interaction
- BC-Ret: Behavioral cloning alignment + retrieval-based interaction
- BC-BC: Behavioral cloning alignment + behavioral cloning interaction
- MT-ACT+: End-to-end multi-task transformer (our adaptation of MT-ACT - see "RoboAgent: Generalization and Efficiency in Robot Manipulation via Semantic Augmentations and Action Chunking" for details)
Tested on: Ubuntu 22.04 LTS (Jammy) with Linux kernel 5.15.0-43-generic
- Docker - Install Docker for Ubuntu
- NVIDIA Container Toolkit (required for GPU support) - Installation guide
# 1. Clone the repository
git clone https://github.com/KamilDre/learning_thousand_tasks.git
cd learning_thousand_tasks
# 2. Download and extract demonstration data
# Download demonstrations.zip from Google Drive
wget --load-cookies /tmp/cookies.txt "https://docs.google.com/uc?export=download&confirm=$(wget --quiet --save-cookies /tmp/cookies.txt --keep-session-cookies --no-check-certificate 'https://docs.google.com/uc?export=download&id=1ZjGuX73LEgmMhVvHuQYwAqkJNFYC-Mwb' -O- | sed -rn 's/.*confirm=([0-9A-Za-z_]+).*/\1\n/p')&id=1ZjGuX73LEgmMhVvHuQYwAqkJNFYC-Mwb" -O demonstrations.zip && rm -rf /tmp/cookies.txt
unzip demonstrations.zip -d assets/
rm demonstrations.zip
# Download inference_example.zip from Google Drive
wget --load-cookies /tmp/cookies.txt "https://docs.google.com/uc?export=download&confirm=$(wget --quiet --save-cookies /tmp/cookies.txt --keep-session-cookies --no-check-certificate 'https://docs.google.com/uc?export=download&id=1TznEjtIi1o-3HR3dOeqbfLnYQgOMY9B_' -O- | sed -rn 's/.*confirm=([0-9A-Za-z_]+).*/\1\n/p')&id=1TznEjtIi1o-3HR3dOeqbfLnYQgOMY9B_" -O inference_example.zip && rm -rf /tmp/cookies.txt
unzip inference_example.zip -d assets/
rm inference_example.zip
# 3. Clone XMem for demonstration preprocessing (required for BC training)
git clone https://github.com/hkchengrex/XMem.git
cd XMem
git checkout v1.0 # Use stable v1.0 release
mkdir -p saves
wget https://github.com/hkchengrex/XMem/releases/download/v1.0/XMem.pth -O saves/XMem.pth
cd ..
# 4. Build the Docker image
make buildNote: If the wget commands for Google Drive downloads fail, you can manually download the files:
- Demonstrations - Extract to
assets/demonstrations/ - Inference Example - Extract to
assets/inference_example/
We provide a Makefile for convenience. All commands automatically handle Docker setup, X11 display forwarding for Open3D visualizations, and volume mounting.
make help # Show all available commands
make build # Build Docker image
make deploy_mt3 # Run MT3 inference example
make preprocess_demos # Preprocess demonstrations for BC training
make create_alignment_dataset # Create dataset for BC alignment training
make create_interaction_dataset # Create dataset for BC interaction training
make create_mtact_dataset # Create dataset for MT-ACT+ training
make train_bc_alignment # Train BC alignment policy
make train_bc_interaction # Train BC interaction policy
make train_mtact # Train MT-ACT+ policy
make debug # Start interactive shell in container
make stop # Stop running containerAfter completing the setup, you will have the following data to run MT3 immediately and showcase how to train the BC policies:
Demonstrations (assets/demonstrations/ - downloaded separately):
- 2× bottle grasping
- 1× shoe grasping
- Each includes RGB-D images, segmentation, bottleneck pose, and end-effector twists
Test Data (assets/inference_example/ - downloaded separately):
- Test RGB-D images:
head_camera_ws_rgb.png,head_camera_ws_depth_to_rgb.png - Pre-computed segmentation:
head_camera_ws_segmap.npy - Camera intrinsics:
head_camera_rgb_intrinsic_matrix.npy
Pre-trained Models (assets/ - included in repository):
geometry_encoder.ckpt- Pre-trained PointNet++ for geometry encodingpose_estimator.ckpt- Pre-trained 4-DOF pose regressor
Camera Extrinsics (assets/ - included in repository):
T_WC_head.npy- Head camera extrinsics (world-to-camera transform) from our experiments
MT3 does not require any task specific training. The supplied deployment script demonstrates the complete pipeline: retrieval, pose estimation, alignment, and interaction.
make deploy_mt3The script runs through 7 steps:
- Load test image - Loads RGB-D image and segmentation from
assets/inference_example/ - Initialize live scene state - Creates point cloud from segmented region
- Retrieve demonstration - Finds most similar demo from
assets/demonstrations/ - Load retrieved demonstration - Loads demo RGB-D and segmentation
- Estimate relative pose - Uses PointNet++ to compute transformation
- Transform bottleneck pose - Computes target pose for robot alignment
- Load end-effector twists - Loads demonstrated velocities for interaction
4 Open3D visualization windows (close each to continue):
- Live scene point cloud
- Live scene vs retrieved demo comparison
- Registration result showing aligned point clouds after PointNet++
- Registration result showing aligned point clouds after refinement
Saved visualizations in assets/example_visualisations/:
test_scene_visualization.png- RGB, depth, and segmentation of test sceneretrieval_visualization.png- Live scene vs retrieved demo (2×2 grid)
Console output showing:
- Retrieved demonstration name
- Estimated 4×4 transformation matrix
- Target bottleneck pose (T_WE) for alignment phase
- End-effector twist dimensions (N×7) for interaction phase
The provided demonstrations include workspace segmentation (first frame) which is sufficient for MT3 deployment. However, training BC policies requires per-timestep segmentation masks for the entire trajectory. This preprocessing step generates these masks using XMem.
In our experiments, we used LangSAM (Language Segment Anything Model) to obtain the initial workspace segmentation from language prompts (e.g., "grey shoe on table"). The preprocessing script assumes you have already segmented the first frame and saved it as head_camera_ws_segmap.npy. The script then:
- Loads the pre-computed workspace segmentation (
head_camera_ws_segmap.npy) - Tracks object through all timesteps using XMem
- Saves outputs:
head_camera_masks.npy- Per-timestep masks (for BC training)demo_video.mp4- RGB trajectory visualizationdemo_segmented_video.mp4- Segmentation overlay visualization
-
Ensure demonstrations have pre-computed workspace segmentation (
head_camera_ws_segmap.npy) for each task directory inassets/demonstrations/. -
Configure task names in
thousand_tasks/demo_preprocessing/preprocess_demos.pyby updating theTASK_NAMESlist with your demonstration directory names. -
Run preprocessing:
make preprocess_demosFor each demonstration task, the script generates:
task_name_0000/
├── head_camera_masks.npy # Per-timestep masks (T, H, W) for BC training
├── demo_video.mp4 # RGB trajectory video
└── demo_segmented_video.mp4 # RGB video with segmentation overlay
Note: If SAVE_PNG_IMAGES = True, an additional segmented_images/ directory will be created with per-frame PNG visualizations.
This section describes how to train the three policy types required for BC-based methods:
The BC alignment policy learns to predict trajectories to reach alignment poses. This policy is used to align the robot to the target pose before executing the interaction phase.
1. Create processed dataset:
make create_alignment_datasetThis will process demonstrations from assets/demonstrations/ and create a preprocessed dataset in assets/demonstrations/bn_reaching_processed/.
2. Train the policy:
make train_bc_alignmentThis will train the BC alignment policy and save checkpoints to assets/checkpoints/bc_alignment/. The best model (lowest combined position + rotation error) is saved as best.pt, and the final model is saved as final.pt.
Configuration: The default configuration generates 10 synthetic trajectories per demonstration (see num_traj_to_bn_train in thousand_tasks/training/act_bn_reaching/config.py). For training on your own demonstrations, you should likely change this to 1000 to generate sufficient training data.
The BC interaction policy learns to predict trajectories (end-effector poses) from the current state during manipulation. This policy replays the demonstrated interaction after alignment.
1. Create processed dataset:
make create_interaction_datasetThis will process demonstrations from assets/demonstrations/ and create a preprocessed dataset in assets/demonstrations/interaction_processed/.
2. Train the policy:
make train_bc_interactionThis will train the BC interaction policy and save checkpoints to assets/checkpoints/bc_interaction/. The best model is saved as best.pt, and the final model is saved as final.pt.
Configuration: The default configuration generates 10 interaction sub-trajectories per demonstration (see num_inter_traj in thousand_tasks/training/act_interaction/config.py). For training on your own demonstrations, you should likely change this to 1000 to generate sufficient training data.
MT-ACT+ is an end-to-end multi-task transformer that directly predicts action sequences from observations without explicit decomposition into alignment and interaction phases.
1. Create processed dataset:
make create_mtact_datasetThis will process demonstrations from assets/demonstrations/ and create a preprocessed dataset in assets/demonstrations/processed/.
2. Train the policy:
make train_mtactThis will train the MT-ACT+ policy and save checkpoints to assets/checkpoints/mtact_plus/. The best model is saved as best.pt, and the final model is saved as final.pt.
Configuration: The default configuration generates 10 full trajectories per demonstration (see num_inter_traj in thousand_tasks/training/act_end_to_end/config.py). For training on your own demonstrations, you should likely change this to 200 to generate sufficient training data.
After training BC policies, you can deploy the various method combinations. Each deployment script demonstrates the complete pipeline with visualization.
This method uses retrieval-based alignment (like MT3) but replaces retrieval-based interaction with a learned BC policy.
Prerequisites: Trained BC interaction policy at assets/checkpoints/bc_interaction/best.pt
make deploy_ret_bcPipeline Steps:
- Load test RGB-D image and segmentation
- Retrieve similar demonstration via hierarchical retrieval
- Estimate relative pose with PointNet++ and refine with ICP
- Apply 4DOF inductive bias and transform bottleneck pose to live scene
- Run BC interaction policy to predict waypoint trajectory
- Visualize predicted trajectory with Open3D
What You'll See:
- 4 Open3D windows showing point clouds and registrations (same as MT3)
- Final trajectory visualization with waypoints
- Console output with instructions for robot deployment
- Saved visualizations in
assets/example_visualisations/
This method uses a learned BC policy for alignment and retrieval-based interaction (like MT3).
Prerequisites: Trained BC alignment policy at assets/checkpoints/bc_alignment/best.pt
make deploy_bc_retPipeline Steps:
- Load test RGB-D image and segmentation
- Retrieve similar demonstration via hierarchical retrieval
- Run BC alignment policy to predict trajectory to bottleneck pose
- Visualize predicted alignment trajectory
- [SIMULATED] Track waypoints until bottleneck pose is reached
- Replay demonstrated end-effector velocities (like MT3)
What You'll See:
- 2 Open3D windows showing point clouds and retrieval comparison
- Alignment trajectory visualization with waypoints leading to target bottleneck
- Console output with instructions for robot deployment
- Saved visualizations in
assets/example_visualisations/
This method uses learned BC policies for both alignment and interaction phases, with no retrieval required.
Prerequisites:
- Trained BC alignment policy at
assets/checkpoints/bc_alignment/best.pt - Trained BC interaction policy at
assets/checkpoints/bc_interaction/best.pt
make deploy_bc_bcPipeline Steps:
- Load test RGB-D image and segmentation
- Run BC alignment policy to predict trajectory to bottleneck pose
- Track waypoints until alignment policy signals termination (terminate_prob > 0.95)
- Run BC interaction policy to predict manipulation trajectory
- Track waypoints with gripper control until interaction policy signals termination (terminate_prob > 0.95)
- Visualize both alignment and interaction trajectories
What You'll See:
- 3 Open3D windows: point cloud, alignment trajectory (green), interaction trajectory (red)
- Console output with instructions for both phases
- Saved visualizations in
assets/example_visualisations/
Key Difference from Other Methods:
- No retrieval required - policies generalize directly to novel objects
- Both phases use learned policies with closed-loop re-inference
- Alignment continues until bottleneck reached (terminate > 0.95)
- Interaction continues until task complete (terminate > 0.85)
This is an end-to-end baseline that directly predicts manipulation trajectories without explicit decomposition.
Prerequisites: Trained MT-ACT+ policy at assets/checkpoints/mtact_plus/best.pt
make deploy_mtactPipeline Steps:
- Load test RGB-D image and segmentation
- Run MT-ACT+ policy to predict end-to-end trajectory
- Track waypoints with gripper control
- Re-run inference in closed-loop until termination (terminate_prob > 0.95)
- Visualize predicted trajectory
What You'll See:
- 2 Open3D windows: point cloud and predicted trajectory (blue)
- Console output with instructions for closed-loop deployment
- Saved visualizations in
assets/example_visualisations/
Key Difference from Decomposition Methods:
- No explicit alignment + interaction phases
- Single policy handles complete task from start to finish
- Directly predicts waypoints without retrieving demonstrations
- Trained on full demonstration trajectories (not decomposed)
Each demonstration is a directory containing RGB-D observations, segmentation masks, robot poses, and interaction trajectories. Example from pick_up_grey_shoe:
| File | Shape | Type | Description |
|---|---|---|---|
| Required for MT3 | |||
head_camera_ws_rgb.png |
(720, 1280, 3) | uint8 | Workspace RGB image (gripper out of view) |
head_camera_ws_depth_to_rgb.png |
(720, 1280) | uint16 | Aligned depth in millimeters |
head_camera_ws_segmap.npy |
(720, 1280) | bool | Binary mask (True=object, False=background) |
head_camera_rgb_intrinsic_matrix.npy |
(3, 3) | float64 | Camera intrinsics [[fx,0,cx],[0,fy,cy],[0,0,1]] |
bottleneck_pose.npy |
(4, 4) | float64 | Target end-effector pose (SE(3) matrix) |
demo_eef_twists.npy |
(T, 7) | float64 | Velocity commands [vx,vy,vz,wx,wy,wz,gripper] |
task_name.txt |
- | text | Task description (e.g., "pick_up_grey_shoe") |
| Additional for BC Training | |||
head_camera_rgb.npy |
(T, 720, 1280, 3) | uint8 | Full RGB trajectory |
head_camera_depth_to_rgb.npy |
(T, 720, 1280) | uint16 | Full depth trajectory in millimeters |
head_camera_masks.npy |
(T, 720, 1280) | bool | Per-timestep masks (generated by preprocessing) |
| Optional | |||
geometry_encoding.npy |
(512,) | float32 | Pre-computed PointNet++ embedding |
workspace_img_T_WE.npy |
(4, 4) | float64 | Robot pose at workspace image capture |
Where T = number of timesteps (e.g., 179 frames at 30Hz = ~6 seconds)
-
Geometry Embeddings: The geometry encoder can be used to generate them. Retrieval should also regenerate them if they are missing.
-
Workspace Images ("ws" prefix): These are frames captured with the gripper moved out of the camera's field of view to avoid occluding the object.
-
End-Effector Twists: Velocities in the end-effector frame, not world frame. Format per row:
[vx, vy, vz, wx, wy, wz, gripper_next]where velocities are in m/s and rad/s, gripper is 0=close, 1=open. -
Bottleneck Pose: This is the alignment pose. It is just a pose suitable for the upcoming manipulation. Typically, the closer to the target object, the better.
-
Segmentation: The workspace segmentation (
head_camera_ws_segmap.npy) must be computed before using the demonstrations. We used LangSAM with language prompts in our experiments, but any segmentation method (SAM, manual annotation, etc.) can be used. For BC training, runmake preprocess_demosto generate per-timestep masks via XMem tracking.
If you find our code useful, please consider citing:
@article{doi:10.1126/scirobotics.adv7594,
author = {Kamil Dreczkowski and Pietro Vitiello and Vitalis Vosylius and Edward Johns},
title = {Learning a Thousand Tasks in a Day},
journal = {Science Robotics},
volume = {10},
number = {108},
pages = {eadv7594},
year = {2025},
doi = {10.1126/scirobotics.adv7594},
url = {https://www.science.org/doi/abs/10.1126/scirobotics.adv7594}
}