中文说明见 README_zh.md
A lightweight experiment launcher for multi-GPU, multi-host research workflows.
multi_task provides a single entrypoint for running training, evaluation, and inference jobs across local and remote GPU machines. It is designed for researchers or individual developers who want a simple, file-based workflow instead of a full cluster scheduler.
The core idea is straightforward:
- submit every job through one command
- prefer an idle GPU on the current machine
- fall back to a remote machine when needed
- record status, logs, summary, and metrics in a stable job directory
- let humans or Claude Code read those artifacts afterward
When experiments are launched ad hoc, it becomes hard to answer basic questions:
- Where did the job run?
- Is it still running?
- What command was actually executed?
- Did it succeed or fail?
- What was the final metric?
multi_task solves that by wrapping job submission in a thin orchestration layer. It does not try to be Kubernetes, Slurm, or Ray. Instead, it keeps the workflow simple:
- a small set of CLI entrypoints
- plain files as the source of truth
- optional SSH-based remote dispatch
- minimal third-party dependencies
- One unified entrypoint for
train,eval, andinfer - Local GPU preferred before remote dispatch
- File-based job state tracking
- Stable, machine-readable stdout protocol after submission
- Per-job output directory with logs and metadata
- Automatic
summary.mdgeneration - Optional
metrics.jsonextraction from logs - Optional notification integration
- Lightweight SSH bootstrap helper for passwordless access
At a high level, a job goes through this flow:
User / Claude Code
|
v
cc_run_experiment
|
+--> inspect local GPU availability
|
+--> local path: launch process on current machine
|
\--> remote path: submit through SSH-based adapter
to another machine
Both paths write to:
jobs/<job_id>/
- meta.json
- status.json
- command.sh
- train.log
- summary.md
- metrics.json (optional)
The result is that local and remote jobs look the same from the outside: they always produce a local job directory that can be inspected later.
multi_task/
├── cc_run_experiment # main unified launcher
├── cc_job_status # print the current status of a job
├── cc_job_wait # wait until a job reaches a terminal state
├── cc_job_notify # send a completion/failure notification
├── gpu_probe.py # inspect local GPU state via nvidia-smi
├── list_idle.py # list idle GPUs across configured hosts
├── submit_experiment.py # remote submission adapter
├── job_status.py # refresh status for remote jobs
├── remote_launch.sh # shell wrapper used on remote workers
├── setup_passwordless_ssh.py # generic SSH bootstrap helper
├── lib/
│ ├── __init__.py
│ ├── gpu_utils.py # GPU discovery, filtering, lease/lock handling
│ ├── notify.py # notification integration
│ ├── scheduler.py # local/remote dispatch and summary orchestration
│ └── status_io.py # status/meta read-write helpers
└── jobs/ # runtime output, ignored by git
The main entrypoint. It:
- validates arguments
- validates the conda environment
- creates a job directory
- picks local vs remote execution
- starts the job
- emits stable key-value output for downstream parsing
Reads jobs/<job_id>/status.json and prints a compact JSON summary.
Polls a job until it finishes, then prints the summary and metrics paths.
Sends a notification after a job succeeds or fails.
Uses nvidia-smi to inspect GPU memory and utilization on the current machine.
Uses SSH plus gpu_probe.py to inspect multiple hosts and show which GPUs are idle.
Handles remote submission and compatibility with the existing SSH-based remote execution path.
Refreshes remote job state and synchronizes it back to the local job record.
Shell wrapper executed on a remote worker. It prepares environment activation, writes status files, and captures exit codes.
A generic helper that installs an SSH public key onto a set of hosts and optionally appends an SSH config block.
Responsible for:
- querying local GPUs
- applying GPU filtering rules
- choosing candidates
- maintaining lease/lock files so concurrent submissions do not grab the same GPU
Responsible for:
- atomic status writes
- metadata writes
- state updates
- generating stable key-value submission output
Responsible for:
- local submission
- remote submission
- remote-to-local status synchronization
- metric extraction
- summary generation
Wraps the optional notification backend.
This repository is intentionally lightweight and expects your environment to provide the basics.
Typical assumptions are:
- Linux-based machines
- Python 3
nvidia-smiavailable on worker machines- Conda or compatible environment manager
- SSH access between machines
- a shared or otherwise coordinated filesystem layout for code and job artifacts
This project is not plug-and-play for every infrastructure setup. You should expect to adapt host inventory, SSH behavior, and directory conventions to your own environment.
./cc_run_experiment \
--cwd "$PWD" \
--task-type train \
--name "demo-train" \
--conda-env your-env \
--async \
-- python train.py --config configs/foo.yaml./cc_run_experiment \
--cwd "$PWD" \
--task-type eval \
--name "demo-eval" \
--conda-prefix /path/to/conda/env \
--async \
-- python eval.py --config configs/foo.yaml./cc_job_status <job_id>./cc_job_wait <job_id> --timeout 3600 --poll-interval 15./cc_job_notify <job_id>After a submission, cc_run_experiment prints stable key-value lines such as:
JOB_ID=exp_20260421_153012_ab12cd
STATE=queued
MODE=local
HOST=my-machine
WORK_DIR=/path/to/repo
JOB_DIR=/path/to/multi_task/jobs/exp_20260421_153012_ab12cd
STATUS_PATH=/path/to/multi_task/jobs/exp_20260421_153012_ab12cd/status.json
LOG_PATH=/path/to/multi_task/jobs/exp_20260421_153012_ab12cd/train.log
SUMMARY_PATH=/path/to/multi_task/jobs/exp_20260421_153012_ab12cd/summary.md
METRICS_PATH=/path/to/multi_task/jobs/exp_20260421_153012_ab12cd/metrics.json
PID=12345
REMOTE_JOB_ID=
ENV_KIND=conda
CONDA_ENV=your-env
CONDA_PREFIX=
PYTHON_BIN=
This makes it easy for other tools, wrappers, or agents to parse job creation results.
Each job writes to jobs/<job_id>/.
Common files include:
meta.json— metadata about the declared taskstatus.json— current state, paths, environment info, and execution modecommand.sh— the business command that was launchedtrain.log— captured stdout/stderr log streamsummary.md— generated human-readable summarymetrics.json— structured metrics, if available or extracted
This directory is the main debugging surface for both humans and automation.
If your training script does not write metrics.json, multi_task can extract simple scalar values from logs.
Examples:
val_accuracy=0.9123
train_loss: 0.321
best_epoch=4
The extracted metrics are written to metrics.json, and a primary metric is selected automatically when possible.
The repository includes setup_passwordless_ssh.py as a generic helper for installing public keys on multiple hosts.
Typical usage pattern:
python3 setup_passwordless_ssh.py \
--hosts-file hosts.txt \
--user your-ssh-user \
--host-pattern 'gpu-*'This helper is intentionally generic. You should review and adapt it before using it in a production or shared environment.
This repository does not include:
- private host inventories
- historical jobs
- runtime lock files
- internal handoff notes
- private experiment logs
- private machine addresses or credentials
- The project assumes a Linux/SSH/Conda style workflow
- Remote execution is adapter-based, not a full scheduler abstraction
- Some infrastructure behavior is environment-specific and may require customization
- It is optimized for simple research workflows, not for multi-tenant production clusters
This project is released under the MIT License. See LICENSE.
This repository is provided as a lightweight research workflow utility, not as a hardened production orchestration system.
You are responsible for reviewing and adapting:
- SSH configuration
- host inventory management
- filesystem layout
- credential handling
- environment activation
- access control and operational safety
Use it only in environments where you understand and accept those tradeoffs.