sbatch, srun, squeue, sinfo, sacct, and scancel work on the login node and in a bare srun. Inside srun --container-image, they return command not found. When you can, submit dependent jobs from the login node. When a job must call sbatch or srun based on results it produces at runtime, mount the host Slurm CLI into the container.
Submit from the login node
Chain jobs with dependencies so the container never needs Slurm binaries:Mount the host Slurm CLI
When a running job must callsbatch or srun based on runtime results, bind-mount the host Slurm binaries, configuration, Munge client library, and Munge socket. The following example is for an x86_64 node. Replace [USERNAME], [IMAGE], and [SCRIPT] with your values.
/usr/lib/aarch64-linux-gnu instead. Confirm the paths on the node before you copy the list:
--container-image job from inside the container, also mount /opt/sunk:/opt/sunk and /etc/enroot:/etc/enroot.
Install slurm-client in the image
A Slurm client works only when its version is inside the range the cluster’s controller accepts. Slurm 24.11 and later accept commands from the current release and the three previous major releases. A client outside that window fails with a protocol version error even when the configuration and the Munge socket are mounted.
Distribution packages are usually too old to qualify. Ubuntu 24.04 ships slurm-client 23.11, Debian 12 ships 22.05, and Ubuntu 22.04 ships 21.08. Run sbatch --version on the login node to read the cluster version, then compare it against the package your base image installs.
When the versions line up, install the package in the image:
/etc/slurm, /etc/munge, and /run/munge so the client can authenticate to the cluster. When the versions don’t line up, bind-mount the host Slurm CLI instead.
Related
- How do I run containers with Pyxis and Enroot in Slurm?
- Submit a training job with PyTorch or TensorFlow
- Connect to the Slurm login node
Workload Scheduling