kubectl exec, depending on your directory service configuration. SSH requires a directory service pre-configured for SSH access, while kubectl exec doesn’t. This page covers SSH (with and without port forwarding), running your first Slurm command to confirm access, using kubectl exec when SSH isn’t available, and troubleshooting a login node that’s slow or drops your session.
For information about initial setup of Slurm login nodes, see Configure Slurm individual login nodes.
Before you connect
Before you can reach a login node, both an administrator and each user must finish provisioning. A missing step here is the most common reason a first connection fails:- Administrator: enable the SCIM API and SUNK User Provisioning, then create the user groups your cluster uses for login and admin access. The Cloud Console creates
slurm-usersandsudo-usersby default, but a cluster set up with custom groups uses whatever names you defined. See Provision users in SUNK andnsscacheconfiguration. - Each user: add at least one SSH public key to the Slurm Attributes section of your Cloud Console settings, then confirm with your administrator that you belong to the required groups. See Provision users in SUNK.
Connect through SSH
SSH is the preferred way to reach a Slurm login node when the cluster exposes the login service on a public IP address. Use this method when your directory service is already configured for SSH access.Accessing the login node through SSH requires a directory service with users configured for SSH access.
kubectl get svc slurm-login command to identify the login service’s IP address or DNS record so you know where to point your SSH client. The EXTERNAL-IP field in the command output contains the IP address.
In the following example, the target IP address is 203.0.113.100:
Obtain the External IP address
Log in with SSH
Connect through port forwarding
Use port forwarding when no public IP address is allocated for the login node, so you can still reach it through SSH from your local machine. If no public IP address is allocated for the node, first port-forward the service with thekubectl port-forward command, then log in through SSH using the port-forwarded address. Each login Pod has an associated headless service, allowing users to refer to the pod by name without specifying a fully qualified domain name (FQDN).
To access an individual login Pod with port-forwarding, use the kubectl port-forward and ssh commands, as shown in the following example:
Log in with port-forwarding
kubectl port-forward svc/slurm-login-slurmuser1 10022:22, works as follows:
- The
kubectl port-forwardcommand creates a port-forward. svc/specifies that the targeted resource is a Service.slurm-login-slurmuser1is the exact name of the targeted Kubernetes Service. Replace this value with the name used within your namespace.10022:22defines the port mapping. In this case, it forwards traffic from local port10022to port22on the target Service.
ssh example-user@localhost -p 10022, then connects to the local port 10022. Because of the port-forwarding performed in the preceding command, Kubernetes sends this traffic to port 22 of the specified Kubernetes Service.
You’re now logged into the Slurm login node and can run Slurm commands.
Run Slurm commands
Once you’re connected to the login node, confirm that the cluster is reachable from your session and that your user has permission to submit work. After you log in, you have access to all normal Slurm operations to submit jobs or manage the cluster. SchedMD provides extensive documentation for Slurm commands and some printable cheat-sheets. To verify that the cluster is working, run a small job. For example, discover the hostname on six nodes, as shown in the following example:[USERNAME] with your Slurm username:
Add a Slurm user on a cluster without SUNK User Provisioning
root, run the preceding command for each account you need to be added to.
Troubleshooting
If SSH is unavailable or you need root access for debugging, you can fall back tokubectl exec to open a shell directly on the login Pod.
When SSH isn’t possible, use kubectl exec to access the Slurm login node as root. This method is useful for debugging and maintenance tasks.
Access the Slurm login node with kubectl exec
Slow shell commands
Slowcd, ls, or git on a login Pod usually comes from memory pressure, shared-storage contention, or a heavy neighbor process on a shared login Pod. The login Pod isn’t a compute node, and a common cause for delay is interactive tooling that quietly consumes the whole pod: security scanners, AI coding agents, package builds such as uv sync or ninja, JVM-based IDEs, rclone transfers, and IDE remote-server processes. Move that work to srun or sbatch. Isolate shared-storage contention by comparing a local path to a shared path:
/tmp is fast but the shared path is slow, the contention is in shared storage, not the pod.
SSH disconnects mid-session
Start by asking whether you’re the only one affected. That narrows the cause faster than any other check.- Resource pressure or OOM, if only sessions on your pod dropped. If the login Pod is OOM-killed, every session on it drops. A non-zero
RESTARTSvalue onkubectl get pod [LOGIN-POD-NAME]around the time of the disconnect indicates that the pod was killed, often for memory. For related guidance on compute-node memory pressure, see Diagnose unexplained memory pressure on a SUNK node. - Node-level CPU contention, if your pod looks healthy but sessions still drop. Workloads from other namespaces scheduled onto the same node can starve
sshduntil keepalives time out. Per-pod CPU can look normal while the node is saturated, so check the node rather than the pod. This is a node-level condition you can’t fix from inside the pod, so escalate. - Authentication backend trouble, if many users lost their sessions at the same moment across different pods. The directory service that authenticates SSH can drop its backend connection and reject every authentication until it reconnects, usually within minutes. Simultaneous disconnects across pods on different nodes point here rather than at any one node.
- A transient network drop, if SSH was unreachable for a minute or two and then recovered on its own with no pod restart and no node alert. Escalate if it recurs on the same node.
tmux, and set ServerAliveInterval in your local ~/.ssh/config. If tmux reports server exited unexpectedly on a pod that has been running for weeks, clear its stale sockets with rm -rf /tmp/tmux-$(id -u)/ and start a new session.