14 tháng 01, 2026

Quản trị cụm Slurm HPC và các thực hành tốt nhất

image

Ở Phần 1 và Phần 2, chúng ta đã xây dựng một cụm Slurm HPC hoàn chỉnh từ một node duy nhất đến một hệ thống đa node sẵn sàng cho production. Bây giờ hãy cùng tìm hiểu cách quản lý, bảo trì và bảo mật cụm này một cách hiệu quả.

Bài viết cuối cùng này bao quát các tác vụ quản trị hằng ngày, xử lý sự cố, tăng cường bảo mật và tích hợp với các framework xử lý dữ liệu.

Tổng quan series

  • Phần 1: Giới thiệu, kiến trúc và thiết lập một node đơn
  • Phần 2: Mở rộng lên production với Ansible
  • Phần 3 (bài viết này): Quản trị và các thực hành tốt nhất

Tổng quan về quản trị

Quản lý một cụm Slurm bao gồm một số lĩnh vực chính:

  • Quản lý cụm: Xây dựng, bảo trì và cập nhật cụm thông qua Ansible
  • Quản lý người dùng: Đồng bộ người dùng trên các node với quyền hạn phù hợp
  • Bảo mật đăng nhập: Tăng cường SSH với 2FA hoặc cặp khóa
  • Quản lý tài nguyên: Áp đặt giới hạn và chính sách fair-share
  • Giám sát: Theo dõi hiệu năng và mức sử dụng tài nguyên
  • Xử lý sự cố: Chẩn đoán và giải quyết vấn đề

Quản lý người dùng và tài nguyên

Thêm người dùng và nhóm

Slurm sử dụng account (nhóm) để tổ chức người dùng và áp dụng các chính sách tài nguyên:

Terminal window
# Add a new account/group
sacctmgr add account research_team Description="Research Team"
# Add a user to an account
sacctmgr add user john account=research_team
# Add user with multiple accounts
sacctmgr add user alice account=research_team,dev_team DefaultAccount=research_team
# View accounts
sacctmgr show account
# View users
sacctmgr show user

Thiết lập giới hạn tài nguyên

Giới hạn ở cấp account

Kiểm soát tài nguyên cho toàn bộ nhóm:

Terminal window
# Limit CPU minutes (prevents monopolizing cluster)
sacctmgr modify account research_team set GrpCPUMins=100000
# Limit memory (in MB)
sacctmgr modify account research_team set GrpMem=500000
# Limit concurrent jobs
sacctmgr modify account research_team set GrpJobs=50
# Limit concurrent running jobs
sacctmgr modify account research_team set GrpJobsRun=20
# Limit number of nodes
sacctmgr modify account research_team set GrpNodes=10

Giới hạn ở cấp người dùng

Kiểm soát hành vi của từng người dùng:

Terminal window
# Limit jobs in queue
sacctmgr modify user john set MaxJobs=10
# Limit running jobs
sacctmgr modify user john set MaxJobsRun=5
# Limit wall time (in minutes)
sacctmgr modify user john set MaxWall=1440 # 24 hours
# Limit CPUs per job
sacctmgr modify user john set MaxCPUs=32
# View user limits
sacctmgr show user john withassoc format=user,account,maxjobs,maxsubmit,maxwall

Quality of Service (QoS)

QoS cho phép bạn tạo các tầng dịch vụ với mức ưu tiên khác nhau:

Terminal window
# Create QoS levels
sacctmgr add qos normal priority=100
sacctmgr add qos high priority=500 MaxWall=2-00:00:00 MaxJobs=5
sacctmgr add qos low priority=50
# Assign QoS to account
sacctmgr modify account research_team set qos=normal,high
# Users can specify QoS when submitting
sbatch --qos=high job_script.sh

Lập lịch fair-share

Đảm bảo phân phối tài nguyên một cách công bằng:

Terminal window
# Set fair-share values (higher = more priority)
sacctmgr modify account research_team set fairshare=100
sacctmgr modify account dev_team set fairshare=50
# View fair-share tree
sshare -a
# View detailed fair-share info
sshare -A research_team --all

Quản lý node

Kiểm tra trạng thái node

Terminal window
# View all nodes
sinfo
# Detailed node information
sinfo -Nel
# Show node states
sinfo -N -o "%N %T %C %m %e %f"
# View specific node details
scontrol show node worker-01

Các trạng thái node bạn sẽ gặp:

  • IDLE: Sẵn sàng nhận job
  • ALLOCATED: Đang chạy job
  • MIXED: Được cấp phát một phần
  • DRAIN: Không nhận job mới (đang drain)
  • DRAINED: Đã drain hoàn toàn
  • DOWN: Không phản hồi

Bảo trì node

Drain một node

Khi bạn cần thực hiện bảo trì:

Terminal window
# Drain node (won't accept new jobs, allows running jobs to finish)
scontrol update NodeName=worker-01 State=drain Reason="Hardware upgrade"
# Force drain (terminate running jobs)
scontrol update NodeName=worker-01 State=drain Reason="Emergency maintenance"
# Check drain reason
sinfo -R

Resume một node

Sau khi bảo trì:

Terminal window
# Resume node
scontrol update NodeName=worker-01 State=resume
# Verify it's back
sinfo -n worker-01

Buộc node chuyển sang trạng thái Down

Nếu một node hoạt động không đúng cách:

Terminal window
# Mark node as down
scontrol update NodeName=worker-01 State=down Reason="Hardware failure"
# When fixed, resume
scontrol update NodeName=worker-01 State=resume

Thêm node tính toán mới

  1. Cập nhật inventory của Ansible (inventories/hosts):
[slurm_worker]
worker-01 ansible_host=192.168.58.11
worker-02 ansible_host=192.168.58.12
worker-03 ansible_host=192.168.58.13 # NEW
  1. Chạy playbook Ansible:
Terminal window
ansible-playbook -i inventories/hosts cluster_slurm.yml --limit worker-03
  1. Cập nhật slurm.conf trên controller và tất cả các node (Ansible sẽ xử lý việc này)

  2. Khởi động lại slurmctld:

Terminal window
sudo systemctl restart slurmctld
  1. Xác minh node mới:
Terminal window
sinfo
scontrol show node worker-03

Giám sát và xử lý sự cố

Kiểm tra log của Slurm

Log là công cụ thiết yếu để chẩn đoán vấn đề:

Terminal window
# Controller logs
sudo tail -f /var/log/slurm/slurmctld.log
# Worker node logs (on compute nodes)
sudo tail -f /var/log/slurm/slurmd.log
# Database logs
sudo tail -f /var/log/slurm/slurmdbd.log
# Filter for errors
sudo grep "error" /var/log/slurm/*.log
# Filter for specific node
sudo grep "worker-01" /var/log/slurm/slurmctld.log
# Last 100 lines with context
sudo tail -100 /var/log/slurm/slurmctld.log

Các vấn đề thường gặp và cách giải quyết

Vấn đề: Node hiển thị trạng thái DOWN

Chẩn đoán:

Terminal window
sinfo
# OUTPUT: worker-01 down ...
scontrol show node worker-01
# Check "Reason" field

Giải pháp:

Terminal window
# 1. Check if slurmd is running
ssh worker-01 "sudo systemctl status slurmd"
# 2. Restart slurmd
ssh worker-01 "sudo systemctl restart slurmd"
# 3. Check network connectivity
ping worker-01
# 4. Check logs
ssh worker-01 "sudo tail -50 /var/log/slurm/slurmd.log"
# 5. Resume the node
scontrol update NodeName=worker-01 State=resume

Vấn đề: Job bị kẹt ở trạng thái Pending

Chẩn đoán:

Terminal window
squeue
# See jobs in PD (pending) state
# Check why job is pending
squeue --start -j JOB_ID
# View detailed job info
scontrol show job JOB_ID

Nguyên nhân thường gặp:

  • Resources: Không đủ tài nguyên khả dụng
  • Priority: Mức ưu tiên thấp hơn các job khác
  • Dependency: Đang chờ một job khác hoàn thành
  • QOSMaxJobsPerUser: Người dùng đang chạy quá nhiều job

Giải pháp:

Terminal window
# 1. Check available resources
sinfo -o "%P %a %l %D %N %C"
# 2. View job requirements
scontrol show job JOB_ID | grep -E "Partition|NumNodes|MinMemory"
# 3. Cancel job if needed
scancel JOB_ID
# 4. Modify pending job
scontrol update JobId=JOB_ID NumNodes=1

Vấn đề: Job thất bại ngay lập tức

Chẩn đoán:

Terminal window
# Check job status
sacct -j JOB_ID
# View job output files
cat slurm-JOB_ID.out
cat slurm-JOB_ID.err

Nguyên nhân thường gặp:

  • Lỗi script (kiểm tra dòng shebang)
  • Thiếu tệp thực thi
  • Vượt quá giới hạn tài nguyên
  • Vấn đề về quyền truy cập

Vấn đề: Cơ sở dữ liệu kế toán không hoạt động

Chẩn đoán:

Terminal window
# Check slurmdbd status
sudo systemctl status slurmdbd
# Test database connection
sudo mysql -u slurm -p slurm_acct_db -e "SHOW TABLES;"
# Check slurmdbd logs
sudo tail -50 /var/log/slurm/slurmdbd.log

Giải pháp:

Terminal window
# 1. Restart slurmdbd
sudo systemctl restart slurmdbd
# 2. Verify database credentials in slurmdbd.conf
sudo cat /etc/slurm-llnl/slurmdbd.conf
# 3. Check database permissions
sudo mysql -e "SHOW GRANTS FOR 'slurm'@'localhost';"
# 4. Restart slurmctld to reconnect
sudo systemctl restart slurmctld

Log hệ thống với rsyslog

Thiết lập Ansible của chúng ta cấu hình ghi log tập trung:

Terminal window
# On controller (rsyslog server)
sudo tail -f /var/log/syslog
# Filter by hostname
sudo grep "worker-01" /var/log/syslog
# Filter by service
sudo grep "slurmd" /var/log/syslog
# Check authentication logs
sudo tail -f /var/log/auth.log

Các thực hành bảo mật tốt nhất

Tăng cường bảo mật SSH

Hãy bảo vệ các node đăng nhập của bạn! Các cụm HPC là mục tiêu hấp dẫn đối với kẻ tấn công.

Để biết chi tiết về thiết lập bảo mật SSH, xem tài liệu SSH Remote Server của chúng tôi.

Các khuyến nghị chính:

  1. Tắt xác thực bằng mật khẩu:
/etc/ssh/sshd_config
PasswordAuthentication no
PubkeyAuthentication yes
  1. Triển khai 2FA với Google Authenticator hoặc tương tự

  2. Sử dụng cặp khóa SSH:

Terminal window
# Generate key on your machine
ssh-keygen -t ed25519 -C "your_email@example.com"
# Copy to cluster
ssh-copy-id user@controller-node
  1. Giới hạn quyền truy cập SSH:
/etc/ssh/sshd_config
AllowUsers alice bob charlie
AllowGroups cluster_users
# Or deny specific users
DenyUsers baduser
  1. Thay đổi cổng mặc định (bảo mật bằng cách che giấu):
/etc/ssh/sshd_config
Port 2222

Xác thực Munge

Munge cung cấp xác thực giữa các thành phần của Slurm:

Terminal window
# Verify munge is running
sudo systemctl status munge
# Test munge
munge -n | unmunge
# Generate new key (do this on controller, then distribute)
sudo /usr/sbin/create-munge-key
# Copy key to all nodes (Ansible does this automatically)
sudo scp /etc/munge/munge.key worker-01:/etc/munge/
# Restart munge on all nodes
sudo systemctl restart munge

Khóa munge phải giống hệt nhau trên tất cả các node và có quyền phù hợp (0400, thuộc sở hữu của munge).

Cấu hình tường lửa

Hạn chế quyền truy cập vào các cổng của Slurm:

Terminal window
# Allow Slurm ports only from cluster network
sudo ufw allow from 192.168.58.0/24 to any port 6817 # slurmctld
sudo ufw allow from 192.168.58.0/24 to any port 6818 # slurmd
sudo ufw allow from 192.168.58.0/24 to any port 6819 # slurmdbd
# Allow SSH from anywhere
sudo ufw allow 22/tcp
# Enable firewall
sudo ufw enable

Các thực hành tốt nhất cho lưu trữ chia sẻ

Tinh chỉnh hiệu năng NFS

Tối ưu NFS cho khối lượng công việc của bạn:

Terminal window
# /etc/fstab on compute nodes
controller-01:/home /home nfs4 rw,soft,rsize=262144,wsize=262144,timeo=14,intr 0 0

Các tham số được giải thích:

  • soft: Hết thời gian chờ sau khi thử lại (so với hard là chờ mãi mãi)
  • rsize/wsize: Kích thước bộ đệm đọc/ghi (lớn hơn = hiệu năng tốt hơn)
  • timeo: Giá trị timeout
  • intr: Cho phép ngắt

Bố trí lưu trữ

Cấu trúc thư mục được khuyến nghị:

Terminal window
/home/ # User home directories (SSD/NVMe)
├─ alice/
├─ bob/
└─ charlie/
/mnt/data/ # Large datasets (HDD or object storage)
├─ shared/ # Common datasets
├─ projects/ # Project-specific data
└─ scratch/ # Temporary data (auto-cleanup)
/opt/ # Shared software/modules
├─ anaconda/
├─ modules/
└─ apps/

Quota

Ngăn người dùng làm đầy dung lượng lưu trữ chia sẻ:

Terminal window
# Set user quotas
sudo setquota -u alice 50G 60G 0 0 /home
sudo setquota -u alice 500G 550G 0 0 /mnt/data
# Check quotas
quota -u alice
# View all quotas
sudo repquota -a

Tích hợp với các framework xử lý dữ liệu

Một trong những điểm mạnh nhất của Slurm là khả năng tích hợp với các framework tính toán hiện đại:

Apache Spark

Bạn có thể cài đặt Spark/PySpark trên môi trường conda-base (pixi). Để chạy một cụm Spark trên nền cụm Slurm, bạn có thể xem thiết lập trong repo này: https://github.com/vieomics/omicslab-kit/tree/main/spark-on-slurm

Gửi job Spark đến Slurm:

#!/bin/bash
#SBATCH --job-name=spark-job
#SBATCH --nodes=4
#SBATCH --ntasks-per-node=1
#SBATCH --cpus-per-task=8
#SBATCH --mem=32G
# Load Spark module
module load spark/3.5.0
# Run Spark application
spark-submit \
--master yarn \
--num-executors 4 \
--executor-cores 8 \
--executor-memory 28G \
my_spark_app.py

Ray (Distributed ML)

#!/bin/bash
#SBATCH --job-name=ray-job
#SBATCH --nodes=2
#SBATCH --ntasks-per-node=1
#SBATCH --cpus-per-task=16
#SBATCH --gpus-per-node=2
# Start Ray cluster
ray start --head --port=6379
srun --nodes=1 --ntasks=1 ray start --address=$HEAD_NODE:6379
# Run Ray application
python ray_train.py

Dask

from dask_jobqueue import SLURMCluster
from dask.distributed import Client
cluster = SLURMCluster(
cores=8,
memory="16GB",
processes=2,
walltime="02:00:00",
queue="compute"
)
cluster.scale(jobs=10) # Request 10 jobs
client = Client(cluster)
# Your Dask code here

Nextflow (tin sinh học)

nextflow.config
process {
executor = 'slurm'
queue = 'compute'
memory = '8 GB'
time = '2h'
}

Chạy với:

Terminal window
nextflow run nf-core/rnaseq -profile test,singularity --outdir test_rnaseq_profile_outdir

Các tác vụ bảo trì

Cập nhật định kỳ

Terminal window
# Update cluster via Ansible
ansible-playbook -i inventories/hosts cluster_slurm.yml --tags update
# Update specific nodes
ansible-playbook -i inventories/hosts cluster_slurm.yml --limit worker-01,worker-02

Sao lưu dữ liệu quan trọng

Terminal window
# Backup Slurm configuration
sudo cp /etc/slurm-llnl/slurm.conf /backup/slurm.conf.$(date +%Y%m%d)
# Backup accounting database
sudo mysqldump -u slurm -p slurm_acct_db > slurm_acct_backup_$(date +%Y%m%d).sql
# Backup user data (use rsync for efficiency)
sudo rsync -av /home/ /backup/home/

Giám sát dung lượng đĩa

Terminal window
# Check disk usage on all nodes
ansible all -i inventories/hosts -m shell -a "df -h"
# Check specific directory
ansible all -i inventories/hosts -m shell -a "du -sh /var/log/slurm"
# Find large files
find /home -type f -size +1G -exec ls -lh {} \;

Mẹo tối ưu hiệu năng

1. Tinh chỉnh tham số bộ lập lịch

/etc/slurm-llnl/slurm.conf
# Increase scheduling frequency
SchedulerTimeSlice=30
# Prioritize recent submitters less
PriorityWeightAge=1000
PriorityWeightFairshare=10000
# Enable backfill scheduling with larger window
SchedulerType=sched/backfill
bf_window=1440 # 24 hours

2. Tối ưu đóng gói job

Terminal window
# Use CR_CPU for CPU-bound jobs
SelectType=select/cons_tres
SelectTypeParameters=CR_CPU
# Or CR_Memory for memory-bound jobs
SelectTypeParameters=CR_Memory
# Or CR_Core for mixed workloads
SelectTypeParameters=CR_Core

3. Tạo nhiều partition

/etc/slurm-llnl/slurm.conf
# Fast partition for short jobs
PartitionName=quick Nodes=worker-[01-02] Default=NO MaxTime=01:00:00 State=UP Priority=100
# Standard partition
PartitionName=standard Nodes=worker-[01-04] Default=YES MaxTime=2-00:00:00 State=UP Priority=50
# Long partition for extended jobs
PartitionName=long Nodes=worker-[03-04] Default=NO MaxTime=7-00:00:00 State=UP Priority=25
# GPU partition
PartitionName=gpu Nodes=gpu-[01-02] Default=NO MaxTime=1-00:00:00 State=UP Priority=75

4. Bật job array cho xử lý theo lô

#!/bin/bash
#SBATCH --array=1-100%10 # 100 tasks, max 10 concurrent
# Process task based on array index
python process.py --input data_${SLURM_ARRAY_TASK_ID}.txt

Kết luận

Xin chúc mừng! Bạn đã có đủ kiến thức để xây dựng, triển khai và quản lý một cụm Slurm HPC production. Hãy cùng nhìn lại hành trình:

Phần 1: Nền tảng

  • Hiểu kiến trúc Slurm
  • Thiết lập một node đơn để học
  • Cấu hình cgroup quan trọng
  • Kiến thức cơ bản về kế toán job

Phần 2: Triển khai production

  • Tự động hóa với Ansible
  • Thiết lập cụm đa node
  • Giám sát với Grafana
  • Cảnh báo qua Slack

Phần 3: Quản trị (bài viết này)

  • Quản lý người dùng và tài nguyên
  • Bảo trì node và xử lý sự cố
  • Tăng cường bảo mật
  • Tối ưu hiệu năng
  • Tích hợp framework

Tài liệu tham khảo

1.Slurm Overview — Tài liệu chính thức cho trình quản lý khối lượng công việc Slurm

2.NVIDIA/deepops — Bộ công cụ triển khai cụm mã nguồn mở (giấy phép BSD-3-Clause)

3.elasticluster — Công cụ cấp phát cụm đàn hồi (giấy phép GPL-3.0)

4.Ansible- Ansible for IT Automation DevOps

5.GitHub Repository-Omicslab HPC- Script Ansible để thiết lập cụm Slurm HPC


Bài viết này khép lại series của Omicslab về xây dựng cụm Slurm HPC. Cảm ơn bạn đã theo dõi! Chúng tôi hy vọng hướng dẫn này giúp bạn xây dựng và quản lý cơ sở hạ tầng HPC hiệu quả.

Bài viết gần đây