Practical: High-Performance Computing System Administration

High-Performance Computing System Administration is essential for managing HPC resources not only as a user but as a cluster administrator. As part of this practical course, you will take part in a hands-on one-week block course, which will introduce the basics of Linux and using HPC resources and then go into depth on HPC system administration. At the end of the block course you will choose a topic in terms of a tool related to HPC system administration, evaluate that tool and hand-in a report at the end of the semester. For this a supervisor will be assigned to you, who is an expert on the assigned tool and is able to guide you.

Contact Julian Kunkel, Jonathan Decker
Location Virtual Main Room Support Room
Time 19.10.26-23.10.26 5-day block course
Language English
Module M.Inf.1831: High-Performance Computing System Administration
SWS 4
Credits 6 (+ 3 with M.Inf.1834)
Contact time up to 84 hours (63 full hours), depending on the course
Independent study up to 186 hours

Please note that we plan to record sessions (lectures and seminar talks) with the intent of providing the recordings via BBB to other students but also to publish and link the recordings on YouTube for future terms. If you appear in any of the recordings via voice, camera or screen share, we need your consent to publish the recordings. See also this Slide.

  • Understanding of Linux basics and having used Linux before and being able to operate a Bash shell is beneficial
  • Discuss theoretic facts related to networking, compute and storage resources
  • Integrate cluster hardware consisting of multiple compute and storage nodes into a “supercomputer“
  • Configure system services that allow the efficient management of the cluster hardware and software including network services such as DHCP, DNS, NFS, IPMI, SSHD.
  • Install software and provide it to multiple users
  • Compile end-user applications and execute it on multiple nodes
  • Analyze system and application performance using benchmarks and tools
  • Formulate security policies and good practice for administrators
  • Apply tools for hardening the system such as firewalls and intrusion detection
  • Describe and document the system configuration
  • Benchmarking Edge Devices for HPC Data Ingestion
  • Benchmarking indexing tools (plocate, mlocate, librer, ratarmount,Katalog,vvv)
  • Benchmarking Streaming Systems for HPC Data Ingestion
  • Benchmarking Time-Series Models for HPC Failure Prediction
  • Building an End-to-End Edge-to-HPC Streaming Pipeline
  • Centralized package management using EESSI and E4S
  • Coding Agents from First Principles
  • Confidential Computing
  • Deep-Dive Workload Analysis (DLIO)
  • Deployment of Lightweight AI Models on Edge Devices
  • Ethical AI Scheduler for HPC Clusters
  • Feature Engineering for HPC Log-Based Anomaly Detection
  • GPU aware Fault Injection: Storage I/O Faults and GPU Memory interaction
  • HPC I/O Workload Characterization & Tool Evaluation
  • Implementation and Evaluation of DAG Scheduling Heuristics for HPC Workflows
  • Interoperability Between Groupware and Collaboration Platforms
  • Kubernetes Multi-Cluster Management
  • Learning-Based Resource Selection Across CPU, GPU, and Emerging HPC Accelerators
  • Managing Edge Devices Connected to HPC Systems
  • Monitoring and Observability of Streaming Systems in HPC
  • Multi-Metric Correlation Analysis for Failure Precursors
  • Neuromorphic Computing with SpiNNaker
  • Performance Evaluation of Jacobi Iterative Solution
  • Performance Evaluation of Physics Mini-Apps
  • Performance Evaluation of Shallow Water Code
  • Rootless Kubernetes
  • Scalable databases with e.g., Elasticsearch, Postgres
  • Secure Data Ingestion Pipelines for HPC Environments
  • Security Mechanism Overhead Micro-benchmarking
  • Service Discovery and Traffic Management in Cloud Applications
  • Time-Series Deep Learning for GPU Failure Prediction
  • Using Large Language Models to Assist Workflow Scheduling Decisions
  • Using Large Language Models to Generate Knowledge Base
  • Visual Analytics for Workflow Scheduling Experiments in Heterogeneous HPC Systems
  • What's new in the Kubernetes ecosystem

This part is attended by BSc/MSc students and GWDG academy participants

Note: There are only breaks for lecture slots in the schedule. You can take a break during exercises as necessary. Preparation sheets: Preparation

Monday 19.10.2026

  • 09:00 - 10:00 Welcome, Organization of the block course and VM setup – Jonathan Decker Slides Exercise
    • Agenda of the week
    • Forming support groups
    • Format of the “group work”
    • Exercise (10 min): Introduce yourself in the “learning groups”
    • Tutorial (10 min): Demo; setting up cloud resources from a fresh account
    • Exercise (20 min): Is your cloud setup working?
    • Plenary (10 min): Discussion of the format, Q&A
  • 10:00 - 12:00 Cluster Management with Warewulf Part 1 – Timon Vogt Slides Exercise
    • “How to boot a thousand nodes”
    • Lecture (20 min): Motivation, components of cluster management (DNS, DHCP, PXE-Boot process, images, resource management, monitoring, hardware-components)
    • Management Demo
    • Exercise (30 min): Describing the responsibility of Warewulf components and the boot process
    • Lecture: Technical details and administration of dnsmasq, DHCP, and investigating logfiles
    • Exercise 1
  • 12:00 - 13:00 Lunch Break
  • 13:00 - 15:00 Cluster Management with Warewulf Part 2 – Timon Vogt Slides Exercise
    • Lecture: Warewulf configuration
    • Demo: Image creation and deployment
    • Exercise 2
  • 15:00 - 17:00 User Management with Warewulf – Freja Nordsiek Slides Exercise

Tuesday 20.10.2026

  • 09:00 - 10:00 Best practices for administrators and documentation – Stefanie Mühlhausen Slides Exercise
    • Lecture (20 min): processes and management, documentation, frameworks: ITIL, PRINCE2
    • Exercise (20 min): Discussion of the best-practices, searching for related work, critical discussion of your own experience with the setup of Warewulf and Slurm
    • Plenary discussion (20 min)
  • 10:00 - 11:00 Recap of Monday – Timon Vogt
    • Recap of cluster and user management with Warewulf
    • Making sure that everyone is ready with the setup
  • 11:00 - 12:00 Slurm administration Part 1 – Timon Vogt Slides Exercise Config
    • Slurm installation, basic configuration, testing
    • Lecture: introduction to Slurm
  • 12:00 - 13:00 Lunch Break
  • 13:00 - 15:00 Slurm administration Part 2 – Timon Vogt
    • Tutorial server installation, basic configuration and testing (flexible break)
    • Exercise: adjustments of the configuration, integration of the cluster nodes, testing
  • 15:00 - 17:00 Monitoring in HPC – Marcus Merz Slides Tutorial
    • Lecture(15 min): Monitoring introduction and software stacks
    • Lecture(5 min): InfluxDB
    • Exercise(20 min): Installing InfluxDB
    • Lecture(5 min): Telegraf
    • Exercise(20 min): Installing Telegraf
    • Lecture(5 min): Grafana
    • Exercise(35 min): Installing Grafana and setting up a dashboard for an example application (Slurm)
    • Plenary discussion (15 min)

Wednesday 21.10.2026

  • 09:00 - 10:00 Provisioning of an Environment for Parallel Computing – Slides Exercise
    • Lecture(15 min): Providing a joint software environment with environment modules and Spack
    • Exercise(45 min): Installing MPI and Gromacs and providing module descriptions (other group members to test)
    • Plenary Discussion(15 min)
  • 10:00 - 11:00 Firewalls – Freja Nordsiek – Slides Exercise NFT Ruleset
    • Lecture(15 min): Introduction to firewalls
    • Exercise(35 min): Exploring firewall rules, port scanning with nmap, internet access for the nodes using NAT
    • Plenary Discussion(10 min)
  • 11:00 - 12:00 Network File System Setup – Patrick Höhn Slides Exercise
    • Lecture(15 min): NFS Introduction
    • Exercise(30 min): Setup of a basic NFS Server and client
    • Plenary Discussion(15 min)
  • 12:00 - 13:00 Lunch Break
  • 13:00 - 14:00 Intelligent Platform Management Interface (IPMI) – Nils Kanning Slides
    • Lecture(15 min): IPMI introduction
    • Exercises(40 min)
    • Plenary discussion (5 min)
  • 14:00 - 14:30 ClusterShell – Slides Exercise
    • Lecture (10 min): Introduction
    • Exercise (15 min): Installation and testing
    • Plenary Discussion(5 min)
  • 14:30 - 15:00 Break
  • 15:00 - 17:00 Documentation Writing (Docathon) – Kevin Lüdemann

Thursday 22.10.2026

  • 09:00 - 11:00 Benchmarking – Aasish Kumar Sharma Slides Tutorial Exercise
    • Lecture(35 min): Benchmarking
    • Exercise(15 min): Real system benchmarking on your VMs
    • Plenary Discussion(10 min)
  • 11:00 - 12:00 Performance Estimation – Zoya Masih Slides Exercise
    • Lecture(20 min): Hardware characteristics and performance estimates in distributed systems
    • Exercise(35 min): Theoretic performance assessment
    • Plenary Discussion(15 min)
  • 12:00 - 13:00 Lunch Break
  • 13:00 - 14:00 Security and security policies – Mojtaba Akbari Slides Exercise Exercise Solution
    • Lecture(30 min): Security introduction + Demo
      • Discussing an existing service and its security implications
    • Exercise(15 min): Theoretical investigation of an existing service (the one from before)
    • Plenary discussion (15 min)
  • 14:00 - 17:00 Use each other's cluster and test the user documentation – Kevin Lüdemann
    • Activate user accounts for someone from your study group
    • Go into someone else's system, explore it, and write a ticket to the admin of the system to fix problems and install new software
  • 17:00 - 17:30 Student Project Assignment and organisational information for students – Lauritz Rasbach slides

Friday 23.10.2026

RzGö live hardware demonstration and Hands-on. If you are a remote participant, we request that you revisit the previous material and prepare questions for Q&A sessions.

On-site is limited to up to 20 participants. For the hands-on sessions and the data center tour the class is split into two halves.

  • All Participants
    • 08:45 Meet at GWDG Burckhardtweg 4, 37077 Göttingen in the lobby - (Bus stop Bruckhardtweg)
    • 09:00 - 10:00 Network interconnects – Sebastian Krey, Freja Nordsiek Slides Exercise Hardware
      • Lecture(20 min): HPC Interconnects, Fabric Manager, RDMA, VLAN, LATP
      • Exercise(20 min): Cable planing
  • First half of the class
    • 10:00 - 12:00 Introduction to our onsite hardware and Hands-on Hardware Exercises – Sebastian Krey, Freja Nordsiek
      Smartboard Group 1
    • 12:00 - 13:00 Lunch Break
    • 13:00 - 14:30 Hands-on Hardware Exercises – Sebastian Krey, Freja Nordsiek
    • 14:30 - 16:30 Tour in the data center
  • Second half of the class
    • 10:00 - 12:00 Tour in the data center
    • 12:00 - 13:00 Lunch Break
    • 13:00 - 14:30 Tour in the data center
    • 14:30 - 18:00 Introduction to our onsite hardware and Hands-on Hardware Exercises – Sebastian Krey, Freja Nordsiek
      Smartboard Group 2

Hands-on Exercises include:

  • Setting up hardware
  • Plugin a small cluster
  • BIOS settings
  • Installation of Warewulf
  • Mounting of Infiniband cards
  • Configuration of Infiniband
  • RMDI performance test
  • 2026-11-01 - Send your requested topic to us until this day
  • 2026-11-08 - We assign a supervisor per student until this day
    • Contact your supervisor
    • Work on your topic
    • Write your reports
    • Get feedback from supervisor
  • 2027-03-31 - Submit final report as PDF per email to jonathan.decker@uni-goettingen.de

The exam is conducted through a report. The report should cover the evaluation of the assigned tool. The report should describe:

  • What the tool is, what it is used for
  • How the tool was set up
  • How you evaluated it
  • The results of your evaluation
  • Discussion of problems and potential of the tool
  • Conclusion

The report should not exceed 15 pages (only counting raw text in the main part, the full report including cover pages and appendix may be longer). It is not sufficient to repeat the documentation of the tool in your own words.

We recommend to use the LaTeX templates provided by us here: https://hps.vi4io.org/teaching/ressources/start#templates

In order to be allowed to take the examination, you have to show that you have taken the majority of the sessions of the block course. To prove this, please send 1-2 pages of notes on the course to us. These can be your personal notes from the course you took during the sessions and does not need to be a formatted document and is just to prove that you took the course. These do NOT need to be complete solutions to the exercises, a few sentences on your takeaways per section are enough.

If you joined the course late or had to miss out on some of the sessions, you can find the recordings on BBB and the materials on this web page. The exercises can be completed on a personal VM.

Student Supervisor Topic Submissions
Your Name Your Supervisor Your Topic Report
  • teaching/autumn_term_2026/hpcsa.txt
  • Last modified: 2026-10-09 17:29
  • by Jonathan Decker