#

Checkpointing and Restarting Computational Workflows Workshop

NOTE: This is an in-person session on campus.

Checkpointing and Restarting Computational Workflows (in-person)

Description: In-person training on checkpointing and restarting computational workflows on FASRC clusters. This workshop will introduce general checkpointing strategies and demonstrate practical techniques on how to make your computational workflows resilient to unexpected node failures or consequences of underestimating your code's runtime. This will include examples for Python, C/C++, and Fortran applications as well as AI and machine-learning workflows using PyTorch, TensorFlow/Keras, and JAX. Enabling checkpointing in your code saves the program's progress at regular intervals thereby resuming the program from its most recent save instead of starting over.

Participants will learn how to make long-running jobs resilient to failures, Slurm time limits, preemption, and planned requeue events. Hands-on exercises will cover saving and restoring application state, writing requeue-aware Slurm jobs, handling termination signals, and creating robust checkpoints for machine-learning training.

Note: This is an in-person session. It will not be available via Zoom and will not be recorded.

Level: Intermediate/Advanced

Time: 11am - 3pm

Location: SEC LL2.221, Science and Engineering Complex, 150 Western Ave, Boston, MA 02134

Presenters: Plamen Krastev

Who can attend this workshop: Anyone with a FASRC cluster account. An active cluster account is required to participate in the hands-on exercises.

What you will learn:

  1. Why checkpointing is important for long-running and failure-prone computational workloads

  2. General checkpointing approaches, design patterns, and restart granularity

  3. How to create checkpoint-and-restart workflows in Python, C/C++, and Fortran

  4. How to write requeue-aware Slurm jobs on the Cannon cluster

  5. How to handle Slurm signals and exit gracefully before a job reaches its time limit

  6. How to checkpoint AI/ML workflows using PyTorch, TensorFlow/Keras, Lightning, and JAX

  7. How to save and restore model parameters, optimizer state, training progress, and random-number-generator state

  8. Strategies for distributed and large-scale ML checkpointing, including rank-0 and sharded checkpoints

  9. Best practices for reliable, portable, and storage-efficient checkpoints

Prerequisites:

  1. A FASRC cluster account. If you do not have an account, see
    Request a FAS Research Computing Account well in advance. See prerequisite 2 and 3.

  2. Previous experience submitting batch jobs on a FASRC cluster. 

  3. Basic familiarity with Slurm job scripts.

  4. Basic programming experience in Python, C/C++, or Fortran. Participants interested in the AI/ML examples should have basic familiarity with at least one supported machine-learning framework.

Registration: * Registration Link *

 
  • 00

    days

  • 00

    hours

  • 00

    minutes

  • 00

    seconds

Date

Oct 27 2026

Time

11:00 am - 3:00 pm
Category
QR Code