BEGIN:VCALENDAR
VERSION:2.0
METHOD:PUBLISH
CALSCALE:GREGORIAN
PRODID:-//WordPress - MECv6.2.5//EN
X-ORIGINAL-URL:https://www.rc.fas.harvard.edu/
X-WR-CALNAME:FAS Research Computing
X-WR-CALDESC:.HARVARD UNIVERSITY | FACULTY OF ARTS &amp; SCIENCES
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-PUBLISHED-TTL:PT1H
X-MS-OLK-FORCEINSPECTOROPEN:TRUE
BEGIN:VEVENT
CLASS:PUBLIC
UID:MEC-5c0321b6b78eecdfcf72e6a44222fef9@rc.fas.harvard.edu
DTSTART:20261027T150000Z
DTEND:20261027T190000Z
DTSTAMP:20260810T142300Z
RDATE;VALUE=PERIOD:20220607T170000Z/20220607T183000Z
CREATED:20260810
LAST-MODIFIED:20261008
PRIORITY:5
TRANSP:OPAQUE
SUMMARY:Checkpointing and Restarting Computational Workflows Workshop
DESCRIPTION:NOTE: This is an in-person session on campus.\nCheckpointing and Restarting Computational Workflows (in-person)\nDescription: In-person training on checkpointing and restarting computational workflows on FASRC clusters. This workshop will introduce general checkpointing strategies and demonstrate practical techniques on how to make your computational workflows resilient to unexpected node failures or consequences of underestimating your code's runtime. This will include examples for Python, C/C++, and Fortran applications as well as AI and machine-learning workflows using PyTorch, TensorFlow/Keras, and JAX. Enabling checkpointing in your code saves the program's progress at regular intervals thereby resuming the program from its most recent save instead of starting over.\nParticipants will learn how to make long-running jobs resilient to failures, Slurm time limits, preemption, and planned requeue events. Hands-on exercises will cover saving and restoring application state, writing requeue-aware Slurm jobs, handling termination signals, and creating robust checkpoints for machine-learning training.\nNote: This is an in-person session. It will not be available via Zoom and will not be recorded.\nLevel: Intermediate/Advanced\nTime: 11am - 3pm\nLocation: SEC LL2.221, Science and Engineering Complex, 150 Western Ave, Boston, MA 02134\nPresenters: Plamen Krastev\nWho can attend this workshop: Anyone with a FASRC cluster account. An active cluster account is required to participate in the hands-on exercises.\nWhat you will learn:\n\n\nWhy checkpointing is important for long-running and failure-prone computational workloads\n\n\nGeneral checkpointing approaches, design patterns, and restart granularity\n\n\nHow to create checkpoint-and-restart workflows in Python, C/C++, and Fortran\n\n\nHow to write requeue-aware Slurm jobs on the Cannon cluster\n\n\nHow to handle Slurm signals and exit gracefully before a job reaches its time limit\n\n\nHow to checkpoint AI/ML workflows using PyTorch, TensorFlow/Keras, Lightning, and JAX\n\n\nHow to save and restore model parameters, optimizer state, training progress, and random-number-generator state\n\n\nStrategies for distributed and large-scale ML checkpointing, including rank-0 and sharded checkpoints\n\n\nBest practices for reliable, portable, and storage-efficient checkpoints\n\n\nPrerequisites:\n\n\nA FASRC cluster account. If you do not have an account, see\nRequest a FAS Research Computing Account well in advance. See prerequisite 2 and 3.\n\n\nPrevious experience submitting batch jobs on a FASRC cluster. \n\n\nBasic familiarity with Slurm job scripts.\n\n\nBasic programming experience in Python, C/C++, or Fortran. Participants interested in the AI/ML examples should have basic familiarity with at least one supported machine-learning framework.\n\n\nRegistration: * Registration Link ( https://harvard.az1.qualtrics.com/jfe/form/SV_3reOVFeSBAY6yV0 ) *\n \n
URL:https://www.rc.fas.harvard.edu/events/checkpointing-workshop/
CATEGORIES:Training
END:VEVENT
END:VCALENDAR
