Keep AI workloads running through failures
Link flaps, GPU errors and node failures no longer crash your workloads. Preemptions and node drains no longer mean starting over. Clockwork AI Fault Tolerance keeps jobs alive, or moves them to healthy resources automatically, so your GPUs spend more time training and less time restarting.
The state of AI training today: one failure stops the whole cluster, forces a restart, and throws away the work since the last checkpoint.
1 Revisiting Reliability in Large-Scale Machine Learning Research Clusters (Meta FAIR) - arxiv.org/abs/2410.21680
Clockwork AI Fault Tolerance stops failures
from impacting running workloads.
LinkPass: a link flaps or fails. Your job doesn’t.
Link flaps happen all the time. Today a single flap can crash a training job, forcing a checkpoint restart, or take down a multi-node inference replica, leaving the replica offline until it restarts. LinkPass catches these failures the moment they happen and reroutes traffic across the node’s other NICs, so collectives don’t time out and your jobs keep running.
Link flaps with and without LinkPass
TorchPass Live GPU Migration: a GPU dies, and training resumes on a replacement as if nothing happened.
When a GPU, HBM or whole node fails, TorchPass captures the affected worker’s state, moves it to a replacement, and resumes from the exact iteration, as though nothing happened. It also handles planned events: a drain, taint, preemption notice or scheduled maintenance triggers a planned migration from the worker, before anything breaks.
How a TorchPass migration works
The TorchPass Orchestrator tracks every job, and a Scheduler Plugin connects it to your cluster scheduler (Kubernetes, Kubeflow, Slurm, Slinky or Soperator). When a worker is tainted, drained or fails, TorchPass brings the job to a consistent point, requests a replacement from the scheduler, captures state (from the worker itself for a planned event, or from a healthy data-parallel replica for a failure), transfers it to the replacement over RDMA, and resumes training rapidly with no lost compute.
TorchPass Snapshots: save state in seconds, without touching a line of model code.
TorchPass migration handles failures and drains when a spare is available or can be preempted from another job. TorchPass snapshots cover everything else: the node you have to drain for maintenance, the spot instance about to be reclaimed, the GPU that is starting to look unhealthy, and the catastrophic failure that takes all workers out at once.
Flexible targets – write direct to storage (GDS) or to DRAM with async. drain.
Industry’s first multinode platform snapshot
Platform Snapshots are an industry first — they capture the full state of a multi-node training job in under 20 seconds without needing any model integration or code changes. Model Checkpoints are also supported in the same framework through lightweight model integration, for teams that want even faster checkpointing.
TorchPass migrations and snapshots run at the platform layer, or with model integration
Planned migration for drains, preemptions or upon failure prediction
Under 20 seconds per snapshot — about 1% overhead at 25-minute frequency.
Planned and unplanned migration. Resilience to catastrophic failures e.g. GPU, node, etc.
Under 5 seconds per checkpoint — about 1% overhead at 5-minute frequency.
For Neoclouds and Platform Teams
Your customers measure you on how many of the GPU-hours they pay for turn into training progress. Clockwork AI Fault Tolerance runs in the platform layer, so every tenant is protected without changing a line of their code.
Prevents link flaps and failures from impacting training and multi-node inference.
Requires zero model integration.
- Triggered, or periodic (~15 min)
- Recover from preemptions and catastrophic failures
- For drains, preemptions or predicted failures
- Training carries on without a restart
For Model Builders and Neocloud Tenants
You are paying for every GPU-hour whether it trains or restarts. With TorchPass integrated into your training loop, hard failures become migrations, planned events become non-events, and checkpoints stop stalling training.
Prevents link flaps and failures from impacting training and multi-node inference.
Requires zero model integration.
- Triggered, or periodic
(< 5 min) - Recover from preemptions and catastrophic failures
- For planned or unplanned events
- Training survives hard failures
See it in four minutes
Stop paying for restarts.
See LinkPass and TorchPass on your own cluster. We will walk you through a failure injection on a live training job and show you what happens next.