Introducing TorchPass Snapshots: Protect training progress. Free up GPU capacity. No training-code changes.

Launching YOCO (You Only Compute Once): Industry’s First Contractual Guarantee To End GPU Waste In AI Training

Replay the Webinar: Navigating Networking Transitions Shaping AI Infra Economics: Scaling Up, Out, and Across

Play SemiAnalysis-Clockwork Webinar: Comparing Fault Tolerance Frameworks & TCO Impact

Fault Tolerance

Keep AI workloads running through failures

Link flaps, GPU errors and node failures no longer crash your workloads. Preemptions and node drains no longer mean starting over. Clockwork AI Fault Tolerance keeps jobs alive, or moves them to healthy resources automatically, so your GPUs spend more time training and less time restarting.

20–50%
Higher GPU utilization
90%
Fewer checkpoint restarts
100%
Software
The problem today

The state of AI training today: one failure stops the whole cluster, forces a restart, and throws away the work since the last checkpoint.

7.9 hrs
Mean time to failure on a 1,024-GPU cluster1
2.3–4.5 hrs
GPU time lost per day to failures and restarts
~$307K / month
Cost of disruption on a 1,024-GPU cluster

1 Revisiting Reliability in Large-Scale Machine Learning Research Clusters (Meta FAIR) - arxiv.org/abs/2410.21680

The solution

Clockwork AI Fault Tolerance stops failures
from impacting running workloads.

LinkPass
Resilience to link flaps and failures
TorchPass
Resilience to GPU, node and software failures
Protection with zero code changes for neoclouds and platform teams. Or integrated into the training code by model builders for the fastest recovery from any failure.
LinkPass · Resilience to link flaps and failures

LinkPass: a link flaps or fails. Your job doesn’t.

Link flaps happen all the time. Today a single flap can crash a training job, forcing a checkpoint restart, or take down a multi-node inference replica, leaving the replica offline until it restarts. LinkPass catches these failures the moment they happen and reroutes traffic across the node’s other NICs, so collectives don’t time out and your jobs keep running.

Link flaps with and without LinkPass

Why LinkPass is easy to adopt

Zero model changes
Standard NCCL network plugin.
Installs in minutes
Plug-and-play install.
Supports training, RL and inference
Protects any multi-node NCCL workload.
TorchPass · Resilience to GPU, node and software failures

TorchPass Live GPU Migration: a GPU dies, and training resumes on a replacement as if nothing happened.

When a GPU, HBM or whole node fails, TorchPass captures the affected worker’s state, moves it to a replacement, and resumes from the exact iteration, as though nothing happened. It also handles planned events: a drain, taint, preemption notice or scheduled maintenance triggers a planned migration from the worker, before anything breaks.

TorchPass use cases

Planned
Health issues with predictable symptoms — drain, taint or cordon
  • Temperature exceeding thresholds
  • ECC memory errors
  • Xid recoverable issues
  • Power or fan issues
Tasks that require migration
  • Maintenance during long training jobs, e.g. security patches, firmware upgrades
  • Workload rebalancing
  • Removing straggler servers
Unplanned
Hard failures
  • GPU failure
  • Uncorrectable HBM errors
  • Node failure
  • Kernel crash

How a TorchPass migration works

The TorchPass Orchestrator tracks every job, and a Scheduler Plugin connects it to your cluster scheduler (Kubernetes, Kubeflow, Slurm, Slinky or Soperator). When a worker is tainted, drained or fails, TorchPass brings the job to a consistent point, requests a replacement from the scheduler, captures state (from the worker itself for a planned event, or from a healthy data-parallel replica for a failure), transfers it to the replacement over RDMA, and resumes training rapidly with no lost compute.

TorchPass architecture: the Scheduler Plugin and TorchPass Orchestrator in the control plane, connected to Worker 0 through Worker n in the data plane
TorchPass · Snapshots

TorchPass Snapshots: save state in seconds, without touching a line of model code.

TorchPass migration handles failures and drains when a spare is available or can be preempted from another job. TorchPass snapshots cover everything else: the node you have to drain for maintenance, the spot instance about to be reclaimed, the GPU that is starting to look unhealthy, and the catastrophic failure that takes all workers out at once.

Flexible targets – write direct to storage (GDS) or to DRAM with async. drain.

Industry’s first multinode platform snapshot

Platform Snapshots are an industry first — they capture the full state of a multi-node training job in under 20 seconds without needing any model integration or code changes. Model Checkpoints are also supported in the same framework through lightweight model integration, for teams that want even faster checkpointing.

1 · Preemption or drain
A preemption or drain triggers a snapshot, so training later resumes with nothing lost.
2 · Health signals
A suspect-node health check triggers a snapshot and training continues — a fallback if the node later fails.
3 · Periodic
Frequent snapshots mean a catastrophic failure costs minutes, not hours.
Deployment

TorchPass migrations and snapshots run at the platform layer, or with model integration

TorchPass on the platform
No model integration. Part of the platform service.
For planned events
Migration

Planned migration for drains, preemptions or upon failure prediction

Recovers from any catastrophic failure
Platform Snapshots

Under 20 seconds per snapshot — about 1% overhead at 25-minute frequency.

TorchPass integrated with the model
Requires tightly scoped code changes. Pre-integrated patches for TorchTitan, DeepSpeed & Megatron-LM.
For planned and unplanned failure events
Migration

Planned and unplanned migration. Resilience to catastrophic failures e.g. GPU, node, etc.

Recovers from any catastrophic failure
Model Checkpoints

Under 5 seconds per checkpoint — about 1% overhead at 5-minute frequency.

For Neoclouds and Platform Teams

Your customers measure you on how many of the GPU-hours they pay for turn into training progress. Clockwork AI Fault Tolerance runs in the platform layer, so every tenant is protected without changing a line of their code.

LinkPass

Prevents link flaps and failures from impacting training and multi-node inference.

Requires zero model integration.

TorchPass
Platform
No model integration. Part of the platform service.
Snapshots
  • Triggered, or periodic (~15 min)
  • Recover from preemptions and catastrophic failures
Migration
  • For drains, preemptions or predicted failures
  • Training carries on without a restart

Why it matters

Protection built into the platform.
Ship resilience as a native capability of your offering, a differentiator at the platform level.
Zero changes to customer model code.
Deploy at the infrastructure layer so tenants are protected without touching a training script.
Higher SLAs, higher GPU utilization.
Fewer GPU-hours lost to failures and restarts means higher utilization and stronger uptime guarantees.
Survive drains and preemptions with no lost progress. Minimal training rollback on catastrophic failure.

For Model Builders and Neocloud Tenants

You are paying for every GPU-hour whether it trains or restarts. With TorchPass integrated into your training loop, hard failures become migrations, planned events become non-events, and checkpoints stop stalling training.

LinkPass

Prevents link flaps and failures from impacting training and multi-node inference.

Requires zero model integration.

TorchPass
Model
Requires model integration.
Checkpoints
  • Triggered, or periodic
    (< 5 min)
  • Recover from preemptions and catastrophic failures
Migration
  • For planned or unplanned events
  • Training survives hard failures

Why it matters

Zero lost progress on any failure.
Handles hardware faults, preemptions, network or software errors without impact to workloads.
Asynchronous checkpoints, effortlessly.
Checkpoint frequently without stalling training or writing custom checkpoint frameworks.
Drop-in framework integration.
Ready-made patches for TorchTitan, Megatron-LM and DeepSpeed.
Survive drains, preemptions and catastrophic failures with no lost progress.
Explainer videos

See it in four minutes

LinkPass
A link flaps mid-training and the job keeps running. See LinkPass reroute traffic in milliseconds.
TorchPass
Migrate a live job to spare GPUs and resume from the exact iteration, with no model code changes.
FleetAudit
Catch bad links, NICs and GPUs before a job starts, so every cluster is production-ready.

Stop paying for restarts.

See LinkPass and TorchPass on your own cluster. We will walk you through a failure injection on a live training job and show you what happens next.