Introducing TorchPass Snapshots: Protect training progress. Free up GPU capacity. No training-code changes.

Launching YOCO (You Only Compute Once): Industry’s First Contractual Guarantee To End GPU Waste In AI Training

Replay the Webinar: Navigating Networking Transitions Shaping AI Infra Economics: Scaling Up, Out, and Across

Play SemiAnalysis-Clockwork Webinar: Comparing Fault Tolerance Frameworks & TCO Impact

AI that never stalls.
GPUs that never sit idle.

Clockwork keeps AI workloads running through failures and shows you exactly where your cluster is slow, unhealthy or misconfigured. 100% software. Runs anywhere.

AI Fault Tolerance
GPU errors, link flaps and node failures no longer take down your AI workloads.
AI Fabric Observability
Detect slow or failing jobs, rule the fabric in or out, and pinpoint the exact link, switch or NIC.
Press release
Clockwork.io Raises $31M as LinkedIn, Together AI and WhiteFiber Adopt Its Resilience Software to Stop Wasting GPU-Hours
LinkedIn prevents tens of thousands of GPU-hours of downtime monthly; new TorchPass innovations preserve AI workload progress without code changes and speed reinforcement learning.
SemiAnalysis
SemiAnalysis confirms Clockwork TorchPass is the only fault-tolerance framework that doesn’t cost you training performance
Their report’s TCO and Goodput calculators demonstrate TorchPass outperforms every competing fault-tolerance approach, recovering millions in wasted compute annually.
“Cluster fault tolerance used to be a training problem. It is now an inference problem too. In our ClusterMAX, TorchPass cuts training goodput loss from 14% to under 3% for a gold-rated neocloud. Reinforcement Learning (RL) ties the two together: inference replicas generate rollouts, the trainer learns from them, and the updated weights go back to the replicas. Clockwork.io keeps replicas serving through link flaps and network failures. Its extremely fast checkpoints accelerate weight transfer back into the rollout fleet, so neither direction stalls the run. One fault-tolerance layer under training, inference, and RL is where this has to be solved.”
Dylan Patel
Founder, CEO, and Chief Analyst, SemiAnalysis
“Our customers grade us on goodput, the share of their GPU-hours that actually move the model forward…. Clockwork.io’s TorchPass and LinkPass…. are designed to keep jobs moving through GPU faults and link failures, preserving progress….”
Pavneet Ahluwalia
Product Lead, Together AI
“At AI infrastructure scale, a single network issue should never sideline healthy GPUs or interrupt running workloads…. Clockwork.io prevents tens of thousands of GPU-hours of downtime per month across our fleet….”
Raghu Hiremagalur
SVP, CTO Infrastructure, LinkedIn
“…we are building the foundation for AI at planetary scale..Clockwork’s approach aligns perfectly with ours, and together we’re creating an AI infrastructure that is not only powerful and reliable, but ready to support the most demanding innovations of the future.”
David Power
CTO, NScale
“….Clockwork.io’s automated fleet audit validates every link and node at once, localizes faults in minutes…. We bring clusters up faster, and a customer’s first training run lands on a fabric validated end-to-end, not just powered on….”
Tom Sanfilippo
Chief Technology Officer, WhiteFiber
“At Broadcom, our focus has always been on delivering Ethernet-centric infrastructure that scales AI with both performance and efficiency. Clockwork’s software-driven AI fabric adds an essential layer of agility and observability that enhances the power of our silicon. With proactive fleet monitoring and seamless failover, Clockwork enables platforms such as our Tomahawk 6 and Jericho4 to realize their full potential in flexibility, uptime, and AI performance. Together, we’re driving open, adaptable fabrics that allow enterprises to build AI infrastructure that is resilient, high-performing, and future-ready.”
Ram Velaga
Senior Vice President and General Manager, Core Switching Group, Broadcom
“Our mission at DCAI is…to not only serve researchers, startups, and enterprises today, but also to build the sovereign foundations of tomorrow’s innovation… Gefion is a game-changing resource driving breakthroughs in quantum computing, drug discovery, advanced weather forecasting and beyond…Clockwork enables us to operate Gefion seamlessly and reliably…The result is a compute-efficient, fault-tolerant infrastructure that researchers and industries can trust — lowering costs, eliminating wasted GPU cycles, and helping us deliver a sovereign AI capability second to none.”
Dr. Nadia Carlsten
CEO, DCAI
The state of AI infrastructure today

The problem today: AI workloads fail constantly, and GPU-hours vaporize with every crash.

Thousands of GPUs run in lock-step, so a single bad link, GPU or switch can stall or crash the entire cluster, and recovery means restarting and recomputing everything since the last checkpoint. Compounding that, when a workload is impacted the fabric is the hardest layer to troubleshoot: gray failures pass health checks and dashboards stay green while GPU-hours are wasted.

MTTF 7.9 hrs
Mean time to failure on a 1,024-GPU cluster1
2.3–4.5 hrs lost per day
GPU time lost per day to failures and restarts
MTTD / MTTR 1–3 hrs
Mean time to detect and repair a fabric issue

1 Revisiting Reliability in Large-Scale Machine Learning Research Clusters (Meta FAIR) · arxiv.org/abs/2410.21680

AI Fault Tolerance

AI Fault Tolerance keeps AI workloads running through failures

Link flaps, GPU errors and node failures no longer crash your workloads. Preemptions and node drains no longer mean starting over. Clockwork keeps jobs alive, or moves them to healthy resources automatically, so your cluster never stops doing useful work.

LinkPass: a link flaps. Your job doesn’t.

LinkPass intercepts link failures the moment they happen and rebalances traffic across the node’s other NICs, so distributed training and multi-node inference are not impacted. When the link recovers, traffic fails back automatically. Delivered as an NCCL plugin: zero model changes, installs in under 30 minutes.

TorchPass Migration: a GPU dies. Training keeps going.

When a GPU or node fails, TorchPass just keeps the job running. It captures the state of the affected worker, live-migrates it to a replacement over RDMA and resumes from the exact iteration with no lost compute.

TorchPass Snapshots: Industry’s first multi-node platform snapshots. Zero code changes.

Multi-node Platform Snapshots capture the state of a distributed training job in under 20 seconds with no model integration or code changes. The same framework also supports asynchronous Model Checkpoints through lightweight code integration, for teams that want even faster checkpointing.

Deploy it your way

For Neoclouds and platform teams
Protection at the platform layer
Without changing a line of model code, every tenant gets proactive migration and automatic snapshots on drains and preemptions.
For model builders and Neocloud tenants
Protection at the model layer
Drop-in patches for TorchTitan, Megatron-LM and DeepSpeed adds recovery from sudden failures.
2× faster
Faster than checkpoint/restart under hourly injected failures
18.2% goodput increase
Goodput loss falls from 28.5% to 10.3% on a 4,096-GPU silver-tier cluster2
Zero lost compute
On failures, preemptions or drains: training resumes from the exact iteration

2 SemiAnalysis ClusterMax TCO and Goodput calculator, default values · Open the calculator

Failures that used to cost hours become events the job never notices.
AI Fabric Observability

FleetLens: where exactly is my fabric problem?

FleetLens detects slow or failing jobs and definitively rules whether the fabric is to blame. When it is, FleetLens pinpoints the exact link, switch or NIC at fault. It surfaces the hidden, workload-impacting issues other tools miss: congestion and micro-congestion, topology-specific degradation, gray failures and unexplained slowdowns.

Illustration: traffic flows across a leaf-spine GPU fabric, one uplink goes hot, and Clockwork names the exact link

The foundational technologies that make this possible

ClockSync
Synchronizes clocks with sub-µs accuracy across the fabric so timestamps are directly comparable.
Out-of-band Probemesh
Continuously measures directional one-way delays.
Path decomposition
Adds rail, leaf, spine, port and link context to each probe.
Triangulation
Finds the narrowest shared element explaining degraded paths.
Workload detections
Core workload detections correlated with underlying problem.
“Is it the network, the host, or our code?” Stop guessing!
Explainer videos

See it in four minutes

LinkPass
A link flaps mid-training and the job keeps running. See LinkPass reroute traffic in milliseconds.
TorchPass
Migrate a live job to spare GPUs and resume from the exact iteration, with no model code changes.
FleetAudit
Catch bad links, NICs and GPUs before a job starts, so every cluster is production-ready.

Stop wasting GPU cycles. Start scaling smarter.

See Clockwork on your own cluster. We’ll inject a failure into a live training job, show you what happens next, and walk you through the observability data behind it.