Engineering blog

A GPU host failed. The fix was a driver version.

How a host fault took down our image service, why our compiled TensorRT models only run on one NVIDIA driver branch, and what an hour of recovery taught us.

·

On 29 September, the service that draws every illustration in Hello Kooper stopped answering for a little over an hour. Nothing in our code changed, and the fix did not involve our code either. This is what happened, and what we are changing.

How the image service runs

Kooper AI’s image service runs SDXL, compiled ahead of time into TensorRT engines (FP8, with DeepCache) so a picture takes about a second on a single NVIDIA L40S. The compiled engines live in object storage. When a GPU pod starts, it downloads the bundle, loads every engine, and only then reports ready.

Apps never call a pod directly. They call a stable address on a small Cloudflare Worker, which forwards each request to whichever pod is current. That makes pods disposable: a deploy creates a new pod, waits for it to report ready, points the router at it, and only then removes the old one. If the new pod never becomes ready, it is thrown away and nothing else changes.

What broke

At about 15:40 UTC the container on our production pod restarted. The provider later flagged the host with a critical error, inside a maintenance window for software upgrades.

The restart itself was routine. The pod downloaded its models again, and the download checked out. Then the engine refused to start:

Fatal error: SafetyChecker: failed to load TRT plan

Ray Serve tried three times and gave up. From the outside, the router first returned 404s, because the app had no routes registered, and then 503s with “failed to deploy”. The model files were unchanged, and they had loaded fine on that same machine for three weeks.

Why does a TensorRT engine depend on the NVIDIA driver version?

A TensorRT engine is not a portable model file. It is compiled for a specific GPU, and in practice it is sensitive to the driver branch it was built on. We measured this in August on two identical L40S machines: engines built on driver 580 (CUDA 13.0) loaded cleanly on 580 hosts, but under driver 570 (CUDA 12.8) they tripped TensorRT’s cross-device check, which our engine treats as fatal.

Since then, our deploys ask the GPU provider only for hosts on the 580 branch. The engine files and the driver branch are one artifact generation, and they move together.

The error on the failing host is exactly what an engine sees on the wrong driver. We cannot see inside the provider’s maintenance, but it fits.

Why couldn’t we just start a new pod?

The fix was the same one a normal deploy performs: create a new pod on a healthy host. The first attempt failed after three tries. There was no L40S on a CUDA 13 host in any data center we allow.

The provider’s capacity data made the gap plain. L40S machines were available, but only on CUDA 12.8 hosts, and our pin correctly excluded them, because the new pod would have failed at the same step. The deploy did the right thing: it created nothing and left everything as it was. It just could not help.

Capacity moves minute to minute. A second attempt a little later found a CUDA 13 host. The new pod downloaded its bundle, passed its readiness check, and the router switched to it. Our deploy then checked the full path through the router before removing the broken pod.

What we are changing

Treat the driver branch as part of the build. The engine bundle, its storage prefix and the driver pin already change together. We will add the driver branch to the bundle’s name and to its health report, so a mismatch is obvious in one line of logs.

Build for more than one driver branch. Pinning to one branch keeps us correct, but it also shrinks the pool of GPUs we can use. Compiling a second bundle on the 570 branch lets a pod pick the bundle that matches its host, so more machines qualify.

Make the standby reachable. We run the same image service on a second provider as a standby, and it passed its checks during this incident. But our app was not configured to reach it, so it could not take traffic. A standby only counts if the client can switch to it.

Detect it before a person does. We noticed this because a request failed. The router already exposes a readiness check; it will page us when it fails.

Takeaways

  • Compiled inference engines are build artifacts. Version them with everything they depend on, including the driver.
  • Deploy in a fail-safe order: create, verify, switch, then remove. It turned a bad hour into a boring one.
  • GPU capacity is a dependency. Every constraint you add, whether region, GPU type or driver, narrows the supply you can recover onto.

← All posts · Markdown · RSS