Engineering blog
Notes from building Kooper AI.
What we learn serving image, language and safety models on infrastructure we run: the numbers, the incidents and the decisions behind them.
· incident · TensorRT · NVIDIA drivers · GPU inference
A GPU host failed. The fix was a driver version.
How a host fault took down our image service, why our compiled TensorRT models only run on one NVIDIA driver branch, and what an hour of recovery taught us.