Inference API with batched GPU serving, the answer
Check out shared "Inference API with batched GPU serving, the answer" scene on
Excalidraw+