Skip to content
Tag

Everything tagged #microservices

1 article, newest first.

InferenceDec 10, 2024 · 2 min read

Deploying AI Models with NVIDIA NIM: Production LLMs as Microservices

Running self-hosted LLMs used to mean stitching vLLM, Triton, and CUDA drivers by hand. NVIDIA NIM packages TensorRT-LLM in standardized, production-ready OCI containers.

002 min read
Deploying AI Models with NVIDIA NIM: Production LLMs as Microservices