NVIDIA Inference Microservices (NIM): package, deploy, and scale AI models in production. 2026 edition — model-free NIM, Rubin profiles (FP4), SMPTE ST 2110 microservices for broadcast, air-gapped deployment, and migration playbooks (GLM-5, Llama-3.1-70B, E5 v5 deprecations).
A tour of the available NIM families: text LLMs, embeddings, reranking, vision, ASR/TTS. When to choose which.