Inference on the GPUs You Already Bought
NVIDIA publishes its model catalogue as NIM microservices, and the same models run either on NVIDIA's hosted endpoints or on the organization's own GPUs. That second option is the reason this provider matters here. An agency that bought GPUs for video analytics already owns the hardware a language model needs, and AI Intelligence Hub can point at it instead of at a hosted API.
How it connects
The model behind an agent or a node is configuration. NVIDIA is selected per deployment or per agent, the same way OpenAI or Anthropic is, and nothing about a workflow changes when the provider does.
Two placements, and the choice is about where content goes rather than which model is better. A hosted NVIDIA endpoint is a call leaving the deployment. A self-hosted NIM runs on the customer's own GPUs, which means no content leaves the boundary at all, and that is the answer that survives a classified, criminal-justice or sovereign requirement rather than merely addressing it.
Self-hosting is not free of consequence, and an engagement should be honest about it. NIM needs GPUs sized for the model chosen, and a model that runs comfortably on the hardware bought for object detection is not necessarily the largest model available. What the deployment gets is decided by the GPU, so name the model and the card together rather than separately.
What you can do together
- Run generation on your own GPUs, so no prompt or document leaves the deployment, which is what an air-gapped installation requires rather than prefers.
- Use hardware already bought for video analytics for language work too, rather than provisioning a second estate.
- Change the model behind an agent as configuration, with no workflow rebuilt.
- Mix placements: a self-hosted model for restricted material and a hosted one for general work, in the same deployment.
A scenario
- SizingThe engagement names the model and the GPU together, since the hardware decides which model is realistic.
- SetupNIM is deployed on the existing GPU servers and configured as the provider in AI Intelligence Hub.
- In useAn investigator asks a question and the answer is generated inside the datacenter, with retrieval running under their own identity.
- An auditThe agency demonstrates that no prompt or document left the boundary, which is the question the accreditation actually asks.
- LaterA newer model is published; the provider setting changes and the workflows do not.
What stays where
NVIDIA supplies the models and the runtime
The catalogue, the NIM images and the licensing stay with NVIDIA.
Where inference happens is your choice
A hosted endpoint is an external call; self-hosted NIM keeps everything inside the deployment.
The GPU decides the model
Self-hosting is bounded by the hardware, so the model and the card are chosen together.
Content governance does not move
Retrieval runs under the asking user's identity with their access list applied, whichever provider generates the answer.
Products and solutions
Next step
See it on your own NVIDIA instance. We will show the connection made, the data moving and the output, then size it for your deployment.
Request a demonstration or write to sales@vidizmo.ai
sales@vidizmo.ai · vidizmo.ai/integrations/catalog/nvidia