Integration brief · NVIDIA ← Back to the page   Print or save as PDF
Integration brief NVIDIAAI and LLM Providers

Inference on the GPUs You Already Bought

NVIDIA publishes its model catalogue as NIM microservices, and the same models run either on NVIDIA's hosted endpoints or on the organization's own GPUs. That second option is the reason this provider matters here. An agency that bought GPUs for video analytics already owns the hardware a language model needs, and AI Intelligence Hub can point at it instead of at a hosted API.

What you can do together

  • Run generation on your own GPUs, so no prompt or document leaves the deployment, which is what an air-gapped installation requires rather than prefers.
  • Use hardware already bought for video analytics for language work too, rather than provisioning a second estate.
  • Change the model behind an agent as configuration, with no workflow rebuilt.
  • Mix placements: a self-hosted model for restricted material and a hosted one for general work, in the same deployment.

How it connects

The model behind an agent or a node is configuration. NVIDIA is selected per deployment or per agent, the same way OpenAI or Anthropic is, and nothing about a workflow changes when the provider does.

Two placements, and the choice is about where content goes rather than which model is better. A hosted NVIDIA endpoint is a call leaving the deployment. A self-hosted NIM runs on the customer's own GPUs, which means no content leaves the boundary at all, and that is the answer that survives a classified, criminal-justice or sovereign requirement rather than merely addressing it.

Self-hosting is not free of consequence, and an engagement should be honest about it. NIM needs GPUs sized for the model chosen, and a model that runs comfortably on the hardware bought for object detection is not necessarily the largest model available. What the deployment gets is decided by the GPU, so name the model and the card together rather than separately.

VIDIZMO and NVIDIA · Integration briefPage 1 of 2
How it works NVIDIAAI and LLM Providers
Nexus your video, audio, images and documents, indexed under your access rules content the asking user may see AI Intelligence Hub agents, chat, workflows, retrieval with citations; model chosen per agent prompts and embeddings over API NVIDIA model endpoint, hosted by the provider or on your own servers answers Users chat, search, workflow output NVIDIA VIDIZMO

A scenario

  1. SizingThe engagement names the model and the GPU together, since the hardware decides which model is realistic.
  2. SetupNIM is deployed on the existing GPU servers and configured as the provider in AI Intelligence Hub.
  3. In useAn investigator asks a question and the answer is generated inside the datacenter, with retrieval running under their own identity.
  4. An auditThe agency demonstrates that no prompt or document left the boundary, which is the question the accreditation actually asks.
  5. LaterA newer model is published; the provider setting changes and the workflows do not.

What stays where

NVIDIA supplies the models and the runtime

The catalogue, the NIM images and the licensing stay with NVIDIA.

Where inference happens is your choice

A hosted endpoint is an external call; self-hosted NIM keeps everything inside the deployment.

The GPU decides the model

Self-hosting is bounded by the hardware, so the model and the card are chosen together.

Content governance does not move

Retrieval runs under the asking user's identity with their access list applied, whichever provider generates the answer.

Products and solutions

Next step

See it on your own NVIDIA instance.

We will show the connection made, the data moving and the output, then size it for your deployment.

Contact VIDIZMO

sales@vidizmo.ai

+1 571-969-2180

vidizmo.ai

Product names and logos are the property of their respective owners.Page 2 of 2