VIDIZMO integration brief · NVIDIA · vidizmo.ai/integrations/catalog/nvidia

Integrations / NVIDIA

NVIDIA

Run models on your own GPUs through NVIDIA NIM. Visit website

NVIDIA publishes model endpoints as NIM microservices, and the same catalogue is available hosted or self-hosted on the organisation's own GPUs.

It is the provider that matters where inference has to stay on the customer's hardware but an open-weight runtime is not enough, which is the usual position for an on-premises or air-gapped public safety deployment with GPUs already bought for video analytics.

Inference on the GPUs You Already Bought

Open as a two-page brief

NVIDIA publishes its model catalogue as NIM microservices, and the same models run either on NVIDIA's hosted endpoints or on the organization's own GPUs. That second option is the reason this provider matters here. An agency that bought GPUs for video analytics already owns the hardware a language model needs, and AI Intelligence Hub can point at it instead of at a hosted API.

Nexus your video, audio, images and documents, indexed under your access rules content the asking user may see AI Intelligence Hub agents, chat, workflows, retrieval with citations; model chosen per agent prompts and embeddings over API NVIDIA model endpoint, hosted by the provider or on your own servers answers Users chat, search, workflow output NVIDIA VIDIZMO

How it connects

The model behind an agent or a node is configuration. NVIDIA is selected per deployment or per agent, the same way OpenAI or Anthropic is, and nothing about a workflow changes when the provider does.

Two placements, and the choice is about where content goes rather than which model is better. A hosted NVIDIA endpoint is a call leaving the deployment. A self-hosted NIM runs on the customer's own GPUs, which means no content leaves the boundary at all, and that is the answer that survives a classified, criminal-justice or sovereign requirement rather than merely addressing it.

Self-hosting is not free of consequence, and an engagement should be honest about it. NIM needs GPUs sized for the model chosen, and a model that runs comfortably on the hardware bought for object detection is not necessarily the largest model available. What the deployment gets is decided by the GPU, so name the model and the card together rather than separately.

What you can do together

  • Run generation on your own GPUs, so no prompt or document leaves the deployment, which is what an air-gapped installation requires rather than prefers.
  • Use hardware already bought for video analytics for language work too, rather than provisioning a second estate.
  • Change the model behind an agent as configuration, with no workflow rebuilt.
  • Mix placements: a self-hosted model for restricted material and a hosted one for general work, in the same deployment.

A scenario

  1. SizingThe engagement names the model and the GPU together, since the hardware decides which model is realistic.
  2. SetupNIM is deployed on the existing GPU servers and configured as the provider in AI Intelligence Hub.
  3. In useAn investigator asks a question and the answer is generated inside the datacenter, with retrieval running under their own identity.
  4. An auditThe agency demonstrates that no prompt or document left the boundary, which is the question the accreditation actually asks.
  5. LaterA newer model is published; the provider setting changes and the workflows do not.

What stays where

NVIDIA supplies the models and the runtime

The catalogue, the NIM images and the licensing stay with NVIDIA.

Where inference happens is your choice

A hosted endpoint is an external call; self-hosted NIM keeps everything inside the deployment.

The GPU decides the model

Self-hosting is bounded by the hardware, so the model and the card are chosen together.

Content governance does not move

Retrieval runs under the asking user's identity with their access list applied, whichever provider generates the answer.

Products and solutions

Next step

See it on your own NVIDIA instance. We will show the connection made, the data moving and the output, then size it for your deployment.

Request a demonstration  or write to sales@vidizmo.ai

sales@vidizmo.ai  ·  vidizmo.ai/integrations/catalog/nvidia

Not what you need?

We build it

Send us your API documentation and we build, test and maintain the connector, at no development cost to you.

Request this integration

Bring an MCP server

If the system publishes a Model Context Protocol server, AI Intelligence Hub connects to it as a client with configuration alone.

Define a REST endpoint

Describe your endpoint and it becomes a node in an agent workflow, without waiting on our roadmap.