VIDIZMO integration brief · Hugging Face · vidizmo.ai/integrations/catalog/hugging-face

Integrations / Hugging Face

Hugging Face

Use embedding models from Hugging Face. Visit website

Hugging Face hosts the largest public catalogue of open models, including the sentence embedding models retrieval depends on.

Embedding models are sourced from it, including for deployments that download once and then run disconnected.

Open Embedding Models, Downloaded Once, Run Disconnected

Open as a two-page brief

Hugging Face hosts a public catalog of open models, including the sentence embedding models that retrieval by meaning depends on. Connected to AI Intelligence Hub, a Hugging Face embedding model turns the text of your transcripts, documents and visual descriptions into vectors on your own hardware. The model is fetched once and then runs inside your deployment, which is what lets an air-gapped installation index by meaning with nothing leaving the boundary. Content stays in Nexus.

Nexus your video, audio, images and documents, indexed under your access rules content the asking user may see AI Intelligence Hub agents, chat, workflows, retrieval with citations; model chosen per agent prompts and embeddings over API Hugging Face model endpoint, hosted by the provider or on your own servers answers Users chat, search, workflow output Hugging Face VIDIZMO

How it connects

The embedding provider is configuration. Set to Hugging Face, the platform fetches the named model from the Hugging Face catalog over its API and runs it inside the deployment's own embedding service. For a disconnected site, the model is downloaded once and then runs with no further connection.

At indexing time, the text of each stream goes to that local model and vectors return to the index: transcripts, OCR text, visual descriptions and document text. Nothing goes to Hugging Face after the download. The language model that writes answers is configured separately, and for a self-hosted deployment that is a vLLM or Ollama server on your own hardware.

Retrieval runs under the asking user's identity with their access list applied before the vector search, whichever embedding model is in use.

What you can do together

  • Search by meaning across AI Intelligence Hub with an open embedding model of your choosing, running on your own servers.
  • Index by meaning in an air-gapped deployment, where a hosted embedding endpoint is unreachable by definition.
  • Pair the model with vLLM or Ollama for answers, so retrieval and generation both stay inside the boundary.
  • Keep every answer inside what the asking user may open, cited to the moment or the page.

A scenario

  1. SetupIT downloads the chosen embedding model from Hugging Face once, installs it with the platform, and sets it as the embedding provider in AI Intelligence Hub. A vLLM server on the same rack serves the language model.
  2. IntakeSeized-device audio, surveillance footage, interview recordings and scanned documents for case 23-0412 are transcribed, described and read by the platform's local Whisper, vision and OCR models. The Hugging Face model then embeds the text. Everything stays in Nexus.
  3. 11:05An analyst asks, "Which recordings in 23-0412 mention a storage unit?" The agent retrieves the passages from the recordings she is assigned to and sends them to vLLM with the question.
  4. Seconds laterThe answer cites three moments across two recordings. She opens a card and the audio plays from that second.
  5. 11:20She asks the same of the scanned documents, and the answer cites each page.
  6. ThroughoutNothing had a route out. Licensing verification is disabled for the disconnected deployment, and the embedding model has not contacted Hugging Face since the day it was downloaded.

What stays where

Hugging Face is the source of the model, not part of the runtime

After the download, nothing goes back to it.

AI Intelligence Hub embeds on your hardware

Text goes to the local model and vectors go to the index. Media never leaves Nexus.

Nothing is replaced

The embedding provider is a configuration setting, and the agents, workflows and permissions around it stay as they are.

Where processing runs

On your servers, on-premises or in your own cloud subscription, connected or disconnected. Licensing verification is the platform's one outbound call, independent of AI processing, and it is disabled for a disconnected deployment.

Products and solutions

Next step

See it on your own Hugging Face instance. We will show the connection made, the data moving and the output, then size it for your deployment.

Request a demonstration  or write to sales@vidizmo.ai

sales@vidizmo.ai  ·  vidizmo.ai/integrations/catalog/hugging-face

Not what you need?

We build it

Send us your API documentation and we build, test and maintain the connector, at no development cost to you.

Request this integration

Bring an MCP server

If the system publishes a Model Context Protocol server, AI Intelligence Hub connects to it as a client with configuration alone.

Define a REST endpoint

Describe your endpoint and it becomes a node in an agent workflow, without waiting on our roadmap.