Open Embedding Models, Downloaded Once, Run Disconnected
Hugging Face hosts a public catalog of open models, including the sentence embedding models that retrieval by meaning depends on. Connected to AI Intelligence Hub, a Hugging Face embedding model turns the text of your transcripts, documents and visual descriptions into vectors on your own hardware. The model is fetched once and then runs inside your deployment, which is what lets an air-gapped installation index by meaning with nothing leaving the boundary. Content stays in Nexus.
How it connects
The embedding provider is configuration. Set to Hugging Face, the platform fetches the named model from the Hugging Face catalog over its API and runs it inside the deployment's own embedding service. For a disconnected site, the model is downloaded once and then runs with no further connection.
At indexing time, the text of each stream goes to that local model and vectors return to the index: transcripts, OCR text, visual descriptions and document text. Nothing goes to Hugging Face after the download. The language model that writes answers is configured separately, and for a self-hosted deployment that is a vLLM or Ollama server on your own hardware.
Retrieval runs under the asking user's identity with their access list applied before the vector search, whichever embedding model is in use.
What you can do together
- Search by meaning across AI Intelligence Hub with an open embedding model of your choosing, running on your own servers.
- Index by meaning in an air-gapped deployment, where a hosted embedding endpoint is unreachable by definition.
- Pair the model with vLLM or Ollama for answers, so retrieval and generation both stay inside the boundary.
- Keep every answer inside what the asking user may open, cited to the moment or the page.
A scenario
- SetupIT downloads the chosen embedding model from Hugging Face once, installs it with the platform, and sets it as the embedding provider in AI Intelligence Hub. A vLLM server on the same rack serves the language model.
- IntakeSeized-device audio, surveillance footage, interview recordings and scanned documents for case 23-0412 are transcribed, described and read by the platform's local Whisper, vision and OCR models. The Hugging Face model then embeds the text. Everything stays in Nexus.
- 11:05An analyst asks, "Which recordings in 23-0412 mention a storage unit?" The agent retrieves the passages from the recordings she is assigned to and sends them to vLLM with the question.
- Seconds laterThe answer cites three moments across two recordings. She opens a card and the audio plays from that second.
- 11:20She asks the same of the scanned documents, and the answer cites each page.
- ThroughoutNothing had a route out. Licensing verification is disabled for the disconnected deployment, and the embedding model has not contacted Hugging Face since the day it was downloaded.
What stays where
Hugging Face is the source of the model, not part of the runtime
After the download, nothing goes back to it.
AI Intelligence Hub embeds on your hardware
Text goes to the local model and vectors go to the index. Media never leaves Nexus.
Nothing is replaced
The embedding provider is a configuration setting, and the agents, workflows and permissions around it stay as they are.
Where processing runs
On your servers, on-premises or in your own cloud subscription, connected or disconnected. Licensing verification is the platform's one outbound call, independent of AI processing, and it is disabled for a disconnected deployment.
Products and solutions
- AI Intelligence Hub
- Digital Evidence Management, Corporate Investigations and Intelligent Document Processing
Next step
See it on your own Hugging Face instance. We will show the connection made, the data moving and the output, then size it for your deployment.
Request a demonstration or write to sales@vidizmo.ai
sales@vidizmo.ai · vidizmo.ai/integrations/catalog/hugging-face