Integration brief · Ollama ← Back to the page   Print or save as PDF
Integration brief OllamaAI and LLM Providers

Language Models That Never Leave the Building

Ollama runs open-weight language models on a server or workstation you own. Connected to AI Intelligence Hub, it serves both the language model that writes answers and the embedding model behind retrieval. Every prompt, passage and vector is processed on your own hardware. The platform's own transcription, vision, OCR and PII models run locally beside it, which is what makes an air-gapped installation workable. Content stays in Nexus, and no inference request leaves the site.

What you can do together

  • Ask AI Intelligence Hub a question across interview recordings, body-camera transcripts and discovery documents and get a cited answer without a single external call.
  • Run the whole AI stack disconnected: Ollama for language and embedding models, with local Whisper transcription, vision, OCR and PII detection alongside.
  • Choose the model per agent, so a small model handles routine questions and a larger one handles case analysis, both on your hardware.
  • Keep every answer inside what the asking user is entitled to open, with the agent scoped further to a case folder or category.
  • Build workflows with a human approval step, with the model node pointed at Ollama.

How it connects

AI Intelligence Hub calls the Ollama server over its REST API on your network, at the address supplied in configuration. The model provider is configuration, set per deployment or per agent. Pointing an agent at Ollama rather than a hosted endpoint changes where inference happens without changing what the workflows do.

Two things flow across your own network. At indexing time the text of transcripts, OCR output, visual descriptions and documents goes to Ollama's embedding model and vectors return to the index. At question time the question and the passages retrieved for the asking user go to the language model and the answer returns, cited to the passages it drew on.

Nothing flows outside. Licensing verification is the one outbound call the platform makes, independent of AI processing, and it is disabled for a disconnected deployment. Model quality follows the models you host, and the server needs the GPU capacity for them.

VIDIZMO and Ollama · Integration briefPage 1 of 2
How it works OllamaAI and LLM Providers
Nexus your video, audio, images and documents, indexed under your access rules content the asking user may see AI Intelligence Hub agents, chat, workflows, retrieval with citations; model chosen per agent prompts and embeddings over API Ollama model endpoint, hosted by the provider or on your own servers answers Users chat, search, workflow output Ollama VIDIZMO

A scenario

  1. SetupIT points AI Intelligence Hub at the Ollama server for both the language model and the embedding model, and scopes the Case Review agent to the discovery folders.
  2. IntakeInterview recordings and body-camera footage for case 24-1187 are transcribed by the local Whisper model on ingest, and discovery documents are read by local OCR. All of that text is embedded through Ollama. Everything stays in Nexus.
  3. 10:20An assistant prosecutor asks, "Across the interviews in 24-1187, where does the witness describe the vehicle?" The agent retrieves the passages from the interviews she is assigned to and sends them to the Ollama model.
  4. Seconds laterThe answer quotes each description and cites the moment in each recording. She opens a card and the interview plays from that second.
  5. 10:35She asks for the statements in the order they were given. The conversation carries the context, and the answer cites the same recordings.
  6. ThroughoutThe deployment has no route to the internet and licensing verification is disabled. Prompts, passages and vectors traveled between the platform and the Ollama server on the county network and nowhere else.

What stays where

Ollama runs on your server

You choose the models, load them and size the GPU behind them.

AI Intelligence Hub sends prompts and text across your own network

The question, the retrieved passages and the text to embed go to Ollama, and answers and vectors come back. Nothing goes to a hosted API.

Nothing is replaced

Ollama is a provider setting. An agent moves between Ollama and a hosted provider by configuration, and the workflows behind it do not change.

Where processing runs

Every inference runs on the hardware you pointed the platform at, in your data center or your own cloud subscription. Air-gapped, the same setup works with licensing verification disabled.

Products and solutions

Next step

See it on your own Ollama instance.

We will show the connection made, the data moving and the output, then size it for your deployment.

Contact VIDIZMO

sales@vidizmo.ai

+1 571-969-2180

vidizmo.ai

Product names and logos are the property of their respective owners.Page 2 of 2