Artificial Intelligence, Video Analytics, Manufacturing, AI Live Insight

Unsafe Behavior You Cannot Put on a Preset List

Every safety manager carries a private catalog of the incidents that taught them something, and the striking thing about those stories is how few of them would have appeared on any list of things to watch for. A crew builds a work platform out of pallets because the scissor lift is booked. An operator reaches past a jammed guard with a broom handle, then, a week later, with a hand. Two workers improvise a way to carry a die that works nine times and drops the die the tenth. None of these are on a detection menu, and that is not a failure of any particular product. It is the nature of the problem, because unsafe behavior is improvised, local, and endlessly inventive, while a trained detector can only find the detection types someone anticipated, labeled, and trained for. This article is about the technology that has started to close that gap, what it honestly can and cannot do, and how it fits into the monitoring stack described in our full guide to AI for workplace safety in manufacturing.

Where preset lists run out

Conventional detection earns its place on the predictable half of the safety problem. Missing Personal Protective Equipment (PPE), zone intrusions, vehicle and pedestrian convergence, a person down, these are recurring patterns with stable visual signatures, and the preceding articles in this series cover how well trained detection types handle them. The hazard assessment duty in 29 CFR 1910.132, which requires employers to determine whether hazards "are present, or are likely to be present," is answerable for that half with a finite list, and cameras can watch the list.

The other half of real-world exposure does not hold still long enough to be a list. Behavioral safety programs have known this for decades: the precursors to serious injury are frequently one-off configurations of ordinary things, a ladder placed against the wrong surface, a load slung in an unusual way, a shortcut invented under schedule pressure that nobody in the safety office has ever seen. A site could commission a trained detection type for each of these only after learning about it, which usually means after the incident, and by then the detection type is a memorial rather than a control. The economics compound the problem, because training a detection type takes labeled examples, and the rarest behaviors, which are often the most dangerous, by definition offer the fewest examples to train on.

For years the honest advice was to accept this boundary: cameras for the enumerable hazards, human observation programs for the inventive ones. That boundary has moved.

Asking questions instead of training detectors

The capability that moved it works on a different principle from detection, and understanding the difference is the key to evaluating it sensibly. Frames are cached from the live camera pipeline, and a multimodal model, one that reasons over images and language together, analyzes them against a question written in plain language. Describe what the workers near the press are doing. Is anyone in this sequence working at height without visible fall protection? Does anything in this loading operation look improvised?

The question is authored, not selected from a menu, and that single property changes the economics of the problem. A new concern becomes a new sentence rather than a new model, which means the safety team's accumulated judgment, all those private catalogs of what almost went wrong, becomes directly expressible as monitoring, the same day it is thought of. Because the unit of analysis is a sequence rather than one frame, the model can describe something that unfolded, which is what behavior is: not a person holding a ladder, but a person walking a ladder while extended, which a single frame would miss and a sequence makes obvious.

The architecture this fits into is a chain, and each link does what it is good at. The conventional detectors run continuously and cheaply across every enrolled camera, because watching everything all the time is what they are for. When one flags something worth interpretation, a person in an unusual posture near machinery, activity in a zone during a window when the area should be empty, the surrounding clip is escalated to the model with its question, and the answer arrives seconds later. Only what survives that second look reaches a person. The chain matters because running open-ended visual reasoning against every frame of every camera would be economically absurd, while running it against the moments the detectors have already found interesting is cheap, and the combination covers both halves of the problem, the enumerable and the inventive. The full mechanics of that escalation pattern are covered in our companion article on how the agent decides to look closer.

The limits buyers should know

This capability comes with boundaries that are structural rather than incidental, and a vendor who does not raise them unprompted is selling past them.

It is not real-time in the strict sense, and nothing about clever engineering changes that. The analysis runs on cached frames on a cadence, downstream of the live pipeline, so it answers what has been happening rather than driving a response that depends on the current frame. For behavioral monitoring this is almost always fine, because the response to "a crew is improvising scaffold" is a supervisor conversation, not an emergency stop. But any workflow where the last half-second matters belongs to the live detectors and to engineered controls, and mixing up the two layers is how deployments end up mistrusted.

There is no accuracy percentage, and there cannot be one. A trained detector has a fixed output vocabulary, so precision and recall against a test set mean something. An open-ended question has no enumerable list of right answers to score against, which means any vendor quoting a single accuracy figure for this kind of analysis is describing something other than what they are selling. The workable substitute is validation on your own footage against your own questions during commissioning, which produces a practical sense of reliability without pretending to a metric the method cannot support.

The wording of the question changes the answer. Two differently phrased prompts against the same clip return different descriptions, which is a property of the method, not a defect awaiting a patch. The operational consequence is a discipline: where an answer will be relied on, in an investigation or a disciplinary process or anything that might be contested, the prompt belongs in the record next to the finding, so the question that produced the answer is reviewable alongside it.

And the flexibility is itself an exposure to manage. A system that can be asked anything about footage of people will, sooner or later, be asked something it should not be, and the governance answer is scope agreed in advance: what questions the safety program runs, who can author new ones, and what stays out of bounds. In plants with works councils this belongs in the same written agreement as the rest of the monitoring program, and our article on worker privacy and the no-discipline commitment covers how sites structure it.

Example questions by department

Because the prompt is the product here, it is worth showing what well-designed questions look like against real plant areas, including where each one fails, since the failure modes teach as much as the successes.

For the press and machining bays: describe any interaction between a person and the machinery in this sequence, noting whether guards appear displaced or bypassed. This catches the broom-handle class of improvisation that no fixed detector was trained for, and its known weakness is legitimate maintenance, which is why the escalation review exists and why the maintenance window should suppress paging while leaving logging on, exactly as zone rules do.

For material handling: does anything about how this load is rigged, carried, or stacked look improvised or unstable? This is the question that finds the pallet-built platform and the doubled-up sling, and it fails toward caution, flagging unusual-but-engineered arrangements, which a supervisor clears in seconds with the clip in front of them.

For work at height: is anyone in this sequence above floor level without visible fall protection or on an improvised platform? Phrased with "visible," the question stays honest about what a camera can know, and the answer that comes back describes what was seen rather than certifying compliance, a distinction the safety file should preserve.

For confined and low-traffic areas: describe what the person in this sequence is doing and whether they appear to be in difficulty. This is the welfare question, the complement to the worker-down detection covered in our article on fall and worker-down reliability, and its value is precisely that it does not need to know in advance what difficulty looks like.

Who gets to author questions deserves the same governance as who gets to draw zones. The pattern that works is a small approval loop, safety leadership plus the worker representatives where a council exists, with a standing register of active prompts, their wording, and the areas they run against. The register does double duty: it is the transparency document that keeps workforce trust, and it is the version control that lets the program improve question wording deliberately instead of drifting.

Scaling your behavior observation program

The most grounded way to think about this capability is as an extension of the behavior-based observation programs most mature safety cultures already run. Those programs send trained observers to watch work and record acts and conditions against a checklist, and their known weakness is coverage: a few observation hours a week, on day shift, in the areas the observers happen to walk, with behavior shifting the moment the observer arrives.

The camera version runs the same checklist without those limits. The questions the observers carry, is the work at height protected, are loads slung correctly, are tools being used for their purpose, become standing prompts against the zones where that work happens, evaluated on every relevant escalation rather than on a sampling schedule, on every shift including the ones no observer ever visits. The output feeds the same review the human program feeds, and the human observers do not disappear, they move to the conversations the data surfaces, which was always the valuable part of their job.

What accumulates over a quarter is a behavioral picture no observation program has ever produced: which improvisations recur, where, on which shifts, trending which way. Read alongside the near-miss record, whose systematic gaps are covered in our article on why self-reported safety data understates exposure, it is the closest thing available to a measurement of the exposure that has not yet become an incident.

Reviewing the program itself

A question-driven program needs its own health check, and a short monthly review keeps it honest. Three numbers do most of the work: how often each standing question fired, what fraction of its findings supervisors confirmed against the clip versus cleared, and how long confirmed findings took to produce a change. A question that fires constantly and clears constantly is badly phrased or badly scoped, and gets rewritten or retired; the register described earlier makes that a versioned edit rather than a quiet drift. A question that never fires is either watching a solved problem, which is worth celebrating and archiving, or watching the wrong cameras. And a question whose confirmed findings never produce fixes is the program's way of showing where the response side is under-resourced, which is information leadership needs whether or not it enjoys receiving it. Run this review with the same worker representatives who approve new questions, and the program stays what it was sold as: a shared instrument, examined in the open.

How VIDIZMO fits

In VIDIZMO's platform this capability spans two products working together, and it is worth being precise about the split. AI Live Insight owns the live pipeline, running detection across the site's existing cameras on the customer's own hardware, on premises, where a continuous surveillance workload belongs, and caching the frames the analysis draws on. AI Intelligence Hub supplies the multimodal analysis and the event-driven workflows that trigger it, so an event from the live pipeline starts the graph that puts the question to the footage, and the answer comes back as a structured event like any other, searchable, alertable, and preserved with its clip in the Nexus portal the deployment works alongside. Neither product delivers this alone, and a site that only needs the enumerable half of the problem can run AI Live Insight without the reasoning layer until the day it wants more.

The sensible pilot is one question, one area, one month: take the unsafe pattern the safety team most wishes it could see, phrase it, and let the escalation chain run. The review at the end is rarely about whether the technology worked. It is about how much was happening that nobody knew.

FAQ

Frequently Asked Questions

Why can't trained detectors catch all unsafe behavior?

A trained detector only finds classes someone anticipated, labeled and trained for, and unsafe behavior is improvised and endlessly inventive. The rarest behaviors, which are often the most dangerous, offer the fewest examples to train on, so a detection type for each one could only be commissioned after learning about it, usually meaning after the incident.

How does prompt-driven video analysis work?

Frames are cached from the live camera pipeline and a multimodal model analyzes them against a question written in plain language, such as whether anyone in a sequence is working at height without visible fall protection. The question is authored rather than selected from a menu, so a new safety concern becomes a new sentence rather than a new model, and the analysis covers a sequence rather than a single frame, which is what behavior actually is.

Is this analysis real-time?

No, and it should not be sold as such. It runs on cached frames on a cadence, downstream of the live pipeline, answering what has been happening rather than driving a response that depends on the current frame. For behavioral monitoring that is almost always fine, because the response is a supervisor conversation rather than an emergency stop. Anything where the last half-second matters belongs to live detectors and engineered controls.

What accuracy rate does open-ended video questioning achieve?

There is no meaningful accuracy percentage, and there cannot be one, because an open-ended question has no enumerable list of right answers to score against. A vendor quoting a single accuracy figure for this kind of analysis is describing something other than what they are selling. The workable substitute is validation on your own footage against your own questions during commissioning.

What products are required for this capability?

Two, working together. VIDIZMO AI Live Insight runs the live pipeline on the site's existing cameras, on premises, and caches the frames the analysis draws on. AI Intelligence Hub supplies the multimodal analysis and the event-driven workflows that trigger it, with answers coming back as structured events that are searchable and alertable. Neither product delivers this alone, and a site that needs only conventional detection can run AI Live Insight by itself.

TopicsArtificial IntelligenceVideo AnalyticsManufacturingAI Live Insight

You may also like

Fire and Smoke Detection as a Second Set of Eyes in Schools

Let the first sentence of this article do the compliance work: nothing described here replaces, modifies, or competes ...

Weapon Detection in Schools: Detection, Verification, Response

No school safety technology carries more emotional weight than weapon detection, and no school safety technology is ...

School Safety Grants: What the Money Can Buy

School safety improvements have a funding problem that is really a sequencing problem: the need is continuous, the ...

See all posts

See it on your own content

Tell us what you are trying to solve and we will show you how it works on your infrastructure.