A typical manufacturing plant runs hundreds of cameras and still learns about most of its safety incidents from paperwork, days after the fact, when the corrective action report needs a picture of something that already went wrong. The cameras record everything, the recorder overwrites itself on a thirty-day cycle, and nobody watches the feeds, because no staffing plan on earth covers three hundred screens around the clock. AI for workplace safety exists to close exactly that gap: software that watches the live feeds continuously, recognizes the conditions that lead to injuries, and tells the right person while there is still time to act. This guide explains what the technology can genuinely do on the cameras a plant already owns, what it takes to deploy well, and where its honest limits sit.
The case for it starts with how safety performance is measured today. A corporate Environmental, Health and Safety (EHS) function runs its program on numbers that arrive after the harm: recordable injury rate, lost-day rate, workers' compensation cost per site. Boards ask for those numbers, and every one of them describes a past nobody can change. Meanwhile the conditions that predict the next injury, a person stepping into a live robot cell, a forklift and a pedestrian sharing a blind aisle, missing protective equipment on a night shift, play out on camera every single day, counted by no one. The plants that get ahead of their injury rate are the ones that start measuring those conditions while they are still conditions, and that is what the rest of this guide is about.
Why safety data arrives too late
The Bureau of Labor Statistics reported that private industry employers logged 2.5 million nonfatal workplace injuries and illnesses in 2024, down 3.1 percent from the previous year and the lowest figure in a data series that goes back to 2003. The total recordable case rate came in at 2.3 cases per 100 full-time equivalent workers, down from 2.4. The fatality census tells the same story at its sharpest point: 5,070 workers died on the job in 2024, 353 of them in manufacturing, and both counts fell against the year before. Those figures, released in early 2026, are the most recent that exist, which is itself worth sitting with for a moment. National safety data runs more than a year behind the floor it describes, so a safety leader reading the newest available numbers is reading about a workforce that has since turned over, reorganized, and changed shift patterns. Read quickly, the trend is good news, and the long trend genuinely is. Read carefully, it also shows the limit of what recordable-injury data can tell a plant manager, because a rate expressed per hundred workers per year is a very coarse instrument for a site trying to work out which of its aisles is dangerous this month.
The deeper problem is that the denominator of safety knowledge is self-reporting. Near misses, the events where the hazard was real and the injury simply did not happen, are the richest available signal about where the next injury comes from, and they reach the safety team only when a person decides to file one. People file fewer of them when the shift is behind schedule, when the reporting form is long, when the crew believes reporting triggers an investigation aimed at them, and when the near miss involved something they were doing to save time. None of those pressures are unusual, and all of them push the same direction. A site with a low near-miss count is often a site where reporting has quietly stopped, not a safe site, and experienced EHS leaders read a sudden decline in near-miss volume as a warning rather than a win.
So the picture a safety team works from is assembled out of injuries that already happened, observations from the hours a supervisor happened to be walking the floor, and voluntary reports filtered through production pressure. It is not that anybody is doing this badly. It is that the instrument is thin, and thin instruments produce programs that react.
What OSHA requires and where cameras help
It helps to be precise about what a US manufacturer is legally obliged to do, because a lot of vendor material implies that video analytics satisfies obligations it does not touch.
The personal protective equipment standard at 29 CFR 1910.132 requires that protective equipment be "provided, used, and maintained in a sanitary and reliable condition" wherever hazards make it necessary. The part that matters for this discussion is paragraph (d), which puts an affirmative duty on the employer to "assess the workplace to determine if hazards are present, or are likely to be present," and then to select PPE that protects against the hazards that assessment identified and that properly fits each affected employee. The standard asks for an assessment and for equipment that gets used. It does not prescribe how an employer verifies that the equipment is being worn on the third shift, which is exactly where continuous monitoring becomes interesting rather than merely surveillant.
Powered industrial trucks are governed by 29 CFR 1910.178, and the paragraphs read like a list of the things that go wrong in a warehouse. Trucks are not to be driven up to anyone standing in front of a fixed object. No person is to stand or pass under the elevated portion of a truck, loaded or empty. Operators are required to look in the direction of travel and keep a clear view of the path. Training must specifically cover pedestrian traffic in the areas where the vehicle operates. Every one of those is a spatial relationship between a vehicle, a person, and a moment in time, which is the category of thing a camera and a tracking model are genuinely good at recognizing.
The pattern holds across the standards manufacturers get cited on. In the Occupational Safety and Health Administration's (OSHA) most frequently cited standards for fiscal year 2025, the control of hazardous energy, better known as lockout/tagout, sits at number four; powered industrial trucks at number eight; and machine guarding, general industry, at number ten. Respiratory protection is at number five. These are not obscure rules. They are the recurring failures of ordinary industrial work, they cluster around equipment and movement and protective gear, and they are visible to anyone standing in the right place at the right time. The trouble has always been that nobody is standing there.
The cost of the citation keeps climbing with inflation adjustments. For penalties assessed after January 15, 2026, OSHA's maximums stand at $16,550 for a serious violation and $165,514 for a willful or repeated one, with failure to abate accruing another $16,550 for each day past the abatement date. A single serious citation is survivable for most manufacturers. The willful finding that follows a documented pattern nobody acted on is a different order of exposure, and it is worth understanding early that a monitoring system's record cuts both ways: the same event log that demonstrates a working safety program to an insurer will document a known, unaddressed hazard with equal clarity. That argues for pairing detection with the resourcing to respond, not for avoiding the record.
Two cautions belong here rather than at the end. First, a detection is not a compliance determination. A model that flags a worker without eye protection has produced an observation, and turning that into a citation-relevant fact requires a human to confirm the zone, the task, and the applicable policy. Second, OSHA's recordkeeping rule at 29 CFR Part 1904 governs what goes on the log, and no camera decides whether a case is recordable. Video is evidence that supports a determination somebody else makes.
What a camera can already tell you
Vendor decks in this market tend to promise everything, so the more useful exercise is to walk through what a standard fixed IP camera pointed at a plant floor can actually detect today, capability by capability, with the conditions each one carries.
Protective equipment detection works because the models are trained on the absence as well as the presence. A detector that recognizes a hard hat is close to useless for safety, since a plant that requires hard hats mostly contains people wearing them, and an alert stream reporting compliance is noise. The useful detection types are the negative ones: head protection absent, eye protection absent, high-visibility gear absent. Detection of missing equipment is what turns a camera into a safety instrument, and it is the difference between a system that produces a count and a system that produces an intervention.
Spatial rules are the second building block, and they matter more than most buyers expect. A detection on its own says a person is in frame. A rule says a person is inside the polygon drawn around the press, or has crossed the line at the mouth of the robot cell in the direction that leads into it, or has remained inside a marked area longer than the dwell time the rule allows, or that the count of people in a space has exceeded what the space is rated for. Each rule is drawn on an image from the camera itself, scoped to a target type, and evaluated against tracked objects as they move. This is what makes a hazard zone expressible: not "there is a person," but "a person is where a person should not be while that machine is energized."
Tracking is the quiet piece that makes the rest work. Without it, a camera produces a detection for every frame, and one worker crossing a bay becomes hundreds of alerts. With tracking, that same crossing is one object with one identity, one entry, one exit, and one event with a start and an end. Everything downstream, the alerting, the counting, the trend analysis, depends on the system knowing that it saw one person once rather than a person two hundred times.
Text recognition rounds it out for a set of cases people rarely anticipate. Markings, plate numbers, and labels visible in the frame are read and emitted as searchable events whether or not they match anything anyone was looking for, which means a coil marking or a vehicle plate is findable later even when nobody set an alert for it in advance.
Detecting what you cannot predefine
Here is where most safety-analytics projects hit their real limit, and where the technology has changed recently enough that a lot of published material is out of date.
Fixed detectors find what they were trained to find. If a plant wants to know about workers reaching into a machine while it is running, or a crew improvising a work platform out of pallets, or someone standing beneath a suspended load, each of those has to exist as a trained detection type before it can be detected. Some of them can be trained, given labeled footage from the site. But the category of unsafe behavior is open-ended in a way that a preset list is not, and every safety manager can describe an incident whose precursor would never have appeared on any list drawn up in advance.
The approach that addresses this works differently, and the distinction is worth understanding before evaluating any vendor's claims. Rather than asking a fixed detector what it sees, the system caches frames from the live pipeline and puts a written question to a multimodal model: describe what is happening in this sequence, or say whether anyone here appears to be working at height without fall protection. The question is authored rather than selected from a menu, so a new kind of concern requires a new prompt and not a new model. Because the unit of analysis is a sequence rather than a single frame, the answer can describe something that unfolded over a period, which is what most unsafe behavior actually does.
The practical architecture that results is a chain rather than a single step. A conventional detector runs continuously and cheaply, flags something worth a second look, and escalates a short clip to the model that can interpret it. What survives that second look reaches a person. That escalation costs seconds, which is the honest description and a meaningful one, but seconds against the alternative of nobody looking at all is not a real objection. A supervisor who learns about a hazard forty seconds after it appeared is in a different world from one who learns about it in the incident review three days later.
Two limits attach to this, and any vendor who does not volunteer them is worth pressing. There is no enumerable list of what a prompt-driven analysis recognizes, which means there is also no accuracy percentage for it in the way a trained detector has one, and a supplier quoting a precision figure for open-ended visual questioning is describing something they have not measured. And the wording of the question changes the answer, so where a result will be relied on in an investigation, the prompt belongs in the record next to the finding.
Alert quality decides adoption
Nearly every deployment of this technology that fails goes down at this exact point. The technology detects correctly, the volume is unmanageable, the supervisors stop looking, and within a quarter the system is running and nobody reads it.
The controls that decide the outcome are unglamorous ones that no demo ever shows. Confidence sets how certain the model must be before a detection counts, and severity sets how much a given detection type matters on a given camera, and both are configured per detection type per camera rather than globally. That granularity is the whole point. A missing-hard-hat detection on a camera covering the maintenance mezzanine is a different event from the same detection on a camera covering a visitor walkway, and a system that cannot express that difference will either flood the queue or suppress the thing you cared about.
Inference cadence is the other dial, and it is commonly misunderstood as a hard limit when it is a tuning decision. Running the detector on every frame produces the finest temporal resolution and consumes the most Graphics Processing Unit (GPU); running it on an interval with tracking filling the gaps serves many more cameras from the same hardware. Neither is correct in the abstract. A camera watching a fast packaging line and a camera watching a chemical store have genuinely different requirements, and the sizing conversation is about which cameras deserve which cadence rather than about a fixed ceiling on what the platform can do.
Budget real calendar time for the tuning itself, not just for the installation. A pilot that starts with every detection type enabled on every camera at default thresholds will generate a bad first impression that the technology did not deserve, and the fix is a few weeks of narrowing rather than a different vendor.
From detection to defensible record
A safety alert, once it has done its job, has a remarkably short life. Somebody sees it, somebody acts, and the moment passes. The durable value shows up later, when an insurer asks why premiums should not rise, when a citation is contested, or when the same corner of the plant produces its third injury in a year and nobody can reconstruct what the first two had in common.
Recording that follows detection is what makes this possible. Rather than keeping continuous footage from every camera and hoping the relevant minute survives the retention window, the system writes a clip around each event: a configurable stretch of footage before the trigger, the event itself, and a stretch after it ends. The library then holds the moments that mattered instead of thousands of hours of empty corridor, which changes both the storage economics and, more importantly, the odds that the clip still exists when somebody finally asks for it.
Where those clips land matters as much as whether they were captured. A clip that sits in a surveillance silo with its own permissions and its own retention rules is an operational artifact. A clip that lands in a content library with access control, a retention policy, and an audit trail over who viewed and exported it is closer to a record, and the distinction becomes sharp the moment an incident turns into a claim or a grievance. Safety teams rarely think about this at purchase and reliably wish they had.
Worker trust and privacy
None of the preceding matters if the workforce refuses, and in a lot of manufacturing environments that is the live risk rather than a hypothetical one. In plants with a works council or a strong union presence, camera-AI deployments stall on consent long before they stall on technology, and a safety leader who arrives with a signed contract and no worker engagement plan has usually lost a year.
The commitments that get these programs approved are consistent across the sites where they succeed. Detections are used to change conditions rather than to discipline individuals, and that commitment is written down rather than promised verbally. Aggregate trend data goes to the safety committee, including the worker representatives on it. Faces are not enrolled against a gallery for safety use cases, because identifying which person was not wearing eye protection is a different project from knowing that the eye-protection rate in that bay is falling. Retention is short and stated. Coverage is agreed, including which areas are excluded, and break areas and changing facilities are excluded as a matter of course.
Framed that way, the system measures the workplace rather than the worker, and that framing is not a communications trick. It is a design decision that shows up in what the system is configured to do, and workers can tell the difference between a program built that way and one dressed up to look like it.
Where this is the wrong tool
A guide that only describes what a technology does is marketing. Several situations are genuinely poor fits, and recognizing them early saves money.
If the hazard is not visible, cameras have nothing to offer. Gas concentration, noise exposure, vibration and thermal stress all need instrumentation of their own, and a video layer is at best a way of seeing whether people are where the sensor says the danger is. If the plant's actual constraint is that supervisors already know about the hazards and lack the budget to fix them, more detection produces a longer list and no improvement, and the safety team ends up defending a system that documented failures it could not resource. If the requirement is inspecting every unit on a fast line for fine surface defects, that is a machine-vision problem with purpose-built optics and lighting, related to this discussion but not the same as it. And if the site has fewer than a dozen cameras and one supervisor who genuinely can see the floor, the honest answer is that the manual process is working.
How VIDIZMO approaches it
VIDIZMO AI Live Insight leverages the camera infrastructure a plant already owns rather than relying on the costlier route of AI-enabled cameras. Any camera or recorder that produces a standard Real-Time Streaming Protocol (RTSP) or ONVIF (Open Network Video Interface Forum) stream becomes an AI-monitored feed through a software layer, which in practice covers the Axis, Hanwha Vision, and Bosch class of fixed cameras and the sites running Milestone XProtect, Genetec Security Center, and their peers, so the intelligence lives on the server rather than in the camera housing, and the cameras, the Video Management System (VMS), and the network segmentation all stay where they are. The detection types, the zone rules, and the confidence and severity thresholds are configured per camera, and detections raise alerts, become events on a timeline, and trigger event-based recording in the same motion.
AI Live Insight works side by side with a VIDIZMO Nexus portal and uses it as the library its clips land in, so a captured clip inherits the portal's access control, retention policy, and audit trail rather than living in a separate surveillance store, which is what allows a safety clip to be handled as a record when an incident becomes a claim. Where a site needs the open-ended questioning described earlier, AI Intelligence Hub supplies the multimodal analysis and the agent workflows that act on events, and the two products are licensed separately because plenty of deployments need the first without the second.
The deployment question that usually decides the shortlist is where the processing happens, and for live surveillance the practical answer is close to the cameras. Streaming dozens of continuous feeds off site and waiting on the round trip adds bandwidth cost and latency that a real-time safety alert cannot afford, so this workload runs on the customer's own hardware as a matter of engineering rather than preference, including fully disconnected. Manufacturers under export control or running segmented Operational Technology (OT) networks arrive at the same answer for their own reasons. That is a longer conversation than this guide can hold, and it belongs with the IT and OT owners rather than with the safety team.
Where to go from here
This guide covers the shape of the problem. The specific decisions live in the articles below, each of which goes deeper than the summary here allows.
For detection scope, start with what protective equipment detection actually covers and which detection types earn their place, then hazard zones and restricted areas for the rule types that make a zone meaningful, and vehicle and pedestrian proximity for the aisle problem that produces so many of the serious injuries. Worker-down and fall detection covers the class of event where response time changes the outcome.
For program design, leading indicators and safety Key Performance Indicators (KPIs) covers what to measure once detections start arriving, the near-miss article covers why the baseline you are comparing against is probably wrong, and alert fatigue covers the tuning work that decides whether the deployment survives its first quarter.
For the harder conversations, works councils and worker privacy covers the consent problem, OSHA recordkeeping covers what video can and cannot support on the log, and investigating a workplace accident from camera footage covers what happens after the event you were trying to prevent happens anyway.