Anyone who has run a Personal Protective Equipment (PPE) program at a manufacturing site knows the strange arithmetic of compliance checks. A supervisor walks the floor twice a shift, sees perhaps forty people each time, and writes up the two violations that happened to be visible in those minutes. The other twenty-two hours of the day, across every aisle the walk did not cover, compliance is whatever it is, and the safety team finds out what it was when an injury report arrives with the phrase "employee was not wearing" in it. PPE detection on camera exists to close that gap, and it is one of the most mature applications of AI to workplace safety, which is exactly why it deserves a more careful look than the vendor slide it usually gets. This article is part of our broader guide to AI for workplace safety in manufacturing, which covers where detection fits in a safety program overall.
Detecting missing PPE is the point
Here is the point most evaluations miss at first, and it changes how you read every demo. A model that detects hard hats is nearly worthless for safety, because a plant that requires hard hats is full of people wearing them, and a system reporting each one produces a stream of confirmations nobody needs. What a safety team actually wants to know is the opposite: who is in the press bay without eye protection, right now.
That is why detection types come in pairs. For each category of protective equipment there is a positive type and a corresponding no-protection type, and the negative half is the one doing the safety work. On VIDIZMO's detection path there are fourteen granular detection types covering head, eye, ear, hand, and foot protection along with safety vests, bulletproof vests, and fire extinguishers, each with its negative form. When an evaluation compares vendors, this is the first question worth asking, because some processing paths offer a much coarser set. A four-type detector that reports "PPE present" and "PPE absent" cannot tell you that the gloves policy is holding while the eye-protection policy is failing, and those are different problems with different fixes.
The regulatory frame makes the same point from the other direction. The PPE standard at 29 CFR 1910.132 requires employers to assess the workplace for hazards, select equipment that protects against them, and ensure it is used. The standard's language is about hazards and the equipment that answers them, which means an employer's obligations are specific: this zone requires eye protection because of that grinder, that bay requires hearing protection because of this press. A monitoring capability that cannot express the same specificity cannot verify the program the law actually requires.
Matching detection types to zones
The second thing that separates a working deployment from a noisy one is scope. A manufacturing site is a patchwork of PPE regimes. The machine shop requires eye and ear protection, the loading dock requires high-visibility gear, the chemical store requires gloves and possibly respirators, and the office corridor that cuts between them requires nothing at all. A detection system configured as if the plant were one uniform space will flag office staff walking to a meeting, and its credibility will not survive the first week.
The mechanics that prevent this are configuration rather than intelligence. Detection types are enabled per camera, so the camera over the grinder watches for missing eye protection while the camera on the dock watches for missing hi-vis, and neither wastes attention on the other's rules. Zones drawn on the camera frame narrow it further, so the rule applies inside the marked area and not on the walkway beside it. Confidence thresholds are set per detection type per camera, which matters because lighting, distance, and angle differ enough between cameras that a threshold tuned for one produces false alarms on another.
This is also where a fair amount of deployment honesty belongs. Detection quality depends on what the camera can actually see. A camera mounted for general surveillance, high and wide, may resolve a person clearly while resolving safety glasses poorly, since glasses occupy a few dozen pixels at that distance. Ear protection is harder than head protection almost everywhere, because earplugs are nearly invisible to any camera and earmuffs are often occluded by hard hats or hair. An honest pilot measures which detection types perform on which cameras before anyone writes a rollout plan, and treats "this camera cannot support this detection type" as a finding rather than a failure.
What happens when a detection fires
A detection that nobody acts on is a statistic, so the pipeline after the detection deserves as much scrutiny as the model. When a no-protection type fires, three things happen in the same motion: an alert goes out, an event lands on the timeline, and the recording around the moment is preserved with configurable footage before and after the trigger. The alert reaches the people subscribed to that camera, carrying the severity, the detection type, and a snapshot from the moment itself, so a supervisor can decide from their phone whether this needs a walk over or a note.
Severity is the underrated control in that chain. Missing eye protection at the visitor entrance and missing eye protection inside the machining cell are the same detection type carrying entirely different urgency, and a system that treats them identically trains supervisors to ignore both. Setting severity per detection type per camera is tedious work during commissioning and it is the work that decides whether the alert channel stays trusted.
Over weeks, the accumulated events become the thing the individual alerts never were: a measurement. Compliance rates by zone, by shift, and by day expose patterns no walkthrough could, and they tend to be uncomfortable in useful ways. Third shift compliance that drops after the second hour is a fatigue and supervision finding. A bay whose rate falls every time a particular product runs is a workflow finding, usually meaning the required equipment interferes with the task. These are the leading indicators our companion article on safety Key Performance Indicators (KPIs) covers in depth, and PPE compliance is typically the first one a site can measure continuously.
Monitor conditions, not individuals
Every PPE monitoring project reaches the same fork, usually in the first planning meeting: does the system tell you that someone is unprotected, or who is unprotected. The technology can support either, since face recognition against an enrolled gallery is a separate capability that could in principle be layered on. Almost every successful safety deployment declines it, and the reasoning is worth spelling out rather than treating as squeamishness.
A program that identifies individuals becomes a discipline instrument the moment the first write-up cites camera evidence, and workers respond to discipline instruments rationally, by finding the camera blind spots and by withdrawing the goodwill that safety programs run on. A program that measures conditions, meaning rates by zone and shift rather than names, keeps the workforce inside the project. The finding "eye protection compliance in bay four drops during changeovers" leads to a conversation about why, and the answer is usually a fogging problem, a supply problem, or a task-interference problem that discipline would never have surfaced. In plants with works councils, this distinction is not optional, and our article on worker privacy and the no-discipline commitment covers the commitments that get these systems approved at all.
None of this prevents a supervisor responding to a live alert from walking over and talking to the person. It prevents the archive from becoming a personnel file, which is a different thing.
What detection cannot verify
The 1910.132 standard contains a requirement that camera detection cannot check, and pretending otherwise would be exactly the overclaiming this series avoids. Since a 2016 revision for construction and long-standing guidance elsewhere, the standard requires employers to "select PPE that properly fits each affected employee," and fit is invisible to a camera at surveillance distance. A hard hat perched loosely, gloves a size too large for the task, a respirator with a broken seal, all of these read as compliant to a detector because the equipment is present, and every one of them is a protection failure in the way that matters. Fit verification stays with the people who hand equipment out and the workers who wear it, and a monitoring program should say so in its own documentation rather than allowing anyone to believe the cameras cover it.
The same honesty applies to a handful of adjacent gaps worth listing during planning. Respirator cartridge selection and change-out schedules are chemistry, not vision. Fall-arrest harnesses can be detected as present, but anchorage and lanyard condition cannot be assessed from a fixed camera. Cut-resistant versus general-purpose gloves look identical at distance, so a zone that requires a specific glove requirement is verifiable only to the level of "gloves present." None of this diminishes what detection does verify, which is the daily, shift-by-shift question of whether required equipment is being worn at all, the question no manual program has ever answered continuously. It just draws the boundary where it belongs, and a safety leader who states that boundary in the rollout communication builds more credibility with the floor than any capability claim would.
Running a pilot that produces a real baseline
The difference between a PPE detection pilot that convinces a leadership team and one that dissolves into argument is almost always methodology rather than technology, so it is worth describing the shape that works. Pick two or three zones with unambiguous PPE rules, one where compliance is believed to be strong and one where the safety team quietly suspects it is not, because the contrast is where the learning lives. Run the first two weeks silent, with no alerts to anyone, letting the system accumulate a baseline nobody has had a chance to react to. That baseline is the only honest one the site will ever collect, since the moment supervisors know the cameras are counting, behavior shifts, and the shift is itself informative but it is a different measurement.
Then review three numbers with the people who own the zones. The compliance rate by shift, which tells you whether this is a training problem or a supervision-coverage problem. The false-positive rate by detection type and camera, which tells you which cameras can support which detection types and which thresholds need moving, and which should be honestly retired from the plan. And the time-of-day pattern, because compliance that sags in the last two hours of a shift is a fatigue finding, while compliance that sags during changeovers is a workflow finding, usually meaning the equipment interferes with the task in a way nobody has said out loud. Sites that walk into the readout with those three numbers, measured over identical periods, tend to leave with budget. Sites that walk in with a vendor's demo reel tend to leave with another pilot.
Questions to ask vendors
PPE detection is offered by nearly everyone in video analytics, so evaluations tend to bog down in claims that all sound alike. A few questions separate the field faster than a bake-off.
Ask for the preset list with the negative forms, in writing, since the difference between four coarse detection types and fourteen granular ones with absence detection decides what your program can verify. Ask where the processing runs, because continuous camera streams are exactly the workload that does not belong on a round trip to an external service, and for latency and bandwidth reasons alone live analysis belongs close to the cameras; a vendor whose architecture requires footage to leave the site has a different answer to worker privacy questions too. Ask how thresholds are set, because "per detection type per camera" and "globally" describe two different products. Ask what happens to the clip, since a detection that cannot be reviewed with footage attached will not survive a contested write-up or an insurance conversation. And ask which detection types degrade on which camera types, because a vendor who answers honestly about ear protection has probably tested the rest.
How VIDIZMO handles it
VIDIZMO AI Live Insight leverages the cameras a plant already owns rather than requiring AI-enabled replacements, reading standard Real-Time Streaming Protocol (RTSP) and ONVIF (Open Network Video Interface Forum) streams and running detection on the customer's own hardware, on site, which is where a continuous surveillance workload practically belongs. The fourteen PPE detection types with their negative forms are configured per camera alongside zone rules, confidence, and severity, and a detection raises the alert, marks the timeline, and preserves the surrounding clip in the same motion. The clips land in the VIDIZMO Nexus portal the deployment works with, where access control and retention policy govern who sees them and how long they exist, which is what keeps the archive a safety record rather than a personnel file.
A pilot on a handful of cameras over the zones with the clearest PPE rules is the sensible start, and it produces the compliance baseline the rest of the program gets measured against.