Enterprise security systems — service, repair, programming and support

Why security servers and storage should be monitored

Access-control and video systems fail loudly at the edge and quietly at the centre. A door that will not unlock generates a phone call within minutes. A recording server writing to a degraded array generates nothing at all.

That asymmetry is the whole argument for monitoring the server and storage layer specifically. The failures there are not dramatic. They are conditions that persist, and the cost only lands at the moment somebody needs the footage or the audit trail, which is invariably the moment when it matters most and when there is no time to fix anything.

Below are the conditions worth watching, and what each one actually exposes you to.

The rebuild window is the real risk, not the failed drive

A single drive failure in a redundant array is a designed-for event. The array continues, the system records, nobody notices. That is exactly the problem: on many installations the degraded state persists for weeks because the alert went to a mailbox that nobody reads, or to a management interface nobody logs into.

The exposure begins when the replacement drive goes in. Rebuilding a modern high-capacity nearline drive is not a quick operation. On a large array under a continuous video write load, a rebuild can run for a day or longer, and it is the most sustained read the array will ever experience, because every remaining drive must be read end to end to reconstruct the missing member. Two things can go wrong during that window. A second drive can fail, which on single-parity RAID 5 ends the array. Or a single unrecoverable read error on any surviving drive can halt the rebuild, which on a large array is not a remote possibility. This is the reason dual-parity layouts exist for large-capacity arrays, and the reason a degraded array left degraded for a month is a far larger risk than it appears.

The practical consequences: alert on degraded state immediately and treat it as urgent rather than informational, keep a spare drive of the correct model on site so the rebuild starts the same day, and know before you need it whether the array is single or dual parity and how long a rebuild takes at your capacity.

What to watch and what each item catches

SignalWhat it prevents
Array stateDegraded, rebuilding, foreign or failed memberWeeks of unprotected operation before anyone notices
Drive healthReallocated and pending sector counts trending upwardA drive that fails during a rebuild rather than before it
Controller cacheBattery or supercapacitor fault on the RAID controllerA silent switch to write-through, which collapses write performance and drops frames
Volume free spacePercentage free and days-to-full at current growthRecording halting, or a database refusing writes
Database sizeEvent or journal growth rate against any engine size ceilingAn Express-edition database hitting its hard limit and stopping history
Retention actualOldest recording per camera against the retention policyDiscovering at retrieval time that thirty days is now eleven
Recording statePer-camera recording status, not just reachabilityA camera that pings, streams live and is writing nothing
Power supplyRedundant PSU failure and input feed lossRunning on one supply for months, then losing the second
UPS runtimeBattery age and a real load test, not a self-test passA UPS that reports healthy and delivers ninety seconds
Time syncNTP offset on servers, controllers and camerasVideo and access events that cannot be correlated in an investigation
CertificatesExpiry dates on server and integration certificatesClients and integrations failing to connect on a date nobody diarised
Backup jobsSuccess, and the size of the resulting fileEight months of jobs that reported success and wrote nothing usable

Each row is a condition that is invisible from the operator interface and obvious from the server. That gap is what monitoring closes.

Retention drift, and why it surprises people

Retention is a calculated outcome, not a setting. It is a function of camera count, resolution, frame rate, codec and compression, scene complexity and motion, and the storage available. Change any input and retention changes, silently and immediately.

The common sequence: cameras are added over three years without a storage review; a handful of cameras are moved from a fixed frame rate to a higher one for a specific investigation and never moved back; a lens is changed and the scene now has constant motion from traffic, which raises the bitrate for that stream permanently. The system does not warn about any of this. It simply keeps the newest footage and discards the oldest, exactly as designed. The written policy still says thirty days. The array now holds eleven.

The check is trivial and almost nobody runs it: for each camera, ask the system for the oldest recorded frame, and compare that date against the policy. Run it monthly. A camera whose oldest footage is materially newer than its peers is either recording at a much higher bitrate than expected or has been recording continuously where it was supposed to record on event.

Server-side conditions that are not storage

Operating system behaviour

  • Automatic updates rebooting a recording or directory server without a window.
  • Antivirus real-time scanning pointed at the recording volume, adding latency to every write.
  • A pagefile or temp directory sharing the recording volume and competing for the same spindles.

Services and accounts

  • Application services set to manual, so a reboot leaves them stopped.
  • A service account with an expiring password, which fails at expiry rather than at change.
  • An administrative account still tied to a departed employee and disabled by the joiner-mover-leaver process.

Hardware management

  • Out-of-band management such as iDRAC or IPMI configured but sending alerts nowhere.
  • Alert destinations pointing at a distribution list that has been dissolved.
  • Firmware on the storage controller and drives left at the shipped version for the life of the machine.

The network path

  • A single uplink where the design assumed two, after a switch replacement.
  • Multicast or QoS configuration lost in a switch swap, which changes stream behaviour under load.
  • Firewall changes that block an archiver or a controller port and produce an intermittent fault, not a clean failure.

Monitoring is not the same as alerting

Collecting all of this and sending it to an unread mailbox is the failure mode most sites already have. Two rules make the difference. First, every alert has a named owner and a stated action, and an alert with neither should be removed rather than tolerated. Second, alerts are tested by causing the condition, not by assuming the configuration works. Pull a drive from a test array and confirm somebody receives the notification and knows what to do with it.

Backup jobs deserve their own treatment because a successful job report is weak evidence. What proves a backup is a restore. That subject is covered separately in why configuration backups matter. Where server work is part of a broader platform move, see preparing for a Genetec or C•CURE upgrade.

Find out what your server layer is not telling you

A health review covers array state, retention actuals, database growth, backup verification and alert routing.


Working through a similar problem?

Describe the platform and what the system is doing. Please do not include passwords, IP addresses, credentials or facility drawings.