Services
Troubleshooting and System Restoration
Diagnosis and repair of faults in enterprise access control, video and security server systems, including the faults that earlier service attempts did not resolve. We find the cause before anything is replaced.
Recurring faults
A system that keeps failing is not mysterious. It is undiagnosed.
Intermittent faults have causes. They are harder to find than constant faults. They are not different in kind.
A door that releases late twice a month, a panel that drops off the network at 3am and is back by shift change, a camera that records for six days and then quietly stops: each of these leaves evidence behind. Access control event history, controller diagnostics, archiver logs, database job records, switch port counters and Windows event logs all hold traces of what happened and when it happened. Restoration is the work of reading that evidence until the pattern is no longer ambiguous.
The alternative method is replacement. Swap the reader. If the fault returns, swap the controller. If it returns again, swap the power supply, then the lock, then the camera. Each swap is reported as progress, and each one costs an attendance, a work order, an outage window and a piece of hardware. Replacement is a legitimate repair once a component has been shown to be at fault. Used as a way of searching for the fault, it is guessing, and the cost of the guess is carried by the site rather than by the vendor making it.
The test is straightforward. If nobody can say what the fault was, the fault was not found. A closed work order that reads “replaced reader, tested OK” records an action, not a cause. Ask what the evidence was and what it ruled out. Where there is no answer, expect the symptom back.
- Repeat attendances on one symptom. Three or more visits to the same door, panel, server or camera with no written cause on any of them.
- Parts changed with no failure explanation. A component was replaced, but nothing on the record describes how it was shown to be faulty.
- “We rebooted it and it cleared.” A restart that restores service tells you the running state was corrupt. It says nothing about what corrupted it, and it will corrupt again.
- Faults that follow a schedule. Overnight, every Monday, during the backup window, at shift change, after patching. A pattern in time is one of the strongest clues available and is usually the one left unexamined.
- Blame moving between trades. The security vendor points at the network, the network team points at the panel, and nobody owns the measurement that would settle it.
- A workaround that has become procedure. Staff who badge twice out of habit, or prop a door, or check a camera manually every morning. The site has absorbed a fault instead of having it fixed.
Faults we are called in to diagnose
Calls almost always arrive as a symptom rather than a diagnosis. These are the symptom groups we see most often in critical environments, and what tends to sit behind them.
Intermittent and time-linked faults
The hardest class of fault to pin down and the one most often mis-repaired, because the symptom is usually gone by the time anyone is standing in front of the device.
- Doors that fail to release, or fail to secure, a few times a month
- Readers that stop responding until power is cycled at the panel
- Faults that appear overnight, on weekends, or inside a backup or patching window
- Symptoms that clear before a technician arrives and return after they leave
- Marginal conditions: voltage drop under load, cable runs near their limit, temperature swings, moisture ingress at exterior devices
Offline devices and communication failures
A device shown as offline is a report, not a cause. The cause is on the bus, the cable, the power supply, the network path or the controller itself.
- Controllers, panels and downstream modules showing offline at the head end
- RS-485 and OSDP bus faults: termination, polarity, addressing conflicts, run length, mixed baud rates, secure channel key mismatch
- Modules dropping while the parent controller stays online, or the reverse
- Devices that recover on their own and drop again hours later
- Network path changes made without notice: VLAN moves, IP re-addressing, DHCP lease behaviour, port security, spanning-tree events, new firewall rules
Server and application faults
Head-end faults are often mistaken for field faults, because the symptom shows up at a door or a camera while the cause sits on the server.
- Services that fail to start after a reboot, or stop silently while the console still looks normal
- Client workstations that cannot connect, or lose connection under load
- Database growth, transaction log growth, failed archiving and purge jobs, index degradation
- Licence and certificate expiry causing partial loss of function rather than an outright stop
- Windows updates, antivirus and endpoint agents interfering with security application services and their file paths
- Delayed alarms, slow event display, backlogged event queues under peak card traffic
Door, reader and field hardware faults
Door faults are rarely one component. They are usually a disagreement between the lock, the position switch, the request-to-exit device and what the controller has been programmed to expect.
- Doors reporting forced or held open when they are neither
- Credentials that work at one reader and not at another
- Request-to-exit, door position and lock feedback contradicting each other
- Electric strikes, mag locks and mortise locks that pass a bench test and fail under real traffic
- Interlocks, mantraps and vestibule sequencing that has drifted away from the operational requirement
- Readers, keypads and biometric devices with intermittent read failures at one location only
Recording and video faults
Video failures are the quietest failures in the building. Live display keeps working while the recording path has stopped, and nothing tells the operator.
- Cameras recording intermittently, or not recording at all, with no alarm raised
- Retention shorter than the configured policy
- Gaps in recorded footage that match no known outage
- Streams that display live but never write to disk
- Archiver throughput, storage and network bandwidth limits reached without an alert
- Motion, analytics or event-driven recording that no longer triggers after a change
Integration and cross-system faults
Integrations break on one side of the link and fail silently on the other. The systems keep running, and the connection between them stops meaning anything.
- Access control to video links that no longer pull the correct camera to an alarm
- Intercom, intrusion, elevator and building system interfaces that stopped passing events
- Directory or HR synchronization failing quietly and leaving stale or orphaned cardholder records
- Interfaces broken by a version upgrade applied to only one system
- Duplicated, delayed or missing events between two systems that both report healthy
Faults previous attempts did not resolve
We are often the second or third company to look at a problem. That is a normal starting point and it changes the approach: the first task is to work out what has already been ruled out and what only appears to have been.
- Repeat calls closed with an action recorded and no cause recorded
- Systems where several components have been replaced and the symptom is unchanged
- Faults attributed to the network with no port-level, packet-level or timing evidence behind the claim
- Sites running a documented workaround in place of a repair
- Systems installed, programmed or previously serviced by another company
How a diagnosis runs
The order matters. Each step either narrows the fault or is discarded, and no component is replaced before the evidence implicates it.
- 01 Reproduce or characterize the fault — Where the fault can be forced, we force it and hold the conditions steady so it can be observed directly. Where it cannot, we characterize it instead: exact times and dates, frequency, which devices, which credentials, which operators, what else was running on site. An intermittent fault that cannot be reproduced on demand can still be bounded, and a bounded fault is a diagnosable fault.
- 02 Gather evidence before changing anything — Access control event history, controller and panel diagnostics, archiver and recording logs, Windows event logs, database and backup job history, switch port counters and error rates, power and UPS records. Evidence is collected first so there is a baseline to compare against. A system that has been reconfigured before it was examined has lost the record of its own failure.
- 03 Build a timeline — Occurrences are aligned against changes: firmware and software updates, Windows patching, network changes, added devices, configuration edits, construction on site, seasonal and temperature patterns. Correlation is not proof, but it discards a large number of hypotheses quickly and it often puts a date on when the system was last known good.
- 04 Isolate the layer — The fault is placed at a layer: field device, cabling and termination, power, network, controller, application, database or server platform. This is done by measurement and by controlled substitution at one layer at a time. It is not done by replacing parts in sequence and waiting to see whether the symptom stops.
- 05 Form one hypothesis — A single statement that accounts for every observed symptom, including the ones that do not fit the convenient answer. A hypothesis that explains four symptoms out of five is not the cause. It is the next thing to disprove.
- 06 Test it in a way that can fail — A real test predicts a result in advance and is capable of coming out wrong. If the hypothesis is a marginal RS-485 run, the test loads or shortens that run and states what should happen. If the prediction fails, the hypothesis is wrong and the work returns to the previous step rather than proceeding on hope.
- 07 Correct the root cause — Repair, reterminate, correct the addressing or the network path, rebuild the service, repair or reindex the database, restore the configuration, or replace the component that has now been shown to be at fault. Replacement still happens. It happens at the end of the process, with a reason attached to it.
- 08 Verify under real conditions — The correction is tested under load and over time rather than once: full card reads at the affected doors, recording verified to disk and against the retention policy, controllers watched through the window in which they previously failed, reconnection and failover exercised where the platform supports it and the site allows a window.
- 09 Document the cause — A written record of the symptom, the evidence gathered, what was ruled out, the cause, the correction and the verification. It is what lets the site hold a vendor to account, what lets the next technician start informed, and what makes a repeat of the same symptom significant rather than mysterious.
Isolating the layer
An enterprise security system is a stack, and symptoms seldom appear at the layer that is actually failing. A door that will not release can be a lock, a relay, a bus, a controller, a rule, a database or a clock. Each layer is tested on its own terms.
Field devices and door hardware
- Reader and keypad behaviour under repeated reads
- Lock release and re-secure under real door traffic
- Door position, request-to-exit and lock feedback agreement
- Physical alignment, wear and strike engagement
- Tamper and enclosure condition
Cabling, termination and power
- Termination quality at both ends of an existing run
- Continuity, shorts, shield and ground reference
- Voltage measured under load rather than at rest
- Standby battery condition and behaviour on transfer
- Shared power supplies loaded beyond their rating
Network
- Switch port errors, discards, duplex and speed mismatches
- VLAN, routing and firewall changes made since the system was last healthy
- Latency, loss and jitter on links to remote buildings
- Multicast, IGMP and bandwidth behaviour on video networks
- DHCP, DNS and time source behaviour affecting device registration
Controllers and panels
- Firmware version against known defects affecting the site
- Memory, database and cardholder capacity at the controller
- Offline mode behaviour and buffered event handling
- Bus loading, addressing and secure channel status
- Communication error and retry counters over time
Application and database
- Service state, dependencies and start-up behaviour
- Event queue depth and processing delay under peak load
- Database size, transaction logs, archiving and index health
- Access rules, schedules, clearances and privilege inheritance
- Licence state and certificate validity
Server and platform
- CPU, memory, disk and storage throughput headroom
- Patch level against what the security application supports
- Antivirus and endpoint agent interference with service paths
- Time synchronization across servers, controllers and cameras
- Virtual host, storage and backup interaction with a live security workload
What we take on, and what we do not
Skyfal is a service company. We do not sell cameras, panels, readers or recorders, so there is no product line to defend when we tell you what is wrong. Where a part is genuinely needed, it is specified, sourced and fitted as part of the repair, and the reason it was needed is written down.
We do not install new cabling, drill, run conduit or perform construction wiring. We work with the infrastructure already in place: existing runs are tested, reterminated and corrected where they are at fault, and where a run has to be replaced outright we scope it and coordinate with the trade doing the physical work.
We make no claim of manufacturer authorization, certification, dealership or partner status on any platform. What stands behind the work is more than 12 years of hands-on service on these systems in live critical environments. See systems we support.
After the repair
Restoration means the system is trusted again
A fault that has been running for months usually leaves more behind it than the original defect. Workarounds get built. Configuration gets changed by people trying to make the symptom stop. Devices get disabled and never re-enabled. Documentation stops matching reality. Correcting the cause is the first part of the job. Putting the system back into a state that operators and auditors can rely on is the rest of it.
Where the diagnosis exposes conditions that will cause the next failure, they are reported rather than left for someone else to find. Those items feed naturally into preventive maintenance and, where the system is at the end of its supportable life, into upgrades and migrations.
- Workarounds removed. Temporary rules, disabled devices, propped doors and manual checks are identified and closed out or documented as deliberate.
- Configuration corrected and backed up. The working configuration is captured and stored so the state is recoverable rather than reconstructed from memory.
- Verification recorded. What was tested, under what conditions, and what the result was.
- Remaining risks listed. Conditions found during diagnosis that were not the cause of this fault but will cause the next one.
- Handover to your team. The cause explained in terms your security staff and your IT group can both act on.
Common questions
A door at our site fails about once a week. It has been attended three times and it still happens. Where do you start?
With the history rather than the door. Three attendances have generated event records, work order notes and probably some replaced hardware, and together those tell us what has already been eliminated and what only appears to have been. We pull the access control event history for that door over the full period and look at what the system recorded each time it failed: whether the controller issued a release, whether the door position switch reported a change, whether request-to-exit fired, whether the event even reached the server.
That alone usually separates the three broad possibilities. If the controller never issued the release, the fault is in the rule, the schedule, the credential or the communication path. If it issued the release and the door did not open, the fault is in the relay, the wiring, the power under load or the lock. If the door opened and the system did not record it correctly, the fault is in the position feedback or the event path.
A weekly frequency is also useful in itself. It suggests something scheduled, something that accumulates, or a specific recurring condition on site, and that narrows the search before anyone opens an enclosure.
Our panels drop offline overnight and are back before staff arrive. Is that a panel problem or a network problem?
It is answerable, and answering it is the whole job. Overnight recovery without intervention points away from hardware failure and towards something that runs, changes or reaches a threshold at that time. Common causes include a nightly backup saturating a link, a switch or firewall applying a schedule, DHCP leases expiring, a maintenance or patch window on the server, a scheduled task on the security application, a UPS self-test, or building HVAC affecting an enclosure in an unconditioned space.
The evidence needed is on both sides. From the security system: the exact offline and online timestamps for each panel over several weeks, and whether they drop together or in sequence. From the network: switch port counters, error and discard rates, and logs for the relevant window. When panels on separate switches drop simultaneously, the cause is upstream or on the server. When one panel drops alone, the cause is local to that path.
Until that comparison is made, “panel problem or network problem” is an opinion. After it, it is a measurement.
What happens if the fault does not occur while you are on site?
That is expected on intermittent faults and it does not stop the diagnosis. The system has been recording while we were not there, and that record is often better than a live observation because it covers weeks rather than hours.
Where the recorded evidence is not enough, we instrument the system and wait: increased logging on the affected devices, monitoring on the relevant switch ports, capture of controller diagnostics, and where appropriate a short-term recording of the specific conditions we suspect. We agree in advance what would confirm and what would rule out the hypothesis, so the next occurrence produces an answer rather than another attendance.
What we do not do is replace a component so that the visit ends with something having been done. If the fault has not been identified, the honest report is that the fault has not yet been identified and here is what the next occurrence will tell us.
What do you need from our IT team to diagnose a network-layer fault?
Usually less than people expect, and never unsupervised access. The most useful items are switch port statistics for the ports serving the security devices, including error, discard and reset counters, plus the switch and firewall logs covering the times the fault occurred. Alongside that: the VLAN and IP addressing used by the security system, whether addresses are static or reserved, the routing and firewall rules between controllers, servers and clients, and a change record for anything altered on the network since the system was last known to be working.
For video, we also need to understand available bandwidth on the links carrying camera traffic, any QoS or rate limiting in place, and multicast configuration where it is used.
Work can be done alongside your team on a screen share, with your staff running the commands and holding the credentials. Many sites in correctional, government and healthcare settings require exactly that, and it does not slow the diagnosis. See ongoing technical support for how that working relationship is normally set up.
Will you work on a system another company installed?
Yes. Most of our restoration work is on systems we did not install, and a large share is on systems that are still under someone else’s support arrangement. We are not there to displace an installer, and we do not require the site to change vendors.
What we ask for is access to the evidence: event history, current configuration, whatever documentation exists, and permission to make the corrections the diagnosis calls for. Where a manufacturer warranty applies to a component, we will say so rather than opening it, and where the correction is properly the installing contractor’s responsibility under an existing contract, that goes in the report too.
We have been told the only fix is to replace the system. Is that ever true?
Sometimes. Hardware reaches end of support, firmware stops being maintained, a controller generation stops receiving fixes, and a platform version eventually stops running on any operating system that a security team will accept. Those are real limits and they justify replacement.
The distinction is whether the recommendation follows a diagnosis or replaces one. “Replace the system” offered without a documented cause is a way of ending an investigation that was not going anywhere. Ask which specific fault cannot be corrected on the current platform and why. If that question has a clear technical answer, the recommendation is sound and the next conversation is about migration planning. If it does not, the fault is still undiagnosed and replacing the equipment may simply carry it into a new system.
Bring us the fault that keeps coming back
Send the symptom, the history and what has already been tried. We will tell you what evidence is needed to find the cause.