Short answer. The SLA makes maintenance measurable: it lists equipment, categorizes incidents by criticality, differentiates response and recovery times, sets support hours, remote diagnostics, on-site visits, preventive maintenance, spare parts, reports, and the responsibilities of both parties.
Promise and service level are not the same
The phrase "we will respond promptly" does not answer when the engineer will confirm the request, when diagnostics will start, and when the function will be restored. SLA — service level agreement — translates expectations into measurable conditions. It can be part of a contract and a service regulation.
The goal of the SLA is not to punish the performer with a penalty but to agree in advance on what the facility considers a failure and what resources are ready from both sides. If the customer does not provide access to the site or remote connection, even the strictest deadline will not be realistic.
First create a hardware inventory
The restoration of a system whose composition is unknown cannot be promised. The registry records the model, serial number, installation location, commissioning date, firmware, network address, warranty, criticality, and connection to other devices. For foreign or old equipment, an incoming diagnosis is performed.
Schematics, configuration backups, credentials in a secure storage, and repair history are collected simultaneously. This reduces dependence on a single specialist and speeds up troubleshooting, especially when a failure occurs at the intersection of automation, network, and power.
- Device and its exact location
- Customer Function Owner
- Severity and Tolerance of Idle Time
- Related systems and available spare parts
Divide incidents by site impact
The failure of a single blocker at the main entrance and the loss of one overview camera should not have the same priority. P1 level can be used for a complete stop of a critical function or a safety threat, P2 — for a serious degradation with temporary bypass, P3 — for a local malfunction that does not stop the process.
Definitions should be specific. For example, 'the archive is unavailable for all cameras' is clearer than 'a serious problem with video.' For each level, they specify the hours for receiving requests, confirmation time, start of diagnostics, the need for on-site visit, frequency of updates, and the recovery goal.
Do not confuse reaction and recovery
Reaction time is the moment when the service registered the incident and started work. It does not mean that the system is already functional. Recovery time depends on diagnostics, access, part availability, delivery, programming, and construction work. If the supplier promises to fix any failure equally quickly, ask them to describe the assumptions.
For critical nodes, it is useful to set a temporary function recovery: manual gate control, backup logger, spare power supply, or alternative inspection line. Full repair can be completed later if the facility continues to operate safely.
Maintenance should have a measurable composition
The formulation "perform maintenance" is too broad. The regulations list operations and frequency: inspection of fastenings, cleaning, checking drives and leaks, testing safety sensors, measuring batteries, log analysis, monitoring free space, trial video export, and checking the emergency mode.
The result is documented in an act with metrics, defect photos, and recommendations regarding timing. Good preventive maintenance detects degradation before failure: rising drive current, reduced battery capacity, disk errors, or an unstable network port.
Spare parts are selected based on the risk of downtime
Storing all equipment on-site is expensive, but having nothing is risky. For each critical function, a minimal set is defined: fuses, sensors, power supplies, batteries, boards, drives, actuators, or a ready-made backup module. Lead times and the possibility of compatible replacement are taken into account.
The reserve must be accounted for and verified. A battery ages on the shelf, the board firmware may not match the system, and the drive may not be supported by the recorder. The SLA records the reserve owner, storage location, and procedure for issuance and replenishment after use.
The customer's responsibility is also recorded
The customer appoints contact persons, ensures passes, safe access to equipment, up-to-date diagrams, and the possibility of system shutdown for testing. For remote diagnostics, a protected connection method and action logging are agreed in advance. Passwords must not be sent in plain text or left without an owner.
The contractor determines the contact channels, escalation, composition of the on-duty team, scope of work, and procedure for involving the manufacturer. Exceptions are listed separately: power outages in the building, damage by third parties, construction work, lack of communication, or equipment without documentation.
Reporting shows what the facility is paying for
Monthly or quarterly reports may include the number of incidents, deadlines met, repeat failures, maintenance performed, parts used, and risks for the next period. This data allows decisions about what to repair, what to modernize, and what inventory to maintain.
A useful indicator is the availability of a critical function, not just the number of closed requests. If the same sensor is replaced three times, the requests are formally closed, but the cause is not eliminated. Analysis of repeats helps move from emergency on-site visits to equipment condition management.
Minimum to include in the SLA
You can start with a one-page table: function, equipment composition, criticality, support hours, response, diagnostics, temporary and full recovery, spare parts, and responsible personnel. Then preventive maintenance regulations and the report form are added.
Before signing, simulate one real failure. Who calls? What do they report? How does the engineer gain access? Is there a diagram and a spare part? Who authorizes manual mode? Such a check quickly reveals conditions that look good in the contract but don’t work on site.
- List of serviced systems and exceptions
- Three to four severity levels with examples
- Reaction, status update, and restoration goal
- Prevention, Spare Parts and Reporting
- Contacts, escalation, and responsibilities of both parties
PRIMARY SOURCES AND METHODS
The material was written by KRONIS GROUP in their own words. The links below are needed so that the customer can verify the original requirements and professional recommendations.
The material is for informational purposes. Final requirements for a specific facility, project, and permitting documents are determined after inspection and review of applicable Kazakh standards.