- [Reliability](#reliability) - [Serviceability](#serviceability) - [Availability](#availability) - [Failover](#failover) - [Cold Standby](#cold-standby) - [Advantages](#advantages) - [Disadvantges](#disadvantges) - [Use Cases](#use-cases) - [Warm Standby](#warm-standby) - [Advantages](#advantages-1) - [Disadvantages](#disadvantages) - [Hot Standby](#hot-standby) - [Advantages](#advantages-2) - [Disatvanges](#disatvanges) - [Questions](#questions) - [Failover](#failover-1) # Reliability [Wikipedia](https://en.wikipedia.org/wiki/Reliability,_availability_and_serviceability): "Reliability can be defined as the probability that a system will produce correct outputs up to some given time t.[5] Reliability is enhanced by features that help to avoid, detect and repair hardware faults. A reliable system does not silently continue and deliver results that include uncorrected corrupted data." Question: can you explain the difference between Availability and Reliability? # Serviceability [Wikipedia](https://en.wikipedia.org/wiki/Reliability,_availability_and_serviceability): "Serviceability or maintainability is the simplicity and speed with which a system can be repaired or maintained; if the time to repair a failed system increases, then availability will decrease." # Availability [Wikipedia](https://en.wikipedia.org/wiki/Reliability,_availability_and_serviceability): "Availability means the probability that a system is operational at a given time, i.e. the amount of time a device is actually operating as the percentage of total time it should be operating. High-availability systems may report availability in terms of minutes or hours of downtime per year." In simpler words, the percentage of time a system, service or app is accessible, operational # Failover Failover (or failover strategy) is the process of switching upon a failure from non-operational system (application, storage, DB, ...) to an operational system or a previous operational state of the same system. It used to be done either automatically or manually. Today, especially in the era of public clouds, this process is often automated. ### Cold Standby In Cold Standby the transition to the backup/standby service happens only after the active service becomes non-operational. The process is described in the image below taking as example a database: 1. The server works against an active database 2. The database becomes non-operational for different reasons 3. There is a switch to the standby database. It's important to note that this switch might not be quick. It might involve a process of restoring the data and it will for sure involve redirecting the traffic to the new database instead of the previous one