| |

The Cloud at Scale

 February 14, 2012, Cloud Connect Conference, Santa Clara, CA—Jesse Robbins from OpSource described the efforts and requirements to operate in the cloud at scale. The systems need to be robust, and the operators need to be prepared for failures.

As operations and software grow in scale, all aspects of the operations increase in complexity, and latent effects become acceptable. These conditions contribute to catastrophic failures, because total availability and reliability is the multiplicative product of all the contributing subsystems. For example, three components each with 99.9 percent availability, together result in a system with 99.7 percent availability.

When things break and systems fail, people respond through a progression of emotions. First, they deny the failure, “this can’t be happening”, then they get angry. Next, they hunker down, followed by some level of depression, and finally accept the fact that the system failed. At this point, they start doing something about the failure.

Instead of developing systems in an ad hoc manner, people need to start designing systems for failure. The people need to have a training day where large-scale faults are injected into the system to get exposure to various failure modes and to identify the resilience layers in both people and technology.

The necessary ingredients for quick and successful responses to failure are preparation, participation, and exposure. Performing training day functions by starting small and increasing scale and complexity of the problems raises awareness and builds confidence in those people task with responding to failures. Over time and through repetition, you build up to a full system failure in a “live fire” exercise. The results of these exercises lead to operating safety standards and expertise. Applying the lessons learned in these exercises builds up trust cases which are needed because there are no perfect systems.

Systems operating with normal processes move towards greater automation, and the infrastructure gets treated like code, that’s written, integrated, and deployed. This leads to believing that there is a golden state of operation, but when changes occur they multiply causing search for discrepancies and complexity issues to increase dramatically. With experience, companies develop better incident management practices, and drive development and apps to be more robust.

Systems have to be designed to fail over with dedicated standard deployment and management functions integrated with emergency management tools. Getting the system administrators lots of levers and knobs helps to accelerate recovery efforts. Compartmentalizing regular processes from its soft failures and minimizes system disruptions. System the Zionist take into consideration the mean time to recovery is much more important than the mean time to failure.

Similar Posts