As part of Pacific Crest’s Mosaic Expert team, I had the opportunity to attend their annual Technology Leadership Forum in Vail last month. I participated in half-a-dozen panels and was fortunate to meet with several contributors in the technology research and investment arena. Three things seemed to rank high on everyone’s agenda: cloud computing and its twin enablers - virtualization and data center automation. The cloud juggernaut is making everyone want a piece of the action – investors want to invest in the next big cloud (pun intended!), researchers want to learn about it and CIOs would like to know when and how to best leverage it.
Interestingly, even “old-world” hosting vendors like Savvis and Rackspace are repurposing their capabilities to become cloud computing providers. In a similar vein InformationWeek recently reported some of the telecom behemoths like AT&T and Verizon with excess data center capacity have jumped into the fray with Synaptic Hosting and Computing as a Service - their respective cloud offerings. And to add to the mix, terms such as private clouds are floating around to refer to organizations that are applying SOA concepts to data center management making server, storage and application resources available as a service for users, project teams and other IT customers to leverage (complete with resource metering and billing) – all behind the corporate firewall.
As already stated in numerous publications, there are obvious concerns around data security, compliance, performance and uptime predictability. But the real question seems to be: what makes an effective cloud provider?
Google’s Dave Girourad was a keynote presenter at Pacific Crest and he touched upon some of the challenges facing Google as they opened up their Google Apps offering in the cloud. In spite of pouring hundreds of millions of dollars on cloud infrastructure, they are still grappling with stability concerns. It appears that size of the company and type of cloud (public or private) is less relevant, and more relevant is the technology components and corresponding administrative capabilities behind the cloud architecture.
Take another example: Amazon. They are one of the earliest entrants to cloud clouding and have the broadest portfolio of services in this space. Their AWS (Amazon Web Services) offering includes storage, queuing, database and a payment gateway in addition to core computing resources. Similar to Google, they have invested millions of dollars, yet are prone to outages.
In my opinion, while concerns over privacy, compliance and data security are legitimate and will always remain, the immediate issue is around scalability and predictability of performance and uptime. Clouds are being touted as a good way for smaller businesses and startups to gain resources, as well as for businesses with cyclical resource needs (e.g., retail) to gain incremental resources at short notice. I believe the current crop of larger cloud computing providers such as Amazon, Microsoft and Google can do a way better job with compliance and data security than the average startup/small business. (Sure, users and CIOs need to weigh their individual risk versus upside prior to using a particular cloud provider.) However for those businesses that rely on the cloud for their bread-and-butter operations whether cyclical or around-the-year, uptime and performance considerations are crucial. If the service is not up, they don’t have a business.
Providing predictable uptime and performance always boils down to a handful of areas. If provisioned and managed correctly, cloud computing has the potential to be used as the basis for real-time business (rather than being relegated to the status of backup/DR infrastructure.) But the key question that CIOs need to ask their vendors is: what is behind the so-called cloud architecture? How stable is that technology? How many moving parts does it have? Can the vendor provide component-level SLA and visibility? As providers like AT&T and Verizon enter the fray, they can learn a lot from Amazon and Google’s recent snafus and leverage technologies that can simplify the environment enabling it to operate in lights-out mode – making the difference behind a reliable cloud offering and one that’s prone to failures.
The challenge however, as Om Malik points out on his GigaOm blog, is that much of cloud computing infrastructure is fragile because providers are still using technologies built for a much less strenuous web. Data centers are still being managed with a significant amount of manual labor. “Standards” merely imply processes documented across reams of paper and plugged into Sharepoint-type portals. No doubt, people are trained to use these standards. But documentation and training doesn’t always account for those operators being plain forgetful, or even sick, on vacation or leaving the company and being replaced (temporarily or permanently) with other people who may not have the same operating context within the environment. Analyst studies frequently refer to the fact that over 80% of outages are due to human errors.
The problem is, many providers while issuing weekly press releases proclaiming their new cloud capabilities, haven’t really transitioned their data center management from manual to automated. They may have embraced virtualization technologies like VMware and Hyper-V, but they are still grappling with the same old methods combined with some very hard-working and talented people. Virtualization makes deployment fast and easy, but it also significantly increases the workload for the team that’s managing that new asset behind the scenes. Because virtual components are so much easier to deploy, it results in server and application sprawl and demands for work activities such as maintenance, compliance, security, incident management and service request management go through the roof. Companies (including the well-funded cloud providers) do not have the luxury of indefinitely adding head-count, nor is throwing more bodies at the problem always a good idea. They need to examine each layer in the IT stack and evaluate it for cloud readiness. They need to leverage the right technology to manage that asset throughout its lifecycle in lights-out mode – right from provisioning to upgrades and migrations, and everything in between.
That’s where data center automation comes in. Data center automation technologies have been around now for almost as long as virtualization and are proven to have the kind of maturity required for reliable lights-out automation. Data center automation products from companies such as HP (on the server, storage and network levels) and Stratavia (on the server, database and application levels) make a compelling case for marrying both physical and virtual assets behind the cloud with automation to enable dynamic provisioning and post-provisioning life-cycle management with reduced errors and stress on human operators.
Data center automation is a vital component of cloud computing enablement. Unfortunately, service providers (internal or external) that make the leap from antiquated assets to virtualization to the cloud without proper planning and deployment of automation technologies tend to provide patchy services giving a bad name to the cloud model. Think about it… Why can some providers offer dynamic provisioning and real-time error/incident remediation in the cloud, while others can’t? How can some providers be agile in getting assets online and keeping them healthy, while others falter (or don’t even talk about it)? Why do some providers do a great job with offering server cycles or storage space in the cloud, but a lousy job with databases and applications? The difference is, well-designed and well-implemented data center automation - at every layer across the infrastructure stack.
Wednesday, September 03, 2008
Clouds, Private Clouds and Data Center Automation
Posted by
Venkat Devraj
at
7:58 AM
2
comments
Labels: Cloud Computing, data center automation, PacificCrest
Sunday, September 02, 2007
Taking a Stab at a Shared IT Industry Definition of "Data Center Automation"
The problem with certain grandiose terms such as "IT automation" and "data center automation" is that they have no shared definition across vendors in the IT industry. The only thing that's common is their repeated reference by multiple sources, all in different contexts and scope. They become part of the hype vernacular generated by different vendors and their marketing machines and eventually a word’s true meaning becomes irrelevant. In such a state, everyone thinks they know what it is, but no one really does and alas, it becomes so ubiquitous that people don't even bother challenging each others' assumptions regarding its scope.
The term “automation” is rapidly free-falling into just such a state in the IT industry. So here's me taking a stab at level-setting the meaning and scope for automation that (in my humble opinion) is capable of bringing both customers and vendors to the same playing field. Even if not, if it generates some cross-vendor discussion and allows people to challenge each other's perception of what data center automation ought to mean, my purpose would be served.
In a prior blog entry, I refer to 15 specific levels of requirements that need to be addressed for any IT automation solution to be effective. So ladies and gents, here’s that requirements stack. If these requirements are satisfied, you would have reached a state of automation nirvana and somewhere along the way, you would have imbibed (and likely, surpassed) the isolated automation capabilities touted by most IT tools vendors.
The accompanying picture shows Levels 1 to 8 as a pre-requisite to automation, and works itself all the way up to level 15 (autonomics) via a model that fosters shared intelligence across disparate capabilities and functions. It is important to get the context right with each of these levels before assuming one has attained them. And until one has got a particular level right, it is often futile to attempt to go to the next level. Short-cuts have an uncanny way of short-circuiting the process. That's why you find so many vendors and organizations with fragmented notions of these requirements just not cutting it in real-world automation deployments. Specifically because such approaches lack a sound methodology to build the proper foundation.
Nuff said. Let’s look at the individual levels now and how they build on one another.
Level 1 pertains to achieving 360-degree monitoring. This capability breaks through the typical siloed monitoring that exists today in many environments and allows problems to be viewed across multiple tiers, applications and service stacks in a cohesive manner –preferably, in the same call sequence utilized by the end-user application. Current monitoring deployments often remind me of the old fable of the seven blind men and the elephant (where each blind man would feel a different part of the elephant and perceive the animal to resemble a familiar item. For instance, one would touch the elephant’s tail and spread the notion that the animal looks like a rope, whereas another person would feel the elephant’s foot and try to convince everyone that the animal resembled a pillar). Lack of a comprehensive and consistent view of the same issue by multiple individuals and groups cause more delays in solving problems than the typical IT manager would admit.
Levels 2 and 3 utilize 360-degree monitoring for proper problem diagnosis, triage and alerting . Rather than leaving the preliminary diagnosis regarding nature and origin of a problem to human hands and eyes (say, a Tier 1 or Help Desk team), the monitoring software should be able to examine the problem end-to-end, carry out sufficient root cause analysis, narrow down the scope based on the analysis and send a ticket to the right silo/individual. Such precision alerting reduces the need for manual decision-making (and chances of error) regarding which team to assign a ticket to and positively impacts metrics such as first-time-right and on-time-delivery.
Level 4 pertains to ad-hoc tasks that administrators often do, such as adding a user or changing the configuration parameters of an application. There are lots of popular tools and point solutions for Levels 1 to 4 ranging from monitoring tools to ad-hoc task GUIs. However their functionality really ends there (or they try to jump all the way from Level 4 to Level 9 (Automation), skipping the steps in between and causing the resultant automation to be of rather limited use.)
Level 5 questions the premise of an “ad-hoc task”. Wisdom from the trenches often tells us that there is no such thing a one-time / ad-hoc task. Everything ends up being repeatable. For instance, when creating a user on one database, one may bring up her favorite ad-hoc task GUI, click here, click there, type in a bunch of command attributes and then hit the Execute button. That works for creating one user. However when the exact same user needs to be rolled out on 20 servers, it involves a lot of pointing, clicking and typing and leaves the environment vulnerable to human errors. Suddenly the ad-hoc GUI ceases to be effective.
Requirement levels 5 to 9 address this by calling for standard operating procedures (SOPs) that can be applied across multiple environments from a central location. Level 6 requires diverse sub-environments to be categorized such that different activities pertaining to them (including all service requests and incident remediation efforts) can all have standard task recipes. This refers to rolling up disparate physical environments into fewer logical components based on policies and usage attributes including service level requirements.
Task recipes, once defined, need to be maintained in a central knowledgebase and need to be in a format wherein they can easily serve as a blue-print for any subsequent automation. Further, the task recipes need to be directly linked to automation routines/workflows in a 1:1 manner such that one can reach the workflows from the SOP and vice versa. In other words, there needs to be a shared intelligence between the SOP and the automation routine. Keeping the two separate (for instance, keeping the task recipes on a SharePoint portal and keeping the automation routines within a scheduling or workflow tool, or worse, as a set of scripts distributed locally on the target servers) without any hard-wired connection and tracking between the two mediums will make it easy for task recipes and automation routines to get out of sync. Such automation tends to be uncontrolled; each user of such automation is left to his/her own devices to leverage it as optimally as possible, and eventually, its utility becomes questionable at best with each user (administrator) customizing the automation routines to their particular environment and their individual preferences. Hitherto noble notions such as shared intelligence and centralized control across task recipes and automation routines (and eventually, consistency in quality of work) go out the window!
With shared intelligence comes the ability to track and enforce standard task recipes across different personnel and environments. Level 10 states this exact requirement - the ability to maintain a centralized audit trail describing when an automation routine ran, on what server, who ran it (or what event triggered it), what were the run-results and so on – for ongoing assessment of SOP quality, pruning the task recipes and corresponding automation code, and finally, to ensure enterprise-wide adherence to best practices.
Level 11 or virtualization allows automation to be applied in a easier manner across the different environment categories defined in level 6. It does so via a “hypervisor” that abstracts and masks the different nuances across multiple categories by having a SOP call wrapper applying a task across different environment types (ideally, done via an expert system). Within that single SOP call, auto-discovery capabilities identify the current state of a target server or application, evoke a decision tree to determine how to best perform a task and then finally, call the appropriate SOP (or sub-SOP, as the case may be) to carry out the task in the pre-defined and pre-approved manner that’s most suited to that environment. In other words, level 11, much like storage virtualization that occurs within a SAN (think EMC Invista, not VMWare), calls for multiple disparate environments to be viewed as a single environment (or at the very least, a smaller subset of environments) and dealt with in an easier manner.
Levels 12 to 15 allow increasingly sophisticated levels of analysis and intelligence to be applied to a target environment, including event correlation, root cause analysis, and predictive analytics to be able to discern problems as they occur or ideally, before they occur and link those back to one or more SOPs which can be automatically triggered to avert an outage, performance degradation or policy violation and retain status quo, thereby making the target environment more autonomic.
Granted that many companies and administrators may not have the appetite to go all the way up to level 15 for most administrative tasks. But regardless, I would hope that this model provides a clear(er) roadmap for automation than just different groups and individuals scrapping together a bunch of proprietary scripts and Word documents, and keeping relevant information on how to exactly execute them within their heads or worse, each group relying on a bunch of disjointed tools, vendors and promises, resulting in a fragmented automation strategy and chaotic (read, unmeasurable) results.
Maybe unmeasurable results are not such a bad thing, especially if you are a large and well-established software vendor, who has little value to deliver, but thrives on FUD to keep the market confused.
If you are an IT manager or administrator and have a favorite software tools vendor, ask them how their automation strategy stacks up against these 15 levels. Drop me a note if you get an answer back.
Posted by
Venkat Devraj
at
3:15 AM
2
comments
Labels: data center automation, IT automation, virtualization