Thursday, June 17, 2010

Are You Over Virtualizing Your Technology Infrastructure?

In a recent key note address at Gartner’s 2010 Infrastructure and Operations Summit, Raymond Paquet, Managing VP at Gartner, made the point that consistently over the last eight years 64% of the typical IT budget has been dedicated to tactical operations such as running existing infrastructure. A driving factor behind this consistent trend is the tendency of technology managers to over deliver on infrastructure services to the business. Paquet makes the point that technology managers should always be focused on providing the right balance between technology solution cost and the quality of service the solution provides to the business. As the quality of a technology service increases the cost of that service will naturally increase. There is often still a gap in the quality of service a business requires and what is architected and built by technology professionals. The key is to understand the businesses requirements in terms of parameters such as uptime, scalability and recoverability and to architect an infrastructure aligned with these requirements. Sounds simple, right? Often however technology professionals find themselves (usually inadvertently) swept up in implementing technology solutions that at specific levels of adoption add value to a business but at unnecessary levels of adoption serve to pull budgets and resources away from more value added activities. Virtualization is a case in point.

The gap between quality of technology services required by the business and the cost to deliver them can be observed today in the tenacious drive by many technology managers to virtualize 100% of their IT infrastructure. Gartner research VP Phillip Dawson points out what he calls the “The Rule of Thirds” with respect to virtualizing IT workloads. The theory is that one third of your IT workload will be “easy” to virtualize, another one third will be “moderately difficult” with the remaining third being “difficult” to virtualize. Most technology organizations today have conquered the easy one third and have began working on the moderately difficult workloads. The cost in terms of effort and dollars to virtualize the final one third of your IT workload increases exponentially. Are you really adding value to your business by driving 100% of your environment into your virtualization strategy? The recommendation from Dawson is to first simplify IT workloads as much as possible by driving towards well run x86 architectures. Once simplified, IT workloads with low business value should be evaluated for outsourcing or retirement. Retirement may consist of moving the service the application provides into an existing ERP system thus leveraging existing infrastructure. Workloads with high business value will likely now (after simplification) be cost effective to bring into your virtualization strategy.



So be selective in the applications you target for virtualization. Be sure that all steps have been taken to simplify the respective applications technology infrastructure so that it is not in the upper third of costs and effort in terms of virtualization effort. Outsource or retire complex applications where applicable rather than force fitting them into your own virtualization strategy. Look for ways to retire low business value applications with complex technology architectures (such as moving them into an existing ERP package). The bottom line is that all technology investments and initiatives must be aligned with your organizations needs and risk tolerance. Virtualization is no different. Making unnecessary efforts to virtualize certain applications may be doing more harm than good.

Saturday, June 5, 2010

Software Licensing Models Must Evolve to Match Innovation in Computing Resource Delivery

Software licensing has never been one my favorite topics. It has grabbed my attention lately however as a result of what I perceive as a gap between the evolution of infrastructure provisioning models and software vendor licensing models. As infrastructure virtualization and server consolidation continue to dominate as the main trends in computing delivery, software vendors seem bent on clinging to historical licensing models based on CPU. In today’s dynamically allocated world of processing power designed to serve SaaS models for software delivery and flexible business user capabilities, software providers still insist that we as providers of computing services be able to calculate exactly how many “processors” we need to license or exactly how many unique or simultaneous users will access our systems. In addition, many software salespeople use inconsistent and confusing language that causes confusion among business people and even among some CIO’s.

Is it a CPU, a Core or a Slot?

This is where I see the most confusion and inconsistency in language, even among software sales reps. Here is a little history of how things have gotten more complex. It used to be (in the bad old days) that one slot equaled one CPU which had one core. So, when software was licensed by the CPU it was simple to understand how many license you needed. You just counted the number of CPU’s in the server that was going to run your software and there you had it. As chip makers continued to innovate, this model started getting slightly more complex. For example, when multi-core technology was introduced it became possible to have a CPU in a slot that actually constituted multiple cores or execution units. This innovation continues today and is evident in the latest six and even eight core Nehalem and Westmere architectures. So now the question had to be asked to software sales reps “what do you mean by CPU?”

Processing Power Gets Distributed

At nearly the same time that the chips inside our servers where getting more powerful and more complex, our server landscapes where getting more complex as well. Corporate enterprises began migrating away from centralized mainframe computing to client/server delivery models. New backend servers utilizing multi-core processors began to be clustered using technologies such as Microsoft Cluster Server in order to provide high availability to critical business applications. Now the question of “how many CPU’s do you have?” became even more complex.

ERP Systems Drove Complexity in Server Landscapes

In the late 1990’s, many businesses introduced Enterprise Resource Planning (ERP) systems designed to consolidate sprawling application landscapes consisting of many different application for functions such as finance, inventory and payroll. The initial server architecture for ERP systems generally consisted of a development system, a quality system and a production system, each running on distinct computing resources. Most business users only accessed the production system while users from IT could potentially access all three systems. At this point, software sales reps began speaking of licensing your “landscape” and many ERP providers offered named user licensing. A named user could access the ERP system running on development, quality or production regardless of the fact that these where distinct systems.

ERP Technical Architecture Options Drive Confusion

ERP systems still run on databases and require appropriate database licensing. ERP packages such as SAP offer technical design options allowing the database used for the ERP development, quality and production systems to be ran separately in a one for one configuration or in a clustered environment with one or two large severs providing the processing power for all the system databases. With the database licensing needed to run the ERP system there was still a breakeven point where licensing the database by “CPU” could come out cheaper than licensing the database by named user. And so as ERP implementations continued, software salespeople now spoke to technology and business managers about “licensing their landscape by named users with options for CPU based pricing based on cores”. The fog was starting to set in. In fact these options became so confusing that when asked what the letters SAP stood for many business managers would reply “Shut up And Pay”. The situation only got worse in the early 2000’s when ERP providers like SAP saw the adoption rate of new core ERP systems begin to slow. At this point, the majority of enterprise organizations already had some sort of ERP system in place. ERP providers branched out into other supporting systems such as CRM, Supply Chain Management or Business Analytics. Now the “landscape” of applications became even more complex with named users crossing development, quality and production systems across multiple sets of business systems.

Virtualization - The Final Straw

In the 2007 timeframe, a major shift in the way computing power was allocated to business systems began to appear in enterprise data centers. Virtualization technology from companies such as VMware began to gain main stream adoption in enterprise class IT operations. Virtualization gave rise to terms such as Cloud Computing meaning that the processing power for any given business application was provided from a pooled set of computing resources, not tied specifically to any one server or set of servers. An individual virtual server now depended on virtual CPU’s which themselves did not have any necessary direct relationship to a CPU or core in the computing “Cloud”.

Despite this colossal shift in the provisioning of IT computing power, many software vendors clung to their licensing models, insisting on knowing how many “CPU’s” you would use to run their software. Even today in 2010 when Virtualization has been well vetted and proven to be a viable computing delivery model at all tiers of an enterprises technical architecture, some software vendors either insist on sticking to outdated license models or simply issuing vague statements of support for virtualization. These vague statements of support often seem to be based more on resistance to change traditional licensing models rather than on any clearly stated technical facts.
After a long and storied evolution of computing power delivery from single core, single CPU machines to Cloud Computing, it seems like software providers ranging from database providers to providers of business analytics have been slow to innovate in terms of licensing models. Enterprise class providers of business software such as SAP or Oracle who claim to have embraced virtualization yet continue to either license certain product sets by CPU or issue only vague statements of virtualization support seem to be struggling to provide innovate pricing structures aligned with the new realities of computing power delivery.

I am sure they will get there, but for now I still get emails from software sales representatives quoting prices for products “by the CPU”. Sigh…..

Friday, May 28, 2010

Your Data Center Hosting Provider is Stealing Your Money

Well, stealing is a strong word. What is happening however is that traditional data center hosting providers are getting in the way of thousands of small to medium sized businesses as it relates to realizing the true energy based cost savings associated with virtualization. Virtualization has changed many aspects of the traditional IT infrastructure and fostered innovations in all areas of traditional infrastructure service provision. What has not kept pace is innovation and investment on the part of data center hosting providers in facilities infrastructure geared towards delivering services such as power and cooling in a fashion that is aligned with the new realities of virtualized IT computing loads. This is directly limiting the ability of small to midsized companies who host their IT environment in these data centers to fully realize the total energy savings that virtualization can provide. In the following paragraphs I will provide a summary of APC white Paper 118 which does an excellent job of explaining how the total energy savings from virtualization is dependent on a realignment of data center infrastructure to meet the needs of a reduced and consolidated IT load. I will point out exactly where small to midsized organizations hosting there IT load in traditional hosting providers facilities are leaving money on the table as a result of their hosting providers in-action.

The following diagram demonstrates the primary sources of energy consumption in a data center. The support power represents energy that is lost due to the inefficiency of data center physical infrastructure such as power and cooling systems. This is energy that is consumed in the operation of the equipment itself rather than being transferred to the IT load and being used for useful computing work.

Data Center Power Sources

After you virtualize your server environment, your IT load will decrease. This decrease will make the PUE of the data center worse due to inefficiencies caused by a physical infrastructure continuing to operate at what is now over capacity for the new virtualized IT load.

PUE Decreases after Virtualization

So while a decrease in energy cost due to IT load consolidation and virtualization is certainly positive, it is only a fraction of the overall savings possible. The total energy savings made possible by virtualization can only be achieved if the data center physical infrastructure is re-architected to be more allinged with the new realities of virtualized IT loads.

Aditional Gains from PUR Optimization

Specific recommendations for changes to data center physical infrastructure to achieve a closer alignment of physical infrastructure services with virtual IT loads can be found in APC white paper 126.

In all fairness to hosting providers, realizing some of the efficiency gains of re-architecting physical infrastructure is a true challenge in an environment inherently designed to provide shared service across many organizations with unique IT loads, peak demand periods and degrees of virtualization and consolidation. Looking from the point of view of a mid-sized IT organization that has diligently virtualized and reduced IT load requirements only to find themselves “trapped” by existing power circuit contracts or by an inflexible hosting provider who has not invested in physical infrastructure innovation reveals a logical degree of frustration. In any market such as data center hosting where the barrier to entry is high due to large capital expenditure requirements, innovation by market leaders tends to be slow. What is needed is a new type of hosting provider built from the ground up to provide modern, flexible solutions such as “pay by the drink” for power services and individualized cooling solutions through innovations like row based cooling while maintaining independence from the overall environment of the data center.

Data Center Zones

Should traditional hosting providers fail to make these innovations then no doubt, a new breed of more agile competitors, unburdened by large historical capital investments in dated infrastructure will emerge and force fundamental change in the hosting industry. Mid-sized customers may also find the additional value hosting services provide such as physical security to longer be enough to prevent them from investing in their own facilities where they can innovate themselves and keep all the gains. Virtualization has changed almost everything with respect to IT infrastructure service delivery, it is time for hosting providers to catch up.

Sunday, May 23, 2010

A Simple Offshoring Model

I recently had the opportunity to travel to India to meet with a large provider of IT services. The purpose of the trip was to understand how best to work with this organization to deliver additional resources and thus business results to my organization. Some of the questions that needed answering where: When does it make sense to utilize offshore IT resources?, What is the best way to run projects using offshore service providers?, What are the criteria for deciding if a particular project is a fit for the utilization of offshore resources? What follows is a model I created based on my observations and peer discussions. This model has in it some implicit lessons learned from this particular trip as well as previous project experiences. The model is fairly self describing. The underlying principals are that projects with a high dependence on institutional knowledge require greater in-house resource involvement whereas projects that involve more standard business processes or technology are better fits for the utilization of offshore resources. Also, larger projects lend themselves better to the utilization of offshore resources due to the inherent overhead involved in managing offshore resources. Most IT managers will likely find this a common sense line of reasoning, the model simply provides a simple and clear representation of this logic.

offshoring model

Sunday, April 18, 2010

You Don't Have to be a Giant to be Green

Last week I was fortunate enough to attend SAP Virtualization Week at the SAP Co-Innovation Labs in Palo Alto, California. The conference consisted of three tracks: Virtualization, Cloud Computing and Green IT. Lots of thought provoking material was presented by SAP, consulting partners and customers. When I attend these type of events I look hard to find a few practical take aways that could be put in place immediately. Of course these sessions help frame my thinking on long-term strategies around SAP virtualization and Data Center design, but what can I go home with and ask my team about next week? One of those take away points for me last week was this: You don’t have to be a giant to be green.

At surface level Green IT seems like a concept relative only to the largest of IT organizations. Organizations such as Intel and Colgate-Palmolive who can tell stories of collapsing 50 or 60 global data centers down to one facility certainly have a green story to tell. Organizations such as NetApp who have pioneered efficient data center designs can go to conferences to show how they received a million dollars in rebates from their energy provider. These are great stories, but as I sat through the first couple of these last week sipping Starbucks coffee I was thinking to myself “This is awesome, but what can I do, how is this relative to me?” How can a < $1 Billion enterprise who leverages a colocation service for data center space devise a green IT strategy? What about very small organizations who have all of their servers hosted at their own facility, all in one rack that sits in a locked (or sometimes not) closet that is doubling as storage and data center space? As the week went on I managed to pick up several practical action items that can be leveraged by smaller IT shops to build a green story. While in smaller IT shops the savings may not be as dramatic and jaw dropping as global IT operations, keep in mind all things are relative. If you can show where you started and demonstrate a thoughtful effort that created a tangible reduction in energy consumption and thus cost then you have executed a successful green IT strategy. It may not be enough to save the planet, but it is your part and shows that you are doing all you can to be a good steward of your organizations IT operations.

Know Where You Are

Your tangible achievements with Green IT thinking will be measured in percentage points. For example, a 10% reduction in power consumption, a 2% reduction in utility costs etc.. Obviously, to show these numbers you have to know what you are consuming today. Even if you have a single rack of servers sitting in a closet, do you know how much power they are consuming? Do you know how much you are paying monthly in energy costs to run the equipment you have? The answer to these types of questions becomes slightly more challenging for mid-sized IT shops leveraging colocation providers for data center space. Do you know exactly how many power circuits and what type (120/280V, 20A/30A) are provisioned to your space? Do you know the current draw on those circuits? In a colocation environment, a very practical first step is to install intelligent Rack Distribution Units (RDU’s). There are lots of vendors who offer these, see this one from APC as an example. Intelligent RDU’s will give you the data you need to establish your baseline. Most of these devices provide an HTTP interface that gives you basic statistics and reporting. You don’t need anything fancy, just a browser and a spreadsheet. Get a total average draw for all of your equipment over a length of time long enough to cover any major fluctuations that may occur in your computing environment.

How Redundant Do You Need To Be?

In a typical server deployment, redundancy is built into the design. We expect a certain degree of equipment failure so we account for that in our capacity planning and design efforts. So ask yourself this, how much redundancy is enough? This has a direct impact on green IT thinking. Most servers have redundant dual power supplies. Think of the environments you have with redundancy built in so that if you lose a single server you will have no downtime. Now ask, why do each of those servers have TWO power supplies plugged in to prevent it from failing in the event a power supply goes bad or you have a problem with an electrical circuit? Doesn’t your design accommodate for losing a server anyway? The real answer to this sort of question is that you plug in both power supplies on each and every server even in a redundant server arrangement because that is how you have always done it. Challenge the notion that this is necessary. You may find that with this simple thought process you cut out 50% of your power consumption in certain environments. Also, categorize your systems into different classes of criticality. This is a natural exercise for disaster recovery planning but is not often thought of when planning for power design. If an application running on a server can be down for a defined period of time with no serious impact to your business, maybe you don’t need to plug in both power supplies and provide redundancy at the power level. Again, challenge the assumption that just because a piece of gear has two power supplies that they MUST both be in use.

Buy Green

Pat of knowing where you are in terms of your level of green IT operations is knowing where your vendors and partners stand. Make energy efficiency a part of your purchasing decisions when it comes to IT equipment and service. As you are gathering quotes for hardware, ask your vendor to provide you with energy rating information for the equipment along with the price. With respect to service such as consulting, ask your provider what they are doing to reduce their carbon footprint and provide services in a green way. For example, how much of the work can be done remotely versus onsite? This not only has a practical implication from a green mindset, it also reduced T&E expense. If you leverage a collocation provider for data center space, insist that your provider be able to tell you the PUE of their facility. Remember, you are paying your provider their cost plus margin. If their operations are not ran efficiently then their costs will be higher and you will pay the price. Leverage the information you are gathering from your intelligent RDU’s to renegotiate the way you are buying power in your colocation space. For example, many colocation providers charge you “by the circuit” for power. Their price often includes their cost for delivering you the energy on that circuit plus their cost for removing the heat generated by the consumption of that energy. If you can tangibly show that you are only consuming a fraction of the energy provided over a circuit and thus generating less than a 1:1 ratio of heat to circuit capacity then you have a strong case to moving to a “pay by the drink” model for power. Your case here is that you should be paying in a manner that covers your providers OPEX, not their CAPEX. Your providers true operational cost to remove heat from the data center is a function of how much heat you actually generate plus how efficient they are at removing it.

These are just a few practical suggestions. All of these things can be done on a modest IT budget. The key is to understand where you are starting from, take tangible steps based on that knowledge and measure your change. Again, success is measured in percentage points, not raw numbers. As manager for a small or mid-sized IT shop, you won’t put up numbers the likes of what you will see from global, Fortune 500 IT shops. Success is demonstrating that you are aware of the need to be a good steward of your organizations IT operations relative to its environmental impact, measuring the results of your efforts and arming your organization with your numbers to add to its overall sustainability efforts. Oh yea, you will likely save money too.

Saturday, April 3, 2010

The Value of an Outside IT Resource: Pointing at Elephants

There are times in the evolution of an IT organization where bringing in a new team member from the outside has significant value. While at key evolutionary points it might make sense for this new injection to come in the form of a new, permanent hire, there are times when simply bringing in a consultant can have the same value. The “value” in these cases is not fully supplied through additional technical expertise. Sometimes, an outside resource can prevent proposed solutions from being constrained by Group Think. IT teams who work together for a while come to know and understand the implied constraints that often surround proposed technical and process architectures. These implied constraints may come in the form of known manager biases, past group experiences, perceived realities of the organization (correct or not) and the simple desire to “fit” and be perceived as a team player. Many organizations do purposefully cultivate a specific corporate culture aimed at helping all members of the company to understand the guard rails within which the organization wishes to operate. New hires to IT teams such as data center operations, network engineering, database administration or software engineering certainly need to learn and understand the corporate culture which may help to define the big picture of the overall IT strategy. At a tactical level however, new hires can bring in new experiences, challenge team assumptions and ask fresh questions. In short, a new hire can point at the elephant in the room that current team members understand is there but also understand that they should not point out.

While new hires will likely cause the existing team to cycle through the typical stages of group development (forming, storming, norming, performing), it may be a move that pays dividends at key evolutionary milestones in an IT departments growth. There may be other times however where an outside resource can add value to a particular decision. Consultants are often sought out for their technical expertise or their real world experiences with other clients. Sometimes, the value of a consultant is also his or her “unbiased” opinion on a particular technical design or process improvement effort.

Whether it is a new hire at a strategic evolutionary milestone in the IT departments growth or a consultant brought in to evaluate a specific thought process, part of the value an outside resource provides is to point at your elephants.

Sunday, March 28, 2010

From Hero’s to Process: The IT Growth Challenge

I recently finished reading How to Castrate a Bull by Dave Hitz, one of the founders of NetApp. I found one concept in particular to stick in my head: as companies get larger they naturally gravitate away from needing hero’s to needing more process. Dave tells the story of how during the early years of NetApp, one support engineer went above and beyond to satisfy the needs of a customer. The support engineer took a call late in the evening and determined the customers NetApp unit needed to be replaced. The engineer went into manufacturing, took a new unit off the line, hopped a flight to the customers’ location and worked all night to install the new unit. Once word got back to Dave about the engineer's heroics, the engineer was nowhere to be found. Turns out, the engineer was asleep in the customers’ parking lot inside the van he had rented. Dave points out that at the time, this engineer was hailed as a hero but that now he hopes this kind of thing never happens again. Why? When NetApp was a small company its customer base was generally small organizations with very little gear. They themselves where very nimble and acted on failures in their enterprise very swiftly. The catch is, given their small size and the amount of IT gear in their enterprise, they usually did not encounter a large number of failures. When failures are rare, organizations can afford to rely on hero’s to step up in those rare occasions where duty calls. As enterprises get larger, the amount of IT gear and the complexity of the environment that gear supports also grow. Failures become more common place and the degree to which you can rely on hero’s is diminished. For example, for simplicity say a piece of IT gear has an average failure rate of once every 365 days. A small organization with only one of these devices can expect a failure once a year. A larger organization with 365 of these devices can expect a failure every day! Larger IT enterprises need process to deal with these repeated occurrences, not hero’s to step up on rare occasions. Hero’s come and go, process is permanent. Dave's point is that the same behavior that won the engineer and NetApp accolades from the small customer years ago, would likely lose business in a large enterprise

This evolution from needing hero’s to needing process is one of the most subtle yet important changes an IT leader must make as his or her enterprise grows larger and more complex. Tearing down technical fiefdoms, redefining reward systems, purposefully slowing down to gain control and potentially even making staffing adjustments is a daunting set of goals. This evolution inside of IT is a natural part of organizational growth and if handled correctly can be a very exciting managerial challenge.