Sunday, April 18, 2010

You Don't Have to be a Giant to be Green

Last week I was fortunate enough to attend SAP Virtualization Week at the SAP Co-Innovation Labs in Palo Alto, California. The conference consisted of three tracks: Virtualization, Cloud Computing and Green IT. Lots of thought provoking material was presented by SAP, consulting partners and customers. When I attend these type of events I look hard to find a few practical take aways that could be put in place immediately. Of course these sessions help frame my thinking on long-term strategies around SAP virtualization and Data Center design, but what can I go home with and ask my team about next week? One of those take away points for me last week was this: You don’t have to be a giant to be green.

At surface level Green IT seems like a concept relative only to the largest of IT organizations. Organizations such as Intel and Colgate-Palmolive who can tell stories of collapsing 50 or 60 global data centers down to one facility certainly have a green story to tell. Organizations such as NetApp who have pioneered efficient data center designs can go to conferences to show how they received a million dollars in rebates from their energy provider. These are great stories, but as I sat through the first couple of these last week sipping Starbucks coffee I was thinking to myself “This is awesome, but what can I do, how is this relative to me?” How can a < $1 Billion enterprise who leverages a colocation service for data center space devise a green IT strategy? What about very small organizations who have all of their servers hosted at their own facility, all in one rack that sits in a locked (or sometimes not) closet that is doubling as storage and data center space? As the week went on I managed to pick up several practical action items that can be leveraged by smaller IT shops to build a green story. While in smaller IT shops the savings may not be as dramatic and jaw dropping as global IT operations, keep in mind all things are relative. If you can show where you started and demonstrate a thoughtful effort that created a tangible reduction in energy consumption and thus cost then you have executed a successful green IT strategy. It may not be enough to save the planet, but it is your part and shows that you are doing all you can to be a good steward of your organizations IT operations.

Know Where You Are

Your tangible achievements with Green IT thinking will be measured in percentage points. For example, a 10% reduction in power consumption, a 2% reduction in utility costs etc.. Obviously, to show these numbers you have to know what you are consuming today. Even if you have a single rack of servers sitting in a closet, do you know how much power they are consuming? Do you know how much you are paying monthly in energy costs to run the equipment you have? The answer to these types of questions becomes slightly more challenging for mid-sized IT shops leveraging colocation providers for data center space. Do you know exactly how many power circuits and what type (120/280V, 20A/30A) are provisioned to your space? Do you know the current draw on those circuits? In a colocation environment, a very practical first step is to install intelligent Rack Distribution Units (RDU’s). There are lots of vendors who offer these, see this one from APC as an example. Intelligent RDU’s will give you the data you need to establish your baseline. Most of these devices provide an HTTP interface that gives you basic statistics and reporting. You don’t need anything fancy, just a browser and a spreadsheet. Get a total average draw for all of your equipment over a length of time long enough to cover any major fluctuations that may occur in your computing environment.

How Redundant Do You Need To Be?

In a typical server deployment, redundancy is built into the design. We expect a certain degree of equipment failure so we account for that in our capacity planning and design efforts. So ask yourself this, how much redundancy is enough? This has a direct impact on green IT thinking. Most servers have redundant dual power supplies. Think of the environments you have with redundancy built in so that if you lose a single server you will have no downtime. Now ask, why do each of those servers have TWO power supplies plugged in to prevent it from failing in the event a power supply goes bad or you have a problem with an electrical circuit? Doesn’t your design accommodate for losing a server anyway? The real answer to this sort of question is that you plug in both power supplies on each and every server even in a redundant server arrangement because that is how you have always done it. Challenge the notion that this is necessary. You may find that with this simple thought process you cut out 50% of your power consumption in certain environments. Also, categorize your systems into different classes of criticality. This is a natural exercise for disaster recovery planning but is not often thought of when planning for power design. If an application running on a server can be down for a defined period of time with no serious impact to your business, maybe you don’t need to plug in both power supplies and provide redundancy at the power level. Again, challenge the assumption that just because a piece of gear has two power supplies that they MUST both be in use.

Buy Green

Pat of knowing where you are in terms of your level of green IT operations is knowing where your vendors and partners stand. Make energy efficiency a part of your purchasing decisions when it comes to IT equipment and service. As you are gathering quotes for hardware, ask your vendor to provide you with energy rating information for the equipment along with the price. With respect to service such as consulting, ask your provider what they are doing to reduce their carbon footprint and provide services in a green way. For example, how much of the work can be done remotely versus onsite? This not only has a practical implication from a green mindset, it also reduced T&E expense. If you leverage a collocation provider for data center space, insist that your provider be able to tell you the PUE of their facility. Remember, you are paying your provider their cost plus margin. If their operations are not ran efficiently then their costs will be higher and you will pay the price. Leverage the information you are gathering from your intelligent RDU’s to renegotiate the way you are buying power in your colocation space. For example, many colocation providers charge you “by the circuit” for power. Their price often includes their cost for delivering you the energy on that circuit plus their cost for removing the heat generated by the consumption of that energy. If you can tangibly show that you are only consuming a fraction of the energy provided over a circuit and thus generating less than a 1:1 ratio of heat to circuit capacity then you have a strong case to moving to a “pay by the drink” model for power. Your case here is that you should be paying in a manner that covers your providers OPEX, not their CAPEX. Your providers true operational cost to remove heat from the data center is a function of how much heat you actually generate plus how efficient they are at removing it.

These are just a few practical suggestions. All of these things can be done on a modest IT budget. The key is to understand where you are starting from, take tangible steps based on that knowledge and measure your change. Again, success is measured in percentage points, not raw numbers. As manager for a small or mid-sized IT shop, you won’t put up numbers the likes of what you will see from global, Fortune 500 IT shops. Success is demonstrating that you are aware of the need to be a good steward of your organizations IT operations relative to its environmental impact, measuring the results of your efforts and arming your organization with your numbers to add to its overall sustainability efforts. Oh yea, you will likely save money too.

Saturday, April 3, 2010

The Value of an Outside IT Resource: Pointing at Elephants

There are times in the evolution of an IT organization where bringing in a new team member from the outside has significant value. While at key evolutionary points it might make sense for this new injection to come in the form of a new, permanent hire, there are times when simply bringing in a consultant can have the same value. The “value” in these cases is not fully supplied through additional technical expertise. Sometimes, an outside resource can prevent proposed solutions from being constrained by Group Think. IT teams who work together for a while come to know and understand the implied constraints that often surround proposed technical and process architectures. These implied constraints may come in the form of known manager biases, past group experiences, perceived realities of the organization (correct or not) and the simple desire to “fit” and be perceived as a team player. Many organizations do purposefully cultivate a specific corporate culture aimed at helping all members of the company to understand the guard rails within which the organization wishes to operate. New hires to IT teams such as data center operations, network engineering, database administration or software engineering certainly need to learn and understand the corporate culture which may help to define the big picture of the overall IT strategy. At a tactical level however, new hires can bring in new experiences, challenge team assumptions and ask fresh questions. In short, a new hire can point at the elephant in the room that current team members understand is there but also understand that they should not point out.

While new hires will likely cause the existing team to cycle through the typical stages of group development (forming, storming, norming, performing), it may be a move that pays dividends at key evolutionary milestones in an IT departments growth. There may be other times however where an outside resource can add value to a particular decision. Consultants are often sought out for their technical expertise or their real world experiences with other clients. Sometimes, the value of a consultant is also his or her “unbiased” opinion on a particular technical design or process improvement effort.

Whether it is a new hire at a strategic evolutionary milestone in the IT departments growth or a consultant brought in to evaluate a specific thought process, part of the value an outside resource provides is to point at your elephants.

Sunday, March 28, 2010

From Hero’s to Process: The IT Growth Challenge

I recently finished reading How to Castrate a Bull by Dave Hitz, one of the founders of NetApp. I found one concept in particular to stick in my head: as companies get larger they naturally gravitate away from needing hero’s to needing more process. Dave tells the story of how during the early years of NetApp, one support engineer went above and beyond to satisfy the needs of a customer. The support engineer took a call late in the evening and determined the customers NetApp unit needed to be replaced. The engineer went into manufacturing, took a new unit off the line, hopped a flight to the customers’ location and worked all night to install the new unit. Once word got back to Dave about the engineer's heroics, the engineer was nowhere to be found. Turns out, the engineer was asleep in the customers’ parking lot inside the van he had rented. Dave points out that at the time, this engineer was hailed as a hero but that now he hopes this kind of thing never happens again. Why? When NetApp was a small company its customer base was generally small organizations with very little gear. They themselves where very nimble and acted on failures in their enterprise very swiftly. The catch is, given their small size and the amount of IT gear in their enterprise, they usually did not encounter a large number of failures. When failures are rare, organizations can afford to rely on hero’s to step up in those rare occasions where duty calls. As enterprises get larger, the amount of IT gear and the complexity of the environment that gear supports also grow. Failures become more common place and the degree to which you can rely on hero’s is diminished. For example, for simplicity say a piece of IT gear has an average failure rate of once every 365 days. A small organization with only one of these devices can expect a failure once a year. A larger organization with 365 of these devices can expect a failure every day! Larger IT enterprises need process to deal with these repeated occurrences, not hero’s to step up on rare occasions. Hero’s come and go, process is permanent. Dave's point is that the same behavior that won the engineer and NetApp accolades from the small customer years ago, would likely lose business in a large enterprise

This evolution from needing hero’s to needing process is one of the most subtle yet important changes an IT leader must make as his or her enterprise grows larger and more complex. Tearing down technical fiefdoms, redefining reward systems, purposefully slowing down to gain control and potentially even making staffing adjustments is a daunting set of goals. This evolution inside of IT is a natural part of organizational growth and if handled correctly can be a very exciting managerial challenge.

Saturday, March 20, 2010

Considerations for Implementing SIP

In a recent Computer World article I discussed how the real cost savings associated with VoIP in the corporate enterprise comes from the implementation of Sesion Initiation Protocol (SIP). With most organizations paying very low long distance rates, justifying VoIP investments on long distance savings becomes a real challenge. SIP on the other hand allows you to get rid of local loop charges through a reduction in the number of PRI's you have dedicated to voice at remote and field locations. There are big savings to be had in SIP but in order to understand what your net savings will be you need to think through some of the following basic factors.

SIP Provider Pricing

Unlike traditional TDMS services, pricing for SIP is far less standardized. Different carriers will charge for the service in different ways. For example, PAETEC charges a flat fee for the bandwidth of the MPLS connection used to provide the SIP service plus $1.50/"virtual DID". On the other hand Time Warner Telecom charges more on an actual call volume basis. There really is no "best" way for a company to price its SIP service, what works best for you will depends on the dynamics of your particular organization. Make sure you understand the pricing scheme of the SIP providers you are evaluating and match that model to your organizations typical call dynamics.

Understand Where You Can Truly Remove Services

The implementation of SIP in your network is not the end to your PRI based TDMS service. Several factors will likely limit your ability to remove PRI's in certain areas. For example, make sure you understand for which of your remote locations your SIP provider can provide the virtual DID's. Not every SIP provider can provide local numbers in all markets, for those markets not covered by your SIP provider you will need to maintain your PRI service to maintain that locations local DID's. Also, consider services like faxing that may be riding on PRI's through FXO ports today. What will you do about this? How many local lines will you need to maintain for systems such as building security, fire alarms etc..? It may be that the need to maintain support for an antiquated technology such as faxing (my opinion) may hamper your ability to leverage modern network designs for financial gains. You may find yourself using your SIP implementation as a good time to rationalize the continued support for older communications mediums.

How Much Bandwidth Will You Need At Remote Sites?

Your current data link at each of your remote sites is running at some percent of utilization today. What will that be once you add voice traffic over the circuit? You need to make some estimates about a locations call volume combined with the type of voice compression you plan to use in order to determine if you will need to increase the bandwidth on your locations data pipe. This is significant because you could find yourself simply shifting costs around, from PRI to Data circuit. In order for SIP to represent a true cost savings at your field site, make sure the overall net spend decreases after PRI removal, additional services for local lines and potential increased data bandwidth.

Understand the Best Technical Architecture for your Organization

The technical details of a network design for SIP can be daunting. At the end of the day, there are really two major types of designs you will likely choose from. In a centralized model, all remote locations receive SIP trunks and thus call trunks through a centralized connection to the SIP provider. The SIP provider will deliver some type of connection to your corporate HQ or data center where you will connect to them through a CUBE. You can see an image of this here..
A distributed model requires each location to peer with the SIP provider. The distributed model provides a slightly more robust design in terms of backup and redundancy but at a higher CAPEX and OPEX cost. You can see the distributed model diagram here..

This is not an exhaustive list of considerations but thinking throh these things will help you to solidify your thinking on both the technical and financial benefits of SIP in your organization. I recommend that for evaluating SIP providers and doing your cost/benefit analysis, you utilize a third party who specializes in Telco analysis. It is critical that for you to make good decisions about SIP that you understand your current state. The SIP market is still young and services are still often tailored to your specific needs. Make sure you truly understand the net of it all before jumping in.

Sunday, March 14, 2010

Zen and the Art of Converged and Efficient Data Centers

What a few weeks it has been. Over the last month I have been fortunate enough to meet with the CTO from Frito-Lay, the CTO from NetApp, attend a joint SAP and NetApp executive briefing at SAP’s North American HQ in Philadelphia and tour two world class IT support centers (PepsiCo in Dallas and Dimension Data in Boston). I have spent days pouring through technical documentation geared towards architecting my organizations next generation data center centered around 10GB Ethernet, Virtualization, Blade Systems and efficient energy practices. So this post is probably as much for myself as anyone else, meant to simply document of few of the key learning’s I have taken away from the flurry of activity over the last few weeks. Hey, maybe someone else will find it interesting too?

PUE & The Green Grid

In a one on one conversation with Dave Robbins, the CTO of NetApp Information Technology, he asked what my data center space providers PUE is. My response was an inquisitive, what? PUE stands for Power Use Efficiency and is a measure of how effectively a data center is using its energy resources. Essentially, PUE is the amount of electricity used by a data center for cooling and mechanics divided by the actual IT load. Efficient Data Centers run at around 1.6. The concept of PUE and its measurement was created by an organization known as The Green Grid and you can find all kinds of great resources at their web site. This is an excellent tool for you to use when negotiating power costs with a Hosting provider. You should know their PUE and insist that you will not pay for their inefficacy. You can also find a cool tool for PUE calcualtion at 42U.com.

It is Time to Converge

The introduction of 10GB Ethernet in Data Centers (and perhaps even more important, lossless Ethernet) has truly created an opportunity to collapse Ethernet and Fiber Channel networks in the Data Center backbone, cutting huge costs in Fiber Channel infrastructure. 10 GB Ethernet and Lossless Ethernet serve as enablers for protocols such as FCoE and FIP which allow Fiber Channel frames to be encapsulated and carried across Ethernet backbones. There are a few watch outs when adopting FCoE that you need to be aware of. First, make sure your storage vendor has a CNA (Converged Network Adapter) that supports BOTH FCoE and other IP based traffic. Some of the early “converged” adapters only support FCoE, not much real convergence there. Put some effort in understanding Cisco’s current support of FCoE and Fiber Channel Initialization Protocol (FIP) in their Nexus line of switches. You will find some good resources here. The details of this are too complex for me to go into here but suffice it to say, you need to think long and hard about your data center switch layout in order to get full FCoE support across your 10GB backbone. Also, remember that lossless Ethernet or data center bridging are keys to FCoE success but are fairly new. So, when you hear people tell you they knew someone who tried FCoE a couple of years ago but found it lacking, take it with a grain of salt.


The FUD around Cisco UCS

Let me get one thing out of the way upfront, the Cisco Unified Computing System (UCS) is sexy. Cisco’s tight relationship with VMware, stateless computing and a seemingly end to end vision for the data center combine for a powerful allure. Competitors such as IBM and HP are quick to point out that their blade center products perform the same functions as Cisco’s UCS but with a proven track record. In general, these claims are true. I have been exposed to some competitive claims against the UCS that where simply meant to plant the seed of Fear Uncertainty and Doubt (FUD) in the mind of technology managers. What if Cisco changes their Chassis design, is your blade investment covered? UCS is meant for VMware only (not true). The list goes on. I have been heavily comparing the Cisco UCS to IBM’s H series Blade Center. I had originally convinced myself that the difference between these two offerings was all about the network. Cisco’s UCS does offer some interesting ways to scale across chassis and provides some great management tools. For a mid-sized organization, the ability to scale across chassis becomes less important however when you can get a concentrated amount of compute power inside one or maybe two chassis. Some new technology coming from IBM in the form of their MAX5 blades is going to allow for some massive compute power inside a two socket blade. If you are a large organization planning on adding many UCS chassis, the networking innovations in the UCS likely will fit your needs well. For a mid-sized company, consider getting more compute power inside fewer chassis by using some hefty blades. This not only reduces your need to scale across many chassis, it also helps lower your VMware costs. VMware is licensed by the socket so fewer sockets with more cores on blades with higher memory capabilities ultimately drives down your VMware licensing needs. Also, before you completely convince yourself that the Cisco UCS has a strong hold on the networking space in the data center, spend some time understanding IBM’s Virtual Fabric technology. This offers similar features to the VIC cards in the Cisco UCS. The point is this, don’t be immediately sucked in by the sexy UCS. Cisco has come to the blade market with some cool innovation and in some circumstances, it will be exactly what you need. Make the investment in time to really understanding competing products. Avoid FUD in all directions.

Saturday, February 13, 2010

Be Consistent & Flexible

I remember an experiment covered in an under graduate psychology class meant to demonstrate that people prefer consistency in thought over randomness even if they disagree with the principal of the consistent thought pattern. Here is the scenario:

Person X states that he dislikes all people from place Y. He then comes to know that his new friend to whom he has grown very close is from place Y. Person X has three options:

1. Change his views about people from place Y.

2. Maintain his view about people from place Y but rationalize an exception.

3. Denounce his friendship given his new knowledge of his friends place of origin.


Most people see option 1 as the "correct” choice, not surprising given the positive connotation this option holds. What is more interesting is that given only a choice between option 2 and option 3, most people choose option 3 as the "correct" choice despite the negative connotation that option holds. My point here is not to delve into the deep rooted psychological reasoning behind this fact but rather to point out the following: Given a choice between consistency and inconsistency, people prefer consistency even if they do not necessarily agree with the principal being consistently adhered to. This is an extremely important point for IT leaders to keep in mind across a wide facet of IT operations.

Consider certain IT policy related to issues such as who in the organization has local administrator rights to their workstation. Users will generally prefer "option 1", giving them free reign and full flexibility as it relates to installing software and making configurations on their own workstation. However, the very practical considerations of security and stability force us as IT leaders to take "option 1" off the table. So an IT workstation policy can really be seen as serving two purposes. First, it clearly outlines to users in a practical and no disputable fashion why "option 1" is not a possibility. Second, the policy tells users how the rules will be applied consistently across the organization. Armed with this information, users will prefer a consistent application of IT policies even if it means they are not given their "option 1". Remember, users will always prefer "option 1" but will accept "option 3" over "option 2" if it is the only alternative. Users will always be on the diligent lookout for the existence of "option 2" in the organization. The moment users perceive the application of policies in an inconsistent fashion the howling will begin.

Consider also the yearly ritual of the performance appraisals. An employee who is rated low on a category will accept that rating if he or she feels the standards for that category are being applied consistently across all staff members. IT team members may not necessarily agree that "number of support cases closed in a 24 hour period" is a relevant measure, however, if all team members are graded consistently with respect to this metric it will become accepted as a goal. The hidden gotcha here of course is to be careful what it is you consistently ask for because that is exactly what you will get!

These are only two examples, many more abound. So as an IT leader, ask yourself if you are being consistent in your actions. Do you users understand the rules and see them as being applied the same to everyone? Do your staff members feel they are being measured consistently against the goals you have set whether they necessarily agree with your goals or not? Be consistent in your actions but do not lose sight of the need to listen and be flexible. As things change, the rules and goals that govern your IT organization may also need to change. Be willing to adjust, inform everyone clearly of how the game has changed and adhere to the new standards in a consistent fashion. This last point reminds me of an adage told to me by a colleague who spent many years in the Navy:

On a foggy evening, a U.S. Navy destroyer is leaving harbor, heading out to sea. The Admiral is setting in his quarters working on paperwork when he notices a light dead ahead. He radios down to the bridge to inform the light to turn 20 degrees port side. The bridge calls back up to the Admiral and tells him the light has refused the order. Hearing this, the Admiral gets livid and asks to be connected directly. Once connected the Admiral roars: "I am an Admiral in the U.S. Navy and this is a U.S. Navy Destroyer". The voice on the other end responds "Understood Sir, I am a third class Navy seaman and this is lighthouse".....

Saturday, January 30, 2010

Don't Spend 80% Of Your Time On 20% Of Your Users

Some people just don't like change. This is especially true when it comes to changes in technology that they use every day to accomplish work tasks that bring their own set of stressors. Folks learn something, fall into a routine and then just want it to stay the same. This is understandable; technology is an enabler to accomplish a goal, not a goal in and of itself. Unfortunately, the only constant with technology is change. New versions of operating systems are released, new mobile platforms, new portal technology and on and on. Vendors drop support of older technologies over time, forcing us as technology leaders to impose change upon our user base. Not all of your users will react to change in the same way and you should therefore not adopt a one size fits all approach to your change management strategy.

Leverage You Champions

Not all of your users will be resistant to change. Just as consumer technologies have a predictable early adopters portion associated with their user adoption curve, so too will your corporate technologies. Some users will be excited about the new features of a technology; some will be excited to be "first" when it comes to something new. Whatever their motives, identify your champions, get the technology in their hands early and most of all, make sure they are happy. Let these folks serve as your sounding board across the organization. Let them go to meetings, present to a group of their peers and show off "cool" new features of your new operating system. Let them sit with their peers in airports, bring up your mobile app and gain access to information their peers don’t have. These are your evangelists, treat them well and let them spread the word.

Take Care of the Masses

The majority of your user base will adopt new technology with only a short period needed to get over the proverbial "hump". Most users won’t be vocal in either direction, positive or negative. It is important to actively generate feedback from the majority to ensure true issues are separated from expected transition grumps. Pay close attention to related tickets in support desk ticketing systems, talk to as many people as you can looking for common themes and clearly document your findings to identify trends. Don’t confuse standard grumping with true wider spread issues. Nearly everyone will be slightly more vocal regarding their dislikes versus their likes. The key is to separate true, constructive feedback from simple "I don’t like this because it is not what I had before" feedback. This leads to the last group of users.

Marginalize the Hold Outs

Some people will not be pleased no matter what. Once you have listened carefully to the majority of your users, addressed the true issues related to your new technology and have started getting wide spread, positive feedback, move on. Don’t let your team spend 80% of their time struggling to please 20% of the people. If you have pleased your early adopters, won the acceptance of your masses and received positive feedback from all levels of your organization, you have succeeded. Again, make sure the few hold outs do not truly have legitimate complaints, have you over looked something specific to their job? Be sure to share your positive feedback in a very public way so as to not let the few remaining complaints become the only remaining voice being heard following a technology change.