Preventing RAID problems

June 9, 2008, 02:39 PM —  ITworld.com — 

It's always better to prevent problems than to try and resolve them when they occur. And one or the primary failure points of a server (or a high-powered workstation) is the storage subsystem. This is because hard drives are mechanical doohickeys that are prone to sudden failure, with either accompanying operating system failure, data loss, or both.

Using RAID storage is supposed to help anticipate such catastrophes by providing a level of fault tolerance for your storage subsystem. But even RAID can have its problems, especially when you use some newer high-capacity hard drives in the near-terabyte range that have less than stellar reputations for reliability. A colleague found this out recently when one of the drives in his RAID 6 array failed after only a year of use. Since RAID 6 provides an extra level of redundancy over RAID 5 and can survive the loss of two drives without data loss, my colleague felt it was safe to contact the manufacturer and request a replacement drive for the one that failed.

Fortunately, by the time the replacement had arrived, no other drive had failed. My colleague then removed what he thought was the failed drive from the RAID array, only to discover he had removed a working drive instead. Now if he had been running RAID 5 at that point, he'd be toast. Fortunately, he was able to resolve the situation by (a) plugging the working drive he had removed back into the array (b) waiting for the array to rebuild the parity info (c) removing the actual failed drive and replacing it with the replacement from the manufacturer and (d) waiting for the rebuild of the second level parity info. Everything worked fine, but he got to bed quite late that night.

What's the solution to avoiding near-disasters like this? Simple: label the drives in your RAID array so you can match them to the ports on your controller card and to drive numbers in your controller's management software. And while you're at it, you may as well label all your drive cables as well, and your drive bays also. A little preventative maintenance like this can be a big time-saver by preventing you from doing something stupid when the heat is on.

ITworld.com

I like it!
Post a comment
The content of this field is kept private and will not be shown publicly.
  • Allowed HTML tags: <a> <em> <strong> <cite> <code> <ul> <ol> <li> <dl> <dt> <dd>
  • Lines and paragraphs break automatically.
Resources
White Paper

Symantec Backup Exec 12 and Backup Exec System Recovery 8 deliver industry leading Windows data protection and system recovery. Download this whitepaper to find out the top reasons to upgrade and how to get continuous data protection and complete system recovery.

Webcast

Data and system loss — from a hard drive failure, malicious attack, natural disaster, or simple human error — can happen anytime. Don’t leave your business vulnerable. Make sure you have a secure recovery strategy in place. Symantec's latest backup and system recovery technology can efficiently restore critical applications, individual emails and documents and even restore your entire system in minutes in the event of a loss.

White Paper

Businesses face a growing challenge to ensure that the IT environment is properly protected. Backup Exec 12 integrates with other applications in the Symantec family of products, to complement your current data protection strategy, keep your data securely backed up and make it recoverable when you need it most.

Free stuff

Enterprise 2.0 Implementation
By Aaron C. Newman, Jeremy Thomas
Published by McGraw-Hill
Learn more!

Deploying Cisco Wide Area Application Services
By Zach Seils, Joel Christner
Published by Cisco Press
Learn more!

Featured Sponsor

AISO founders envisioned a Web hosting company that was environmentally friendly. While the company employed energy-efficient innovations like solar panels, its infrastructure produced unacceptable power and cooling requirements. Find out how AISO leveraged AMD technology to overcome their challenge in this case study white paper.

In this whitepaper, Scalar explores the opportunity to change the landscape with respect to mission critical databases built around Oracle. Leveraging technologies such as Linux, high-end commodity processing power and Oracle RAC technology to architect, design, build and maintain database infrastructure that delivers maximum availability, reliability and performance at a fraction of traditional cost.

On a typical day, weather.com, the Web site for The Weather Channel in Atlanta, serves up between 15 million and 20 million page views. But in September 2004, when back-to-back hurricanes ransacked Florida, the peak traffic on one day more than tripled: over 70 million page views by more than 7 million unique visitors. Read the full success story now.

More Resources