Mike Ault's thoughts on various topics, Oracle related and not. Note: I reserve the right to delete comments that are not contributing to the overall theme of the BLOG or are insulting or demeaning to anyone. The posts on this blog are provided “as is” with no warranties and confer no rights. The opinions expressed on this site are mine and mine alone, and do not necessarily represent those of my employer.
Monday, March 02, 2009
God and Technology
I prefer to think of myself as more a Jeffersonian Christian rather than a Paulian Christian. It amazes me that many Christians swallow hook-line-and-sinker every word penned by the one Apostle that never actually met Jesus face-to-face. Many scholars feel that Paul was sent on so many missions not because he was good at them but because the other Apostles really couldn’t stand him and hoped that he wouldn’t return. Many believers take Paul’s letters out of context and usually completely incorrectly as their meaning is understood by true bible scholars. In fact many articles of faith were added in after the fact to the regular gospels as can be proved by stylistic differences and from going back to the earliest known translations. The fact is that no matter how current your translation you are still starting from flawed beginnings. Many of the books that “didn’t make the cut” when the first bibles were compiled were destroyed as heresy thus removing them from possible future examination.
With any scientific field of study you reach a certain point and you can go no further, from that point on you have to accept things on theories and faith. Even with the less than whole cloth parts of the Bible’s New Testament removed, what is left is still an amazing history of a real man who lived, and died for his faith and his friends. Is Jesus the Son of God? Yes, but then we all are the sons and daughters of God. Did Jesus die for our sins? Yes. Was he raised from the dead? This is where there is some contention over what was added after the fact and what is whole cloth. But let’s examine this.
Two men: One knows he is the physical Son of God, he knows that no matter what evil painful things happen on Earth he has a place at the right hand of God. Second man, a man, with man’s frailties, and doubts. Now, both give up their Earthly lives for what they believe in, which one required greater faith? Would Jesus be less or more of an inspiration if he was a frail human or the anointed Son of God? Would you believe it was a vote by a group of flawed humans (the Nicene Council) that decided Jesus was a deity and it was actually a very close vote.
Unfortunately the only documents that provide “proof” of Jesus’ deity are in the Bible and using the Bible to prove the Bible is circular logic and therefore flawed. It is like using the Dianetics text to prove L. Ron Hubbard’s qualifications as a deity. Since most Islamic accounts of Jesus are actually taken from the Bible then related texts quoting them are not relevant. While Jesus is mentioned in some historical texts none go into great detail as to his birth, (a virgin was a young maiden, not someone who had never had sex) life (he existed and taught and was hated by Rome), death (he was probably crucified) or resurrection. These accounts of the resurrection were actually added after the original text in the gospels, it was felt that the Mithran belief, having a virgin birth, life and resurrection was a big spur to add these passages. (see: http://www.near-death.com/experiences/origen048.html) Among non-Christian historians, Pliny the Younger, Suetonius and Tacitus refer to Jesus, as does Josephus (Joseph ben Matthias).
So, do I believe Jesus is my savior? Yes, his teachings show the way to the father as he himself said “There is no way to the father but through me” meaning through his teachings we find the way. Do I believe in the resurrection? That is more complex to answer. Unfortunately the resurrection is one of the parts added after the original text, that makes it suspect in my eyes. The key question is “Would I believe without the resurrection?” The answer is yes, I would, so whether I believe in the resurrection or not is moot. You are free to believe as you wish and I would never dream of pushing my beliefs onto you, after all, we have free will.
Do I believe in the life everlasting? Yes. There is enough anecdotal evidence to show that something of us exists after death, that the spirit left gets rewarded or punished based on a set of criteria created within its own belief structure is not that far of a reach. After all Jesus also said “In my Father’s house are many mansions, I go there now to prepare a place for you.”
So what have I attested to? I believe in God and I believe in Jesus. I believe Jesus died for my sins. I believe Jesus’ teachings show the way to true belief in God. If this diminishes me in some folk’s eyes, so be it. However, it is not for people that I live, I live for God, my family and myself.
Does God reject technology? No, he gave us technology to better ourselves. As with any tool, how we use it determines whether the tool is good or bad. Do modern teachings contradict the Bible? No. If you realize that most of the creation story is metaphor, used to explain something we don’t have full understanding of even today, to ignorant herders. When looked at as metaphor it actually parallels what we know. Look at the theory of vacuum fluctuations and compare it to the story of genesis. As to what timelines are used in the Bible, again, try to explain millions of years to someone who barely understands how to count his herd of sheep.
It is odd that those that insist on a literal interpretation of their favorite passages that damn certain behaviors or exalt other behaviors they profess to believe in themselves but then they tell us other parts are metaphor. Remember that the true test of a prophet is that what he prophesizes comes true. I am afraid many of the added texts in the Bible fail this test as do many of the founders of many splinter religions who used imperfect understanding to make prophesies of specific dates for events such as the “rapture” and the second coming. Of course instead of applying the test to these false prophets and rejecting them, they merely allowed that they were mistaken but that their prophesies would come true eventually.
Limiting God to a simplistic creation story is demeaning to God. That God could put into motion such marvelous mechanisms as those behind vacuum fluctuations and evolution is a testament to his greatness, not a detractor from it. That we cannot understand everything is a testament to God’s greatness . God guides technology, giving us tools to better understand his universe.
Burying our heads in simplistic beliefs because we cannot understand God’s plan as implemented in his universe is an affront to God.
Tuesday, February 17, 2009
RMOUG Notes
My only complaint about both presentations was that when they presented the user test results they neglected to show the full (or even partial) configurations of the servers and disk systems they had tested against. Rather like saying my car is 10 times faster than Joe’s and telling you mine is a 1995 Dodge Avenger and failing to mention Joe’s is a Stanley Steamer. Be that as it may, I still enjoyed the presentations and the best take away was from Kevin’s presentation when he said that “If your current system is fully tuned, has adequate disk resources, and is performing well, the Exadata has nothing to offer you.” An example from kevin would be a 128 CPU Superdome with 128 4GFC HBAs that were being fed by ample XP storage as that would be 51GB/s ingest-capable. Also during Tom’s presentation he admitted the primary target of the Exadata was those shops with row-after-row of Oracle servers followed by a single Netezza or Teradata server or servers.
Essentially the Exadata Database Machine is targeted at the larger (several terabytes) data warehouse that would otherwise be placed on a Netezza or Teradata machine and I couldn’t agree more. However, it would be a fun test to replace the disks in an Exadata cell with a RamSan-500 and see what (if any) additional performance could be gained. After all, the disks are still the limiting factor in the performance of the system. For example, a single Exadata cell tops out at around 2,700 IOPS, according to white papers on the Oracle site; a single RamSan-500 can sustain 100,000 mixed read/write IOPS and 25,000 pure write IOPS with minimal response times. As far as I can tell, no additional smarts are built into the Exadata disk drives in the place of special firmware, such as is supposedly done with EMC systems, so replacing the drives with a single RamSan-500, either set up as 12 LUNs, or as a single large LUN, should be easy.
Another interesting discussion I had during this time frame was with our (Texas Memory Systems) own Matt Key, one of our Storage Applications Engineers, about why adding the Enterprise Flash Drives (EFDs) to arrays produces little if any benefit for large levels of writes. Turns out there is an upward limit on the bandwidth a single disk tray can handle and with the EFDs instead of disk drives the disk tray tops out at around 3000 (between 1600 and 3200) or so IOPS (based on a 64K stripe) so you actually need several trays (with a max of only 4 drives to a tray because of other limits) to get significant write IOPS. For comparison, the RamSan-500 can handle 25,000 sustained write IOPS. Now don’t get me wrong, the EFDs can improve the performance of certain types of loads when compared to a standard array with no EFDs, but if you are write-heavy you may wish to consider other technologies. Note: The calculations are based on a 200 megabyte/second FC-AL bandwidth with 64K writes, since RAID6 is used there are 2-64K writes for each write, 200MBS/64K=3200 IOPS, 200MBS/128K=1600 IOPS. These limitations apply to all array-based EFDs.
The RamSan-500 makes an excellent complement to any enterprise array, especially if you use the preferred read technology to read from the RamSan-500 while writing to both, for example, when you are using array-based replication, such as SRDF, to provide geo-mirroring of the frame to a remote site. By offloading the reads, the number of writes that can be supported by the array can be increased as a factor of the percent of reads in the work load, thus increasing the performance of the entire system. As an example, if you have an 80/20 read/write workload and you offload the 80 percent of reads to the RamSan, this frees up the array to handle a factor of 4 more writes, up to the actual maximum IOPS of the array. This is a 4X increase in I/O with 0-impact to infrastructure or BCVs.
Oh, on February 24-25 I’ll be in Charlotte, NC presenting at the Southeast Oracle Users Convention (SEOUC). My two presentations are: “My Ideal Data Warehouse System” and “Going Solid: Use of Tier Zero Storage in Oracle Databases.” I hope I see you there!
As I digest more of the information I obtained this week, I will try to write more blog entries. So for now I will sign off. Good bye from 37,000 feet over Colorado!
Mike
Wednesday, February 04, 2009
Do You Need Solid State Technology?
SSD Advantages
SSD technology has one big advantage over your typical hard disk based storage array: SSD does not depend on physical movement for retrieval of data. Being non-dependent on physical movement for data retrieval means that you can significantly reduce the latency involved with each data retrieval operation, usually on the order of a factor of 10 (for Flash-based technology) to over 100 (for RAM-DDR -based technology.) Of course cost increases as latency decreases with SSD technology, with Flash running about a quarter of the cost of RAM-DDR technology.
SSD Costs
Flash and DDR-based SSD technology are usually on a par with, or can be cheaper than, IO equivalent SAN based technology. Due to the much lower latency of SSD technology you can get many more input-output operations per second (IOPS) from them than you can from a hard disk drive system. For example, from the “slow” Flash-based technology you can get 100,000 IOPS with an average latency of 0.20 milliseconds worse case. From the fastest DDR based technology you can achieve 600,000 IOPS with a latency of .015 milliseconds.
To achieve 100,000 IOPS from hard drive technology you would need around 500 or more 15K rpm disks at between 2 and 5 milliseconds latency, giving a peak IOPS of around 200 per disk drive for random IO, regardless of their storage capacity. A 450 gigabyte 15K rpm disk drive may have up to 4 or more individual disk platters with 12 read-write heads (one for each side of the disks); however, these read-write heads are mounted on a single armature and are not capable of independent positioning. This limits the latency and IOPS to that of a single disk platter with two heads, so an 146 gigabyte 15K rpm drive will have the same IOPS and latency as a 450 gigabyte 15K rpm drive from the same manufacturer (http://www.seagate.com/docs/pdf/datasheet/disc/ds_cheetah_15k_6.pdf.)
Given that the IOPS and latency are the same regardless of storage capacity for a 146 to 450 gigabytes array of disk drive sizes why not pick the smallest drive and save money? The reason is that to get the best latency you need to be sure not to fill the various disks in the disk drive (from 2 to 4) more than 30% hence the need for so many drives. So to get high performance from your disk based IO subsystem you need to throw away 60-70% of your storage capacity!
Do I Need That Many IOPS?
Many critics of SSD technology state that most systems will never need 100,000 IOPS, and in many cases they are correct. However, in testing using a 300 gigabyte TPC-H (data warehouse) type test load using SSD I was able to get peak loads of over 100,000 IOPS using a simple 4 node Oracle11g Real Application Clusters-based setup. Since many systems are considerably larger than 300 gigabytes and have more users than the 8 users with which I reached 100,000 IOPS, it is not inconceivable that given the capability to achieve 100,000 IOPS of throughput many current databases would easily exceed that value. It must also be realized that the TPC-H system I was testing utilized highly optimized indexing, partitioning, and parallel query technology; eliminate any of these capabilities and the IOPS required increases, sometimes dramatically.
Questions to Ask
So now we reach the heart of the question, do you need SSD for your system? The answer depends on several questions which only you can answer:
1. Is my performance satisfactory? If yes then why are you asking about SSD?
2. Have you maximized use of memory and optimization technologies built into your database system? If no, then do this first.
3. Has my disk IO subsystem been optimized? (Enough disks and HBAs?)
4. Is my system spending an inordinate amount of time waiting on the IO subsystem?
If the answer to question 1 is no, and questions 2, 3 and 4 are yes then you are probably a candidate for SSD technology. Don’t get me wrong, if I had my choice I would skip disk based systems altogether and use 100% SSD in any system I bought, given the choice. However, you are probably locked into a disk-based setup with your existing system until you can prove it doesn’t deliver the needed performance. Let’s look closer at the 4 questions.
Question 1 is sometimes hard to answer quantitatively. Usually the answer to 1 is more of a gut reaction than anything that can be put on paper. The users of the system can usually tell you if the system is as fast as they need it to be. Another consideration in question 1 is: performance is fine now, but what if you grow by 25-50%? If your latency is at 3-5 milliseconds on the average now, adding more load may drive it much higher.
Question 2 will require analysis of how you are currently configured. An example from Oracle is that a wait on db file sequential reads can indicate that not enough memory has been allocated to cache data blocks read based on index reads. So, even if the indexes are cached, the data blocks are not and must be read into the cache on each operation. A sequential read is an index-based read followed by a data table read and usually should be cached if there is sufficient memory. Another Oracle wait, db file scattered reads indicates full table scans are occurring. Usually full table scans can be mitigated by use of indexes or partitioning. If you have verified your memory is being used properly (perhaps everything that can be allocated has been) and you have utilized the proper database technologies and performance is still bad, then it is time to consider SSD.
A key source of wait information is of course the Statspack or AWR report for Oracle based systems. One additional benefit to the Statspack or AWR reports is that they both can contain a Cache Advisory sub-section that is used to actually determine if adding memory will help your system. By examining waits and looking at the cache advisory section of the report you can quickly determine if adding memory will help your performance. Another source of information about the database cache is the V$BH dynamic performance view. The V$BH view contains an entry for every block in the cache and with a little SQL against the view you can easily determine if there are any free blocks or if you have used all available and are in need of more. Of course use of the automated memory management features in 10g and 11g limit the usefulness of the V$BH view. In Oracle Grid and Database control interfaces (providing you have the proper licenses) you also get performance advisories which will tell you when you need more memory. Of course if you have already maximized the size of your physical memory, most of this is moot.
Question 3 may have you scratching your head. Essentially if your disk IO subsystem has reached its lowest latency and the number of IO channels (as determined by the number and type of host bus adapters) is such that no channel is saturated, then your disk-based system is optimized. Usually this is shown by latency being in the 3-5 millisecond range and still having high IO waits with low CPU usage.
Question 4 means you look at you system CPU statistics and you see that your CPUs are being under utilized and the IO waits are high, indicating the system is waiting on IO to complete before it can continue processing. A Unix or Linux based system in this condition may show high values for runqueue even when the CPU is idle.
SSD Criticisms
Other critics of SSD technology cite problems with reliability and possible loss of data with Flash and DDR technologies. In some forms of Flash and DDR they are correct; if the Flash isn’t wear leveled properly or the DDR is not properly backed-up. However, as long as the Flash technology utilizes proper wear leveling and the RAM-DDR system uses proper battery backup with permanent storage on either a Flash or hard disk based subsystem, then those complaints are groundless.
The final criticism of SSD technology is usually that the price is still too high compared to disks. I look at an advertisement from a local computer store and I see a terabyte disk drive for $99.00; it is hard for SSD to compete with that low base cost. Of course, I can’t run a database on a single disk drive. Given our 300 gigabyte system, I was hard-pressed to get reasonable performance placing it on 28 – 15K high performance disk drives; most shops would use over 100 drives to get performance. So on a single disk to SSD comparison yes, this cost would appear to be an issue; however, you must look at other aspects of the technology. To achieve high performance most disk-based systems utilize specialized controllers and caching technology and spread IO across as many disk drives as possible. This is known as short-stroking the drive so that only 20-30% of each disk drive is ever actually used. The disks are rarely used individually, instead they are placed in a RAID array (usually RAID 5, RAID 10, or some exotic RAID technology). Once the cost of additional cabinets, controllers, and other support technology is added to the base cost of the disks, not to mention any additional firmware costs added by an OEM, the costs soon level between SSD and standard hard drive SAN systems.
In a review of benchmark results the usual ratio between needed capacity and capacity utilized to achieve performance is 40-50 to 1, meaning our 300 gigabyte TPC-H system would require at least 12 terabytes of storage to provide adequate performance spread over at least 200 or more disk drives. To contrast that, an SSD based system would only need a factor of 2 to 1 (to allow for the indexes and support files).
In addition to the base equipment costs, most disk arrays consume a large amount of electricity which then results in larger heat loads for your computer center. In many cases the SSD technology only consumes a fraction of the energy and cooling costs of regular disk based systems, providing substantial electrical and cooling cost savings over their lifetimes. SSD by its very nature is green technology.
When Doesn’t SSD help?
SSD technology will not help CPU bound systems. In fact, SSD may increase the load on overworked CPUs by reducing IO based waits. Therefore it is better to resolve any CPU loading issues before considering a move to SSD technology.
In Summary
The basic rule for determining if your system would benefit from SSD technology is that if your system is primarily waiting on IO then SSD technology will help mitigate the IO wait issue.
Saturday, January 31, 2009
Scuba Diving New York City
Perhaps my most wild eyed suggestion was to place a solar shield between the Earth and the sun to reduce the amount of solar energy reaching the Earth. Oddly the most wild eyed suggestion seems to be the only one that would make a difference at all. Unless we find a way to reduce the amount of solar energy reaching the Earth’s surface we can expect global temperatures to increase steadily until there are no ice caps, Arctic or Antarctica.
The net effect of the melting of the melting of the Arctic ice cap would be negligible since it is in effect floating so the change in sea levels would be near zero, however, the polar bears would have a few issues. The biggest problem would be the Antarctic ice sheet which is resting on the continent of Antarctica. As it melts and adds to the water levels the weight compressing the Antarctic land mass decreases causing rebound. Between the added water level and the displacement from rebound we are talking over 100 feet of additional water levels around the world. Say good bye to New York, London, Hong Kong, Tokyo, almost all of Florida, heck, most of our seaports. You think Bangladesh has problems with flooding now, just wait.
We need to start planning and doing something now. By reducing the amount of sunlight we receive by 10% we could nip this issue in the bud before 42nd street is considered an advanced level scuba dive. This would be done by orbiting a single large sunscreen or multiple smaller sunscreens in the L1 Lagrange point. Maybe by making these sunscreens actually into thermionic generators or simple solar cell arrays we could also provide large amounts of energy which could be sent back to Earth via microwave beams for use as a green energy source(http://gltrs.grc.nasa.gov/reports/2004/TM-2004-212743.pdf). To block 10 percent of the energy of the sun that reaches the Earth sounds a bit crazy but it could be done.
Even by spreading large clouds of metallic debris (how about all those aluminum cans we see lining the roadways?) or using a few large asteroids pushed into place using low thrust ion drives (http://www.grc.nasa.gov/WWW/ion/) .
It is time to look out for all of us and put aside petty (in the scheme of global warming scale disaster) differences and pull together to really do something to fix this problem. Driving a fuel efficient car and turning your thermostat down in the winter and up in the summer may give you warm and fuzzy (or cold and fuzzy depending on the season) feeling but it won’t amount to a hill of beans when it comes to helping fix global warming.
Wednesday, January 07, 2009
Database Entomology in Oracle11g Land
As a pre-test I had created a small TPC-H (not more than 30 gigabytes) and on a single-server rig I could run the complete 22 query TPC-H set of queries in parallel query. Of course I wasn’t using high levels of partitioning and parallel query simultaneously with the small data set.
I knew it would be interesting when I got to query 9 on the 300 gigabyte dataset and on query 9 I got an ORA-00600:
ORA-00600: internal error code, arguments [kxfrGraDistNum3],[65535],[4]
When I ran an 8 stream randomized order query run I also periodically received:
ORA-12801: error signaled in parallel query server P009, instance dpe-d70:TMSTPCH3 (3)
ORA-00600: internal error code, arguments: [qks3tStrStats4], [], [], [], [], [], [], []
On query 18 once in a while but not every run. Due to not being a full customer (only having a partner CSI number) I was unable to report these as possible bugs. I did completely check out OTN and Metalink as well as Google and no one else seems to be having these issues, of course how many folk are running Oracle11g 11.1.0.6 or 11.1.0.7 with RAC, cross instance parallel query and heavy partitioning and sub-partitioning?
I had a quick look at the Oracle11g 11.1.0.7 release notes and saw a load of bug fixes and hoped mine were covered, even though a text search didn’t show up the ORA-00600 arguments I received. So I bit the Oracle bullet and performed an upgrade.
Well, actually 2 sets of upgrades. First I upgraded my home, 32-bit (this is important later) servers and other than the usual documentation gotchas and needing to set the database as exclusive, start it, stop it and then reset it as a cluster before the dbua program would run properly, I was successful and now have an Oracle11g 11.1.0.7 instance running on my home RAC setup. A quick run against the 30 GB database showing no really stellar improvements against my test setup using JBOD arrays for TPC-H and it successfully ran Query 9 against the non-partitioned, smaller 30 gigabyte data set.
Feeling pleased with the success I immediately set out the next morning to update my large test environment, my 64 bit cluster. The CRS update went smoothly, and other than some space issues (you may want to add a datafile to your SYSTEM tablespace) and some package problems (for some reason DBMS_SQLTUNE and DBMS_ADVISOR where missing) the database upgrade went fine, right up to the point of starting the instances under 11.1.0.7. It seems there is just a small bug with the use of the new MEMORY_TARGET parameter and release 11.1.0.7…you can’t go above 3 gigabytes! This is why I said that the upgrade on 32 bits was important to remember, in 32 bit systems you will rarely get above a 3 gigabyte SGA size once you allow for user logins, process space and operating system memory needs. However, one of the major reasons for going to 64 bit is to have SGA sizes in Oracle greater than 4 gigabytes. Now, if you go back to using the SGA_MAX_SIZE and SGA_TARGET or the full manual specifications such as SHARED_POOL_SIZE and DB_CACHE_SIZE you can get above the 3 gigabyte setting.
Another annoying thing with 11g and the MEMORY_MAX_SIZE setting is that you cannot exceed MEMORY_MAX_SIZE with the sum of your SGA settings plus PGA_AGGREGATE_TARGET. Now for those of you with small sort sizes this isn’t really a problem and you probably won’t have any issues. However, with a TPC-H you need a large sort size so you need a large PGA_AGGREGATE_TARGET but, you quickly get into trouble with the limits on MEMORY_MAX_SIZE and large PGA_AGGREGATE_TARGETS. With my 16 gigabytes of memory per server I was only able to allocate about 8 gigabytes to Oracle (actually about 7.5 in 11.1.0.6 and 3 in 11.1.0.7) anything larger and I would get errors. So needless to say, I turned off the total memory management and did it the old fashioned way.
Finally, I had my instances up, about a 7 gigabyte SGA with 5 gigabytes of DB_CACHE_SIZE and 5.5 gigabytes of PGA_AGGREGATE_TARGET. Oh, did I mention, your SHARED_POOL_SIZE must be at least 600 megabytes for Oracle11g 11.1.0.7? If it isn’t you will get 4030 errors on startup, I ended up with 750 megabytes worth.
So after 8 hours of upgrade time for my main set of instances I was finally ready to run a TPC-H. Guess what, with the upgrade in place, more cache and bigger sort areas, I seem to be getting worse performance than with the sub-optimal query resolving, bug-ridden 11.1.0.6 version. Looks like they fed the bugs instead of killed them. Oh well, back to the tuning bench.
Thursday, December 04, 2008
Happy Holiday Shopping - Not!
Sears and Others:
http://money.cnn.com/2008/11/28/technology/bc.apfn.tec.holidayshop.ap/index.htm
http://tasquatch-sentinelling.blogspot.com/2008/11/blog-post_7519.html
Dr. Pepper:
http://www.bizjournals.com/atlanta/stories/2008/12/01/daily45.html
In order to understand what is happening you must understand what occurs when a user attempts to access a website, let alone when they attempt a complex transaction. When a user logs in to a website the user identification must be validated from a database using several database queries. For example: Is the user ID valid? Is the password correct? Has the password expired? What type of user is this? etc. It is even worse if the user has to create a login and password as well as enter other data such as address or credit data. All of this transactional traffic causes a flurry of underlying IO subsystem activity and web traffic across the networks. And of course all of this is magnified when actual query and sales transactions are also being performed.
All of the traffic to the system database can overwhelm the underlying IO subsystem, especially when it is disk array based which is the case with a majority of databases. Disks generally can only respond within a 5 millisecond window. Now, with on-disk caching and large caches in the disk arrays this response time can sometimes be reduced to 1 millisecond but as the caches flood with high activity performance generally drops to 5 milliseconds per IO or more. As the time to respond increases the number of users which can be served drops in a direct example of Little’s Law.
Unfortunately, while you can increase the IOPS (input output per second) by increasing the number of disks in an array, you cannot decrease the latency beyond that of any one disk. In fact, for a large read you will suffer from convoy effect where the slowest disk in the array involved in the IO operation will drive down the performance of the IO itself.
Several websites have found that by placing the user tables and other transaction dependent tables on low latency storage such as solid state SAN replacements like the RamSan 400 or 500 series from Texas Memory Systems they can dramatically increase the capability to support many more transactions (read clients) than before. An example of the dramatic improvements that can be achieved is shown in the recent press release from The Container Store group:
http://www.superssd.com/pressrelease/2008-12-02.htm
Realizing that their underlying IO subsystem couldn’t support the expected peak load generated from a positive blurb about the company on The Oprah Winfrey Show, the folks at The Container Store turned to TMS RamSan technology.
Another area that must support high concurrent logins and maintain strict inventories is in the online gaming community. Eve Online was able to go from 15,000 online users to over 17,000 with a 40x performance improvement:
http://www.superssd.com/success/ccpgames.htm
As a final example, IC Source first tried doubling the number of disks in their infrastructure; the net result? Zero improvement only to solve the problem with SSD (RamSan) technology:
http://www.superssd.com/success/icsource.htm
What all of these examples show is that many times your problem cannot be solved by throwing more disks at it, you must get to the ultimate problem, latency, to fix the issue.
With latency numbers of 15 microseconds (.015 milliseconds) and the resulting ability to support 600,000 IOPS and 4.5 GB/sec the RamSan-440 is the heavy hitter in the TMS line as far as throughput however, it is DDR RAM based and currently limited to 0.5 terabytes of storage capacity per unit. Compared to normal disk latency of 5 milliseconds the RamSan-440 shows a factor of 333 decrease in latency, even if you get 1.0 millisecond latency due to short-stroking and aggressive caching the 440 is still a factor of 67 times faster.
http://www.superssd.com/products/RamSan-440/
The RamSan-500 series utilizing Flash memory technology tops out at 2 terabytes (with promises to go to 8 terabytes in the near future) of storage capacity per unit and 200 microsecond ( 0.2 milliseconds ) peak latency with a minimal IOPS rating of 100,000 IOPS. Just to put this in perspective, EMC recently achieved 100,000 IOPS using 2-CX30 racks and 391 disk drives in a RAID0 configuration, the RamSan-500 does it with a single 4-U unit, and the RamSan-440 beats it by a factor of 6 in a similar footprint as the RamSan-500.
http://www.superssd.com/products/RamSan-500/
Being solid state (except for the cooling fans) means the RamSan technology is inherently more reliable and less prone to crashes. Current estimates of MTBF show a value of at least 500,000 hours per RamSan before a critical failure. With built in RAID, write leveling for Flash and the use of ECC memory as well as ChipKill the RamSans have built in redundancy. Utilizing Flash drives for backup, the 440 also provides unparalleled data persistence with triple battery backup ensuring that all data is written to Flash before shutdown. The RamSan-500 also uses battery backup to ensure that the 64 gigabytes of DDR cache is written to Flash on shutdown.
Another big movement is the green technology push we see today. I’ll leave the math on figuring the amount of electrical and cooling costs disk arrays to you, but at under 300 watts for the RamSan-500 and 600 watts for the RamSan-440 it is easy to see the cost savings from the energy footprint reduction.
So what should you take away from all of this? Essentially, if you have reached the latency limit of your IO subsystem, increasing the number of disks will not help. The only way to improve the performance of the system is to reduce overall latency. If you are an online retailer looking at the coming holiday season with dread because of performance issues, look at using SSD technology such as the TMS RamSan 400/500 series to slay the latency monster and achieve stellar website performance.
Saturday, November 08, 2008
From Mike at 30,000 feet
I have also been busy creating a demonstration database to show a side-by-side comparison of the differences in performance between disk and RamSan, after all seeing is believing as the old saw goes. I am writing this entry while flying over the mountains of Virginia and North Carolina on my way home from giving a one day tuning seminar in Stirling, Virginia at Oracles offices there under the auspices of the NatCap Oracle group. I leave again next week to give the same seminar to the Dallas Oracle User Group at the Oracle office in Plano Texas and from there go on to the Super Computing conference in Austin and then finally home again to Alpharetta for the Thanksgiving holiday.
I am also compiling a set of statistics to show the IOPS/gigabyte needed by Oracle databases of various sizes, in this task I will be gathering what historical data I have from past clients and if any of you have that type of data for your systems I would love to have it to add to the mix. If I have time I will try to come up with a query that shows this for a system and will post it. I imagine it will involve a rather simple sum of gigabytes (maybe just the active ones from dba_segments) and a v$sysstat capture of read and write IOPS. Here is a simple first cut, no doubt someone can make it simpler:
set serveroutput on
SET FEEDBACK OFF
col tod new_value now noprint
select to_char(sysdate, 'ddmonyyyyhh24mi') tod from dual;
TTITLE 'IOs Per Second'
col null noprint
select null from dual;
spool io_sec&&now
declare
cursor get_io is select
nvl(sum(a.phyrds+a.phywrts),0) sum_io1,nvl(sum(b.phyrds+b.phywrts),0) sum_io2
from sys.gv_$filestat a,sys.gv_$tempstat b;
cursor get_gig is select
sum(bytes)/(1024*1024*1024) from dba_segments;
gig number;
now date;
elapsed_seconds number;
sum_io1 number;
sum_io2 number;
sum_io12 number;
sum_io22 number;
tot_io number;
tot_io_per_sec number;
fixed_io_per_sec number;
temp_io_per_sec number;
iopsgig number;
begin
open get_io;
fetch get_io into sum_io1, sum_io2;
open get_gig;
fetch get_gig into gig;
close get_io;
close get_gig;
select sum_io1+sum_io2 into tot_io from dual;
select sysdate into now from dual;
select ceil((now-max(startup_time))*(60*60*24)) into elapsed_seconds from gv$instance;
fixed_io_per_sec:=sum_io1/elapsed_seconds;
temp_io_per_sec:=sum_io2/elapsed_seconds;
tot_io_per_sec:=tot_io/elapsed_seconds;
iopsgig:=tot_io_per_sec/gig;
dbms_output.put_line('Elapsed Sec :'to_char(elapsed_seconds, '9,999,999.99'));
dbms_output.put_line('Fixed IO/SEC :'to_char(fixed_io_per_sec,'9,999,999.99'));
dbms_output.put_line('Temp IO/SEC :'to_char(temp_io_per_sec, '9,999,999.99'));
dbms_output.put_line('Total IO/SEC :'to_char(tot_io_Per_Sec, '9,999,999.99'));
dbms_output.put_line('Total Used Gig:'to_char(gig, '9,999,999.99'));
dbms_output.put_line('Total IOPS/Gig:'to_char(iopsgig, '9,999,999.99'));
end;
/
spool off
ttitle off
set feedback on
Here is what the output show look like and as you can see it will generate a report as well:
Fri Nov 07 page 1
IOs Per Second
Elapsed Sec : 132,900.00
Fixed IO/SEC : 3.93
Temp IO/SEC : .01
Total IO/SEC : 3.94
Total Used Gig: 1.34
Total IOPS/Gig: 2.94
Well, I will sign off for now, hope to see some of you at the Dallas Bootcamp seminar and more at the Supercomputing conference.
Monday, October 06, 2008
Two Tales
Tale 1:
In 1979 or so a man who called himself R.C Christian (not his real name) came to the office of the Pyramid Quarry in Elberton, Georgia and contracted them, as the representative to several private individuals, create a large granite monument (the town of Elberton is the self proclaimed Granite capital of the world.) The design and location of the monument was provides and in March of 1980 it was erected by the Elberton Granite Finishing Company at a cost of approximately $41 million. Known as The Georgia Guidestones it sits outside of Elberton, Georgia on one of the tallest hills in Elbert county Georgia.

The Georgia Guidestones
The monument consists of 4-standing stones with a central Gnome stone and capstone. The faces of the standing stones are all inscribed with the same set of guidelines, each in a different “modern” language:
Maintain humanity under 500,000,000 in perpetual balance with nature
Guide reproduction wisely — improving fitness and diversity
Unite humanity with a living new language
Rule passion — faith — tradition and all things with tempered reason
Protect people and nations with fair laws and just courts
Let all nations rule internally resolving external disputes in a world court
Avoid petty laws and useless officials Balance personal rights with social duties.
Prize truth — beauty — love —seeking harmony with the infinite
Be not a cancer on the earth — leave room for nature — leave room for nature
Note that the last guideline has “leave room for nature” repeated twice on each of the 4 monoliths eight total sides with the exception of the one in Russian, where it looks like they ran out of space. The monoliths have the message in the following languages, moving clockwise around the structure from the north: English, Spanish, Swahili, Hindi, Hebrew, Arabic, ancient Chinese, and Russian.
On the capstone is the message, in Babylonian Cuneiform (north), Classical Greek (east), Sanskrit (south), and Egyptian Hieroglyphs (west), and the stone embedded in the Earth on the wesern exposure provides what is supposed to be an English translation: "Let these be guidestones to an age of reason." It is rumored to be engraved on the top of the capstone but as I am not 20 years old anymore I didn't try to get up there and see!
The guidestones messages and capstone message give a tip of the hat to Thomas Paine and other “radical” authors who espoused these types of rules for life in ages past.
The monument itself is aligned on several astronomical axis, it has a slit for viewing the sunset, a hole to look at the North Star (if you move to the left and squint as the monument is supposedly two degrees off proper placement) and a hole in the capstone that marks noon throughout the year. The width of the monument supposedly aligns with the migration of the moon throughout the year.

The setting sun through the slit
The sight of the setting sun through the aperture provided is quite fetching. Like I said, the North Star (“Celestial North”) is also visible through a hole in the center slab, if you move the left and look carefully through the hole due to a misalignment of the monument. I wasn’t there for noon so I couldn’t verify the noon time affect but I am assured it works well.
Now, tale 2.
As I was wandering around the site video taping and taking pictures an older gentleman and his wife drove up, he got out and his wife waited in the truck. Tall (about 6 foot) and a little heavy although not overly so, he reminded me of several rural Georgia farmers I have met in my wanderings. He told me that he had heard someone had vandalized the stones and wanted to see how badly (you can see what appears, upon close inspection, to be some type of clear polymer resin splashed on two of the stones.) We struck up a conversation and I asked him about the mystery surrounding who had them erected. He asked if I had heard the story and I said I read it on the web. He then stated:
“All of that is bullshit.” And smiled at me. “I worked at the quarry when this was built, Mr. Fendley himself had it built and made up that story” Mr. Fendley is/was the owner of the Pyramid Quarry. “He liked publicity and Lord knows this gave it to him.”
He also said he confronted Mr. Fendley about the monument and “He didn’t deny it, he looked mad, but kind of half smiled at me.”
Seems Joe F. Fenley Sr., Wayne Mullinex (the man who donated the land) and Wyatt C. Martin, President of the Granite City Bank involved with the financing were all Shriners (and therefore Masons) and agreed to do this together.
Now even if you take out having to pay for the granite itself (119 tones give or take) and the cost of the land, the cost of paying the workers to extract, cut, shape and carve the custom sized stones as well as place them is still rather high just for a publicity stunt, however, I wouldn’t put it past some folks. Also, the alignments and other parts of the monument are intriguing, why go to such effort when just the slabs themselves and their message would have been enough?
To add to the message there is supposedly a time capsule buried (or will be buried) under the slab that is on the west side of the monument. The slab explains the monuments message, who supposedly built it and details of the stones themselves. The “six feet under” makes me wonder if it isn’t someone who will be buried, hence the missing dates.

Inscription about the time capsule, notice missing dates.
All in all the stones looked like they were rather hurriedly inscribed because some of the inscriptions overrun the finished part of the slabs, there are translation errors and character errors on some messages and the lack of the repeat on the Russian stone due to running out of space. On the stone showing noon for several different parts of the year, the final marker overruns the finished part of the stone, was this a mistake, or is that marker more (or less) significant than the others?. Also, the North Star viewing hole appears to have been done as two holes drilled to run into each other, unfortunately they are off access to each other giving the hole a curved appearance and the noon hole looks like it was added as an afterthought.
Some have commented that if you wanted people to see the stones why not place them in a more significant location than a rural Georgia county? Well, for one thing, the sunset view and the North Star views probably would be difficult to guarantee in the heart of Atlanta or in areas where development might throw a skyscraper in the way. Besides if this really is a message for post-Armageddon survivors you don’t want it near any ground –zero targets.
Final “noon” marker overruns finished stone (circled in red)
Now, do the “mistakes” add up to rushed workers or are they deliberate? Are they sending us a message? I’ll leave that to the “DiVinci Code” fans to determine.
So, which tale is the true one? I know I have my opinion although the romantic side of me leans towards the mysterious group of strangers. For those close enough, go see this “America’s Stonehenge” yourself, for those who can’t, look them up on the net. I’ll be posting a bunch of other photos from my visit at my http://www.scubamage.com/ site as soon as I finish processing them, I’ll provide a more substantial link at that time.
Monday, September 29, 2008
Oracle's Data(Warehouse)base Machine
The biggest news at the conference was Larry Ellison’s announcement of the Exadata storage concept and the Oracle Database Machine both developed jointly with HP. These new storage and database devices offer up to 168 terabytes of raw storage with 368 gigabytes of caching and 64 main CPUs in 8 stacked DL 360 G5 servers and each Exadata unit has a HP Proliant DL 180 G5 with dual quadcore CPUs, 8 gigabytes of memory and 12 SAS 300 GB or SATA 1 terabyte drives. The entire HP Oracle Database Machine contains 14 Exadata blocks and 8 – dual quadcore servers in a full configuration. The Exadata blocks can be purchased separately. There are 4-24 port Infiniband switches provided in the Database Machine. The entire device provides a throughput of 10.5 (SATA) to 14 (SAS) GB/second.
Now, each Exadata block can only provide 1 terabyte if the 300 GB drives are utilized and 3.3 terabytes if the 1 terabyte drives are used unless Oracle compression is also used. This space calculation (from Oracle documentation) is based on mirroring of all the drives and subtracting space for logs, undo and temp space. The usual “your mileage may vary” warning applies to this available space. ASM with what appears to be high redundancy storage is being used to manage the drives. So while raw storage appears to be 3.3 TB to 12 TB the actual space that ends up being usable is only 1/3 of those amounts. Each Exadata has 2 – 20 Gigabit Infiniband interfaces. However, the blocks can only support 1 GB per second of output with the SAS configuration and 750 MB per second in the SATA configuration.
The Oracle Database Machine was actually designed for large data warehouses but Larry assured us we could use it for OLTP applications as well. Performance improvements of 10X to 50X if you move your application to the Database Machine are promised. This dramatic improvement over existing data warehouse systems is provided through placing an Oracle provided parallel processing engine on each Exadata building block so instead of passing data blocks, results are returned. How the latency of the drives is being defeated wasn’t fully explained.
The HP Oracle Database Machine must run Oracle11g, 11.1.0.7, RAC and Linux and each Exadata block must have the new Oracle 11g parallel query engine installed. So in a full configuration you are on the tab for a 64 CPU Oracle and RAC license and 112 Oracle parallel query licenses (assuming it is per CPU, if it is per Exadata block then it will be 14) as well as any Grid control licenses you may need. The base cost of the full Database Machine is around $650K which seems quite a bargain for 14-46 terabytes of usable storage and a 64 processor stack, however, you will also need over a million dollars in licenses even with aggressive reductions from your sales representative.
The HP-Oracle Database Machine only works with Oracle databases (just thought I should throw that in.)
Whew! The HP-Oracle database Machine offers quite an impressive array of facts, figures, promises and price tags. It will be interesting to see how this all sorts out over the coming months. Will there be enough profit in the HP-Oracle Database Machine to keep HP interested? Or is this another network computer? For those too young to remember Larry’s last hardware foray was the Network Computer, a device that would replace all the desktops and centralize application and data storage, it failed. [Note: don’t forget about Pillar Data another Ellision investment that would seem to be hurt by this announcement]. I believe this system is designed to help Oracle protect turf from Netezza and promote growth in the analytics market. Targeting the product to OLTP environments is just sloppy marketing as the system will not offer the latency needed in real OLTP transaction intensive shops. These applications do not need parallel queries, they need low latency database writes and reads.
So for nearly 2 million dollars (licenses plus hardware) you get a dedicated Oracle server in a rack with 64 CPUs of central processing and 46 terabytes of usable storage managed by 112 block resident CPUs and a wee bit less than 224 gigabytes of cache area (168 gigabytes of cache were promised after processing overhead was subtracted). However, you must throw away your existing infrastructure, upgrade to Oracle11g and marry your future to the HP Oracle Database Machine to do so.
What might be an alternative? Well, how about keeping your existing hardware, keep your existing licenses, and just purchase solid state disks to supplement your existing technology stack? For that same amount of money you will shortly be able to get the same usable capacity of Texas Memory Systems RamSan devices. By my estimates that will give you 600,000 IOPS, 9 GB/sec bandwidth (using fibre Fibre Channel , more withor Infiniband), 48 terabytes of non-volatile flash storage[S1] , 384 GB of DDR cache and a speed up of 10-50X depending on the query (based on tests against the TPCH data set using disks and the equivalent Ram-San SSD configuration). More importantly, this performance can be delivered with sub-millisecond response time. At Oracle World I was presenting on the importance of latency to Oracle databases. The Exadata is massive and offers great bandwidth but will have nearly awful disk access tedencies due to the massive and slow disk drives included in the system.
Of course, if you don’t need 48 terabytes you can purchase Ram-San SSD technology from 32 gigabytes up to whatever you actually need. Ram-San SSD technology works with all Oracle versions and requires no special licensing or changes to your system. Ram-San uses fibre channel or Infiniband for connection to your infrastructure and looks identical to a disk drive once configured (it takes about 10 minutes.)
Oh, and the Ram-San doesn’t care if you are on Oracle, SQL Server, MySQL, or cousin Joe’s Ozark Mountain special database.
So let’s recap:
Purchase the complex, tied to Oracle with a golden chain, HP-Oracle Database Machine and end up throwing away your existing technology stack, and spend up to 2 million dollars for the privilege, for a speed up of 10x to 50x on data warehouse type queries.
Or
Purchase just as much Ram-San SSD technology as you need ($43K base price for 32 GB mirrored), keep your existing hardware and license structure (or possibly reduce it) and get a 10x to 50x speed up on data warehouse type queries, with the freedom to change databases as you need to.
Call me simple, but I think I see the proper choice.
[S1]Note that the 8TB systems are projected to have 1.5GB/second of bandwidth sustained reads or writes (to Flash). This can be accomplished with four 4Gbit FC ports or two 4x IB ports per unit.
Tuesday, September 16, 2008
Is there a DBA in the House?
All of the problem solving that occurs on House reminds me of trouble shooting in the Oracle world. In many cases there are a number of symptoms with database problems and some of them are contradictory just as with problems in the human health areas. Oh, and everybody lies. “No, there weren’t any changes”, “No, nothing is different between these two test runs”, ”Yes, we used the same data/transactions/parameters”.
In almost every episode of House they go through at least three different “cures” before they find the real problem and solution. Many times in the Oracle universe we apply a fix for a problem, only to find it wasn’t really the issue, or, in fixing it we transfer the problem to another area of the database. Another similarity is that many times House and his team will take a shotgun approach when there isn’t a clear solution, applying two or more “cures” at the same time, much like a DBA will apply multiple fixes in a single pass, thus not really knowing what was fixed but just breathing a sigh of relief when performance improves.
I think the character portrayed as Dr. House would make a great DBA, but somehow I can’t see folks glued to their TV screens hoping that next index will fix the query…
Thursday, September 04, 2008
SSD is Green Technology
But just how much can be saved? In comparisons to state of the art disk based systems (sorry, I can’t mention the company we compared to) at 25K IOPS, 50K IOPS and 100K IOPS with redundancy, SSD based technology saved from a low of $27K per year at 6 terabytes of storage and 25K IOPS to a high of $120K per year at 100K IOPS and 2 terabytes of storage using basic electrical and cooling estimation methods. Using methods documented in an APC whitepaper the cost savings varied from $24K/yr to $72K/yr for the same range. The electrical cost utilized was 9.67 cents per kilowatt hour (average commercial rate across the USA for the last 12 months) and cooling costs were calculated at twice the electrical costs based on data from standard HVAC cost sources. It was also assumed that the disks were in their own enclosures separate from the servers while the SSD could be placed into the same racks as the servers. For rack space calculations it was assumed 34U of a 42U rack was available for the SSD and its required support equipment leaving 8U for the servers/blades.
Even figuring in the initial cost difference, the SSD technology paid for itself before the first year was over in all IOPS and terabyte ranges calculated. In fact, based on values utilized at the storage performance council website and the tpc.org website for a typically configure SAN from the manufacturer used in the study, even the cost for the SSD was less for most configurations in the 25K-100K IOPS range.
Obviously, from a green technology standpoint SSD technology (specifically the RamSan 500) provides directly measurable benefits. When the benefits from direct electrical, space and cooling cost savings are combined with the greater performance benefits the decision to purchase SSD technology should be a no brainer.
Sunday, August 24, 2008
A Tale of Two Databases
In my own laboratory (no, I don’t have an Igor, but he would come in handy to help dispose of dead computers now and then) I am running a two-node RAC cluster with each cluster node being a single 3 gigahertz 32 bit CPU with hyperthreading and 3 gigabytes of memory. The cluster interconnect is a single gigE Ethernet. As to storage, I utilize two JBOD arrays, a NexStor 18F and a NexStor 8F each fully loaded with 72 gigabyte 10K Seagate SCSI drive and connected to the servers via a Brocade 2250 with dual QLA2200 1gb HBAs on each server. The arrays themselves are connected through single 1 gb fibre channel connections. Oh, it is also running RedHat 4.0 Linux.
On the system in the remote lab, let’s call it TMSORCL for ease of identification, I have 4-3.6 Ghz CPUS, 4-2 port QLA2642 4 Gb HBAs (I used 1) and 8 gigabytes of memory. I have Oracle10g 10.1.0.2 installed, on the home lab (let’s call it AULTDB) I have Oracle11g, 11.1.0.2 installed.
AULTDB is utilizing ASM on two diskgroups. One diskgroup consists of 16 drives, 8 on the 18F and 8 on the 8F using normal redundancy and failgroups, the other diskgroup uses 8 drives on the 8f and uses external redundancy. I have placed the application data files and index files on the datagroup with the 16 drives and the rest (system, temp, user, undotbs, redo logs) on the 8 disk externally redundant datagroup. Using ORION (Oracle’s disk IO simulator) I was able to achieve nearly 6000 IOPS using 24 drives, so I hope that I can get near that using this configuration.
TMSORCL is only using the RAMSAN 400 through its 4 gb fibre channel connection for all database files. The RAMSAN should be able to produce enough IOPS to provide at least 100,000 through the single HBA (if it doesn’t saturate that is.)
Using Benchmark Factory from Quest Software I created two identical databases (you guessed it, TMSORCL and AULTDB.) These are TPCH type environments with a scale factor of 18, the largest I could support and still have enough temporary space to build indexes and run queries.
Now, some may be saying that I have set up an apples to oranges environment for comparison, and they would be correct, however, many folks will be facing just such a choice very soon, that is, to stick with an Oracle10g environment or upgrade to an 11g environment. Another question that many folks have, should I go to RAC? So this test is not as useless as you may have first thought.
I set up both TMSORCL and AULTDB to be as near in effective size (memory settings wise) allowing that AULTDB was spread across two 3 gigabyte memory areas allowing for a total memory footprint of about 3-4 GB, I set up the TMSORCL environment to have a maximum size of 4 GB with a target of 3 GB for memory.
I had some interesting problems with the number 5 query in the TPCH query set until I added a few additional indexes, it kept blowing out the temporary tablespace even though I had it as large as 50 gigabytes. If you want the script for the additional indexes, email me. I also had issues with parallel query slaves not releasing their temporary areas. Anyway, after adding the indexes and incorporating periodic database restarts into the test régime I was able to complete multiple full TPCH power runs on each machine (just the single user stream 0 power run for now).
So, before we get to the results, let’s recap:
AULTDB – 32 bit RedHat 4 Linux RAC environment with two single 3 Ghtz CPUS running hyperthreaded to simulate 4 CPU total in 6 gigabytes of memory utilizing 4-1gb QLA2200 HBAs to access 24-10K 72 gigabyte drives using ASM. Effective SGA 4 gigabytes.
TMSORCL – 64 bit RedHat 4.0 Linux single server with 4 3.6 Ghtz CPUS with 8 gigabytes of memory utilizing 1-4 gb HBA port to access a single 128 GB RAMSAN 400. Effective SGA constrained to 4 gigabytes.
I ran the TPCH on each database (after getting stable runs) a total of 4 times. I will use the best run, as measured by total time to complete the 22 queries, for each database for the comparison runs. For AULTDB the best time to complete the run was 1:37:55 (One hour, thirty seven minutes and fifty-five seconds.) For TMSORCL the best time was 0:15:15 (zero hours, fifteen minutes and 15 seconds.) So on just raw, total time elapsed for identical data volumes, identical data contents and identical queries the TMSORCL database completed the runs 6.42 times faster (642%). The actual query timings in seconds are shown in the following chart. Based on summing the given query times the performance improvement factor from AULTDB to TMSORCL is 6.68 or 668% faster.

As you can see, TMSORCL beat out AULTDB on all queries with a range of 79% up to a whopping 4,474% improvement being shown based on the individual query times in seconds.
During the tests I monitored vmstat output at 5 second intervals, at no time did run queue length get over 3 and IO wait was less than 5% on both servers. This indicates that the IO subsystem never became over burdened, which of course was more of a concern with AULTDB rather than TMSORCL.
Now the TPCH benchmark is heavy on IOPS, so we would expect the database using the RAMSAN to perform better, and that is in fact what we are seeing, in spite of only having a single HBA and being on an older, less performing version of Oracle. So what conclusions can we draw from this test? Well, there are several:
For IOP heavy environments RAMSAN technology can improve performance by up to 700% against a JBOD array properly sized for the same application, depending on number and type of queries.
Use of a RAMSAN can delay moving to larger, more expensive sets of hardware or software if the concern is IO performance related.
Now there are several technologies out there that offer query acceleration, most of them place a large data cache in front of the disks to “virtualize” the data into what are essentially memory disks. The problem with these various technologies (including TimesTen from Oracle) is that there are coding issues, placement issues (what gets cached and what is left out?) and management issues, for example, with TimesTen there are logging and backup issues to contend with. In addition, utilities that use the hosts memory such as TimesTen add CPU as well as memory burden to what is probably an overloaded system.
What issues did I deal with using RAMSAN? Well, using the provided management interface GUI (via a web browser) I configured two logical units (LUNS), assigned them to the HBA talking to my Linux host and then refreshed the SCSI interface to see the LUNS. I then created a single EXT3 partition on each LUN and pointed the database creation with DBCA at those LUNs. Essentially the same exact things you would do with a disk you had just assigned to the system. The RAMSAN LUNs are treated exactly as you would a normal disk LUN (well, you have to grant execute permission to the owner, but other than that…) Now, if you don’t place the entire database on the RAMSAN then you have to make the choice of what files to place there, usually a look at a Statspack or AWR report will head you in the correct direction.
An interesting artifact from the test occurred on both systems, after a couple of repetitive runs the times would degrade on several queries, if I restarted the database, the times would return to near the previous good values. This artifact probably points to temporary space or undo tablespace cleanup and management issues.
I next intend to run some TPCC and maybe if I can get the needed infrastructure in place, a TPCE on each system, watch here for the results.
Saturday, July 19, 2008
Atoms to Plowshears (Literally and figuratively)
After the conference I had some time to kill so I decided to visit Magnuson State Park. It wasn’t an arbitrary decision just selected at random from the Seattle area map, I had heard (and seen some pictures) of a sculpture there and decided to visit there if I could as a result. My first set of directions actually took me to the artist’s house, I didn’t stop in and say hi, but just backtracked and went to the park.
Partial Shot of Sculpture In the picture shown above, if you are not sure what you are looking at, let me explain. From the 1950’s to the current date the USA has been building and using Nuclear Submarines, starting with the USS Nautilus, SSN 571 commissioned in 1954. Both search and destroy (fast attack) and stealth missile deployment (Ballistic Missiles) submarines utilize stern planes and sail planes (the “sail” is what landlubbers would call the conning tower.) The sail planes are also called the dive planes as they are used to cause the submarine to dive and surface when it is neutrally buoyant.
Seattle Fins: SSN 669 Seahorse, SSBN 641 Simon Bolivar, SSN 652 Puffer, SSN 615 Gato, SSBN 620 John Adams, SSN 595 Plunger, SSN 638 Whale, SSN 667 Bergall, SSN 673 Flying Fish, SSN 597 Tullibee, SSN 650 Pargo, SSN 662 Gurnard.
Miami Fins: Sea Devil SSN 664, Pogy SSN 647, Sand Lance SSN 660, Pintado SSN 672, Trepang SSN 674, Billfish SSN 676, Archerfish SSN 678, Tunny SSN 682, Von Steuben SSBN 632, Sculpin SSN 590, Cavalla SSN 684.
I served on two nuclear submarines during the period 1976-1979, the USS John Adams, SSBN 620 and the USS Bergall, SSN 667, so you can see my interest in the Seattle sculpture. The SSBN on the Adams number means she was a ballistic missile boat, we carried up to 16 Poseidon missiles with MIRV warheads (multiple independent re-entry vehicle, meaning each missile of the Poseidon class could hit multiple targets) with the nuclear capability that exceeded the explosive power of all the munitions used in WWII. All this was used to carry out the MAD (mutually assured destruction) doctrine between the USA, USSR and at times Communist (Red) China, although the main targets were predominantly in the USSR. The SSN means the Bergall was a fast attack submarine used to hunt and kill other ships, including hostile submarines.
The MAD concept was that the SSBN type submarines, being undetectable, would be unstoppable launch platforms that would be used to respond to any nuclear aggression from anywhere in the world. Thus assuring we could utterly destroy Russian civilization should they launch a first attack that succeeded in taking out our land based missile systems. The Russians spent a great deal of time, money and resources trying to find ways to beat the SSBN submarines, in no small part they were one of the key technologies that kept the Russians and Chinese from launching a first strike during the worst part of the cold war.
ComSubLAnt (Commander Submarine Atlantic) could communicate with us using radio and LFT (Low Frequency Transmissions.) They kept the encrypted traffic going 24X7 replacing any actual command traffic with 15 word family grams, news and other items to not allow the Russians the ability to sense something was happening by seeing increased communications traffic. Each sailor was only allowed a limited number of family grams per patrol, no reverse communication, from the sailors back to the families was allowed. A patrol lasted 3 months with most of that spent underwater on patrol and the balance in such sun-fun spots as Holy Loch, Scotland repairing what the other crew broke on their patrol. The SSBNs had two crews, the Golds and Blues, I was on the gold crew. You usually spent about 70-80 days underwater with no fresh air, no outside views and no females! Your biggest enemies where boredom and doing qualifications, you didn’t think about the hundreds of pounds per square inch of pressure that were striving to snuff out your life every second of every day while you were on patrol or you would go mad.
We were the warriors of the cold war. The cold war was officially over (at least most felt it was) when the Berlin wall was taken down in 1989, the submarine fleet was as much responsible for that as any president. We were away 3 months out of every 6 from our families, for these patrols, I did 5 patrols and a DASO run for a total of 18 months out of the 33 I spent on the Adams. At just about any time during those 18 months a worn seal, a broken valve, a busted pipe could have killed us all, as it did for the sailors on the two nuclear submarines that didn’t come back, the USS Thresher, SSN 593 and the USS Scorpion, SSN 589. Believe me, listening to the pings and squeals as we went to test and one time to crush depth was a bit unnerving when you realized how much pressure it took to do that to several inches of stainless steel pressure hull. The Russians lost several submarines during that time as well and now most of their fleet lies in ruins silently rusting away at the piers in Vladivostok and other Russian ports.
As I stood there and placed my hand against the only surviving part of the submarine that had guarded my (and your) life both directly when I was aboard her and indirectly through the MAD concept when I wasn’t I couldn’t help but feel a bit nostalgic and melancholy that such a fine ship met such an ignoble end as becoming feed stock for John Deere tractors except for one sail plane in this sculpture garden. Of course the transition from a ship of war to farm implements maybe has greater cosmic import that I realize. The Bergall had both of her sail planes here, but since I only spent a few months and never went to sea on her, I didn’t feel the connection I did with the Adams.
As I wandered the sculpture garden taking pictures I heard and watched a group of children playing on a nearby hill. Later from that same hill I watched them walk down through the sculpture garden toward the beach and right past the last intact piece of the USS John Adams. I wondered if any of them truly understood what that piece of steel really meant? Of course maybe it’s true purpose was so that they never again would have to live under the threat of nuclear annihilation of the entire planet. I hope someone explains it to them, so that the meaning is not forgotten.

The Author beside the USS John Adams SSBN 620 Sail Plane
Wednesday, July 02, 2008
Lies, Damn Lies and SSD Technology
Let’s look at some of the highlights:
1. Solid state drive technology is very expensive
2. Solid state devices are best when directly attached to the internal bus architecture
3. Solid state drives will only be niche players
4. You can get the same IO rate from disks as from SSD
First, the myth that solid state drives are expensive was, like many myths involving Oracle and computers, true at one time, however, times change. The huge leap in demand for flash memory with the advent of I-pods, digital cameras and video recorders has created a memory glut. You can get a 4 gigabyte flash memory stick or card for under a hundred dollars for your camera or other flash device. In fact memory prices promise to plunge even farther as mass production techniques and miniaturization technology improves. The cost for a gigabyte of enterprise class disk storage is around $84 at last count, for the most current version of the Texas Memory System RAMSAN SSD technology, using flash memory and regular memory, the cost is around $100 per gigabyte, with further decreases in memory costs, RAMSAN SSD prices will fall even further.
Second, in a recent article a producer of both disk and solid state technology seemed to indicate it works best when hooked directly into the internal bus for the computer and really wasn’t efficient when attached as a SAN would be attached. I am not sure where he is getting his information (other than his company is trying to shoe-horn solid state drive technology into their existing SAN infrastructure) but it has been my experience that rarely if ever do users flood the fibre channels, they may overload a couple of the disk drives, but generally the SAN connections are not the source of the bottleneck when it comes to SAN technology. Using standard fibre channel connections and standard host bus adapters Texas memory Systems achieves over 400,000 IOPS from a single 4U RAMSAN SSD. To get the equivalent IOPS using regular disk technology you would need over 6000 or more individual disk drives, the racks to hold them and the controllers to control them, not to mention the air conditioning and electrical power needed for that many disks.
Next, solid state drives will only be niche players, this is a ridiculous statement. Most clients of RAMSAN SSD technology use them just as they would disk arrays. The RAMSAN SSD technology will replace disks as we know it in the near future and disks will be relegated to second tier storage and backup duties, replacing tapes. Many experts are talking of the tier 0 level of storage and specifically mentioning SSD when they do so. When you can place a single 4U sized RAMSAN SSD into your system and replace literally hundreds or thousands of disks the idea that they will only be niche devices is foolish. This is especially true when you consider the decreasing costs, the ease of administration and the performance gains that you get when SSD technology is properly deployed.
Finally, the myth that with disks you can get the same IOPS as with SSD. Yes, you can, however, you would need X/(IOPS/disk) number of disks where X is the desired IOPS to achieve it, double that number for RAID10 or RAID01. Even high speed 15K drives can only deliver around 100 to 130 IOPS per second of random reads due to the mechanical nature of disk drives, as the late Scotty on the Federation Starship Enterprise used to say (about every other episode): “Ya cannot change the laws of physics.” Disks, without prohibitive cooling technologies, cannot exceed certain maximum rotational speeds, read heads, mounted on mechanical arms can only move so fast and the magnetic traces can only be packed so close on the disk surface. To get 400,000 IOPS you would need at least (400,000/130)*2= 6153 drives in a RAID10 array. At 18 drives per tray that is 341 trays of disk drives at 8 trays per rack that is almost 43 racks needed to hold the drives. Now, even with the largest caches available you still require anywhere from a millisecond to several (up to 5 with minimal loads, higher with large loads or more than single block reads) milliseconds to do each IO, this latency will always be there in a disk based system, the latency on SSD based systems such as RAMSAN are in the hundreds of nanoseconds range (fractional milliseconds).
So, what have we determined? We have found that SSD technology is comparable in cost with enterprise level disk systems and will soon beat the cost of enterprise level disk systems. We have also seen that SSD technology when properly designed and implemented (not shoe-horned into a disk-based SAN) will fulfill the promise of fibre technology and allow use of the bandwidth currently squandered by disk technology. We have also seen that far from being a niche technology, SSD is becoming the tier 0 storage for many companies and will soon supplant disks as the primary storage medium in many applications. Finally, while it is possible to achieve the same level of IOPS using disk technology that SSD technology provides, it would be cost prohibitive to do so, and, even if you did achieve the same level of IOPS, each IO would still be subject to the same disk based latencies.
I am not afraid to say it: SSD technology is here, it is ready for prime time and it is only a matter of time before disks are relegated to second tier storage. Disks are dead, they just don’t know it yet.
Friday, June 20, 2008
Is Hydrogen Really Green?
Since there is relatively no free hydrogen in nature, hydrogen has to be produced through electro-chemical or catalytic means. The most common means of hydrogen production is the electrolysis of common water, you take two water molecules (2-H2O) and combine them with a jolt of electricity and you get 2 hydrogen molecules (2H2) and one oxygen molecule (O2). In a loss-less system you could then take the hydrogen and recombine it with the oxygen and get back the energy used to separate the molecules back in the form of heat, (the reaction is exothermic, as in gives off heat) or in the case of a fuel cell electricity, and have a waste product of pure water. However, there is no such thing as a loss-less conversion going either way so it takes more energy to produce hydrogen gas than you get from burning it or using it to produce electricity in a fuel cell.
Hydrogen is much less dense (in liquid form) than gasoline, this means that while hydrogen provides more btu of energy per pound than gasoline, a pound of hydrogen takes up more volume. This density difference means that to take advantage of the increased BTU per pound you have to burn or convert a larger volume of hydrogen. How about a 60 gallon tank for your SUV? Hydrogen also must be kept compressed and/or insulated to prevent losses. Hydrogen tends to cause hydrogen embrittlement of most metals so long term storage is also an issue. Liquid natural gas lines could possibly be used to transport the liquefied hydrogen, however, better insulation and the embrittlement issues would have to be addressed before the existing infrastructure could be used safely.
One promising hydrogen storage technology uses metal hydrides such as zirconium hydride that allow storage of hydrogen in interstitial sites in the crystal lattice, some tests show that storage densities exceeding liquid densities by several fold are possible. The use of metal hydride storage would put the fuel tank back at the current size in your SUV and would provide added safety since the hydrogen is not in liquid or gas form.
It should be obvious that using fossil fuels such as natural gas, oil, or coal to generate electricity to create hydrogen would be a huge mistake. The amount of carbon dioxide (CO2), the major greenhouse gas, which would be created from using any fossil fuels to produce hydrogen would cause far more damage from CO2 emissions than any benefits gained from using the hydrogen thus produced. However, it waits to be seen if some enterprising third-world country (no doubt financed by mainline energy companies) doesn’t use massive coal burning to produce hydrogen for sale to the more industrialized countries.
This leaves us with wind, solar, tidal or nuclear power to provide the needed energy to produce hydrogen in sufficient quantities to make it a viable energy source. What are the economic considerations of each of these energy sources?
Use of wind power
On the surface wind power looks good. You put up a tower (or two, or a hundred) with a wind generator and get electricity and dump the electricity into an electrolytic cell that produces hydrogen. Of course some of the electricity needs to go into compressors to store the gaseous hydrogen, some needs to go into pumping and purifying the water being fed into the electrolytic cell. At current manufacturing costs electricity from wind runs 4-6 cents per kilowatt hour, when the wind blows, it isn’t raining to hard, freezing or being repaired because of lightening strikes. Wind also ties up huge amounts of real estate, causes noise pollution and has reliability issues. Plus to produce the nearly 250 gigawatts of energy needed to produce the amount of hydrogen gas to support just the needs of the USA to replace fossil fuels in transportation alone would require 12,500 2 megawatt wind turbines and a mere 5000 of the new 5 megawatt mega-turbines. Want one over the top of your house?
Use of Solar Power
Let’s look at solar power, it is quiet, produces no waste, perfect right? Not quite. Even if the new high efficiency cells pan out where we double or quadruple the efficiency of existing cells by using new technologies (to 40-60 percent conversion efficiency) we will still only have a cost of 8-10 cents per kw-H, more expensive than wind generation technology. At a solar constant of 1,395 watts per acre and a 60% conversion efficiency yields a back of the envelope calculation of 837 watts per acre. To provide the 250 gigawatts using solar we would need to clear 30 million acres of land and put in high efficiency solar panels, how does that grab you?
Use of Nuclear Power
Nuclear power can produce immense amounts of energy while requiring only small amounts of space. The 250 gigawatts required for the hydrogen gas production would require 150-160 new reactors to be built. At 100 acres per plant this is only 16,000 acres of land, as a comparison, Ted Turner’s ranch properties are estimated at over 1.9 million acres. As to the nuclear waste issues, the new designs for reactors promise to reduce waste and make better use of recycling of fuel. This won’t completely eliminate the nuclear waste problem but may make it more manageable. At 11.1 to 14.5 cents per kw-H it is currently one of the more expensive options until you consider the environmental costs of other technologies. However, using the heat from the nuclear process to facilitate steam methane reforming, biomass gasification or coal gasification to produce hydrogen more efficiently we could reduce the cost, increase the output and reduce the number of needed nuclear plants thus further mitigating the nuclear waste problem.
Summary
So, before we all leap upon the hydrogen technology bandwagon we need to step back and examine the real costs and the real technologies needed to make it a reality. Is hydrogen really a green technology? As with all things technology the answer is a fully qualified maybe.
Friday, May 30, 2008
Opportunities Part Two
For those not sure what Texas Memory Systems does, check out their website at http://www.texmemsys.com/. One of my duties will be to manage/monitor the StatsPack Analyzer website at http://www.statspackanalyzer.com/, if you have a statspack (or AWR) report you want analyzed, log on and upload it! Also, swing by the forums and have a go at improving our rules for statspack/AWR evaluation. We want to improve the Statspack Analyzer application and we value your feedback.
So, I was actually unemployed for 15 days (11 workdays) but I did have the offer within 36 hours a new record for me. I can’t imagine how I would feel being out of work for several weeks or months as 15 days was disturbing enough (after all, how much Oprah or Dr. Phil can one person take?) To all of those job hunting right now, stick with it, good luck and I hope you have as good a luck as I did in my search.
I am really looking forward to getting in the new equipment down in the dungeon where I can torture it to my hearts content with various tests, benchmarks and other procedures my devious mind can come up, I am looking forward to also publishing the results so everyone can see how really amazing is this SSD technology. Imagine a 146 times improvement in query performance. How about virtually no latency for redo log operations? Take a look at some of the client histories and white papers on the http://www.texmemsys.com/ to see real data.
Anyway, it is great to be part of the working masses once again, unemployment is hard work!
Sunday, May 18, 2008
Opportunity
During my 10 years as a Nuclear Chemist I watched the Nuclear industry grow from a growth industry, to a stable, to a declining one. Now of course with oil prices rising and CO/CO2 emissions on everyone’s mind the Nuclear industry is making a come back. No, I am not returning to the Nuclear industry as my skill set is nearly 20 years out of date. However, for the first time in 18 years I have been noticing certain indicators in the Oracle DBA marketplace that show it may be time for a change.
The indicators of course deal with the increased automation of the Oracle DBA job by Oracle coupled with the large supply of (at least on the low end) DBA services as outsource resources. It is difficult to compete with DBA resources that are happy to receive a small fraction of the salary you need to live on in the USA. I predict that the DBA job for Oracle will cease to exist as we know it within 3-5 years. Perhaps rather than being part of the large out flux of skilled but overpriced Oracle talent at this future date it is time to evaluate where you want to be in 3-5 years.
You may know (or you may not as they are keeping things rather under raps) Quest just went through a large round of layoffs in their Oracle and database areas. Yes, I was caught in the lay off and was officially put in the ranks of the unemployed on May 15, 2008. Perhaps it is a signal that it is time to move to another area of expertise. Don’t worry, I am landing on my feet and already have an excellent opportunity I am considering and unless something else shows up that is absolutely stellar, I will probably take it. The opportunity provides a path to stop feeding from the Oracle trough and move into an area that is more future-proof. I’ll keep you posted as I move forward with this new prospect.
Of course this is a record for me, in 35 years of working it is the longest I have gone without a job (so far 2 full work days) however, I did have the offer within 36 hours of being laid off. So, I plan to use this as an opportunity to step back and really consider where the computer industry is going and look to a position that will enable me to make full use of the skills I have acquired while moving forward into new and exciting areas. Opportunity is knocking, I think I’ll answer.
