{"id":2044,"date":"2009-07-03T08:25:45","date_gmt":"2009-07-03T13:25:45","guid":{"rendered":"http:\/\/www.daniel-lemire.com\/blog\/?p=2044"},"modified":"2016-02-12T01:00:04","modified_gmt":"2016-02-12T01:00:04","slug":"column-stores-and-row-stores-should-you-care","status":"publish","type":"post","link":"https:\/\/lemire.me\/blog\/2009\/07\/03\/column-stores-and-row-stores-should-you-care\/","title":{"rendered":"Column stores and row stores: should you care?"},"content":{"rendered":"<p>Most database users know row-oriented databases such as Oracle or MySQL. In such engines, the data is organized by rows.\u00a0Database researcher and guru <a href=\"http:\/\/en.wikipedia.org\/wiki\/Michael_Stonebraker\">Michael Stonebraker<\/a> has been advocating column-oriented databases. The idea is quite simple: by organizing the data into columns, we can compress it more efficiently (using simple ideas like <a href=\"https:\/\/en.wikipedia.org\/wiki\/Run-length_encoding\">run-length encoding<\/a>). He even founded a company, <a href=\"http:\/\/www.vertica.com\/\">Vertica<\/a>, to sell this idea.<\/p>\n<p>Daniel Tunkelang is <a href=\"http:\/\/thenoisychannel.com\/2009\/07\/02\/the-wild-world-of-sigmod\/\">back from SIGMOD<\/a>:\u00a0he reports that column-oriented databases have grabbed much mindshare. While I did not attend SIGMOD, I am not surprised. <a href=\"http:\/\/cs-www.cs.yale.edu\/homes\/dna\/\">Daniel Abadi<\/a> was  awarded the\u00a02008 SIGMOD Jim Gray Doctoral Dissertation Award for his excellent thesis on Column-Oriented Database Systems. Such great work supported by influential people such as Stonebraker is likely to get people talking.<\/p>\n<p>But are column-oriented databases the <strong>next big thing<\/strong>? No.<\/p>\n<ul>\n<li>Column stores have been around for a long time in the form of bitmap and projection indexes. Conceptually, there is little difference. (See my <a href=\"http:\/\/arxiv.org\/abs\/0901.3751\">own work on bitmap indexes<\/a>.)<\/li>\n<li>While it is trivial to change or delete a row in a row-oriented database, it is harder in column-oriented databases. Hence, applications are limited to data warehousing.<\/li>\n<li>Column-oriented databases are faster for some applications. Sometimes faster by two orders of magnitude, especially on low selectivity queries. Yet, part of these gains are due to the recent evolution in our hardware. Hardware configurations where reading data sequentially is very cheap favor sequential organization of the data such as column stores. What might happen in the world of storage and microprocessors in the next ten years?<\/li>\n<\/ul>\n<p>I believe\u00a0Nicolas Bruno said it best in\u00a0<a href=\"http:\/\/research.microsoft.com\/apps\/pubs\/default.aspx?id=74156\">Teaching an Old Elephant New Tricks<\/a>:<\/p>\n<blockquote><p>(&#8230;) some C-store proponents argue that C-stores are fundamentally dif<span>ferent from traditional engines, and therefore their bene\ufb01ts cannot<span> be incorporated into a relational engine short of a complete<span> <\/span>rewrite<span> (&#8230;) we (&#8230;) show that many of the<span> bene\ufb01ts of C-stores can indeed be simulated in traditional engines<span> with no changes whatsoever.\u00a0\u00a0Finally, we predict that traditional relational engines will<span> eventually leverage most of the bene\ufb01ts of C-stores natively, as is<span> currently happening in other domains such as XML data.<span> <\/span><\/span><\/span><\/span><\/span><\/span><\/span><\/span><\/p><\/blockquote>\n<p>That is not to say that you should avoid Vertica&#8217;s products or do research on column-oriented databases. However, do not bet your career on them. The hype will not last.<\/p>\n<p>(For a contrarian point of view, read Adabi and Madden&#8217;s blog post on why column stores are fundamentally superior.)<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Most database users know row-oriented databases such as Oracle or MySQL. In such engines, the data is organized by rows.\u00a0Database researcher and guru Michael Stonebraker has been advocating column-oriented databases. The idea is quite simple: by organizing the data into columns, we can compress it more efficiently (using simple ideas like run-length encoding). He even &hellip; <a href=\"https:\/\/lemire.me\/blog\/2009\/07\/03\/column-stores-and-row-stores-should-you-care\/\" class=\"more-link\">Continue reading <span class=\"screen-reader-text\">Column stores and row stores: should you care?<\/span><\/a><\/p>\n","protected":false},"author":56,"featured_media":0,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[8],"tags":[],"class_list":["post-2044","post","type-post","status-publish","format-standard","hentry","category-data-warehousing-and-olap"],"_links":{"self":[{"href":"https:\/\/lemire.me\/blog\/wp-json\/wp\/v2\/posts\/2044","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/lemire.me\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/lemire.me\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/lemire.me\/blog\/wp-json\/wp\/v2\/users\/56"}],"replies":[{"embeddable":true,"href":"https:\/\/lemire.me\/blog\/wp-json\/wp\/v2\/comments?post=2044"}],"version-history":[{"count":5,"href":"https:\/\/lemire.me\/blog\/wp-json\/wp\/v2\/posts\/2044\/revisions"}],"predecessor-version":[{"id":10984,"href":"https:\/\/lemire.me\/blog\/wp-json\/wp\/v2\/posts\/2044\/revisions\/10984"}],"wp:attachment":[{"href":"https:\/\/lemire.me\/blog\/wp-json\/wp\/v2\/media?parent=2044"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/lemire.me\/blog\/wp-json\/wp\/v2\/categories?post=2044"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/lemire.me\/blog\/wp-json\/wp\/v2\/tags?post=2044"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}