{"id":20257,"date":"2023-03-01T23:01:30","date_gmt":"2023-03-01T23:01:30","guid":{"rendered":"https:\/\/lemire.me\/blog\/?p=20257"},"modified":"2023-05-03T18:52:11","modified_gmt":"2023-05-03T18:52:11","slug":"arm-vs-intel-on-amazons-cloud","status":"publish","type":"post","link":"https:\/\/lemire.me\/blog\/2023\/03\/01\/arm-vs-intel-on-amazons-cloud\/","title":{"rendered":"ARM vs Intel on Amazon&#8217;s cloud: A URL Parsing Benchmark"},"content":{"rendered":"<p>Twitter user opdroid1234 remarked that <a href=\"https:\/\/twitter.com\/opdroid1234\/status\/1631041253843382274\">they are getting more performance out of the ARM nodes than out of the Intel nodes on Amazon&#8217;s cloud (AWS)<\/a>.<\/p>\n<p><a href=\"http:\/\/lemire.me\/blog\/wp-content\/uploads\/2023\/03\/Capture-decran-le-2023-03-01-a-17.48.42.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-medium wp-image-20258\" src=\"http:\/\/lemire.me\/blog\/wp-content\/uploads\/2023\/03\/Capture-decran-le-2023-03-01-a-17.48.42.png\" alt=\"\" width=\"600\" height=\"206\" srcset=\"https:\/\/lemire.me\/blog\/wp-content\/uploads\/2023\/03\/Capture-decran-le-2023-03-01-a-17.48.42.png 1166w, https:\/\/lemire.me\/blog\/wp-content\/uploads\/2023\/03\/Capture-decran-le-2023-03-01-a-17.48.42-300x103.png 300w, https:\/\/lemire.me\/blog\/wp-content\/uploads\/2023\/03\/Capture-decran-le-2023-03-01-a-17.48.42-1024x351.png 1024w, https:\/\/lemire.me\/blog\/wp-content\/uploads\/2023\/03\/Capture-decran-le-2023-03-01-a-17.48.42-768x263.png 768w\" sizes=\"auto, (max-width: 600px) 100vw, 600px\" \/><\/a><\/p>\n<p>I found previously that the <a href=\"https:\/\/lemire.me\/blog\/2022\/06\/07\/memory-level-parallelism-intel-ice-lake-versus-amazon-graviton-3\/\">Graviton 3 processors had less bandwidth than comparable Intel systems<\/a>.\u00a0However, I have not done much in terms of raw compute power.<\/p>\n<p>The Intel processors have the crazily good AVX-512 instructions: ARM processors have nothing close except for dedicated accelerators. But what about more boring computing?<\/p>\n<p>We wrote a fast <a href=\"https:\/\/github.com\/ada-url\/ada\">URL parser<\/a> in C++. It does not do anything beyond portable C++: no assembly language, no explicit SIMD instructions, etc.<\/p>\n<p>Can the ARM processors parse URLs faster?<\/p>\n<p>I am going to compare the following node types:<\/p>\n<ul>\n<li><tt>c6i.large<\/tt>: Intel Ice Lake (0.085 US$\/hour)<\/li>\n<li><tt>c7g.large<\/tt>: Amazon Graviton 3 (0.0725 US$\/hour)<\/li>\n<\/ul>\n<p>I am using Ubuntu 22.04 on both nodes. I make sure that cmake, ICU and GNU G++ are installed.<\/p>\n<p>I run the following routine:<\/p>\n<ul>\n<li><tt>git clone https:\/\/github.com\/ada-url\/ada<\/tt><\/li>\n<li><tt>cd ada<\/tt><\/li>\n<li><tt>cmake -B build -D ADA_BENCHMARKS=ON<\/tt><\/li>\n<li><tt>cmake --build build<\/tt><\/li>\n<li><tt>.\/build\/benchmarks\/bench --benchmark_filter=Ada<\/tt><\/li>\n<\/ul>\n<p>The results are that the ARM processor is indeed slightly faster:<\/p>\n<table>\n<tbody>\n<tr>\n<td>Intel Ice Lake<\/td>\n<td>364 ns\/url<\/td>\n<\/tr>\n<tr>\n<td>Graviton 3<\/td>\n<td>320 ns\/url<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>The Graviton 3 processor is about 15% faster. It is not the 20% to 30% that\u00a0opdroid1234 reports, but the Graviton 3 nodes are also slightly cheaper.<\/p>\n<p>Please note that (1) I am presenting just one data point, I encourage you to run your own benchmarks (2) I am sure that opdroid1234 is being entirely truthful (3) I love all processors (Intel, ARM) equally (4) I am not claiming that ARM is better than Intel or AMD.<\/p>\n<p><strong>Note<\/strong>: I do not own stock in ARM, Intel or Amazon. I do not work for any of these companies.<\/p>\n<p><strong>Further reading<\/strong>: <a href=\"https:\/\/aws.amazon.com\/fr\/blogs\/machine-learning\/optimized-pytorch-2-0-inference-with-aws-graviton-processors\/\">Optimized PyTorch 2.0 inference with AWS Graviton processors<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Twitter user opdroid1234 remarked that they are getting more performance out of the ARM nodes than out of the Intel nodes on Amazon&#8217;s cloud (AWS). I found previously that the Graviton 3 processors had less bandwidth than comparable Intel systems.\u00a0However, I have not done much in terms of raw compute power. The Intel processors have &hellip; <a href=\"https:\/\/lemire.me\/blog\/2023\/03\/01\/arm-vs-intel-on-amazons-cloud\/\" class=\"more-link\">Continue reading <span class=\"screen-reader-text\">ARM vs Intel on Amazon&#8217;s cloud: A URL Parsing Benchmark<\/span><\/a><\/p>\n","protected":false},"author":56,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[84],"tags":[],"class_list":["post-20257","post","type-post","status-publish","format-standard","hentry","category-84"],"_links":{"self":[{"href":"https:\/\/lemire.me\/blog\/wp-json\/wp\/v2\/posts\/20257","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/lemire.me\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/lemire.me\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/lemire.me\/blog\/wp-json\/wp\/v2\/users\/56"}],"replies":[{"embeddable":true,"href":"https:\/\/lemire.me\/blog\/wp-json\/wp\/v2\/comments?post=20257"}],"version-history":[{"count":5,"href":"https:\/\/lemire.me\/blog\/wp-json\/wp\/v2\/posts\/20257\/revisions"}],"predecessor-version":[{"id":20516,"href":"https:\/\/lemire.me\/blog\/wp-json\/wp\/v2\/posts\/20257\/revisions\/20516"}],"wp:attachment":[{"href":"https:\/\/lemire.me\/blog\/wp-json\/wp\/v2\/media?parent=20257"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/lemire.me\/blog\/wp-json\/wp\/v2\/categories?post=20257"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/lemire.me\/blog\/wp-json\/wp\/v2\/tags?post=20257"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}