{"id":13632,"date":"2023-09-08T11:11:39","date_gmt":"2023-09-08T10:11:39","guid":{"rendered":"https:\/\/telecomkh.info\/?p=13632"},"modified":"2023-09-08T11:11:39","modified_gmt":"2023-09-08T10:11:39","slug":"architecting-the-future-of-supercomputing","status":"publish","type":"post","link":"https:\/\/telecomkh.info\/?p=13632","title":{"rendered":"Architecting the future of supercomputing"},"content":{"rendered":"<p><strong>As chief architect and principal investigator for the Aurora supercomputer at Argonne National Laboratory in Illinois, Olivier Franza plays a leading role in bringing one of the most ambitious scientific instruments \u2013 not to mention the world\u2019s largest GPU cluster \u2013 into existence<\/strong><\/p>\n<p>&nbsp;<\/p>\n<p><span style=\"color: #999999;\"><em>Intel chief architect Olivier Franza on how the Aurora supercomputer came into existence<\/em><\/span><\/p>\n<p>Aurora is among the most anticipated and highly visible projects Intel has been a part of in recent memory \u2013 a bold bet on Intel\u2019s entire system portfolio. The machine is expected to be the first supercomputer with a peak performance reaching 2 exaflops, or 2&#215;1018, floating point operations per second.<br \/>\nThat puts a bit of pressure on Franza, a 22-year Intel veteran who joined the Aurora project as system hardware architect in 2016, oversaw the pivot to a GPU-based machine and became chief architect in 2021.<br \/>\n\u201cThe chief architect is responsible for defining the overall system architecture of the supercomputer, according to the customer\u2019s high-level requirements,\u201d Franza explains. \u201cThere are fundamental ones like general performance metrics and power envelope, but also inherent features like RAS \u2013 reliability, availability, serviceability \u2013 that are essential to building a scalable system.\u201d<br \/>\nHis responsibilities also encompass the details of the system topology from a node to a rack to the complete system, including its networking fabric and storage components.<\/p>\n<p><strong>A Roadmap Pivot Opens Opportunity to Shape Future Products<\/strong><br \/>\nWhen initial planning began for Aurora, a U.S. Department of Energy-sponsored system, the design consisted of a collection of Intel technologies. However, changes to Intel\u2019s product roadmap, notably the end of the Xeon Phi and Omnipath product families, required a restart. As Intel made plans to build data center GPUs, Franza became enmeshed in discussions on the design of the Intel\u00ae Data Center GPU Max Series (code-named Ponte Vecchio).<br \/>\nIn this way, Aurora isn\u2019t just a one-off system. Rather, it helped inform the Intel-wide strategy and product portfolio to address scale and performance at the highest level.<br \/>\n\u201cWe infused all the Aurora system-level requirements down to the components\u2019 level,\u201d Franza says.<br \/>\nThe architecture and concept for the Intel\u00ae Xeon\u00ae CPU Max Series with high bandwidth memory, for instance, was spawned by some features from the Intel Xeon Phi platform, the first product to integrate an innovative memory architecture for high bandwidth and high capacity on package.<br \/>\nAdditionally, the need for high performance drove further advances across all subsystems, from the compute blade\u2019s thermo-mechanical solution to its dense physical integration, to storage.<br \/>\n\u201cIntel ended up architecting a completely new storage concept, DAOS (distributed asynchronous object storage),\u201d Franza says. It\u2019s an open source software ecosystem to enable high-speed storage on traditional hardware. \u201cAurora will be among the first systems to use it, and by far the largest.\u201d<\/p>\n<p><strong>From Designing Components to Bolting Together Thousands of Systems<\/strong><br \/>\nThe Aurora project drove system-level thinking and broad collaboration across various business units inside Intel, as well as with Argonne scientists and engineers at Hewlett Packard Enterprise, the project\u2019s other main partner.<br \/>\n\u201cGetting the whole team to align and deliver a machine like Aurora is, for many of us, a once-in-a-lifetime experience,\u201d Franza says.<br \/>\nAlthough engineers installed the final blade in June, the project continues to keep Franza up at night as the system passes through the stages of testing, stabilization and validation at scale.<br \/>\nHe provides guidance to a large team working on system bring-up, validation, stabilization, optimization and enablement of full-system performance workloads. Most notable is the High Performance Linpack (HPL) benchmark that determines the top systems in the world, as certified by the bi-annual Top500 list.<br \/>\nEach morning, Franza joins the daily standup meeting to scrutinize nightly runs on every single node and makes a game plan for the next day\u2019s work and beyond. Each afternoon, a daily closeout meeting summarizes progress and hurdles. The work never stops; the machine always runs.<br \/>\n\u201cWe have a step-by-step approach to methodically validate and stabilize at scale,\u201d he explains. \u201cYou start with the blade, then move to the rack, then multiple racks, and you scale from there.\u201d<br \/>\nAurora is made up of 10,624 compute blades, boasting 63,744 Intel Max Series GPUs \u2013 more GPUs than any other system in the world \u2013 and 21,248 Intel Xeon Max CPUs across 166 racks.<br \/>\n\u201cIt\u2019s the size of four tennis courts, which sounds like a lot, right?\u201d he says. \u201cBut it&#8217;s only when you actually go see it that you just realize the sheer magnitude of the project.\u201d<br \/>\nFranza must ensure the vast system is stable, functional and performing. It\u2019s a daunting task, but the end is within reach.<br \/>\n\u201cWalking through the aisles, with all the lights on, and feeling that the machine is running is impressive and obviously extremely rewarding,\u201d he says. \u201cIt&#8217;s a very tangible achievement that speaks for itself.\u201d<\/p>\n<p><strong>A \u2018Once-in-a-Lifetime\u2019 Effort, a Science-Shaping Supercomputer<\/strong><br \/>\nWhat keeps him going, through engineering hurdles and unexpected roadblocks, is the opportunity to build \u201can extraordinary machine\u201d that will power impactful research. He cites Aurora\u2019s enormous potential for cancer research as an area where the project will benefit us all.<br \/>\n\u201cI think that&#8217;s something that is going to make us very proud,\u201d he says.<br \/>\nNot only will Aurora work on solving some of the most complex scientific and engineering problems in the world, it will also be an ideal platform for running generative AI and applying it to research. \u201cIt will enable one of the biggest large language models planned to date, the 1 trillion parameter Aurora GenAI project, enhancing, enabling and easing the lives of scientists,\u201d Franza says.<br \/>\nBut it\u2019s the teamwork and camaraderie he enjoys more than anything else.<br \/>\n\u201cIt&#8217;s an extended effort, and it requires a lot of perseverance,\u201d he says. \u201cThe core team has maintained a marathon mentality where it&#8217;s not over until it&#8217;s over. We needed the kind of people that can effectively focus for a long time on something immensely challenging. And in the end, the accomplishment is something that very few can say they have achieved.\u201d<\/p>\n<p><span style=\"color: #999999;\"><em>Above, Intel chief architect Olivier Franza \/ Image credited to Intel\u00a0<\/em><\/span><\/p>\n","protected":false},"excerpt":{"rendered":"<p>As chief architect and principal investigator for the Aurora supercomputer at Argonne National Laboratory in Illinois, Olivier Franza plays a leading role in bringing one of the most ambitious scientific instruments \u2013 not to mention the world\u2019s largest GPU cluster \u2013 into existence &nbsp; Intel chief architect Olivier Franza on how the Aurora supercomputer came &hellip; <\/p>\n<p class=\"link-more\"><a href=\"https:\/\/telecomkh.info\/?p=13632\" class=\"more-link\">Continue reading<span class=\"screen-reader-text\"> \u00abArchitecting the future of supercomputing\u00bb<\/span><\/a><\/p>\n","protected":false},"author":1,"featured_media":13634,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":[],"categories":[41,69],"tags":[],"_links":{"self":[{"href":"https:\/\/telecomkh.info\/index.php?rest_route=\/wp\/v2\/posts\/13632"}],"collection":[{"href":"https:\/\/telecomkh.info\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/telecomkh.info\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/telecomkh.info\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/telecomkh.info\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=13632"}],"version-history":[{"count":1,"href":"https:\/\/telecomkh.info\/index.php?rest_route=\/wp\/v2\/posts\/13632\/revisions"}],"predecessor-version":[{"id":13635,"href":"https:\/\/telecomkh.info\/index.php?rest_route=\/wp\/v2\/posts\/13632\/revisions\/13635"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/telecomkh.info\/index.php?rest_route=\/wp\/v2\/media\/13634"}],"wp:attachment":[{"href":"https:\/\/telecomkh.info\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=13632"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/telecomkh.info\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=13632"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/telecomkh.info\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=13632"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}