概述
日期
2024年03月29日
09:00 - 10:00
所在
运动行、Bilibili

Scheduling Deep Learning Workloads at Scale in GPU Data Centers

首页- 优德官网集团(中国)有限公司

  对人工智能 日益增添的问题解决能力和泛化能力需求,,,,,,,现代深度学习模子变得越来越重大且重大,,,,,,,需要消耗大宗盘算资源和时间。。。。。。 。使用大规模GPU数据中心举行模子训练和推理优化已成为常见做法。。。。。。 。然而,,,,,,,由于深度学习使命的高盘算需求和底层硬件的异构性,,,,,,,GPU数据中心治理和调理使命面临多重挑战。。。。。。 。

第十三期优德官网-TNSE团结优异讲座系列运动,,,,,,,我们有幸约请到南洋理工大学的文勇刚教授先容GPU数据中心大规模深度学习负载调理,,,,,,,并分享他在这个领域内的相关研究效果与有趣发明。。。。。。 。

    

优德官网-TNSE Joint Distinguished Seminar Series is co-sponsored by IEEE Transactions on Network Science and Engineering (TNSE) and Shenzhen Institute of Artificial Intelligence and Robotics for Society (优德官网), with joint support from The Chinese University of Hong Kong, Shenzhen, Network Communication and Economics Laboratory (NCEL), and IEEE. This series aims to bring together top international experts and scholars in the field of network science and engineering to share cutting-edge scientific and technological achievements.

Join the seminar through Bilibili (http://live.bilibili.com/22587709).

  • 首页- 优德官网集团(中国)有限公司
    Jianwei Huang
    Vice President, 优德官网; Presidential Chair Professor, CUHK-Shenzhen; Editor-in-Chief, IEEE TNSE; IEEE Fellow; AAIA Fellow
    Executive Chair
  • 首页- 优德官网集团(中国)有限公司
    Yonggang Wen
    Professor and President's Chair in Computer science and Engineeringat Nanyang Technological University; Editor in Chief of lEEE Transactions on Multimedia; lEEE Fellow
    Professor and President's Chair in Computer science and Engineeringat Nanyang Technological UniversityEditorin Chief of lEEE Transactions on Multimedia lEEE Fellow

    文勇刚,,,,,,,南洋理工大学盘算机科学与工程学院校长讲席教授,,,,,,,于2008年在美国剑桥的麻省理工学院获得电子工程和盘算机科学博士学位(辅修西方文学),,,,,,,现在担当新加坡南洋理工大学副教务长(研究生教育)和研究生院院长。。。。。。 。此前,,,,,,,他曾担当新加坡南洋理工大学校长办公室协理副校长(能力建设)(2023年)、工程学院副院长(研究)(2018-2023年)、南洋科技创业中心署理主任(2017-2019年)和盘算机科学与工程学院助理主席(立异)(2016-2018年)。。。。。。 。文教授在顶级期刊和著名聚会上揭晓了300多篇论文。。。。。。 。他的系统研究获得了全球认可,,,,,,,他在多屏云社交电视方面的事情曾受到全球媒体的关注(来自29个国家的1600多篇新闻文章),,,,,,,并获得2013年东盟ICT奖(金奖)。。。。。。 。他在数据中心认知数字孪生方面的事情,,,,,,,获得了2015年数据中心动力学奖- APAC(数据中心行业的“奥斯卡”奖)、2016年东盟ICT奖(金奖)、2020年IEEE TCCPS工业手艺卓越奖、2021年W.Media APAC云与数据中心手艺首脑奖,,,,,,,以及2022年新加坡盘算机学会数字成绩手艺首脑奖。。。。。。 。他是2019年南洋研究奖获得者和2016年南洋立异创业奖唯一获得者,,,,,,,这两个奖项都是南洋理工大学的最高声誉。。。。。。 。他曾获得多个最佳论文奖,,,,,,,包括2019年IEEE TCSVT和2015年IEEE Multimedia的最佳论文奖,,,,,,,以及多个国际聚会的最佳论文奖,,,,,,,包括2023年ASPLOS、2016年IEEE Globecom、2016年IEEE Infocom MuSIC Workshop、2015年EAI Chinacom、2014年IEEE WCSP、2013年IEEE Globecom和2012年IEEE EUC。。。。。。 。他是IEEE Transactions on Multimedia (TMM)的主编,,,,,,,担当或曾担当多个IEEE和ACM Transactions的编辑委员会成员,,,,,,,并中选为IEEE ComSoc多媒体通讯手艺委员会主席(2014-2016)。。。。。。 。文教授的主要研究偏向为云盘算、绿色数据中心、大数据剖析、多媒体网络和移动盘算。。。。。。 。他是IEEE会士、新加坡工程院院士,,,,,,,也是ACM的优异成员。。。。。。 。

    To meet the ever-growing demand of problem-solving capability and generalizability via artificial intelligence, modern deep learning models are becoming larger and more sophisticated, while at the cost of huge amounts of computing resources (e.g., GPU) and prolonged training time. it has become a common practice to leverage large-scale GPU data centers (i.e., AI data centers) to optimize and accelerate model training and inference. However, the management and scheduling of these deep learning workloads in the GPU data centers present numerous challenges, due to their high computational requirements, distinct and diverse runtime characteristics, and heterogeneous nature of the underlying hardware.

    In this talk, we will investigate deep learning workload scheduling accelerating, training execution over GPU datacenters, with a multifold objective of improving resource utilization, enhancing users’ experience, and easing operators’ management. Specifically, we will introduce novel and practical methodologies and system designs to achieve those goals. These solutions are highly integrated to tackle different challenges, paving the way for optimal utilization of GPU resources and accelerated progress in deep learning applications.