Artificial Intelligence with 2nd Gen Intel® Xeon® Scalable Processor
The 2nd Gen Intel® Xeon® Scalable processor provides scalable performance for the widest variety of datacenter workloads – including deep learning. The new 2nd Gen Intel® Xeon® Scalable processor platform offers built-in Return on Investment (ROI), potent performance and production-ready support for AI deployments.
In our smart and connected world, machines are increasingly learning to sense, reason, act, and adapt in the real world. Artificial Intelligence (AI) is the next big wave of computing, and Intel uniquely has the experience to fuel the AI computing era. AI will let us accelerate solutions to large-scale problems that would otherwise take months, years, or decades to resolve.
AI will unleash new scientific discoveries, automate undesirable tasks and extend our human senses and capabilities. Today, machine learning (ML) and deep learning (DL) are two underlying approaches to AI, as are reasoning-based systems.
Deep learning is the most rapidly emerging branch of machine learning, in many cases supplanting classic ML, relying on massive labeled data sets to iteratively “train” many-layered neural networks inspired by the human brain. Trained neural networks are used to “infer” the meaning of new data, with increased speed and accuracy for processes like image search, speech recognition, natural language processing, and other complex tasks.
The 2nd Generation Intel® Xeon® Scalable processors take AI performance to the next level with Intel® Deep Learning Boost (Intel® DL Boost), a new set of embedded processor technologies designed to accelerate AI deep learning use cases such as image recognition, object detection, speech recognition, language translation, and others. It extends Intel® Advanced Vector Extensions 512 (Intel® AVX-512) with a new Vector Neural Network Instruction (VNNI) that significantly increases deep learning inference performance over previous generations. With 2nd Gen Intel® Xeon® Platinum 8280 processors and Intel® Deep Learning Boost (Intel® DL Boost), we project that image recognition with Intel optimized Caffe ResNet-50 can perform up to 14x1 faster than on prior generation Intel® Xeon® Scalable processors (at launch, July 2017).
Here we present some AI workloads showing advanced inference performance with the 2nd Gen Intel® Xeon® Scalable processors family.
All performance measurements are accurate as of April 2, 2019.1 2 3 4
ResNet-50 performance with Intel® Optimization for Caffe*
Designed for high performance computing, advanced artificial intelligence and analytics, and high density infrastructures Intel® Xeon® Platinum 9200 processors deliver breakthrough levels of performance. Using Intel® Deep Learning Boost (Intel® DL Boost) combined with Intel optimized Caffe, new breakthrough levels of performance can be achieved. Here we show the throughput on an image classification topology – ResNet-50 on the 2nd Generation Intel® Xeon Scalable processor.
ResNet-50 Inference Throughput Performance
Inference () generally happens instantaneously at the edge or in the data center, such as when a new photo is uploaded for inspection. For inference throughput TCO is really critical. Inference output can be fed into a number of different usages including â a dashboard for visualization or a decision tree for automatic decision making. Here we show inference throughput on an image database using multiple popular deep learning frameworks such as Caffe, TensorFlow, Pytorch, and MxNet with the ResNet-50 topology.
Inference Throughput Performance
The 2nd Gen Intel® Xeon® Scalable processors are built specifically to run high-performance AI and IoT workloads on the same hardware as other existing workloads. Intel® Deep Learning Boost (Intel® DL Boost) can benefit many inference applications ranging from recommendation systems, Object detection and image recognition and classification. Here we show inference throughput for image classification, object detection and a recommendation system. Multiple frameworks are used including TensorFlow*, Caffe2 and MxNet* and multiple topologies such as ResNet-101, Inception v3, RETINANET*, SSD-VGG16 and Wide and Deep.
Intel® OpenVINO™ Inference Throughput Performance
AI at the edge is opening up new possibilities in every industry, from predicting machine failures to personalizing retail. With the OpenVINO™ toolkit, businesses can take advantage of near real-time insights to help make better decisions, faster. The OpenVINO™ toolkit allows your business to implement computer vision and deep learning solutions quickly and effectively across multiple applications.
Информация о продукте и производительности
Повышение скорости обработки логических выводов в 30 раз по результатам тестирования системы на базе процессора Intel® Xeon® Platinum 9282 с использованием технологии Intel® Deep Learning Boost (Intel® DL Boost): протестировано в корпорации Intel 26 февраля 2019 года. Платформа: 2-разъемный процессор Intel® Xeon® Platinum 9282 Dragon Rock (56 ядер на разъем), технология HT включена, режим turbo включен, общий объем памяти 768 ГБ (24 модуля/ 32 ГБ/ 2933 МГц), BIOS: SE5C620.86B.0D.01.0241.112020180249, CentOS 7 с ядром 3.10.0-957.5.1.el7.x86_64, платформа глубинного обучения: оптимизация Intel® для Caffe* версии: https://github.com/intel/caffe d554cbf1, ICC 2019.2.187, MKL DNN версии: v0.17 (хэш фиксации: 830a10059a018cd2634d94195140cf2d8790a75a), модель: https://github.com/intel/caffe/blob/master/models/intel_optimized_models/int8/resnet50_int8_full_conv.prototxt, BS=64, нет syntheticData уровня данных: 3x224x224, 56 экземпляров/2 разъема, тип данных: INT8 по сравнению с конфигурацией, протестированной Intel на 11 июля 2017 г: 2-разъемный процессор Intel® Xeon® Platinum 8180 с тактовой частотой 2,50 ГГц (28 ядер), технология HT отключена, режим turbo отключен, для управления масштабированием установлен режим «производительность» с помощью драйвера intel_pstate, ОЗУ DDR4-2666 ECC емкостью 384 ГБ. ОС Linux* CentOS версии 7.3.1611 (основная), ядро Linux* 3.10.0-514.10.2.el7.x86_64. Твердотельный накопитель: твердотельный накопитель Intel® для ЦОД серии S3700 (800 ГБ, 2,5 дюйма, SATA, 6 Гбит/с, 25 нм, MLC). Измерение производительности: переменные среды: KMP_AFFINITY='granularity=fine, compact‘, OMP_NUM_THREADS=56, CPU freq set with CPU Power frequency-set -d 2.5G -u 3.8G -g performance. Caffe: (http://github.com/intel/caffe/), версия f96b759f71b2281835f690af267158b82b150b5c. Логические выводы измерялись с помощью команды «caffe time --forward_only», обучение измерялось с помощью команды «caffe time». Для топологий «ConvNet» использовался синтетический набор данных. Для других топологий данные были сохранены в локальном хранилище и закэшированы в памяти до начала обучения. Спецификации топологий https://github.com/intel/caffe/tree/master/models/intel_optimized_models(ResNet-50). Компилятор Intel® C++ версии 17.0.2 20170213, малые библиотеки Intel® Math Kernel Library (Intel® MKL) версии 2018.0.20170425. ПО Caffe выполнялось с «numactl -l».
Результаты тестов производительности основаны на тестировании по состоянию на момент времени, указанный в конфигурации, и могут не отражать всех общедоступных обновлений безопасности. Подробная информация представлена в описании конфигурации. Ни один продукт или компонент не может обеспечить абсолютную защиту.
Производительность зависит от вида использования, конфигурации и других факторов. Дополнительная информация — по ссылке: www.Intel.ru/PerformanceIndex.