Intel Releases OpenVINO Toolkit 2024.2 with GenAI NPU Acceleration and Llama 3 Support
Intel released OpenVINO Toolkit 2024.2, introducing direct Intel NPU execution pipelines for generative AI workloads alongside support for Meta Llama 3, Microsoft Phi-3, and Google Gemma models. The update adds micro-scaling (MX) data format optimizations and enhances PyTorch model export capabilities.
Verified State Diff
Impact & Verification Analysis
AI developers, software engineers, and enterprise ISVs deploying deep learning models and local LLMs onto Intel Core Ultra CPUs, Intel Arc GPUs, and Intel Xeon processors.
Significantly lowers memory usage and improves token generation latency for local generative AI applications, enabling enterprise software to run LLMs on end-user PCs and edge hardware without relying on cloud APIs.
Full Fact Overview
Intel officially released OpenVINO Toolkit 2024.2, focusing on local generative AI performance across Intel client and server architectures. This release adds runtime support for popular model architectures including Meta Llama 3, Microsoft Phi-3, and Google Gemma. Additionally, developers can leverage micro-scaling (MX) data formats to reduce memory footprints and accelerate execution on Intel Arc GPUs and Intel Core Ultra processors equipped with integrated Neural Processing Units (NPUs).