Joinwin Electronics HK Limited518031About Join WinJoinwin Electronics HK Limited FLAT/RM 29, 22/F, YAN’S TOWER, 25-27 WONG CHUK HANG ROAD,ABERDEEN HONGKONG Shenzhen ZhongYiYingTong Technology Limited Room 1610, Block A, Overseas Decorative Building, Zhenhua Road,Huaqiang North Street, Shenzhen, 518031Join-Win will be your one-stop purchasing assistant. Join together, achieve win-win!Joinwin Electronics HK Limited518031About Join WinJoinwin Electronics HK Limited FLAT/RM 29, 22/F, YAN’S TOWER, 25-27 WONG CHUK HANG ROAD,ABERDEEN HONGKONG Shenzhen ZhongYiYingTong Technology Limited Room 1610, Block A, Overseas Decorative Building, Zhenhua Road,Huaqiang North Street, Shenzhen, 518031Join-Win will be your one-stop purchasing assistant. Join together, achieve win-win!Joinwin Electronics HK LimitedJoinwin Electronics HK LimitedJoinwin Electronics HK LimitedJoinwin Electronics HK LimitedJoinwin Electronics HK LimitedJoinwin Electronics HK LimitedJoinwin Electronics HK LimitedJoinwin Electronics HK LimitedJoinwin Electronics HK LimitedJoinwin Electronics HK LimitedJoinwin Electronics HK LimitedJoinwin Electronics

  1. Home
  2. News
  3. Moving From CPU Acceleration To Central Computing, AMD's Second-Generation Versal SoC Has A Very Different Identity

Moving From CPU Acceleration To Central Computing, AMD's Second-Generation Versal SoC Has A Very Different Identity

0
Have to mention ACAP
When it comes to Versal, it is important to recall the new product category that Xilinx launched in 2018 - ACAP(Adaptive Compute Acceleration Platform), after all, Versal is the industry's first ACAP architecture product.
ACAP is a highly integrated multi-core heterogeneous computing platform that can be flexibly modified from the hardware layer according to the requirements of various applications and workloads. With a four-year R&D cycle, a cumulative R&D investment of more than $1 billion, and more than 1,500 software and hardware engineers involved in the design of the project, the company has high expectations for it.
According to the structural block diagram published at the time, The ACAP platform combines distributed memory, multi-core SOCs, highly integrated programmable I/O, SerDes transceiver technology, cutting-edge RF-ADC/DAC, integrated high bandwidth memory (HBM), and one or more software-programmable and hardware-flexible computing engines (DSP/AI, etc.). All are interconnected via Network on Chip (NoC). Software developers can use software tools such as C/C++, OpenCL and Python to apply ACAP systems, as well as FPGA tools to program from the RTL level.
The name Versal is derived from two words, one for diversity and the other for versatility. The first generation portfolio includes the Versal Basic series (Versal Prime), the Versal Flagship series (Versal Premium) and the HBM series. In addition, it also includes the AI Core series, the AI Edge series, and the AI RF series.
Some of the benchmark hardware architectures include the use of TSMC's 7nm FinFET process, the integration of a dual-ARM Cortex-A72 application processor and a dual-ARM Cortex-R5 real-time processor, among others. In addition, Xilinx also introduced an innovative engine, the platform management controller, which can control the entire device, which can meet the top-down design and realize the programmable software.
Launched in 2020, Versal Premium is the industry's most bandwidth and computant-dense adaptive platform at the time. Its system logic unit from a minimum of 1.6 million to a maximum of 7.4 million, the number of adaptive engine LUTs from a minimum of 720,000 to a maximum of 3.4 million, can provide 3 times higher throughput and 2 times higher computing density than the mainstream FPGA, and built-in Ethernet, Interlaken and encryption engines, Designed for running the highest bandwidth networks in thermal and space-constrained environments, and for cloud providers who require scalable, flexible application acceleration.
Unveil the 2nd generation Versal
Manuel Uhm divides the processing phase of AI-driven embedded systems into three segments: pre-processing - AI inference - post-processing, noting, "AI brings more demanding workloads to highly constrained systems, so true system-wide performance can only be achieved if all three phases are accelerated in high-performance embedded systems."
However, the current practical system construction idea is to use FPGA and SoC for optimization in the "pre-processing" stage, use vector processor SoC in the "inference" stage, and use high-performance embedded CPU in the "post-processing" stage. That said, "no processor can be optimized for all three phases," and this multi-chip solution comes with significant overhead - from higher power requirements, board area, and memory requirements, to more security vulnerabilities, component obsoletions, design time, and effort.
This explains why AMD chose to introduce the second generation of the Versal AI Edge and Versal Prime series at this time - that is, to use the next generation of AI engines, new high-performance integrated cpus, and AMD programmable logic to bring "single chip intelligence" to embedded systems, or, The hope is to provide "end-to-end acceleration in a single device." The following diagram clearly illustrates this idea.
It is not difficult to see that AMD has integrated the AIE-ML v2 AI engine in the second generation Versal Adaptive SoC, which can achieve up to three times the TOPS performance per watt compared to the previous generation; Programmable logic enables flexible real-time preprocessing, especially in the face of sensor fusion, data conditioning, hard image/video processing; In terms of CPU performance, scalar computing power is increased by 10 times by integrating 8X Arm Cortex-A78AE application processor and 10X Arm Cortex-R52 real-time processor; At the same time, considering that edge applications have very strict requirements for information security and functional safety, new products have increased support for functional safety and information security, and increased support for standards such as ASIL D/SIL 3.
Figure 2 illustrates the higher level of system performance that the new products bring to embedded applications:
In L2+/L3 ADAS applications, due to the addition of hard image processing capabilities, the second-generation AI Edge series has increased its image processing capacity by four times with similar power resources.
In the smart city scenario, the second-generation AI Edge series, while reducing the edge AI device's board area by 30%, supports twice the video stream, meaning that each video stream accounts for 65% of the board area.
In video streaming, the second generation Versal Prime series provides twice the video processing power for multi-port encoding and streaming compared to the efficiency of Zyng MPSoC, reducing the board area per stream by 35%.
"When it comes to preprocessing, adaptation equals flexibility." Manuel Uhm pointed out that for customers, the most important thing about programmable logic is that it can program the hardware in real time, adapt to different sensors, IO interfaces, data types, and enable hardware customization. Processors, by contrast, are limited by their instruction sets, making it difficult to achieve such flexibility.
Acceleration from CPU to system central computing
"We want the second generation of Versal Adaptive SoC to be central computing for AI-driven and classic embedded systems, rather than more CPU acceleration, which is the biggest difference from the first generation." Manuel Uhm said.
Take the "pre-processing" segment for example. If the processor-based approach is used, in the face of different sensors and different types of data, fixed I/O with interfaces and hard ISPs are limited in the number of processes, lack of flexibility, and sometimes have to be stored and cached through external memory, resulting in high latency and low efficiency. In contrast, when a programmable logic approach is adopted, these disadvantages are transformed into advantages.
The same goes for "AI reasoning." Different from the first generation of AI engine control mainly through programmable logic, the control processor of the new generation of products is included in the AI engine array and has been hardened, and the work of AI engine control does not need to be processed by programmable logic in the future, and the surplus programmable logic resources will be used for the processing of sensors and other data.
To better address the throughput and accuracy challenges faced during AI inference, the Dense TOPS situation in the second generation Versal AI Edge family of devices has also been enhanced: When the data type is MX6/INT8, the highest end can achieve 370 TFLOPS and 184 TOPS, respectively, with the former providing up to 60% improvement in TOPS per watt with similar or greater accuracy. If the sparsity index is used, the performance can be doubled.
At the same time, in order to achieve better and faster model deployment, AMD provides Vitis™ AI development environment to help developers use familiar open source tools, such as PyTorch, TensorFlow, etc., to optimize and reason in Vitis.
Finally, take a look at the second-generation Versal Adaptive SoC's performance in the "post-processing" phase. As mentioned earlier, the new product can achieve up to 8x Arm Cortex-A78AE cores with a maximum frequency of 2.2GHz per core, and has up to 200.3K of DMIPS computing power, laying the foundation for complex post-processing with up to 10x scalar computing power. For real-time processing units for control functions, the RPU can have up to 10 times the Arm Cortex-R52 core, with a maximum frequency of 1.05GHz per core, and up to 28.5K of DMIPS computing power. In addition, enhanced functional safety significantly reduces the need for external security microcontrollers.
The Subaru EyeSight Vision System is a prime example of using the second generation Versal™ AI Edge family of products. Through the cooperation, the next-generation EyeSight vision system has been further improved in pre-collision braking, lane departure warning, adaptive cruise control and lane keeping assistance. In addition, using programmable logic, Subaru can also modify the stereo camera processing algorithm in real time, further strengthening the vehicle safety performance.
Camera-based 3D perceptual visual processes are another example. According to the presentation, during the entire mode process, the pre-processed data will be transmitted to the AI engine with 3D performance models (such as BEVFormer), and then use the processor to plan the behavior pattern or other real-time sensing, so that the camera sensor alone can achieve the visual effect of overlooking, without the use of Lidar.
According to the plan, early trial programs for the second generation Versal™ AI Edge series and the second generation Versal Prime series have been launched, early access documents have been published, and key customers, including Subaru, are currently being approached. Chip prototypes will be released in the first half of 2025, evaluation kits and System modules (SOM) will be introduced in mid-2025, and production chips will be available in late 2025.
TOP
RFQ List ( 0 items)