Ampere Computing Buys An AI Inference Performance Leap
Summary
Ampere Computing is the current best bet one can make for an independent vendor of Arm-based server CPUs, albeit they are only aimed at hyperscale and cloud builder workloads. On X86 machines, the speedup of the OnSpecta Deep Learning Software (DLS) inference engine is around a factor of 2X, and on the Ampere Computing Altra and AltraMax CPUs the speedup is around 4X, according to Wittich, and specifically he is comparing a 28-core “Cascade Lake” Xeon SP processor with the DLBoost mixed precision enhancements for the AVX512 vector engines to the an 80-core Altra processor with much more modest math units, but a lot more of them. A lot of this kind of tuning is done by hand at hyperscalers and cloud builders and AI researchers, and the whole idea of OnSpecta is that it this is done seamlessly and automagically for those who are not necessarily wanting to get down into machine code to drive performance. This is for everyone else, and if the performance is good enough with DLS, then they won’t have to worry about doing it by hand on a new Arm architecture – which is an important point and which is why Ampere Computing is just buying the company. What he did admit is that the OnSpecta stack makes the job of porting inference codes from X86 CPUs and Nvidia GPUs a whole lot easier, which is why it has been available on the A1 instances on the Oracle Cloud for some time now.