| |

Commonalities Between High Performance Computing and Gaming

March 19, 2013, GPU Technology Conference, San Jose, CA—Sarah Tariq from Nvidia described the intersection of high-performance computing and gaming. At first glance, there appears to be no common areas between the two areas, but the reality is a lot of functional overlap.

Parallel computers emerged as an alternative to higher clock speeds in the ’80s. Now, the latest top-end supercomputer, called Titan, has been installed in Oak Ridge, Tennessee. This machine has over 18,000 CPU nodes, and as many Tesla processors. The problem for this, and similar machines, is that the power density for the processor arrays is climbing at an exponential rate with every generation of supercomputer. As a result, all computer systems must be concerned with power consumption, so the appropriate metrics are now increasing performance per watt rather than trying to increase clock frequency.

For a gaming computer, ongoing requirements are for better graphics and higher processor throughput for better gaming performance. In addition to being a part of an $80 billion market, gaming machines represent a large installed base of GPUs and the highest performance possible within a single chassis.

Drilling down below this top level, both types of systems must manage intensive graphics functions. They are handling millions of triangles, use tessellation engines, rotate, translate, rasterize, and shade images. All of these functions must be completed at video frame rates. The gaming machines are doing this for higher realism and more immersive game play, while the high-performance computers are helping researchers visualize issues like Earth-scale weather.

Within these machines, the CPUs have from 4 to 16 cores and work in a single threaded mode. Most of the time and energy within a CPU is used for data flow management, such as scheduling, out of order, and branch prediction versus a small amount for actual program execution. In comparison, the GPU can have thousands of cores that are optimized for execution at the cost of increased latency for out of order or missed branch conditions. Nevertheless, the GPUs have much greater bandwidth per watt than the CPUs. Over time, GPUs have increased performance and added more general purpose programmability. Now, GPUs can form as massively parallel general purpose compute engines.

When we look at workloads, there a lot of areas of commonality. For example, in image processing, one important process is particle simulation. Particles have constraints, interactions, or are independent from other particles. In games, hair uses a lot of particles with one-dimensional distance constraints so the hair doesn’t get stretched. Cloth requires a 2-D mesh because their differences between the vertical, diagonal, horizontal constraints of the cloth. In high-performance computing, molecular dynamics look at complex molecular systems to help researchers understand chemical reactions or to develop new types of pharmaceutical systems. The constraints are bonds between particles, other forces, distance etc. from surrounding atoms.

Another common function is convolution over any number of dimensions. In games, convolution is one of the post-processing actions to modify depth of field, and add in film bloom to highlight central characters and action. Another application for games is subsurface scattering for materials like skin. Because skin is translucent, sunlight penetrates the surface and is diffusely reflected back to the surface, resulting in a blurred glow. In the same pain, high-performance computing uses convolution for processes like reverse time migration in seismic imaging. The AxRTm program performs a time domain process to identify the convolution intervals.

Partial differential equations are also used in both areas. Partial differential equations are used in flow simulations to show movement of a fluid over a range of velocity and pressure gradients. In high-performance computing, these functions in simulating airflow, ocean currents, etc. for physical phenomena over a very large areas. An alternative to the differential equations is a fast Fourier transform. In gaming, FFT’s are used for post production functions such as creating lens flare and simulating oceans. These functions were first developed for cinema and are now applied to games. In high-performance computing FFT’s are used to simulate turbulence, develop cryptography, etc.

Laplace transforms are applied to situations where spherical harmonics are involved. In gaming, this might be applied for scenes with indirect lighting and multiple reflective surfaces. In high-performance computing, these transforms are used for magnetized solar ejection simulations, thermal dynamics, and numerical weather predictions with localized turbulence.

Some functions are memory bandwidth bound For scenes in games, this can be seen as ambient occlusion and in high-performance computing for sparse matrix vector multiplication. Others are mass bound. In gaming, performance is generally pixel shader limited while in high-performance computing, the representation of molecular dynamics runs into the same problem.

In the future, both of these areas are looking for improved quality and performance. Gaming will benefit from better algorithms developed for the high-performance computing areas and other hardware and software improvements will increase performance for volumetric lighting. High performance computing will be able to simulate weather over a larger space, using smaller grade spacings. Research into climate change requires modeling the entire earth, so performance improvements are necessary. At the same time, this performance is power limited, so new hardware configurations will enable increase performance per watt and greater throughput in both of these areas.

Similar Posts