AES – Gaming Audio Focuses on Spatial Sound
October 2012 – At the 133rd AES Convention in San Francisco, they held a multi-day track program on gaming audio. One of the recurring themes and topics was the incorporation of spatial sound. Just like 3D has brought a deeper level of immersion into the graphics of the video games, surround sound and spatial audio are in the process of bringing the same shift to gaming sound.
A workshop was held that reviewed the structure of object based sound development. This focused on several efforts being made to define a common interface and terminology set for the creation of sound pieces that are assignable to objects that appear in a game. This allows for dynamic game play with different outcome scenarios and they will have correspondingly unique sound tracks and mixes based on the existence and actions of the objects.
One of the changes with respect to spatial sound, is the definition of the magnitude and phase scaling for the position of the object in the visual domain tracking to the sound domain. A challenge for the sound domain is the method of sound presentation. Planar surround sound and dimensional surround sound, do not necessarily have the physiological interpretation when played back through a true multi-speaker surround sound system, a front facing 2 speaker arrangement and with either single speaker or multi-speaker headphones. As a result, the uniformity of experience that is possible to create for a known gaming console on a single big forward viewed screen, is hard to duplicate with sound.
The creation of the common definitions and standardized spatial phase mixing techniques will help the audio engineer at least get an idea of the sound experience the game design is seeking, and there is a chance to create it properly. As the games continue to be ported across more platforms and higher performance games are moving to the mobile platform, this uniformity of experience will become a requirement.
At this time there are two major types of spatial audio – single plane and dimensional. The single plane sound includes front, back and sides all located at the approximate level of the ears of the listener. The dimensional sound includes front, back and sides and adds the z-axis for up and down. This allows sounds to be interpreted as behind you on the ground or over the top of your head. At this time, there are many different systems for implementing these dimensional sound processing solutions. The drawback, is there is a “sound imaging training curve” that needs to be initiated per session to “calibrate” the sound location. This training curve can be incorporated into the standard game play, however it has some impact on the story lines, and also resulting dynamic range of the playout devices. Some suppliers, such as Dolby Labs are working on dimensional spatial sound that does not require a training curve, or dynamic range impact.
The issue of playout of these systems on headphones was the subject of a special session on its own. In this session, the headphones were defined as standard multi-use listening headphones that supported an audio back channel for a microphone as is used for most multi-player games. These have a particular challenge on where to place the voice communication for the chat with other players in the context of the spatial sound mix. This problem is analogous to the placement of closed captioning in the 3D depth field on the visual side. The session reviewed several mix options and the associated tradeoffs and benefits from the solutions. The goal was to provide full sound response over a large dynamic range with headphones that primarily cover only the 40hz-16Khz range rather than the full 20hz-20KHz band. In this session, ear buds with passive noise cancellation were also discussed as being compatible with “headphone” solutions.


