| |

MPEG-H Audio: Next Generation System for Interactive and Immersive Sound

October 10, 2014, Audio Engineering Society Convention, Los Angeles—Fraunhofer IIS and their partners techicolor and Qualcomm have released the latest specification and standard for sound. The standard is gearing up for the more immersive 4k and 8k TVs and cinema productions.

The new standard changes sound from a baked-in mix to a 2-D output format like stereo or 5.1 to an object-oriented package geared towards a rendered sound field more like 7.1 + 4 channels for height content. Putting the various sound components into objects allows for greater interactivity and user control of the mix.

The listener can change the volume of any of the objects while keeping the average level constant, can change the peak-level ratio, or change the level of any object such as changing the volume of the dialog in a program. The flexibility also allows for changes in language or even for content like sporting events, change the announcer. Another alternative is to have only ambient sounds from the game come though your TV.

The use of objects requires the end device to render the sound for the environment, but this allows the broadcaster to only have one stream for any device configuration, such as full 11.1, headphones, or a mobile device without any changes. Today, the broadcaster has to generate separate channel-based audio for each platform, momo, stereo, 5.1, and 7.1 +4H. The issues here is that this technology is not compatible with existing infrastructure. The benefit for the user is the ability to get richer 3-D sound and higher-order ambisonics for a more immersive sound experience.

The organizing group of companies is working with other organizations like ATSC on the ATSC 3.0 broadcast standards. The forklift upgrade to the audio-video chain allows for new functions and capabilities in the system to create an environment of envelopment and immersion for viewers. Although the move to mobile devices is going to be slow, Qualcomm is planning to add the capabilities to their next generation chipsets.

Amir Iljazociv, a Fraunhofer spokesperson, noted that the new systems will face high inertia, so adoption is expected to be staged. The first change is to replace the existing AC-3 codecs with the new codec in the MPEG-H. This change will allow for a 50 percent audio bit-rate reduction. Mobile devices will just need updated software with the new codec while fixed platforms can use an external decoder-encoder to use the newer signal type. The newer software can be an upgrade to an existing app, or a new app for the phone or tablet.

Next, systems will add objects that offer interactivity. Interactive menus will let the user adjust a few object parameters for a simple change to the defaults, or fuller access to most objects for fine tuning. This change will allow for different mixes and greater individualization of the whole A-V experience.

The third stage will add 3-D sound with height channels and higher-order ambisonics. Here, the end user needs to have a renderer in the device to generate the sound field and adjust volume, phase, and timing across the entire speaker array and the listening environment.

Finally, the systems will add dynamic objects which will allow the audio objects t track the video action. The system implementation will have to sync broadcast and broadband streams and maintain a constant latency. The data volume for a dynamic object plus metadata can be as large or larger than the video if it is not highly compressed.

The full system may require new sound frames for TV sets. These frames will go around the display to replace the 2-D sound bars or internal speakers with a wrap around frame that uses audio processing technologies and other software to drive the sound for proper psychoacoustics. Fraunhofer demonstrated a sound frame and a new codec on a tablet to show the relative ease for system upgrades to the new MPEG-H formats. See www.mpeghaa.com for more details.
 

Similar Posts