ETIA 2014 Audio Panel
June 18, 2014, SMPTE Entertainment Technology in the Internet Age Conference, Stanford, CA—The ability to deliver full 3-D audio greatly enhances the AV experience. Sunil Bharitkar from Dolby moderated this panel. Panelists were Sripal Mehta from Dolby, Roger Charlesworth from Charlesworth Media, and Philip Hilmes from Lab 126 and Amazon.
Mehta opened with a short tutorial on object-based audio. This technology was first developed for cinema where the space and environment are relatively fixed. The increased interest in sound quality is driving the technology to other platforms and media, and the advances in processor performance now make it practical to use the technology in many other places. The changes in sound quality lead to better user experiences and give the user control over balance and mix as well as choices in language.
The ability of objects in sound make audio adaptable to any device and any environment. In a demonstration, he showed a hockey game clip with 3-D audio. The sound capture was standard microphones plus a couple of extra mikes for crowd sounds. Some post processing enabled the objects and also made the audio choices easy to use.
In one example, he changed the announcer during play. Decoupling the audio from the video and speakers requires more metadata, but the audio can be rendered in the cloud to minimize the CPU loading on the display device. The different presentations all sync to the video, with the default being crowd and commentary in a 5.1 audio image. The user can use the raw audio feed or redefine at the point of use.
To reduce complexity and improve performance, the system uses presets for mix and balance, although the settings are still adjustable. This system enables broadcast by uplinking to the head-end, through the transmission channel to the consumer. Render is in the cloud for mobile or render the whole package at a fixed end point.
Charlesworth described the challenges for object-based audio personalization. For live content, there is no time for the processing compared to cinema where render time is a part of the finishing steps. For cinema or games, the non-real-time aspects are not an issue and all of the objects can be automated for the best user experience.
A live workflow for broadcast generates a single, flat mix and other flat streams for other languages. There are no breakout handles and no metadata for other mixes, resulting in a complete change in workflows. Even with new workflows, the objects still have to work with existing 5.1, 7.1, or 2.1 systems. Isolating the elements leaves rights issues for the broadcast. The mix and monitor tools for this type of audio work are in place, so a system with table-based metadata and static objects is easier to implement.
In live audio, the dynamic needs change the metadata. In addition, there is a cultural issue in production; many live feeds have freelancers with minimal training, older tools for mix and monitoring are not able to handle the complexities of object-based audio, and the industry is highly based on tradition. The problem with tradition is that it is an open-loop process with inertia.
On top of the technology issues are the problems with music rights. The technology is not the biggest issue in broadcast, but the legacy equipment is probably limited to 16 audio channels. In comparison, a digital object audio system is modular and flexible. The resulting cost efficiency will drive adoption.
A digital system enables automation, reduces workers, increases flexibility and can render at any point in the flow. All of the processing can be done on existing infrastructure. The transition to IP with organic metadata removes the signal count limits and allows virtualization for the production processes. The various input and output files can be managed with standard communications protocols. An object-based flow enables algorithmic mix optimizations, automated QA, and error checking. Real-time content algorithms can perform metadata checks.
Hilmes commented on the challenges of high-quality sound on mobile devices. The issues are small speakers, ear buds, noisy environments, and battery life. The quality issues are due to the variability, but the latest advanced sensors can lead to greater personalization and a better user experience. A user ID can indicate preferences about sound; volume, equalization, and the drivers for the I/O devices-speakers or headset. This ID can create new metadata about personal preferences before re-render against the 3-D head model
For hearing-impaired people, 80 percent don’t have hearing aids, and many have significant high-frequency losses. An audiogram estimation can provide changes in equalization for these people. Location and reference of head position relative to the screen as well as head tracking can enable on-the-fly corrections to the sound.
The user hears the environment, the content, and the qualities of the playback device. The personalization could take into account the brand of headphones for even more customization. Furthermore, the equalization could change if the user changed from a mobile device to streams on a TV or into a sound system.
Developments in improving bandwidth and reducing latencies permit cloud processing into the MEMS speakers on mobile devices. Integrated sensors and content metadata can provide a personalized optimized mix.
Machine learning and optimization?
Hilmes suggested a personalized record of volume settings and content playback preferences for this function. Lots of data exist but most are unused. A generic template is easily generated by averaging over 1M users. The system can incorporate both active and passive feedback, with passive feedback on an opt-in basis.
Do people like this audio?
Hilmes said yes, they will appreciate the better user experience. Noise canceling headphones are OK in a stable environment, but cannot adapt well to a highly variable one.
Charlesworth declared that a personalized system could allow for a different balance for voice and effects. The audio environments for a producer and active user are very different and a need for changing the mix is very real.
Experiments and data rates?
Mehta noted that existing productions separate voice and effects, so 160 Kbps is sufficient. There is a question on delay.
Charlesworth stated that the channel count easily allows for 16 objects into a phone. In production, the NFL has 32 channels for the different mikes.
Codecs?
Mehta declared that SDI and other formats allow modular updates, so it is practical to integrate new technologies.
Charlesworth suggested that it is easier to encode metadata than an IP export.
Hilmes added that IP solves many problems by decoupling the audio and video.


