Sounding Great Over the Web
June 17, 2015, SMPTE Entertainment Technology in the Internet Age, Stanford, CA—A panel reviewed the issues and technologies for audio in Internet content. Sunil Bharitkar from Dolby moderated, with panelists including Tim Carroll from Linear Acoustics on Skype, Jeffery Reidmiler from Dolby, Robert Fisher from Warner Bros. and Phil Hilmes from Amazon Lab126.
Bharitkar opened with background information. Audio is key to engagement for any entertainment. Consumption depends on types of content, delivery modes, and playback devices. The quality of the sound is influenced by the noise contributions which are affected in turn by such factors as codec settings, multiple encode and decode processes, lost metadata, playback equipment, and overall signal path.
There are lots of break points in this signal path since quality is a function of the ecosystem, that includes the capture quality, transcoding, metadata, and playback devices. One critical issues is to maintain the artistic intent.
Fisher noted that an emerging issue is the need to create audio for small devices. The digital transforms require a near-field reference equipment room per SMPTE 222M, that is flat to 4 kHz then rolls off. This reference room is very different from that of a theater. The workflow starts with a mix, dupe, restore, clean-up, and conform including fixing dropouts and re-mix. Mostly they use Dolby AC3 or DTS for streams. In the future, more work will be for mobile and Dolby ATMOS to add high dynamic range audio. Still, there is a great need for standards and systems to support all devices.
Reidmiller commented that a key to personalized audio is some standard file interchange with metadata, standard real-time interchange with metadata, loudness measurements, and some confidence monitor and quality control system. The audio industry has to change its mindset from packaged audio to standards for interchange. the ITU-R BWAV defines metadata and interoperability for synchronizing the audio and mandatory metadata for personalized audio. This latest standard also changes to IP per AES 3-ST337 and ETSI TS102 366 V1.3.1 annex H.
Ultimately, the efforts need to preserve the original work. The newest audio standards are adding channel count, defining audio objects, and other functions. The first standards will be out later this year. When the audio is associated with video, the focus is on timing. The new standards decouple quality from features, so data bandwidth, channel count, etc. are now independent to take advantage of the improved infrastructure. The integration of audio and metadata are essential and enable greater personalization.
Hilmes suggested that the consumer, or last quarter mile, are creating new issues due to the myriad devices being connected. For example, Blu-Tooth from tablet to headphones creates a non-synchronous stream over WiFi. To make matters worse, the playback equipment and other acoustics can be very noisy. Time synchronization between audio and video, and audio channel to channel has to be within a few milliseconds or users become uncomfortable. The detection threshold for synchronization is about 1 ms.
Few standards exist for audio synchronization, and the situation gets worse with more equipment. Blu-Tooth is the most common interface, and also causes the most problems. Latency can be from 20-400 ms, and can vary across the whole range. There are no protocols for synchronization and WiFi solutions are incompatible with many other protocols. In addition, addressing the issues may take highly trained technical support people.
The acoustic issues make matters worse. Many devise have poor speakers and the users are using their devices in non-ideal settings with lots of noise. Some of the devices can barely achieve 2-channel performance. Some improvements on the device side provide post-processing and better designs can aim the sound to the listener. On the cloud side, the metadata standards enable 2-way communications to facilitate the optimization of the whole playback chain. There are many opportunities for standards.
Define immersive and next-generation audio?
Reidmiller offered a system view to produce more life-like 3-D audio. Ideally, it is interactive and personalized to permit changes in language or mix, etc. As content moves from cinema to broadcast and homes, Hollywood and the creative organizations will embrace object-based flows for native work, and will mix down for broadcast.
Challenges and hurdles for distribution?
Carroll claimed that the greater number of channels have to get through the old, small pipes. Internet protocols may provide the necessary time for 128 channels so the technology itself is only an incremental change. The move to object-oriented audio will create a better user experience at the cost of re-architecting the whole processing path and processing chain. To preserve artistic intent, the content should promote any change for convenience, but cannot detract from the original intent. It is possible to do multi-language flow in 5.1 audio for broadcast today with existing tools and ecosystems.
Retaining the metadata for 3-D is possible if the codecs change, otherwise it is not that hard to set up if the metadata identifies changes in other metadata. It all becomes a part of the encoded audio. It won’t work if the audio is not integrated and synchronized. The graphics industry has used metadata for decades in the background. Audio should adopt similar processes.
Creation and the move to mobile changes cinema and home masters?
Fisher stated that the best way is to generate a 5.1 for home and theaters and convert it to binaural for mobile. You still need some control of the content and quality. A 5.1 or Atmos in headphones requires headphone model information to reconstruct the audio sound field. Mobile and headphones is the only way. If you add haptics you need to synchronize another transducer.
Although the focus is on cinema, the work flow for phones and tablets just needs to compare the audio to the original for accuracy. Then you just render while keeping the original mix from the dub stage. Check for quality and auto render.
Amazon is delivering content, services, and hardware. How do you ensure quality at the consumer end?
Hilmes responded that technologies like variable bit rate etc. need more work to become seamless. The object-oriented audio formats acknowledge bandwidth. Dialog is sent first, then the background, which can drop out. The technologies need more automation to detect bandwidth, especially with the high variability in mobile devices.
The privacy issues due to the cameras, speakers, and mikes on the mobile devices is a challenge, since voice control is one of the best user interfaces. They limited the vocabulary and require a keyword to activate the listener function. It is important to have clear communications with the consumer about the various functions and how to protect privacy in this environment.


