|

Context-Aware Voice Activity Detector

February 25, 2015, ISSCC, San Francisco—K. Badam from Katholieke Universiteit in Belgium described context-aware hierarchical information sensing voice activity detector. The intent is to replace always-on sensors and their related processing with a more power efficient approach.

The main challenges in user interfaces is that the many sensors are always on and sending data to the processor. The massive quantities of data are mostly blind, and of little use to the user. This operating mode applies full signal processing to all signals and maximizes power consumption. An alternative is to only spend resources on those data that carry relevant information.

Some devices use a very coarse power management scheme that adds a wake-up state between the sleep and active modes. Still, much of the time is spent in high-power modes as the device turns on all functions as soon as a signal is detected.

Energy proportional sensing offers scalable information extraction by turning on more processing as the signal is determined to be of greater interest. With no signals, the system is mostly off. When some signals are present, a classifier determines if the signals are of interest and if so enables more advanced signal statistics processing. upon determination that the signals are relevant, the full signal processing chain starts operating.

The key enablers for energy proportional systems are hierarchical and scalable sensing, and adaptive and contact-aware information extraction. Upon signal detection, the various sub-systems start functioning. Each mode provides more utility than the previous mode and uses more power. Therefore, the hierarchical activation scales power with information.

In addition, the system has to be context aware. The context drives a dynamic feature activation or deactivation as the machine learns which active features are relevant. In the case of audio, voice detection is becoming more common, and most devices use about 50-100 uW for the processing. A low power enables a reactive sensor interface.

In this system, a wakeup detector uses a controllable gin amplifier to feed into a clocked comparator. This system watchdog generates a high false alarm rate but all signals over threshold go to a classification stage. The analog feature extraction is a scalable function block that is controlled by a context-aware control register.

The nice ting about voice recognition is that the signals share the frequency band for many non-speech signals, but the sound power levels are only useful in a limited range, and in fairly narrow frequency windows. As a result, the analog feature extractor uses 16 frequency bands and measures the average signal value in each band. If the energy profiles are not consistent with speech, the decision tree classifies the signal as noise and prevents further processing.

If the classifier determines ths signal is voice, it turns on the applications processor for further processing. The classifier detects context changes and the decision tree relearns for context changes as a function of power use. The information in each signal band is translated to DC for low sampling rates for additional power reduction.

The implementation chip does a good job of discriminating car, babble, and exhibition noise from speech with almost 90 percent classifier accuracy at 6 uW for feature extraction and classification. This energy proportional computing technique is promising for other sensor interfaces.
 

 

Similar Posts