Automatic Mood Labeling
May 17, 2013, Music Meets Science Forum, SwissNex, San Francisco—Peter DiMaria and Ching-Wei Chen from Gracenote described their efforts in aiding discovery through music metadata and automated analysis tools. The company’s approach is to provide a sonic mood and curation service.
The challenges of discovery are in the numbers. There are over 10M songs available, and most discovery methods depend upon some prior knowledge as the seed for further investigations. Search assume you know what you are looking for. Browsing by artist, genre, or song title use an existing collection paradigm. These seed-based recommendations all need some starting point.
To address this issue, they added metadata on mood to 30M recordings. The data was machine processed to get a sonic mood profile. The overall goal is to develop a personal recommendation engine that creates a mood-based user experience. The software uses individual descriptors and not a set of standard terms, since users have different descriptions for the same emotions. To do this, they made a scalable classifier model.
It all starts with people for the taxonomy, mood definition, and a training set. They developed a 101 vector mood descriptor base. The taxonomy is new and iterative, and enables the end user to get to a desired state. The high granularity is necessary to capture nuances. The biggest challenge is that their work has to cover all genres, global origins, over all time.
In combination with cultural associations, the mood database must address many subtleties in addition to pure emotion. The dimensions for moods range from positive to dark, calm to energetic. In addition, the international nature of music requires local inputs. The user interface needs proper labeling, so direct translations of moods is not possible. Local editors define moods and descriptors at the local level.
Mood and similarities in the music can be mapped to multiple classes, and changing associations can expand the choices to similar content. The sonic mood and genre are automatically recognized through their mapping engine. Sonic versus lyrical moods or instrumental versus vocal styles are categorized. The training library for recognition is a 10,000 recording set that has been evaluated by listeners. The mood annotation may use only a part of a song, since many songs have variable moods.
The supervised machine learning starts with a song and mood metadata. The recognition engine extracts sonic features and levels for a new model. Some of the parameters include spectral-frequency and frequency changes over time-tonal, rhythmic features, harmonic, and percussiveness. They train the engine and make changes to the models for beter distribution of the classes. Then the engine predicts a probability to make a score. The highest score becomes the primary mood. All of the outputs are put into the user interface classifier to allow the selection of a range or a specific term.
A new application is in some of the latest Mercedes show cars. A mood grid is displayed on the infotainment center to allow the driver to pick music to match a desired mood, or to evoke a different mood. For example, when stuck in traffic, a calming mood can be picked to help reduce rush-hour frustration. Other aftermarket infotainment systems are also offering these capabilities.


