Multimedia Search UI
February 7, 2013, Electronic Imaging Conference, Burlingame, CA—Marti Hearst from UC Berkeley described the requirements for future search user interfaces for multimedia. The growing volume of non-text entries on the Web demands a different search UI.
Many of the search principles for non-text search are highlighted in the website http://searchuserinterfaces.com/book/, which identifies the issues associated with images and video and a little audio. The capabilities for search in a text-oriented format are well established. In fact, if you want to compare the appearance of a ’76 Mustang and a ’07 one, the images will come up in almost any search engine. The search for car images are independent of the search engine, and a have similar look and UI. The differences between the search engines are invisible to the user.
The view of a search tool and its results have not changed much over time. A search from ’97 looks like a search from ’07, with ranked lists of results. The underlying engines and cosmetic differences still have a generic look and feel, and are accepted by users as standard. The changing nature of the content is being ignored by the search engines, which still favor text-based results. Image search is poor, and is mostly popular and commonly referenced images.
Text search uses faceted navigation: first define categories such as shopping, then allow the user to drill down to specific items. The basic presentation is good because is reduces mental work and offers alternatives, previews, and histories of the search. The greatest benefit is that people are better at recognition than at recall, so the common shopping metaphor is a meaningful organization.
The faceted flow is much better for navigating known categories, especially when compared to linear organizations like Wikipedia. Clustering also helps. Putting objects into a well structured format with integrated search improves the visual search operations. For images, small details really matter and it is important that the search engine not give empty results.
One good example of a good UI, is Getty Images, which facets the information, adds similarity and statistical information to the search results. One operational issue is how to determine the role of the metadata in defining similarity. A good search will provide relevant feedback, but a text-based interface is easy for computers and hard for people, and an image-based UI is better for people and very hard for computers.
One way to improve the feedback is to improve relevancy by adding user-generated tags to the computer-based retrieval. Combining text and other features eases the task for both people and computers. Combining the images and text provides both parties with useful context. Updating the relevance requires meaningful labels. One way to improve the labels is to crowd-source the content and create multi-dimensional labels that address the facets and the similarities and create the metadata.
Another item is to get real-time auto-suggest. The increasing search times for standard text-based queries using keywords implies that the searches need to change to some other input method. Longer, natural language queries offer greater user satisfaction, because they people prefer natural language over keywords. How many times have you tried to identify the key word that gets you your desired results? Most people are not good at “keywordese”.
Natural language queries are also able to link to other people’s searches and can enable a social response in user-based language. When a keyword search fails, a natural language UI can check for similar queries and provide answers and links that others have found useful. Ultimately, this can be like a blend of “sloppy” commands. The query would be like a command language with user flexibility and visual feedback. The concatenation of recognized components, highlights, rich data, and limited graph search would be much more useful to users. Natural language can be combined with auto-suggest to provide the best of both paradigms.
So how do you apply these concepts to multimedia? Integrate auto-suggest with faceted metadata and similarity. The context-based retrieval and touch or point selection will allow the natural input and improve responsiveness, resulting in a more meaningful experience.
In other areas, it might be better to solve specialized but high value search problems. The simple and generic problems are already well solved, but complex issues are out of reach of most search engines. For example, categories like travel and legal require satisfying many requirements that may be partially self-contradictory. With the compute power behind the search tools, is should be easy to ask ” illustrate my slide presentation” because the images only have to be close enough.
The ability to search without requiring expertise is important. Say you needed to replace your car windshield. How do you find all of the correct parts without knowing what they are called? In the patent area, one issue is to detect differences in documentation as well as find common points. The mis-identification of the features can be costly. Adding extraction capabilities for equations and graphic elements still requires a person to oversee the process.
These examples are a subset of all search, and show the value of specialization. You don’t see these features in the widely adopted UIs and the major search tools cannot map similarities or show clusters in 3-D. users will benefit form 2.5-d and 3-d hierarchies and more color regions. Even text clusters don’t exist in the existing tools. The image-centric search needs a change in visualization per person and per query.
fork browser ui
Promising approaches for new search tools are starting to appear. Forkbrowser creates linkages that show nearest neighbors, history, and similarities. See figure. The tools iterates a user-centered feedback string and shows changes in searches. There are some layout issues in the tree-map and with the hierarchical clusters, and more interaction sis needed to become more highly adopted by average users.


