| |

Using fMRI to Reverse Engineer the Human Vision System

February 4, 2014, Electronic Imaging Conference, San Francisco—Jack Gallant from the University of California, Berkeley described his research into the human vision system. These efforts are important because computer vision and human vision are similar in some ways and very different in most others. One quarter of our brain is associated with vision.

Comparing brains and computers, we see that the computer is relatively hardware and software independent where as the brain is configured so that the hardware constraints the software. It takes about 20 minutes to form a new synapse. Current models of vision system in humans are centered around 40 visual areas within the brain. These areas constitute a complicated network with over 1600 connections and lots of feedback.

The visual areas are also built in three levels of hierarchy. The lowest level looks for simple features like edges, the next areas address grouping and segmentation, and the top levels identify semantic categories. Researchers are starting to use fMRI to measure blood volume as an indicator or indirect at go of activity. Unfortunately however, there is only general correspondence between blood volume and actual brain activity, so work continues in mapping these active areas.

One challenge is that the interconnections a very complicated pattern, with over 60 thousand point on the brain map. The system has an identification problem to map 1-0 pixels to brain activity in a non-linear structure. The response is to assume a nonlinear transform and find the pixel-feature space through a linear transform. Then map the feature space into the activity space.

Voxel-wise modeling facilitates this transform, by reducing individual pixels into structures. The resulting movies show the brain activity in the expected functional areas; a motion engine, semantics, and categories. Running the activity patterns though a finite impulse response filter show that 4-5 k features are sufficient to generate reasonably accurate predictions and allow the evaluation of activity.

A test focusing on the low-level vision areas has a movie as the input and neural activity as the output. The test results show that their configuration can predict the outputs of the neurons. The results show good correlation in the early and intermediate areas indicating that the choice of a voxel of 1M neurons is realistic. The high-level semantic object and action models still need refinement.

Although the process currently uses manual transforms and keyword generation, they now can take a movie as input and predict features and responses. They generated a matrix of 1785 nouns and verbs to define categories and structures at the individual voxel level. The subjective preferences agreed with the first 5 principal components across 90 percent of all test subjects.

The next step is to develop groupings of semantics and map them the areas on the brain. It appears that text is on a separate vector from other semantic evaluations. The existing maps of object and action categories tends to be centered in the back of the brain, but is dynamic due to the taxonomy hierarchy and the focus of attention. If a test subject is passive, you get one type of response, if the defined target is a human, you get a different response, which differs from a inanimate object like a vehicle.

This shift of activity is based on the attention or object and the voxel activity shifts accordingly. This dynamic response indicated the need for a general model of mental standards. Some of the characteristics include a scene category model, a natural scene study, a word network and language tree, and others.

To date, they have had to hand label all objects within an image leading to category features the modified weights to predict voxel responses. Current efforts have started out with 25 categories, which are a good fit for natural image statistics. Some work is ongoing to map scene selection to brain areas.

The mid-level vision and segmentation and grouping is hard reproduce. There are no good models for this level of brain activity. The closest models integrate a new feature space that combines the high-and low-level feature sets. As a result, the new mappings are just an engineering problem in encoding and decoding images.

One issue is that comparisons require prior images for the structural and semantic decoding. They used 50 M Flickr images as a proxy for the prior image baseline. The MRI takes about 1-2 seconds for capture, and 10 seconds for decay. Using 5,000 hours of YouTube videos, test subjects were imaged and did well in showing good agreement across subjects and responses. the semantic decoder plus a filter recovers much of the brain functions. Now, they think that the state of the brain can be decoded with about 1825 photo images.

The next stage will add audio to the images and will try to get greater resolution. A single neuron is not interesting, and the MRI has too little resolution. A resolution of between 0.5 and 0.75mm is about the size of a brain column. The issue of locality is already addressed by the brain structure as a set of columns. There is some difference in tuning with different languages, because the neuron diversity makes the imaging that of an ensemble of complex functions.
 

Similar Posts