New research shows that reorganizing a model’s visual representations can make it more helpful, robust and reliable
“Visual” artificial intelligence (AI) is everywhere. We use it to sort our photos, identify unknown flowers and steer our cars. But these powerful systems do not always “see” the world as we do, and they sometimes behave in surprising ways. For example, an AI system that can identify hundreds of car manufacturers and models might still fail to capture the commonalities between a car and an airplane, i.e. both are large vehicles made primarily of metal.
To better understand these differences, today we’re publishing a new paper in Nature analyzing the important ways AI systems organize the visual world differently from humans. We present a method for better aligning these systems with human knowledge, and show that addressing these discrepancies improves their robustness and ability to generalize.
This work is a step towards building more intuitive and trustworthy AI systems.
When you see a cat, your brain creates a mental representation that captures everything about the cat, from basic concepts like its color and furriness to high-level concepts like its "cat-ness." AI vision models also produce representations, by mapping images to points in a high-dimensional space where similar items (like two sheep) are placed close together, and different ones (a sheep and a cake) are far apart.
To understand the differences in how human and model representations are organized, we used the classic "odd-one-out" task from cognitive science, asking both humans and models to pick which of three given images does not fit in with the others.
Source link







