Learning to represent scenes as collections of discrete entities with properties rather than as raw pixels.