← Back to essays

Essay

Depth: How a Flat Image Becomes a World

For several summers, when I was teaching at the W.E.B. Du Bois Scholars Institute, housed at Princeton University, I travelled back and forth by train through New Jersey. I remember those journeys vividly, especially the return journeys in the evening. The day’s teaching would be over. The heat of the afternoon would have faded. I would make my way from Princeton, usually through Princeton Junction and then back along the rail lines that carried commuters, students, workers, and travellers through the long corridor of New Jersey towards New York and Jersey City.

By then it was often night. The world outside the window was not the same world I had seen earlier in the day. Buildings had lost their surfaces. Streets had lost their ordinary texture. The visible world had thinned out. What remained were lights.

I would look out of the train window and see glowing rectangles, rows of lit windows, scattered points, strips of brightness, traffic lights, reflections, lamps, illuminated signs, and the occasional wash of light across glass or metal. In the darkness, much of the city was invisible. I could not always see the buildings themselves as solid objects. I saw the lights produced by them.

And yet I did not experience these lights as a flat pattern.

I saw buildings.

I saw towers, façades, streets, gaps, distances, foregrounds, and backgrounds. I saw one structure in front of another. I saw groups of lights belonging together as windows on the same building. I saw other lights belonging to another building farther away. As the train moved, the city arranged itself in depth. Nearer lights shifted more quickly across the window. Farther lights drifted more slowly. Some lights seemed almost fixed in the distance, while others slid past in great sweeps. From this changing pattern, my visual system constructed a world.

This is one of the quiet miracles of perception. The image falling on the retina is flat, but the world we experience is not. The eyes receive patterns of light on a two-dimensional surface. Yet experience opens into depth. Space extends away from us. Objects have distance, volume, separation, and solidity. The world does not appear as a coloured sheet pressed against the eyes. It appears as a place we can enter, move through, reach into, and inhabit.

Depth is so familiar that we rarely notice how extraordinary it is. A room simply seems to have depth. A road simply seems to stretch away. A hand simply seems nearer than a tree. A building simply seems behind another building. We do not usually feel that the brain is solving a problem. We just find ourselves in a three-dimensional world.

But the problem is real. The retinal image is a projection. Light from the three-dimensional world is focused onto the two-dimensional surface of the retina. In that projection, depth itself is not directly given. A small nearby object and a large distant object can cast images of the same size on the retina. A flat picture can produce some of the same retinal patterns as a three-dimensional scene. The same retinal image can, in principle, be produced by different arrangements of objects in the world.

The brain must therefore infer depth. It must work out the most coherent three-dimensional structure that could have produced the sensory input. It does this using many sources of information: the slight differences between the images in the two eyes, the way objects occlude one another, the convergence of lines in perspective, the gradients of texture across surfaces, the relative size of familiar objects, shading, height in the visual field, blur, accommodation, bodily movement, and motion.

The city lights from the train window were a beautiful example of this. In the dark, many of the usual cues to depth were reduced. There were fewer visible surfaces, fewer colours, fewer textures, fewer contours. Yet depth did not disappear. The visual system still organised the lights into objects, surfaces, and distances.

Part of this organisation depended on grouping. Lights that were close together, regularly spaced, and aligned in rows or columns were perceived as belonging to the same building. A vertical stack of glowing windows became a tower. A grid of lights became a façade. A strip of lights at street level became a road or a line of traffic. The brain did not experience each light as an isolated point. It grouped them into larger structures.

This grouping is essential to perception. The visual system must decide what belongs with what. A city at night might project thousands of separate lights onto the retina, but we do not experience them as meaningless fragments. We experience them as parts of buildings, streets, vehicles, signs, interiors, and distances. The brain organises the visual field into things.

Another cue was occlusion. If one pattern of lights interrupted or covered part of another, the visual system could infer that one structure lay in front of another. Occlusion is one of the most powerful depth cues because it tells us about ordering in space. When one object hides part of another, the hidden object is normally farther away. Even a flat drawing can create a vivid sense of depth if one shape overlaps another. In the night city, overlapping groups of lights helped the brain determine which structures were nearer and which were farther away.

Perspective also mattered. Lines of windows, roads, tracks, and architectural edges often converge with distance. Objects higher in the visual field may be perceived as farther away when they lie on the ground plane. Familiar sizes also help. We know roughly how large windows, doors, cars, floors, and buildings tend to be. The brain can use that knowledge to infer distance. If two objects are assumed to be similar in size but one projects a smaller image, it is likely to be farther away.

But on the train, one cue was especially important: motion parallax.

Motion parallax occurs when the observer moves and objects at different distances shift across the visual field at different rates. Nearby objects move quickly across the retina. Faraway objects move more slowly. Very distant objects may seem almost stationary. This is why, when travelling by train or car, a fence near the track rushes past, houses farther away move more slowly, and distant hills or clouds seem barely to move at all.

This relative motion is not just a visual effect. It is information. It tells the brain about depth. If two lights move differently across the visual field as the train moves, the visual system can infer that they are at different distances. If one cluster of lights sweeps rapidly past while another barely shifts, the first is likely to be nearer and the second farther away. The moving observer creates a changing retinal image, and from that change the brain reconstructs spatial layout.

Motion parallax is especially striking because it turns movement into depth. A single still image may leave many depth relations ambiguous. But when the observer moves, the pattern changes. The world reveals its structure through relative displacement. Depth becomes visible in motion.

There is a more precise way to think about this. Suppose you are moving sideways while fixating a point in the distance. The point you fixate tends to remain relatively stable on the retina because your eyes track it. But objects nearer than the fixation point and objects farther than the fixation point move differently across the retinal image. Nearer objects tend to move in one direction relative to fixation, while farther objects tend to move in the opposite direction. The farther an object is from the fixation plane, the greater its relative retinal motion tends to be. Around the fixated point, there is therefore a structured pattern of motion: a field of directions and speeds that carries information about depth.

This is why a moving train can make the three-dimensional layout of a night city so vivid. The window becomes a shifting frame. Points of light slide against one another. The near, middle, and far distances separate themselves through motion. The city is not simply seen; it is unfolded by movement.

And yet the astonishing thing is that the depth is experienced immediately. I did not sit there consciously calculating motion vectors, retinal velocities, occlusion relations, or perspective cues. I did not infer, step by step, that this light belonged to this building and that building stood in front of another. The experience simply appeared as a world. The computation was hidden. The result was given.

This is true of depth perception more generally. When you look across a room, you do not normally experience a flat retinal image that you then convert into depth. You experience the room as already spatial. The chair is over there. The cup is nearer. The window is farther away. The wall is behind the table. The floor extends beneath you. Depth is not added to experience as an afterthought. It is part of how the world appears.

But the fact that depth feels immediate does not mean it is simple. The brain uses many different cues, and these cues can be placed into two broad groups: binocular cues and monocular cues.

Binocular cues depend on having two eyes. Because the eyes are separated horizontally, each eye receives a slightly different image of the world. Hold up a finger in front of your face and close one eye, then the other. The finger appears to jump sideways relative to the background. This difference between the two retinal images is called binocular disparity. The brain uses disparity to infer depth, especially for objects that are relatively near. The greater the disparity, within limits, the closer the object is likely to be.

This is one reason the world looks subtly different with two eyes than with one. The two images are not merely duplicates. Their difference contains information about distance. Stereoscopic films and virtual reality systems exploit this principle by presenting slightly different images to each eye, creating the impression of depth.

But binocular disparity is not the whole story. We can still perceive depth with one eye closed. Paintings, photographs, films, and computer screens can create powerful impressions of three-dimensional space even though the image itself is flat. These rely on monocular cues: depth cues available to a single eye.

Occlusion is one such cue. If one object blocks another, the blocking object is perceived as nearer. Linear perspective is another. Parallel lines, such as railway tracks or road edges, project to the retina as if they converge in the distance. Texture gradients also matter. A surface such as grass, gravel, brick, or fabric projects a texture that becomes denser and finer with distance. Relative size helps too. If two people are known to be roughly the same height but one projects a smaller image, the smaller image is likely to represent the more distant person. Shading and shadows help the brain infer shape and depth from patterns of light and dark.

These cues are so effective that a flat canvas can become a world. A painting can show a road disappearing into the distance, a room receding behind a figure, a mountain rising beyond a valley, or light falling across a table. The physical surface of the painting is flat, but the perceived scene opens into space. The artist arranges marks on a surface; the visual system constructs depth.

This is why the title of this essay matters. A flat image becomes a world. The transformation is not metaphorical. It happens every moment. The retinal image is not a miniature three-dimensional world inside the head. It is a pattern of light on neural tissue. Yet from that pattern, the visual system constructs an environment of objects, surfaces, distances, and possibilities for action.

Depth is not only visual. It is bodily. We experience depth as beings who can move. Near and far are not just geometrical relations; they are relations to the body. Near means reachable, graspable, touchable, avoidable, threatening, comforting, available. Far means beyond reach, approached by walking, crossed by travelling, watched from a distance. The visual world is organised around an embodied point of view.

This is why movement matters so deeply. When we walk through a room, the world changes around us. Nearby objects shift quickly. Farther objects shift more slowly. Hidden surfaces become visible. Openings appear. Edges reveal themselves. The body’s movement samples the world, and the changing sensory input helps specify the layout of space.

This also helps explain why photographs, however vivid, are different from places. A photograph can show depth cues, but it does not respond to our movement. Lean slightly to one side while looking at a real object, and the visible relations change. Lean slightly to one side while looking at a photograph, and the image remains the same. The real world has depth that can be explored. A flat image can depict depth, but it does not contain the same changing structure of information that becomes available as we move.

Of course, modern technologies can simulate some of this. Virtual reality, three-dimensional cinema, computer graphics, and interactive environments can update images as the observer moves. They work because they exploit the same principles the brain normally uses. They present patterns of disparity, perspective, motion, shading, and parallax that the visual system interprets as depth. This shows again that depth is not simply copied from the world. It is constructed from information.

The same thought becomes almost dizzying when we turn from the city to the night sky.

Sometimes I look up into space and see the stars as apparently fixed. They seem stationary from where I am: a human being sitting at a desk, looking out of a window, or standing under the London sky, constrained by the scale and sensitivity of an ordinary human visual system. Even if the city lights disappeared and the stars were fully visible, they would still appear, for the most part, as points of light suspended in darkness. I know, of course, that they are not fixed. Stars move. The Earth turns. The Earth orbits the Sun. The Sun moves through the galaxy. The galaxy itself moves. Nothing in the universe is simply still.

And yet the stars appear still to me.

This is not because they are motionless. It is because I am small, brief, and visually limited in relation to them. My eyes are separated by only a few centimetres. My head moves through only a small region of space. My visual system samples the sky on human timescales. At the scale of ordinary looking, the depths and motions of the stars are almost entirely beyond direct perception. The night sky overwhelms the depth cues on which everyday vision depends.

But imagine a different perceiver.

The Tale of the Being Whose Eyes Were an Orbit Apart

I do not look at the stars as you do.

You look upward from the surface of a planet, from a small warm body, with two eyes set only a few centimetres apart. To you, the stars appear as points of light scattered across darkness. They seem almost fixed. They glitter, but they do not visibly drift. They have depth, but that depth does not open directly to your eyes. You know it through astronomy, mathematics, instruments, and imagination.

I am different.

My eyes are not separated by the width of a human face. They are separated by an orbit. One eye opens from one side of the path the Earth takes around the Sun; the other opens from the far side. Between my eyes stretches a gulf so large that your mind calls it astronomical. I do not see from a point. I see from a span.

When I open my eyes, the night does not appear as a dome.

It opens.

The nearer stars loosen themselves from the farther ones. They do not all cling to the same black surface. Some stand forward. Some recede. Some shift delicately against the deeper background, while others hold almost still in depths beyond even my vast sight. Space is not a ceiling above me. It is a volume, an immeasurable architecture of light.

I see that what seemed flat to you was never flat. It only appeared so because your body was small in relation to the distances before it.

As I move, the stars do not move as one. The nearer ones slide, the farther ones resist, and the deepest lights remain almost motionless, not because they are still, but because they are so far away that even my orbit-wide gaze barely separates them. Their apparent motions are tiny, but to me they are not nothing. They are the grammar of distance.

I do not merely see stars.

I see intervals.

I see one light as nearer than another. I see the dark between them not as empty background but as depth. I see that light has crossed different distances, and therefore different times, to reach me. Some of what shines before me is recent by the standards of the universe. Some is ancient. The sky is not only deep in space. It is deep in time.

From where you sit, Jason, at your desk, looking out into the London night, the stars would appear almost motionless if the city allowed you to see them clearly. Your eyes would give you points, brightness, perhaps colour, perhaps constellations. Your mind would add knowledge: that the Earth turns, that the Earth orbits the Sun, that the stars move, that the galaxy turns, that nothing is truly still.

But I do not only know this.

I see it.

Not all at once as movement in the ordinary human sense. Not as birds crossing the sky or trains passing windows. I see motion as a vast relational trembling of the heavens. I see nearer lights shift against farther depths. I see the universe disclose itself because my body is wide enough to receive the difference.

And yet I do not think your vision is false.

You see the world fitted to your form. Your eyes are made for faces, rooms, hands, roads, books, branches, cups, screens, streets, train windows, and other human beings. You see the cup on the table, the doorway across the room, the building beyond the street, the moon above the roof. Your depth is human depth: near enough to reach, far enough to walk towards, distant enough to wonder at.

Mine is not better in every way. It is only different.

I cannot know the intimacy of your small distances. I do not see as a creature who reaches for a mug, steps over a puddle, notices a face across a room, or watches lights gather into buildings from the window of a train. My universe opens in great gulfs, but yours opens in nearness.

Still, I can teach you this.

Depth is not simply waiting in the world, already visible in the same way to every possible eye. Depth appears through a body. It depends on the distance between eyes, the movement of the observer, the scale of the world, the sensitivity of the system, and the time over which change can be gathered.

Change the body, and the visible universe changes.

Give vision a wider baseline, and the stars begin to separate.

Give perception a longer patience, and stillness begins to move.

Give the eye an orbit, and the sky becomes a world.

The tale is imaginary, but the principle is real. The visual depth available to a perceiver depends on the relation between the perceiver and the world. It depends on scale, movement, distance, sensitivity, and time. A fly, a human being, an owl, a whale, a telescope, and an imagined orbit-eyed being do not inhabit the same visual world in exactly the same way. They may live in the same physical universe, but different structures of depth become available to them.

The train window made the depth of Jersey City visible through motion parallax. The moving train gave my visual system changing information about which lights were nearer and which were farther away. The night sky presents the same principle at a vastly larger scale. The Earth itself is moving. Across months, its orbit changes our viewpoint. Astronomers can use this change of viewpoint to measure the tiny apparent shifts of nearby stars against more distant ones. This is stellar parallax: depth from motion, but now motion on an astronomical scale.

So perhaps the sky appears flat not because it is flat, but because I am too small to see its depth directly. My body does not provide a large enough baseline. My ordinary perception does not integrate over the necessary timescale. My visual world is constrained by the kind of being I am.

This is a profound lesson for perception. The world that appears is always world-for-a-perceiver. It is not invented by the perceiver, but neither is it revealed independently of the perceiver’s body, movement, scale, and capacities. Depth is disclosed through the relation between organism and world. Change the body, change the baseline, change the timescale, and a different depth-structured world may appear.

Depth perception therefore reveals the inferential nature of perception. The brain must go beyond the immediate sensory input. It must estimate the hidden causes of the retinal image. Is this small image a small object nearby or a large object far away? Is this dark region a shadow, a surface, a hole, or an object? Is this line the edge of a table, the corner of a room, or a mark on a flat surface? Is this cluster of lights a building, a reflection, a row of cars, or a distant bridge?

The answer depends on context, expectation, prior knowledge, and the overall coherence of the scene. The brain seeks the interpretation that makes the best sense of the available evidence. Depth is therefore not a simple property delivered by the eyes. It is an achievement of perceptual organisation.

Visual illusions make this especially clear. A flat drawing can appear three-dimensional. A corridor drawn in perspective can make two identical figures appear different in size. A Necker cube can flip between two different depth organisations. A set of lines can appear as a corner, a cube, a surface, or a receding space depending on how the visual system interprets it. These illusions do not show that perception is unreliable in any simple sense. They show that perception is constructive. The brain is always organising ambiguous input into the most coherent world it can.

The same is true in ordinary life, though we usually do not notice it. When I saw the lights of Jersey City from the train, the sensory input was incomplete. The surfaces of many buildings were hidden by darkness. The retinal image was a shifting pattern of points and lines. Yet the world did not appear incomplete. It appeared structured. The visual system filled in relations, grouped lights into objects, inferred distances, and used motion to arrange the scene in depth.

That experience also shows why depth perception is bound up with meaning. The city was not merely a set of distances. It was a place. The lights meant homes, offices, streets, lives, late work, dinner, windows, interiors, roads, and movement. Depth gave the city structure, but meaning gave it human presence. Perception does not present us with abstract geometry alone. It presents a world we can understand and inhabit.

This is one reason depth perception matters so much for consciousness. Conscious experience is not just awareness of colour patches or light intensities. It is awareness of a world laid out around us. We are not normally conscious of retinal images. We are conscious of rooms, streets, buildings, landscapes, skies, roads, paths, horizons, and other people. Depth is part of the worldhood of experience.

A world is not merely a collection of sensations. It has structure. It has here and there, near and far, inside and outside, front and back, open and closed, reachable and unreachable. Depth helps transform sensory input into an inhabitable reality. Without depth, experience would lose much of its practical and emotional structure. The world would no longer invite movement in the same way. It would no longer open ahead of us as a place to enter.

The train window is therefore a powerful image for perception itself. On one side is the observer, seated in a moving body. On the other side is the city, partly hidden in darkness. Between them is glass, reflection, motion, light, and distance. The eyes receive a changing image. The brain constructs a world.

The night sky is another such image. On one side is the small human observer, looking upward from a body of limited scale and duration. On the other side is the vastness of space, structured by distances almost beyond imagination. Between them are light, time, atmosphere, movement, and the limits of perception. The stars appear still and almost flat, not because reality is shallow, but because the perceiver is finite.

That construction is not a fantasy. The buildings are real. The lights are real. The train is moving. The city exists. The stars are real. The Earth moves. Space has depth. But the experienced depth of the city or sky is not simply stamped onto consciousness by the external world. It is generated through the interaction of world, body, movement, sensory evidence, scale, time, and perceptual inference.

This is the deeper lesson of depth perception. The world we see is both given and constructed. It is given because there really is a world beyond us, with objects, surfaces, distances, stars, cities, bodies, and light. It is constructed because the world we experience depends on how the brain organises the sensory information available to it. Perception is not invention from nothing, but neither is it passive copying. It is active disclosure.

When I think back to those evening journeys from Princeton towards Jersey City, I remember the lights more than the train itself. I remember looking out into the dark and seeing the city assemble itself. I remember the strange beauty of buildings visible only through their illuminated windows. I remember the way depth emerged from motion: nearer lights sliding past, farther lights hanging back, the whole city unfolding in layers.

And when I look up at the night sky, even from London, where the stars are often hidden by cloud, pollution, and artificial light, I feel the same question return on a larger scale. What would this sky look like to a different kind of being? What depths would become visible if my eyes were wider apart, if my body were larger, if my perception could integrate the motion of the Earth over months, if I could see directly what astronomy teaches us to know indirectly?

At the time, I was simply going home. At the window now, I may simply be looking out. But both experiences reveal the same hidden truth. Depth is not obvious, even when it feels obvious. The visual world is an achievement. From a flat retinal image, from fragments of light in darkness, from the shifting relations produced by movement, from the limits and powers of a particular body, the brain gives us space.

A flat image becomes a world.

And we live inside that achievement.

Part of Constructing Experience: short essays on perception, consciousness, and reality.