PSYU2239 Week 6 Notes, Depth Perception

What is Depth Perception?

Depth perception = the ability to perceive the world in 3D and judge how far away objects are.

Images are projected onto each retina in 2D, but our perceptual experience is 3D.

So the visual system has to use depth cues to infer distance and spatial relationships.

Why do we need it?

Depth perception allows us to:

  • reach accurately for objects
  • navigate around obstacles
  • judge distances
  • drive
  • catch/throw objects
  • understand the layout of a scene

Depth Cues

A depth cue = information the visual system can use to estimate depth or distance.

There are several different ways of classifying depth cues.

This is probably the most important organisational framework for the module:

ClassificationQuestion
Monocular vs binocularDo you need one eye or two?
Pictorial vs non-pictorialDoes it work in a flat picture?
Visual vs oculomotorDoes the information come from the retinal image or eye muscles?
Ordinal vs metricDoes it tell you just the order, or actual distance?

Important: these aren’t four completely separate groups. One depth cue can have several properties at once.

For example:

Occlusion = monocular + pictorial + visual + ordinal

Monocular vs Binocular Cues

Monocular depth cues

Monocular cues work with just one eye.

Cover one eye and you can still use them.

Examples include:

  • occlusion
  • familiar size
  • linear perspective
  • texture gradient
  • height in the visual field
  • blur
  • motion parallax
  • accommodation

Binocular depth cues

Binocular cues require information from both eyes.

The major binocular processes covered here include:

  • binocular disparity
  • stereopsis
  • convergence
  • fusion

Pictorial vs Non-Pictorial Cues

Pictorial cues

Pictorial cues can create depth in a completely flat, 2D image.

That’s why a photograph can look three-dimensional even though the actual photograph is flat.

Examples:

  • occlusion
  • familiar size
  • linear perspective
  • texture gradient
  • height in the visual field

Easy test

Ask:

“Could this depth cue work in a photograph?”

If yes β†’ probably pictorial.

Non-Pictorial Cues

These require something that a static 2D image cannot reproduce, such as:

  • eye movements
  • changes in lens shape
  • movement of the observer
  • differences between the two eyes

Examples include:

  • accommodation
  • motion parallax
  • binocular disparity
  • convergence

Visual vs Oculomotor Cues

Visual cues

Visual cues come from information contained in the image falling on the retina.

Examples include:

  • familiar size
  • linear perspective
  • texture gradient
  • occlusion
  • blur

Oculomotor cues

Oculomotor cues come from information associated with the eyes themselves, such as muscle activity or changes in the lens.

Examples:

  • accommodation
  • convergence

These cues are particularly useful for nearby objects.

Ordinal vs Metric Cues

Ordinal depth cues

Ordinal = tells you the ORDER.

It tells you:

A is closer than B.

But not exactly how much closer.

Example: Occlusion

If:

Apple is in front of Orange.

you know:

apple is closer than orange

But occlusion alone doesn’t tell you:

apple is exactly 30 cm closer.

So occlusion is an ordinal cue.

Metric depth cues

Metric cues provide information about the amount of distance/depth.

They allow more precise estimates of:

how far apart things are

rather than simply which is closer.

Examples can include:

  • familiar size
  • linear perspective
  • motion parallax
  • blur

Easy distinction

Ordinal:

“A is closer than B.”

Metric:

“A is this much closer than B.”

Major Monocular Depth Cues

Occlusion

Occlusion = when one object blocks part of another object.

The object doing the blocking is perceived as closer.

For example:

If a tree blocks part of a house:

tree β†’ closer
house β†’ further away

Classification

Monocular + pictorial + visual + ordinal

It only gives relative depth order, not exact distance.

Familiar Size

We use our existing knowledge of an object’s usual size to estimate how far away it is.

If you know two people are roughly the same physical size but one creates a much smaller retinal image, you assume:

the smaller-looking person is further away.

This requires previous knowledge about the object’s normal size.

Example

You don’t normally think:

“Wow, that car in the distance is the size of a hamster.”

Your brain assumes:

“Cars are roughly this big β†’ therefore that car must be far away.”

Linear Perspective

Parallel lines appear to converge as they get further away.

For example:

Railway tracks

The tracks are physically parallel:

| |

But in the distance they appear to move toward one another:

\ /
\ /
\ /

The greater the apparent convergence, the greater the perceived depth.

This is heavily used in:

  • photography
  • drawing
  • architecture
  • paintings

to create the illusion of 3D space.

Texture Gradient

Textures appear:

large + detailed + spread out when close

and

smaller + denser + less detailed when far away.

Imagine standing in a field of flowers.

Nearby

🌼 🌼 🌼

You can distinguish individual flowers.

Far away

β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’

They become smaller and more densely packed.

The gradual change in texture provides information about depth and distance.

Height in the Visual Field

For objects resting on a ground plane, their position in the visual field can indicate distance.

Generally:

objects closer to the horizon β†’ perceived as further away

while

objects further from the horizon β†’ perceived as closer.

It’s another reason artists can create convincing depth in flat images.

Blur

The amount of blur in an image can also provide information about depth.

Objects at different distances from the point of focus can appear more or less blurred.

The visual system can therefore use:

relative sharpness vs blur

as information about distance.

Blur can provide metric depth information.

Accommodation

Accommodation is different from the pictorial cues because it comes from the eye itself.

Accommodation = the lens changing shape to focus objects at different distances.

The brain can use information about this lens adjustment to estimate distance.

It’s particularly useful for nearby objects.

Classification

Monocular + non-pictorial + oculomotor + metric

Don’t mix this up with convergence

Accommodation β†’ LENS changes shape

Convergence β†’ EYES rotate inward

Motion Parallax

Motion can also provide depth information.

Motion parallax = when you move, objects at different distances appear to move at different speeds across your visual field.

Imagine looking out a car window:

Nearby objects

Trees beside the road seem to:

WHOOOOSH past

Distant objects

Mountains seem to:

move very slowly

So:

near objects β†’ faster apparent movement

far objects β†’ slower apparent movement

Importantly, this cue can work with one eye.

Why Have Two Eyes?

Having two eyes provides several advantages.

Increased sensitivity

There are essentially two opportunities to detect a stimulus.

If something is very faint, information from two eyes can improve detection.

Larger/overlapping visual fields

Eye placement affects how much of the environment can be seen.

Depth perception

Most importantly for this module, two eyes provide slightly different images of the world.

The brain can compare these images to extract depth information.

Eye Placement: Predators vs Prey

Prey animals

Animals such as rabbits often have eyes positioned more toward the sides of the head.

This provides:

huge visual field β†’ good predator detection

BUT

less binocular overlap

Predator Animals

Animals such as cats and owls tend to have more forward-facing eyes.

This produces:

smaller total visual field

BUT

greater binocular overlap β†’ better depth perception

That’s useful when you need to accurately judge:

how far away the thing you’re about to pounce on is.

Binocular Disparity

Because our eyes are separated horizontally, each eye views the world from a slightly different position.

Therefore:

left eye image β‰  right eye image

The difference between these two retinal images is called:

Binocular disparity

The brain compares the disparity between the images to estimate depth.

Disparity and Distance

The amount of disparity changes depending on an object’s depth relative to where you’re looking.

In general:

Larger disparity

β†’ greater difference between the two retinal images

β†’ stronger information that the object lies at a different depth from fixation.

Smaller disparity

β†’ images are more similar

β†’ object is closer to the fixation depth.

This information contributes to the vivid perception of 3D depth.

Stereopsis

Stereopsis = the perception of depth produced by binocular disparity.

So don’t treat these as identical concepts:

Binocular disparity

= the difference between the two retinal images

↓

Brain processes disparity

↓

Stereopsis

= the resulting experience/perception of depth

Think:

Disparity = information
Stereopsis = depth perception created from it

Convergence

When you look at something nearby, your eyes rotate inward toward one another.

This is called:

Convergence

The closer an object is:

β†’ the more the eyes must turn inward.

The brain can use information from the muscles controlling the eyes to estimate distance.

Therefore convergence is an oculomotor depth cue.

Fusion

Each eye receives a slightly different image.

Fusion = combining the left-eye and right-eye images into one unified percept.

But fusion only works when the difference between the images isn’t too large.

Horopter & Panum’s Fusional Area

Horopter = Objects on/around the horopter can generally be combined into single vision.

Thankfully, the images don’t need to correspond perfectly.

There is a small region around the horopter where slightly different retinal images can still be fused.

This is:

Panum’s fusional area

If disparity is within this area:

two retinal images β†’ fusion β†’ one perceived object

If disparity becomes too large:

fusion fails β†’ diplopia

Diplopia

Diplopia = double vision.

It occurs when the disparity between the two retinal images becomes too large for the brain to fuse them.

Example:

Holding your finger close to your face and focusing on something behind it. Your finger will look like it has doubled.

Binocular Vision Process

Two eyes view from different positions

↓

Slightly different retinal images

↓

Binocular disparity

↓

Brain compares disparity

↓

If disparity is within Panum’s fusional area

↓

Fusion

↓

One unified percept + depth information

↓

Stereopsis

If disparity becomes too large:

↓

Diplopia

The Cue Classification Table

CueMonocular / BinocularPictorial?Information sourceDepth information
OcclusionMonocularYesVisualOrdinal
Familiar sizeMonocularYesVisualMetric
Linear perspectiveMonocularYesVisualMetric
Texture gradientMonocularYesVisualDepth/distance
Height in visual fieldMonocularYesVisualOrdinal
BlurMonocularCan beVisualMetric
AccommodationMonocularNoOculomotorMetric
Motion parallaxMonocularNo/static image can’t reproduce itVisualMetric
Binocular disparityBinocularNoVisualMetric
ConvergenceBinocularNoOculomotorMetric

Hi, I’m Daisy!

I created The Psych Diaries to make studying psychology a little less overwhelming. Here you’ll find study guides, Stata tutorials, psychology resources, and everything I’m learning along the way!

About Me β†’

Popular Guides


🀍 Thanks for stopping by. I hope you find something here that helps!

Discover more from The Psych Diaries

Subscribe now to keep reading and get access to the full archive.

Continue reading