Cluster D

CELPIP Speaking Task 3: Describe a Scene Like a 9

A photo appears on your screen. A kitchen, a park, a street, a living room. You get about 30 seconds to look at it and roughly 60 seconds to describe what's happening. Most test takers do the same thing: they inventory. "There is a man. He is cooking. There is a boy. The boy is eating." Every sentence true, every sentence flat, and the score lands at 7 or below. The people who score 9 do something the photo actually rewards — they interpret. This post shows you the difference in real sentences, gives you the speculation language that makes interpretation easy, and walks through a full band-9 kitchen-scene response, annotated.

Curious what your own scene description scores? Record one against any photo on your phone, then take a free mock test to see → celpipuni.com/free-mock-test.

TL;DR

CELPIP Speaking Task 3 shows you a photo and asks you to describe the scene — about 30 seconds of prep, roughly 60 seconds of speaking. Band 7 describes what's visible. Band 9 describes what's visible plus what it means: who these people probably are, what's likely going on, what the mood is. Structure: broad setting → foreground details with interpretation → background and mood. The key fact people miss: there's no answer key — you cannot lose points for guessing wrong about the photo.

What Task 3 actually wants

The task sounds like a description exercise. It isn't. It's a thinking-out-loud exercise with a picture as the prompt. Compare two sentences about the same kitchen photo:

  • Inventory: "There is a woman. She is cooking. There is a pot on the stove."
  • Interpretation: "She's mid-stir but holding the spoon kind of away from her body, which is what people do right after tasting something — I'd guess whatever's in that pot isn't quite right yet."

Both are accurate. Only one of them shows a mind at work. The rater's four criteria — content/coherence, vocabulary, listenability, task fulfillment — all reward the second sentence, because the second sentence demonstrates development, precise vocabulary ("mid-stir"), rhythm, and an actual response to the task's real demand: not list the objects but make me see the scene.

Interpretation sounds risky to a lot of test takers — "what if I guess wrong about the photo?" Here's the thing: there's no answer key. The boy reaching into the jar is not secretly documented as "eating cookies." The photo cannot tell on you. "I'd guess he does this often" is not a wrong answer, because the photo doesn't have a right one. The only risk in Task 3 is playing it safe.

How to run the 60 seconds

Structure: camera moves from wide to close, then back out.

Beat 1 — the broad setting (10 seconds). Where are we, what time of day, what's the overall event. "This looks like a weekday-evening kitchen, somewhere around dinner time — the light's gone that warm gold it gets before six." One or two sentences. The setting sentence earns everything after it, because a scene described in context is a scene the rater can picture.

Beat 2 — foreground details, each with interpretation (30–35 seconds). Pick two or three things near the front of the photo. For each: what it looks like, plus what it probably means. "The kid's got one hand in a jar on the counter and his eyes locked on his mother's back — frozen mid-reach, which tells me he knows exactly what the rules are."

Beat 3 — background and mood (10–15 seconds). What the edges of the photo say about the atmosphere. "The table's half-set, a school bag's dumped on a chair — nobody's rushed, nobody's fighting. This is the calm part of the evening."

That's three beats, and the middle beat is where the band lives. Two details interpreted beat six details listed — same reason two pieces of advice beat five in Task 1: the rubric scores depth, not coverage.

Prep (the ~30 seconds): don't memorize objects. Pick your two foreground details and decide your setting sentence. The guesswork — who, why, what's about to happen — should happen live, while you're speaking. Rehearsed interpretations sound rehearsed; improvised ones sound like fluency, which is the entire point of the task.

One test-room detail: the photo stays on screen while you speak, so you can glance back — but test-takers who keep looking at the photo keep finding new objects and run out of time to interpret anything. Decide your beats during prep, then trust them. (Formats get revised periodically — confirm current timings on the official CELPIP site.)

The numbers: Task 3 at a glance

Task 3 (describe a scene)
Prep time~30 seconds (photo on screen)
Speaking time~60 seconds
RegisterCasual, observational — like narrating a photo to a friend
StructureSetting → foreground + interpretation → background/mood
Core tensePresent continuous ("she's stirring," "he's reaching")
Common capInventory mode; "I can see" on repeat; silence at 40 seconds

Why silence at 40 seconds is the inventory-mode giveaway: a photo contains maybe eight describable facts. List them and you're done in 40. Interpret them and each fact becomes three sentences — what it looks like, what it means, what's probably about to happen. Sixty seconds fills itself.

The mistakes that cap Task 3 at 7

Inventory mode. "There is a man. There is a table. There are two chairs." No meaning, no rhythm, no person behind the voice. This is the most common Task 3 failure, and it feels safe, which is why it survives.

"I can see" twelve times. Every sentence starts with the same three words, so every sentence sounds the same. Swap in "there's," "front and left," or just start with the object: "A kid's frozen mid-reach, hand in a jar…"

Guessing nothing. Some test takers, terrified of being "wrong," describe only what's provably visible. But a description with zero speculation has no development, and content/coherence pays for development.

Describing the edges first. Starting with "in the background there are some buildings" buries your best material. The foreground people are where the story is — background is for the mood beat at the end.

What good looks like: a band-9 sample, annotated

The photo: a home kitchen. A woman cooks at the stove; a boy of about eight stands at the counter; the table is partially set.

"This looks like a weekday-evening kitchen, somewhere around six — you can tell from the light, which has that warm gold it gets just before dinner. The whole scene is organized around one person: a woman at the stove, mid-stir. She's holding the spoon slightly away from her body, which is what people do right after tasting something — I'd guess whatever's in that pot isn't quite right yet. And there's steam, so it's been cooking a while. This isn't a five-minute meal.

Front and left, a kid — eight, maybe — is doing something he believes is unobserved. One hand deep in a jar on the counter, eyes locked on his mother's back, frozen mid-reach. I'd give him three seconds before she turns around. And he's clearly done this before — there's a chair pulled up to the counter, and that's not a one-time arrangement.

The background tells you the mood: the table half-set, a school bag dumped on a chair, a phone face-down by the sink. Nobody's rushed, nobody's fighting. This is the calm part of the evening — everyone in the same room, doing their own thing."

Why this scores 9:

  • Task fulfillment: a full description — setting, people, action, atmosphere — delivered like a scene, not a list.
  • Content/coherence: every detail gets interpreted ("which is what people do right after tasting"), so ideas develop instead of accumulating. The camera moves wide → close → wide, and the listener can follow.
  • Vocabulary: "organized around," "frozen mid-reach," "a one-time arrangement," "the calm part of the evening." Ordinary words, placed with intent. Zero exotic vocabulary.
  • Listenability: varied sentence lengths — a long observational sentence, then a three-word verdict ("This isn't a five-minute meal"). Pauses implied, rhythm alive.
  • The interpretive voice: "I'd guess," "I'd give him three seconds," "clearly done this before." The speaker sounds like they're genuinely figuring the photo out in real time — which is precisely what a 9 sounds like.

Your Task 3 checklist

  • Setting sentence first: where, when, what's happening overall
  • Two foreground details, each with an interpretation ("which tells me…")
  • At least three speculation phrases used naturally
  • Zero "I can see" repeats — vary or drop the frame
  • Background saved for the mood beat, last 10–15 seconds
  • Practice on three real photos this week, timed, out loud

Want a grader to tell you whether your description reads as inventory or interpretation? Get an evaluation with band-level feedback → celpipuni.com/evaluation.

The deeper dive: the speculation toolkit

The difference between wanting to interpret and being able to interpret, under a clock, is having the phrases ready. Here's the kit — grouped by confidence:

Safe deductions (you're basically reading the photo):

  • "This looks like…" / "…which tells me…" / "you can tell from the light that…"
  • "She's clearly been cooking a while" (clearly + present perfect = evidence-based)

Mid-confidence guessing:

  • "probably" / "I'd guess…" / "there's a good chance he…"

Full speculation (and this is where 9s live):

  • "My money's on her noticing in the next few seconds"
  • "I wouldn't be surprised if this is a regular routine"
  • "If I had to bet, that pot's been on the stove for an hour"

Notice what all of them have in common: they attach to a reason. "I'd guess she just tasted it — she's holding the spoon away from her body." Speculation with visible evidence reads as intelligence. Speculation without evidence reads as filler.

And one habit worth stealing from news photographers: describe what's about to happen. The frozen kid, the mid-stir spoon, the half-set table — photos capture moments mid-motion, and a 9 speaker narrates the motion. "I'd give him three seconds" is six words doing the work of twenty.

Tools and resources

  • The official CELPIP site — free sample prompts and the current task format; practice against real prompt phrasing. (Formats get revised; confirm there.)
  • CELPIP Speaking: the complete guide — all 8 tasks and the four criteria explained: celpipuni.com/celpip-speaking-guide
  • CELPIP Speaking samples with scores — band 7 vs 9 comparisons across three task types: celpipuni.com/celpip-speaking-samples

Frequently asked questions

How long is CELPIP Speaking Task 3?

About 30 seconds of prep with the photo on screen, and roughly 60 seconds of speaking. Confirm current timings on the official CELPIP site — formats get revised periodically.

Can I be wrong about what's happening in the photo?

No — there's no answer key. The rater scores your language, not your detective work. A reasonable, well-argued guess ("I'd guess he's sneaking a snack") scores better than a fact stated with no development.

What tense should I use?

Present continuous mostly ("she's stirring," "he's reaching"), present simple for states ("the table looks half-set"), and "must be" or "looks like" for deductions. A past-tense story about the photo is a mismatch — describe what's in front of you.

What kinds of photos come up?

Ordinary daily life: kitchens, markets, streets, parks, offices, classrooms, parties. Never technical, never current events. Your own photo library is a legitimate practice set.

What if I finish early?

You left the mood beat out. Return to the background: the light, the mess level, the overall atmosphere. That beat exists partly so you never hit dead air at 45 seconds.

What to do next

  1. Today: pick three photos from your phone and describe one, timed, 60 seconds, out loud. Record it.
  2. Listen back and count your speculation phrases — fewer than three means inventory mode.
  3. Re-record the same photo using the three-beat structure: setting, two interpreted details, mood.
  4. This week: do one photo a day, and get one evaluation to check whether your interpretations read as band 9.

The bottom line

Task 3 doesn't score your eyes; it scores your inferences. Anyone can see the woman and the pot. The 9 sees the spoon held away from her body and says what it means, then says what's probably about to happen — and does it all in a rhythm a listener can settle into. Inventory gets you 7. Interpretation, in natural spoken English, gets you the rest of the way.

Find out what your scene descriptions actually score — get a trained grader's band-level feedback → Get your speaking evaluated.