CELPIP Speaking Task 3 is Describing a Scene: 30 seconds of preparation and 60 seconds of recording. You see an illustration and describe what is happening.
The official score report does not produce a Task 3 score on its own — this recording sits alongside the other tasks as evidence for your Speaking Overall performance.
Content / Coherence
This dimension asks whether your description has an overall picture, clear focus points, and sensible organisation. CELPIP official advice is to summarise the scene in one sentence and then develop a selection of details — you do not need to describe everything in the image.
You can organise left to right, foreground to background, or group by group. Quickly listing woman, child, dog, and car is far weaker than explaining where people are, what they are doing, and how they relate to one another.
Specific detail is the evidence
Beyond actions, describe appearance, clothing, and probable feelings: Near the entrance, a man in a blue jacket is carrying several boxes, and he appears to be in a hurry.
Words like seems, appears, and might separate observation from inference, so you avoid stating something you cannot know as fact.
Vocabulary
Vocabulary here is mainly about accurate, natural language for position, action, appearance, and setting: next to, in the foreground, behind the counter, is reaching for, appears frustrated.
A high level does not require you to know the technical name of every object. When a word will not come, describe shape, purpose, and location — continuing to communicate beats silence.
Listenability
Listenability is about whether a listener can rebuild the image in their head. Clear pronunciation, a steady pace, and natural phrase boundaries all help.
Restarting with There is for every new object makes your rhythm monotonous and caps your sentence range. Use while, who, which, and prepositional phrases to connect people with actions — while keeping accuracy first.
Task Fulfillment
The task is to describe the current scene — not to tell an unrelated story, and not to move early into predictions. Your answer should be relevant, complete, delivered as if explaining the scene to someone who cannot see it, and long enough to provide sufficient evidence.
There is no rule about how many people you must describe, and no official algorithm awarding a point per object named.
How to practise
In your 30 seconds, choose one overview sentence and three or four focus points, then record for 60 seconds.
When you replay, ask: could someone without the picture tell what kind of place this is, where the main people are, and what is happening? If not, add relationships and positions — not more objects. That check matches the four dimensions much better than counting nouns.
FAQs
Must I describe everything? No. Overview first, then selected details.
Why is listing objects weak? It gives no relationships and builds no picture.
What if I cannot name something? Describe its shape, purpose, and position and keep going.
Next, read how Task 4 is scored and how to answer Tasks 1-4.
