Cluster D

CELPIP Speaking Scoring Criteria: The 4 Things Raters Score

Most candidates prepare CELPIP Speaking against a rubric they've never read. Ask them what gets scored and you'll hear "fluency, grammar, vocabulary, accent" — one criterion right, one half right, and one that isn't on the paper at all. Then they wonder why the score doesn't match the feeling.

There are four dimensions: content/coherence, vocabulary, listenability, and task fulfillment. This post takes each one apart — what a 6, an 8, and a 10 sound like in each, how the rubric connects to the 1–12 scale, and how to score yourself against it honestly. If you want your self-scores checked by someone trained on the real thing, start with a free mock test → celpipuni.com/free-mock-test.

TL;DR

CELPIP Speaking is scored on four dimensions: content/coherence, vocabulary, listenability, and task fulfillment — roughly equally weighted, on a 1–12 scale. A 6 is competent but thin; an 8 is effective and developed; a 10 is precise and effortless. Accent isn't a criterion — intelligibility is. You can't game a rubric you haven't read, and most candidates haven't read it.

What the rubric actually is

The rater listening to your eight recordings isn't grading impressions. They're holding a rubric — level descriptors for four dimensions, each describing what a speaker at that level can do with English in the real world. The descriptors aren't about test tricks. They describe function: how well you develop an idea, how precisely you choose words, how easily a listener follows you, how completely you did the job the task set out.

Two structural facts matter for how you prepare.

First, the four dimensions are roughly equally weighted. That kills the most common misallocation in speaking prep — drilling grammar for a month because it feels like "the English part," when three of the four dimensions are about content, coherence, and listenability. Grammar lives inside vocabulary and listenability as accuracy; it doesn't get its own line.

Second, the 1–12 scale describes levels of function, not percentages of correctness. A 12 isn't "zero errors" — it's "operating comfortably in demanding situations." Your score is where your overall performance sits against those descriptions, which is why one bad task doesn't sink you and one lucky one doesn't save you.

How a score actually gets built

Here's the process from your eight recordings to the number on your report.

  1. Each task gets scored against the four dimensions. The rater listens for whether you addressed the prompt, developed it with reasons and detail, chose words precisely, and stayed easy to follow.
  2. The dimensions interact. A speaker who fully addresses the task with thin vocabulary can't score like one who does both — the rubric expects the dimensions to line up at each level.
  3. Your eight tasks together set the level. One weak task among eight is a wobble, not a verdict. A pattern across all eight is your score.
  4. The level lands on the 1–12 scale and maps 1:1 to CLB from level 4 up — a CELPIP 8 in Speaking is CLB 8, and each level step is real CRS territory.

(Score descriptors get revised periodically — confirm current details on the official CELPIP site.)

The four criteria at a glance

CriterionThe question the rater is askingWhat pushes it up
Content/coherenceDid you develop real ideas, and does the answer hold a logical thread?Reasons, consequences, one idea developed instead of four listed
VocabularyAre you choosing the right words, not just many words?Precision and natural collocation — "wall-to-wall meetings" over "very busy"
ListenabilityHow easily can a listener follow you?Rhythm, pausing, stress on the words that matter
Task fulfillmentDid you do the specific job — advise, describe, compare, persuade?Answering the actual prompt in full, not the one you prepared for

Preparing for the wrong rubric

The mistakes below all come from one root: practicing against an imagined rubric.

Drilling grammar as if it's a fifth criterion. Accuracy counts inside vocabulary and listenability, but nobody is tallying your errors. A grammatically flawless answer with no development scores below a slightly imperfect one that goes somewhere.

Chasing rare vocabulary. The vocabulary dimension rewards the right word, not the fanciest one. "The queue exhibited considerable length" is worse than "the line was dragging on" — the second is precise and alive, and precision is the descriptor.

Doing accent work when the problem is coherence. Accent isn't scored. Intelligibility is. If the rater can follow you — and if you're scoring 7s, they can — your pronunciation isn't the ceiling. Your development and rhythm are.

Answering a prepared answer, not the prompt. Task fulfillment is the quiet killer. The candidate who bends their memorized market description onto a kitchen photo loses points not for vocabulary but for missing the task.

Treating all four dimensions as one thing called "English." "My speaking is weak" is not a diagnosis. Four numbers is. One of them is your bottleneck, and it's almost never the one you'd guess.

What a 6, an 8, and a 10 sound like

Illustrative descriptors, written for this post to show the shape of each level. Same advice prompt — "a friend is overwhelmed at work" — three levels:

CriterionA 6 sounds likeAn 8 sounds likeA 10 sounds like
Content/coherence"She should talk to her manager. Also make a list. Lists help with stress." — points listed, nothing developed"I'd tell her to pick the one task that actually has a deadline this week, and say no to the rest — because right now she's treating everything as urgent, and it isn't." — one idea, developed with a reason"I'd tell her to pick the one task with a real deadline and say no to the rest — and to say it exactly that way to her manager, too, because the overload probably started when she said yes to everything and meant none of it." — developed, connected, building
Vocabulary"She has a lot of stress and too much work to do." — adequate but generic"Her calendar is wall-to-wall meetings; she's underwater." — precise, natural"Her calendar is wall-to-wall meetings, she's underwater, and the worst part is half of it is self-inflicted — nobody asked her to say yes." — precise, layered, effortless
ListenabilityEven pace, even melody, a "umm" before each new pointNatural pausing; stress lands on "one task" and "no"Deliberate pacing; slows before the point, lands it, moves on
Task fulfillmentAdvice given, but generic — could apply to anyoneAdvice given, specific to this friend's situationAdvice, tailored, with what to say and why it'll land

Read the 6 column honestly. It's not bad English. It's the answer most unprepared candidates give — and most stuck-at-7 answers sit between the 6 and 8 columns. The 10 column has no rare words. The difference is development, precision, and rhythm.

Your rubric checklist for this week

  • Score your last three recordings on all 4 criteria, 1–12, in writing
  • Find your lowest number — that's your bottleneck criterion, not "English"
  • Redo one task fixing only that criterion
  • Check each answer for one interpretation moment ("which tells me…")
  • Check each answer against the prompt, not your prepared version
  • Compare two self-scored recordings with a trained grader's scoring once

Scoring yourself against the real rubric is learnable — but one professional calibration keeps your self-scores honest for months. Get a speaking evaluation with band-level feedback → celpipuni.com/evaluation.

Most candidates prepare against an imagined rubric

Here's the deeper point. Go find ten CELPIP Speaking posts at random. Most teach task structures and sample answers without ever stating the four dimensions. So candidates absorb a folk rubric — "speak longer, use big words, don't make grammar mistakes, lose the accent" — and train against that.

The folk rubric isn't just incomplete; parts of it are anti-rubric. Longer isn't better when the extra twenty seconds are empty. Big words actively hurt the vocabulary dimension when they replace precise ones. And the accent point sends uncountable study hours toward the one thing the rubric doesn't measure, while coherence — the dimension most 7s are actually failing — goes untouched.

The test is beatable precisely because the rubric is public. Read it, score against it, and you're preparing for the actual test. That's also why self-scoring works better than people expect: the rubric's questions are concrete enough to ask of your own recording. Where's your interpretation? Where's your precise word? Where does your rhythm move? The rubric tells you what to listen for. That's the entire game.

Tools and resources

Frequently asked questions

What are the CELPIP speaking scoring criteria?

Four dimensions: content/coherence, vocabulary, listenability, and task fulfillment. They're roughly equally weighted, and each is scored against level descriptors on the 1–12 scale.

Is accent scored in CELPIP speaking?

No. Accent is not a criterion on the rubric. Intelligibility is — the rater must follow you. Plenty of strong accents score 9+ because clarity of ideas and rhythm are what the descriptors measure.

Is grammar scored separately in CELPIP speaking?

No. Grammar doesn't get its own dimension. It counts inside vocabulary and listenability as accuracy, which is why drilling grammar in isolation moves a speaking score less than content and rhythm work.

How does the rubric connect to the 1–12 scale?

The scale describes levels of function. Your performance across the eight tasks is matched to level descriptors for each dimension, and the overall level lands somewhere on 1–12 — equal to your CLB level from 4 up.

Can I score a 10 with a simple vocabulary?

Yes, mostly. The vocabulary dimension rewards precision and naturalness over rarity. "The line was dragging on" outperforms "the queue exhibited considerable length." A 10's words are right, not fancy.

How do I score my own speaking against the rubric?

Record a task, play it back, and score each dimension 1–12 in writing using the official descriptors. Then check your scoring once against a trained grader — self-scoring works when it's calibrated.

What to do next

  1. Today: read the official score descriptors once, slowly, with a recording of yours playing.
  2. Tonight: score your last recording on all four dimensions — four numbers, in writing.
  3. This week: redo three tasks, each fixing only your lowest criterion.
  4. This month: one evaluation to confirm your self-scoring matches a trained grader's.
  5. Ongoing: log four numbers per session — watch which line drags and which climbs.

The bottom line

The rubric is public, four dimensions, roughly equal weight — and mostly ignored by the people preparing for the test. Read it, score yourself against it in writing, and aim every practice session at one criterion with a number on it. You can't game a rubric you haven't read. Once you've read it, you don't need to game it.

Find out which of the four criteria is actually costing you points — get band-level feedback from a trained grader → Get my speaking evaluated.