+1 (415) 555-0142

The board

Assessment research

The evidence behind the test: what we measure, why we measure it that way, and what we are still working on.

Validity

A test is valid when it measures what it claims to measure. CELP tasks are built from situations candidates actually meet: a notice with opening hours, a platform-change announcement, an email to a landlord, a question about how you prefer to work.

We avoid tasks that reward test technique over English. A candidate who can read a timetable and write a clear complaint has demonstrated something real about their English; a candidate who has memorised a rubric has not.

Reliability

Reading and Listening are marked against a fixed key, so they are perfectly repeatable. The risk sits in Writing and Speaking, where two examiners could reasonably differ.

We manage that by sampling: a proportion of marked responses is second-marked blind, and disagreement beyond half a band is investigated rather than averaged away.

Fairness

Topics are chosen to be answerable by anyone, regardless of country, profession or background knowledge. A question that rewards familiarity with one country's institutions is testing general knowledge, not English.

Invigilation is applied identically to every candidate and plays no part in marking. It does not identify individuals, draw demographic inferences or score behaviour, and an examiner never sees a candidate before assigning a band.

Open questions

  • How short can a speaking sample be before a band becomes unreliable?
  • Does automatic marking of Writing reach examiner agreement on short functional tasks?
  • What is the effect of a candidate's own device and connection on measured performance?

Last updated 23 August 2026