Two annotators agree on most items. What can that agreement tell you about quality, and what can it miss?

Instruction: Calculate exact-match agreement on the four supplied items, then check both annotators against the written rule. Explain the limits of the result.

Context:

Separate agreement from correctness using a small labeled fixture. Identify a shared error and a disagreement without inferring wider performance.

Fictional practice task

The written rule labels a request as urgent only if it explicitly says "today"; all other requests are routine.

Item Text Annotator A Annotator B
1 Please finish today. urgent urgent
2 Please finish soon. urgent urgent
3 Complete it next week. routine routine
4 Can this be ready tomorrow? routine urgent

Use exact matching of the two annotators' labels to calculate raw agreement. Then compare the labels with the stated rule. Explain what agreement does and does not establish in this small example. Do not infer population performance from four items.

Updated

Prepare a stronger answer

They agree on items 1, 2 and 3, so their raw agreement is three out of four, or 75%. That tells me how often their labels match in this example. It doesn’t tell me that those matching labels are correct...

This member answer includes:

  • • A complete, copyable sample answer
  • • Guidance for adapting the answer to your experience
  • • Common mistakes and how to avoid them
  • • Answered interviewer follow-ups
  • • Strong, adequate and weak assessment criteria
Unlock the full answer and preparation guide

One payment for one year of full access. No automatic renewal.

See pricing and everything included

Your preparation path

Choose the track that matches the role. Work through its questions in order, then explain each answer in your own words.

Try the 20 minute mock assessment. Use the fictional cases to practice; the self-check is not an employer's hiring benchmark.

Related Questions