Two annotators agree on most items. What can that agreement tell you about quality, and what can it miss?
Instruction: Calculate exact-match agreement on the four supplied items, then check both annotators against the written rule. Explain the limits of the result.
Separate agreement from correctness using a small labeled fixture. Identify a shared error and a disagreement without inferring wider performance.
Fictional practice task
The written rule labels a request as urgent only if it explicitly says "today"; all other requests are routine.
| Item | Text | Annotator A | Annotator B |
|---|---|---|---|
| 1 | Please finish today. | urgent | urgent |
| 2 | Please finish soon. | urgent | urgent |
| 3 | Complete it next week. | routine | routine |
| 4 | Can this be ready tomorrow? | routine | urgent |
Use exact matching of the two annotators' labels to calculate raw agreement. Then compare the labels with the stated rule. Explain what agreement does and does not establish in this small example. Do not infer population performance from four items.
Updated
Prepare a stronger answer
They agree on items 1, 2 and 3, so their raw agreement is three out of four, or 75%. That tells me how often their labels match in this example. It doesn’t tell me that those matching labels are correct...
This member answer includes:
- • A complete, copyable sample answer
- • Guidance for adapting the answer to your experience
- • Common mistakes and how to avoid them
- • Answered interviewer follow-ups
- • Strong, adequate and weak assessment criteria
One payment for one year of full access. No automatic renewal.
See pricing and everything includedYour preparation path
Choose the track that matches the role. Work through its questions in order, then explain each answer in your own words.
1. Entry level annotation
Apply guidelines, label text and spans, and explain a small practice project.
- How would you explain the data annotator role and the kind of work you would expect to do? Free sample
- How would you learn a new annotation guideline before starting your first batch? Member answer
- Apply a sentiment guideline to four short comments. Which labels would you choose, and why? Free sample
- Mark two location mentions using the exact character-offset contract. How would you check your result? Free sample
- Walk me through an annotation or quality-checking project you can discuss, including your own contribution and limits. Member answer
2. AI response evaluation
Compare responses using separate criteria for correctness, instruction following, and writing quality.
- Compare two AI responses against a supplied fact sheet. Which response is better under the rubric? Free sample
- One response is accurate but breaks the required format; another follows the format but contains a false claim. How would you rate them? Member answer
- Write a short rating rationale that identifies the decisive error without restating both responses. Member answer
- An AI response includes a factual claim you cannot verify from the supplied sources. What would you do? Member answer
- Two responses have different strengths and neither clearly wins. How would you apply the ranking rules? Member answer
3. Senior review and quality
Work through disagreement, missed critical cases, changing guidelines, review capacity, and reviewer calibration.
- Two experienced reviewers disagree repeatedly, and the deadline leaves little time for adjudication. What would you recommend? Free sample
- A batch has 98% accuracy against reviewed references but misses every critical item. Would you accept it? Member answer
- A labeling rule changes halfway through a delivery. Would you relabel old work, split the dataset or delay the release? Member answer
- The delivery requires review of every item, but the available reviewers cannot finish by the deadline. What would you change? Member answer
- A reference answer appears to contradict the written rule, and workers are being penalized for disagreeing with it. What would you do? Member answer
Try the 20 minute mock assessment. Use the fictional cases to practice; the self-check is not an employer's hiring benchmark.
Related Questions
-
easy
-
easy
-
easy
-
easy
-
medium
-
medium