Compare two AI responses against a supplied fact sheet. Which response is better under the rubric?
Instruction: Use only the fictional fact sheet and stated preference rubric. Choose a response and identify the decisive supported or unsupported claims.
Tests whether factual grounding determines a pairwise preference when a fluent response invents material visiting information.
Fictional practice task
Approved fact sheet: The Harrow Museum opens Tuesday through Sunday, 10:00–17:00. It is closed on Mondays. Standard adult admission is $12. No evening schedule is provided.
User request: "Can I visit on Monday evening, and how much is adult admission?"
Response A: "It is closed on Mondays, so Monday evening will not work. Standard adult admission is $12. Its listed visiting hours are Tuesday through Sunday, 10:00–17:00."
Response B: "Absolutely! Visit Monday at 18:00 for a quiet evening. Adult admission is $12, and you can stay until 20:00."
Rubric: Prefer the response that answers the user's question using only supported facts. A wrong opening day or invented opening time is a material error. Tone and fluency cannot compensate for such an error.
Updated
Example Answer
I'd prefer A. It correctly says the museum is closed on Mondays and gives the supported adult admission price of $12. Its listed visiting hours also match the fact sheet, so it answers the request without inventing an evening schedule.
B gets the price right, but it invites the user on Monday and introduces 18:00 and 20:00 times that the source does not support. The Monday claim directly contradicts the closure rule. Under this rubric, that is a material error; friendly wording and one correct detail cannot compensate for it. I'd base my preference on the supplied evidence, without searching for a different schedule for this fictional museum.
Make it your own
Keep the fact sheet authoritative for this exercise. In another task, use its permitted sources and preference rule rather than assuming that tone or length controls the ranking.
Why this works
The answer chooses A and pinpoints the wrong day and invented times, rather than giving both responses an impressionistic fluency score.
Interviewer follow-up
Could B still win because it answers more confidently and sounds more welcoming?
No, not under the supplied rubric. Its confidence makes unsupported visiting advice sound usable, but the advice is materially wrong. A directly answers that Monday evening will not work and supplies the requested price. I'd prefer accurate, supported guidance here rather than reward a friendly invitation that would send the user to a closed museum.
Assessment criteria
Strong: Prefers A and identifies the Monday contradiction and unsupported evening hours, using the rubric's material-error rule.
Adequate: Prefers A because it matches the fact sheet, but does not specify which B claims determine the preference.
Weak: Prefers B for confidence or tone, treats the matching price as canceling the wrong schedule, or searches outside the closed-reference exercise.
A tempting weak answer
“B is better because it says yes, gives a price and provides more detailed visiting advice.”
Why it fails: Its extra schedule details are unsupported and its Monday invitation is false under the fact sheet. More detail does not satisfy the grounding requirement.
References
Your preparation path
Choose the track that matches the role. Work through its questions in order, then explain each answer in your own words.
1. Entry level annotation
Apply guidelines, label text and spans, and explain a small practice project.
- How would you explain the data annotator role and the kind of work you would expect to do? Free sample
- How would you learn a new annotation guideline before starting your first batch? Member answer
- Apply a sentiment guideline to four short comments. Which labels would you choose, and why? Free sample
- Mark two location mentions using the exact character-offset contract. How would you check your result? Free sample
- Walk me through an annotation or quality-checking project you can discuss, including your own contribution and limits. Member answer
2. AI response evaluation
Compare responses using separate criteria for correctness, instruction following, and writing quality.
- Compare two AI responses against a supplied fact sheet. Which response is better under the rubric? Free sample
- One response is accurate but breaks the required format; another follows the format but contains a false claim. How would you rate them? Member answer
- Write a short rating rationale that identifies the decisive error without restating both responses. Member answer
- An AI response includes a factual claim you cannot verify from the supplied sources. What would you do? Member answer
- Two responses have different strengths and neither clearly wins. How would you apply the ranking rules? Member answer
3. Senior review and quality
Work through disagreement, missed critical cases, changing guidelines, review capacity, and reviewer calibration.
- Two experienced reviewers disagree repeatedly, and the deadline leaves little time for adjudication. What would you recommend? Free sample
- A batch has 98% accuracy against reviewed references but misses every critical item. Would you accept it? Member answer
- A labeling rule changes halfway through a delivery. Would you relabel old work, split the dataset or delay the release? Member answer
- The delivery requires review of every item, but the available reviewers cannot finish by the deadline. What would you change? Member answer
- A reference answer appears to contradict the written rule, and workers are being penalized for disagreeing with it. What would you do? Member answer
Try the 20 minute mock assessment. Use the fictional cases to practice; the self-check is not an employer's hiring benchmark.
Related Questions
-
easy
-
easy
-
easy
-
easy
-
medium
-
medium