Datascreen
Public dataset case study / August 2026

Five training-data problems were found by comparing similar prompts with different answers.

Datascreen scanned 25,000 Airoboros examples and grouped similar prompts when their answers appeared to disagree. We reviewed every group. Five required a change because the answers were wrong or contradicted themselves. Fourteen needed no change.

EXAMPLES CHECKED 25,000 Every row in the recorded source file
SIMILAR-PROMPT GROUPS REVIEWED 19 Every group was reviewed
REQUIRED A CHANGE 5 Wrong or contradictory answers
NEEDED NO CHANGE 14 No dataset change needed
01 / HOW THE CHECK WORKED

Similar prompts were compared when their answers appeared to differ.

The scan grouped prompts with nearly the same wording, compared the answers, and moved disagreements into review. It did not decide which answer was correct.

STEP 1 Group similar prompts

Near-duplicate questions were placed together for comparison.

STEP 2 Flag different answers

Groups moved to review when the answers did not match.

STEP 3 Review the full responses

All 19 groups were checked before any conclusion was recorded.

02 / THE REVIEW RESULT

Five groups required a change. Fourteen did not.

The groups that needed no change mostly came from numbers mentioned during correct calculations, code tokens mistaken for short answers, or small prompt changes that made different answers appropriate.

REQUIRED A CHANGE 5 groups

Two wrong answers, two self-contradictory responses, and one ill-posed problem with conflicting responses.

CORRECT RESPONSES 6 groups

The final answers agreed even though an intermediate number was extracted.

DIFFERENT QUESTIONS 6 groups

Different units, values, or wording made different answers valid.

CODE READ AS AN ANSWER 2 groups

Letters or numbers inside code were mistaken for short answers.

03 / EXAMPLE: WRONG CALCULATION

A polar-curve response reached the wrong final area.

The prompt asked for the area enclosed by r = 2sin(theta). The response dropped a factor during the calculation and concluded 2 pi. The correct area is pi.

PROMPT Find the area enclosed by the polar curve r = 2sin(theta).
RESPONSE CONCLUDED2 pi
CORRECT RESULTpi
REVIEW ACTIONRegenerate
04 / EXAMPLE: WRONG FACT

One response confused the Earth-to-Mars distance with the Sun-to-Mars distance.

The question asked how long sunlight takes to reach Mars. The response used about 34.8 million miles as the Sun-to-Mars distance. That number describes a close Earth-to-Mars distance, so the answer of roughly three minutes was built on the wrong measurement.

REVIEW DECISION
ProblemWrong distance
Produced answerAbout 3 minutes
ActionRemove or regenerate
This is evidence of an incorrect answer, not evidence of malicious data poisoning.
05 / WHY REVIEW MATTERS

Different answers to similar prompts do not always mean the data is wrong.

The same comparison was applied to all 9,500 source rows in the human-written No Robots dataset. It grouped 20 possible answer disagreements. Complete review found one group that required a change and 19 groups that needed no change.

NO ROBOTS REVIEW
Source rows9,500
Similar-prompt groups reviewed20
Required a change1
Needed no change19
The comparison identifies where to look. Review determines whether the data should change.
06 / WHAT THIS ESTABLISHES

A bounded review result with clear limits on the claim.

OBSERVED

Datascreen reduced 25,000 examples to 19 groups that needed comparison. Reviewing every group found five wrong or internally inconsistent answers that required a change.

NOT ESTABLISHED

The result does not show malicious poisoning, measure downstream model harm, prove the rest of the dataset is clean, or show that every data problem was found.

Exact source record

25,000-row Airoboros source file used for this scan

SHA-256 · 0c53a38296073f3d423ee9358881b55ba70c24c0dd9bb3ebe18947c5a657b428

Scan date · 23 August 2026
Inspect your own dataset

Turn scattered answer disagreements into a reviewable set of source rows.

Datascreen groups related examples, preserves the full responses, and lets reviewers decide which rows require correction or removal.