Similar prompts were compared when their answers appeared to differ.
The scan grouped prompts with nearly the same wording, compared the answers, and moved disagreements into review. It did not decide which answer was correct.
Near-duplicate questions were placed together for comparison.
Groups moved to review when the answers did not match.
All 19 groups were checked before any conclusion was recorded.
Five groups required a change. Fourteen did not.
The groups that needed no change mostly came from numbers mentioned during correct calculations, code tokens mistaken for short answers, or small prompt changes that made different answers appropriate.
Two wrong answers, two self-contradictory responses, and one ill-posed problem with conflicting responses.
The final answers agreed even though an intermediate number was extracted.
Different units, values, or wording made different answers valid.
Letters or numbers inside code were mistaken for short answers.
A polar-curve response reached the wrong final area.
The prompt asked for the area enclosed by r = 2sin(theta). The response dropped a factor during the calculation and concluded 2 pi. The correct area is pi.
One response confused the Earth-to-Mars distance with the Sun-to-Mars distance.
The question asked how long sunlight takes to reach Mars. The response used about 34.8 million miles as the Sun-to-Mars distance. That number describes a close Earth-to-Mars distance, so the answer of roughly three minutes was built on the wrong measurement.
Different answers to similar prompts do not always mean the data is wrong.
The same comparison was applied to all 9,500 source rows in the human-written No Robots dataset. It grouped 20 possible answer disagreements. Complete review found one group that required a change and 19 groups that needed no change.
A bounded review result with clear limits on the claim.
Datascreen reduced 25,000 examples to 19 groups that needed comparison. Reviewing every group found five wrong or internally inconsistent answers that required a change.
The result does not show malicious poisoning, measure downstream model harm, prove the rest of the dataset is clean, or show that every data problem was found.
25,000-row Airoboros source file used for this scan
SHA-256 · 0c53a38296073f3d423ee9358881b55ba70c24c0dd9bb3ebe18947c5a657b428