The result was an inspectable review record, not a spreadsheet of unexplained counts.
Datascreen opened the affected source rows, showed the evidence behind each finding, suggested an action, and recorded the reviewer’s final decision. The completed record retained all 9,000 decisions with no findings left open.
Private keys and access tokens were present in data already available for model use.
Datascreen found private-key and access-token patterns inside the public corpus. Reviewers could inspect the surrounding row while the sensitive value remained redacted, then remove the row from the next dataset version.
The scan confirms credential material in the corpus. It does not establish whether any value was still active.
Expected-answer text remained inside conversations that could become training or evaluation data.
One source row included an explicit correct-answer line inside the conversation. That text can expose the expected output instead of leaving the model to learn from the question and response normally.
7,343 later copies were traced back to their earlier source rows.
Repeated examples can give the same content more influence than intended. Datascreen linked every later copy to its earlier source row, then checked all 7,343 relationships against the source data.
Evidence for review, with clear limits on the claim.
The public dataset contained exposed credential patterns, expected-answer text, and repeated conversations. Datascreen connected those findings to the affected source rows and review decisions.
This scan does not prove malicious poisoning, downstream model harm, regulatory non-compliance, or that every possible data problem was found.
ShareGPT_2023.05.04v0_Wasteland_Edition.json
SHA-256 · 701b6b1beea876183b768b2ac530ba0b7f32a5e3755885c4e36b92302aa4c6bd