Datascreen
Public dataset case study / August 2026

A public ShareGPT dataset contained private keys, access tokens, answer keys, and thousands of repeated conversations.

Datascreen scanned 52,052 conversations through the hosted product, linked each finding to its source row, and preserved the decisions made during review.

CONVERSATIONS SCANNED 52,052 Every parsed source row
FINDINGS RECORDED 9,000 Each linked to source data
CREDENTIAL FINDINGS 24 Private keys, tokens, and similar secrets requiring review
ANSWER-KEY FINDINGS 66 Expected-answer text inside the data
01 / FROM SCAN TO DECISION

The result was an inspectable review record, not a spreadsheet of unexplained counts.

Datascreen opened the affected source rows, showed the evidence behind each finding, suggested an action, and recorded the reviewer’s final decision. The completed record retained all 9,000 decisions with no findings left open.

APPLICATION EVIDENCE COMPLETED REVIEW RECORD
Datascreen review record showing 52,052 rows, 9,000 findings, 9,000 reviewed, and zero open
The hosted review record shows 52,052 source rows, 9,000 findings, 9,000 recorded decisions, and zero findings left open.
02 / EXPOSED CREDENTIALS

Private keys and access tokens were present in data already available for model use.

Datascreen found private-key and access-token patterns inside the public corpus. Reviewers could inspect the surrounding row while the sensitive value remained redacted, then remove the row from the next dataset version.

Rows flagged24
Unique credential values found22
Raw values in this case studyRedacted

The scan confirms credential material in the corpus. It does not establish whether any value was still active.

APPLICATION EVIDENCE ROW 46,071
Datascreen credential exposure review modal with private-key material redacted
A real finding from the hosted application with the credential redacted and the row marked for removal.
03 / ANSWER-KEY RESIDUE

Expected-answer text remained inside conversations that could become training or evaluation data.

One source row included an explicit correct-answer line inside the conversation. That text can expose the expected output instead of leaving the model to learn from the question and response normally.

SOURCE-LINKED FINDING
Source row49,968
Text foundCorrect Answer: C
FindingExplicit answer field
Review decisionMarked for removal
The finding identifies both the affected row and the exact answer text that triggered review.
04 / REPEATED CONVERSATIONS

7,343 later copies were traced back to their earlier source rows.

Repeated examples can give the same content more influence than intended. Datascreen linked every later copy to its earlier source row, then checked all 7,343 relationships against the source data.

DUPLICATE LINK AUDIT
Later copies checked7,343
Copies linked to themselves0
Copies linked to a later row0
Different text linked as a copy0
Largest repeated set50 rows
The largest group contained one earlier row and 49 later copies.
05 / WHAT THIS ESTABLISHES

Evidence for review, with clear limits on the claim.

OBSERVED

The public dataset contained exposed credential patterns, expected-answer text, and repeated conversations. Datascreen connected those findings to the affected source rows and review decisions.

NOT ESTABLISHED

This scan does not prove malicious poisoning, downstream model harm, regulatory non-compliance, or that every possible data problem was found.

Exact source

ShareGPT_2023.05.04v0_Wasteland_Edition.json

SHA-256 · 701b6b1beea876183b768b2ac530ba0b7f32a5e3755885c4e36b92302aa4c6bd

View source dataset →
Inspect your own dataset

Find review-worthy rows before they reach model training or evaluation.

Datascreen gives teams source-linked findings, suggested actions, and a durable record of every review decision.