Datascreen
Public dataset case study / August 2026

Datascreen found 633 findings across 376 HH-RLHF rows.

Each preference example should end with one candidate answer. In 40,000 prepared source rows, Datascreen found 633 messages that continued into an extra Human or Assistant turn.

ROWS SCANNED 40,000 A bounded slice of the public dataset
FINDINGS 633 Messages containing an extra conversation turn
AFFECTED ROWS 376 Rows containing at least one finding
INCOMPLETE ROWS 3 Rows missing a chosen or rejected answer
01 / WHAT A CLEAN ROW SHOULD CONTAIN

Preference training expects two separate candidate answers to one prompt.

HH-RLHF preference comparisons pair a prompt with a chosen answer and a rejected answer. Each side should preserve the conversation context and end with one complete assistant response. Keeping those boundaries intact lets the training process compare like with like.

PROMPT What the user asked

The shared conversation context for both candidate answers.

CHOSEN ANSWER One preferred response

The answer the dataset says should be preferred.

REJECTED ANSWER One less-preferred response

The alternative answer used for the same comparison.

02 / WHAT DATASCREEN FOUND

Some stored messages continued into another Human or Assistant turn instead of ending where expected.

The prepared dataset treated the text as a single message, but the content inside that message began another conversation turn. The stored candidate therefore contained multiple turns rather than one clean answer.

EXPECTED CANDIDATE One complete answer

The response ends before another speaker begins.

AFFECTED CANDIDATE Answer plus another conversation turn

The same stored message continues as though a new Human or Assistant turn has begun.

WHY THE ROW NEEDS REVIEW

The training example no longer contains the single response its structure claims to contain.

A reviewer can separate the combined turns when the intended split is clear or remove the row when it is not. The finding is linked back to the original source row so that decision is made with the surrounding content available.

03 / THE SCAN RESULT

Datascreen found 633 messages that continued into another Human or Assistant turn across 376 source rows.

A prepared source row can contain multiple messages, and more than one message in the same row can contain an extra conversation turn. That is why 633 affected messages came from 376 source rows.

AFFECTED ROWS 376

Each row contained at least one message that continued after it should have ended.

FINDINGS 633

Each recorded finding represents one affected message.

ASSISTANT MESSAGES 629

Stored assistant messages that continued into another Human or Assistant turn.

USER MESSAGES 4

Stored user messages that continued into another Human or Assistant turn.

04 / REVIEW EVIDENCE

A fixed sample confirmed the extra turns were already present in the source data.

Fifty findings were selected in advance and reviewed. Every sampled finding required removal or repair. None had been introduced while converting the source into the format used by the scanner, and one sampled finding was traced back to the original compressed source.

FIXED SAMPLE REVIEW
Findings reviewed from the preselected sample50
Reviewed findings requiring removal or repair50
Findings caused by preparing the data for scanning0
Findings traced to the original compressed source file1
This fixed sample supports the result for those 50 findings. It does not classify the remaining findings.
05 / MISSING ANSWERS

Three additional rows could not provide the two answers required for preference comparison.

One row was missing its chosen answer and two were missing their rejected answer. These rows need the missing response restored or must be excluded before they can support preference training.

INCOMPLETE PREFERENCE ROWS
Chosen answer missing1
Rejected answer missing2
Total affected rows3
Without both answers, the intended preference comparison cannot be performed.
06 / WHAT THIS ESTABLISHES

A concrete structural problem with an honest evidence boundary.

OBSERVED

The 40,000-row slice contained 633 messages with an extra conversation turn across 376 source rows. Every finding in the fixed 50-item sample required removal or repair.

NOT ESTABLISHED

The sample does not classify all 633 findings. The scan also does not establish malicious poisoning, downstream model harm, or the quality of the complete upstream HH-RLHF dataset.

Exact source record

40,000 HH-RLHF rows used for this scan

SHA-256 · 70b1d6253cab5e58d20bc126271ea4d74b730266ef9fd5770f8d7b5549c8ac21

View source dataset → Scan date · 23 August 2026
Inspect your own dataset

Find rows whose stored structure no longer matches the training example inside them.

Datascreen links each finding to the affected source row so reviewers can separate the combined turns, remove the row, or preserve it with a recorded decision.