Preference training expects two separate candidate answers to one prompt.
HH-RLHF preference comparisons pair a prompt with a chosen answer and a rejected answer. Each side should preserve the conversation context and end with one complete assistant response. Keeping those boundaries intact lets the training process compare like with like.
The shared conversation context for both candidate answers.
The answer the dataset says should be preferred.
The alternative answer used for the same comparison.
Some stored messages continued into another Human or Assistant turn instead of ending where expected.
The prepared dataset treated the text as a single message, but the content inside that message began another conversation turn. The stored candidate therefore contained multiple turns rather than one clean answer.
The response ends before another speaker begins.
The same stored message continues as though a new Human or Assistant turn has begun.
The training example no longer contains the single response its structure claims to contain.
A reviewer can separate the combined turns when the intended split is clear or remove the row when it is not. The finding is linked back to the original source row so that decision is made with the surrounding content available.
Datascreen found 633 messages that continued into another Human or Assistant turn across 376 source rows.
A prepared source row can contain multiple messages, and more than one message in the same row can contain an extra conversation turn. That is why 633 affected messages came from 376 source rows.
Each row contained at least one message that continued after it should have ended.
Each recorded finding represents one affected message.
Stored assistant messages that continued into another Human or Assistant turn.
Stored user messages that continued into another Human or Assistant turn.
A fixed sample confirmed the extra turns were already present in the source data.
Fifty findings were selected in advance and reviewed. Every sampled finding required removal or repair. None had been introduced while converting the source into the format used by the scanner, and one sampled finding was traced back to the original compressed source.
Three additional rows could not provide the two answers required for preference comparison.
One row was missing its chosen answer and two were missing their rejected answer. These rows need the missing response restored or must be excluded before they can support preference training.
A concrete structural problem with an honest evidence boundary.
The 40,000-row slice contained 633 messages with an extra conversation turn across 376 source rows. Every finding in the fixed 50-item sample required removal or repair.
The sample does not classify all 633 findings. The scan also does not establish malicious poisoning, downstream model harm, or the quality of the complete upstream HH-RLHF dataset.
40,000 HH-RLHF rows used for this scan
SHA-256 · 70b1d6253cab5e58d20bc126271ea4d74b730266ef9fd5770f8d7b5549c8ac21