# LitSynth untouched screening holdout v1 — preregistration

Status: **frozen before labels**.

This evaluation locks one unseen CLEF TAR 2019 Task 2 Intervention topic (`CD000996`, 281 records), the `title-abstract-v3` policy, the DeepSeek model/configuration, the review protocol, and two independent acceptance gates before reading qrels.

## Locked outcomes

- Safety: report false exclusions and protected recall; pass at protected recall >= 95%.
- Workload: report UNSURE/manual-review count and rate; pass at manual review <= 80%.
- The run passes only if both gates pass. There is no composite “accuracy” score.

## Blinding contract

The runner must validate the topic and PMID-list hashes, screen all 281 records, write and hash a label-free blinded prediction artifact, and only then fetch qrels. It records the preregistration commit, blinded-prediction hash, qrels hash, and unblinding time.

If a gate fails, LitSynth publishes the failure. This corpus cannot be tuned and rerun as untouched evidence.

The earlier 277-record benchmark remains downloadable only as a post-error-analysis regression archive; its results are not used for current performance promotion.

