Loading…
Loading…
On the untouched 281-record holdout, LitSynth made zero false exclusions and protected every relevant record. But 82.9% still required manual review, missing the preregistered 80% workload ceiling. The holdout result is therefore a failure.
TOPIC / CD000996
TASK / Intervention
RECORDS / 281
PROTOCOL VERSION / 1
POLICY VERSION / title-abstract-v3
PREREG COMMIT / 8fd1e19
BLIND HASH / 282a24a4830f…
STATUS / FAIL — NO RETEST
Protected recall
100.0%
9/9 relevant records protected · PASS
False exclusions
0
0.0% of relevant records
Manual review
82.9%
233/281 records · FAIL (limit ≤80%)
Auto exclusion
15.7%
44/281 records recommended for exclusion
Overall result
FAIL
both gates had to pass; no composite accuracy score
01 / Locked before unblinding
Safety passed. Workload did not. The workload gate failed by 2.9 percentage points. We did not revise the locked policy or rerun the corpus after seeing labels.
LOCK 01
CD000996: Inhaled corticosteroids for bronchiectasis. The pinned source commit, topic hash, PMID-list hash, and 281-record count are public.
LOCK 02
title-abstract-v3, deepseek-v4-flash, temperature 0.1, batch size 5, and source hashes are fixed.
LOCK 03
All 281 label-free decisions must be written and hashed before the runner is allowed to fetch qrels.
LOCK 04
False exclusion and protected recall are published beside UNSURE count and manual-review rate. A safe but unusable review queue cannot pass.
02 / Inspect the complete audit trail
Only one untouched holdout has been scored: one CLEF TAR topic, 281 records, and 9 gold-relevant records, using titles and abstracts. Protected recall counts relevant records marked INCLUDE or UNSURE; it does not mean every relevant record was correctly included.
Manual review here means UNSURE / all records (233/281). It is a queue-size measure, not measured reviewer time or total human work: INCLUDE and EXCLUDE recommendations still need reviewer oversight. Auto exclusion is the model recommendation rate, not permission to discard records without review.
Its labels informed policy revision, so it remains only as a transparent post-error-analysis regression archive. The untouched holdout was scored once. Its failed workload gate is published and the corpus will not be recycled as untouched evidence.
Open historical development archive