LitSynthLitSynth
  • Features
  • Pricing
  • FAQ
  • Literature Review
  • Blog

Loading…

LitSynthLitSynth

Transform your research workflow with AI-powered literature analysis and synthesis

DiscordEmail
Product
  • Features
  • Pricing
  • FAQ
Resources
  • Blog
  • Changelog
Friends
  • Find Papers
  • LitFigure – Scientific Figures
Company
  • Contact
Legal
  • Privacy Policy
  • Terms of Service

Research workflows

Explore focused guides for literature reviews, PubMed workflows, citation audits, and PRISMA-lite review planning.

AI literature review generatorsystematic review AI toolPRISMA litesystematic review screeningsystematic review data extractionsystematic review protocol generatorsemaglutide systematic review exampleAI in education literature review exampleclimate change health literature review example
©️ 2026 LitSynth. All Rights Reserved.
SCORED ONCE · OVERALL FAIL

Screening safety and workload, measured on untouched holdouts

On the untouched 281-record holdout, LitSynth made zero false exclusions and protected every relevant record. But 82.9% still required manual review, missing the preregistered 80% workload ceiling. The holdout result is therefore a failure.

TOPIC / CD000996

TASK / Intervention

RECORDS / 281

PROTOCOL VERSION / 1

POLICY VERSION / title-abstract-v3

PREREG COMMIT / 8fd1e19

BLIND HASH / 282a24a4830f…

STATUS / FAIL — NO RETEST

Protected recall

100.0%

9/9 relevant records protected · PASS

False exclusions

0

0.0% of relevant records

Manual review

82.9%

233/281 records · FAIL (limit ≤80%)

Auto exclusion

15.7%

44/281 records recommended for exclusion

Overall result

FAIL

both gates had to pass; no composite accuracy score

01 / Locked before unblinding

The failure is part of the result.

Safety passed. Workload did not. The workload gate failed by 2.9 percentage points. We did not revise the locked policy or rerun the corpus after seeing labels.

  1. LOCK 01

    Corpus lock

    CD000996: Inhaled corticosteroids for bronchiectasis. The pinned source commit, topic hash, PMID-list hash, and 281-record count are public.

  2. LOCK 02

    Policy lock

    title-abstract-v3, deepseek-v4-flash, temperature 0.1, batch size 5, and source hashes are fixed.

  3. LOCK 03

    Predictions-before-labels

    All 281 label-free decisions must be written and hashed before the runner is allowed to fetch qrels.

  4. LOCK 04

    Safety plus workload

    False exclusion and protected recall are published beside UNSURE count and manual-review rate. A safe but unusable review queue cannot pass.

02 / Inspect the complete audit trail

Reproduce the order, counts, and failure.

Predictions sealed (UTC)
2026-09-01T02:31:50.490Z
Labels fetched / unblinded (UTC)
2026-09-01T02:31:50.492Z
Raw blinded predictions · SHA-256
282a24a4830f59e1ce7e1b91c449e7dcc9171464784a3075f0cfea827693f486
Full scored predictionspredictions.csvFalse-exclusion ledgerfalse-excludes.csvManual-review ledgermanual-review.csvMetrics and gatesmetrics.jsonRun order and blind hashrun.jsonCorpus and source hashesmanifest.jsonBlinded predictionsblinded-predictions.jsonMachine-readable preregistrationpreregistration.jsonHuman-readable preregistrationpreregistration.md

Limitations

Only one untouched holdout has been scored: one CLEF TAR topic, 281 records, and 9 gold-relevant records, using titles and abstracts. Protected recall counts relevant records marked INCLUDE or UNSURE; it does not mean every relevant record was correctly included.

Manual review here means UNSURE / all records (233/281). It is a queue-size measure, not measured reviewer time or total human work: INCLUDE and EXCLUDE recommendations still need reviewer oversight. Auto exclusion is the model recommendation rate, not permission to discard records without review.

What this does NOT prove

  • Guaranteed recall or workload savings on other topics, corpora, models, or policy versions.
  • Full-text eligibility accuracy, data-extraction accuracy, risk-of-bias validity, or claim-level cited synthesis quality.
  • Superiority to Rayyan or another tool: this is not a head-to-head comparison.
  • A replacement for dual independent screening or the reviewer’s final inclusion and exclusion decisions.

The 277-record corpus is no longer a performance claim

Its labels informed policy revision, so it remains only as a transparent post-error-analysis regression archive. The untouched holdout was scored once. Its failed workload gate is published and the corpus will not be recycled as untouched evidence.

Open historical development archive