grantscience.com · evidence synthesis
The shape of the evidence on grant peer review
How reliable is the expert review that decides who gets funded? This site presents a systematic synthesis of the research literature on peer-review reliability, grant review first, with journal review and neighbouring judgment contexts as the comparison: what has been studied, where the gaps sit, and what the measured agreement between reviewers actually looks like.
9,783 records identified → 9,242 screened (dual AI screen, kappa 0.866; recovered arm 0.890) → 1,627 included → 1,123 in the corpus after de-duplication and retrieval
Scoping funnel: PRISMA-ScR scoping search, kappe-scoping screening records. Provisional.
01The field at a glance
What the corpus is made of. The two facets, review context and research construct, re-express the screening themes around the decomposition of judgment error into bias and noise; study design is classified from titles and abstracts by the coding panel.
02Does the field practise what it studies?
Every paper is coded for the transparency of its own reporting: preregistration, data and code availability, statement by statement. Preregistration is counted among papers that could have been preregistered (empirical studies), open data among papers reporting their own data. The literature on peer review adopts these practices late, but the recent shift is clear.
PreregisteredOpen data
03Explore the evidence
Four views over one dataset. Every count, label and colour derives from the generated data files, and every view opens on the grant subset with the other contexts one click away.
Evidence gap map
Process stage crossed with construct: bubbles sized by number of studies, coloured by the share that are controlled experiments. 422 papers in the Grant subset, mapped over 7 stages of the review process.
Reference network
All 1,123 papers as an interactive map: 5 layouts, 8,422 in-corpus citation links, recolourable by any facet, with per-paper summaries and references out into the wider literature.
Reliability, coded
Every reliability coefficient located in the corpus, coded to its exact form so like compares with like. Median single-reviewer ICC 0.28 against a panel-average 0.55.
Browse the corpus
Search and filter all 1,123 papers by context, construct, study design and process stage, each with an AI-generated summary and a link to the source.
04Read this with the caveats
Not yet source-verified. Every number on this site is provisional until checked against the primary source's full text.
Machine-classified. Context, construct and study design are the majority vote of a five-model coding panel reading titles and abstracts, with ties adjudicated by hand; process stages come from a separate AI pass. The full human validation pass is still pending, so labels and counts may change.
A modest corpus for topic modelling. Inductive topics and time-sliced trends rest on a corpus that is small by topic-modelling standards; early time slices are thin.
Original summaries. Paper summaries are AI-generated original text, never the publisher's abstract, and are flagged for verification before citation.