The SciTrust™ Score
A transparent, reproducible composite metric that quantifies confidence in the quality of a scientific study, based on independent expert assessment.
A single number for study quality
The SciTrust™ Score is a 0 to 10 composite metric assigned to each scientific publication reviewed through SciPinion’s certified peer review platform. It distills the judgments of a panel of independent subject-matter experts into one interpretable value that reflects overall confidence in a study’s methodology, findings, and conclusions.
Each reviewer evaluates the study across three dimensions: Methods, Results, and Discussion/Conclusions. These component ratings are aggregated using a two-step process that balances within-reviewer consistency against across-reviewer agreement, producing a score that is both mathematically principled and practically meaningful.
Each component rating is anchored at three reference points: 0 marks a fundamental flaw, 5 medium confidence, and 10 the highest confidence. Studies that score well must demonstrate quality across all three dimensions and earn strong marks across the review panel.
Score Components
Peer review without measurement is incomplete
The limitation of traditional review
Conventional peer review produces a binary outcome: accept or reject. It does not generate a quantitative record of study quality, nor does it allow systematic comparison across studies within a body of evidence. Decision-makers reviewing dozens or hundreds of studies have no standardized way to weight one study against another.
What SciTrust addresses
The SciTrust Score fills this gap by providing a numerical quality metric that is transparent in its construction, reproducible across reviews, and comparable across studies, endpoints, and evidence streams. Regulatory bodies, expert panels, and systematic reviewers can use SciTrust Scores to incorporate study quality directly into weight-of-evidence assessments, rather than relying on subjective narrative judgments alone.
Key properties
The score is designed to reflect several principles that matter for defensible scientific evaluation. It penalizes inconsistency within a study: strong results cannot compensate for fundamentally flawed methods. It treats reviewer disagreement as legitimate discourse rather than error. And it operates under an objective process whereby reviewers are selected based on expertise and not relationships with an editor. All reviewers are blinded to each other and remain anonymous so as to provide psychological safety to reviewers.
Transparent. Every input is recorded, and the aggregation formula is fully disclosed.
Reproducible. Given the same component scores, anyone can compute the same result.
Comparable. Scores are on a common scale across studies, endpoints, and evidence streams.
Defensible. Reviewer independence and anonymity guard against identity bias in the evaluation.
Wondering how SciTrust Scores would apply to your body of evidence?
Let’s DiscussTwo-step aggregation, by design
The SciTrust Score uses different aggregation methods at each level for principled reasons. The geometric mean within each reviewer penalizes inconsistency. The arithmetic mean across reviewers preserves the equal weight of independent expert judgment.
Component Scoring
Each member of an independent expert panel rates the study’s Methods, Results, and Discussion/Conclusions on a 0 to 10 scale. Reviewers remain anonymous and do not see one another’s scores.
Reviewer Score
Each reviewer’s three component scores are combined using the geometric mean: the cube root of Methods × Results × Discussion. This ensures a fatal flaw in any one dimension dominates the result.
Study Score
The SciTrust Score is the arithmetic mean of all reviewer scores. Each reviewer’s independent assessment carries equal weight. Disagreement is averaged, not penalized.
From component scores to SciTrust
Consider a study reviewed by a panel of three experts with the component scores shown in the table. Each reviewer’s score is the geometric mean (cube root of the product) of their three ratings.
Reviewer 1 scored the study 8, 7, and 6 across Methods, Results, and Discussion. Their geometric mean is (8 × 7 × 6)1/3 = 6.95. By contrast, Reviewer 3 gave lower and more variable scores (3, 2, 4), yielding a geometric mean of just 2.88. The geometric mean ensures that the low Results score of 2 substantially pulls down that reviewer’s overall assessment.
The final SciTrust Score is the arithmetic mean of the three reviewer scores: 4.95. Reviewer scores are rounded for display; the calculation carries full precision. The same calculation extends to panels of any size.
Want to see SciTrust Scores in a real review? Request a sample report.
| Methods | Results | Discussion | Reviewer Score | |
|---|---|---|---|---|
| Reviewer 1 | 8 | 7 | 6 | 6.95 |
| Reviewer 2 | 5 | 5 | 5 | 5.00 |
| Reviewer 3 | 3 | 2 | 4 | 2.88 |
| SciTrust Score | 4.95 | |||
Reviewer scores rounded to two decimals for display; the SciTrust Score is computed from unrounded values.
Formulas
Why the geometric mean?
The geometric mean is chosen for within-reviewer aggregation because it reflects a core principle of scientific evaluation: a study is only as strong as its weakest component. A brilliantly designed study (Methods = 9) with unreliable results (Results = 1) should not receive a moderate score. The geometric mean ensures it does not.
If any single component is rated 0, the geometric mean is 0 regardless of the other ratings. The geometric mean is always less than or equal to the arithmetic mean, with equality only when all three components receive the same score. This property naturally rewards consistency across methodological, analytical, and interpretive quality.
The arithmetic mean is appropriate at the second aggregation step because each reviewer’s independent judgment should carry equal weight. Unlike within-reviewer aggregation, where a fatal flaw in any component should dominate, disagreement between reviewers represents legitimate scientific discourse. Averaging preserves the signal from each independent assessment without disproportionately penalizing the outlier.
This two-step structure, geometric then arithmetic, produces a score that is sensitive to both internal study quality and the breadth of expert opinion. It is a deliberate design choice, not a default.
Fit for purpose, when the decision demands it
The SciTrust Score measures how well a study was conducted. Many decisions also turn on whether a study is suited to the specific question being asked. For those applications, the score can be extended with additional fit-for-purpose components that assess a study’s relevance to the decision at hand, alongside its quality.
A rigorous study of the wrong exposure, population, or endpoint may warrant less weight than its quality alone would suggest. Fit-for-purpose components capture that distinction while leaving the underlying quality score intact.
Ready to quantify study quality?
Learn how SciPinion’s certified peer review platform can bring transparent, reproducible quality metrics to your evidence evaluation.
Get in Touch Explore Our Process