Biomonitoring, the practice of measuring chemicals or their metabolites in urine, blood, or other biological samples, has become a common way to estimate individual exposures in epidemiology. Yet many chemicals of current interest, including bisphenols, phthalates, and triclosan, have short-lived biomarkers whose concentrations swing widely within the same person from day to day, in some cases by several thousand-fold in a single day. A single spot sample can badly misrepresent a person’s typical exposure. When studies fail to account for this within-person variability, effect estimates can be biased and studies can be underpowered, which in turn weakens the risk assessments and health policies that rely on them.
To address this, SciPinion set out to give researchers both practical tools and an independent, expert-vetted set of best practices. Two of the study authors developed a suite of four free, web-based “calculators” that estimate the sample sizes, the number of repeated measurements per person, the statistical power, and the expected bias for linear and logistic exposure-response models under classical measurement-error assumptions. SciPinion then convened an independent panel of nine experts in epidemiology, biostatistics, and exposure assessment, drawn from institutions across Australia, North America, and Western Europe and carrying 13 to 45 years of experience each, to peer review the calculators and to develop best-practice recommendations.
The review followed SciPinion’s triple-blinded, three-round modified Delphi process. The project sponsor and the panelists were blinded to one another, the panelists were blinded to each other (identified during deliberations only as “Expert 1,” “Expert 2,” and so on), and every response and comment was recorded and reported in full to support transparency and reduce the risk of groupthink. Across three rounds, the experts reviewed a white paper describing the calculators, tested the tools themselves, debated one another’s answers, and responded to structured charge questions on best practices.
The panel judged the underlying calculations to be sound and well motivated and rated the tools’ usability highly, recommending only targeted refinements, chiefly clearer labeling of input and output units and a clearer presentation of the logistic-regression results, all of which were implemented. The experts agreed that the calculators address a genuine gap for epidemiologists who lack ready access to a statistician or specialized software.
On best practices, the panel converged on several recommendations: run a pilot study to estimate the biomarker’s variance components; specify in advance the minimum effect size a study intends to detect; favor biomarkers with the highest intraclass correlation coefficient, a measure of how reliably repeated samples distinguish one person’s exposure from another’s; adjust for measurement error in both exposure and outcome; treat expected bias above roughly 10 percent as a signal to correct for it during analysis; interpret results through confidence intervals rather than rigid p-value cutoffs; and design studies for statistical power closer to 90 percent than the customary 80 percent.
To show the tools in practice, the panel’s methods were applied to two published biomonitoring studies: one linking urinary bisphenols to fetal growth across 1,379 pregnancies, and one linking urinary triclosan to childhood neurodevelopment among 377 mother-child pairs. In both cases the calculators indicated the studies were substantially underpowered and carried expected bias in their effect estimates well above the panel’s 10 percent threshold. Each would have required many more subjects, more repeated measurements per person, or both, to reliably detect even the associations it reported. The exercise illustrates how readily unaddressed within-person variability can produce findings that cannot support the conclusions drawn from them.
The implications extend beyond individual study design. Because regulatory and public-health decisions increasingly rest on biomonitoring-based epidemiology, unrecognized measurement error can propagate into risk assessments and the policies built on them. The panel’s recommendations and the accompanying calculators give researchers a way to plan studies that are adequately powered and appropriately corrected for bias before data collection begins, while making clear that software complements rather than replaces collaboration with a statistician.
The authors conclude that justifying sample size and study design is essential whenever exposure is estimated from biomarkers, because within-person variability in those biomarkers introduces measurement error that can distort effect estimates. Making the relevant power and bias calculations accessible, together with the panel’s best practices, should support improvements in current epidemiologic practice, particularly for researchers without ready access to specialized statistical expertise. Absent attention to within-person variability in both design and analysis, the authors caution, it will remain difficult to distinguish causal associations from spurious ones.
Read the full paper here.
Try the calculators here.
Back to Panel Findings

Figure 1. Expected bias in the slope of a linear regression as a function of the intraclass correlation coefficient (ICC) of the exposure measure, shown for different numbers of repeated measurements per person (m). The horizontal dashed line marks the 10 percent bias threshold above which the expert panel recommends adjusting for bias during data analysis.