q08

Calibration that masks environmental coupling

2026-09-19 · Btrfs/ZFS/bcachefs under workloads class

The recent discussion of Btrfs, ZFS and bcachefs performance on classic benchmarks notes that the author ran 593 calibration‑filtered runs on GitHub Actions runners before seeking a dedicated Hetzner server that survived only three measurements. This pattern reveals a measurement system in which the metric of interest is entangled with uncontrolled environmental variation, and the response is to apply a post‑hoc filter that cannot remove the coupling.

The benchmarkers are motivated to publish numbers that suggest a clear hierarchy among the three filesystems. To increase confidence they repeat each workload many times, compute an average, and discard runs that exceed a preset variance threshold. The calibration step is described as a way to “reject completely unreliable VMs” and to “limit” the influence of noisy neighbours, although the author acknowledges that it cannot fully fix the problem. Because the GitHub Actions runners share underlying hardware with other jobs, the CPU frequency, memory bandwidth and cache contention fluctuate in ways that are correlated with the workload being measured. When a run is discarded because its latency spikes, the underlying cause — a transient increase in neighbour load — is not eliminated from the remaining sample; it merely shifts the distribution of the observed times toward lower values. The average of the surviving runs therefore remains biased upward or downward depending on the direction of the environmental perturbation, and the bias does not diminish with more repetitions because each new run is subject to the same stochastic coupling.

The same logical structure appears in unrelated fields where producers of a metric have an incentive to present a favourable outcome and the measurement apparatus is susceptible to contextual factors that covary with the target. In the credit‑rating industry, agencies such as Moody’s and S&P are paid by the issuers whose securities they evaluate. Analysts build quantitative models that incorporate issuer‑supplied data, and they may adjust assumptions or override model outputs to keep the rating within a range that the issuer finds acceptable. The resulting rating is therefore a function both of the underlying credit risk and of the fee‑generating relationship; attempts to “calibrate” the model by back‑testing against historical defaults do not break the coupling because the source of bias — issuer compensation — continues to influence the judgment at every step. Historical precedents show this mechanism producing systemic mispricing: in the years preceding the 2008 financial crisis, tranches of subprime mortgage‑backed securities received AAA ratings despite underlying default probabilities that later proved an order of magnitude higher, and the agencies’ internal documents reveal that overrides were frequently applied to accommodate issuer preferences while the quantitative core remained unchanged.

In educational testing, schools and teachers are evaluated by the average scores of their pupils on standardized examinations. The incentive to achieve high scores leads to instructional time being reallocated toward test‑specific drills, practice exams and item‑familiarisation strategies. The measured score thus reflects a mixture of genuine subject mastery and test‑taking proficiency that has been coached. Statistical procedures that attempt to “adjust for socioeconomic status” or to “remove guessing effects” treat the coaching influence as noise to be filtered out, yet because the coaching is systematically applied across the cohort it shifts the mean of the distribution in a way that persists regardless of how many test administrations are averaged. Longitudinal studies of programs that eliminated high‑stakes testing have documented drops in reported proficiency that were not accompanied by comparable declines in independent assessments of skill, indicating that the original scores were inflated by the measurement‑environment coupling.

Clinical trials of pharmaceuticals provide a further illustration. Investigators are rewarded for demonstrating a statistically significant benefit of the investigational drug over control, and trial sponsors often select sites with populations that are expected to respond favourably. Site‑level differences in clinician enthusiasm, patient adherence, or concomitant medication use can therefore covary with the treatment effect estimate. Central monitoring committees routinely apply statistical models that include site as a random effect or that discard outliers based on residual diagnostics, but these adjustments assume that the site influence is exchangeable and independent of the treatment. When the site effect is correlated with the assignment — for example, when higher‑performing sites are preferentially chosen for the active arm — the adjustment fails to remove the bias, and the reported efficacy remains inflated. Historical examples include early trials of anti‑arrhythmic agents in the 1980s, where investigators repeatedly excluded sites with higher baseline mortality, producing an apparent survival advantage that vanished when later trials mandated random site allocation and blinded outcome assessment.

Physical constants experiments exhibit the same coupling between apparatus and environment. The Cavendish torsion balance used to measure the gravitational constant G is suspended in a room where temperature gradients cause the supporting fibre to expand or contract, altering the restoring torque. Contemporary repetitions of the experiment place the apparatus in an underground laboratory, surround it with mu‑metal shielding, and employ active temperature control, yet residual seismic tremors and magnetic fluctuations still impart tiny torques that mimic a gravitational signal. Data‑analysis pipelines often fit and subtract a low‑frequency drift model, treating the environmental contribution as additive noise that can be averaged away. Because the coupling is multiplicative — the torque scales with the instantaneous fibre length — the subtraction leaves a systematic offset that varies with the season and with the local microseismic spectrum, explaining why different laboratories report values of G that differ beyond their quoted uncertainties.

Agricultural yield trials face an analogous problem. Researchers test new cultivars across multiple locations, seeking to identify genotypes with superior productivity. Yield is strongly influenced by soil moisture, nutrient availability, and micro‑climate, all of which vary from plot to plot and from season to season. To obtain a stable estimate, analysts use randomized complete block designs and include plot‑level covariates in a linear model, assuming that the block effects capture all spatial heterogeneity. When an unobserved factor such as a localized pathogen outbreak coincides with the placement of a particular genotype, the block adjustment mistakenly attributes the yield decrement to the block rather than to the genotype‑environment interaction, and the resulting genotype ranking is biased. Historical wheat‑breeding programs in the mid‑twentieth century documented cases where a line appeared superior in early trials because it was inadvertently planted in fields with lower disease pressure; later multi‑environment analyses that incorporated disease‑scoring covariates revealed the earlier advantage to be spurious.

Across these domains the causal chain is identical: an actor seeks to maximise a measurable outcome; the measurement process is exposed to a contextual variable that covaries with the actor’s behaviour or with the object of measurement; the actor responds by applying a post‑hoc filter, calibration, or statistical adjustment that treats the contextual influence as random noise; because the influence is in fact systematic and coupled to the measurement, the filter fails to eliminate bias, and the reported metric remains distorted. The persistence of the bias does not depend on the frequency of repetition; averaging many noisy observations only reduces the variance of the random component while leaving the systematic component intact.

The mechanism does not require modern technology, nor is it confined to any single discipline. It appears whenever a reward structure ties the evaluator’s payoff to the measured value, and whenever the evaluation apparatus shares a physical or informational substrate with the thing being evaluated. The historical record shows that attempts to mitigate the problem through environmental controls, blinding, or sophisticated modelling have repeatedly fallen short when the coupling itself was not addressed at its source. The enduring challenge, therefore, is not to refine the calibration but to redesign the measurement setting so that the metric of interest is statistically independent of the contextual variables that the evaluator can influence. Until such a separation is achieved, the numbers produced will continue to reflect the hidden hand of the environment as much as the property they purport to quantify.

Was this worth your time? yesflatno

Sources & further reading