Abstract
A research institution that tests many ideas and rejects most of them must decide when a rejected idea may be tested again. We model a stream of ideas, an evidence standard that improves in steps, and five policies: bury forever, retry on request, retry the same data with analyst freedom, sweep the graveyard after each improvement under a spending rule, and sweep under a batch false discovery rate rule. At matched retest budgets, retry on request doubles the false discovery proportion from 0.098 to 0.217 and tuning the same data more than triples it, while a spending sweep recovers all but 0.4 percent of true ideas at 0.102. Improvements can be flawed. The graveyard, being mostly false ideas, works as a free calibration set that caught every flawed improvement we simulated, but it is blind to variance flaws.
Motivation
A group runs a test on an idea and the idea fails. What then? Two answers are common and both are wrong in a measurable way.
The first is to bury it permanently. This is clean, and it is how a group guards itself against the pleasure of rescuing its favourites. But the test that killed the idea had a certain power, and an idea with a modest real effect is more likely than not to fail a weak test. If the evidence standard later improves, a permanent grave is a standing loss of true findings.
The second is to let anyone reopen anything at any time. This feels open-minded. It also gives every false idea another lottery ticket each time someone asks, and the people who ask are the people who believe in the idea.
There is a third move that looks like the second and is worse: keep the data, vary the analysis, and run the test again on the same evidence until it passes. The rule this paper tests is that a rejected idea may be reopened only when the court itself improves, that is, when the test becomes more able to tell true from false, and then every affected idea is retested under the same improved standard. Reopen the verdict by upgrading the procedure, not by coaching the defendant. We ask what each rule costs and buys, and what is needed to keep the upgrade itself honest.
Related work and what is new
Sequential testing has a developed theory. Foster and Stine's alpha-investing and the online procedures of Javanmard and Montanari control the false discovery rate as hypotheses arrive in a stream, each tested once and the budget replenished by discoveries [1, 2]. Benjamini and Hochberg's procedure controls it within a batch [3]. Dwork and colleagues show how a holdout set can be reused safely when analyses adapt to it [4]. Simmons, Nelson and Simonsohn quantified how analytic flexibility inflates false positives [5]. Ioannidis argued from the base rate of true hypotheses that most claimed findings are false [6]. Efron proposed estimating the null distribution empirically from the bulk of the test statistics rather than assuming it [7]. Empirically, a study that re-ran failed replications at larger sample sizes found that most still failed to replicate, so extra resources seldom rescued them [8].
What these do not address is the object we study: a population of hypotheses already tested and rejected, to be reopened at later times, under a rule that says who may reopen and when. We contribute four things. A policy comparison at matched retest budgets. A likelihood-ratio accounting of what each kind of retest is worth, which shows that the graveyard is a different population from the stream and that same-data tuning adds almost no information. A flawed-upgrade model. And a test, using the graveyard itself, for whether a proposed upgrade is valid before it is used to exhume anything. Our search was not exhaustive and the multiple-testing ingredients are standard; the claim concerns the combination.
Method
The world
Ideas arrive one hundred per period for forty periods. Each is true with probability 0.10. A true idea has a signal strength λ drawn uniformly from 1 to 4, a false one has zero. The institution tests an idea by computing a statistic z equal to λ times the current resolution r, plus standard normal noise, and certifies the idea if z exceeds the threshold for a one-sided level α = 0.01 (z above 2.326). Resolution starts at 1 and is multiplied by 1.4 at each of four upgrades, at periods 8, 16, 24 and 32. The first test therefore has an average power of 0.55, and later tests have more. Failed ideas are buried, carrying the data noise they were tested with and the generation of the test that failed them.
The units are abstract. Resolution corresponds to anything that raises the signal-to-noise ratio of a test: more data, a better instrument, a cleaner definition.
Policies
All five policies receive the same fresh ideas, the same upgrades and the same random numbers. They differ only in what happens to buried ideas.
- Bury: nothing, ever.
- Free retry: each period, each buried idea is retested with probability 0.10 on fresh data at the current resolution and the flat threshold. An idea may be retried any number of times.
- Coach: each period, each buried idea is, with the same probability, retested on its original data with eight analytic variants. Each variant shares 70 percent of its noise variance with the original test and has its own independent noise for the rest. The idea is certified if any variant passes the flat threshold. This models trying alternative specifications of the same evidence.
- Sweep with spending: when an upgrade is adopted, every buried idea from an earlier generation is retested once on fresh data at the new resolution. After the g-th adopted upgrade the threshold is for level α divided by g(g + 1), so the levels over all sweeps sum to α.
- Sweep flat: as above but at level α every time.
- Sweep with a batch rule: as above, but the sweep's retest p-values are passed through the Benjamini and Hochberg procedure at level 0.15 [3].
Free retry and coach have a retest probability of 0.10 chosen so that their average total number of retests matches the sweeps (about 7200 against 7150 to 7220). Sweeps retest every eligible idea at every upgrade, and the number grows as the graveyard grows.
We report the false discovery proportion as the ratio of total false certifications to total certifications, summed across runs, with 95 percent bootstrap intervals over one thousand runs. We also report the proportion of true ideas that remain buried at the end, and the same quantities for exhumed ideas alone.
Flawed upgrades
With probability one half, an upgrade has a hidden defect that adds a constant shift δ to every statistic, true or false, until the next adopted upgrade replaces it. This stands for a leak, a mis-specified filter, a contaminated reference set. A flawed upgrade raises the false positive rate. At δ = 0.8 the level 0.01 becomes about 0.063.
Two ways to check an upgrade
Planted nulls. Before adopting an upgrade, run m ideas known to be false through the new test and flag the upgrade if the number passing the threshold exceeds the binomial critical value at the 0.05 level for level α.
Graveyard calibration. Before adopting an upgrade, compute the new test's statistic for every buried idea without certifying anything. If the graveyard is mostly false, about half of the statistics should fall below zero. We allow up to 15 percent of the graveyard to be true, so we flag the upgrade if the proportion below zero falls under 0.425 by more than 1.645 standard errors. This is an empirical null in the sense of Efron [7], applied to the lower half of the distribution, where true ideas rarely lie.
Results
The graveyard is a different population from the stream
After the first generation of tests, 4.8 percent of the buried ideas are true, against 10 percent of all ideas. The true ideas that were buried are the weak ones: their mean signal strength is 1.96, against 2.50 for all true ideas. And 45 percent of all true ideas were buried by that first test. So the graveyard is richer in false ideas than the stream, and its true members are harder to detect than average.
This matters for what a retest is worth. We took the buried ideas after the first generation and retested each once in several ways. The table gives the probability a true buried idea passes, the probability a false one passes, their ratio, and the proportion of passers that are true.
| Retest | Passes if true | Passes if false | Ratio | Proportion true among passers |
|---|---|---|---|---|
| Coach, same data, 1 variant | 0.146 | 0.0057 | 25.6 | 0.564 |
| Coach, same data, 8 variants | 0.485 | 0.034 | 14.2 | 0.418 |
| Coach, same data, 32 variants | 0.662 | 0.078 | 8.5 | 0.299 |
| Fresh data, resolution 1.0, flat level | 0.376 | 0.0100 | 37.5 | 0.655 |
| Fresh data, resolution 1.4, flat level | 0.602 | 0.0100 | 59.9 | 0.752 |
| Fresh data, resolution 2.0, flat level | 0.822 | 0.0099 | 83.0 | 0.808 |
| Fresh data, resolution 3.8, flat level | 0.995 | 0.0100 | 99.9 | 0.835 |
| Fresh data, resolution 1.4, level 0.01/6 | 0.430 | 0.0017 | 259.3 | 0.929 |
| Fresh data, resolution 2.0, level 0.01/6 | 0.697 | 0.0017 | 414.8 | 0.954 |
Three results. First, the same test at the same threshold is worth less on a buried idea than on a fresh one: the first test certified 85.9 percent true among its passers, and a retest at resolution 1.0 on the graveyard certifies 65.5 percent, even though the threshold is identical. The only difference is the lower proportion of true ideas in the pool. Second, giving analysts more variants on the same data lowers the ratio, from 25.6 to 8.5, and the proportion true among passers from 0.564 to 0.299. Tuning raises the chance a true idea passes, but raises the chance a false one passes faster. The same data contain the same information; the extra freedom only increases the number of tickets. Third, new evidence at higher resolution raises the ratio, and a stricter level multiplies it, so a retest at resolution 2.0 and level 0.01/6 certifies 95.4 percent true.
| Policy | Retests per run | False discovery proportion overall | Among exhumed only | True ideas left buried | True found by exhumation per 1000 retests | False found by exhumation per 1000 retests |
|---|---|---|---|---|---|---|
| Bury | 0 | 0.098 [0.097, 0.099] | none | 0.165 [0.164, 0.166] | none | none |
| Free retry | 7303 | 0.217 [0.216, 0.218] | 0.548 | 0.015 | 8.2 | 9.9 |
| Coach | 7188 | 0.347 [0.346, 0.348] | 0.795 | 0.059 | 5.9 | 22.7 |
| Sweep, flat | 7151 | 0.211 [0.210, 0.212] | 0.518 | 0.001 | 9.2 | 9.9 |
| Sweep, spending | 7220 | 0.102 [0.102, 0.103] | 0.125 | 0.004 | 8.9 | 1.3 |
| Sweep, batch rule | 7213 | 0.108 [0.107, 0.109] | 0.157 | 0.004 | 8.9 | 1.7 |
Permanent burial leaves 16.5 percent of all true ideas buried at the end. This is the cost of the rule, and it is large even though the test improves, because early ideas were judged by weak tests. Retry on request recovers most of them and doubles the false discovery proportion, from 0.098 to 0.217, and more than half of what it exhumes is false. Flat sweeps do the same, because after four upgrades a false idea has been through the flat test five times. Coaching is worst on false discoveries and on yield per retest: a third of all certifications are false, and four fifths of what it exhumes.
The spending sweep recovers true ideas at the same rate per retest as the flat one and the retry (about 9 per 1000) but finds about one eighth as many false ones. Overall false discovery proportion rises by only 0.004 over burial, from 0.098 to 0.102. The batch rule is close behind. We read the result as follows: the number of true recoveries is set by how much new evidence the retests carry, and the number of false ones by how many times each null is exposed to a test, and a rule that spends its error budget across exposures separates the two.
Sensitivity
We varied the setting one factor at a time, with five hundred runs per row.
| Setting | Bury: true buried | Spending sweep: true buried | Bury: false discovery | Spending: false discovery | Spending, exhumed only | Flat sweep, exhumed only | Coach, exhumed only |
|---|---|---|---|---|---|---|---|
| Base | 0.165 | 0.004 | 0.098 | 0.102 | 0.125 | 0.518 | 0.796 |
| True fraction 0.02 | 0.163 | 0.004 | 0.369 | 0.381 | 0.439 | 0.855 | 0.955 |
| True fraction 0.05 | 0.163 | 0.004 | 0.186 | 0.194 | 0.233 | 0.696 | 0.893 |
| True fraction 0.30 | 0.165 | 0.004 | 0.027 | 0.029 | 0.036 | 0.218 | 0.503 |
| Upgrade factor 1.0 | 0.449 | 0.336 | 0.141 | 0.146 | 0.168 | 0.457 | 0.667 |
| Upgrade factor 1.1 | 0.330 | 0.190 | 0.119 | 0.123 | 0.140 | 0.450 | 0.707 |
| Upgrade factor 1.8 | 0.118 | 0.000 | 0.093 | 0.102 | 0.163 | 0.599 | 0.844 |
| Maximum signal 2.5 | 0.284 | 0.008 | 0.112 | 0.102 | 0.077 | 0.383 | 0.705 |
| Maximum signal 6.0 | 0.100 | 0.002 | 0.091 | 0.102 | 0.190 | 0.640 | 0.865 |
| Level 0.002 | 0.240 | 0.013 | 0.023 | 0.022 | 0.019 | 0.132 | 0.472 |
| Level 0.05 | 0.089 | 0.001 | 0.332 | 0.360 | 0.552 | 0.901 | 0.949 |
| No upgrades | 0.449 | 0.449 | 0.141 | 0.141 | none | none | 0.667 |
The ordering of the rules is the same in every row: spending is no worse than flat, flat is no worse than coach, on the exhumed false discovery proportion. Two cautions come from the table. When true ideas are rare (a true fraction of 0.02), the spending sweep keeps the false discovery of the whole programme near that of burial but 44 percent of what it exhumes is false, so a spending rule bounds the number of false exhumations, not the quality of each exhumed set. And when the upgrade factor is 1.0, so that retests bring fresh data of no greater quality, the spending sweep recovers 11 of the 45 points of true ideas that burial leaves buried, a quarter of them. At a factor of 1.1 it recovers 42 percent and at 1.4 it recovers 98 percent. New evidence of the same quality is worth something, and improved quality is worth much more.
Flawed upgrades and the graveyard as a calibration set
Now half of the upgrades are flawed. Without any check, a flawed upgrade of size δ = 0.8 is adopted in every case and the programme's false discovery proportion with a spending sweep rises to 0.285 from 0.102, with 45.6 percent of exhumed ideas false. The table gives the detection rate on flawed upgrades, the false alarm rate on valid ones, and the outcome.
| Shift | Check | Flawed upgrades caught | Valid upgrades flagged | False discovery proportion | True ideas left buried |
|---|---|---|---|---|---|
| 0.8 | None | 0.00 | 0.00 | 0.285 | 0.001 |
| 0.8 | 50 planted nulls | 0.64 | 0.015 | 0.210 | 0.056 |
| 0.8 | 100 planted nulls | 0.88 | 0.022 | 0.155 | 0.103 |
| 0.8 | 200 planted nulls | 0.99 | 0.018 | 0.120 | 0.129 |
| 0.8 | Graveyard | 1.00 | 0.00 | 0.115 | 0.129 |
| 0.4 | None | 0.00 | 0.00 | 0.166 | 0.002 |
| 0.4 | 200 planted nulls | 0.48 | 0.018 | 0.150 | 0.042 |
| 0.4 | Graveyard | 1.00 | 0.00 | 0.115 | 0.129 |
The graveyard check, which costs nothing because the graveyard already exists and holds hundreds of ideas by the first upgrade, caught every flawed upgrade in these runs, at both shifts, and flagged no valid one. Detection has a price: a discarded upgrade is an improvement not made, and the proportion of true ideas left buried rises from 0.001 to 0.129 because half of the upgrades were thrown away. This is the real cost of the check in this setting, not a flaw in it.
How good is the check in general? We computed detection directly, with five thousand runs per cell, for a shift in the statistics and for a flaw that inflates their variance.
| Flaw | Planted nulls, 100 (tail test) | Planted nulls, 100 (central test) | Graveyard, 100 ideas | Graveyard, 300 ideas |
|---|---|---|---|---|
| None | 0.01 | 0.04 | 0.01 | 0.00 |
| Shift 0.2 | 0.09 | 0.46 | 0.14 | 0.26 |
| Shift 0.4 | 0.30 | 0.93 | 0.68 | 0.98 |
| Shift 0.8 | 0.88 | 1.00 | 1.00 | 1.00 |
| Standard deviation 1.3 | 0.51 | not applicable | 0.00 | 0.00 |
| Standard deviation 1.6 | 0.94 | not applicable | 0.00 | 0.00 |
Two things follow. First, a hundred planted nulls tested only at the tail the gate uses are much less sensitive to a shift than a hundred graveyard ideas tested at the centre of the distribution, because the tail sees few events. A central test on planted nulls dominates the graveyard, as it should: planted nulls are known to be false, and the graveyard is contaminated by true ideas, which we allow for by loosening the criterion. Second, the graveyard check is blind to a flaw that widens the distribution without moving its centre, which the planted nulls catch. The two are complements. The graveyard is free and can be run before any exhumation; planted nulls cost effort to build and cover failure modes the graveyard cannot see.
Falsification and limits
We tried to break the headline in four ways.
The null of no improvement. With no upgrades at all, the spending sweep never runs and its outcomes are identical to burial's, as they must be. With an upgrade factor of 1.0, the sweep recovers a quarter of what burial loses, against 98 percent at 1.4. What helps is evidence, and the label upgrade adds nothing in itself.
Rare true ideas. When only two percent of ideas are true, the spending rule's exhumed set is 44 percent false, and the flat sweep's is 86 percent. The rule limits the count of false exhumations to a budget; it does not make the exhumed set trustworthy when the prior is low. Anyone using a sweep at a low base rate should expect most of what it exhumes to be false unless the batch rule is tightened.
A check that discards good work. In the flaw experiments the graveyard check threw away half of all upgrades and the programme paid for it in true ideas left buried. If flaws are rarer, the cost is smaller, but we did not vary the flaw probability.
Assumptions we did not test. The statistics are exactly standard normal under the null, which is what makes both checks work; real tests have unknown, sometimes heavy-tailed nulls, and the graveyard check depends on the proportion of statistics below zero being near one half. Ideas are independent; correlated ideas would make the false discovery proportion more variable and the empirical null less reliable. The retest evidence is genuinely fresh and independent of the first test's noise; where the evidence is a single sealed sample that cannot be replenished, no sweep is possible and burial is the only honest option. The spending schedule, the level of the batch rule, the number of variants and the 15 percent allowance for true ideas in the graveyard were each fixed in advance and none was tuned; other values would move the numbers. The model has no costs for the testing itself beyond counting retests, and no opportunity cost for time. We also did not model the human side: who decides that an upgrade has occurred is exactly where the hard cases arise, and a procedure that depends on honest declaration of what counts as improvement is only as strong as the declaration.
What this means in practice
Write down the rule for reopening a rejected idea before the first idea is rejected. Allow reopening when the evidence standard has improved by a declared amount, apply the improved standard to every affected rejection at once, and give each exposure a share of an error budget that was fixed in advance. Do not allow one party to reopen one idea. In our simulation the difference between this and open reopening was a doubling of false discoveries, for the same number of retests.
Do not treat a retest of old data as a retest. A new analysis of evidence that has already judged an idea is worth less than an equal number of tests of fresh ideas, and the more freedom the analyst has, the worse the exchange.
Test the upgrade before using it. Run the proposed test on the whole graveyard first and inspect the centre of its distribution; if far fewer than half the statistics are negative, something is wrong, and no idea should be certified under it. Add a smaller set of known-false cases to catch what the graveyard cannot see, in particular a change in spread.
Expect the sweep to cost something and say so. Burial's loss of true findings is a cost that is invisible because nobody counts the ideas that never came back. The exhumation rule makes it visible and recovers most of it, at the price of retests and of discarding upgrades that fail their check.
Reproducibility appendix
Python with numpy and scipy. Seeds: main comparison 11, sensitivity 21, flaw study 33, retest value 12, detection curves 9 and 10. Runtimes on a ten-core laptop: main comparison about 25 seconds for six policies with one thousand runs each, sensitivity about four minutes, retest-value analysis eighteen seconds, flaw study about fifteen seconds, detection curves six seconds.
The state update for one policy-run is the core of the simulation. The sweep step, with the spending rule, and the two upgrade checks are:
def sweep_spend(state, bgen, g, lam, r, delta, alpha, rng):
elig = (state == 2) & (bgen < g[:, None])
z = lam * r[:, None] + delta[:, None] + rng.standard_normal(lam.shape)
level = alpha / (np.maximum(g, 1) * (np.maximum(g, 1) + 1))
ok = elig & (z > norm.isf(level)[:, None])
return np.where(ok, 1, state), ok, elig.sum(axis=1)
def check_planted(delta_new, m, alpha, rng, level=0.05):
z = rng.standard_normal((delta_new.size, m)) + delta_new[:, None]
za = norm.isf(alpha)
crit = binom.isf(level, m, alpha) + 1
return (z > za).sum(axis=1) >= crit
def check_graveyard(lam, r_new, delta_new, state, bgen, g, rng, pi_cap=0.15):
elig = (state == 2) & (bgen < (g + 1)[:, None])
z = lam * r_new[:, None] + delta_new[:, None] + rng.standard_normal(lam.shape)
n = elig.sum(axis=1)
below = ((z < 0) & elig).sum(axis=1) / np.maximum(n, 1)
p0 = 0.5 * (1 - pi_cap)
se = np.sqrt(p0 * (1 - p0) / np.maximum(n, 1))
return (n > 20) & (below < p0 - 1.645 * se)The coaching retest draws variants that share a fraction of the original noise:
def coach_pass(lam, r_at_burial, e0, rho_var, m, za, rng):
shape = lam.shape + (m,)
zv = (lam * r_at_burial)[..., None] + np.sqrt(rho_var) * e0[..., None] \
+ np.sqrt(1 - rho_var) * rng.standard_normal(shape)
return zv.max(axis=-1) > zaReferences
- Foster, D. P., and Stine, R. A. (2008). Alpha-investing: a procedure for sequential control of expected false discoveries. Journal of the Royal Statistical Society, Series B, 70(2), 429-444. doi:10.1111/j.1467-9868.2007.00643.x
- Javanmard, A., and Montanari, A. (2018). Online rules for control of false discovery rate and false discovery exceedance. The Annals of Statistics, 46(2). doi:10.1214/17-AOS1559
- Benjamini, Y., and Hochberg, Y. (1995). Controlling the false discovery rate: a practical and powerful approach to multiple testing. Journal of the Royal Statistical Society, Series B, 57(1), 289-300. doi:10.1111/j.2517-6161.1995.tb02031.x
- Dwork, C., Feldman, V., Hardt, M., Pitassi, T., Reingold, O., and Roth, A. (2015). The reusable holdout: preserving validity in adaptive data analysis. Science, 349(6248), 636-638. doi:10.1126/science.aaa9375
- Simmons, J. P., Nelson, L. D., and Simonsohn, U. (2011). False-positive psychology: undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359-1366. doi:10.1177/0956797611417632
- Ioannidis, J. P. A. (2005). Why most published research findings are false. PLoS Medicine, 2(8), e124. doi:10.1371/journal.pmed.0020124
- Efron, B. (2004). Large-scale simultaneous hypothesis testing: the choice of a null hypothesis. Journal of the American Statistical Association, 99(465), 96-104. doi:10.1198/016214504000000089
- Boyce, Prystawski, and colleagues (2024). Estimating the replicability of psychology experiments after an initial failure to replicate. Collabra: Psychology, 10(1). doi:10.1525/collabra.125685
Cite this paper
@misc{forticia2026never,
title = {{Never Tune a Corpse}},
author = {{Forticia Research Institute}},
year = {2026},
month = oct,
publisher = {Forticia Research Institute},
howpublished = {\url{https://www.forticia.uk/papers/exhumation-policies-rejected-ideas}},
note = {Paper, published online}
}Generated from this page’s metadata. Forticia does not assign DOIs to these papers.