Skip to content
Forticia
Research
Quantitative FinanceEquities, FX, futures. Factors and backtests, every run replayable.Computational BiologySequence, folding, simulation. Versioned, reproducible labs.Cultural IntelligencePhilosophy, governance, ethics. How institutions decide and answer for it.AI InstrumentationPrivate models. Multi-agent orchestration. Guardrails on write.View all research
Quantitative Finance
  • Equities, FX, futures
  • Factors and backtests
  • Every run logged and replayable
InfrastructurePapersPolarisLink™About
Sign inRequest access
Request access

Research

Quantitative FinanceEquities, FX, futures. Factors and backtests, every run replayable.Computational BiologySequence, folding, simulation. Versioned, reproducible labs.Cultural IntelligencePhilosophy, governance, ethics. How institutions decide and answer for it.AI InstrumentationPrivate models. Multi-agent orchestration. Guardrails on write.

Platform

InfrastructurePapersPolarisLink™About
Request accessSign in
Forticia

A private institute for computational research. It publishes original research and runs a governed environment where every run is logged and replayable.

Sign inSystem status

Research

Quantitative FinanceComputational BiologyCultural IntelligenceAI Instrumentation

Platform

InfrastructurePapersPolarisLink™StatusSign in

Institute

AboutRequest accessContactGitHubPolarisLink repository

Forticia publishes research and simulations. Nothing on this site is investment advice.

© 2026 ForticiaPrivacyTerms
Papers/Cultural Intelligence
CulturePaper

Never Tune a Corpse

Author
Forticia Research Institute
Published
5 October 2026
Last updated
5 October 2026
Reading time
16 min
Cite this paper

On this page 0%

  1. Abstract
  2. Motivation
  3. Related work and what is new
  4. Method
  5. The world
  6. Policies
  7. Flawed upgrades
  8. Two ways to check an upgrade
  9. Results
  10. The graveyard is a different population from the stream
  11. Policy comparison at matched budgets
  12. Sensitivity
  13. Flawed upgrades and the graveyard as a calibration set
  14. Falsification and limits
  15. What this means in practice
  16. Reproducibility appendix
  17. References
  18. Cite this paper

Abstract

A research institution that tests many ideas and rejects most of them must decide when a rejected idea may be tested again. We model a stream of ideas, an evidence standard that improves in steps, and five policies: bury forever, retry on request, retry the same data with analyst freedom, sweep the graveyard after each improvement under a spending rule, and sweep under a batch false discovery rate rule. At matched retest budgets, retry on request doubles the false discovery proportion from 0.098 to 0.217 and tuning the same data more than triples it, while a spending sweep recovers all but 0.4 percent of true ideas at 0.102. Improvements can be flawed. The graveyard, being mostly false ideas, works as a free calibration set that caught every flawed improvement we simulated, but it is blind to variance flaws.

Motivation

A group runs a test on an idea and the idea fails. What then? Two answers are common and both are wrong in a measurable way.

The first is to bury it permanently. This is clean, and it is how a group guards itself against the pleasure of rescuing its favourites. But the test that killed the idea had a certain power, and an idea with a modest real effect is more likely than not to fail a weak test. If the evidence standard later improves, a permanent grave is a standing loss of true findings.

The second is to let anyone reopen anything at any time. This feels open-minded. It also gives every false idea another lottery ticket each time someone asks, and the people who ask are the people who believe in the idea.

There is a third move that looks like the second and is worse: keep the data, vary the analysis, and run the test again on the same evidence until it passes. The rule this paper tests is that a rejected idea may be reopened only when the court itself improves, that is, when the test becomes more able to tell true from false, and then every affected idea is retested under the same improved standard. Reopen the verdict by upgrading the procedure, not by coaching the defendant. We ask what each rule costs and buys, and what is needed to keep the upgrade itself honest.

Related work and what is new

Sequential testing has a developed theory. Foster and Stine's alpha-investing and the online procedures of Javanmard and Montanari control the false discovery rate as hypotheses arrive in a stream, each tested once and the budget replenished by discoveries [1, 2]. Benjamini and Hochberg's procedure controls it within a batch [3]. Dwork and colleagues show how a holdout set can be reused safely when analyses adapt to it [4]. Simmons, Nelson and Simonsohn quantified how analytic flexibility inflates false positives [5]. Ioannidis argued from the base rate of true hypotheses that most claimed findings are false [6]. Efron proposed estimating the null distribution empirically from the bulk of the test statistics rather than assuming it [7]. Empirically, a study that re-ran failed replications at larger sample sizes found that most still failed to replicate, so extra resources seldom rescued them [8].

What these do not address is the object we study: a population of hypotheses already tested and rejected, to be reopened at later times, under a rule that says who may reopen and when. We contribute four things. A policy comparison at matched retest budgets. A likelihood-ratio accounting of what each kind of retest is worth, which shows that the graveyard is a different population from the stream and that same-data tuning adds almost no information. A flawed-upgrade model. And a test, using the graveyard itself, for whether a proposed upgrade is valid before it is used to exhume anything. Our search was not exhaustive and the multiple-testing ingredients are standard; the claim concerns the combination.

Method

The world

Ideas arrive one hundred per period for forty periods. Each is true with probability 0.10. A true idea has a signal strength λ drawn uniformly from 1 to 4, a false one has zero. The institution tests an idea by computing a statistic z equal to λ times the current resolution r, plus standard normal noise, and certifies the idea if z exceeds the threshold for a one-sided level α = 0.01 (z above 2.326). Resolution starts at 1 and is multiplied by 1.4 at each of four upgrades, at periods 8, 16, 24 and 32. The first test therefore has an average power of 0.55, and later tests have more. Failed ideas are buried, carrying the data noise they were tested with and the generation of the test that failed them.

The units are abstract. Resolution corresponds to anything that raises the signal-to-noise ratio of a test: more data, a better instrument, a cleaner definition.

Policies

All five policies receive the same fresh ideas, the same upgrades and the same random numbers. They differ only in what happens to buried ideas.

  • Bury: nothing, ever.
  • Free retry: each period, each buried idea is retested with probability 0.10 on fresh data at the current resolution and the flat threshold. An idea may be retried any number of times.
  • Coach: each period, each buried idea is, with the same probability, retested on its original data with eight analytic variants. Each variant shares 70 percent of its noise variance with the original test and has its own independent noise for the rest. The idea is certified if any variant passes the flat threshold. This models trying alternative specifications of the same evidence.
  • Sweep with spending: when an upgrade is adopted, every buried idea from an earlier generation is retested once on fresh data at the new resolution. After the g-th adopted upgrade the threshold is for level α divided by g(g + 1), so the levels over all sweeps sum to α.
  • Sweep flat: as above but at level α every time.
  • Sweep with a batch rule: as above, but the sweep's retest p-values are passed through the Benjamini and Hochberg procedure at level 0.15 [3].

Free retry and coach have a retest probability of 0.10 chosen so that their average total number of retests matches the sweeps (about 7200 against 7150 to 7220). Sweeps retest every eligible idea at every upgrade, and the number grows as the graveyard grows.

We report the false discovery proportion as the ratio of total false certifications to total certifications, summed across runs, with 95 percent bootstrap intervals over one thousand runs. We also report the proportion of true ideas that remain buried at the end, and the same quantities for exhumed ideas alone.

Flawed upgrades

With probability one half, an upgrade has a hidden defect that adds a constant shift δ to every statistic, true or false, until the next adopted upgrade replaces it. This stands for a leak, a mis-specified filter, a contaminated reference set. A flawed upgrade raises the false positive rate. At δ = 0.8 the level 0.01 becomes about 0.063.

Two ways to check an upgrade

Planted nulls. Before adopting an upgrade, run m ideas known to be false through the new test and flag the upgrade if the number passing the threshold exceeds the binomial critical value at the 0.05 level for level α.

Graveyard calibration. Before adopting an upgrade, compute the new test's statistic for every buried idea without certifying anything. If the graveyard is mostly false, about half of the statistics should fall below zero. We allow up to 15 percent of the graveyard to be true, so we flag the upgrade if the proportion below zero falls under 0.425 by more than 1.645 standard errors. This is an empirical null in the sense of Efron [7], applied to the lower half of the distribution, where true ideas rarely lie.

Results

The graveyard is a different population from the stream

After the first generation of tests, 4.8 percent of the buried ideas are true, against 10 percent of all ideas. The true ideas that were buried are the weak ones: their mean signal strength is 1.96, against 2.50 for all true ideas. And 45 percent of all true ideas were buried by that first test. So the graveyard is richer in false ideas than the stream, and its true members are harder to detect than average.

This matters for what a retest is worth. We took the buried ideas after the first generation and retested each once in several ways. The table gives the probability a true buried idea passes, the probability a false one passes, their ratio, and the proportion of passers that are true.

Retest Passes if true Passes if false Ratio Proportion true among passers
Coach, same data, 1 variant 0.146 0.0057 25.6 0.564
Coach, same data, 8 variants 0.485 0.034 14.2 0.418
Coach, same data, 32 variants 0.662 0.078 8.5 0.299
Fresh data, resolution 1.0, flat level 0.376 0.0100 37.5 0.655
Fresh data, resolution 1.4, flat level 0.602 0.0100 59.9 0.752
Fresh data, resolution 2.0, flat level 0.822 0.0099 83.0 0.808
Fresh data, resolution 3.8, flat level 0.995 0.0100 99.9 0.835
Fresh data, resolution 1.4, level 0.01/6 0.430 0.0017 259.3 0.929
Fresh data, resolution 2.0, level 0.01/6 0.697 0.0017 414.8 0.954

Three results. First, the same test at the same threshold is worth less on a buried idea than on a fresh one: the first test certified 85.9 percent true among its passers, and a retest at resolution 1.0 on the graveyard certifies 65.5 percent, even though the threshold is identical. The only difference is the lower proportion of true ideas in the pool. Second, giving analysts more variants on the same data lowers the ratio, from 25.6 to 8.5, and the proportion true among passers from 0.564 to 0.299. Tuning raises the chance a true idea passes, but raises the chance a false one passes faster. The same data contain the same information; the extra freedom only increases the number of tickets. Third, new evidence at higher resolution raises the ratio, and a stricter level multiplies it, so a retest at resolution 2.0 and level 0.01/6 certifies 95.4 percent true.

Figure 1. What a retest is worth on the graveyard: proportion of true ideas among those that pass one retest of every idea buried after the first generation of tests; 4.8 percent of that graveyard is true. Variants share 70 percent of their noise with the original test. The number beside each bar is the ratio of the chance a true idea passes to the chance a false one passes. Dashed line: the first test on fresh ideas.

Policy comparison at matched budgets

Figure 2. Two costs of each policy for reopening rejected ideas: each policy at a matched budget of about 7,200 retests per run (none for burial), 1,000 simulated runs. 95 percent bootstrap intervals are narrower than the markers. Hover or focus a point for the retest count and the false share of what it exhumed.
Policy Retests per run False discovery proportion overall Among exhumed only True ideas left buried True found by exhumation per 1000 retests False found by exhumation per 1000 retests
Bury 0 0.098 [0.097, 0.099] none 0.165 [0.164, 0.166] none none
Free retry 7303 0.217 [0.216, 0.218] 0.548 0.015 8.2 9.9
Coach 7188 0.347 [0.346, 0.348] 0.795 0.059 5.9 22.7
Sweep, flat 7151 0.211 [0.210, 0.212] 0.518 0.001 9.2 9.9
Sweep, spending 7220 0.102 [0.102, 0.103] 0.125 0.004 8.9 1.3
Sweep, batch rule 7213 0.108 [0.107, 0.109] 0.157 0.004 8.9 1.7

Permanent burial leaves 16.5 percent of all true ideas buried at the end. This is the cost of the rule, and it is large even though the test improves, because early ideas were judged by weak tests. Retry on request recovers most of them and doubles the false discovery proportion, from 0.098 to 0.217, and more than half of what it exhumes is false. Flat sweeps do the same, because after four upgrades a false idea has been through the flat test five times. Coaching is worst on false discoveries and on yield per retest: a third of all certifications are false, and four fifths of what it exhumes.

The spending sweep recovers true ideas at the same rate per retest as the flat one and the retry (about 9 per 1000) but finds about one eighth as many false ones. Overall false discovery proportion rises by only 0.004 over burial, from 0.098 to 0.102. The batch rule is close behind. We read the result as follows: the number of true recoveries is set by how much new evidence the retests carry, and the number of false ones by how many times each null is exposed to a test, and a rule that spends its error budget across exposures separates the two.

Sensitivity

We varied the setting one factor at a time, with five hundred runs per row.

Setting Bury: true buried Spending sweep: true buried Bury: false discovery Spending: false discovery Spending, exhumed only Flat sweep, exhumed only Coach, exhumed only
Base 0.165 0.004 0.098 0.102 0.125 0.518 0.796
True fraction 0.02 0.163 0.004 0.369 0.381 0.439 0.855 0.955
True fraction 0.05 0.163 0.004 0.186 0.194 0.233 0.696 0.893
True fraction 0.30 0.165 0.004 0.027 0.029 0.036 0.218 0.503
Upgrade factor 1.0 0.449 0.336 0.141 0.146 0.168 0.457 0.667
Upgrade factor 1.1 0.330 0.190 0.119 0.123 0.140 0.450 0.707
Upgrade factor 1.8 0.118 0.000 0.093 0.102 0.163 0.599 0.844
Maximum signal 2.5 0.284 0.008 0.112 0.102 0.077 0.383 0.705
Maximum signal 6.0 0.100 0.002 0.091 0.102 0.190 0.640 0.865
Level 0.002 0.240 0.013 0.023 0.022 0.019 0.132 0.472
Level 0.05 0.089 0.001 0.332 0.360 0.552 0.901 0.949
No upgrades 0.449 0.449 0.141 0.141 none none 0.667

The ordering of the rules is the same in every row: spending is no worse than flat, flat is no worse than coach, on the exhumed false discovery proportion. Two cautions come from the table. When true ideas are rare (a true fraction of 0.02), the spending sweep keeps the false discovery of the whole programme near that of burial but 44 percent of what it exhumes is false, so a spending rule bounds the number of false exhumations, not the quality of each exhumed set. And when the upgrade factor is 1.0, so that retests bring fresh data of no greater quality, the spending sweep recovers 11 of the 45 points of true ideas that burial leaves buried, a quarter of them. At a factor of 1.1 it recovers 42 percent and at 1.4 it recovers 98 percent. New evidence of the same quality is worth something, and improved quality is worth much more.

Flawed upgrades and the graveyard as a calibration set

Now half of the upgrades are flawed. Without any check, a flawed upgrade of size δ = 0.8 is adopted in every case and the programme's false discovery proportion with a spending sweep rises to 0.285 from 0.102, with 45.6 percent of exhumed ideas false. The table gives the detection rate on flawed upgrades, the false alarm rate on valid ones, and the outcome.

Shift Check Flawed upgrades caught Valid upgrades flagged False discovery proportion True ideas left buried
0.8 None 0.00 0.00 0.285 0.001
0.8 50 planted nulls 0.64 0.015 0.210 0.056
0.8 100 planted nulls 0.88 0.022 0.155 0.103
0.8 200 planted nulls 0.99 0.018 0.120 0.129
0.8 Graveyard 1.00 0.00 0.115 0.129
0.4 None 0.00 0.00 0.166 0.002
0.4 200 planted nulls 0.48 0.018 0.150 0.042
0.4 Graveyard 1.00 0.00 0.115 0.129

The graveyard check, which costs nothing because the graveyard already exists and holds hundreds of ideas by the first upgrade, caught every flawed upgrade in these runs, at both shifts, and flagged no valid one. Detection has a price: a discarded upgrade is an improvement not made, and the proportion of true ideas left buried rises from 0.001 to 0.129 because half of the upgrades were thrown away. This is the real cost of the check in this setting, not a flaw in it.

How good is the check in general? We computed detection directly, with five thousand runs per cell, for a shift in the statistics and for a flaw that inflates their variance.

Flaw Planted nulls, 100 (tail test) Planted nulls, 100 (central test) Graveyard, 100 ideas Graveyard, 300 ideas
None 0.01 0.04 0.01 0.00
Shift 0.2 0.09 0.46 0.14 0.26
Shift 0.4 0.30 0.93 0.68 0.98
Shift 0.8 0.88 1.00 1.00 1.00
Standard deviation 1.3 0.51 not applicable 0.00 0.00
Standard deviation 1.6 0.94 not applicable 0.00 0.00
Figure 3. Which flawed upgrades each check catches: probability that a proposed upgrade is flagged, 5,000 simulated runs per point. The central test applies only to planted nulls and to shift flaws, so it has no curve on the right.

Two things follow. First, a hundred planted nulls tested only at the tail the gate uses are much less sensitive to a shift than a hundred graveyard ideas tested at the centre of the distribution, because the tail sees few events. A central test on planted nulls dominates the graveyard, as it should: planted nulls are known to be false, and the graveyard is contaminated by true ideas, which we allow for by loosening the criterion. Second, the graveyard check is blind to a flaw that widens the distribution without moving its centre, which the planted nulls catch. The two are complements. The graveyard is free and can be run before any exhumation; planted nulls cost effort to build and cover failure modes the graveyard cannot see.

Falsification and limits

We tried to break the headline in four ways.

The null of no improvement. With no upgrades at all, the spending sweep never runs and its outcomes are identical to burial's, as they must be. With an upgrade factor of 1.0, the sweep recovers a quarter of what burial loses, against 98 percent at 1.4. What helps is evidence, and the label upgrade adds nothing in itself.

Rare true ideas. When only two percent of ideas are true, the spending rule's exhumed set is 44 percent false, and the flat sweep's is 86 percent. The rule limits the count of false exhumations to a budget; it does not make the exhumed set trustworthy when the prior is low. Anyone using a sweep at a low base rate should expect most of what it exhumes to be false unless the batch rule is tightened.

A check that discards good work. In the flaw experiments the graveyard check threw away half of all upgrades and the programme paid for it in true ideas left buried. If flaws are rarer, the cost is smaller, but we did not vary the flaw probability.

Assumptions we did not test. The statistics are exactly standard normal under the null, which is what makes both checks work; real tests have unknown, sometimes heavy-tailed nulls, and the graveyard check depends on the proportion of statistics below zero being near one half. Ideas are independent; correlated ideas would make the false discovery proportion more variable and the empirical null less reliable. The retest evidence is genuinely fresh and independent of the first test's noise; where the evidence is a single sealed sample that cannot be replenished, no sweep is possible and burial is the only honest option. The spending schedule, the level of the batch rule, the number of variants and the 15 percent allowance for true ideas in the graveyard were each fixed in advance and none was tuned; other values would move the numbers. The model has no costs for the testing itself beyond counting retests, and no opportunity cost for time. We also did not model the human side: who decides that an upgrade has occurred is exactly where the hard cases arise, and a procedure that depends on honest declaration of what counts as improvement is only as strong as the declaration.

What this means in practice

Write down the rule for reopening a rejected idea before the first idea is rejected. Allow reopening when the evidence standard has improved by a declared amount, apply the improved standard to every affected rejection at once, and give each exposure a share of an error budget that was fixed in advance. Do not allow one party to reopen one idea. In our simulation the difference between this and open reopening was a doubling of false discoveries, for the same number of retests.

Do not treat a retest of old data as a retest. A new analysis of evidence that has already judged an idea is worth less than an equal number of tests of fresh ideas, and the more freedom the analyst has, the worse the exchange.

Test the upgrade before using it. Run the proposed test on the whole graveyard first and inspect the centre of its distribution; if far fewer than half the statistics are negative, something is wrong, and no idea should be certified under it. Add a smaller set of known-false cases to catch what the graveyard cannot see, in particular a change in spread.

Expect the sweep to cost something and say so. Burial's loss of true findings is a cost that is invisible because nobody counts the ideas that never came back. The exhumation rule makes it visible and recovers most of it, at the price of retests and of discarding upgrades that fail their check.

Reproducibility appendix

Python with numpy and scipy. Seeds: main comparison 11, sensitivity 21, flaw study 33, retest value 12, detection curves 9 and 10. Runtimes on a ten-core laptop: main comparison about 25 seconds for six policies with one thousand runs each, sensitivity about four minutes, retest-value analysis eighteen seconds, flaw study about fifteen seconds, detection curves six seconds.

The state update for one policy-run is the core of the simulation. The sweep step, with the spending rule, and the two upgrade checks are:

python
def sweep_spend(state, bgen, g, lam, r, delta, alpha, rng):
    elig = (state == 2) & (bgen < g[:, None])
    z = lam * r[:, None] + delta[:, None] + rng.standard_normal(lam.shape)
    level = alpha / (np.maximum(g, 1) * (np.maximum(g, 1) + 1))
    ok = elig & (z > norm.isf(level)[:, None])
    return np.where(ok, 1, state), ok, elig.sum(axis=1)

def check_planted(delta_new, m, alpha, rng, level=0.05):
    z = rng.standard_normal((delta_new.size, m)) + delta_new[:, None]
    za = norm.isf(alpha)
    crit = binom.isf(level, m, alpha) + 1
    return (z > za).sum(axis=1) >= crit

def check_graveyard(lam, r_new, delta_new, state, bgen, g, rng, pi_cap=0.15):
    elig = (state == 2) & (bgen < (g + 1)[:, None])
    z = lam * r_new[:, None] + delta_new[:, None] + rng.standard_normal(lam.shape)
    n = elig.sum(axis=1)
    below = ((z < 0) & elig).sum(axis=1) / np.maximum(n, 1)
    p0 = 0.5 * (1 - pi_cap)
    se = np.sqrt(p0 * (1 - p0) / np.maximum(n, 1))
    return (n > 20) & (below < p0 - 1.645 * se)

The coaching retest draws variants that share a fraction of the original noise:

python
def coach_pass(lam, r_at_burial, e0, rho_var, m, za, rng):
    shape = lam.shape + (m,)
    zv = (lam * r_at_burial)[..., None] + np.sqrt(rho_var) * e0[..., None] \
         + np.sqrt(1 - rho_var) * rng.standard_normal(shape)
    return zv.max(axis=-1) > za
  • epistemics
  • research governance
  • multiple testing
  • replication
  • simulation
  • falsification

References

  1. Foster, D. P., and Stine, R. A. (2008). Alpha-investing: a procedure for sequential control of expected false discoveries. Journal of the Royal Statistical Society, Series B, 70(2), 429-444. doi:10.1111/j.1467-9868.2007.00643.x
  2. Javanmard, A., and Montanari, A. (2018). Online rules for control of false discovery rate and false discovery exceedance. The Annals of Statistics, 46(2). doi:10.1214/17-AOS1559
  3. Benjamini, Y., and Hochberg, Y. (1995). Controlling the false discovery rate: a practical and powerful approach to multiple testing. Journal of the Royal Statistical Society, Series B, 57(1), 289-300. doi:10.1111/j.2517-6161.1995.tb02031.x
  4. Dwork, C., Feldman, V., Hardt, M., Pitassi, T., Reingold, O., and Roth, A. (2015). The reusable holdout: preserving validity in adaptive data analysis. Science, 349(6248), 636-638. doi:10.1126/science.aaa9375
  5. Simmons, J. P., Nelson, L. D., and Simonsohn, U. (2011). False-positive psychology: undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359-1366. doi:10.1177/0956797611417632
  6. Ioannidis, J. P. A. (2005). Why most published research findings are false. PLoS Medicine, 2(8), e124. doi:10.1371/journal.pmed.0020124
  7. Efron, B. (2004). Large-scale simultaneous hypothesis testing: the choice of a null hypothesis. Journal of the American Statistical Association, 99(465), 96-104. doi:10.1198/016214504000000089
  8. Boyce, Prystawski, and colleagues (2024). Estimating the replicability of psychology experiments after an initial failure to replicate. Collabra: Psychology, 10(1). doi:10.1525/collabra.125685

Cite this paper

@misc{forticia2026never,
  title        = {{Never Tune a Corpse}},
  author       = {{Forticia Research Institute}},
  year         = {2026},
  month        = oct,
  publisher    = {Forticia Research Institute},
  howpublished = {\url{https://www.forticia.uk/papers/exhumation-policies-rejected-ideas}},
  note         = {Paper, published online}
}

Generated from this page’s metadata. Forticia does not assign DOIs to these papers.

NewerWhen a failure record lies: the economics of shared negative results in a research swarmOlderWhen Approval Stops Carrying Information
All papers
0%25%50%75%100%Proportion true among passersFirst test on fresh ideasSame data, analytic variantsFresh data, level 0.01Fresh data, level 0.01 divided by 61 variant56.4%ratio 25.62 variants52.9%ratio 22.24 variants48.1%ratio 18.38 variants41.8%ratio 14.216 variants35.3%ratio 10.832 variants29.9%ratio 8.5Resolution 1.065.5%ratio 37.5Resolution 1.475.2%ratio 59.9Resolution 2.080.8%ratio 83.0Resolution 2.782.6%ratio 94.1Resolution 3.883.5%ratio 99.9Resolution 1.086.3%ratio 124.9Resolution 1.492.9%ratio 259.3Resolution 2.095.4%ratio 414.8Resolution 2.796.4%ratio 529.6Resolution 3.896.6%ratio 566.5
Bar chart of the proportion of true ideas among those that pass a single retest of every idea buried after the first generation of tests, 4.8 percent of which are true. Retests on the same data with 1 to 32 analytic variants fall from 56 to 30 percent true. Retests on fresh data at a flat level rise from 65 to 83 percent as resolution grows, and at a stricter level reach 86 to 97 percent. A reference line marks 85.9 percent, the proportion true among passers of the first test.
ItemValueLowerUpperNote
1 variant56%ratio 25.6
2 variants53%ratio 22.2
4 variants48%ratio 18.3
8 variants42%ratio 14.2
16 variants35%ratio 10.8
32 variants30%ratio 8.5
Resolution 1.065%ratio 37.5
Resolution 1.475%ratio 59.9
Resolution 2.081%ratio 83.0
Resolution 2.783%ratio 94.1
Resolution 3.883%ratio 99.9
Resolution 1.086%ratio 124.9
Resolution 1.493%ratio 259.3
Resolution 2.095%ratio 414.8
Resolution 2.796%ratio 529.6
Resolution 3.897%ratio 566.5
0%10%20%30%40%False discovery proportion0%5%10%15%20%True ideas still buried at the endLower left is betterBuryFree retryCoachSweep, flatSweep, spendingSweep, batch rule
Point
Hover or focus a point to read its values
Scatter plot of six policies for reopening rejected ideas. Horizontal axis: share of true ideas still buried at the end. Vertical axis: false certifications as a share of all certifications. Bury sits at 16.5 percent buried and 9.8 percent false. Free retry, coach and the flat sweep have 21 to 35 percent false. The spending sweep and the batch-rule sweep sit lowest left, at about 10 to 11 percent false and under half a percent buried.
ItemTrue ideas still buried at the endFalse discovery proportionDetail
Bury17%10%No retests
Free retry2%22%7,303 retests per run; 54.8% of exhumed ideas are false
Coach6%35%7,188 retests per run; 79.5% of exhumed ideas are false
Sweep, flat0%21%7,151 retests per run; 51.8% of exhumed ideas are false
Sweep, spending0%10%7,220 retests per run; 12.5% of exhumed ideas are false
Sweep, batch rule0%11%7,213 retests per run; 15.7% of exhumed ideas are false
A flaw that shifts every statistic
0%25%50%75%100%0.00.20.40.60.8Shift added
A flaw that widens the spread
0%25%50%75%100%1.01.31.6Standard deviation multiplierGraveyard: under 0.5%
  • Planted nulls, 100, tail test
  • Planted nulls, 100, central test
  • Graveyard, 100 ideas
  • Graveyard, 300 ideas
Two line charts of the probability that a check flags a proposed upgrade as flawed. Left: a flaw that shifts every statistic by 0 to 0.8. The graveyard check at 300 ideas and the central planted-null test reach near certain detection at a shift of 0.4, while the tail test on 100 planted nulls needs a shift of 0.8. Right: a flaw that widens the spread by a factor of 1.3 or 1.6. The tail test on planted nulls flags 51 and 94 percent, while both graveyard checks flag under half a percent.
PanelSeriesPoints
A flaw that shifts every statisticPlanted nulls, 100, tail test0.0: 1%; 0.1: 4%; 0.2: 9%; 0.3: 16%; 0.4: 30%; 0.6: 60%; 0.8: 88%
A flaw that shifts every statisticPlanted nulls, 100, central test0.0: 4%; 0.2: 46%; 0.4: 93%; 0.8: 100%
A flaw that shifts every statisticGraveyard, 100 ideas0.0: 1%; 0.1: 3%; 0.2: 14%; 0.3: 38%; 0.4: 68%; 0.6: 98%; 0.8: 100%
A flaw that shifts every statisticGraveyard, 300 ideas0.0: 0%; 0.1: 3%; 0.2: 26%; 0.3: 75%; 0.4: 98%; 0.6: 100%; 0.8: 100%
A flaw that widens the spreadPlanted nulls, 100, tail test1.0: 1%; 1.3: 51%; 1.6: 94%
A flaw that widens the spreadGraveyard, 100 ideas1.0: 1%; 1.3: 0%; 1.6: 0%
A flaw that widens the spreadGraveyard, 300 ideas1.0: 0%; 1.3: 0%; 1.6: 0%