Skip to content
Forticia
Research
Quantitative FinanceEquities, FX, futures. Factors and backtests, every run replayable.Computational BiologySequence, folding, simulation. Versioned, reproducible labs.Cultural IntelligencePhilosophy, governance, ethics. How institutions decide and answer for it.AI InstrumentationPrivate models. Multi-agent orchestration. Guardrails on write.View all research
Quantitative Finance
  • Equities, FX, futures
  • Factors and backtests
  • Every run logged and replayable
InfrastructurePapersPolarisLink™About
Sign inRequest access
Request access

Research

Quantitative FinanceEquities, FX, futures. Factors and backtests, every run replayable.Computational BiologySequence, folding, simulation. Versioned, reproducible labs.Cultural IntelligencePhilosophy, governance, ethics. How institutions decide and answer for it.AI InstrumentationPrivate models. Multi-agent orchestration. Guardrails on write.

Platform

InfrastructurePapersPolarisLink™About
Request accessSign in
Forticia

A private institute for computational research. It publishes original research and runs a governed environment where every run is logged and replayable.

Sign inSystem status

Research

Quantitative FinanceComputational BiologyCultural IntelligenceAI Instrumentation

Platform

InfrastructurePapersPolarisLink™StatusSign in

Institute

AboutRequest accessContactGitHubPolarisLink repository

Forticia publishes research and simulations. Nothing on this site is investment advice.

© 2026 ForticiaPrivacyTerms
Papers/Quantitative Finance
QuantPaper

When a failure record lies: the economics of shared negative results in a research swarm

Author
Cayden Richards
Published
5 October 2026
Last updated
5 October 2026
Reading time
17 min
Cite this paper

On this page 0%

  1. Abstract
  2. Motivation
  3. Related work and what is new
  4. Model
  5. The idea space and the truth
  6. Tests
  7. Researchers
  8. Policies
  9. Measures
  10. Results
  11. Sharing exact records is an economics result about herding
  12. Reaching into neighbours buys speed and costs recall
  13. What the stamp does
  14. When to test again
  15. Falsification and limits
  16. What this means in practice
  17. Reproducibility appendix
  18. References
  19. Cite this paper

Abstract

We simulate sixteen researchers searching 6,400 related ideas with two-stage tests of known power, and ask when a shared record of falsified ideas pays. Recording exact failures lets the swarm reach in about 5,300 compute units what an unshared swarm needs 8,000 for, and the saving grows with how much researchers herd onto the same ideas. Letting a record close its neighbours buys speed early and costs recall later: a strict rule with radius 2.5 closes 85% of the true ideas untested and idles the swarm after 6% of its budget. Stamping each record with the power of its test separates a recoverable graveyard from one that closes off real ideas. Under a work cutoff it keeps 69.5 true ideas where the unstamped version keeps 15.7.

Motivation

Two researchers, a week apart, test the same intuitive idea on the same history. Both record a failure. The register that should have stopped the second one says only that the idea failed. It does not say how hard it was tested, and a test with a 45% chance of missing a real effect fails often enough to fill a graveyard with live ideas.

Sharing falsified results is cheap to say and expensive to get right. A shared record saves the compute that would be spent re-testing the same cell, which is the argument usually made for it. But a record is also a prior: researchers read it and avoid the idea, its neighbours, and the region around it. If the original test was noisy, that prior is wrong in a way no one can see, because the idea is never tested again. A false positive gets re-tested by everyone who hears of it and is corrected. A false negative is quietly final.

We ask when the graveyard pays. The question has parts: how much compute a record saves, how that depends on how much researchers crowd onto the same ideas, how far a record should reach into related ideas, what happens when the records stop a real effect from being found, and what a record needs to carry so that it can be trusted. We answer them in a simulation of a research swarm searching a shared idea space, with every number produced by the code in the appendix. The simulation is of a process. Nothing in it is a result about markets, and no real research output is used.

Related work and what is new

Publication bias and the file drawer are old topics [5]. Models of the community-level effect include replication dynamics, where communicating failed replications can support discovery even when they have low power and where suppressing negative novel findings may sometimes help [1], and a Markov model in which a claim is canonised unless enough negative results are published [2]. For autonomous agents, a shared bank of structured failure records has been shown to cut tokens and raise pass rates on real tasks [3]. Fiedler and colleagues argue that false negatives are often the more serious problem, since no rigor applied to the hypotheses that were tested can repair the ones that never were [4].

What none of these model, and what we do, is the economics inside one organisation: a fixed compute budget, researchers who herd onto the same ideas, an idea space in which related ideas share truth, tests with known and unequal power, and one underlying history that makes a repeated test of the same idea only partly independent. We measure the compute a record saves, define and measure the damage a record does to ideas it never tested, and test one design change: stamping each record with the power of the test that produced it and letting that, not a fixed rule, set how far the record reaches.

Model

The idea space and the truth

Ideas are cells of an 80 by 80 torus, 6,400 in all. Related ideas are neighbours. A fraction of 3% of cells are true, and the truth is spatially clustered: we smooth white noise with a Gaussian of standard deviation ℓ cells and keep the top 3%. At ℓ = 3 this gives about 192 true cells in blobs; at ℓ = 0 the true cells are scattered independently. Clustering is the assumption that related ideas share truth, and it is the only reason a record about one idea can say anything about another. We vary it, down to nothing.

Tests

A test is two-stage and costs compute. A screen costs 1 unit and passes a true cell with probability 0.55 and a false cell with probability 0.10. A cell that passes the screen goes to confirmation at 10 units, which passes a true cell with probability 0.90 and a false cell with probability 0.01. A cell that passes both is a discovery. A cell that fails either stage gets a failure record. Both stages are z-tests; the signal strengths follow from the stated power and size. The noise in each test has two parts. A fixed per-cell component, with variance share ρ_d, is drawn once and shared by every test of that cell: it is the one history. The remainder is fresh. At ρ_d = 1 a repeated test returns the same answer and teaches nothing; at ρ_d = 0 repeated tests are independent.

Researchers

Sixteen researchers work in rounds. In each round each picks one cell, tests it, and the results are shared at the end of the round. Researchers are attracted to ideas by a fixed salience, lognormal with log standard deviation h, unrelated to the truth. At h = 0 they pick uniformly; at h = 1 the most salient 10% of cells draw about 39% of the attention; larger h means more herding. The budget is in compute units, and every policy spends the same.

Policies

Policy What a failure record does
none nothing; no memory, so cells are re-tested freely
private each researcher avoids only their own tested cells
exact everyone avoids every tested cell
hard, radius r exact, and every cell within distance r of a failure record is closed until no open cell remains, after which the swarm falls back to untested cells
hard, stamped as hard, but only failure records from the 10-unit confirmation, which have power 0.90, close neighbours
strict hard with no fallback: a closed cell is never tested, and researchers go idle when no open cell remains
soft each failure record lowers the log-odds of its neighbours by the log likelihood ratio of the failure, ln((1 − power)/(1 − size)), spread over a Gaussian of width 2 cells; researchers sample cells by salience times posterior probability of being true
soft plus positive records as soft, and each discovery raises the log-odds of its neighbours
soft with retest as soft, but tested cells keep a quarter of their weight, so they can be tried again
cutoff soft, with a work cutoff: a cell whose posterior falls below 1% is closed for good

The cutoff policies come in four forms. Evidence from several neighbouring failure records is either summed, as if the records were independent, or pooled, taking only the strongest. And a record is either stamped, so a screen failure counts as the weak evidence it is, or unstamped, so every failure counts as a failed confirmation. Pooling is the conservative reading, since neighbouring tests of related ideas are not independent evidence.

The soft policy is the stamped graveyard: a screen failure at power 0.55 has a likelihood ratio of 0.5 and a confirmation failure at power 0.90 one of 0.1, so the same failure record counts for less when the test was weaker. Its prior is the true prevalence, 3%, and the kernel width of 2 is near the truth's scale; we examine both choices in the results.

Measures

True discoveries, false discoveries and the false discovery rate at a given compute. Duplicate fraction, the share of screens on a cell tested before. Collateral loss, the number of true cells that were closed by a neighbour's failure record before anyone tested them. The sensitivity tables (herding, clustering, screen power, kernel width, latency) use 120 seeds and a budget of 4,000 units, each row its own set of runs. Table 1, the strict and cutoff tables, the retest table and the figure curves all come from one set of runs: the same 60 worlds, run to 24,000 units, read at checkpoints. Within a table the worlds are shared across policies and intervals are 95% intervals on paired differences.

Results

Sharing exact records is an economics result about herding

Figure 1. What a shared graveyard buys, and what it can cost: true discoveries against compute units for a swarm of sixteen researchers, with clustering 3, herding 1 and shared-noise share 0.5. Tabs group the policies; lines run only through the checkpoints that were run. On the first tab the dashed line marks what the unshared swarm finds at 8,000 units.

Table 1 sets every policy against the same 60 worlds at the 8,000-unit checkpoint, with clustering ℓ = 3, herding h = 1, shared-noise share ρ_d = 0.5 and screen power 0.55. The truth has 192 true cells.

Policy True discoveries False FDR Screens on tested cells Change vs exact True cells closed untested
none 40.6 (39.2, 42.1) 16.9 0.294 39.2% -19.9 0
private 43.9 (42.5, 45.3) 19.1 0.303 37.2% -16.6 0
exact 60.5 (58.9, 62.1) 19.6 0.245 0 0 0
hard, r = 1.5 64.3 (62.8, 65.7) 18.5 0.224 0 +3.8 73.1
hard, r = 2.5 59.9 (58.2, 61.5) 19.2 0.243 0 -0.6 79.1
hard, r = 3.5 60.8 (59.3, 62.2) 19.5 0.243 0 +0.3 79.0
hard, r = 3.5, stamped 68.3 (66.1, 70.5) 18.4 0.212 0 +7.8 64.8
soft 79.7 (77.4, 81.9) 18.2 0.186 0 +19.2 0
soft plus positive records 100.6 (98.7, 102.6) 17.1 0.145 0 +40.1 0

A swarm with no memory spends 39% of its screens on cells already tested. Private memory barely helps, because the duplicates come from different researchers. Sharing exact records removes them and turns 8,000 units into 60.5 true discoveries against 40.6; read the other way, the sharing swarm reaches the unshared swarm's 40.6 at about 5,300 units, a saving of a third. False discoveries do not fall, since removing a duplicate removes the same chance of a false pass as of a true one; the false discovery rate falls because the true count rises.

The size of that saving depends on how much researchers crowd onto the same ideas.

Figure 2. The gain from sharing grows with herding: share of screens on already-tested cells without sharing (left) and true discoveries at 4,000 units with and without exact records (right), as researchers crowd onto the same ideas. The label is the gain from sharing and its size relative to no sharing.
Herding h Screens on tested cells, no sharing True discoveries, no sharing True discoveries, exact Gain from sharing
0 13.2% 27.4 30.1 +2.7 (10%)
0.5 16.1% 26.4 30.6 +4.2 (16%)
1.0 25.4% 24.0 30.9 +6.9 (29%)
1.5 39.6% 20.4 29.9 +9.5 (47%)
2.0 55.1% 16.2 30.2 +14.0 (87%)

This table uses a budget of 4,000 units. The shared-record swarm is flat in h, as it should be, since it never duplicates. The unshared swarm loses a growing share of its compute to repeats, and the gain from a graveyard goes from a tenth to nearly a doubling. The number of researchers in the range 4 to 64, which changes how stale a record is when it is read, moved the gain only between +5.1 and +6.9 at h = 1. Where nobody crowds, a shared graveyard is a small convenience. Where they do, it is most of the efficiency.

Reaching into neighbours buys speed and costs recall

A record can say more than "this cell failed". It can close the cells around it. Table 2 gives the change in true discoveries against exact records for hard rules of several radii, at a budget of 4,000 units, as the truth goes from independent to strongly clustered.

Figure 3. Reaching into neighbours helps only when truth is clustered: change in true discoveries against exact records at 4,000 units, by policy and by size of the true clusters. Pink cells are gains and pale cells losses. With independent truth every cell is within a fifth of a discovery of zero.
Clustering ℓ exact (true found) hard r = 1.5 hard r = 2.5 hard r = 3.5 hard, stamped, r = 3.5 soft
0 (independent) 29.0 +0.2 -0.1 -0.1 -0.1 +0.1
1.5 30.3 +5.7 +1.1 +0.2 +2.9 +4.2
3 30.9 +8.0 +1.7 +0.6 +5.6 +8.7
6 30.4 +8.9 +3.5 +1.5 +9.8 +13.0

Three features matter. With independent truth nothing helps, to within a fifth of a discovery, because a failure record says nothing about its neighbours. With clustered truth the gain is real and grows with the cluster size. And the gain shrinks as the radius grows, at every ℓ we tried: a record that reaches further is worse, not better, because it closes the space faster. The reason is arithmetic. Nine in ten tests fail, each failure closes about πr² cells, and the cells it closes overlap, so the map is full after a number of tests close to the cell count divided by the footprint. For r = 2.5 and 6,400 cells that is several hundred screens, and 634 in our strict runs.

The gain also fades with budget. In the long runs, at 8,000 units the best hard rule, r = 1.5, is ahead of exact records by 3.8 discoveries where at 4,000 it was ahead by 8.5, and the larger radii have already fallen to the exact figure. What the rule bought was order: it sent the swarm to unexplored regions first. A fallback that reopens closed cells when the open ones run out is what keeps this from being a loss, and a graveyard rule that is meant to be binding has no such fallback.

Strict rule True found at 4,000 at 8,000 at the end Screens before idle True cells closed untested
exact records 30.9 60.5 103.6 6,488 0
hard, r = 1.5, strict 32.3 32.3 32.3 1,299 133
hard, r = 2.5, strict 14.7 14.7 14.7 634 164
hard, r = 3.5, stamped, strict 36.0 52.5 52.5 2,415 94

This table runs to 24,000 units with 60 seeds. A strict rule with radius 2.5 closes 164 of the 192 true cells, 85%, before anyone tests them, and the swarm goes idle after about 6% of its budget with 14.7 discoveries where the exact swarm finishes with 103.6. The idle swarm is not an error in the simulation. It is the consequence of a rule that says closed means closed, applied to a graveyard that fills with noisy failures.

What the stamp does

A record can carry the power of the test that produced it. We tested the idea two ways.

First, as a rule for which failure records may close neighbours. Table 3 varies the power of the screen, at 4,000 units and ℓ = 3. Columns two and three are discoveries found; the rest are changes against exact records.

Screen power none (found) exact (found) hard, r = 2.5 hard, stamped, r = 3.5 soft
0.30 13.9 16.7 +0.9 +6.3 +4.1
0.55 24.0 30.9 +1.7 +5.6 +8.7
0.80 32.7 41.7 +9.0 +4.5 +27.0

When screens are weak, letting only confirmations close neighbours is worth +6.3 against +0.9 for letting everything do it. When screens are strong the order reverses, +9.0 against +4.5, because the stamp now throws away good information. A fixed rule cannot be right at both ends. A weight that follows the stamp, the soft policy, is ahead of exact records at all three and is far ahead when the tests are good: +4.1, +8.7 and +27.0. Its weakness is that it uses a prior and a kernel width we gave it correctly. At width 0.5 its gain is +0.1, at 2 and 4 it is +8.7 and +8.6, at 8 it is +4.0 (ℓ = 3).

Second, as the protection against a work cutoff. Suppose, as in most organisations, that nobody works an idea whose posterior has fallen below a threshold, here 1%. Then the posterior does the closing. Table 4 runs it to 24,000 units over 60 seeds.

Cutoff policy True found at 4,000 at 8,000 at the end True cells closed untested
exact records 30.9 60.5 103.6 0
summed, unstamped 7.8 7.8 7.8 178
summed, stamped 16.1 16.1 16.1 163
pooled, unstamped 15.7 15.7 15.7 164
pooled, stamped 34.6 69.5 69.5 61
pooled, stamped, cutoff 0.5% 32.5 66.7 96.0 13
pooled, stamped, cutoff 2% 20.7 20.7 20.7 154

Pooling evidence and stamping each matter, and together they are the difference between 15.7 and 69.5. The best of them is ahead of exact records at 4,000 and 8,000 units, 34.6 against 30.9 and 69.5 against 60.5, and finishes 34 discoveries behind, because it has closed itself. And it is a knife edge in the cutoff: moving the threshold from 1% to 0.5% takes the end figure from 69.5 to 96.0, and to 2% takes it to 20.7. With independent truth the same policies are ahead nowhere (pooled and stamped finishes at 51.3 against 101.2): they only close.

When to test again

A failure record fixes a cell for good only if the history is frozen. When the per-cell noise is shared completely, a second test returns the first answer, so a false negative is permanent whatever the policy. When it is partly fresh, a retest can recover one.

Shared-noise share ρ_d exact, at the end soft with retest, at the end Change False, exact False, retest
0 95.6 (13,850 units) 167.3 (24,000) +71.7 6.0 10.6
0.5 103.6 148.7 +45.1 33.0 47.8
1.0 106.7 103.7 -3.0 63.1 56.9

A retest costs compute and is worth it only once the fresh cells run out. At 12,000 units, before they do, soft with retest has 105.5 against 100.8 for soft without it and 90.2 for exact records (ρ_d = 0.5). After that it keeps going where the others have idled, and at ρ_d = 0 it recovers 75% more true cells at a false discovery rate of 0.06, the same as exact records. At ρ_d = 1 it returns nothing. A record that does not say which data it consumed cannot say whether a retest would be independent.

Falsification and limits

Two experiments were designed to find no effect. With independent truth (ℓ = 0), no policy that reaches into neighbours changes the number of discoveries by more than 0.2 in Table 2, which is what must happen if the gain comes from structure. And in a world with no true cells, the 120 seeds produce no true discoveries under any policy, and the number of false discoveries per 4,000 units is 9.9 for no memory, 10.5 for exact records and 10.8 for the hard and soft rules. No record changes the false pass rate of a test. The difference between none and the rest is the repeats.

Some of what we expected did not happen. We expected a larger radius to pay more on larger clusters, and it paid less at every size. We expected an over-confident negative prior to hurt, and over a finite budget it helped: scaling the log-odds update by 4 raised the gain over exact records at ℓ = 3 from +8.7 to +15.4. The cost of that confidence shows only at exhaustion and under a cutoff, which is where Table 4 sits. And the soft rule's lead over exact records does not survive to the end of an exhaustive search: at 24,000 units it has 102.2 true discoveries against 103.6. It buys speed, not reach.

The limits are large. The idea space is a grid, and real ideas have no coordinates; the premise that related ideas share truth, which is the only source of gain for any rule that reaches into neighbours, is an assumption we vary and not a fact we measure. The soft policy's prior equals the true prevalence and its kernel is near the true cluster scale. Researchers are identical except for a shared salience and have no skill; writing a record costs nothing; and no one tests ideas for reasons other than the chance of success. The power of each test is known and fixed. Real graveyards also rot: a record that was right when written is wrong after the data or the market changes, and we do not model expiry. Finally, none of this uses a real research record.

What this means in practice

Four things follow for anyone running a shared register of failed ideas. Record the power of the test with every entry, as the smallest effect it would have caught at 80%, because the same word, failed, means very different things for a screen and a full test, and the stamp was the largest single factor in whether a closing rule did harm. Record which data the test consumed, so a reader can tell whether a retest is independent: at ρ_d = 0 a retest quota recovers three quarters more of the true ideas in an exhausted space, and at ρ_d = 1 it recovers none. Prefer a weight to a wall. A posterior that lowers an idea's priority was never measurably worse than exact records in our runs, where a rule that closes it was worse on every budget long enough to matter, and where it was better it was better early. And if a wall is wanted, such as a work cutoff, test its threshold, because it will be a knife edge, and keep a retest quota open so that the closed region is sampled at a small fixed rate.

A shared graveyard also needs the least glamorous number: how crowded the swarm is. The gain from exact records was 10% where nobody herded and 87% where researchers piled on the same ideas. A swarm can estimate its own herding from the fraction of its screens that repeat a cell, and price its graveyard accordingly.

Reproducibility appendix

Python 3.14, NumPy 2.4 and SciPy 1.17. The 4,000-unit sensitivity runs use seeds 0 to 119 (120 worlds, shared across policies); Table 1, the strict, cutoff and retest tables and the curves use seeds 0 to 59 run to 24,000 units. In Table 1 the none, private, hard and positive-record rows were run to 8,000 units with the final round completed, which reproduces the 8,000-unit checkpoint of a long run exactly (checked on three seeds); the exact and soft rows are read from the long runs. Runtimes on a shared 10-core machine with three workers: first batch 801 s, second 564 s, third 89 s. The core of the simulation:

python
def make_truth(rng, p_true, ell):
    w = rng.standard_normal((G, G))
    kx = np.fft.fftfreq(G) * G
    k2 = kx[:, None] ** 2 + kx[None, :] ** 2
    f = np.fft.ifft2(np.fft.fft2(w) * np.exp(-0.5 * k2 * (2 * np.pi * ell / G) ** 2)).real
    return (f >= np.quantile(f, 1 - p_true)).ravel()

def test_cell(rng, cell, truth, xi_cell, P):
    c1 = norm.isf(P["a1"]); mu1 = c1 - norm.ppf(1 - P["w1"])
    c2 = norm.isf(P["a2"]); mu2 = c2 - norm.ppf(1 - P["w2"])
    rho = P["rho_d"]
    def z(mu):
        return mu * truth[cell] + np.sqrt(rho) * xi_cell[cell] + np.sqrt(1 - rho) * rng.standard_normal()
    if z(mu1) <= c1:
        return "screen_fail", P["w1"], P["a1"], P["cost1"]
    if z(mu2) <= c2:
        return "confirm_fail", P["w2"], P["a2"], P["cost1"] + P["cost2"]
    return "discovery", None, None, P["cost1"] + P["cost2"]

def update_soft(logodds, negev, cell, power, size, kern, base_lo, pooled, unstamped):
    llr = np.log(0.10 / 0.99) if unstamped else np.log((1 - power) / (1 - size))
    k, dx, dy = kern
    i, j = divmod(cell, G)
    idx = ((i + dx) % G) * G + (j + dy) % G
    if pooled:
        negev[idx] = np.maximum(negev[idx], k.ravel() * (-llr))
        logodds[idx] = base_lo - negev[idx]
    else:
        logodds[idx] += k.ravel() * llr

def weights(sal, logodds, known_pos, pmin):
    p = 1.0 / (1.0 + np.exp(-logodds))
    w = sal * p
    if pmin > 0:
        w[p < pmin] = 0.0
    w[known_pos] = 0.0
    return w

Each round, every researcher draws a cell from their normalised weights, the tests run, and the updates are applied together at the end of the round. A strict policy draws nothing when the weights sum to zero.

  • Research methodology
  • Negative results
  • Multi-agent search
  • False negatives
  • Simulation

References

  1. McElreath, R., Smaldino, P. E., Replication, Communication, and the Population Dynamics of Scientific Discovery, PLOS ONE 10(8), e0136088, 2015.
  2. Nissen, S. B., Magidson, T., Gross, K., Bergstrom, C. T., Publication bias and the canonization of false facts, eLife 5, e21451, 2016.
  3. Wang, H., Negative Knowledge as Failure-aware Shared Memory for AutoResearch, arXiv:2606.21024, 2026.
  4. Fiedler, K., Kutzner, F., Krueger, J. I., The Long Way From α-Error Control to Validity Proper: Problems With a Short-Sighted False-Positive Debate, Perspectives on Psychological Science 7, 661-669, 2012.
  5. Rosenthal, R., The file drawer problem and tolerance for null results, Psychological Bulletin 86, 638-641, 1979.

Cite this paper

@misc{forticia2026when,
  title        = {{When a failure record lies: the economics of shared negative results in a research swarm}},
  author       = {Richards, Cayden},
  year         = {2026},
  month        = oct,
  publisher    = {Forticia Research Institute},
  howpublished = {\url{https://www.forticia.uk/papers/graveyard-economics-shared-negative-results}},
  note         = {Paper, published online}
}

Generated from this page’s metadata. Forticia does not assign DOIs to these papers.

NewerGate stacks are not products: measuring the rubber-stamp risk of a research pipelineOlderNever Tune a Corpse

Related

  • Feed honesty before alpha: findings from the Yash Desk options research programmeQuantitative Finance / 11 min
  • Canaries in the approval queue: what injected known-bad requests buy a fatigued human reviewerAI Instrumentation / 15 min
All papers
04080120160True discoveries04,0008,00012,00016,00020,00024,000Compute unitsNo sharing at 8,000 units: 40.6
  • No sharing
  • Exact records
  • Soft weights
  • Soft with retest
Compute units
Move across the chart, or focus it and use the arrow keys
Line charts of true discoveries against compute units, drawn only at the compute checkpoints that were run, with three tabs. Sharing records: with no sharing the swarm finds 40.6 true ideas at 8,000 units, where exact records have 60.5; at the end exact records have 103.6, soft weights 102.2 and soft weights with retest 148.7. Strict rules: radius 1.5 stops at 32.3, radius 2.5 at 14.7 and radius 3.5 with stamps at 52.5. Work cutoff: pooled and stamped records end at 69.5, pooled and unstamped at 15.7, summed and unstamped at 7.8, and with a 0.5 percent cutoff at 96.0.
Compute unitsNo sharingExact recordsSoft weightsSoft with retest
1,000788
2,00013151819
4,00023313939
6,000334661
8,00041618077
12,00090101105
16,000104102125
24,000104102149
Screens on cells already tested, no sharing
0%20%40%60%0.00.51.01.52.0Herding
True discoveries at 4,000 units
0102030400.00.51.01.52.0Herding+2.7 (10%)+14.0 (87%)
  • Screens on tested cells
  • No sharing
  • Exact records
Two charts against herding from 0 to 2. Left: the share of screens spent on cells already tested, without sharing, rises from 13.2% to 55.1%. Right: true discoveries at 4,000 units. With exact records they stay near 30 at every herding level, without sharing they fall from 27.4 to 16.2, so the gain from sharing grows from +2.7 (10%) to +14.0 (87%).
PanelSeriesPoints
Screens on cells already tested, no sharingScreens on tested cells0.0: 13%; 0.5: 16%; 1.0: 25%; 1.5: 40%; 2.0: 55%
True discoveries at 4,000 unitsNo sharing0.0: 27; 0.5: 26; 1.0: 24; 1.5: 20; 2.0: 16
True discoveries at 4,000 unitsExact records0.0: 30; 0.5: 31; 1.0: 31; 1.5: 30; 2.0: 30
Cluster size, in cells (0 is independent)01.536Hard, radius 1.5Hard, radius 2.5Hard, radius 3.5Hard, stamped, radius 3.5Soft+0.2+5.7+8.0+8.9−0.1+1.1+1.7+3.5−0.1+0.2+0.6+1.5−0.1+2.9+5.6+9.8+0.1+4.2+8.7+13.0−14.0+14.0Change in true discoveries
Cell
Move across the matrix, or focus it and use the arrow keys
Matrix of the change in true discoveries against exact records at 4,000 units, for 5 policies (rows) and clustering 0, 1.5, 3, 6 cells (columns). With independent truth (clustering 0) every cell is within 0.2 of zero. At the largest clustering the gains run from +1.5 to +13.0.
01.536
Hard, radius 1.5+0.2+5.7+8.0+8.9
Hard, radius 2.5−0.1+1.1+1.7+3.5
Hard, radius 3.5−0.1+0.2+0.6+1.5
Hard, stamped, radius 3.5−0.1+2.9+5.6+9.8
Soft+0.1+4.2+8.7+13.0