Skip to content
Forticia
Research
Quantitative FinanceEquities, FX, futures. Factors and backtests, every run replayable.Computational BiologySequence, folding, simulation. Versioned, reproducible labs.Cultural IntelligencePhilosophy, governance, ethics. How institutions decide and answer for it.AI InstrumentationPrivate models. Multi-agent orchestration. Guardrails on write.View all research
Quantitative Finance
  • Equities, FX, futures
  • Factors and backtests
  • Every run logged and replayable
InfrastructurePapersPolarisLink™About
Sign inRequest access
Request access

Research

Quantitative FinanceEquities, FX, futures. Factors and backtests, every run replayable.Computational BiologySequence, folding, simulation. Versioned, reproducible labs.Cultural IntelligencePhilosophy, governance, ethics. How institutions decide and answer for it.AI InstrumentationPrivate models. Multi-agent orchestration. Guardrails on write.

Platform

InfrastructurePapersPolarisLink™About
Request accessSign in
Forticia

A private institute for computational research. It publishes original research and runs a governed environment where every run is logged and replayable.

Sign inSystem status

Research

Quantitative FinanceComputational BiologyCultural IntelligenceAI Instrumentation

Platform

InfrastructurePapersPolarisLink™StatusSign in

Institute

AboutRequest accessContactGitHubPolarisLink repository

Forticia publishes research and simulations. Nothing on this site is investment advice.

© 2026 ForticiaPrivacyTerms
Papers/Quantitative Finance
QuantPaper

Gate stacks are not products: measuring the rubber-stamp risk of a research pipeline

Author
Cayden Richards
Published
5 October 2026
Last updated
5 October 2026
Reading time
19 min
Cite this paper

On this page 0%

  1. Abstract
  2. Motivation
  3. Related work and what is new
  4. Method
  5. The population
  6. Tuning and noise
  7. The gates
  8. Measures
  9. Results
  10. Pipelines compared
  11. Stacking gates against chance
  12. The permutation baseline
  13. Where the stack is redundant by construction
  14. A placebo judged on one draw is a coin flip
  15. The degradation ratio punishes tuning
  16. The audit is only as independent as the auditor
  17. How many looks does a blind window survive
  18. Noise misspecified
  19. Prevalence
  20. Falsification and limits
  21. What this means in practice
  22. Reproducibility appendix
  23. References
  24. Cite this paper

Abstract

We simulate a research pipeline of the kind used to certify trading ideas: a sample-size gate, an in-sample gate, an air-gapped blind window, a degradation ratio, a placebo control and a second-researcher audit. The population mixes true edges with five ways of looking like one. Defining rubber-stamp risk as the probability that a pipeline passes an input known to be false, we find that the product of pass rates understates the risk of the five gates that read the data by a factor of 18 on pure-chance ideas, because the gates share noise. A placebo gate that asks one control draw to be unprofitable passes half of everything and halves power. The residual risk sits in the mechanism with a single defender. We give an estimator and the code.

Motivation

A team has five gates between an idea and a certificate. Under a pure-noise idea the gates pass 67%, 10%, 1%, 4% and 5% of the time. Multiplying gives about one in 700,000, and the team reports that its pipeline is that much more selective than no pipeline. In the simulation below the measured figure for those same gates is about one in 40,000, eighteen times worse, and the same arithmetic hides a bigger problem: the product is a statement about one failure mechanism, chance. A research process fails in several other ways, and for each one the question is how many independent gates are looking at it.

We care because the process we describe has this shape. A modern in-sample window is used for exploration. A separate, much longer window is kept blind and consulted only to certify. Controls are run that the effect must fail to survive, costs are charged, a small-sample rule keeps thin evidence out, and a second researcher tries to break the result before anything is accepted. Each addition answers a known way for research to go wrong. What we had not done is measure whether the additions compose. This paper does that on a simulated population, with every number coming from code in the appendix. No real strategy, instrument or return series is involved, and nothing here reports the outcome of any trading.

Related work and what is new

Selection bias under multiple testing is well covered. The deflated Sharpe ratio corrects a selected Sharpe ratio for the number of trials [2], the probability of backtest overfitting measures how often the in-sample winner is below median out of sample [3], the reality check tests whether the best of many rules beats a benchmark [5], and Harvey, Liu and Zhu argue for a higher t-statistic hurdle in factor research [4]. All of them evaluate one statistic against one selection process. The intersection-union test is the classical result for requiring several tests to pass: size is bounded by the largest individual size, not the product [1]. Pollanen has recently studied r-of-k release gates for machine-learning checkpoints under an equicorrelated latent model, and finds that correlation between tests can matter more than the number of tests [6]. Adaptive reuse of a holdout is analysed in the reusable holdout and the Ladder [7, 8].

Four things here are new. First, correlation between gates is not a parameter: it follows from which data each gate reads, and we measure it per failure mechanism. Second, we define a stack efficiency that can be estimated from a null ensemble for any black-box pipeline. Third, we show two design defects in common gates: a control that is judged on a single draw, and a degradation ratio that penalises tuned candidates. Fourth, we quantify how many looks at a blind window it takes before it stops acting as one test.

Method

The population

Each idea has an annualised Sharpe ratio gross of costs in two windows, an in-sample window of 6 years and a blind window of 18 years, a cost drag in Sharpe units, and an event rate. An idea belongs to one of six classes, constructed so that exactly one gate family is the natural defender of each.

Class Construction Share
chance no edge; costs make net Sharpe negative 77%
real net Sharpe 0.6 to 1.6 in both windows 3%
costly gross edge that costs consume 5%
fragile edge in the in-sample window only 5%
confounded exposure to a market-wide factor whose drift looks like edge 7%
buggy an error that inflates both windows equally, such as lookahead 3%

Cost drag is uniform on 0.05 to 0.4 Sharpe units for every class except costly ideas, where it is uniform on 0.5 to 1.2 and the gross edge is 50% to 100% of it. Fragile ideas have a gross in-sample edge of 0.65 to 2.0 and none in the blind window. Confounded ideas have exposure β uniform on 0.5 to 1.2 to the market factor. Buggy ideas have an inflation uniform on 0.8 to 1.8 in both windows. Event rates are log-uniform from 10 to 400 per year. Only the real class is a true positive. The confounded class draws its drift from windows of a public daily market factor series [9]: for each simulated world we pick a start date, take 18 years as the blind window and the following 6 as the in-sample window, and use the realised annualised Sharpe of that stretch. The shares are assumptions, chosen so that positives are rare, and the results section varies them.

Tuning and noise

Observed Sharpe ratios are the true value plus noise with standard deviation s = sqrt(φ/Y), where Y is the window length in years and φ inflates the textbook variance for dependence and tails. We measured φ on 47 public industry portfolios over 1969 to 2026 (14,537 daily observations), demeaned and block-bootstrapped: φ = 1.10 with block length 21 days (standard error 0.02) against 1.01 for independent resampling. We use 1.10.

Before certification the developer tries K = 20 variants of an idea in the in-sample window and keeps the best. Variants share a common noise component with weight ρ_v = 0.5. The kept variant's in-sample estimate is therefore inflated by selection, which is the winner's curse, and its blind-window estimate is a fresh draw. For the control draws we need the maximum of K normals, which we sample exactly as the inverse normal of U^(1/K).

The gates

All thresholds are illustrative and are not those of any institution.

Gate Passes when Reads
N events in the in-sample window ≥ 200 event rate only
IS tuned in-sample net Sharpe ≥ 0.8 in-sample noise
BL blind net Sharpe ≥ 0.4 independent blind noise
DEG blind / in-sample ≥ 0.6 both
PL control, three designs below in-sample noise, fresh control noise
AUD a second researcher finds no defect the idea's implementation

The deflated-Sharpe-only procedure does not use these gates: it tunes on all 24 years, deflates by the expected maximum of K trials as in [2], and passes at a one-sided probability of 0.95. The placebo has three designs. PL0 draws one control and passes if its Sharpe is not positive. PL1 draws 19 controls and passes if the candidate beats all of them, but the controls are untuned single draws. PL2 draws 19 controls and gives each the same K-variant tuning as the candidate. The control is exposure-matched: it keeps the market drift of a confounded idea and has no edge, so it separates edge from exposure.

The audit models the case where two reviewers share blind spots. A buggy idea has a latent difficulty d, and the author's own review already caught everything with d below zero, so the buggy ideas that reach the pipeline have d above zero. The auditor perceives κd plus independent noise and catches the bug if the perceived value is below zero. With κ = 0 the auditor catches half of surviving bugs; at κ = 0.8, 21%. Clean ideas are falsely flagged 5% of the time. The bug share and the catch rule are assumptions about human and machine reviewers that we cannot calibrate from data. We vary both.

Measures

For a pipeline and an idea class m, the rubber-stamp risk R_m is the probability that the pipeline passes an idea of that class. Power is R for the real class. The false discovery rate (FDR) is the fraction of passes that are not real, which depends on the shares. For a set of gates with marginal null pass rates R_g and joint rate R, we define stack efficiency GSE = ln R / Σ ln R_g. It is 1 for independent gates and falls toward 1/n for n gates that fail together. The redundancy factor R / Π R_g says by how much the product understates the risk.

Every figure is an average over simulated worlds, each with its own market path, and confidence intervals come from resampling worlds.

Results

Pipelines compared

Figure 1. Six pipelines, two measures: power and false discovery rate (top), then the probability that each pipeline passes an idea of each class on a log colour scale (bottom). Population of 10 million ideas in 500 worlds; the paper gives intervals within 0.003 for power and false discovery rate.

Table 1 compares a single-holdout procedure (tune in-sample, confirm once in the holdout, gross of costs), a deflated-Sharpe-only procedure (tune on all 24 years, deflate for K), a purely statistical stack, and the stack with a placebo and an audit. Population of 10 million ideas in 500 worlds.

Pipeline Power FDR Chance Costly Fragile Confounded Buggy
Single holdout, gross 0.988 0.761 1.30e-2 0.672 5.2e-2 0.277 0.988
Deflated Sharpe only 0.977 0.559 2.1e-4 4.4e-4 3.3e-2 0.097 0.961
N, IS, BL, DEG 0.400 0.513 1.7e-4 3.2e-4 1.0e-4 1.3e-2 0.388
plus PL2 0.369 0.493 2.8e-5 3.0e-4 3.2e-5 8.5e-4 0.358
plus PL2 and audit (κ=0) 0.351 0.339 2.6e-5 2.9e-4 3.2e-5 8.0e-4 0.178
plus PL2 and audit (κ=0.8) 0.351 0.448 2.6e-5 2.9e-4 3.2e-5 8.0e-4 0.284

The 95% intervals on power and FDR, from resampling worlds, are within ±0.003. Three things stand out. The single holdout is a fair filter against chance (1.3% pass) and almost none against costs (67% pass) or errors (99%). The deflated Sharpe ratio is the opposite: it handles chance and costs, about 2e-4, because it is applied to net returns over the whole sample, and it passes 96% of buggy ideas and 10% of confounded ones, because a drift or a bug that inflates both windows is a real Sharpe ratio. And the most elaborate stack has a false discovery rate of 0.34 to 0.45, nearly all of it buggy ideas: 3% of the population, passing at 18% to 28%. The statistical gates cannot see that mechanism because by construction it corrupts both windows equally. Only the audit can, so the audit's independence sets the pipeline's floor.

Stacking gates against chance

Figure 2. Gates are not independent: pass rate of pure-chance ideas as gates are added one at a time (15 million ideas), against the product of each gate's own pass rate. The multiple above each point is observed divided by product.

For pure-chance ideas we ran 15 million per mechanism and added gates one at a time.

Gates Observed Product of marginals Observed / product
N 0.674 0.674 1.0
N, IS 6.69e-2 6.70e-2 1.0
N, IS, BL 9.29e-4 6.33e-4 1.5
N, IS, BL, DEG 1.63e-4 2.77e-5 5.9
N, IS, BL, DEG, PL2 2.53e-5 1.38e-6 18.3
all six 2.39e-5 1.31e-6 18.2

Stack efficiency for IS, BL and DEG is 0.82 (95% interval 0.82 to 0.83); for the full stack 0.79 (0.78 to 0.79). The N gate is independent of everything else and, with it included, the product is exact: it thins a population without selecting from it. The in-sample gate and the degradation ratio, and the placebo and the in-sample gate, share the same tuned noise, so a null idea that got through the first is far more likely to get through the second. The same stack applied to real edges has a redundancy factor of 1.03: positives pass the gates together because they do pass, and nothing is shared that matters. Correlation between gates therefore shrinks what the stack removes without shrinking what it keeps. That asymmetry is why it is easy not to see.

The permutation baseline

We checked the measurement itself by permuting each gate's pass vector across ideas within a class, which keeps every marginal and destroys all dependence. The measured redundancy then returns to one: for chance ideas through IS and BL, 9.56e-4 permuted against 9.39e-4 for the product, and through IS, BL and DEG, 4.2e-5 against 4.1e-5. Redundancy is a property of the data lineage of the gates, not an artefact of how we count it.

Where the stack is redundant by construction

Pipeline Power FDR Chance Confounded Buggy Fragile
full 0.352 0.343 2.7e-5 7.1e-4 0.180 3.2e-5
drop N 0.520 0.343 4.0e-5 1.1e-3 0.266 5.0e-5
drop IS 0.353 0.350 4.0e-5 8.7e-4 0.180 4.0e-5
drop BL 0.352 0.343 2.7e-5 7.1e-4 0.180 3.2e-5
drop DEG 0.571 0.366 2.8e-4 5.5e-3 0.296 5.5e-3
drop PL2 0.381 0.373 1.6e-4 1.2e-2 0.194 8.0e-5
drop AUD 0.370 0.496 2.6e-5 7.7e-4 0.360 3.5e-5

Dropping the blind gate changes nothing, to the last digit. The reason is arithmetic: the in-sample gate requires 0.8, the ratio requires 0.6 of that, so any idea reaching the ratio has a blind Sharpe of at least 0.48, above the 0.4 blind threshold. A gate whose threshold is implied by two others is dead weight and no simulation is needed to find it, but we did not see it until the marginal-contribution table was printed. The same table shows what each remaining gate is for. Dropping the ratio gate raises fragile ideas by a factor of about 160. Dropping the placebo raises confounded ideas by 17. Dropping the audit raises the false discovery rate from 0.34 to 0.50. Dropping N leaves the false discovery rate unchanged and raises power by 48%, which is the thinning at work.

A placebo judged on one draw is a coin flip

Figure 3. Three placebo designs: share of ideas of each class that pass. A rule that judges one control draw passes about half of everything with no confound. Controls that are not tuned as the candidate was pass it too often. Controls with the same tuning return the nominal five percent on chance ideas.
Placebo design Chance Real Costly Confounded
PL0: one control, pass if not profitable 0.500 0.499 0.500 0.181
PL1: candidate beats 19 untuned controls 0.293 0.985 0.818 0.294
PL2: candidate beats 19 equally tuned controls 0.050 0.916 0.496 0.050

A control with no edge has a Sharpe that is positive half the time, so a rule requiring it to be non-positive passes half of any idea that has no confound. It does catch confounded ideas, because the exposure-matched control keeps the drift and fails 82% of the time, but it halves the pass rate of real edges to 0.499. Adding it to the statistical stack takes power from 0.400 to 0.200 and changes the false discovery rate from 0.513 to 0.502. PL1 fixes the arithmetic and introduces a different error: the candidate was selected as the best of 20 variants and the controls were not, so the candidate beats them 29% of the time against a nominal 5%. PL2 gives the controls the same search and returns exactly 1/(B+1) = 5% on chance ideas, with 92% of real edges passing. The rule is that a control must be subject to the same selection as the thing it controls.

The degradation ratio punishes tuning

A ratio of blind to in-sample Sharpe is intended to detect over-fitting. For a true edge the in-sample estimate is a selected maximum, so the ratio is below one by construction and falls as the search widens.

Variants K Median ratio for real edges 10th percentile Power, N-IS-BL-DEG Chance, IS and BL only Chance, full stack
1 0.88 0.54 0.407 1.4e-4 1.6e-5
5 0.74 0.45 0.450 5.7e-4 3.4e-5
20 0.66 0.40 0.401 1.4e-3 2.7e-5
100 0.59 0.36 0.320 2.9e-3 8.7e-6
500 0.54 0.33 0.248 4.6e-3 3.0e-6

With a fixed threshold of 0.6 the gate rejects a growing share of real edges as the search widens, and from K = 100 the median real edge that passes the in-sample gate fails it. Meanwhile the rubber-stamp risk of the in-sample and blind gates alone grows 32-fold from K = 1 to K = 500, which is the multiple-testing problem the deflated Sharpe ratio addresses. A threshold fixed in advance is wrong at every K but one. The threshold that keeps 90% of real edges ranges from 0.54 to 0.33 across this range, so it should be set as a function of the declared search size.

The audit is only as independent as the auditor

Bug share κ = 0 κ = 0.4 κ = 0.8 κ = 0.95
1% 0.151 0.181 0.216 0.237
3% 0.344 0.396 0.452 0.482
10% 0.634 0.686 0.733 0.757

Entries are the false discovery rate of the full pipeline. Chance of catching a surviving bug is 0.50, 0.37, 0.21 and 0.10 for those κ. At κ = 0.95 the audit nearly vanishes. If the second researcher uses the same tools, the same model family or the same checklist as the first, the audit is closer to a repeat of the author's own review than to a second opinion, and its contribution to the stack is a fraction of what its pass rate suggests. The statistical gates do not compensate, since they cannot see this mechanism.

How many looks does a blind window survive

A blind window stays blind only while it is consulted once. Each time the developer revises a failed idea and resubmits, the blind estimate is a new look, and the question is how many independent tests those looks are worth. We model L successive variants of a chance idea submitted until one passes, where the blind noise of two variants has correlation ρ_b. Variants of one idea on one history are strongly correlated, so ρ_b near 1 is the realistic end. The single-look pass rate is 0.0095. We define the effective number of independent looks as n_eff = ln(1 − R_L) / ln(1 − R_1).

Figure 4. How many looks a blind window survives: pass rate of chance ideas after repeated looks at the blind window (left) and the effective number of independent looks (right), for four correlations between successive variants. Dashed lines: the single-look pass rate (left) and a raw count of looks (right). The multiples on the left are the pass rate at 20 looks divided by the single-look rate.
Looks L ρ_b = 0 ρ_b = 0.5 ρ_b = 0.8 ρ_b = 0.95
2 0.0186 (1.98) 0.0175 (1.85) 0.0151 (1.59) 0.0125 (1.32)
5 0.0457 (4.91) 0.0363 (3.88) 0.0255 (2.71) 0.0166 (1.76)
20 0.160 (18.4) 0.090 (9.9) 0.047 (5.1) 0.023 (2.5)
50 0.320 (40.5) 0.144 (16.4) 0.065 (7.1) 0.028 (3.0)

Entries are the chance-idea pass rate of the blind gate after L looks, with n_eff in brackets; each cell used 2 million ideas. Risk doubles after about 2 looks at ρ_b = 0 and 0.5, 3 looks at 0.8 and 8 looks at 0.95. At 20 looks the risk is 17 times the single-look figure for independent variants and 2.5 times for ρ_b = 0.95. The practical reading is that the raw count of looks always overcharges, and by more the more alike the variants are: 50 looks are worth 40.5 independent tests at ρ_b = 0 and 3.0 at ρ_b = 0.95. The number to charge is the effective one, which needs an estimate of ρ_b from the variants' own blind returns. Holding the risk at its single-look level after 20 looks requires raising the blind threshold from 0.40 to 0.65, 0.63, 0.57 or 0.49 for the four ρ_b, which in this population of strong edges leaves power at 0.974 at ρ_b = 0.95, against 0.971 for a single look, and at 1.00 for independent variants, because re-submission also gives real edges extra chances. Weaker edges would pay more. A ledger that records every submission and sets the threshold from n_eff keeps the window usable; one that records only passes does not.

Noise misspecified

Figure 5. Fixed thresholds under misspecified noise: pass rate of chance ideas as the noise in the world is inflated against thresholds that stay fixed, for four pipelines. The multiple at the right end of each line is its growth from an inflation factor of 1 to 3.

We held thresholds fixed and raised φ in the world from 1.0 to 3.0, which stands for returns with more dependence and heavier tails than the thresholds were set for. Rubber-stamp risk on chance ideas:

φ Single holdout IS and BL N, IS, BL, DEG Full stack
1.0 9.5e-3 8.9e-4 9.6e-5 1.8e-5
1.5 3.0e-2 4.9e-3 6.8e-4 4.1e-5
2.0 5.3e-2 1.2e-2 1.8e-3 5.1e-5
3.0 9.7e-2 3.2e-2 5.4e-3 6.0e-5

Fixed thresholds degrade by a factor of 36 for IS and BL over this range and 10 for the single holdout. The full stack degrades by 3.3 because the placebo is built from the same series and inherits its noise: it self-calibrates, where a Sharpe threshold does not. That holds only if the placebo preserves the dependence structure of the data, which a naive shuffle does not.

Prevalence

The false discovery rates in Table 1 are conditional on 3% real ideas. At 0.5% the full pipeline's FDR is 0.76 and the single holdout's is 0.95; at 25% they are 0.05 and 0.23. The pipeline ordering does not change over this range. Power does not depend on prevalence.

Falsification and limits

We tested the simulator against what can be computed in closed form. The blind-gate pass rate on chance ideas is 0.00950 in simulation and 0.00946 by numerical integration over the cost distribution. The permutation baseline above is a null experiment for the measurement. The placebo on chance ideas lands on its nominal level for PL2 and on one half for PL0, both to three digits. These checks show the code does what we say; they do not show the model is the world.

The Gaussian summary-statistic approximation was checked against real returns in two ways. Annualised Sharpe ratios of block-bootstrapped industry portfolios over 6 and 18 years are close to Gaussian (kurtosis 3.06 and 3.08) with φ of 1.10. And for event-level data, the false-positive rate of a t ≥ 3 rule on resampled standardised daily returns, relative to the Gaussian 0.00135, is 5.3 times at 8 events, 2.6 at 15, 1.5 at 30, 1.25 at 50, 1.1 at 100, and 1.0 from 200 up. A sample-size gate therefore protects against heavy-tailed small samples only below about a hundred events. At the 200-event threshold we used, it is a pure thinning gate, as the model finds. This is a result about the estimator, not about trading: where event counts proxy something else, such as outlier dominance or capacity, the gate has another job that we have not modelled.

Three limits cannot be removed with this design, and we want to be plain about them. The size of the false discovery rates depends on the class shares and on the bug mechanism, both assumed; we report relative orderings and mechanism-level risks, not a forecast. The audit model is two numbers, a share and a correlation, and real review is richer. And the blind-window look model assumes each re-submission is a new variant; a researcher who steers on the score, rather than on pass or fail, extracts more from each look than we have modelled, and the adaptive-data-analysis bounds [7, 8] are the right tool for that case. Thresholds, window lengths and class constructions are ours; a pipeline with different numbers will give different magnitudes, though we expect the same qualitative ranking of mechanisms. We also did not model drift between the windows other than the fragile class.

What this means in practice

We would change four things in any pipeline of this shape. Report rubber-stamp risk per failure mechanism, from a null ensemble for each, instead of one pass rate. For each mechanism, ask which gates have noise the others do not share, and count only those. Treat a mechanism with one defender as a single point of failure and make that defender as independent of the author as possible: different tools, different data path, adversarial incentive. Check threshold arithmetic for entailment before running anything, since a gate implied by two others is free to delete and free to forget.

Gates against selection should scale with the declared search size. A control should undergo the same tuning as its candidate and be judged by rank against many draws, not by a single draw's sign. The blind window should be treated as a budget: record each look, charge it against the pass threshold with the effective number of independent looks, not the raw count, and do not let the score leave the gate when pass or fail is enough.

The estimator is short. Given any function that maps an input to pass or fail and a generator of inputs known to be false for a chosen mechanism, run a few hundred batches, report the pass fraction with a batch-resampled interval, and for every gate subset report the redundancy factor. A pipeline whose redundancy factor on chance ideas is above ten is not as selective as its gates suggest.

Reproducibility appendix

All simulations are in Python 3.14 with NumPy 2.4 and SciPy 1.17. Seeds: 11 (pipelines), 31 (leave-one-out), 41 plus K (variant sweep), 100 plus class index (stack efficiency), 404 (looks). The real-data inputs are the daily Fama-French five-factor file and the 49-industry file from the public data library [9]; the market factor drives the confound windows and the industries give φ. Runtimes on a 10-core laptop shared with other jobs: pipelines 44 s, stack efficiency 128 s, sweeps 175 s.

The tuned-noise sampler and the gates:

python
def tuned_noise(rng, M, K, rho):
    zc = rng.standard_normal(M)
    mx = norm.ppf(rng.random(M) ** (1.0 / K))
    return np.sqrt(rho) * zc + np.sqrt(1 - rho) * mx

def observe(rng, W, P):
    s1, s2 = np.sqrt(P["phi"] / P["Y1"]), np.sqrt(P["phi"] / P["Y2"])
    G1 = W["g1"] + s1 * tuned_noise(rng, W["M"], P["K"], P["rho_v"])
    S1 = G1 - W["c"]
    S2 = W["g2"] + s2 * rng.standard_normal(W["M"]) - W["c"]
    pl = W["drift1"][:, None] + s1 * tuned_noise(rng, W["M"] * P["B"], P["K"], P["rho_v"]).reshape(W["M"], P["B"])
    return G1, S1, S2, pl

def gates(W, G1, S1, S2, pl, P):
    return {
        "N": W["tau"] * P["Y1"] >= P["nmin"],
        "IS": S1 >= P["th1"],
        "BL": S2 >= P["th2"],
        "DEG": S2 / np.where(S1 > 0, S1, np.nan) >= P["ratio"],
        "PL2": G1 > pl.max(1),
    }

The estimator:

python
def rubber_stamp_risk(pipeline, null_batch, n_batches, seed=0):
    rng = np.random.default_rng(seed)
    per_batch = np.array([pipeline(null_batch(rng)).mean() for _ in range(n_batches)])
    r = per_batch.mean()
    se = per_batch.std(ddof=1) / np.sqrt(n_batches)
    return r, (max(r - 1.96 * se, 0.0), r + 1.96 * se)

def stack_efficiency(gate_matrix):
    r_stack = gate_matrix.all(axis=1).mean()
    r_marg = gate_matrix.mean(axis=0)
    gse = np.log(r_stack) / np.log(r_marg).sum()
    return gse, r_stack / np.prod(r_marg), gse * gate_matrix.shape[1]
  • Research methodology
  • Multiple testing
  • Backtest validation
  • Placebo controls
  • Simulation

References

  1. Berger, R. L., Multiparameter hypothesis testing and acceptance sampling, Technometrics 24, 295-300, 1982.
  2. Bailey, D. H., López de Prado, M., The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality, Journal of Portfolio Management, 2014 (SSRN 2460551).
  3. Bailey, D. H., Borwein, J. M., López de Prado, M., Zhu, Q. J., The Probability of Backtest Overfitting, Journal of Computational Finance 20(4), 39-70, 2017.
  4. Harvey, C. R., Liu, Y., Zhu, H., ... and the Cross-Section of Expected Returns, Review of Financial Studies 29(1), 5-68, 2016.
  5. White, H., A Reality Check for Data Snooping, Econometrica 68(5), 1097-1126, 2000.
  6. Pollanen, M., The Price of Correlated Tests: How Strict Should a Model Release Gate Be?, arXiv:2610.00993, 2026.
  7. Dwork, C., Feldman, V., Hardt, M., Pitassi, T., Reingold, O., Roth, A., Generalization in Adaptive Data Analysis and Holdout Reuse, arXiv:1506.02629, 2015.
  8. Blum, A., Hardt, M., The Ladder: A Reliable Leaderboard for Machine Learning Competitions, ICML 2015, arXiv:1502.04585.
  9. French, K. R., Data Library: Fama/French 5 factors (2x3) daily and 49 industry portfolios daily, Dartmouth Tuck School of Business, https://mba.tuck.dartmouth.edu/pages/faculty/ken.french/data_library.html (files dated August 2026).

Cite this paper

@misc{forticia2026gate,
  title        = {{Gate stacks are not products: measuring the rubber-stamp risk of a research pipeline}},
  author       = {Richards, Cayden},
  year         = {2026},
  month        = oct,
  publisher    = {Forticia Research Institute},
  howpublished = {\url{https://www.forticia.uk/papers/gate-stack-rubber-stamp-risk}},
  note         = {Paper, published online}
}

Generated from this page’s metadata. Forticia does not assign DOIs to these papers.

NewerWhat an agent audit log can answer: measuring incident answerability across logging schemasOlderWhen a failure record lies: the economics of shared negative results in a research swarm

Related

  • Feed honesty before alpha: findings from the Yash Desk options research programmeQuantitative Finance / 11 min
  • Canaries in the approval queue: what injected known-bad requests buy a fatigued human reviewerAI Instrumentation / 15 min
All papers

Power and false discovery rate

0%25%50%75%100%Single holdout, gross98.8%76.1%Deflated Sharpe only97.7%55.9%N, IS, BL, DEG40.0%51.3%plus PL236.9%49.3%plus PL2 and audit (κ=0)35.1%33.9%plus PL2 and audit (κ=0.8)35.1%44.8%
  • Power
  • False discovery rate
Grouped bars of power and false discovery rate for six pipelines.
CategoryPowerFalse discovery rate
Single holdout, gross99%76%
Deflated Sharpe only98%56%
N, IS, BL, DEG40%51%
plus PL237%49%
plus PL2 and audit (κ=0)35%34%
plus PL2 and audit (κ=0.8)35%45%

Probability the pipeline passes an idea of each class

ChanceCostlyFragileConfoundedBuggySingle holdout, grossDeflated Sharpe onlyN, IS, BL, DEGplus PL2plus PL2 and audit (κ=0)plus PL2 and audit (κ=0.8)0.0130.6720.0520.2770.9882.1e−44.4e−40.0330.0970.9611.7e−43.2e−41.0e−40.0130.3882.8e−53.0e−43.2e−58.5e−40.3582.6e−52.9e−43.2e−58.0e−40.1782.6e−52.9e−43.2e−58.0e−40.2841.0e−51.000Risk, log scale
Cell
Move across the matrix, or focus it and use the arrow keys
Matrix of rubber-stamp risk by pipeline and idea class.
ChanceCostlyFragileConfoundedBuggy
Single holdout, gross0.0130.6720.0520.2770.988
Deflated Sharpe only2.1e−44.4e−40.0330.0970.961
N, IS, BL, DEG1.7e−43.2e−41.0e−40.0130.388
plus PL22.8e−53.0e−43.2e−58.5e−40.358
plus PL2 and audit (κ=0)2.6e−52.9e−43.2e−58.0e−40.178
plus PL2 and audit (κ=0.8)2.6e−52.9e−43.2e−58.0e−40.284
10⁻⁶10⁻⁴10⁻²10⁰Pass rate of pure-chance ideasNN+ISN+IS+BL+DEG+PL2+AUDGates added, left to rightObservedProduct of marginals1.5×5.9×18.3×18.2×
Line chart on a log axis of the pass rate of pure-chance ideas as gates are added one at a time: N, N+IS, N+IS+BL, +DEG, +PL2, +AUD. The observed rate falls from 0.674 to 2.4e−5, while the product of the marginal pass rates falls to 1.3e−6. The observed rate is 18.3× the product at its largest gap, and 18.2× with all six gates.
Gates added, left to rightObservedProduct of marginals
110⁰10⁰
210⁻¹10⁻¹
310⁻³10⁻³
410⁻⁴10⁻⁵
510⁻⁵10⁻⁶
610⁻⁵10⁻⁶
0%25%50%75%100%Share of ideas that passChance50.0%29.3%5.0%Real49.9%98.5%91.6%Costly50.0%81.8%49.6%Confounded18.1%29.3%5.0%
  • One control, pass if not profitable
  • Beat 19 untuned controls
  • Beat 19 equally tuned controls
Grouped bars of the pass rate of three placebo designs across four idea classes. On chance ideas the designs pass 50.0%, 29.3%, 5.0%. On real edges they pass 49.9%, 98.5%, 91.6%. On costly ideas 50.0%, 81.8%, 49.6%, and on confounded ideas 18.1%, 29.3%, 5.0%.
CategoryOne control, pass if not profitableBeat 19 untuned controlsBeat 19 equally tuned controls
Chance50%29%5%
Real50%98%92%
Costly50%82%50%
Confounded18%29%5%
Pass rate of chance ideas
0%10%20%30%152050Looks at the blind window17×2.5×
Effective independent looks
125102050152050Looks at the blind window
  • Correlation 0
  • Correlation 0.5
  • Correlation 0.8
  • Correlation 0.95
  • Raw count of looks
Two charts against the number of looks at the blind window on a log axis, for variant correlations 0, 0.5, 0.8, 0.95. Left: the pass rate of chance ideas rises from 0.95% at one look to 32% at 50 looks for independent variants and only 2.8% at correlation 0.95. Right: the effective number of independent looks is 40.5 at 50 looks for independent variants and 3.0 at correlation 0.95.
PanelSeriesPoints
Pass rate of chance ideasCorrelation 01: 1%; 2: 2%; 3: 3%; 5: 5%; 10: 9%; 20: 16%; 50: 32%
Pass rate of chance ideasCorrelation 0.51: 1%; 2: 2%; 3: 2%; 5: 4%; 10: 6%; 20: 9%; 50: 14%
Pass rate of chance ideasCorrelation 0.81: 1%; 2: 2%; 3: 2%; 5: 3%; 10: 4%; 20: 5%; 50: 7%
Pass rate of chance ideasCorrelation 0.951: 1%; 2: 1%; 3: 1%; 5: 2%; 10: 2%; 20: 2%; 50: 3%
Effective independent looksRaw count of looks1: 1; 50: 50
Effective independent looksCorrelation 01: 1; 2: 2; 3: 3; 5: 5; 10: 10; 20: 18; 50: 41
Effective independent looksCorrelation 0.51: 1; 2: 2; 3: 3; 5: 4; 10: 6; 20: 10; 50: 16
Effective independent looksCorrelation 0.81: 1; 2: 2; 3: 2; 5: 3; 10: 4; 20: 5; 50: 7
Effective independent looksCorrelation 0.951: 1; 2: 1; 3: 2; 5: 2; 10: 2; 20: 2; 50: 3
10⁻⁵10⁻⁴10⁻³10⁻²10⁻¹Pass rate of chance ideas1.01.52.02.53.0Noise inflation factorSingle holdoutIS and BLN, IS, BL, DEGFull stack10×36×56×3.3×
Log-log line chart of the pass rate of chance ideas against the noise inflation factor from 1 to 3, with thresholds held fixed. Single holdout grows by a factor of 10×; IS and BL grows by a factor of 36×; N, IS, BL, DEG grows by a factor of 56×; Full stack grows by a factor of 3.3×.
Noise inflation factorSingle holdoutIS and BLN, IS, BL, DEGFull stack
1.010⁻²10⁻³10⁻⁴10⁻⁵
1.110⁻²10⁻³10⁻⁴10⁻⁵
1.510⁻²10⁻²10⁻³10⁻⁴
2.010⁻¹10⁻²10⁻³10⁻⁴
3.010⁻¹10⁻¹10⁻²10⁻⁴