Abstract
We simulate a research pipeline of the kind used to certify trading ideas: a sample-size gate, an in-sample gate, an air-gapped blind window, a degradation ratio, a placebo control and a second-researcher audit. The population mixes true edges with five ways of looking like one. Defining rubber-stamp risk as the probability that a pipeline passes an input known to be false, we find that the product of pass rates understates the risk of the five gates that read the data by a factor of 18 on pure-chance ideas, because the gates share noise. A placebo gate that asks one control draw to be unprofitable passes half of everything and halves power. The residual risk sits in the mechanism with a single defender. We give an estimator and the code.
Motivation
A team has five gates between an idea and a certificate. Under a pure-noise idea the gates pass 67%, 10%, 1%, 4% and 5% of the time. Multiplying gives about one in 700,000, and the team reports that its pipeline is that much more selective than no pipeline. In the simulation below the measured figure for those same gates is about one in 40,000, eighteen times worse, and the same arithmetic hides a bigger problem: the product is a statement about one failure mechanism, chance. A research process fails in several other ways, and for each one the question is how many independent gates are looking at it.
We care because the process we describe has this shape. A modern in-sample window is used for exploration. A separate, much longer window is kept blind and consulted only to certify. Controls are run that the effect must fail to survive, costs are charged, a small-sample rule keeps thin evidence out, and a second researcher tries to break the result before anything is accepted. Each addition answers a known way for research to go wrong. What we had not done is measure whether the additions compose. This paper does that on a simulated population, with every number coming from code in the appendix. No real strategy, instrument or return series is involved, and nothing here reports the outcome of any trading.
Related work and what is new
Selection bias under multiple testing is well covered. The deflated Sharpe ratio corrects a selected Sharpe ratio for the number of trials [2], the probability of backtest overfitting measures how often the in-sample winner is below median out of sample [3], the reality check tests whether the best of many rules beats a benchmark [5], and Harvey, Liu and Zhu argue for a higher t-statistic hurdle in factor research [4]. All of them evaluate one statistic against one selection process. The intersection-union test is the classical result for requiring several tests to pass: size is bounded by the largest individual size, not the product [1]. Pollanen has recently studied r-of-k release gates for machine-learning checkpoints under an equicorrelated latent model, and finds that correlation between tests can matter more than the number of tests [6]. Adaptive reuse of a holdout is analysed in the reusable holdout and the Ladder [7, 8].
Four things here are new. First, correlation between gates is not a parameter: it follows from which data each gate reads, and we measure it per failure mechanism. Second, we define a stack efficiency that can be estimated from a null ensemble for any black-box pipeline. Third, we show two design defects in common gates: a control that is judged on a single draw, and a degradation ratio that penalises tuned candidates. Fourth, we quantify how many looks at a blind window it takes before it stops acting as one test.
Method
The population
Each idea has an annualised Sharpe ratio gross of costs in two windows, an in-sample window of 6 years and a blind window of 18 years, a cost drag in Sharpe units, and an event rate. An idea belongs to one of six classes, constructed so that exactly one gate family is the natural defender of each.
| Class | Construction | Share |
|---|---|---|
| chance | no edge; costs make net Sharpe negative | 77% |
| real | net Sharpe 0.6 to 1.6 in both windows | 3% |
| costly | gross edge that costs consume | 5% |
| fragile | edge in the in-sample window only | 5% |
| confounded | exposure to a market-wide factor whose drift looks like edge | 7% |
| buggy | an error that inflates both windows equally, such as lookahead | 3% |
Cost drag is uniform on 0.05 to 0.4 Sharpe units for every class except costly ideas, where it is uniform on 0.5 to 1.2 and the gross edge is 50% to 100% of it. Fragile ideas have a gross in-sample edge of 0.65 to 2.0 and none in the blind window. Confounded ideas have exposure β uniform on 0.5 to 1.2 to the market factor. Buggy ideas have an inflation uniform on 0.8 to 1.8 in both windows. Event rates are log-uniform from 10 to 400 per year. Only the real class is a true positive. The confounded class draws its drift from windows of a public daily market factor series [9]: for each simulated world we pick a start date, take 18 years as the blind window and the following 6 as the in-sample window, and use the realised annualised Sharpe of that stretch. The shares are assumptions, chosen so that positives are rare, and the results section varies them.
Tuning and noise
Observed Sharpe ratios are the true value plus noise with standard deviation s = sqrt(φ/Y), where Y is the window length in years and φ inflates the textbook variance for dependence and tails. We measured φ on 47 public industry portfolios over 1969 to 2026 (14,537 daily observations), demeaned and block-bootstrapped: φ = 1.10 with block length 21 days (standard error 0.02) against 1.01 for independent resampling. We use 1.10.
Before certification the developer tries K = 20 variants of an idea in the in-sample window and keeps the best. Variants share a common noise component with weight ρ_v = 0.5. The kept variant's in-sample estimate is therefore inflated by selection, which is the winner's curse, and its blind-window estimate is a fresh draw. For the control draws we need the maximum of K normals, which we sample exactly as the inverse normal of U^(1/K).
The gates
All thresholds are illustrative and are not those of any institution.
| Gate | Passes when | Reads |
|---|---|---|
| N | events in the in-sample window ≥ 200 | event rate only |
| IS | tuned in-sample net Sharpe ≥ 0.8 | in-sample noise |
| BL | blind net Sharpe ≥ 0.4 | independent blind noise |
| DEG | blind / in-sample ≥ 0.6 | both |
| PL | control, three designs below | in-sample noise, fresh control noise |
| AUD | a second researcher finds no defect | the idea's implementation |
The deflated-Sharpe-only procedure does not use these gates: it tunes on all 24 years, deflates by the expected maximum of K trials as in [2], and passes at a one-sided probability of 0.95. The placebo has three designs. PL0 draws one control and passes if its Sharpe is not positive. PL1 draws 19 controls and passes if the candidate beats all of them, but the controls are untuned single draws. PL2 draws 19 controls and gives each the same K-variant tuning as the candidate. The control is exposure-matched: it keeps the market drift of a confounded idea and has no edge, so it separates edge from exposure.
The audit models the case where two reviewers share blind spots. A buggy idea has a latent difficulty d, and the author's own review already caught everything with d below zero, so the buggy ideas that reach the pipeline have d above zero. The auditor perceives κd plus independent noise and catches the bug if the perceived value is below zero. With κ = 0 the auditor catches half of surviving bugs; at κ = 0.8, 21%. Clean ideas are falsely flagged 5% of the time. The bug share and the catch rule are assumptions about human and machine reviewers that we cannot calibrate from data. We vary both.
Measures
For a pipeline and an idea class m, the rubber-stamp risk R_m is the probability that the pipeline passes an idea of that class. Power is R for the real class. The false discovery rate (FDR) is the fraction of passes that are not real, which depends on the shares. For a set of gates with marginal null pass rates R_g and joint rate R, we define stack efficiency GSE = ln R / Σ ln R_g. It is 1 for independent gates and falls toward 1/n for n gates that fail together. The redundancy factor R / Π R_g says by how much the product understates the risk.
Every figure is an average over simulated worlds, each with its own market path, and confidence intervals come from resampling worlds.
Results
Pipelines compared
Table 1 compares a single-holdout procedure (tune in-sample, confirm once in the holdout, gross of costs), a deflated-Sharpe-only procedure (tune on all 24 years, deflate for K), a purely statistical stack, and the stack with a placebo and an audit. Population of 10 million ideas in 500 worlds.
| Pipeline | Power | FDR | Chance | Costly | Fragile | Confounded | Buggy |
|---|---|---|---|---|---|---|---|
| Single holdout, gross | 0.988 | 0.761 | 1.30e-2 | 0.672 | 5.2e-2 | 0.277 | 0.988 |
| Deflated Sharpe only | 0.977 | 0.559 | 2.1e-4 | 4.4e-4 | 3.3e-2 | 0.097 | 0.961 |
| N, IS, BL, DEG | 0.400 | 0.513 | 1.7e-4 | 3.2e-4 | 1.0e-4 | 1.3e-2 | 0.388 |
| plus PL2 | 0.369 | 0.493 | 2.8e-5 | 3.0e-4 | 3.2e-5 | 8.5e-4 | 0.358 |
| plus PL2 and audit (κ=0) | 0.351 | 0.339 | 2.6e-5 | 2.9e-4 | 3.2e-5 | 8.0e-4 | 0.178 |
| plus PL2 and audit (κ=0.8) | 0.351 | 0.448 | 2.6e-5 | 2.9e-4 | 3.2e-5 | 8.0e-4 | 0.284 |
The 95% intervals on power and FDR, from resampling worlds, are within ±0.003. Three things stand out. The single holdout is a fair filter against chance (1.3% pass) and almost none against costs (67% pass) or errors (99%). The deflated Sharpe ratio is the opposite: it handles chance and costs, about 2e-4, because it is applied to net returns over the whole sample, and it passes 96% of buggy ideas and 10% of confounded ones, because a drift or a bug that inflates both windows is a real Sharpe ratio. And the most elaborate stack has a false discovery rate of 0.34 to 0.45, nearly all of it buggy ideas: 3% of the population, passing at 18% to 28%. The statistical gates cannot see that mechanism because by construction it corrupts both windows equally. Only the audit can, so the audit's independence sets the pipeline's floor.
Stacking gates against chance
For pure-chance ideas we ran 15 million per mechanism and added gates one at a time.
| Gates | Observed | Product of marginals | Observed / product |
|---|---|---|---|
| N | 0.674 | 0.674 | 1.0 |
| N, IS | 6.69e-2 | 6.70e-2 | 1.0 |
| N, IS, BL | 9.29e-4 | 6.33e-4 | 1.5 |
| N, IS, BL, DEG | 1.63e-4 | 2.77e-5 | 5.9 |
| N, IS, BL, DEG, PL2 | 2.53e-5 | 1.38e-6 | 18.3 |
| all six | 2.39e-5 | 1.31e-6 | 18.2 |
Stack efficiency for IS, BL and DEG is 0.82 (95% interval 0.82 to 0.83); for the full stack 0.79 (0.78 to 0.79). The N gate is independent of everything else and, with it included, the product is exact: it thins a population without selecting from it. The in-sample gate and the degradation ratio, and the placebo and the in-sample gate, share the same tuned noise, so a null idea that got through the first is far more likely to get through the second. The same stack applied to real edges has a redundancy factor of 1.03: positives pass the gates together because they do pass, and nothing is shared that matters. Correlation between gates therefore shrinks what the stack removes without shrinking what it keeps. That asymmetry is why it is easy not to see.
The permutation baseline
We checked the measurement itself by permuting each gate's pass vector across ideas within a class, which keeps every marginal and destroys all dependence. The measured redundancy then returns to one: for chance ideas through IS and BL, 9.56e-4 permuted against 9.39e-4 for the product, and through IS, BL and DEG, 4.2e-5 against 4.1e-5. Redundancy is a property of the data lineage of the gates, not an artefact of how we count it.
Where the stack is redundant by construction
| Pipeline | Power | FDR | Chance | Confounded | Buggy | Fragile |
|---|---|---|---|---|---|---|
| full | 0.352 | 0.343 | 2.7e-5 | 7.1e-4 | 0.180 | 3.2e-5 |
| drop N | 0.520 | 0.343 | 4.0e-5 | 1.1e-3 | 0.266 | 5.0e-5 |
| drop IS | 0.353 | 0.350 | 4.0e-5 | 8.7e-4 | 0.180 | 4.0e-5 |
| drop BL | 0.352 | 0.343 | 2.7e-5 | 7.1e-4 | 0.180 | 3.2e-5 |
| drop DEG | 0.571 | 0.366 | 2.8e-4 | 5.5e-3 | 0.296 | 5.5e-3 |
| drop PL2 | 0.381 | 0.373 | 1.6e-4 | 1.2e-2 | 0.194 | 8.0e-5 |
| drop AUD | 0.370 | 0.496 | 2.6e-5 | 7.7e-4 | 0.360 | 3.5e-5 |
Dropping the blind gate changes nothing, to the last digit. The reason is arithmetic: the in-sample gate requires 0.8, the ratio requires 0.6 of that, so any idea reaching the ratio has a blind Sharpe of at least 0.48, above the 0.4 blind threshold. A gate whose threshold is implied by two others is dead weight and no simulation is needed to find it, but we did not see it until the marginal-contribution table was printed. The same table shows what each remaining gate is for. Dropping the ratio gate raises fragile ideas by a factor of about 160. Dropping the placebo raises confounded ideas by 17. Dropping the audit raises the false discovery rate from 0.34 to 0.50. Dropping N leaves the false discovery rate unchanged and raises power by 48%, which is the thinning at work.
A placebo judged on one draw is a coin flip
| Placebo design | Chance | Real | Costly | Confounded |
|---|---|---|---|---|
| PL0: one control, pass if not profitable | 0.500 | 0.499 | 0.500 | 0.181 |
| PL1: candidate beats 19 untuned controls | 0.293 | 0.985 | 0.818 | 0.294 |
| PL2: candidate beats 19 equally tuned controls | 0.050 | 0.916 | 0.496 | 0.050 |
A control with no edge has a Sharpe that is positive half the time, so a rule requiring it to be non-positive passes half of any idea that has no confound. It does catch confounded ideas, because the exposure-matched control keeps the drift and fails 82% of the time, but it halves the pass rate of real edges to 0.499. Adding it to the statistical stack takes power from 0.400 to 0.200 and changes the false discovery rate from 0.513 to 0.502. PL1 fixes the arithmetic and introduces a different error: the candidate was selected as the best of 20 variants and the controls were not, so the candidate beats them 29% of the time against a nominal 5%. PL2 gives the controls the same search and returns exactly 1/(B+1) = 5% on chance ideas, with 92% of real edges passing. The rule is that a control must be subject to the same selection as the thing it controls.
The degradation ratio punishes tuning
A ratio of blind to in-sample Sharpe is intended to detect over-fitting. For a true edge the in-sample estimate is a selected maximum, so the ratio is below one by construction and falls as the search widens.
| Variants K | Median ratio for real edges | 10th percentile | Power, N-IS-BL-DEG | Chance, IS and BL only | Chance, full stack |
|---|---|---|---|---|---|
| 1 | 0.88 | 0.54 | 0.407 | 1.4e-4 | 1.6e-5 |
| 5 | 0.74 | 0.45 | 0.450 | 5.7e-4 | 3.4e-5 |
| 20 | 0.66 | 0.40 | 0.401 | 1.4e-3 | 2.7e-5 |
| 100 | 0.59 | 0.36 | 0.320 | 2.9e-3 | 8.7e-6 |
| 500 | 0.54 | 0.33 | 0.248 | 4.6e-3 | 3.0e-6 |
With a fixed threshold of 0.6 the gate rejects a growing share of real edges as the search widens, and from K = 100 the median real edge that passes the in-sample gate fails it. Meanwhile the rubber-stamp risk of the in-sample and blind gates alone grows 32-fold from K = 1 to K = 500, which is the multiple-testing problem the deflated Sharpe ratio addresses. A threshold fixed in advance is wrong at every K but one. The threshold that keeps 90% of real edges ranges from 0.54 to 0.33 across this range, so it should be set as a function of the declared search size.
The audit is only as independent as the auditor
| Bug share | κ = 0 | κ = 0.4 | κ = 0.8 | κ = 0.95 |
|---|---|---|---|---|
| 1% | 0.151 | 0.181 | 0.216 | 0.237 |
| 3% | 0.344 | 0.396 | 0.452 | 0.482 |
| 10% | 0.634 | 0.686 | 0.733 | 0.757 |
Entries are the false discovery rate of the full pipeline. Chance of catching a surviving bug is 0.50, 0.37, 0.21 and 0.10 for those κ. At κ = 0.95 the audit nearly vanishes. If the second researcher uses the same tools, the same model family or the same checklist as the first, the audit is closer to a repeat of the author's own review than to a second opinion, and its contribution to the stack is a fraction of what its pass rate suggests. The statistical gates do not compensate, since they cannot see this mechanism.
How many looks does a blind window survive
A blind window stays blind only while it is consulted once. Each time the developer revises a failed idea and resubmits, the blind estimate is a new look, and the question is how many independent tests those looks are worth. We model L successive variants of a chance idea submitted until one passes, where the blind noise of two variants has correlation ρ_b. Variants of one idea on one history are strongly correlated, so ρ_b near 1 is the realistic end. The single-look pass rate is 0.0095. We define the effective number of independent looks as n_eff = ln(1 − R_L) / ln(1 − R_1).
| Looks L | ρ_b = 0 | ρ_b = 0.5 | ρ_b = 0.8 | ρ_b = 0.95 |
|---|---|---|---|---|
| 2 | 0.0186 (1.98) | 0.0175 (1.85) | 0.0151 (1.59) | 0.0125 (1.32) |
| 5 | 0.0457 (4.91) | 0.0363 (3.88) | 0.0255 (2.71) | 0.0166 (1.76) |
| 20 | 0.160 (18.4) | 0.090 (9.9) | 0.047 (5.1) | 0.023 (2.5) |
| 50 | 0.320 (40.5) | 0.144 (16.4) | 0.065 (7.1) | 0.028 (3.0) |
Entries are the chance-idea pass rate of the blind gate after L looks, with n_eff in brackets; each cell used 2 million ideas. Risk doubles after about 2 looks at ρ_b = 0 and 0.5, 3 looks at 0.8 and 8 looks at 0.95. At 20 looks the risk is 17 times the single-look figure for independent variants and 2.5 times for ρ_b = 0.95. The practical reading is that the raw count of looks always overcharges, and by more the more alike the variants are: 50 looks are worth 40.5 independent tests at ρ_b = 0 and 3.0 at ρ_b = 0.95. The number to charge is the effective one, which needs an estimate of ρ_b from the variants' own blind returns. Holding the risk at its single-look level after 20 looks requires raising the blind threshold from 0.40 to 0.65, 0.63, 0.57 or 0.49 for the four ρ_b, which in this population of strong edges leaves power at 0.974 at ρ_b = 0.95, against 0.971 for a single look, and at 1.00 for independent variants, because re-submission also gives real edges extra chances. Weaker edges would pay more. A ledger that records every submission and sets the threshold from n_eff keeps the window usable; one that records only passes does not.
Noise misspecified
We held thresholds fixed and raised φ in the world from 1.0 to 3.0, which stands for returns with more dependence and heavier tails than the thresholds were set for. Rubber-stamp risk on chance ideas:
| φ | Single holdout | IS and BL | N, IS, BL, DEG | Full stack |
|---|---|---|---|---|
| 1.0 | 9.5e-3 | 8.9e-4 | 9.6e-5 | 1.8e-5 |
| 1.5 | 3.0e-2 | 4.9e-3 | 6.8e-4 | 4.1e-5 |
| 2.0 | 5.3e-2 | 1.2e-2 | 1.8e-3 | 5.1e-5 |
| 3.0 | 9.7e-2 | 3.2e-2 | 5.4e-3 | 6.0e-5 |
Fixed thresholds degrade by a factor of 36 for IS and BL over this range and 10 for the single holdout. The full stack degrades by 3.3 because the placebo is built from the same series and inherits its noise: it self-calibrates, where a Sharpe threshold does not. That holds only if the placebo preserves the dependence structure of the data, which a naive shuffle does not.
Prevalence
The false discovery rates in Table 1 are conditional on 3% real ideas. At 0.5% the full pipeline's FDR is 0.76 and the single holdout's is 0.95; at 25% they are 0.05 and 0.23. The pipeline ordering does not change over this range. Power does not depend on prevalence.
Falsification and limits
We tested the simulator against what can be computed in closed form. The blind-gate pass rate on chance ideas is 0.00950 in simulation and 0.00946 by numerical integration over the cost distribution. The permutation baseline above is a null experiment for the measurement. The placebo on chance ideas lands on its nominal level for PL2 and on one half for PL0, both to three digits. These checks show the code does what we say; they do not show the model is the world.
The Gaussian summary-statistic approximation was checked against real returns in two ways. Annualised Sharpe ratios of block-bootstrapped industry portfolios over 6 and 18 years are close to Gaussian (kurtosis 3.06 and 3.08) with φ of 1.10. And for event-level data, the false-positive rate of a t ≥ 3 rule on resampled standardised daily returns, relative to the Gaussian 0.00135, is 5.3 times at 8 events, 2.6 at 15, 1.5 at 30, 1.25 at 50, 1.1 at 100, and 1.0 from 200 up. A sample-size gate therefore protects against heavy-tailed small samples only below about a hundred events. At the 200-event threshold we used, it is a pure thinning gate, as the model finds. This is a result about the estimator, not about trading: where event counts proxy something else, such as outlier dominance or capacity, the gate has another job that we have not modelled.
Three limits cannot be removed with this design, and we want to be plain about them. The size of the false discovery rates depends on the class shares and on the bug mechanism, both assumed; we report relative orderings and mechanism-level risks, not a forecast. The audit model is two numbers, a share and a correlation, and real review is richer. And the blind-window look model assumes each re-submission is a new variant; a researcher who steers on the score, rather than on pass or fail, extracts more from each look than we have modelled, and the adaptive-data-analysis bounds [7, 8] are the right tool for that case. Thresholds, window lengths and class constructions are ours; a pipeline with different numbers will give different magnitudes, though we expect the same qualitative ranking of mechanisms. We also did not model drift between the windows other than the fragile class.
What this means in practice
We would change four things in any pipeline of this shape. Report rubber-stamp risk per failure mechanism, from a null ensemble for each, instead of one pass rate. For each mechanism, ask which gates have noise the others do not share, and count only those. Treat a mechanism with one defender as a single point of failure and make that defender as independent of the author as possible: different tools, different data path, adversarial incentive. Check threshold arithmetic for entailment before running anything, since a gate implied by two others is free to delete and free to forget.
Gates against selection should scale with the declared search size. A control should undergo the same tuning as its candidate and be judged by rank against many draws, not by a single draw's sign. The blind window should be treated as a budget: record each look, charge it against the pass threshold with the effective number of independent looks, not the raw count, and do not let the score leave the gate when pass or fail is enough.
The estimator is short. Given any function that maps an input to pass or fail and a generator of inputs known to be false for a chosen mechanism, run a few hundred batches, report the pass fraction with a batch-resampled interval, and for every gate subset report the redundancy factor. A pipeline whose redundancy factor on chance ideas is above ten is not as selective as its gates suggest.
Reproducibility appendix
All simulations are in Python 3.14 with NumPy 2.4 and SciPy 1.17. Seeds: 11 (pipelines), 31 (leave-one-out), 41 plus K (variant sweep), 100 plus class index (stack efficiency), 404 (looks). The real-data inputs are the daily Fama-French five-factor file and the 49-industry file from the public data library [9]; the market factor drives the confound windows and the industries give φ. Runtimes on a 10-core laptop shared with other jobs: pipelines 44 s, stack efficiency 128 s, sweeps 175 s.
The tuned-noise sampler and the gates:
def tuned_noise(rng, M, K, rho):
zc = rng.standard_normal(M)
mx = norm.ppf(rng.random(M) ** (1.0 / K))
return np.sqrt(rho) * zc + np.sqrt(1 - rho) * mx
def observe(rng, W, P):
s1, s2 = np.sqrt(P["phi"] / P["Y1"]), np.sqrt(P["phi"] / P["Y2"])
G1 = W["g1"] + s1 * tuned_noise(rng, W["M"], P["K"], P["rho_v"])
S1 = G1 - W["c"]
S2 = W["g2"] + s2 * rng.standard_normal(W["M"]) - W["c"]
pl = W["drift1"][:, None] + s1 * tuned_noise(rng, W["M"] * P["B"], P["K"], P["rho_v"]).reshape(W["M"], P["B"])
return G1, S1, S2, pl
def gates(W, G1, S1, S2, pl, P):
return {
"N": W["tau"] * P["Y1"] >= P["nmin"],
"IS": S1 >= P["th1"],
"BL": S2 >= P["th2"],
"DEG": S2 / np.where(S1 > 0, S1, np.nan) >= P["ratio"],
"PL2": G1 > pl.max(1),
}The estimator:
def rubber_stamp_risk(pipeline, null_batch, n_batches, seed=0):
rng = np.random.default_rng(seed)
per_batch = np.array([pipeline(null_batch(rng)).mean() for _ in range(n_batches)])
r = per_batch.mean()
se = per_batch.std(ddof=1) / np.sqrt(n_batches)
return r, (max(r - 1.96 * se, 0.0), r + 1.96 * se)
def stack_efficiency(gate_matrix):
r_stack = gate_matrix.all(axis=1).mean()
r_marg = gate_matrix.mean(axis=0)
gse = np.log(r_stack) / np.log(r_marg).sum()
return gse, r_stack / np.prod(r_marg), gse * gate_matrix.shape[1]References
- Berger, R. L., Multiparameter hypothesis testing and acceptance sampling, Technometrics 24, 295-300, 1982.
- Bailey, D. H., López de Prado, M., The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality, Journal of Portfolio Management, 2014 (SSRN 2460551).
- Bailey, D. H., Borwein, J. M., López de Prado, M., Zhu, Q. J., The Probability of Backtest Overfitting, Journal of Computational Finance 20(4), 39-70, 2017.
- Harvey, C. R., Liu, Y., Zhu, H., ... and the Cross-Section of Expected Returns, Review of Financial Studies 29(1), 5-68, 2016.
- White, H., A Reality Check for Data Snooping, Econometrica 68(5), 1097-1126, 2000.
- Pollanen, M., The Price of Correlated Tests: How Strict Should a Model Release Gate Be?, arXiv:2610.00993, 2026.
- Dwork, C., Feldman, V., Hardt, M., Pitassi, T., Reingold, O., Roth, A., Generalization in Adaptive Data Analysis and Holdout Reuse, arXiv:1506.02629, 2015.
- Blum, A., Hardt, M., The Ladder: A Reliable Leaderboard for Machine Learning Competitions, ICML 2015, arXiv:1502.04585.
- French, K. R., Data Library: Fama/French 5 factors (2x3) daily and 49 industry portfolios daily, Dartmouth Tuck School of Business, https://mba.tuck.dartmouth.edu/pages/faculty/ken.french/data_library.html (files dated August 2026).
Cite this paper
@misc{forticia2026gate,
title = {{Gate stacks are not products: measuring the rubber-stamp risk of a research pipeline}},
author = {Richards, Cayden},
year = {2026},
month = oct,
publisher = {Forticia Research Institute},
howpublished = {\url{https://www.forticia.uk/papers/gate-stack-rubber-stamp-risk}},
note = {Paper, published online}
}Generated from this page’s metadata. Forticia does not assign DOIs to these papers.