Skip to content
Forticia
Research
Quantitative FinanceEquities, FX, futures. Factors and backtests, every run replayable.Computational BiologySequence, folding, simulation. Versioned, reproducible labs.Cultural IntelligencePhilosophy, governance, ethics. How institutions decide and answer for it.AI InstrumentationPrivate models. Multi-agent orchestration. Guardrails on write.View all research
Quantitative Finance
  • Equities, FX, futures
  • Factors and backtests
  • Every run logged and replayable
InfrastructurePapersPolarisLink™About
Sign inRequest access
Request access

Research

Quantitative FinanceEquities, FX, futures. Factors and backtests, every run replayable.Computational BiologySequence, folding, simulation. Versioned, reproducible labs.Cultural IntelligencePhilosophy, governance, ethics. How institutions decide and answer for it.AI InstrumentationPrivate models. Multi-agent orchestration. Guardrails on write.

Platform

InfrastructurePapersPolarisLink™About
Request accessSign in
Forticia

A private institute for computational research. It publishes original research and runs a governed environment where every run is logged and replayable.

Sign inSystem status

Research

Quantitative FinanceComputational BiologyCultural IntelligenceAI Instrumentation

Platform

InfrastructurePapersPolarisLink™StatusSign in

Institute

AboutRequest accessContactGitHubPolarisLink repository

Forticia publishes research and simulations. Nothing on this site is investment advice.

© 2026 ForticiaPrivacyTerms
Papers/Cultural Intelligence
CulturePaper

When Approval Stops Carrying Information

Author
Forticia Research Institute
Published
5 October 2026
Last updated
5 October 2026
Reading time
18 min
Cite this paper

On this page 0%

  1. Abstract
  2. Motivation
  3. Related work and what is new
  4. Method
  5. The gate
  6. Informativeness
  7. The collapse point
  8. Dynamics
  9. Estimating Φ from an audit
  10. Results
  11. Usual metrics cannot see the problem
  12. Informativeness collapses below a computable defect rate
  13. Where the zone sits
  14. After the rate jumps
  15. Levers
  16. Planted defects: a meter that works only if it is hidden
  17. How much audit is enough
  18. Falsification and limits
  19. What this means in practice
  20. Reproducibility appendix
  21. References
  22. Cite this paper

Abstract

An approval gate can look healthy on every usual metric while its decisions carry no information beyond what the proposer already supplied. We define gate informativeness as the conditional mutual information between the approver's decision and the proposal's true status, given the proposer's own grade, normalised by the uncertainty the grade leaves open. We derive the defect rate below which a rational approver stops reviewing, show by simulation that informativeness collapses there while approval rate, override rate and review time stay unremarkable, and compare four ways of restoring it. Increasing the approver's exposure to the cost of a missed defect, or cutting the cost of reviewing, beats random mandatory review. Planted known-bad items measure vigilance accurately only if the approver cannot recognise them. A permutation test detects collapse from a modest audited sample.

Motivation

Consider a queue of proposals generated by an automated process, each graded by the generator for risk, each requiring a human to approve before it takes effect. After a quarter the dashboard reads: ninety percent approved, three percent of the generator's recommendations overridden, median review time short, no incidents attributed to the gate. Everyone concludes the gate works.

Now suppose the defect rate of the generator is very low, one proposal in five hundred. The approver is busy, the generator's grade already flags the worst items, and a review costs minutes. Reading every proposal closely would be irrational for the approver, because the expected loss they personally bear from a defect they miss is small next to the time spent. The approver stamps. Every number on the dashboard stays the same. What has changed is that the approval no longer depends on whether the proposal is defective, except through the grade the generator supplied. The gate has become a function of the generator, and the institution is exposed to exactly the failure the gate was installed to prevent, at the moment the generator looks best.

This is a measurement problem before it is a motivation problem. We want a quantity that is zero when the approver adds nothing, grows as the approver's judgement tracks the truth beyond the grade, can be estimated from logs plus a modest audit, and does not move when someone games the usual metrics. We want to know at what defect rate it falls to zero, and what an institution can do about it at what price.

Related work and what is new

The difficulty is old. Bainbridge argued in 1983 that automating most of a task leaves the human the rare, hard residue of monitoring, with less practice at it [1]. Parasuraman and Manzey reviewed the evidence on complacency and automation bias and gave an integrated account in which attention plays the central role [2]. A recent longitudinal study of human review of machine-written code reports rising approval and falling scrutiny with experience, which the authors read as habituation under workload rather than rational calibration alone [3]. In economics, rational inattention formalises attention as costly information processing priced by Shannon information [4], and costly state verification asks when a principal should pay to check a report [5]. In machine learning the selective labels problem names the fact that outcomes are observed only for cases that received a given decision [6]. Airport screening has used planted fictional threats to measure screener performance on the job, and a study of how realistic those images are finds that unrealistic images are easier to detect, so that the measured hit rate overstates the real one [7].

What we could not find is a definition of rubber-stamping as a quantity: the information the approver's decision carries about the proposal's true status after conditioning on the proposer's own signal. We state it, derive the base rate at which a rational approver's value of reviewing falls below its cost, and use the gap between the approver's threshold and the organisation's to define a rubber-stamp zone. We then test, in one common simulated world, which levers shrink the zone, how fast an adaptive approver recovers after the defect rate jumps, whether planted defects are an honest meter, and how large an audit is needed to detect collapse. Our search was not exhaustive; the claim is about this combination, not about any single ingredient.

Method

The gate

Each proposal has a hidden status T, equal to one if it is defective. The proposer attaches a grade G from ten ordered classes. We model the grade as the decile of a latent score that is standard normal for sound proposals and shifted by d for defective ones; d is the proposer's discriminability and the decile edges are fixed from the marginal distribution at a baseline defect rate of 0.002 and are not recomputed when the defect rate changes. A real proposer's grading scale does not move because the world did. The approver sees G, may review, and decides D, equal to one for approve. A review flags a defective proposal with probability s and a sound one with probability f, both fixed at 0.9 and 0.05 unless stated. Approving a sound proposal is worth 1. Approving a defective one costs the organisation L and costs the approver Λ, the exposure they personally bear. Reviewing costs c. Rejecting is worth zero.

If the approver does not review, they approve when the grade's posterior defect probability q(G) satisfies (1 − q) − qΛ > 0, otherwise they reject.

Informativeness

We define the gate informativeness

Φ = I(D ; T | G) / H(T | G),

the mutual information between decision and status given grade, in bits, divided by the entropy the grade leaves unresolved. Φ is zero exactly when, within every grade, the decision is independent of the truth. It is one when the decision resolves the truth completely. The conditioning is the point. The unconditional quantity I(D ; T) credits the gate for everything the proposer's grade already told us, so a gate that copies a good proposer scores well on it. The conditional version credits only what the approver added.

The collapse point

For a proposal with grade posterior q small enough that the default is approval, reviewing beats not reviewing when q·s·Λ − (1 − q)·f > c, that is, when q exceeds

q* = (c + f) / (sΛ + f).

With our defaults and Λ = 5 this is about 0.022. If the largest posterior risk the proposer's grading leaves anywhere is below q*, a deterministic approver never reviews and Φ is exactly zero. The defect rate at which the highest-risk grade has posterior q* is the onset b*. We find it by scanning the defect rate with the grading edges held fixed. To model a noisy approver we let the review probability be logistic in the review's net value with temperature 0.02.

The organisation's own threshold uses L in place of Λ. Between the two lies the rubber-stamp zone: proposals the organisation would want reviewed and the approver rationally does not.

Dynamics

For drift we simulate a long stream of decisions. For the first three thousand steps the defect rate is 0.002 and the proposer's discriminability is 2.5. Then the defect rate jumps to 0.03 and discriminability falls to 1.5, so the proposer's grades become both less informative and miscalibrated relative to what the approver believes. The approver does not know the defect rate. They estimate it with an exponentially weighted average of defects they observe: those their own reviews flag, and those that escape and are discovered later. An escaped defect is discovered with probability 0.2, after a fixed lag of fifty steps. Each step's review probability is computed from the estimated rate and the current posterior. We report outcomes over the three thousand post-shift steps, with L = 50, Λ = 5, c = 0.05, and one thousand independent runs per condition.

Estimating Φ from an audit

In practice truth is unknown for most decisions. We assume a random subsample of decisions is audited so that T is revealed regardless of D; auditing only approved items would reintroduce the selective labels problem. From the audited counts we compute the plug-in conditional mutual information, its permutation null by shuffling the decision within grade, and a debiased estimate equal to the plug-in value minus the mean permutation value. The null for hypothesis "Φ = 0" is tested with 199 permutations.

Results

Usual metrics cannot see the problem

We built five gates on the same proposer at a defect rate of 0.05 and discriminability 2.5, and computed every quantity exactly from the joint distribution rather than by sampling.

Gate Approval Override Review rate I(D;T) / H(T) Φ Defects approved per 1000
A: follows the grade, never reviews 0.862 0.000 0.00 0.392 0.000 5.7
B: informative review of 30% of items 0.876 0.042 0.30 0.424 0.092 5.5
C: same approvals and overrides as B, decision unrelated to truth 0.876 0.042 0.30 0.253 0.000 14.5
D: reviews everything, review carries no information 0.900 0.210 1.00 0.000 0.000 45.0
E: reviews everything, informative 0.908 0.141 1.00 0.521 0.575 5.0

Override is defined against the proposer's own recommendation, which rejects only the top grade. Three things stand out. Gate A scores 0.392 on the unconditional measure, 92 percent of gate B's 0.424, while adding nothing: its Φ is zero. Gates B and C are identical on approval rate, override rate and review rate, as they must be, because C was constructed to have the same approval probability in every grade as B but independent of the truth. Any measure computed from decisions and grades alone cannot separate them. Φ separates them, 0.092 against zero. And gate D spends the most review effort of all and carries no information, which is what a time-spent metric rewards.

Informativeness collapses below a computable defect rate

Figure 1. Informativeness collapses below an onset defect rate: noisy approver, exposure 5, review cost 0.05, grade edges fixed at their baseline values, proposer discriminability 1.5, 2.5 and 3.5. Each ring marks the closed-form onset b*, where the top grade's posterior risk reaches q* (0.022). The fall above about 10 percent defect rate is a second effect, described in the paper: the approver rejects the riskiest grades without reviewing them.

For three discriminabilities and two values of the approver's exposure, we computed the onset b* and Φ at fractions of it, using the noisy approver.

Discriminability, Λ q* Onset b* Φ at 0.5 b* Φ at b* Φ at 2 b* Φ at 4 b* Defects escaping at 0.5 b*
1.5, 5 0.022 0.0039 0.014 0.101 0.252 0.313 95.4%
1.5, 20 0.0055 0.00095 0.011 0.082 0.214 0.269 95.6%
2.5, 5 0.022 0.0025 0.021 0.164 0.411 0.446 93.5%
2.5, 20 0.0055 0.00062 0.017 0.132 0.343 0.381 93.8%
3.5, 5 0.022 0.0022 0.024 0.187 0.489 0.538 93.0%
3.5, 20 0.0055 0.00056 0.020 0.158 0.403 0.451 93.0%

In every configuration, at half the onset rate, Φ is at most 0.024 and more than 92 percent of defects pass. At twice the onset it has recovered between 60 and 88 percent of its value at eight times the onset. The transition is sharp, and its location is set by q* and by how extreme the proposer's top grade is, not by anything about how hard the approver is working. A better proposer moves the onset only slightly in these runs, from 0.0039 to 0.0022 across the range of discriminability we tried at Λ = 5. A larger exposure moves it a lot: quadrupling Λ lowers the onset by a factor of about four.

The figure plots Φ against the defect rate on a log axis for discriminability 1.5, 2.5 and 3.5 at Λ = 5, with the noisy approver and the grade edges fixed at their baseline values, and marks each curve's onset b*. Each curve rises steeply through its onset and then flattens. Beyond a defect rate of about 0.1 the curves for discriminability 2.5 and 3.5 fall again, to about 0.3 at 0.15 for 2.5 and about 0.06 for 3.5, and the curve for 1.5 does the same near 0.2. That is a second, real effect and not the one this paper is about. With fixed grades, a rising defect rate eventually makes the top grades so likely to be defective, above about 0.62 posterior risk at these settings, that the approver rejects them without review, because even a favourable review would not make approval worthwhile. Within that grade the decision no longer depends on the truth, and Φ drops when the top grade crosses that ceiling. It occurs where defects are common, well above the onset, and we do not pursue it.

Where the zone sits

At the baseline defect rate of 0.002 and discriminability 2.5, the organisation's threshold with Λ replaced by L = 50 is 0.0022 and the approver's, with Λ = 5, is 0.022. The top grade, one tenth of all proposals, has posterior defect risk 0.018, inside the zone, and contains 88.7 percent of all defects. The approver, rationally, stamps the tenth of proposals that hold nearly nine tenths of the defects. The grade is doing its job. It concentrates risk in one class and the approver declines to look there because even concentrated risk is below their personal break-even.

After the rate jumps

An approver who learns the defect rate from what they observe is slow to notice a jump when they have been stamping, because stamping generates few observations.

Condition Φ before the jump Defects escaping, first 1500 steps Defects escaping, 3000 steps Total organisational cost Steps to half-recover review rate
Adaptive approver 0.025 [0.020, 0.032] 24.0 [23.7, 24.3] 44.4 [44.0, 44.8] 2256 [2235, 2276] 249 [237, 262]
Approver told the true rate 0.087 [0.076, 0.097] 19.0 [18.7, 19.3] 37.8 [37.4, 38.2] 1934 [1915, 1954] 12 [11, 14]

Intervals are 95 percent bootstrap intervals over one thousand runs. Total cost is L times escapes plus review cost plus the value of sound proposals wrongly rejected. Not knowing the rate costs about 17 percent of total cost relative to knowing it, and takes about 250 steps to half-recover the review rate. That is a real cost but a modest one next to what the exposure gap costs, below.

We had expected a stronger result. We hypothesised that the better the proposer had been before the jump, the longer the recovery would take, because a lower prior defect rate means less review and thus fewer defects observed. It does not hold. Sweeping the pre-jump defect rate from 0.001 to 0.02, half-recovery time stays between 212 and 349 steps with no monotone trend, and rises only at the highest rate (349 against 229 at the baseline). What drives recovery is how often escaped defects become visible. With a discovery probability of 0.05 half-recovery takes 598 steps and the excess escapes over an informed approver in the first 1500 steps are 10.8 against 4.7 at 0.2; with 1.0 they fall to 1.2.

Levers

We compare the ways an institution can push the zone out, first analytically at the baseline defect rate 0.002, then in the adaptive dynamic simulation. Costs in the static table are per 1000 decisions, including the organisation's cost of review and of rejecting sound proposals.

Lever Setting Φ Review rate Defects approved per 1000 Cost per 1000
None Λ = 5, c = 0.05 0.079 0.034 1.56 81.3
Random mandatory review 5% of items 0.092 0.082 1.49 82.7
Random mandatory review 20% of items 0.133 0.227 1.29 87.0
Random mandatory review 100% of items 0.419 1.000 0.20 109.9
Approver exposure Λ = 10 0.344 0.102 0.48 33.9
Approver exposure Λ = 50 0.372 0.118 0.39 31.3
Cheaper review c = 0.025 0.178 0.079 1.08 60.1
Cheaper review c = 0.0125 0.234 0.111 0.85 49.5
Cheaper review c = 0.003 0.272 0.143 0.71 42.9
Figure 2. What each lever buys, and what it costs: static comparison at defect rate 0.002 and discriminability 2.5, grade edges fixed at that baseline. Every line starts from the same baseline (exposure 5, review cost 0.05) and moves one lever. Cost is the organisation's cost per 1,000 decisions, including review and wrongly rejected sound proposals.

Raising the approver's exposure from 5 to 10 lifts Φ from 0.079 to 0.344 and more than halves the organisation's cost, because it moves the approver's threshold below the grade posteriors that hold the defects. Random mandatory review is the weakest lever per unit of Φ: five percent forced review buys 0.013 of Φ. This is because it samples proposals uniformly, and almost all of the uniform sample is sound. Halving the cost of a review more than doubles Φ, from 0.079 to 0.178.

In the dynamic simulation, with the post-jump window of three thousand steps, the ranking is the same.

Condition Φ before the jump Defects escaping, 3000 steps Total cost
Adaptive, no lever 0.025 44.4 [44.0, 44.8] 2256 [2235, 2276]
5% random mandatory review 0.041 42.1 [41.7, 42.5] 2153 [2135, 2174]
Cheaper review, c = 0.0125 0.067 37.6 [37.3, 38.0] 1917 [1898, 1935]
Exposure raised, Λ = 10 0.057 37.4 [37.0, 37.8] 1920 [1900, 1941]
Exposure raised, Λ = 50 0.370 18.9 [18.6, 19.2] 1146 [1132, 1158]

Exposure equal to the organisation's own loss halves the cost. We have not priced it. In the simulation the exposure is a free parameter. In a real institution, raising the penalty an approver bears for an undetected defect has costs we do not model: willingness to approve at all, defensive rejection, and fairness to someone blamed for a rare event. The honest reading is not that penalties are free, but that the gap between Λ and L is the largest single contributor to the zone, and that every lever should be judged by how much of it it closes.

Planted defects: a meter that works only if it is hidden

An institution can inject known-defective proposals at a small rate and measure the fraction the approver catches. We simulated a two percent injection rate with the approver penalised for missing an injected item. If the approver cannot tell injected items from real ones, the injected catch rate is an unbiased meter of the real one.

Probability the approver recognises an injected item Catch rate on injected items Catch rate on real defects
0.00 0.155 [0.144, 0.167] 0.153 [0.132, 0.176]
0.25 0.337 [0.327, 0.348] 0.132 [0.114, 0.153]
0.50 0.524 [0.516, 0.533] 0.132 [0.114, 0.152]
0.90 0.826 [0.821, 0.832] 0.132 [0.113, 0.150]

At zero recognition the two agree. At 0.9 the meter reads 0.83 while the real catch rate is 0.13, a factor of six. Real defects at this base rate are two per thousand, so a two percent injection gives roughly ten times as many measurement events as the real defects supply. The meter is valuable precisely because it is plentiful, and it is only valid if the injected items are indistinguishable. Airport screening has met the same issue: if projected threats look unrealistic, the hit rate on them overstates the real one [7].

How much audit is enough

We simulated audited samples of various sizes from gates whose true Φ was set by forcing a fraction of items to receive an informative review, and ran the permutation test at the 0.05 level.

Defect rate True Φ n = 1000 n = 5000 n = 20000 Error of the debiased estimate at n = 5000
0.002 0.1 not run 0.77 0.99 0.090
0.005 0.1 not run 0.96 1.00 0.056
0.02 0.1 0.83 1.00 1.00 0.030
0.05 0.1 1.00 1.00 1.00 0.019
0.05 0.05 0.86 1.00 1.00 0.014
Figure 3. How large an audit detects collapse: rejection rate of the permutation test for informativeness above zero, at the 0.05 level, discriminability 2.5 with grade edges fixed at baseline, against the number of audited decisions. 200 runs per point for defect rates below 0.02 and 300 from 0.02 up. Tabs choose the true informativeness.

Cells are rejection rates with three hundred runs per cell for rates of 0.02 and above and two hundred below. Detection is cheap when defects are not rare and expensive when they are: at 0.002, twenty thousand audited decisions detect a Φ of 0.05 with power 0.94, and ten thousand only 0.78. Estimating Φ's size, as opposed to detecting that it is positive, is much harder: the root mean square error at five thousand audited decisions is 0.09 at the lowest defect rate, as large as the effect. The test controls its error: with an approver whose decisions are random within grade, the rejection rate under the null was at most 0.060 across defect rates 0.02, 0.05 and 0.1 and sample sizes 500 to 50000, in one thousand runs per cell, which is within Monte Carlo error of 0.05 (the standard error of a rate near 0.05 over one thousand runs is about 0.007).

Falsification and limits

We tried to break the claims in four ways. First, the null: an approver whose decisions are noise within grade has a mean estimated Φ within 0.0007 of zero in every cell, and its rejection rate stayed at or below 0.060. Second, the effort-without-information case, gate D, reviews everything and has Φ of zero, so the measure does not reward effort. Third, the sensitivity of the dynamics: we varied the estimator's window, the discovery probability, the lag, the size of the jump, the post-jump discriminability and the noise in the approver's choice, and the quantities that survive are the ranking of levers and the existence of a sharp onset. What does not survive is the hysteresis hypothesis above, which we report as a negative result. Fourth, the collapse point is verified against the closed form: with a deterministic approver the simulated onset coincides with the point where the top grade's posterior reaches q*, as it must by construction, so this checks the code, not the economics.

Several limits are structural. Truth is binary and known after audit; real proposals are graded by degree and some defects are never discovered. The approver's rule is a stylised rational model with noise, and real approvers also respond to norms, fatigue and the interface, which we do not model. The grades are exactly calibrated in the baseline and miscalibrated only through the drift experiment. The audit must be random across decisions; if only approved items are audited, the estimator is biased in a way we have not characterised. Φ is a measure of informativeness, not correctness: a gate can be informative and still wrongly thresholded. And we normalise by the residual uncertainty after the grade, which is tiny at low defect rates, so the absolute information in bits per decision should be read alongside it: 0.015 bits for gate B and 0.096 bits for gate E in the first table.

What this means in practice

An institution that places a human approval behind an automated proposer should, once, define what the approver is expected to add beyond the proposer's own signal, and measure that, with a randomly drawn audit that reveals truth regardless of the decision. Approval rate, override rate and time spent do not measure it, by construction.

It should compute, for its own proposer, the defect rate below which its approver's break-even leaves the highest-risk class unreviewed, and ask whether that rate is below where the proposer currently operates. The calculation needs only an estimate of review cost, of the approver's personal exposure, and of how well the proposer's grades separate defects. When the answer is that the approver would rationally stamp, the remedy is to close the gap between the approver's and the organisation's valuation of a miss, or to lower the cost of looking. Random forced review is a weak substitute.

If it plants known defects as a meter, it should spend as much effort on making them indistinguishable as on generating them, and should treat any meter whose items the approver can recognise as an upper bound on vigilance, not an estimate.

Reproducibility appendix

All code is Python with numpy and scipy. Seeds: dynamics 2024, sensitivity 31, policy comparisons 2024, canary study 77, estimator power 7, null study 11. Runtimes on a ten-core laptop: static sweeps under one second, the policy comparison with one thousand runs per condition about half a minute, the estimator power study about twenty seconds, the low-defect-rate extension six seconds, the sensitivity sweep about two and a half minutes.

The approver's review value, the threshold, and the information measures:

python
def review_value(q, s, f, lam, c):
    u0 = np.maximum(0.0, (1 - q) - q * lam)
    u1 = (1 - q) * (1 - f) - q * (1 - s) * lam
    return u1 - c - u0

def q_star(s, f, lam, c):
    return (c + f) / (s * lam + f)

def entropy(p):
    p = np.asarray(p, float); p = p[p > 0]
    return -(p * np.log2(p)).sum()

def cmi_from_joint(j):            # j[g, t, d] joint probabilities
    j = j / j.sum(); pg = j.sum(axis=(1, 2)); tot = htot = 0.0
    for g in range(j.shape[0]):
        if pg[g] <= 0: continue
        w = j[g] / pg[g]; pt = w.sum(axis=1); pd = w.sum(axis=0)
        tot += pg[g] * (entropy(pt) + entropy(pd) - entropy(w.ravel()))
        htot += pg[g] * entropy(pt)
    return tot, htot               # Phi = tot / htot

The audit estimator and its permutation test, written for count tables of shape grade by truth by decision:

python
def cmi_batch(cnt):
    cnt = np.asarray(cnt, float)
    n = cnt.sum(axis=(-1, -2, -3), keepdims=True)
    ng = cnt.sum(axis=(-1, -2), keepdims=True)
    nt = cnt.sum(axis=-1, keepdims=True)
    nd = cnt.sum(axis=-2, keepdims=True)
    with np.errstate(divide='ignore', invalid='ignore'):
        term = np.where(cnt > 0, cnt * np.log(cnt * ng / (nt * nd)), 0.0)
    return term.sum(axis=(-1, -2, -3)) / n[..., 0, 0, 0] / np.log(2)

def perm_null(cnt, B, rng):       # shuffle D within each grade
    K = cnt.shape[0]; out = np.zeros((B, K, 2, 2))
    for g in range(K):
        n11, n10, n01, n00 = cnt[g, 1, 1], cnt[g, 1, 0], cnt[g, 0, 1], cnt[g, 0, 0]
        ng = int(n11 + n10 + n01 + n00); nt1 = int(n11 + n10); nd1 = int(n11 + n01)
        if ng == 0: continue
        x = rng.hypergeometric(nt1, ng - nt1, nd1, size=B)
        out[:, g, 1, 1] = x; out[:, g, 1, 0] = nt1 - x
        out[:, g, 0, 1] = nd1 - x; out[:, g, 0, 0] = ng - nt1 - nd1 + x
    return out

The test rejects when the observed value is not exceeded by the permutation values at the 0.05 level, using (1 + count of permutation values at least as large) divided by (B + 1).

  • governance
  • oversight
  • rubber-stamping
  • information theory
  • automation
  • simulation

References

  1. Bainbridge, L. (1983). Ironies of automation. Automatica, 19(6), 775-779. doi:10.1016/0005-1098(83)90046-8
  2. Parasuraman, R., and Manzey, D. H. (2010). Complacency and bias in human use of automation: an attentional integration. Human Factors, 52(3), 381-410. doi:10.1177/0018720810376055
  3. Yu, H., Liu, L., Jiang, X., Jia, Y., Wang, S., Qian, P., and Chen, Y. (2026). Habituation at the gate: rising approval and declining scrutiny in human review of AI agent code. arXiv:2606.22721.
  4. Sims, C. A. (2003). Implications of rational inattention. Journal of Monetary Economics, 50(3), 665-690. doi:10.1016/S0304-3932(03)00029-1
  5. Townsend, R. M. (1979). Optimal contracts and competitive markets with costly state verification. Journal of Economic Theory, 21(2), 265-293. doi:10.1016/0022-0531(79)90031-0
  6. Lakkaraju, H., Kleinberg, J., Leskovec, J., and Ludwig, J. (2017). The selective labels problem: evaluating algorithmic predictions in the presence of unobservables. Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 275-284. doi:10.1145/3097983.3098066
  7. Riz à Porta, Sterchi, and Schwaninger (2022). How realistic is threat image projection for X-ray baggage screening? Sensors, 22(6), 2220. doi:10.3390/s22062220

Cite this paper

@misc{forticia2026when,
  title        = {{When Approval Stops Carrying Information}},
  author       = {{Forticia Research Institute}},
  year         = {2026},
  month        = oct,
  publisher    = {Forticia Research Institute},
  howpublished = {\url{https://www.forticia.uk/papers/rubber-stamp-gate-informativeness}},
  note         = {Paper, published online}
}

Generated from this page’s metadata. Forticia does not assign DOIs to these papers.

NewerNever Tune a Corpse
All papers
Informativeness
0.00.10.20.30.40.50.60.1%1.0%10.0%Defect rate b, log scale
Share of defects that escape
0%25%50%75%100%0.1%1.0%10.0%Defect rate b, log scale
  • Discriminability 1.5
  • Discriminability 2.5
  • Discriminability 3.5
Two charts against defect rate on a log axis from 0.03 to 40 percent, for proposer discriminability 1.5, 2.5 and 3.5, with a ring on each curve at its closed-form onset: 1.5 (0.39 percent), 2.5 (0.25 percent), 3.5 (0.22 percent). Left: informativeness stays near zero below the onsets, rises steeply through them and flattens. Above a defect rate of about 10 percent the curves fall again, reaching 0.41, 0.22, 0.04 at 40 percent. Right: the share of defects that escape is above 99 percent at the lowest defect rate and falls to between 0.4 and 4.3 percent at 40 percent.
PanelSeriesPoints
InformativenessDiscriminability 1.50.0%: 0.0; 0.0%: 0.0; 0.0%: 0.0; 0.0%: 0.0; 0.0%: 0.0; 0.1%: 0.0; 0.1%: 0.0; 0.1%: 0.0; 0.1%: 0.0; 0.1%: 0.0; 0.1%: 0.0; 0.1%: 0.0; 0.1%: 0.0; 0.1%: 0.0; 0.2%: 0.0; 0.2%: 0.0; 0.2%: 0.0; 0.2%: 0.0; 0.3%: 0.0; 0.3%: 0.0; 0.3%: 0.1; 0.4%: 0.1; 0.4%: 0.1; 0.5%: 0.2; 0.6%: 0.2; 0.6%: 0.2; 0.7%: 0.2; 0.8%: 0.3; 0.9%: 0.3; 1.0%: 0.3; 1.2%: 0.3; 1.3%: 0.3; 1.5%: 0.3; 1.7%: 0.3; 1.9%: 0.3; 2.1%: 0.4; 2.4%: 0.4; 2.7%: 0.4; 3.1%: 0.4; 3.5%: 0.4; 3.9%: 0.4; 4.5%: 0.5; 5.0%: 0.5; 5.7%: 0.5; 6.4%: 0.5; 7.3%: 0.5; 8.2%: 0.5; 9.3%: 0.5; 10.5%: 0.5; 11.8%: 0.5; 13.3%: 0.6; 15.1%: 0.6; 17.0%: 0.6; 19.2%: 0.5; 21.7%: 0.4; 24.6%: 0.4; 27.7%: 0.4; 31.3%: 0.4; 35.4%: 0.4; 40.0%: 0.4
InformativenessDiscriminability 2.50.0%: 0.0; 0.0%: 0.0; 0.0%: 0.0; 0.0%: 0.0; 0.0%: 0.0; 0.1%: 0.0; 0.1%: 0.0; 0.1%: 0.0; 0.1%: 0.0; 0.1%: 0.0; 0.1%: 0.0; 0.1%: 0.0; 0.1%: 0.0; 0.1%: 0.0; 0.2%: 0.0; 0.2%: 0.1; 0.2%: 0.1; 0.2%: 0.1; 0.3%: 0.2; 0.3%: 0.3; 0.3%: 0.3; 0.4%: 0.4; 0.4%: 0.4; 0.5%: 0.4; 0.6%: 0.4; 0.6%: 0.4; 0.7%: 0.4; 0.8%: 0.4; 0.9%: 0.4; 1.0%: 0.4; 1.2%: 0.5; 1.3%: 0.5; 1.5%: 0.5; 1.7%: 0.5; 1.9%: 0.5; 2.1%: 0.5; 2.4%: 0.5; 2.7%: 0.5; 3.1%: 0.5; 3.5%: 0.5; 3.9%: 0.5; 4.5%: 0.5; 5.0%: 0.5; 5.7%: 0.5; 6.4%: 0.5; 7.3%: 0.5; 8.2%: 0.5; 9.3%: 0.5; 10.5%: 0.5; 11.8%: 0.5; 13.3%: 0.5; 15.1%: 0.3; 17.0%: 0.2; 19.2%: 0.2; 21.7%: 0.2; 24.6%: 0.2; 27.7%: 0.2; 31.3%: 0.2; 35.4%: 0.2; 40.0%: 0.2
InformativenessDiscriminability 3.50.0%: 0.0; 0.0%: 0.0; 0.0%: 0.0; 0.0%: 0.0; 0.0%: 0.0; 0.1%: 0.0; 0.1%: 0.0; 0.1%: 0.0; 0.1%: 0.0; 0.1%: 0.0; 0.1%: 0.0; 0.1%: 0.0; 0.1%: 0.0; 0.1%: 0.0; 0.2%: 0.1; 0.2%: 0.1; 0.2%: 0.2; 0.2%: 0.2; 0.3%: 0.3; 0.3%: 0.4; 0.3%: 0.4; 0.4%: 0.5; 0.4%: 0.5; 0.5%: 0.5; 0.6%: 0.5; 0.6%: 0.5; 0.7%: 0.5; 0.8%: 0.5; 0.9%: 0.5; 1.0%: 0.5; 1.2%: 0.6; 1.3%: 0.6; 1.5%: 0.6; 1.7%: 0.6; 1.9%: 0.6; 2.1%: 0.6; 2.4%: 0.6; 2.7%: 0.6; 3.1%: 0.6; 3.5%: 0.6; 3.9%: 0.6; 4.5%: 0.6; 5.0%: 0.6; 5.7%: 0.6; 6.4%: 0.6; 7.3%: 0.6; 8.2%: 0.6; 9.3%: 0.6; 10.5%: 0.6; 11.8%: 0.5; 13.3%: 0.3; 15.1%: 0.1; 17.0%: 0.0; 19.2%: 0.0; 21.7%: 0.0; 24.6%: 0.0; 27.7%: 0.0; 31.3%: 0.0; 35.4%: 0.0; 40.0%: 0.0
Share of defects that escapeDiscriminability 1.50.0%: 99%; 0.0%: 99%; 0.0%: 99%; 0.0%: 99%; 0.0%: 99%; 0.1%: 99%; 0.1%: 99%; 0.1%: 99%; 0.1%: 99%; 0.1%: 99%; 0.1%: 98%; 0.1%: 98%; 0.1%: 98%; 0.1%: 97%; 0.2%: 97%; 0.2%: 96%; 0.2%: 94%; 0.2%: 92%; 0.3%: 89%; 0.3%: 85%; 0.3%: 79%; 0.4%: 72%; 0.4%: 63%; 0.5%: 56%; 0.6%: 51%; 0.6%: 48%; 0.7%: 47%; 0.8%: 46%; 0.9%: 45%; 1.0%: 44%; 1.2%: 42%; 1.3%: 40%; 1.5%: 38%; 1.7%: 35%; 1.9%: 32%; 2.1%: 30%; 2.4%: 28%; 2.7%: 26%; 3.1%: 24%; 3.5%: 23%; 3.9%: 21%; 4.5%: 20%; 5.0%: 18%; 5.7%: 17%; 6.4%: 16%; 7.3%: 15%; 8.2%: 14%; 9.3%: 14%; 10.5%: 13%; 11.8%: 12%; 13.3%: 12%; 15.1%: 12%; 17.0%: 11%; 19.2%: 11%; 21.7%: 8%; 24.6%: 5%; 27.7%: 5%; 31.3%: 4%; 35.4%: 4%; 40.0%: 4%
Share of defects that escapeDiscriminability 2.50.0%: 99%; 0.0%: 99%; 0.0%: 99%; 0.0%: 99%; 0.0%: 98%; 0.1%: 98%; 0.1%: 98%; 0.1%: 98%; 0.1%: 97%; 0.1%: 97%; 0.1%: 96%; 0.1%: 95%; 0.1%: 93%; 0.1%: 91%; 0.2%: 87%; 0.2%: 82%; 0.2%: 74%; 0.2%: 64%; 0.3%: 52%; 0.3%: 40%; 0.3%: 31%; 0.4%: 25%; 0.4%: 22%; 0.5%: 21%; 0.6%: 20%; 0.6%: 20%; 0.7%: 20%; 0.8%: 20%; 0.9%: 20%; 1.0%: 20%; 1.2%: 20%; 1.3%: 20%; 1.5%: 20%; 1.7%: 20%; 1.9%: 20%; 2.1%: 19%; 2.4%: 19%; 2.7%: 19%; 3.1%: 18%; 3.5%: 17%; 3.9%: 16%; 4.5%: 15%; 5.0%: 15%; 5.7%: 14%; 6.4%: 14%; 7.3%: 14%; 8.2%: 13%; 9.3%: 13%; 10.5%: 12%; 11.8%: 12%; 13.3%: 11%; 15.1%: 8%; 17.0%: 4%; 19.2%: 2%; 21.7%: 2%; 24.6%: 2%; 27.7%: 2%; 31.3%: 2%; 35.4%: 2%; 40.0%: 2%
Share of defects that escapeDiscriminability 3.50.0%: 99%; 0.0%: 99%; 0.0%: 99%; 0.0%: 98%; 0.0%: 98%; 0.1%: 98%; 0.1%: 98%; 0.1%: 97%; 0.1%: 97%; 0.1%: 96%; 0.1%: 94%; 0.1%: 93%; 0.1%: 90%; 0.1%: 86%; 0.2%: 81%; 0.2%: 73%; 0.2%: 62%; 0.2%: 48%; 0.3%: 35%; 0.3%: 24%; 0.3%: 17%; 0.4%: 14%; 0.4%: 12%; 0.5%: 11%; 0.6%: 11%; 0.6%: 11%; 0.7%: 11%; 0.8%: 11%; 0.9%: 11%; 1.0%: 11%; 1.2%: 11%; 1.3%: 11%; 1.5%: 11%; 1.7%: 11%; 1.9%: 11%; 2.1%: 11%; 2.4%: 11%; 2.7%: 11%; 3.1%: 11%; 3.5%: 11%; 3.9%: 11%; 4.5%: 11%; 5.0%: 11%; 5.7%: 11%; 6.4%: 11%; 7.3%: 11%; 8.2%: 11%; 9.3%: 11%; 10.5%: 11%; 11.8%: 11%; 13.3%: 8%; 15.1%: 3%; 17.0%: 1%; 19.2%: 1%; 21.7%: 1%; 24.6%: 1%; 27.7%: 1%; 31.3%: 0%; 35.4%: 0%; 40.0%: 0%
0.00.10.20.30.40.5Informativeness20406080100120Organisational cost per 1,000 decisionsNo lever20% reviewedEvery item reviewedReview cost 0.025Review cost 0.003Exposure 10Exposure 20 and 50
  • Random mandatory review
  • Cheaper review
  • Raise approver exposure
Line chart of informativeness against organisational cost per 1,000 decisions at a defect rate of 0.002. All three levers start from the same point, informativeness 0.079 at a cost of 81. Random mandatory review moves right and slowly up, reaching 0.42 at a cost of 110 when every item is reviewed. Cheaper review moves left and up, reaching 0.27 at a cost of 43. Raising the approver exposure moves left and up fastest, reaching 0.34 at a cost of 34 and 0.37 at a cost of about 31.
Organisational cost per 1,000 decisionsRandom mandatory reviewCheaper reviewRaise approver exposure
310.4
310.4
340.3
430.3
450.3
500.2
600.2
810.10.10.1
820.1
820.1
830.1
840.1
870.1
960.2
1100.4
True informativeness
0%25%50%75%100%Rejection rate of the permutation test5001,0002,0005,00010k20k50k100kAudited decisions
  • Defect rate 0.2%
  • Defect rate 0.5%
  • Defect rate 2%
  • Defect rate 5%
  • Defect rate 10%
Audited decisions
Move across the chart, or focus it and use the arrow keys
Line chart of the rejection rate of the permutation test against the number of audited decisions on a log axis, for defect rates of 0.2, 0.5, 2, 5 and 10 percent, with tabs for a true informativeness of 0.05, 0.1 and 0.2. Power rises with sample size and reaches 100 percent first for the highest defect rates. At a defect rate of 0.2 percent and a true informativeness of 0.05, power is 78 percent at 10,000 audited decisions and 93.5 percent at 20,000.
Audited decisionsDefect rate 0.2%Defect rate 0.5%Defect rate 2%Defect rate 5%Defect rate 10%
50056%84%95%
1,00083%100%100%
2,00044%72%97%100%100%
5,00077%96%100%100%100%
10k95%100%100%100%100%
20k100%100%100%100%100%
50k100%100%100%100%100%
100k100%100%100%100%100%