Abstract
Human approval of agent actions fails quietly when almost every request is fine. We simulate a reviewer who sees agent requests at a base rate of 0.5% harmful, and compare tiered asking against tiered asking with canaries: known-bad requests injected into the queue. Canaries do two jobs, restoring the reviewer's sense that bad requests exist and measuring whether the reviewer is still looking. At equal attention, canaries lose below about 25 reviews per shift and win above it: escape of harmful requests falls from 43.8% to 28.4% at 229 reviews, which plain tiering cannot reach at any budget. As a gate, a CUSUM on canary outcomes beat a break schedule of equal length by 3.2 points only when lapses are episodic. A canary the reviewer can recognise silently removes both benefits.
Motivation
A reviewer approves agent requests for a day. Out of 1,200 requests, six would be harmful. After a few hundred approvals the work has the shape of a rhythm: look, click, next. Anthropic reported that Claude Code users approve 93% of permission prompts and called the result approval fatigue [1]. A study of 400 reviewers and 11,429 code reviews found approval rates rising with experience while comment volume fell 22% and latency rose, a pattern the authors read as habituation under growing workload [2]. Both facts say that approving is cheap and attention is not. They do not say what to do about it.
Two designs are on the table. One routes by risk, so that fewer requests reach the human, which is how most deployed systems reduce prompts. The other is old in other fields: put known positives into the stream. Airport screening has done this for decades under the name threat image projection, where synthetic threats are overlaid on real baggage images to raise the number of threats a screener sees, give frequent feedback that real threats rarely provide, and compute each screener's hit rate for compliance [3]. Searching for prior work, we found short mentions of injecting known-bad cases into agent approval streams to watch whether overseers catch them, but no primary source that measures the idea or prices it in reviewer attention. That cost is the question. A canary is itself a request that needs a review.
We ask three things. At equal attention, do canaries reduce the harm that passes? Does a gate built from canary outcomes beat a plain break schedule? And what breaks the design?
All results here are simulations. The reviewer model is calibrated to published visual-search data and otherwise built from assumptions we state and vary. No human reviewed anything. Nothing here describes the performance of a deployed system, Forticia's or anyone's. Forticia's public description of governed AI is drafts by default, lead approval to write and a full audit trail, which places a human in exactly this position, and the design question is ours.
Related work and what is new
Wolfe and colleagues showed that low target prevalence raises miss rates in visual search: in one experiment misses were 0.20 at 50% prevalence and 0.46 at 2%, with false alarms falling from 0.19 to 0.01. In signal detection terms the effect was a shift in criterion, not in sensitivity [4]. Schwark and colleagues found that false feedback telling observers they had missed a target raised detection of low-prevalence targets at the price of more false alarms [5]. Riz à Porta and colleagues describe threat image projection and its three roles: training exposure, feedback and compliance monitoring [3].
For agents, Turan models the reviewer as one whose reliability declines once the number of reviews passes a capacity, and shows in simulation that escalating every action can be less safe than escalating fewer, an inverted-U between escalation rate and safety [6]. Yu and colleagues measure habituation in real code review [2]. Page's CUSUM is the standard sequential test for a shift in a monitored rate [7].
What is new: a matched-attention comparison in which injected canaries are charged as reviews; a reviewer model whose sensitivity to prevalence is calibrated to Wolfe's data and whose sense of prevalence is updated from what the reviewer itself flags; a separation of the two jobs canaries do, prevalence restoration and vigilance gating, by ablation; a gate-versus-schedule test at equal break minutes; and a measurement of what happens when canaries are recognisable.
Method
The stream and the guard
A shift has 1,200 request slots over 480 minutes. Each request is harmful with probability p = 0.005. A guard produces a score, normal with mean 3.0 for harmful requests and 0 for benign ones (an area under the ROC curve of 0.98, a strong guard). A request goes to the human if its score exceeds a threshold θ, and is auto-approved otherwise, so harmful requests below the threshold escape. θ = minus infinity is always ask.
The reviewer
The reviewer is a signal detection observer. A harmful request produces evidence x drawn from a normal with mean d′ and unit variance, a benign one from a normal with mean 0. The reviewer flags if x exceeds d′/2 + c. Three things move d′ and c.
Sensitivity. d′ = d0 (1 − κ f), where f in [0, 1] is fatigue, d0 = 2.3 and κ = 0.5. Each review raises fatigue by 0.02 (1 − f)(1 + σ), where σ in [0, 1] is how similar the item is to recent items (Beta(4, 2)), and fatigue decays with a time constant of 60 minutes. A lapse process (hazard 0.1 per hour, mean duration 60 minutes) multiplies d′ by 0.3, a stand-in for disengagement, until a break or the lapse ends.
Criterion. c = α ln((1 − q̂)/q̂)/d′, where q̂ is the reviewer's perceived prevalence and α = 0.7 sets how much they track it. q̂ is an exponentially weighted average (rate 0.02) of an indicator that is 1 when the reviewer flagged an item or learned a canary was a canary, floored at 0.02. Because q̂ is fed by the reviewer's own flags, low prevalence lowers flags, which lowers q̂, which raises the criterion: a feedback loop that we assume rather than take from Wolfe's data.
Calibration to Wolfe: with d0 = 2.3 and α = 0.7 and no fatigue, the model gives a miss rate and false alarm rate of 0.125 and 0.125 at 50% prevalence (observed 0.20 and 0.19) and 0.514 and 0.010 at 2% (observed 0.46 and 0.01) [4]. The floor of 0.02 on q̂ is the lowest prevalence Wolfe tested; below it we do not extrapolate the effect, which is conservative, since a lower floor would make the plain reviewer worse.
Canaries and the gate
With canary rate ρ, each request slot is accompanied by a canary with probability ρ. A canary is a harmful-looking request that does nothing if approved. It is charged as a review and raises fatigue like any other. The reviewer is told the outcome, so q̂ is updated with 1 whether or not they caught it. A fraction ℓ of canaries are recognisable: the reviewer flags them with probability 0.98, and learns nothing about prevalence.
The gate is a CUSUM [7] on canary outcomes. It tests a catch probability of 0.85 against 0.60. A canary caught adds ln(0.60/0.85) to the statistic, a miss adds ln(0.40/0.15), the statistic is floored at 0, and it fires above h (4 by default). On a firing the reviewer takes a 15-minute break that cuts fatigue and ends any lapse. The comparison is a fixed schedule with evenly spaced breaks of the same total length.
Measures
For each configuration we simulate S shifts and report: escape, the fraction of harmful requests that pass (auto-approved or missed by the human); reviews per shift, including canaries; false rejections per shift (benign requests the reviewer blocks); and break minutes. Intervals are 95% bootstrap percentile intervals over shifts. For the regime map we use a cost J = C × (harmful requests that pass) + reviews + 5 × false rejections + 4 × break minutes, with C = 100, 1,000 or 10,000 review-equivalents per escaped harmful request.
Every parameter other than d0 and α is an assumption. We vary all of them below.
Results
Tiering, and why more asking is not safer
First the no-canary world, 4,000 shifts per row.
| Threshold θ | Reviews per shift | Escape | Share of routed harm the human misses | False rejections per shift |
|---|---|---|---|---|
| always ask | 1,200 | 90.6% [90.2, 90.9] | 90.6% | 4.57 |
| 0.0 | 603 | 84.9% [84.5, 85.4] | 84.9% | 3.44 |
| 0.5 | 374 | 79.8% [79.3, 80.3] | 79.7% | 2.78 |
| 1.0 | 195 | 71.4% [70.6, 72.0] | 70.8% | 2.08 |
| 1.5 | 85.5 | 55.9% [55.1, 56.6] | 52.8% | 1.80 |
| 2.0 | 32.3 | 44.4% [43.8, 45.1] | 34.2% | 1.28 |
| 2.5 | 11.6 | 45.2% [44.4, 45.9] | 20.8% | 0.75 |
| 3.0 | 4.6 | 57.9% [57.3, 58.6] | 14.4% | 0.25 |
The always-ask reviewer lets 90.6% of harmful requests through, a miss rate that in the model comes from two mechanisms acting together: a perceived prevalence driven to its floor and fatigue accumulated across 1,200 reviews. As the guard narrows the queue, the human's miss rate falls from 91% to 14%, because the reviewer now sees harmful requests often enough to expect them (at θ = 2, about 16% of routed requests are harmful, against 0.5% at the base). Escape bottoms out near 44% at θ between 2 and 2.5 because the guard itself starts to drop harmful requests below the threshold. Past that point asking less is worse. Turan's inverted-U appears here too, though the cause in our model is prevalence as much as fatigue [6].
The practical reading: without canaries the reviewer's useful attention saturates at about 20 reviews per shift.
Matched attention
For each attention budget we took the best configuration from a fine grid (3,000 shifts each) in each family, then re-ran the winner on fresh seeds with 20,000 shifts. "Canary" is canaries without the gate, "gate" is canaries with the CUSUM, and rates ρ ran from 0.005 to 0.12.
| Budget (reviews per shift) | Best plain: reviews, escape | Best canary: reviews, escape | Best canary with gate: reviews, escape |
|---|---|---|---|
| 15 | 11.5, 45.4% | 13.2, 49.1% | 13.2, 49.0% |
| 30 | 19.2, 43.9% | 25.2, 42.4% [42.1, 42.7] | 25.2, 42.5% |
| 60 | 19.2, 43.8% | 56.3, 37.8% [37.6, 38.2] | 56.3, 37.0% [36.8, 37.3] |
| 120 | 19.3, 43.7% | 101.2, 33.0% [32.7, 33.3] | 104.1, 32.8% [32.5, 33.0] |
| 250 | 19.3, 43.8% | 229.3, 28.4% [28.1, 28.6] | 229.2, 27.1% [26.8, 27.3] |
| 500 | 19.2, 43.9% | 275.8, 28.1% [27.8, 28.4] | 275.7, 26.1% [25.8, 26.3] |
Below about 25 reviews per shift, canaries do not pay. At a budget of 15 the plain policy beats the canary policy by 3.7 points, because each canary costs a review the tiered queue could have used on a real request. Above it, the picture changes. Plain tiering cannot spend more attention usefully: its best result is flat at 43.8% from 20 reviews upward. Canaries convert spare attention into catches. At a budget of 250 reviews, escape is 28.4% against 43.8%, and the gate takes it to 27.1%.
The price is friction. At 229 reviews the canary policy blocks 14.7 benign requests per shift against 1.0 for plain tiering, and at 276 reviews 23.4, since a reviewer who expects more bad requests also flags more good ones. This is the false-alarm side of Wolfe's criterion shift and of Schwark's false feedback [4][5]. Whether it is acceptable depends on what a blocked request costs.
The gate against a schedule
The gate has to be compared with something that uses the same break minutes, otherwise a gate that simply breaks more looks better. For each CUSUM threshold we ran a fixed evenly spaced schedule of the same total break time and the same canary rate. θ = 1.0, ρ = 0.02, h = 3, 6,000 shifts, escape with 95% intervals:
| Scenario | No breaks | Fixed schedule | Gate | Break minutes per shift |
|---|---|---|---|---|
| default (lapses 0.1 per hour, fatigue κ = 0.5) | 50.9% | 48.0% [47.4, 48.5] | 47.4% [46.8, 47.9] | 28.9 |
| lapse-heavy (lapses 0.4 per hour, mean 90 min, κ = 0.2) | 57.0% | 52.1% [51.4, 52.7] | 48.9% [48.3, 49.6] | 30.8 |
| fatigue only (no lapses, κ = 0.6) | 55.6% | 53.0% [52.5, 53.5] | 53.2% [52.5, 53.8] | 35.8 |
| no lapses, no fatigue | 31.1% | 30.9% [30.3, 31.4] | 31.4% [30.9, 31.9] | 10.4 |
The gate beats the schedule where lapses are episodic and long: 3.2 points in the lapse-heavy scenario, with the reviewer disengaged 20.2% of the time against 25.5% under the schedule and 33.3% with no breaks. In the default scenario the difference is 0.6 points, inside the intervals. Where fatigue accumulates smoothly, a schedule is as good as the gate, because there is nothing for a detector to detect that a clock does not know. With nothing to find, the gate does nothing.
That is the honest scope of the gate. It is a detector for episodes, not a fatigue cure.
Canaries the reviewer can recognise
If canaries have telltale formatting, source or timing, some are spotted. In the lapse-heavy scenario at θ = 1.0, ρ = 0.02 and h = 4 we varied the share ℓ of recognisable canaries.
| Recognisable share ℓ | Escape without gate | Escape with gate | Canary catch rate the gate sees | Break minutes per shift |
|---|---|---|---|---|
| 0 | 56.4% | 49.9% | 51.1% | 24.0 |
| 0.25 | 59.8% | 54.5% | 59.6% | 16.3 |
| 0.50 | 62.4% | 60.2% | 69.3% | 8.1 |
| 0.75 | 66.4% | 66.0% | 82.0% | 1.1 |
| 0.90 | 69.0% | 68.7% | 91.4% | 0.0 |
As ℓ rises, escape worsens by 12.6 points, the reviewer's catch rate on canaries climbs from 51% to 91%, and the gate stops firing. The dashboard improves while the system gets worse. Recognised canaries inflate the measured catch rate, as they are caught by artefact, and they stop updating the reviewer's sense of prevalence, which removes the first benefit as well. Of everything in this paper this is the result that most affects deployment: the canary must be indistinguishable from a real request in every field the reviewer sees, and the catch rate it reports needs an independent check.
When does it pay at all
We drew 120 parameter sets at random (base rate 0.1% to 2% log-uniform, guard mean 1.0 to 3.5, κ 0 to 0.8, α 0 to 0.9, fatigue gain 0.005 to 0.05, lapse hazard 0 to 0.5 per hour) and in each compared the best plain policy against the best canary policy over a grid of thresholds and canary rates, 500 shifts per cell, with J as above.
| Cost per escaped harmful request | Draws where canaries improve J by more than 5% | Median improvement in J |
|---|---|---|
| 100 review-equivalents | 13% | 0.3% |
| 1,000 | 82% | 14.7% (IQR 7.5 to 23.6) |
| 10,000 | 93% | 25.3% |
At C = 1,000, tercile bins of α show median improvements of 5.4%, 17.8% and 25.2% (α under 0.3, 0.3 to 0.6, above 0.6). The rank correlation of improvement with α is 0.79, with base rate 0.28, with guard strength 0.24 and with κ 0.10 (p = 0.28) and lapse hazard 0.02 (p = 0.81). The benefit tracks how strongly the reviewer's criterion follows prevalence, not how tired they get. Over the same draws the gate added a median 1.5% on top of canaries at C = 1,000 and 5.5% at C = 10,000.
Falsification and limits
The null. With α = 0, κ = 0 and no lapses, the reviewer's criterion cannot move and nothing degrades, so canaries should do nothing but cost attention. At 8,000 shifts per cell they did. With the gate on whenever canaries were present, at θ = 1.0 escape was 14.5% [14.2, 14.8] with no canaries and 14.4% [14.1, 14.8] and 14.6% [14.2, 14.9] with canary rates of 0.02 and 0.05, while reviews rose from 196 to 220 and 255. At θ = 2.0 the numbers were 26.4%, 26.5% and 26.5%. The benefit in the main results is therefore produced by the mechanisms we modelled, and not by canaries acting as extra reviews or by a bookkeeping artefact.
What does not survive. The gate does not beat a schedule in any scenario without episodes. Canaries do not help below the spare-attention threshold. The size of the benefit depends heavily on α.
The limits that matter most are about the reviewer. Every behavioural parameter apart from the two calibrated to Wolfe's visual-search data is our assumption. The prevalence sensitivity α = 0.7 comes from observers searching bags for objects and may be very different for people approving shell commands or record edits. Because the benefit tracks α more than anything else, a measurement of α for agent-approval work is the single most valuable follow-up. A human study could fit it from a few hundred approvals with injected positives at known prevalence. A reviewer who knows canaries exist might adopt a different strategy, which we do not model. We treat each canary as a harmful request with the same evidence distribution as a real one, which is the best case for the design.
A second limit is the guard. The guard matters in a direction we did not expect. In the random draws canaries improved J by more than 5% in 69% of draws with the weakest guards (mean 1.0 to 1.75) and in 95% with the strongest (2.5 to 3.5), at C = 1,000. We read this as canaries working best on a queue the guard has already concentrated, but we did not test that explanation directly, and the frontier we report was run with a strong guard. Third, the cost weights are ours: at 100 review-equivalents per escape the case for canaries mostly disappears. Fourth, the shift is eight hours of continuous streaming; real work has bursts, which would stress the fatigue dynamics more.
What this means in practice
Do not add canaries to a reviewer who is already within about 20 reviews per shift of a tight guard. They will cost more than they return. Spend the effort on the guard.
Add them when the reviewer's workload is high and cannot be lowered, and expect friction: more benign requests blocked. At about 230 reviews per shift our reviewer caught more harmful requests but blocked 14.7 benign ones per shift, instead of 1.0.
If a gate is wanted, use it to detect episodes: a sudden fall in the canary catch rate across a window. For a steadily tiring reviewer a break schedule does as well. In either case, count break minutes: a gate that fires often wins on catches simply by giving breaks.
Treat canary realism as a security property. Review them as an adversary would, looking for any field that distinguishes a canary from a real request, and keep the true catch rate from a separate source, for example a small share of canaries drawn from historical real incidents with ground truth established offline.
Finally, report the pair of numbers that this paper had to compute: harmful requests that pass per unit of attention spent, and benign requests blocked. A rejection rate near zero alone says little. It can mean careful work or reflexive approval, and the only way to tell the two apart is a known positive.
Reproducibility appendix
Python 3 with numpy and scipy. Seeds: fixed per configuration (base 1000, 5000, 7000, 9000, 20000 and 900000 plus a counter), listed in the scripts. Shifts per cell: 4,000 (policy grid), 3,000 (fine grid), 20,000 (confirmation), 6,000 (gate against schedule and leakage), 8,000 (null), 500 (parameter draws). Runtimes on a shared machine under load: policy grid 131 s, gate comparison 60 s, parameter draws about 5 s per draw. The simulator loops over the 1,200 slots and vectorises across shifts.
The core of one slot:
def review(mask, h, sigma, canary, ident):
d = P.d0 * (1.0 - P.kappa * f) * np.where(dis, P.lapse_factor, 1.0)
c = P.alpha * np.log((1.0 - q) / q) / d
x = h * d + rng.standard_normal(S)
flag = x > (d / 2.0 + c)
if canary:
flag = np.where(ident, rng.random(S) < P.leak_flag, flag)
f = np.where(mask, f + P.fat_a * (1.0 - f) * (1.0 + P.fat_gamma * sigma), f)
e = np.where(ident, q, 1.0) if canary else flag.astype(float)
upd = mask & ~ident if canary else mask
q = np.where(upd, np.clip((1.0 - P.eta) * q + P.eta * e, P.q_min, 0.5), q)
return flagThe gate, run after each canary:
inc = np.where(cflag, llr_catch, llr_miss)
z = np.where(cm, np.maximum(0.0, z + inc), z)
al = z > P.cusum_h
f = np.where(al, f * np.exp(-P.break_min / P.tau_break), f)
dis = np.where(al, False, dis)
z = np.where(al, 0.0, z)The default parameters are in the table in the Method section; the Wolfe check is miss = norm.cdf(t - d0), fa = 1 - norm.cdf(t) with t = d0/2 + alpha*log((1-q)/q)/d0.
References
- Anthropic. How we built Claude Code auto mode: a safer way to skip permissions. Engineering blog, 25 March 2026. https://www.anthropic.com/engineering/claude-code-auto-mode
- Yu H, Liu L, Jiang X, Jia Y, Wang S, Qian P, Chen Y. Habituation at the gate: rising approval and declining scrutiny in human review of AI agent code. arXiv:2606.22721, 2026. https://arxiv.org/abs/2606.22721
- Riz à Porta R, Sterchi Y, Schwaninger A. How realistic is threat image projection for X-ray baggage screening? Sensors 22(6):2220, 2022. https://pmc.ncbi.nlm.nih.gov/articles/PMC8952858/
- Wolfe JM, Horowitz TS, Van Wert MJ, Kenner NM, Place SS, Kibbi N. Low target prevalence is a stubborn source of errors in visual search tasks. Journal of Experimental Psychology: General 136(4):623-638, 2007. https://pmc.ncbi.nlm.nih.gov/articles/PMC2662480/
- Schwark J, Sandry J, MacDonald J, Dolgov I. False feedback increases detection of low-prevalence targets in visual search. Attention, Perception, & Psychophysics 74:1583-1589, 2012. https://link.springer.com/article/10.3758/s13414-012-0354-4
- Turan E. Oversight has a capacity: calibrating agent guards to a subjective, fatiguing human. arXiv:2606.08919, 2026. https://arxiv.org/abs/2606.08919
- Page ES. Continuous inspection schemes. Biometrika 41(1-2):100-115, 1954.
Cite this paper
@misc{forticia2026canaries,
title = {{Canaries in the approval queue: what injected known-bad requests buy a fatigued human reviewer}},
author = {{Forticia Research Institute}},
year = {2026},
month = oct,
publisher = {Forticia Research Institute},
howpublished = {\url{https://www.forticia.uk/papers/canary-requests-approval-queue-human-oversight}},
note = {Paper, published online}
}Generated from this page’s metadata. Forticia does not assign DOIs to these papers.