<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en-GB">
  <title>Forticia Research Institute: papers and research notes</title>
  <subtitle>New papers and research notes from Forticia Research Institute (forticia.uk).</subtitle>
  <link rel="self" type="application/atom+xml" href="https://www.forticia.uk/papers/feed.xml"/>
  <link rel="alternate" type="text/html" href="https://www.forticia.uk/papers"/>
  <id>https://www.forticia.uk/papers</id>
  <updated>2026-10-05T23:30:00.000Z</updated>
  <author><name>Forticia Research Institute</name><uri>https://www.forticia.uk/</uri></author>
  <icon>https://www.forticia.uk/apple-touch-icon.png</icon>
  <entry>
    <title>How many independent proteins are in a benchmark? Error correlation across sequence similarity and the cost to error bars</title>
    <link rel="alternate" type="text/html" href="https://www.forticia.uk/papers/effective-number-of-independent-proteins-benchmark-error-bars"/>
    <id>https://www.forticia.uk/papers/effective-number-of-independent-proteins-benchmark-error-bars</id>
    <published>2026-10-05T23:30:00.000Z</published>
    <updated>2026-10-05T23:30:00.000Z</updated>
    <author><name>Forticia Research Institute</name></author>
    <category term="Benchmarking"/>
    <category term="Protein sequences"/>
    <category term="Effective sample size"/>
    <category term="Homology"/>
    <category term="Uncertainty"/>
    <summary type="text">Model errors on related proteins are correlated even at the weakest significant hits. We measure it and what it does to benchmark confidence intervals.</summary>
  </entry>
  <entry>
    <title>Where identical code gives different answers: a reproducibility budget for a protein-sequence pipeline</title>
    <link rel="alternate" type="text/html" href="https://www.forticia.uk/papers/reproducibility-budget-protein-sequence-pipeline"/>
    <id>https://www.forticia.uk/papers/reproducibility-budget-protein-sequence-pipeline</id>
    <published>2026-10-05T23:22:00.000Z</published>
    <updated>2026-10-05T23:22:00.000Z</updated>
    <author><name>Forticia Research Institute</name></author>
    <category term="Reproducibility"/>
    <category term="Computational biology"/>
    <category term="Protein sequences"/>
    <category term="Numerical stability"/>
    <category term="scikit-learn"/>
    <summary type="text">One protein pipeline, 192 runs, one factor changed at a time: seeds dominate, thread and BLAS noise is invisible, and one library release flips a verdict.</summary>
  </entry>
  <entry>
    <title>Canaries in the approval queue: what injected known-bad requests buy a fatigued human reviewer</title>
    <link rel="alternate" type="text/html" href="https://www.forticia.uk/papers/canary-requests-approval-queue-human-oversight"/>
    <id>https://www.forticia.uk/papers/canary-requests-approval-queue-human-oversight</id>
    <published>2026-10-05T22:22:00.000Z</published>
    <updated>2026-10-05T22:22:00.000Z</updated>
    <author><name>Forticia Research Institute</name></author>
    <category term="Human oversight"/>
    <category term="Approval fatigue"/>
    <category term="Agents"/>
    <category term="Canaries"/>
    <category term="Signal detection"/>
    <category term="Simulation"/>
    <summary type="text">In simulation, injected known-bad requests help a fatigued reviewer only above a spare-attention threshold, and recognisable canaries hide the failure.</summary>
  </entry>
  <entry>
    <title>What an agent audit log can answer: measuring incident answerability across logging schemas</title>
    <link rel="alternate" type="text/html" href="https://www.forticia.uk/papers/audit-trail-answerability-agent-approvals"/>
    <id>https://www.forticia.uk/papers/audit-trail-answerability-agent-approvals</id>
    <published>2026-10-05T21:23:00.000Z</published>
    <updated>2026-10-05T21:23:00.000Z</updated>
    <author><name>Forticia Research Institute</name></author>
    <category term="Audit trail"/>
    <category term="Agents"/>
    <category term="Approvals"/>
    <category term="Tamper-evident logging"/>
    <category term="Forensics"/>
    <category term="Governance"/>
    <summary type="text">Simulated approval incidents and a hash-chained log show which fields answer which questions, what defaults cost, and how anchoring sets tamper detection.</summary>
  </entry>
  <entry>
    <title>Gate stacks are not products: measuring the rubber-stamp risk of a research pipeline</title>
    <link rel="alternate" type="text/html" href="https://www.forticia.uk/papers/gate-stack-rubber-stamp-risk"/>
    <id>https://www.forticia.uk/papers/gate-stack-rubber-stamp-risk</id>
    <published>2026-10-05T21:03:00.000Z</published>
    <updated>2026-10-05T21:03:00.000Z</updated>
    <author><name>Cayden Richards</name></author>
    <category term="Research methodology"/>
    <category term="Multiple testing"/>
    <category term="Backtest validation"/>
    <category term="Placebo controls"/>
    <category term="Simulation"/>
    <summary type="text">In a simulated research pipeline, correlated gates protect up to 18 times less than their pass rates imply, and a one-draw placebo is a coin flip.</summary>
  </entry>
  <entry>
    <title>When a failure record lies: the economics of shared negative results in a research swarm</title>
    <link rel="alternate" type="text/html" href="https://www.forticia.uk/papers/graveyard-economics-shared-negative-results"/>
    <id>https://www.forticia.uk/papers/graveyard-economics-shared-negative-results</id>
    <published>2026-10-05T21:03:00.000Z</published>
    <updated>2026-10-05T21:03:00.000Z</updated>
    <author><name>Cayden Richards</name></author>
    <category term="Research methodology"/>
    <category term="Negative results"/>
    <category term="Multi-agent search"/>
    <category term="False negatives"/>
    <category term="Simulation"/>
    <summary type="text">In a simulated swarm, sharing falsified ideas saves a third of compute, but neighbourhood rules close off real effects unless records carry their test power.</summary>
  </entry>
  <entry>
    <title>Never Tune a Corpse</title>
    <link rel="alternate" type="text/html" href="https://www.forticia.uk/papers/exhumation-policies-rejected-ideas"/>
    <id>https://www.forticia.uk/papers/exhumation-policies-rejected-ideas</id>
    <published>2026-10-05T20:16:00.000Z</published>
    <updated>2026-10-05T20:16:00.000Z</updated>
    <author><name>Forticia Research Institute</name></author>
    <category term="epistemics"/>
    <category term="research governance"/>
    <category term="multiple testing"/>
    <category term="replication"/>
    <category term="simulation"/>
    <category term="falsification"/>
    <summary type="text">When may a rejected idea be reopened? Simulated burial and exhumation rules, their false discovery cost, and how the graveyard can audit the court.</summary>
  </entry>
  <entry>
    <title>When Approval Stops Carrying Information</title>
    <link rel="alternate" type="text/html" href="https://www.forticia.uk/papers/rubber-stamp-gate-informativeness"/>
    <id>https://www.forticia.uk/papers/rubber-stamp-gate-informativeness</id>
    <published>2026-10-05T20:14:00.000Z</published>
    <updated>2026-10-05T20:14:00.000Z</updated>
    <author><name>Forticia Research Institute</name></author>
    <category term="governance"/>
    <category term="oversight"/>
    <category term="rubber-stamping"/>
    <category term="information theory"/>
    <category term="automation"/>
    <category term="simulation"/>
    <summary type="text">A conditional-information measure of rubber-stamping, the defect rate where review rationally collapses, and which levers restore it. A simulation study.</summary>
  </entry>
</feed>
