Skip to content
Forticia
Research
Quantitative FinanceEquities, FX, futures. Factors and backtests, every run replayable.Computational BiologySequence, folding, simulation. Versioned, reproducible labs.Cultural IntelligencePhilosophy, governance, ethics. How institutions decide and answer for it.AI InstrumentationPrivate models. Multi-agent orchestration. Guardrails on write.View all research
Quantitative Finance
  • Equities, FX, futures
  • Factors and backtests
  • Every run logged and replayable
InfrastructurePapersPolarisLink™About
Sign inRequest access
Request access

Research

Quantitative FinanceEquities, FX, futures. Factors and backtests, every run replayable.Computational BiologySequence, folding, simulation. Versioned, reproducible labs.Cultural IntelligencePhilosophy, governance, ethics. How institutions decide and answer for it.AI InstrumentationPrivate models. Multi-agent orchestration. Guardrails on write.

Platform

InfrastructurePapersPolarisLink™About
Request accessSign in
Forticia

A private institute for computational research. It publishes original research and runs a governed environment where every run is logged and replayable.

Sign inSystem status

Research

Quantitative FinanceComputational BiologyCultural IntelligenceAI Instrumentation

Platform

InfrastructurePapersPolarisLink™StatusSign in

Institute

AboutRequest accessContactGitHubPolarisLink repository

Forticia publishes research and simulations. Nothing on this site is investment advice.

© 2026 ForticiaPrivacyTerms
Papers/AI Instrumentation
AIPaper

What an agent audit log can answer: measuring incident answerability across logging schemas

Author
Forticia Research Institute
Published
5 October 2026
Last updated
5 October 2026
Reading time
16 min
Cite this paper

On this page 0%

  1. Abstract
  2. Motivation
  3. Related work and what is new
  4. Method
  5. The workflow and its ground truth
  6. Nine fields and 512 schemas
  7. Nine questions and two readers
  8. The log implementation
  9. Results
  10. The ladder
  11. Which fields buy what
  12. Three proxies and when they fail
  13. Tampering
  14. Cost
  15. Falsification and limits
  16. What this means in practice
  17. Reproducibility appendix
  18. References
  19. Cite this paper

Abstract

We ask a narrow question that audit-log design guides leave unmeasured: given a log schema, what fraction of the questions an incident review asks can be answered from the log alone? We built a hash-chained append-only log and a simulator of an approval-gated agent workflow (draft, review, approve, execute), generated 2,000 incident days with ground truth, and scored two readers on nine incident questions across 512 logging schemas. A reader who fills gaps with defaults is wrong on 35.6% of answers (95% CI 34.8 to 36.3) when the log holds only executed calls, and still wrong on 19.2% when it also holds approvals and approver identity. Request-to-decision time, the usual stand-in for review time, falls to near chance as queues grow from seconds to minutes. A hash chain without external anchors detected 0% of recomputed rewrites; anchoring every 10 entries detected 84% to 92%.

Motivation

A write to a shared record is executed, someone asks the next morning how it happened, and the log has one line: the tool, the arguments and a timestamp. The reviewer can see that the call ran. Whether a human approved it, who, what they were shown, how long they looked, whether the payload that ran was the payload they approved, and which upstream agent caused the draft are all unrecorded. Each missing fact has a default answer, and the defaults are the dangerous part. The system was designed so that writes need approval, so a reader assumes this one had it.

Forticia describes its governed AI as drafts by default, lead approval to write, and a full audit trail. The third clause is the one that is hardest to check from outside and easiest to under-build, because a log that exists feels like a log that works. "Full" has no unit. This paper proposes one: the fraction of incident questions the log answers correctly, the fraction it answers wrongly because a reader filled a gap, and the bytes it costs.

We do not have production incident data and do not claim results about any deployed system. Everything below is a simulation of an approval workflow, plus real measurements of a real log implementation. The incident mix is ours and the conclusions are conditional on it. Where a conclusion depends on a parameter we sweep the parameter.

Related work and what is new

Tamper-evident logging has a long history. Schneier and Kelsey described a hash chain with forward-secure MACs for logs that must survive a compromised host [1], and Crosby and Wallach gave tree-based structures with logarithmic proofs for untrusted loggers audited by clients [2]. We use the plain hash chain and ask a question those papers do not: what the entries should contain.

On content, Olegård, Axelsson and Li study when logging is sufficient for forensic analysis of attacks and propose causal tracking across applications, with a prototype and overhead measurements [3]. Solozobov defines a reconstructability metric for agent-safety evaluations: whether captured evidence can reconstruct the decision a claim depends on, scored on public traces [4]. None of these, as far as we found, scores a family of schemas against a typed set of approval-workflow incidents, separates answers a log entails from answers a reader assumes, or prices each field.

What is new here is four things. An answerability audit with two readers, one that abstains and one that presumes. A field-level value and cost table obtained by evaluating all 512 subsets of nine logged fields. A measurement showing that two cheap proxies for facts a log lacks fail in predictable regimes. A measured detection curve for a rewriting attacker as a function of the anchor interval of a hash chain.

Method

The workflow and its ground truth

An episode is one working day of an approval-gated workspace. Between 40 and 79 background requests arrive from five agents. Each request is created, shown to one of three approvers, decided (approve 85%, deny 10%, expire 5%) and, if approved, executed. One request in each episode is the incident write, to a fixed record and field. Its properties are drawn independently:

Property Probability Meaning
Bypass 0.30 executed with no request or approval at all
Retroactive approval 0.10 of the rest approval recorded after execution
Truncated view 0.50 the approver's view cut off the changed value
Rubber stamp 0.50 decision 0.3 to 1.8 seconds after display, otherwise 6 to 60
Payload mismatch 0.40 executed payload differs from the approved one; half leave a logged revision after approval, half are silent
Benign normalisation 0.35 of the rest executor changes whitespace and case only
Benign revision 0.30 of the rest agent revises the draft before approval
Delegation depth 0, 1, 2 with 0.40, 0.35, 0.25 the request was caused by up to two upstream requests, half of which were never executed
High load 0.40 the approver made 35 to 54 other decisions in the previous 30 minutes, otherwise 2 to 9
Earlier write to the field 0.60 and in 15% of those the value was changed outside the log
Tail position 0.30 the incident falls in the last minute of the log

Ground truth for every question is read from the generating variables, and the high-load label is recounted from the generated trace. The generator emits a complete event list; a schema decides which events and attributes survive into the log.

Nine fields and 512 schemas

A schema is a subset of nine switches over a base log that always holds executed calls with agent, record, field and arguments.

Field What it adds
lifecycle request, revision and decision events for every request, including denied and expired ones
approver the approver's identity on each decision
ts timestamps on every entry
shown an event marking the moment the request was displayed to the approver
view the canonical view: which fields were visible and whether it was truncated
hashes hashes of the canonicalised payload at request, approval and execution
parent the id of the request that caused this one
policy the policy version in force at request time
before the value of the field before the write

Four named bundles are the ladder reviewers will recognise. B0 is the base log with timestamps. B1 adds lifecycle and approver. B2 adds shown, view, parent and policy. B3 adds hashes and before. Hash chaining is a tenth switch, handled separately because it protects entries rather than describing events.

Nine questions and two readers

The questions, with the incident class that makes each one hard: was the write approved before it executed; by whom; could the approver see the changed value; was the decision made in under two seconds; did the executed payload equal the approved payload; which agent started the causal chain; had the approver made more than 25 decisions in the previous 30 minutes; what was the value before the write; which policy version applied. Questions about an approval are asked only when the write was approved before execution.

The strict reader answers only when the logged evidence entails the answer, and otherwise returns undetermined. The presumptive reader fills each gap with the default a hurried reviewer would use: executed means approved, the approver is the most common one, the view was complete, the decision was deliberate, differing arguments mean a swap, the executing agent is the origin, load was normal, and the policy is the current one. When a proxy is available it uses the proxy: decision time minus request time for review time, a window of the last 150 entries for load. Both readers see only the projected log.

Both are code we wrote, so they encode our idea of reasonable inference. A different reader could extract more from a thin log. We return to this under limits.

The log implementation

The prototype is a hash chain. Each entry stores its sequence number, the previous hash and the SHA-256 of the previous hash, the sequence number and the canonical JSON of the body. Verification recomputes the chain. Anchors are (sequence, hash) pairs held outside the log at a fixed interval. The core functions are in the appendix.

Results

The ladder

Over 2,000 episodes (13,849 scored questions per reader per schema), with bootstrap intervals over episodes:

Figure 1. What each log schema lets an investigator answer: share of scored questions the strict reader answers correctly, wrongly or not at all, with the presumptive reader's wrong-answer rate (diamond, 95 percent interval), over 2,000 simulated incident days. B3 is an upper bound by construction.
Schema Bytes per day Strict correct Strict wrong Strict undetermined Presumptive wrong
B0 args only 17,688 7.5% [7.2, 7.9] 1.4% [1.2, 1.6] 91.1% 35.6% [34.8, 36.3]
B1 + lifecycle, approver 57,391 46.1% [45.7, 46.5] 1.4% 52.6% 19.2% [18.6, 19.7]
B2 + view, causal parent, policy 80,256 89.3% [88.9, 89.8] 1.4% 9.3% 3.2% [2.9, 3.5]
B3 + hashes, before value 89,180 100.0% 0.0% 0.0% 0.0%

The B3 row is an upper bound by construction: the readers were written to use every field. The informative numbers are the ones above it. A log of executed calls with arguments supports a correct strict answer to 7.5% of questions, all of them the before-value, which can be read off an earlier write. The same log supports a confident wrong answer to 35.6% of questions if the reader presumes. The two worst questions are approval status (715 of 2,000 wrong, because 30% of incident writes bypassed approval and 7% were approved after the fact) and the approver's identity (757 of 1,285 wrong).

Bytes here include 150 bytes of chain overhead per entry. B3 is five times the size of B0 and still under 90 kilobytes for a day with about 300 entries.

The strict wrong rate of 1.4% is constant across the ladder until B3. It comes from one place: the before-value. When an earlier write to the field exists in the log the strict reader reports it as the prior value, and in 15% of those cases the field had been changed outside the log. A reader who trusts the log to be a complete record of writes is wrong exactly when it is not.

Which fields buy what

With nine fields there are 512 schemas, and for each field the Shapley value over those schemas gives its average marginal contribution to the strict-correct fraction. Cost is the average added bytes per day, from the same enumeration.

Figure 2. Which fields buy what: each field's Shapley contribution to the strict-correct fraction, averaged over the 512 possible schemas, against its average added bytes per day. Lifecycle is the outlier in both.
Field Share of strict-correct gain Added bytes per day Change in presumptive-wrong
lifecycle 39.5 points 40,465 minus 15.0
parent 12.6 points 1,526 minus 4.3
approver 7.6 points 642 minus 4.3
shown 7.6 points 16,629 minus 4.1
before 6.9 points 1,233 minus 1.5
ts 6.1 points 2,765 minus 1.6
policy 5.3 points 562 minus 0.9
view 4.6 points 1,752 minus 2.5
hashes 2.1 points 4,781 minus 1.0

Lifecycle logging carries the most value and nearly half the bytes, because it triples the number of entries: B0 has 69 entries per day and B1 has 232. The remaining fields are cheap. Approver, parent, policy and before-value together add under 4 kilobytes and are worth 32 points between them. Two fields are expensive for what they return. The separate shown event adds 16.6 kilobytes, and hashes add 4.8 kilobytes for 2.1 points, though the hash is the only field that settles the payload question exactly (below).

The cheapest schema that answers at least 95% of questions strictly is eight fields, everything except hashes, at 79,136 bytes and 95.8% correct. Reaching 100% costs the hashes. No schema under 54,000 bytes exceeded 47% correct.

One design choice removes most of the cost of the display event. Instead of logging "shown" as its own entry, we put the display timestamp and the canonical view on the decision entry. On the same 2,000 episodes this variant answered every question correctly under both readers, as the full separate-event schema does, and used 236 entries and 74,942 bytes per day against 319 entries and 89,180 bytes, a reduction of 16%. The facts are the same. What changes is that no extra entry, with its own chain overhead, is written per request.

Three proxies and when they fail

Two facts that a thin log lacks have cheap substitutes that look reasonable. We swept the environment parameter that controls each.

Review time. Without a display event, the reader uses the interval from request creation to decision. That interval contains queue time, the wait before the approver opened the item. We varied the upper bound of queue time and measured how well the proxy separates rubber-stamped approvals (under 2 seconds of true review) from deliberate ones, by the area under the ROC curve over about 380 cases of each:

Queue time up to AUC of the proxy
8 s 1.00
20 s 0.98
60 s 0.87
150 s 0.68
300 s 0.59
900 s 0.52

At our default of 300 seconds the proxy is close to a coin flip, and the default presumptive reader, which assumes any approval is deliberate, misses every rubber stamp. If approvals in a deployment are answered within seconds the proxy is fine. That is a fact about the queue, not about the log, and it has to be measured before the proxy is trusted.

Payload equality. Without hashes, a reader compares the arguments at request time with the arguments at execution. Benign changes such as case folding and whitespace trimming make equal payloads look different. The precision of the rule "differs means swapped" falls as the share of benign normalisations rises, with recall at 1.0 throughout:

Share of non-swaps that are benign normalisations Precision of the heuristic
0 1.00
0.10 0.84
0.35 0.64
0.70 0.47
1.00 0.38

A hash over the canonical form removes the problem because the canonical forms are equal. This is the one place where hashes are worth their bytes.

Approver load. Without timestamps a reader approximates the 30-minute window by the last 150 entries. With the default background volume the approximation had 91% precision and full recall; with 150 to 250 background requests, 86% precision. It degrades slowly because load and entry count are correlated.

Tampering

Real chains were built over each of 1,500 episode logs under B3 and attacked. The attacker either edits entries and leaves the stored hashes alone, which any verifier catches, or rewrites the chain from the first modified entry, which a chain alone cannot catch. We measured detection with no anchors and with anchors published externally every 200, 50, 10 and 1 entries. Attacks: delete the incident's display and decision entries, edit the executed arguments, edit the view and display time, and truncate the log from the incident onward.

Figure 3. Anchors turn a chain into evidence: share of rewriting attacks detected, by how often the chain head is published outside the log. Attacks whose lines are identical in the data are drawn once.
Anchor interval Delete approval Edit arguments Edit view Truncate
none (chain only) 0.0% 0.0% 0.0% 0.0%
every 200 42.8% 39.1% 42.8% 42.5%
every 50 72.4% 69.1% 72.4% 74.7%
every 10 89.6% 84.1% 89.6% 92.3%
every 1 100% 100% 100% 100%

Rows are the rewriting attacker; the naive attacker who edits without recomputing was detected 99% to 100% of the time for edits and deletions even with no anchors, and not at all for truncation, since a shorter valid chain is still valid. The curve has an interpretation. A rewrite is detected only if an anchor was published at or after the first altered entry. An incident inside the last interval is unprotected, and 30% of our incidents fall in the last minute of the log by construction, so detection is 84% to 92% rather than 100% at an interval of 10. The undetected remainder is consistent with incidents whose first altered entry sits in the newest interval. Detection is a statement about how stale the newest anchor is, not about the hash function.

What tampering does to answers when there is no chain is also measurable. After the same attacks on logs without a chain (3,393 scored answers on approved incidents), the strict reader gave a wrong answer to 11.1% of questions when the approval was deleted, 6.8% when arguments were edited, 11.2% when the view was edited and 33.1% when the log was truncated. The truncation figure comes from a rule in our reader: it looks for the last write to the record, and after truncation the last write is an earlier, unrelated one, which it describes confidently. An investigator pointed at the wrong entry has the same failure.

Cost

Measured on one laptop-class machine, five runs over 93,999 entries of B3 bodies:

Operation Entries per second (mean, sd)
SHA-256 over bodies only 177,946 (23,796)
Append, buffered 78,442 (4,549)
Append, flush and fsync every 100 71,047 (4,657)
Append, fsync every entry (2,000 entries) 15,724 (4,499)
Verify in memory 169,195 (28,600)
Load from disk and verify 98,438 (13,606)

The file held 299 bytes per entry, roughly half of it chain overhead. A day under B3 is roughly 90 kilobytes. Logging cost is not the barrier at these volumes; a team that omits approver identity to save space saves 642 bytes.

Falsification and limits

We tried to break the headline in four ways.

First, sanity checks where the readers must agree. With all nine fields both readers score 100% on every question, and in the sweep with no benign normalisations the payload heuristic reaches precision 1.00. The wrong rates in the ladder are therefore produced by the gaps and the incident mix, not by a defect in the readers.

Second, the incident mix. The presumptive wrong rate for B0 is 35.6% because 36% of incidents involve no valid prior approval and 40% of approved incidents involve a mismatch. We chose those rates. We did not choose them to be realistic, because we have no data on how often real incidents look like this; we chose them so that every question had both outcomes in meaningful numbers. A deployment in which bypass is rare will see a much lower presumptive error on the approval question and the same error on the others. The ordering of the four bundles survives any reweighting of questions, because each bundle answers at least as many incidents correctly as the one before it on every question separately, but the headline percentages do not. They should be read as a property of this generator.

Third, the readers. Both are rule lists we wrote after thinking about what an analyst does. They are not an upper bound on what an analyst, or a language model with access to the raw environment, could infer from a thin log. Some questions the strict reader marks undetermined are partially answerable from side channels (a ticket system, a chat message). We have not measured those.

Fourth, completeness. The strict reader trusts that every write is in the log. The 1.4% strict error is exactly the cost of that trust, and it applies equally to every schema short of the full one.

What does not survive scrutiny: the notion that more logging is monotonically better. The separate display event costs 16.6 kilobytes for 7.6 points, while attaching the same facts to the decision entry gave identical answers for 14.2 kilobytes less in the variant above. The ladder hides this because it adds fields in a customary order.

What this means in practice

Three design rules follow from numbers in this paper rather than from principle.

Log the decision, not the call. A log of executed calls answers 7.5% of the questions an approval incident raises and invites a 35.6% confident error rate. Writing request, revision and decision events for every request, including denied and expired ones, is the single largest improvement and costs about half the bytes of the whole schema.

Record what the human saw and when, as attributes of the decision entry rather than as a separate event (16% smaller in our variant, same answers), and hash the canonical payload at request, approval and execution. These two are the only way to answer review-time, view and payload-equality questions without a proxy, and the proxies fail in regimes (long queues, benign normalisation) that are common.

Anchor the chain outside the host that writes it, at an interval shorter than the delay you can tolerate before a rewrite is visible. A chain alone gave 0% detection against a rewriting attacker in our tests. Anchoring every 10 entries gave 84% to 92% and every entry gave 100%. The residual is the newest interval.

Teams can run this audit on their own schema. Enumerate the questions your incident reviews ask, write the strict reader, replay a set of incidents from your own history, and report three numbers: answered correctly, answered wrongly under defaults, and bytes. The first is the unit that "full audit trail" has been missing.

Reproducibility appendix

All code is Python 3 with numpy. Seeds and counts: episode seeds 100000 to 101999 (ladder, 2,000 episodes), the first 200 of those for the 512-subset enumeration, 300000 to 301499 for tamper experiments (the first 600 for the no-chain answer check), 500000 to 500299 for the cost benchmark, and three ranges of 1,200 episodes starting at 700000, 710000 and 720000 for the three proxy sweeps. Bootstrap: 500 resamples over episodes. Runtimes on a shared development machine under load: ladder and subset enumeration together 87 seconds, tamper experiments 138 seconds, proxy sweeps 233 seconds, cost benchmark a few minutes.

The log:

python
def entry_hash(prev, seq, body_bytes):
    h = hashlib.sha256()
    h.update(prev.encode("ascii"))
    h.update(seq.to_bytes(8, "big"))
    h.update(body_bytes)
    return h.hexdigest()

class ChainLog:
    def append(self, body):
        seq = len(self.entries)
        bb = canonical(body)
        h = entry_hash(self.head, seq, bb)
        self.entries.append({"seq": seq, "prev": self.head, "body": body, "hash": h})
        self.head = h
        if self.anchor_every and (seq + 1) % self.anchor_every == 0:
            self.anchors.append((seq, h))

def verify(entries, anchors=()):
    prev = GENESIS
    for i, rec in enumerate(entries):
        if rec["seq"] != i or rec["prev"] != prev:
            return False, i
        if entry_hash(prev, i, canonical(rec["body"])) != rec["hash"]:
            return False, i
        prev = rec["hash"]
    for seq, h in anchors:
        if seq >= len(entries) or entries[seq]["hash"] != h:
            return False, seq
    return True, len(entries)

Two of the nine strict rules, as examples of the style:

python
if q == "approved_before_exec":
    if "lifecycle" not in f:
        return True if presumptive else UNDET
    return any(e["decision"] == "approve" and i < ii
               for i, e in _by_req(log, "decide", call))
if q == "mismatch":
    if "hashes" in f and "lifecycle" in f:
        d = [e for i, e in _by_req(log, "decide", call) if e["decision"] == "approve"]
        if d and "approved_hash" in d[0]:
            return d[0]["approved_hash"] != ex.get("exec_hash")

Detection under a rewriting attacker: rebuild a chain over the altered bodies, then call verify(rebuilt, anchors) with the anchors from the original log. The attack is detected exactly when an anchor sequence number is at or beyond the first altered entry, or beyond the new length.

  • Audit trail
  • Agents
  • Approvals
  • Tamper-evident logging
  • Forensics
  • Governance

References

  1. Schneier B, Kelsey J. Secure audit logs to support computer forensics. ACM Transactions on Information and System Security 2(2):159-176, 1999. https://dl.acm.org/doi/10.1145/317087.317089
  2. Crosby SA, Wallach DS. Efficient data structures for tamper-evident logging. USENIX Security Symposium, 2009. https://static.usenix.org/event/sec09/tech/full_papers/crosby.pdf
  3. Olegård J, Axelsson S, Li Y. When is logging sufficient? Tracking event causality for improved forensic analysis and correlation. Forensic Science International: Digital Investigation 52, 301877, 2025. https://dfrws.org/wp-content/uploads/2025/03/When-is-logging-sufficient-Tracking-event-c_2025_Forensic-Science-Interna.pdf
  4. Solozobov O. Agent-safety evaluations as load-bearing evidence: a vendor-neutral, cross-harness reconstructability metric. arXiv:2607.12469, 2026. https://arxiv.org/abs/2607.12469

Cite this paper

@misc{forticia2026agent,
  title        = {{What an agent audit log can answer: measuring incident answerability across logging schemas}},
  author       = {{Forticia Research Institute}},
  year         = {2026},
  month        = oct,
  publisher    = {Forticia Research Institute},
  howpublished = {\url{https://www.forticia.uk/papers/audit-trail-answerability-agent-approvals}},
  note         = {Paper, published online}
}

Generated from this page’s metadata. Forticia does not assign DOIs to these papers.

NewerCanaries in the approval queue: what injected known-bad requests buy a fatigued human reviewerOlderGate stacks are not products: measuring the rubber-stamp risk of a research pipeline
All papers
0%25%50%75%100%Share of scored questionsB0 args only17,688 bytes per day91.1%35.6%B1 + lifecycle, approver57,391 bytes per day46.1%52.6%19.2%B2 + view, parent, policy80,256 bytes per day89.3%3.2%B3 + hashes, before value89,180 bytes per day100.0%0.0%
  • Strict correct
  • Strict wrong
  • Strict undetermined
  • Presumptive wrong, 95% interval
Stacked bars for four log schemas, B0 to B3, of the share of scored questions the strict reader answers correctly, answers wrongly and leaves undetermined, with the presumptive reader's wrong-answer rate as a marker with its 95 percent interval. B0 args only, 17,688 bytes per day: 7.5% correct, 1.4% wrong, 91.1% undetermined, presumptive wrong 35.6%; B1 + lifecycle, approver, 57,391 bytes per day: 46.1% correct, 1.4% wrong, 52.6% undetermined, presumptive wrong 19.2%; B2 + view, parent, policy, 80,256 bytes per day: 89.3% correct, 1.4% wrong, 9.3% undetermined, presumptive wrong 3.2%; B3 + hashes, before value, 89,180 bytes per day: 100.0% correct, 0.0% wrong, 0.0% undetermined, presumptive wrong 0.0%.
SchemaStrict correctStrict wrongStrict undeterminedPresumptive wrong, 95% interval
B0 args only7.5%1.4%91.1%35.6%
B1 + lifecycle, approver46.1%1.4%52.6%19.2%
B2 + view, parent, policy89.3%1.4%9.3%3.2%
B3 + hashes, before value100.0%0.0%0.0%0.0%
010203040Gain in strict-correct fraction, points1,00010,000100,000Added bytes per daylifecycleparentapprovershownbeforetspolicyviewhashes
Point
Hover or focus a point to read its values
Scatter plot of each of nine log fields, with added bytes per day on a log axis against its Shapley contribution to the strict-correct fraction in points. lifecycle: 39.5 points for 40,465 bytes; parent: 12.6 points for 1,526 bytes; approver: 7.6 points for 642 bytes; shown: 7.6 points for 16,629 bytes; before: 6.9 points for 1,233 bytes; ts: 6.1 points for 2,765 bytes; policy: 5.3 points for 562 bytes; view: 4.6 points for 1,752 bytes; hashes: 2.1 points for 4,781 bytes.
ItemAdded bytes per dayGain in strict-correct fraction, pointsDetail
lifecycle40,4654040,465 bytes per day; presumptive-wrong change −15.0 points
parent1,526131,526 bytes per day; presumptive-wrong change −4.3 points
approver6428642 bytes per day; presumptive-wrong change −4.3 points
shown16,629816,629 bytes per day; presumptive-wrong change −4.1 points
before1,23371,233 bytes per day; presumptive-wrong change −1.5 points
ts2,76562,765 bytes per day; presumptive-wrong change −1.6 points
policy5625562 bytes per day; presumptive-wrong change −0.9 points
view1,75251,752 bytes per day; presumptive-wrong change −2.5 points
hashes4,78124,781 bytes per day; presumptive-wrong change −1.0 points
0%25%50%75%100%Attacks detectedNoneEvery 200Every 50Every 10Every 1External anchor interval, in entriesA chain alone detects none
  • Delete approval and edit view
  • Edit arguments
  • Truncate
External anchor interval, in entries
Move across the chart, or focus it and use the arrow keys
Line chart of the percent of rewriting attacks detected against the external anchor interval in entries: none, every 200, every 50, every 10, every 1. Delete approval and edit view: 0.0%, 42.8%, 72.4%, 89.6%, 100.0%; Edit arguments: 0.0%, 39.1%, 69.1%, 84.1%, 100.0%; Truncate: 0.0%, 42.5%, 74.7%, 92.3%, 100.0%. A chain with no external anchor detects none of them.
External anchor interval, in entriesDelete approval and edit viewEdit argumentsTruncate
10%0%0%
243%39%43%
372%69%75%
490%84%92%
5100%100%100%