Products
Solutions
Company
Enterprise
Sign inCreate your network
AI safety · cybersecurity · risk governance · Bayesian networks · Theorem 3

Risk thresholds need a measuring mechanism, not a cutoff

A Berkeley white paper argues AI cyber-risk thresholds fail because they are capability cutoffs. The fix is a Bayesian network that measures harm continuously — exactly what Theorem 3 demands.

Risk thresholds need a measuring mechanism, not a cutoff

A UC Berkeley Center for Long-Term Cybersecurity white paper argues that AI cyber-risk thresholds drawn from frontier-lab safety frameworks (OpenAI, Anthropic, DeepMind, Meta) fail because they are capability-based, vague, and detached from real harm. The authors propose Bayesian networks that turn "does this model cross a line?" into "how likely is it to cause harm under real conditions?" — and that move, from a static cutoff to a continuously measured mechanism, is exactly the shift a defensible threshold requires.

The threshold is not a line you cross; it is a measurement you maintain

Frontier AI safety frameworks converge on a handful of "threshold elements" — automation of multi-stage attacks, zero-day discovery, enabling low-skilled attackers — and then attach capability benchmarks to each. The Berkeley paper's central critique is that these cutoffs are deterministic in a system that is fundamentally probabilistic. A model that "can" discover a zero-day in a benchmark does not, by that fact, cause harm; the same capability is benign or catastrophic depending on access, defender posture, and attacker economics. A threshold stated as a capability line ignores every mediating variable.

This is the same failure mode we name with Theorem 3 across the 21 papers: a property is guaranteed exactly when its mechanism is implemented and measuring. [ORIGINAL DATA] Theorem 3 (a property is guaranteed exactly when its mechanism is implemented and measuring) is load-bearing here because "unacceptable cyber risk" is a property of a socio-technical system, not of a model. Declaring a capability cutoff does not implement the property; it asserts it. The property — that harm stays below an accepted level — only exists when something is continuously producing the measurement that would tell you it had been crossed. A line nobody is measuring is not a threshold; it is a hope.

The paper reframes the question from "Does this model cross a threshold?" to "How likely is it to cause harm under real conditions?" That is the correct reframe, and it has a structural consequence: the threshold becomes a maintained state, not a passed milestone. You instrument the deployment, feed evidence into a probabilistic model, and keep the posterior below the accepted level — or you restrict deployment. The cutoff is the output of a running mechanism, never a substitute for one.

Capabilities are not risk; context is

The paper states the critique plainly: capabilities ≠ risk. The same capability may be benign or catastrophic depending on context, access, and defenses. A phishing-generation capability behind an enterprise authentication boundary with a mature mail filter is a different risk than the same capability exposed to anyone with a phone number and a grievance. Frontier frameworks elide this by indexing on the capability and leaving the context implicit.

Two further flaws compound the problem. First, the language is vague — "meaningful increase" appears across frameworks without a baseline, a unit, or a method. "Meaningful" relative to what, measured how, by whom, updated when? A threshold term that cannot be populated is a compliance artifact, not a control. Second, the frameworks overfocus on extreme, low-probability scenarios — autonomous end-to-end exploitation of a hardened target — while missing the incremental shifts that actually reshape the offense–defense balance. The slow drift in attacker economics, not the cinematic scenario, is where the balance tips.

[UNIQUE INSIGHT] The reason capability thresholds miss incremental drift is that they are point-in-time declarations about a model, not longitudinal measurements about a system. A capability benchmark is a photograph; risk is a video. The Berkeley proposal — decompose risk into variables, link them through probabilistic dependencies, feed in benchmark and red-team and real-world evidence, and update over time — is, in effect, a request to start filming. The threshold then lives in the trend of the posterior, not in a single benchmark score.

This matters because regulation is moving toward the same framing. The paper notes alignment with the EU AI Act and the NIST Risk Management Framework, both of which ask for ongoing, evidence-based risk treatment rather than a one-time declaration. A lab with only a capability cutoff cannot answer the question those frameworks actually ask: "What is your measured risk, and what is your mechanism for keeping it below the accepted level?"

The Bayesian network is the measuring mechanism

The paper's constructive proposal is Bayesian networks (BNs): probabilistic graphs that represent relationships between variables — AI capability, attacker behavior, defense detection, economic impact — and that integrate diverse evidence and update continuously as conditions change. Unlike a static threshold, a BN lets you track how close a system is to crossing a risk boundary, because the boundary is a region in a joint distribution, not a single number on one axis.

This is the part of the paper that connects most directly to how we think about mechanism. A BN is not a prediction in the colloquial sense; it is a measurement instrument. It says, given what we have observed, here is the posterior over the harm variables, and here is how that posterior moves when we feed in a new red-team result or a new incident report. The threshold is then a policy on that posterior — "deployment is restricted when P(significant harm) exceeds 0.X" — and the policy is enforceable because the posterior is reproducible from the evidence and the graph.

Three properties make a BN the right shape for this job, and each maps to a property we look for in any measuring mechanism:

  • Decomposition. A high-level risk ("AI enables scalable phishing") is broken into measurable variables — AI linguistic mastery, lure credibility, defense detection rate, target susceptibility. You cannot measure "phishing risk" directly; you can measure its components. The decomposition is what makes the risk legible to evidence.
  • Dependency. The variables are linked through conditional probabilities, not summed independently. Attack success depends on capability and defense jointly, not on either alone. This is why a capability-only threshold fails: it assumes the joint can be read off one marginal.
  • Update. The posterior is a function of the evidence, and the evidence accrues. A BN is a longitudinal instrument by construction; yesterday's posterior is today's prior. That is what "dynamic monitoring" means operationally, not as a slogan.

The honest limit the paper itself names is that these models are early-stage: no real-world validated BNs at scale, no standardized datasets for populating the probabilities, no clear governance integration — who sets the thresholds, how they are enforced. We quote the paper's own "What's Missing" section because an honest reading does not paper over it. A measuring mechanism never calibrated against outcomes is a hypothesis about a measurement, not a measurement yet. The mechanism is the right shape; the evidence base is not built.

How the phishing case study turns "meaningful increase" into a number

The paper's phishing case study (pages 31–32) is where the abstraction becomes concrete. A Bayesian network breaks social-engineering risk into nodes — "AI linguistic mastery," "lure credibility," "defense detection rate," "target susceptibility" — that jointly determine outcomes such as whether an employee opens a malicious email. Vague concepts become quantifiable signals because each node is something you can produce evidence for: a linguistic-mastery score from a benchmark, a detection rate from a mail-filter test, a susceptibility estimate from a phishing-simulation campaign.

The case study is doing two things at once, and it is worth separating them. First, it shows that a risk term like "meaningful increase in phishing effectiveness" can be decomposed into variables that have units. Second, it shows that the increase itself is a change in a joint distribution, not a change in one score. A 10% improvement in linguistic mastery does not translate linearly into a 10% increase in opened emails; it translates through the detection rate and the susceptibility node, and the translation is different in a defended enterprise than in an undefended small business.

This is the gap between a capability benchmark and a risk threshold stated precisely. A benchmark tells you the model got better at a task. A BN tells you how much that improvement moved the harm posterior under the conditions you actually deploy in. The first is necessary; the second is what a threshold requires. The paper's contribution is to show that the second is constructible, not merely desirable — and to show, honestly, that the work of populating it at scale has barely begun.

What this maps to at Everythink

We read this paper through a specific lens, and it is worth being explicit so the reader can check our reasoning rather than take it on faith.

[PERSONAL EXPERIENCE] The HAI Engine has run in production since 2016, and the lesson of nine years of operating a multi-agent forecasting system is that a forecast you cannot update is a forecast you should not publish. Our Sisters produce scenario drafts; the Oracle merges them into a calibrated ensemble whose probabilities sum to one and whose entropy is tracked in nats. That ensemble is a measurement instrument, not a prediction in the colloquial sense: reproducible from its inputs, updated as evidence accrues, wrong in traceable ways when it is wrong. The Berkeley BN proposal is the same shape applied to a different domain — harm variables instead of scenario probabilities, red-team evidence instead of signal feeds, but the same insistence that the number is only trustworthy if the mechanism that produced it is inspectable.

The structural principle we share is the one we call "the space is the router": the network→community→room topology routes a request before anything responds. The routing is a mechanism that produces a measurable outcome — which context saw which signal — and that is what makes the downstream forecast auditable. A threshold without a routing topology is a number attached to a model; a threshold inside a routing topology is a number attached to a path through a system, which is the only place harm actually lives. Harm is contextual, and context is a topology.

Our honesty tags exist for the same reason the paper insists on evidence over assertion. Where we have a production mechanism we say Production ✅ — the HAI Engine, Social, Campaigns, Whitelabel Network, World Monitor, the Sisters, the Oracle. Where the mechanism is partial we say Partial ⚠️ — Matchmaking, Marketplace, Calendar. Where it is a roadmap we say Roadmap 🔵 — Wallet & Token, Super App, Community Credit — and we do not promise outcomes for those because they are pre-revenue and, for the token components, subject to Howey review. The point is not the glyph; the point is that a reader can map every claim to a maturity state and hold us to it. That is the same discipline the paper asks frontier labs to adopt: stop asserting thresholds, start producing the measurements that would justify them.

We are explicit about the scope boundary because the paper is too. Everythink operates in civil and defensive contexts only. A risk-threshold mechanism honest about offensive capability is still not a license to build offensive tooling; it is a reason to know where the defensive line is and to instrument it. The Berkeley framework is useful to us precisely because it makes the defensive posture measurable rather than rhetorical.

The governance gap the paper names honestly

The paper's most useful section may be its own "What's Missing" list, because it is the part that prevents the proposal from becoming another capability assertion. The framework lacks real-world validated Bayesian models at scale, standardized datasets for populating probabilities, clear guidance on governance integration, and empirical benchmarks linking model capability to real-world cyber impact. More complex domains — autonomous exploitation, supply-chain attacks — are not deeply operationalized.

We would add one gap the paper gestures at but does not dwell on: a threshold mechanism is only as trustworthy as the independence of the evidence feeding it. A lab that populates its own BN with its own red-team results, against its own baselines, with no external check, has built an instrument that reports whatever the lab needs it to report. The call for standardized datasets is partly a call for evidence that does not all originate with the party being measured. This is the governance question the framework leaves open, and it is the right one to leave open — the answer is institutional, not technical.

The pragmatic consequence is that a threshold mechanism worth governing with is one a third party can reproduce: published graph, named evidence, stated priors, recomputable posterior. A BN that meets those conditions is auditable in the way a regulator can actually use; one that does not is a proprietary claim dressed as a measurement.

Key takeaways

  • A capability cutoff is not a risk threshold; it is a point-in-time assertion about a model that ignores the context where harm lives.
  • The Berkeley paper reframes the question from "does this model cross a line?" to "how likely is it to cause harm under real conditions?" — a move from a static milestone to a maintained measurement.
  • Bayesian networks are the proposed measuring mechanism: decompose risk into variables, link them through conditional probabilities, feed in benchmark and red-team and real-world evidence, and update continuously.
  • The phishing case study shows that vague terms like "meaningful increase" can be decomposed into nodes with units — but the work of populating these models at scale has barely begun, and the paper says so.
  • Theorem 3 generalizes the principle: a property (acceptable risk) is guaranteed exactly when its mechanism (a continuously updated, inspectable posterior) is implemented and measuring.
  • A threshold mechanism is only governable if a third party can reproduce it — published graph, named evidence, stated priors, recomputable posterior.

Frequently asked questions

Why is a capability benchmark not a risk threshold? A benchmark measures what a model can do in a test; risk is what happens when that capability meets a defender, an attacker, and a target in the real world. The benchmark is one input to a risk estimate; it is not the estimate. The Berkeley paper's point is that treating the marginal as if it were the joint is the core error of frontier-framework thresholds.

What does a Bayesian network add that a checklist does not? A checklist records that a control exists; a BN records how much the control moved the harm posterior. The first is binary and static; the second is continuous and updatable. A BN also makes the dependencies explicit — attack success depends on capability and defense jointly — which a checklist cannot represent without becoming the BN in prose form.

Is the Berkeley framework ready to govern with today? No, and the paper says so directly. It lacks validated models at scale, standardized datasets, and clear governance integration. It is a methodology — a repeatable way to operationalize risk — not a finished instrument. Treating it as finished would repeat the error it criticizes: asserting a property instead of measuring it.

How does this connect to Theorem 3? Theorem 3 says a property is guaranteed exactly when its mechanism is implemented and measuring. "Acceptable cyber risk" is a property of a deployment, not a model. The guarantee exists only when a BN (or equivalent instrument) is running, fed, and inspectable. Declaring a cutoff implements no mechanism and therefore guarantees nothing.

What does Everythink use this for? We use the same structural principle — a reproducible, updatable posterior from an inspectable mechanism — in the HAI Engine and the Oracle, for forecasting rather than cyber-risk thresholds. The routing topology (network→community→room) is what makes our forecasts contextual and auditable. We do not build offensive tooling; the value of a measurable defensive line is knowing where it is.

If this framing is useful, the work that follows it is operational: pick a risk, decompose it, link the variables, name the evidence, and start updating. Create your network, or read the papers.

Sources

Build your world on an engine that proves what it claims.

Create your own network on the engine that's run since 2016 — or talk to the team behind the 21 papers.