Evidence

This page exists because “we take security seriously” is not a measurement.

Below are the real results behind our work. We keep an internal register that traces every number we publish back to the run that produced it, and we re-test across multiple random seeds before quoting a figure — because a single good run is not a result.

None of the research on this page is for sale. It’s here so you can judge whether our scanners are built by people who check their own work.


Guardian 1.7B — what we actually measured

A 1.7B on-device security model, trained to classify threats and emit a structured decision.

MetricResult
Accuracy89.9%
False positive rate~1.6% (mean of 5 seeds)
Threat detection96.1% — so ~3.9% of threats were missed
Output validity99.8%
Model size922 MB

Measured across five random seeds on a clean held-out set. Per-seed false-positive rates: 0.6 / 1.8 / 4.0 / 0.9 / 0.6% — every seed usable, none collapsed.

We quote the mean of five seeds, not the best one. The result worth having here isn’t the headline number, it’s the spread: all five landed in the same basin. An earlier and smaller model of ours did not, and that instability is exactly what a single lucky run hides.


Why we don’t sell a Guardian

A separate, smaller model of ours — 419M, trained for wallet transaction screening — scored 100% on its held-out test set. It looked shippable.

Then we tested it on 11 attack types it had never been trained on. It collapsed:

Familiar attacksNovel attacks
Overall accuracy1.000.46
False blocks0.000.29
Asked instead of guessing1.000.00

The concrete failure: it blocked “revoke an approval” — a security-positive action — because the wording pattern-matched a phishing template it had memorised.

The two properties that made it look ready — rarely blocking good transactions, and knowing when to ask a human — were both memorised per template, and evaporated the moment the attack was unfamiliar. It had learned the templates, not the skill.

A security model that only holds against attacks that already existed is not a security model. So we didn’t ship it, and we don’t list it as a product. This is a solved-in-principle architecture waiting on much broader training data, and we’ll say it’s ready when it survives that test — not before.

(This is a different model and a different task from the 1.7B above. We keep them separate because conflating them would flatter both.)


The 740-repository code study

Behind Axiom Scan’s waste score is an empirical study run in March 2026:

The question was whether structurally wasteful code actually attracts more bug fixes.

Waste scoreFilesAvg. bug-fix commits
0–2015,0351.83
20–509,1763.74
50–802,9284.26
80+1,3094.38

The trend is real — the worst band averages about 2.4× the bug fixes of the cleanest. But the Spearman correlation is ρ = 0.13, which on the scale we set before running it is below “moderate”. And the share of files with any bug fix at all barely moves across bands (42.2% at the cleanest, 42.1% at the worst).

So we don’t claim the waste score predicts bugs. It’s a maintainability signal: it tells you which code will be expensive to live with. Anyone selling you a structural score as a bug oracle is over-reading a weak correlation — we ran the study, and this is what it supports.