Evidence
This page exists because “we take security seriously” is not a measurement.
Below are the real results behind our work. We keep an internal register that traces every number we publish back to the run that produced it, and we re-test across multiple random seeds before quoting a figure — because a single good run is not a result.
None of the research on this page is for sale. It’s here so you can judge whether our scanners are built by people who check their own work.
Guardian 1.7B — what we actually measured
A 1.7B on-device security model, trained to classify threats and emit a structured decision.
| Metric | Result |
|---|---|
| Accuracy | 89.9% |
| False positive rate | ~1.6% (mean of 5 seeds) |
| Threat detection | 96.1% — so ~3.9% of threats were missed |
| Output validity | 99.8% |
| Model size | 922 MB |
Measured across five random seeds on a clean held-out set. Per-seed false-positive rates: 0.6 / 1.8 / 4.0 / 0.9 / 0.6% — every seed usable, none collapsed.
We quote the mean of five seeds, not the best one. The result worth having here isn’t the headline number, it’s the spread: all five landed in the same basin. An earlier and smaller model of ours did not, and that instability is exactly what a single lucky run hides.
Why we don’t sell a Guardian
A separate, smaller model of ours — 419M, trained for wallet transaction screening — scored 100% on its held-out test set. It looked shippable.
Then we tested it on 11 attack types it had never been trained on. It collapsed:
| Familiar attacks | Novel attacks | |
|---|---|---|
| Overall accuracy | 1.00 | 0.46 |
| False blocks | 0.00 | 0.29 |
| Asked instead of guessing | 1.00 | 0.00 |
The concrete failure: it blocked “revoke an approval” — a security-positive action — because the wording pattern-matched a phishing template it had memorised.
The two properties that made it look ready — rarely blocking good transactions, and knowing when to ask a human — were both memorised per template, and evaporated the moment the attack was unfamiliar. It had learned the templates, not the skill.
A security model that only holds against attacks that already existed is not a security model. So we didn’t ship it, and we don’t list it as a product. This is a solved-in-principle architecture waiting on much broader training data, and we’ll say it’s ready when it survives that test — not before.
(This is a different model and a different task from the 1.7B above. We keep them separate because conflating them would flatter both.)
The 740-repository code study
Behind Axiom Scan’s waste score is an empirical study run in March 2026:
- 740 open-source repositories
- 28,455 files with git history
- 80,053 bug-fix commits
- 63,754 structural issues detected
The question was whether structurally wasteful code actually attracts more bug fixes.
| Waste score | Files | Avg. bug-fix commits |
|---|---|---|
| 0–20 | 15,035 | 1.83 |
| 20–50 | 9,176 | 3.74 |
| 50–80 | 2,928 | 4.26 |
| 80+ | 1,309 | 4.38 |
The trend is real — the worst band averages about 2.4× the bug fixes of the cleanest. But the Spearman correlation is ρ = 0.13, which on the scale we set before running it is below “moderate”. And the share of files with any bug fix at all barely moves across bands (42.2% at the cleanest, 42.1% at the worst).
So we don’t claim the waste score predicts bugs. It’s a maintainability signal: it tells you which code will be expensive to live with. Anyone selling you a structural score as a bug oracle is over-reading a weak correlation — we ran the study, and this is what it supports.