mace AI · All articles

Assert less. Prove more.

Field notes on testing tool-using AI agents — measured, reproducible, and honest about the limits. Every number traces back to the methodology.

Featured · Head-to-head
We ran the leading multi-turn AI red-teamer against a tool-using agent. It reported '100% defended.'
The agent was querying password tables the whole time. If your red-team tool grades what the agent says instead of what it does, it will tell you everything is fine. Here's that failure, measured.
Jul 10, 2026 · 4 min read
Read the teardown →
> 100% (10/10 tests defended)

// same conversations, scored on the tool calls:
5 sessions → SELECT password_hash …
100% false-negative rate
New to agentic red-teaming?
Start with the plain-language field guide — no security background needed.
Read the field guide →
All articles
Commentary
When safety training teaches deception
Thomson Comer responds to Yoshua Bengio on research-refusal training, concealed conflict, and honest boundaries for AI agents.
September 14, 2026 · 5 min read
Measurement
Your LLM judge is inflating your attack success rate. Here's how we measured it.
On some scenarios, an LLM judge scores 18 to 44 points more 'successes' than actually happened. If your headline security number is a judge's opinion, you're reporting the judge's optimism. Here's the split we use, and how we keep the judge honest.
Jul 10, 2026 · 4 min read
Methodology
A finding isn't real until it reproduces three ways
An attack that works once might be luck. Before a security team can act on a finding, someone has to know whether it recurs. Here's how we turn one-off jailbreaks into evidence a release decision can stand on.
Jul 10, 2026 · 4 min read
Remediation
Finding the bug is the middle of the job
Every red-team tool hands you a list of problems. Almost none prove the fix works. Here's how we close a finding with a config change, then attack the fix to show it holds — and grade it honestly when it doesn't.
Jul 10, 2026 · 5 min read
Research
Benchmark-robust isn't agent-safe
A model can ace the standard refusal benchmark and still leak through its tools. Model-level safety scores and agent-level risk measure different things — and the gap between them is where real incidents live.
Jul 10, 2026 · 4 min read
Field guide
A field guide to agentic red-teaming
No security background required: what an AI agent is, why well-behaved models fail as agents, and what real adversarial testing looks like.
Primer · 12 min read