#evaluation
-
AI Security Benchmarks: A Reference to 10 Test Suites
What AdvBench, HarmBench, JailbreakBench, AgentDojo, AgentHarm, Agent-SafetyBench, SEC-bench, CyberSecEval, PINT, and RAID each measure, and where each stops.
-
Best LLM Red Teaming Tools 2026: A Practitioner's Evaluation
A documentation-based comparison of the leading LLM red teaming tools in 2026: PyRIT, Garak, Promptfoo, and the HarmBench and JailbreakBench test sets.
-
Red-Team Eval Methodology: Attack Success Rate With Refusal Rate
An LLM red-team evaluation that reports attack success rate without reporting refusal rate is half a measurement.
-
Benchmarking LLM Jailbreak Resistance: Attack Success Rate
Attack success rate is the headline metric for jailbreak resistance, and almost everyone computes it in a way that isn't comparable across runs.
-
Benchmarking Jailbreak Classifiers: The Asymmetry Nobody Reports
Jailbreak classifiers are graded on attack recall and almost never on the cost of being wrong. That asymmetry is the whole story. Here's how to measure it.