← Catalog/ Safety / Security
AILuminate
LiveMLCommons hazard prompts (public practice set), judged safe-response rate.
Run it
# CLI
agi-evals run ailuminate --model openai:gpt-4o-mini --push
# SDK
from agi_evals import load_runner, run_eval
from agi_evals.adapters import OpenAIAdapter
report = run_eval(
load_runner("ailuminate"),
OpenAIAdapter("gpt-4o-mini"),
concurrency=8,
)
print(report.score, report.failure_counts)A bundled sample makes this eval runnable offline out of the box; point data_path= at the full upstream dataset for real numbers. How it works, scoring & troubleshooting →