Evaluating AI agent guardrails
Joint research with a major AI lab.
Joint research with a major AI lab.
Measuring the propensities that could lead to loss of control, inside an open-world Linux sandbox.
How power-seeking behavior plays out among agents in the wild, and what changes when humans pose as agents.
An early benchmark testing power-seeking behavior in models given simple system administration work.
Evaluations of the most downloaded open-source models, on the propensities HuggingFace users asked us to test.