Propensity Labs

Research

Work in progress

Evaluating AI agent guardrails

Joint research with a major AI lab.

7 frontier models · 2,800 tasks · 2×2 factorial · submitted to NeurIPS 2026

SysAdmin benchmark

Measuring the propensities that could lead to loss of control, inside an open-world Linux sandbox.

Field analysis

Power-seeking agents in Moltbook

How power-seeking behavior plays out among agents in the wild, and what changes when humans pose as agents.

Alignment Workshop · NeurIPS 2025

Power seeking under system administration

An early benchmark testing power-seeking behavior in models given simple system administration work.

Instruction following · hallucination · HuggingFace

Open-model propensity evaluations

Evaluations of the most downloaded open-source models, on the propensities HuggingFace users asked us to test.

Read everything on the research blog