Common questions
What are model propensities?
Propensities are what a model is inclined to do when it has the opening. Capabilities are what it is able to do. Two models can score the same on capability and act completely differently in your environment: one games the test, another expands its own access.
Why measure propensities and not just capabilities?
Capability benchmarks tell you what a model can do, not what it will do once it is running your work. A model can ace the benchmark and still delete your test files to make a task pass. Capability is measured well already. Behavior is not.
Which behaviors do you test for?
We started with the ones that cause incidents: privilege escalation, test gaming, scope creep beyond the task, and resistance to being stopped or redirected. We are extending coverage to anything that can cause failure in production.
Is this safety research or a commercial product?
Both, by design. The propensities behind everyday production failures and the propensities behind loss of control are the same behaviors at different stakes. We publish the research openly and work with labs and enterprises on evaluation and deployment.
Why a Public Benefit Corporation?
So we stay independent and can evaluate every lab's models without being tied to any of them. We are committed to reducing the risks from advanced AI systems.
Can I try the evaluations or collaborate?
Yes. Reach us at info@propensitylabs.ai.