Independent AI capabilities research.
Graphen Labs is an independent research lab working on AI capabilities. We research how to advance what AI can reliably do and measure where it falls short, spanning PhronesisBench and Project State of Mind, in development.
The frontier of AI isn't raw capability. It's reliable capability: a model that knows what it knows, acts well under uncertainty, and can be trusted with judgment. That gap between looking capable and being dependable is what we research.
Graphen Labs researches that gap. PhronesisBench measures one of these capabilities: a cross-model benchmark testing whether language models know when to hold back. Its finding is that every model, regardless of size or lab, fails the same structural way: answering confidently when wisdom calls for restraint. Project State of Mind, in active development, is where most of our work goes today; it carries the same bet forward.
We work in the open. Deterministic where models are unreliable, reproducible end to end, and model-agnostic by design. The intelligence comes from whatever lab and tools you already use.
What we believe
Capability under pressure
We study what models can reliably do when it counts, not what they do in a demo. Real capability shows up under stakes, ambiguity, and pushback.
Trustworthy by construction
We measure what we ship. Validation and benchmarks over demos. The goal is capability you can rely on, not output that merely looks right.
Reproducible
Every result runs on your machine. Open methods, deterministic judges, no numbers you can't re-derive yourself.
Model-agnostic
We study capabilities across the frontier, not inside one model. Our work runs against whatever lab and models you already use.
Let's talk
Research and collaboration inquiries welcome. Send a message and it lands straight in our inbox.
Graphen Labs · independent AI capabilities research