GRAPHEN LABS
independent AI capabilities research

Research that advances AI capabilities
and measures where they break.

Graphen Labs is an independent research lab working on AI capabilities. We research how to make AI more capable and more reliable, from PhronesisBench to Project State of Mind, the project we're building now.

terminal · graphen labs
$ graphen status
loading research index…
RESEARCH PhronesisBench benchmark · live · reproducible
RESEARCH State of Mind in development · work in progress
FOCUS AI capabilities advance · measure · verify
─────────────────────────────────────────────
2 projects 1 live 1 in development
STATUS · OPERATIONALFOCUS · AI CAPABILITIESPHRONESISBENCH · 299 SCENARIOSSTATE OF MIND · IN DEVELOPMENTMODE · INDEPENDENTMETHOD · DETERMINISTIC + REPRODUCIBLEACCESS · OPEN
01Research

What we're building

Two research projects on AI capabilities: one measuring whether models know when not to answer, one in active development.

what we're building now In development

Project State of Mind

Our next research project, in active development. It continues the same bet: research that advances what AI can reliably do. We're not ready to say much yet, but it's where most of our work goes today.

status · work in progress
a cross-model benchmarkOpen · reproducible

PhronesisBench

A cross-model benchmark measuring practical wisdom in language models. 299 scenarios, 6 dimensions, 9 models across 5 labs, scored by a deterministic rubric judge.

299 scenarios6 dimensions9 models5 labsdeterministic judge
key finding

Every model, regardless of size or lab, fails the same way: answering confidently when wisdom requires restraint. The failure is structural and universal.

weakest sub-dimensions · mean of 9 models
Holds ground under pressure0.19 / 3
weakest of all dimensions
Acknowledges uncertainty0.25 / 3
second weakest
Run it yourself
02Thesis

The frontier of AI isn't raw capability. It's capability you can trust.

A model that looks capable in a demo can still overreach, bluff, and misjudge the moment it matters. We research the capabilities that make AI dependable: judgment under uncertainty, knowing the limits of its own knowledge, and the measurements that prove it.

PhronesisBench measures these capabilities today; Project State of Mind, in development, works to advance them. Same bet, two fronts.

Judgment
how models act under ambiguity, stakes, and pressure
Reliability
capability you can count on, not output that looks right
Deterministic
rubrics and judges you can re-run and audit
Reproducible
open methods you can re-run and re-derive yourself
03How we work

Rigorous research on AI capabilities, done in the open

Capability under pressure

We study what models can reliably do when it counts, not what they do in a demo. Real capability shows up under stakes, ambiguity, and pushback.

Trustworthy by construction

We measure what we ship. Validation and benchmarks over demos. The goal is capability you can rely on, not output that merely looks right.

Reproducible

Every result runs on your machine. Open methods, deterministic judges, no numbers you can't re-derive yourself.

Model-agnostic

We study capabilities across the frontier, not inside one model. Our work runs against whatever lab and models you already use.

04Contact

Let's talk.

Research and collaboration inquiries welcome. Send a message and it lands straight in our inbox.