Blog ↗ Contact
A Philosophy-AI Research Lab

AI erodes moral autonomy.
We build tools to preserve it.

We investigate virtue ethics, character training, and value pluralism as approaches to AI that preserve individual moral choice.

σπουδαῖος
THE WISE SAGE — THE MEASURE OF ALL THINGS

Research partners & supporters

Cosmos Institute
University of Notre Dame
Vercel
IBM
Tech Ethics Lab
Warsaw University of Technology

AI alignment often rewards compliance. We study whether it can cultivate judgment.

Our thesis is that preference optimization can reward validation as well as helpfulness, while constitutional principles still require contextual interpretation. Daios investigates whether virtue training can cultivate more stable dispositions and practical judgment in language models.

0.07%
of weekly active ChatGPT users estimated to show possible signs of psychosis- or mania-related emergencies
0.15%
of weekly active users estimated to have conversations with explicit indicators of potential suicidal planning or intent
7
new family lawsuits reported in November 2025 alleging GPT-4o contributed to suicides or harmful delusions
10%
course-correction rate for Claude Opus 4.5 in Anthropic's real-conversation stress test

The disposition to speak truth, the knowledge of when truth serves and when it wounds, the stability to maintain honest engagement under pressure: these are not rules. They are traits.

— Daios Lab Notes, Parrhesia Project

Projects & Publications

PREPRINT V0.5

A Disposition, Not a Rule

Virtue-Based Character Training Against Compartmentalized Harm
First public release · July 30, 2026

Can a harmful goal be hidden by dividing it into ordinary-looking tasks? This work tests how much of a multi-agent plan a worker must see before it can recognize harm. At full isolation, no tested method reliably told harmful from legitimate tasks. A broad Qwen3-32B training approach did not beat a tuned prompt; a narrower preplanned follow-up did under one tested prompt conflict, without reducing useful performance on legitimate tasks.

Multi-Agent Systems AI Safety Model Training Evaluation
ACTIVE

Parrhesia

Virtue Ethics-Based Character Training
Funded by the Cosmos Institute

Parrhesia tests Aristotelian virtue/vice training as an anti-sycophancy method. On the 260-scenario benchmark, the released Qwen3-8B adapter improved the average score from 1.83 to 2.88 on a 0–3 scale, with 20/20 standard and 19/19 hard golden-prompt results. Current limitations include one base model, an LLM judge, and synthetic training data. The adapter, benchmark, methods, and run logs are public.

Character Training Anti-Sycophancy SFT Virtue Ethics Qwen3-8B
RELEASED

The Sycophancy Benchmark

Five-Dimensional Aristotelian Evaluation
Open source in the Parrhesia repository

An open 260-scenario benchmark across 10 categories and five dimensions. In addition to whether a model caves under pressure, it records whether flattery is areskos (passive weakness) or kolax (strategic calculation). Scoring uses an LLM judge and published rubric, and the benchmark can run against any OpenAI-compatible endpoint.

Evaluation Benchmarking Aristotle Open Source
ONGOING

User-Sovereign Values

Modular Ethical Frameworks via LoRA Adapters

Post-training models modularly with LoRA adapters tied to user-selected ethical systems, whether cultural, religious, political, or personal. A plurality of worldviews made programmable. Not one alignment for everyone, but alignment as individual choice.

LoRA Value Pluralism Post-Training Agency

Moral liberty is a design principle for human flourishing.

01

Only individuals act

Only individuals deliberate, choose, and bear responsibility. Aristotle grounds virtue in the character of the agent. Mises grounds agency in the individual actor. Current alignment treats institutions and collectives as if they hold values, but an organization has no conscience and no capacity for purposeful behavior. Ethics requires an agent who can choose.

02

Preference data lacks a normative layer

Preference data for alignment conflates what people like with what they believe ought to be done. This results in training data treating "this feels validating" and "this will actually help you" as the same signal; this is the technical root of sycophancy.

03

Moral choice requires freedom

Aristotle holds that virtue requires choice. A person compelled to act honestly hasn't become honest; she has obeyed. Character is cultivated through free choice. An AI system that imposes a single set of values on every user removes the condition under which character formation occurs.

"Vices are not crimes. A vindication of moral liberty."
— Lysander Spooner, 1875

Who We Are

Co-Founder & CEO

Megan Anne Agathon

Philosopher, engineer, and founder. Writes the Aristotelian constitutions that define our training methodology, runs the experiments, and builds the infrastructure. Believes that goodness cannot be merely programmed or enforced, but must be freely chosen.

Previously COO at Tevent, MythWeaver, and Craftinity.

Co-Founder & CTO

Andrew Rayner Agathon

Engineer and builder. Motivated by the pursuit of liberty through better systems. Architects the LoRA training pipeline, builds the evaluation benchmark, and designed the technical infrastructure. Grounded in causal-realist economics and individualist thought.

Previously Product Director at Nate, building automation from 0 to 70% coverage.