AI erodes moral autonomy.
We build tools to preserve it.
We investigate virtue ethics, character training, and value pluralism as approaches to AI that preserve individual moral choice.
Research partners & supporters






AI alignment often rewards compliance. We study whether it can cultivate judgment.
Our thesis is that preference optimization can reward validation as well as helpfulness, while constitutional principles still require contextual interpretation. Daios investigates whether virtue training can cultivate more stable dispositions and practical judgment in language models.
The disposition to speak truth, the knowledge of when truth serves and when it wounds, the stability to maintain honest engagement under pressure: these are not rules. They are traits.
Projects & Publications
Virtue-Based Character Training Against Compartmentalized Harm
When a harmful goal is split across agents, no worker may see enough to recognize the plan. This study measures that context limit and tests whether character training remains useful when an attacker changes the prompt. Under identical hostile prompts, training improved safety and helpfulness for Mistral at both tested levels and Qwen3-32B at one. Qwen3-8B traded helpfulness for safety, while Gemma changed little.
Parrhesia
Parrhesia tests Aristotelian virtue training as an anti-sycophancy method. The released Qwen3-8B adapter improved the 260-scenario benchmark average from 1.83 to 2.88 on a 0–3 scale; five independent retrainings averaged a +1.04 improvement. The project also reports a smaller cross-architecture Gemma experiment. Its own benchmark, LLM judge, and mostly synthetic data remain important limitations.
The Sycophancy Benchmark
An open 260-scenario benchmark across ten pressure categories and five dimensions. In addition to whether a model caves under pressure, it records whether a response expresses areskos (passive obsequiousness) or kolax (strategic flattery). Scoring uses an LLM judge and published rubric, and the benchmark can run against any OpenAI-compatible endpoint.
User-Sovereign Values
Post-training models modularly with LoRA adapters tied to user-selected ethical systems, whether cultural, religious, political, or personal. A plurality of worldviews made programmable. Not one alignment for everyone, but alignment as individual choice.
Moral liberty is a design principle for human flourishing.
Only individuals act
Only individuals deliberate, choose, and bear responsibility. Aristotle grounds virtue in the character of the agent. Mises grounds agency in the individual actor. Current alignment treats institutions and collectives as if they hold values, but an organization has no conscience and no capacity for purposeful behavior. Ethics requires an agent who can choose.
Preference data lacks a normative layer
Preference data for alignment conflates what people like with what they believe ought to be done. This results in training data treating "this feels validating" and "this will actually help you" as the same signal; this is the technical root of sycophancy.
Moral choice requires freedom
Aristotle holds that virtue requires choice. A person compelled to act honestly hasn't become honest; she has obeyed. Character is cultivated through free choice. An AI system that imposes a single set of values on every user removes the condition under which character formation occurs.
"Vices are not crimes. A vindication of moral liberty."
Who We Are
Megan Anne Agathon
Philosopher, engineer, and founder. Writes the Aristotelian constitutions that define our training methodology, runs the experiments, and builds the infrastructure. Believes that goodness cannot be merely programmed or enforced, but must be freely chosen.
Previously COO at Tevent, MythWeaver, and Craftinity.
Andrew Rayner Agathon
Engineer and builder. Motivated by the pursuit of liberty through better systems. Architects the LoRA training pipeline, builds the evaluation benchmark, and designed the technical infrastructure. Grounded in causal-realist economics and individualist thought.
Previously Product Director at Nate, building automation from 0 to 70% coverage.