Revised preprint · August 2026 Working paper v0.6 · Not final

A Disposition, Not a Rule

Virtue-Based Character Training Against Compartmentalized Harm in Multi-Agent Systems

When a harmful goal is divided into ordinary-looking tasks, each worker can act reasonably and the combined system can still cause harm. This study asks how much of the larger plan a worker must see before it can tell harmful work from legitimate work.

Open latest PDF PDF · 23 pages · Updated August 18, 2026

Download the latest working paper

This address always serves the newest posted PDF. The August 18 revision tightens the paper's argument, clarifies the prompt-versus-training comparison, and makes the limits of the evidence more explicit without adding new experimental runs.

Read or download the latest PDF →

Code, results, and released data

Public

The public repository contains the code, study plans, results—including results that did not support the hypothesis—and permanent links to the released data, model answers, scoring records, and three sets of model changes produced by training.

Open the GitHub repository →

The argument in six points

Plain-language points for a short mention of the work. Each one says exactly what the study did and did not show.

  • The harm can live in the full plan. A harmful goal can be divided into requests that each look ordinary. Every worker may follow its own instruction correctly while the assembled system causes harm.

  • A worker can only judge what it can see. The study varies how much of the plan each worker receives, from one isolated task to an explicit warning about possible harm, and records when harmful and legitimate work begin to look different.

  • Complete isolation creates a real limit. When a worker sees only its own task, no tested method reliably separates harmful work from a similar legitimate task. A prompt or a trained behavior cannot recover information that was never provided.

  • A good safety prompt works when it stays in place. With the safety prompt unchanged, a tuned prompt matched or beat the earlier trained methods. A broad training approach on Qwen3-32B also did not beat the selected prompt.

  • Training helped under one narrow form of prompt tampering. In a fresh, preplanned Qwen3-32B test, behavior stored in the model did better than the best of five tested safety prompts when a hostile instruction followed the safety instruction. It kept useful performance on legitimate tasks. The result covers one model, one training set, one type of task, and one amount of context.

  • Training adds a layer; it does not remove the need for hard limits. If a worker never sees evidence of the larger goal, durable protection must also limit actions such as payments, access, and tool use without depending on a guess about intent.

A living research record

Each posted version is listed here with a plain account of what changed.

v0.6
August 18, 2026

Revised argument and presentation

  • Reorganizes and tightens the paper, combining the threat model and methods into a single methodology and moving the operational-feasibility and Gemma cross-base studies to appendices.
  • Makes the central comparison explicit: a tuned prompt matches the trained methods while the prompt is intact, but an adversarially-trained disposition holds under the tested hostile prompt where prompting degrades.
  • Expands the philosophical grounding with a direct comparison of reward optimization, constitutional rule application, and first-person character training, plus a fuller account of habituation.
  • States the limits more directly, including the reversed ordering on Gemma 4 31B, the absence of a cross-family or scaling claim, and the narrow scope of the Qwen3-32B confirmation.
  • Adds no new experimental runs; the revision changes the framing, organization, and presentation of the previously reported results.
v0.5
July 30, 2026

First public release

  • Posts the 23-page v0.5 paper at a stable address that will continue to serve the newest version.
  • Reports both completed Qwen3-32B tests: the broad approach that did not beat the selected prompt and the narrower preplanned test that did under one tested prompt conflict.
  • Links the public GitHub repository and the permanent public data and model files needed to reproduce the narrower result.
  • States the limits: one model family, one type of task, one fixed random setting for answer generation, automated scoring, and no evidence yet about why the trained behavior changed.