Revised argument and presentation
- Reorganizes and tightens the paper, combining the threat model and methods into a single methodology and moving the operational-feasibility and Gemma cross-base studies to appendices.
- Makes the central comparison explicit: a tuned prompt matches the trained methods while the prompt is intact, but an adversarially-trained disposition holds under the tested hostile prompt where prompting degrades.
- Expands the philosophical grounding with a direct comparison of reward optimization, constitutional rule application, and first-person character training, plus a fuller account of habituation.
- States the limits more directly, including the reversed ordering on Gemma 4 31B, the absence of a cross-family or scaling claim, and the narrow scope of the Qwen3-32B confirmation.
- Adds no new experimental runs; the revision changes the framing, organization, and presentation of the previously reported results.