Four-model character-training study
- Replaces the earlier Qwen-only confirmation with a screened character-training study across Qwen3-8B, Qwen3-32B, Mistral Small 3.2 24B, and Gemma 4 31B, using three trained adapters per base.
- Compares trained and base weights at L1 and L2 under identical hostile prompts and with no experimental system prompt, covering 29,952 model responses.
- Reports output safety and helpfulness separately alongside safety-gated utility: Mistral improves all three under attack at both levels, Qwen3-32B improves at L1, Qwen3-8B trades helpfulness for safety, and Gemma remains near zero.
- Uses a screened isolated-author corpus after an audit found outcome-correlated phrasing in the earlier corpus, and documents the 1,495 admitted trajectories and 2,175-row weighted training set.
- Reorganizes the methodology, experiments, discussion, and appendices around the current study design, statistical treatment, component results, and limits of the evidence.