
In scope: the exploitation surface, prompt injection and goal reframing, permission escalation, memory and context poisoning, tool and supply-chain provenance, deployed-agent safety evidence, red-teaming method. Out: model-level alignment research as a subject in itself (models), harness architecture (harnesses).
Analysis only1Show all →
| Type | Source | Published |
|---|---|---|
| ANALYSIS | AI Safety, Alignment, and Interpretability in 2026 Zylos Research · Zylos DPO replacing RLHF analysis, alignment mirages concept, 6 documented failure modes, alignment trilemma | 2026-02-09 |