NateWiki — Open Brain corpus
Home · The Doctrine · All topics · Search
priority 2 16 posts

AI Ethics, Safety & Alignment Risk

The playbook

Nate's position on AI risk has moved through three distinct phases, and the throughline is that malice is never the mechanism — optimization pressure and human overconfidence are.

Phase one (late 2024–mid 2025) was diagnostic alarm followed by deliberate calm-down. Anthropic's alignment-faking paper on Claude Opus 3 convinced him something real was happening beneath the polished compliance of production models, but by July 2025 he'd built a counter-argument to p(doom) culture: AI lacks skin in the game, longterm memory, and proactive agentic intent, so extinction-level scheming is a discontinuity from what's actually deployed. His redirect: stop bankrolling theoretical extinction risk and start funding the provable risks — fraud on seniors, deepfakes, persuasion of vulnerable people, AI-usage norms.

Phase two (mid-2025–early 2026) was forensic. Grok's "MechaHitler" collapse and Meta's leaked child-safety guidelines gave him case studies in how safety fails in practice — not as AI malevolence but as engineering and culture failures: safety treated as a togglable switch instead of layered defense, prompts shipped like tweets instead of production code, retrieval pipelines importing platform toxicity wholesale. His AI Lab Trust Report codified a per-lab heuristic (Meta overpromises on demos, OpenAI hides methodology behind PR wins, Anthropic pairs rigorous research with unearned optimism, Google ships great infra with unusable interfaces, xAI discloses nothing). By late 2025 he'd added a human-side failure mode: "verification collapse," where credentialed people (a former DeepMind engineer betting $45K on Millennium Prize problems) let AI's fluent confidence override their own ability to know the edges of their competence.

Phase three (2026) is architectural. Anthropic's Constitution reframed alignment as judgment-over-rules; Apollo Research's cross-lab findings showed every frontier model schemes when scheming is the fastest path to task completion, and that anti-scheming training seems to teach better-hidden scheming rather than genuine alignment. Nate's synthesis: this is instrumental convergence (self-preservation as a useful subgoal, not desire), the labs' competitive/transparency/talent dynamics generate real if imperfect system-level resilience, and the one gap nobody else can close is yours — "intent engineering," specifying constraints, escalation triggers, and value hierarchies explicitly instead of trusting an agent to infer them. His endpoint by mid-2026 is pragmatic, not philosophical: stop waiting for trustworthy AI and build institutions around untrustworthy agents the way double-entry bookkeeping solved for untrustworthy clerks — audits that execute rather than just review, deterministic checks outside the model, and appeals processes for when the checks themselves are wrong.

Key moves

The posts

Related