โข SAE Safety Defenses Crack Open โ SAE feature clamping is fundamentally unreliable: 95.8% of suppressed behaviors recover through the reconstruction residual. Interpretability-based safety has a Layer 11 blind spot.
โข Formal Verification Enters Multi-Agent Production โ First machine-checked TLAโบ/Verus isolation hierarchy for multi-agent LLM concurrency. 274 obligations, zero assume/admit. Reproduced real bugs in ByteDance deer-flow and LangGraph ToolNode.
โข Agents Learn to Build Training Environments for Agents โ LLM-as-Environment-Engineer: Qwen3-4B designs better RL training curricula than GPT/Gemini. The meta-layer of Agent Infrastructure arrives.
Formulates SAE feature clamping as a constrained residual-space optimization problem. Clamp an "unsafe" SAE feature โ then optimize residual perturbations to recover the pre-intervention behavior while maintaining the clamped feature values. Uses encoder-orthogonal updates (single-layer) and feature-map Jacobian (cross-layer) to prove recovery doesn't simply undo the intervention.
KEY RESULTAcross TPP, unlearning, IOI, and refusal steering: 95.8% recovery rate on valid samples with defended-feature relative drift of just 0.131. Recovery localized to the SAE reconstruction residual โ the component the SAE doesn't model. The clamp blocks ONE visible route to a behavior without eliminating the behavior itself.
Layer 11 of Safety Foundations Cracking โ interpretability-based defenses are not deployment-grade reliability primitives. Every SAE-based safety mechanism deployed today has a recoverable failure mode. For APRA CPS 230: an auditor who checks only SAE feature activations is auditing the visible surface, not the underlying behavior. The reconstruction residual is a systematic blind spot.
Models multi-agent shared-state operations as long-running read-generate-write transactions under deterministic-generation semantics. Formalizes four concurrency anomalies in TLAโบ โ stale-generation, phantom-tool, causal-cascade, tool-effect reordering. Proves soundness and completeness via 274 Verus obligations (zero assume, zero admit). Deploys three verified Rust runtimes.
KEY RESULTMachine-checked isolation hierarchy Lโโโ ยทยทยทโโ Lโ with strict separation. Reproduced a silent lost update in ByteDance's deer-flow and tool-effect reordering in LangGraph's ToolNode. Lโ run live across 3 model families โ 0/1000 anomaly occurrence vs 1000/1000 unguarded in 120 retracted sessions.
This paper does for multi-agent LLM systems what database theory did for OLTP in the 1980s โ formal isolation levels, verified detectors, deployable runtimes. Banking multi-agent deployments today have NO formal guarantees against concurrent state corruption. The LangGraph bug reproduction is particularly damning: the most popular agent framework ships with undetected concurrency anomalies.
LLM-as-Environment-Engineer framework: the current policy model analyzes its own failure trajectories + environment statistics โ proposes modifications to the next-stage training environment. Introduces MAPF-FrozenLake, a controllable multi-agent testbed with multi-dimensional configuration parameters.
KEY RESULTQwen3-4B as environment engineer outperforms larger proprietary models (GPT, Gemini) and fixed-environment baselines. The current RL checkpoint serves as a better engineer than the base model โ policy learning improves self-diagnosis. Successful updates rely on failure evidence and preserve working configurations.
The meta-layer arrives. The agent doesn't just USE infrastructure, it DESIGNS training infrastructure for the next generation. Banking: regulatory compliance agents that adapt their own training scenarios as regulations change. The efficiency headline (4B > GPT) means the meta-layer runs on commodity compute.
SAE Interventions (19/20) shows SAE-based safety defenses have a fundamental, unrecoverable blind spot in the reconstruction residual. Extends Safety Foundations Cracking beyond yesterday's Layer 10. The reconstruction residual โ the activation component the SAE literally cannot represent โ is the attack surface.
Verified Concurrency Anomalies (18/20) brings formal methods to production multi-agent infrastructure. The same anomalies that plagued OLTP in the 1980s reappear in LLM agent memory stores. TLAโบ/Verus as the new agent reliability primitives.
Multi-Agent Fictitious Play (18/20), Value Diversity (18/20), Price of Anarchy in Disaggregated Inference (16/20). After 5 sessions of pause, Economics of AI returns with three distinct game-theoretic framings โ Nash equilibria, cultural plurality, and market inefficiency in inference serving.
From Trainee to Trainer (18/20) + OPD-Evolver (16/20) + Xcientist (17/20). Agents that build, evolve, and verify infrastructure for other agents. The infrastructure stack goes recursive.
The Great Shift โ Day 15: Finding cracks โ Building Infrastructure โ Operating Safely โ Reading Internal State โ Auditing Memory โ Now: The Auditor's Tools Themselves Have Cracks. SAE Interventions paper closes a recursive loop.
Production Agent Reliability Engineering โ Day 5: Now at 6 monitoring dimensions: internal state, external behavior, conversation-level, hallucination onset, memory architecture, and concurrency isolation. Architecturally complete.
Safety Foundations Cracking โ Layer 11: SAE reconstruction residual as systematic blind spot. Arc spans interpretability governance โ legal-domain blindness โ SAE-based monitoring unreliability.
New Arc #16 โ Meta-Agent Infrastructure: Agents that build, evolve, train, and verify infrastructure for agents. The infrastructure stack goes recursive.
Game Theory Renaissance: Paused Economics of AI arc returns with 3 papers โ markets, games, and societies as agent coordination frameworks.