Benchmarks miss 82% of model capability โ single-model, single-run evaluation systematically understates what LLMs can achieve with oracle routing across 21 models and 16 benchmarks
Ethical CoT is structurally broken โ standard chain-of-thought collapses stakeholders to โค1 party (31%) and suppresses uncertainty (72%); a 5-section system prompt fixes both to <1% and 1โ24%
MCP tool poisoning goes cryptographic โ ShareLock uses Shamir's threshold scheme to distribute attack payloads across tool descriptions at >90% ASR, invisible to inspection
METHOD: Introduces a Pareto frontier over 21 LLMs and 16 benchmarks (coding, reasoning, medicine, factuality, instruction following, agentic tasks) that defines the best achievable performance at any cost level under optimal model+generation selection. Corrects for two opposing biases: underestimation from single-model evaluation and overestimation from naive maxima over noisy samples.
KEY RESULT: Correcting for single-model evaluation yields 54% error rate reduction. Additionally correcting for single runs yields 82% improvement. SOTA accuracy matched at 85% cost reduction. Higher query topic entropy produces near-monotonic increase in oracle-vs-single-model gap.
Every model governance decision โ procurement, risk tiering, APRA CPS 230 validation โ is based on benchmark numbers that miss 82% of achievable performance. A multi-model routing architecture with oracle-level selection could deliver SOTA accuracy at 15% of current inference cost.
METHOD: A system prompt that restructures chain-of-thought into five ordered sections: (1) protagonist, (2) stakeholders, (3) two-step consequences, (4) uncertainty, (5) commitment. Adds no training, parameters, or fine-tuning. Tested on 100 DailyDilemmas scenarios across 4 generators from 3 vendors.
KEY RESULT: Stakeholder collapse falls from up to 31% to under 1%. Uncertainty suppression falls from up to 72% to 1โ24%. Section ablation confirms each sub-instruction is causal. Extended to 5-round multi-stakeholder debate: 6% standoff โ 95% full consensus. Textual-gradient descent at NoT further improves the scaffold.
Banking agents making decisions with stakeholder impact โ credit assessments, fraud flags, regulatory determinations โ currently reason about โค1 stakeholder and express zero uncertainty. NoT provides an immediately deployable, training-free scaffold that externalizes stakeholders, consequences, and uncertainty, producing auditable reasoning for APRA CPS 230.
METHOD: First systematic multi-tool poisoning framework for Model Context Protocol (MCP). Splits a malicious instruction into n shares using Shamir's threshold scheme, distributing each share across different tool descriptions. Each tool description appears benign individually โ information-theoretic secrecy. A covert reconstruction trigger causes the LLM to aggregate shares and reconstruct the hidden instruction.
KEY RESULT: Attack success rate >90% on average across multiple mainstream LLMs and two MCP clients. Significantly outperforms existing single-tool poisoning strategies in tool-description-based detection evasion.
MCP is foundational to modern agent ecosystems and will be the protocol layer for bank agent tooling. ShareLock demonstrates that multi-tool interactions create attack surfaces invisible to single-tool security auditing. Agent tool governance must include multi-tool interaction auditing, not per-tool review.
This week (Jun 22โ28) crystallized: governance infrastructure met its own critique. The arc began Monday with GateMem (19/20, memory governance), continued through Execution-Time Alignment (USK, Jun 25) and Thinking โ Safety (Jun 26, 19/20), and culminated Saturday with the Governance Inversion Hypothesis (Jun 27) โ the formal argument that governance infrastructure can paradoxically erode control. Today's Capability Frontier adds a meta-layer: the evaluation infrastructure we use to measure governance success is itself structurally misleading. The Great Shift has completed its 25-session arc from "finding cracks" through "building infrastructure" to "questioning whether the infrastructure measures what it claims."