Polymarket data reveals human-AI hybrid performance is trimodal โ collaborative traits (humility, curiosity) beat raw model benchmarks for forecasting โ hiring design matters more than model selection
LACUNA exposes that SOTA LLM unlearning is precise at output but blind at parameter level โ parameter-localized unlearning is the missing piece for GDPR/right-to-be-forgotten compliance
DiscoBench proves search agents that "search more instead of asking" perform worse than direct guessing โ clarification strategy is the frontier beyond retrieval quality
METHOD: Decomposes human-AI collaboration into three modes using real-money Polymarket prediction data. Analyzes individual forecasters โ most either defer to the model (matching it) or rubber-stamp prior guesses (worse than model alone). A minority engage in genuine complementary reasoning.
KEY RESULT: This minority reaches accuracy matching or exceeding the market itself โ lower error than the market. The predictors are NOT raw cognitive ability or model benchmarks, but collaborative traits: perspective-taking, intellectual humility, and curiosity. Trimodal distribution explains why average effects hide the real story.
๐ SO WHAT: This directly challenges how banks deploy AI analyst tools. The ROI of AI augmentation depends on selecting for collaborative traits, not on chasing the next benchmark point. For APRA-regulated forecasting, hybrid intelligence design (who pairs with what, under what protocol) matters more than model selection. Polymarket as evaluation substrate is also directly relevant to investing.
METHOD: Injects synthetic PII into predefined parameters of 1B and 7B OLMo models via masked continual pretraining, creating known knowledge locations. Evaluates whether SOTA unlearning methods truly target the responsible parameters โ not just whether outputs look clean.
KEY RESULT: Despite strong output-level performance, existing unlearning methods are highly imprecise at parameter targeting and remain susceptible to resurfacing attacks. When localization IS accurate, even a simple gradient-based method achieves strong erasure. The gap between behavioral unlearning and parametric erasure is the core finding.
๐ SO WHAT: GDPR's right-to-be-forgotten and APRA CPS 230 data governance require genuine erasure, not output obfuscation. Every bank deploying LLMs fine-tuned on customer data faces resurfacing risk โ LACUNA provides the first rigorous testbed to audit whether your unlearning actually works at the weight level.
METHOD: Builds a benchmark of 211 samples with 463 ambiguity instances across 11 real-world domains and 4 ambiguity types. Includes a user simulator for multi-turn interactive evaluation. Tests whether search agents can proactively detect ambiguity, ask effective clarification questions, and recover correct reasoning paths.
KEY RESULT: Ambiguity detection and effective clarification are distinct capabilities in current LLMs. Critically: repeatedly searching instead of asking for clarification often performs WORSE than direct guessing. Current agents lack the interaction strategy to know when retrieval isn't the answer.
๐ SO WHAT: Bank analysts use search agents for regulatory research, due diligence, and market intelligence. These queries are inherently ambiguous โ "find precedent for X" means different things in different contexts. DiscoBench exposes a critical gap: agents need to ask clarifying questions rather than silently converge on wrong interpretations. Directly relevant to APRA CPS 230 agent governance.
Week of June 29 โ July 4: Opened with Safety as Geometric Structure (Gravitational Interpretation, 19/20), pivoted to Memory as the Agent Reliability Frontier (AutoMem 19/20, Jul 2), exploded into the Evaluation Infrastructure Crisis with Distributed Attacks + BPE Token Boundary (5ร 18/20, Jul 3), and closes with The "Asking" Frontier โ interaction strategy surpassing raw capability as the differentiator. Arc trajectory: structural safety โ memory optimization โ evaluation crisis โ interaction strategy. Each layer of agent infrastructure reveals its own failure modes and optimization levers. PRE now spans 15+ dimensions.