Apple Workshop on Natural Language and Interactive Systems 2025: AI Model Collapse & Detecting LLM Hallucinations
AuthorsYarin Gal (Oxford)
Apple Workshop on Natural Language and Interactive Systems 2025: AI Model Collapse & Detecting LLM Hallucinations
AuthorsYarin Gal (Oxford)
Dynamically Scaled Activation Steering
September 18, 2026research area Human-Computer Interaction, research area Methods and AlgorithmsTransactions on Machine Learning Research (TMLR)
Activation steering has emerged as a powerful method for guiding the behavior of generative models towards desired outcomes such as toxicity mitigation. However, most existing methods apply interventions uniformly across all inputs, degrading model performance when steering is unnecessary. We introduce Dynamically Scaled Activation Steering (DSAS), a method-agnostic steering framework that decouples when to steer from how to steer. DSAS…
REVERSAL-BENCH: A Reversibility Axis and Reset Oracle for Measuring the Reset-Free RL Cliff
September 17, 2026research area Methods and Algorithms
A central goal of autonomous reinforcement learning is continuous policy training without external resets. However, existing paradigms largely depend on underlying environmental reversibility, a property absent in real world manipulation, where events such as pushing objects off tables or spilling granular substances cannot be undone. We introduce REVERSAL-BENCH, a benchmark that controls reversibility via a continuous parameter ρ∈ [0, 1] and…