Cognitive Behavioral Modeling with Activation Steering

Published in NeurIPS CogInterp Workshop, 2025

Large language models (LLMs) frequently exhibit cognitive behaviors that vary unpredictably across prompts, layers, and contextual settings, posing challenges for both diagnosis and control. We introduce CBMAS, a diagnostic framework for continuous activation steering, which advances cognitive bias analysis beyond discrete pre/post interventions to interpretable behavioral trajectories. By integrating steering vector construction with dense coefficient sweeps, logit-lens bias curves, and layer-specific sensitivity analysis, CBMAS identifies critical tipping points at which minor interventions can invert model behavior and traces how steering effects propagate through layer depth. We argue that such continuous diagnostic methodologies bridge the gap between high-level behavioral assessment and low-level representational dynamics, thereby enhancing the cognitive interpretability of LLMs. Finally, we release a command-line interface and datasets covering multiple cognitive behaviors, available at https://github.com/shimamooo/CBMAS