Probing Layerwise Sensitivity for LLM Security Analysis
Published in NeurIPS Lock-LLM Workshop, 2025
PLSLSA is a framework for diagnosing and interpreting LLM security-domain behaviors by continuously steering activations and analyzing how intervention strength and layer depth affect model biases.
