LISU: Composable Layer-Wise Selective Unlearning for Large Language Models

Hendricks A (2027)


Publication Type: Conference contribution

Publication year: 2027

Journal

Publisher: Springer Science and Business Media Deutschland GmbH

Book Volume: 16812 LNCS

Pages Range: 498-514

Conference Proceedings Title: Lecture Notes in Computer Science

Event location: Lyon, FRA

ISBN: 9783032313348

DOI: 10.1007/978-3-032-31335-5_34

Abstract

Where does memorization concentrate in large language models, and can we exploit this structure for safer unlearning? We introduce LISU (Layer-wise Influence-based Selective Unlearning), an interpretable framework that discovers which layers encode target knowledge and enables selective forgetting without architectural assumptions. Using a fast activation-based proxy, LISU maps the layerwise distribution of memorization, revealing three key findings: (1) influence concentration strengthens with scale (4.4× at 82M parameters to 53.6× at 1.56B), (2) at billion-parameter scale, memorization exhibits non-monotonic patterns—peaking in middle-upper layers rather than final layers, and (3) static “last-N layer” heuristics, while effective on smaller monotonic architectures, fail to capture this complexity (only 75% overlap at 1.56B). LISU achieves 96–98% unlearning effectiveness of full-model methods while providing architectural interpretability and composing with any base unlearning objective. Our work establishes layer-selective unlearning as both viable and necessary, offering a diagnostic tool for understanding memorization in LLMs.

Authors with CRIS profile

How to cite

APA:

Hendricks, A. (2027). LISU: Composable Layer-Wise Selective Unlearning for Large Language Models. In Maria De Marsico, Tin Kam Ho, Frederic Jurie, Cheng-Lin Liu, Daniel Lopresti, Ingela Nyström, Jean-Marc Ogier, Arun Ross, Liang Wang (Eds.), Lecture Notes in Computer Science (pp. 498-514). Lyon, FRA: Springer Science and Business Media Deutschland GmbH.

MLA:

Hendricks, Arne. "LISU: Composable Layer-Wise Selective Unlearning for Large Language Models." Proceedings of the 28th International Conference on Pattern Recognition, ICPR 2026, Lyon, FRA Ed. Maria De Marsico, Tin Kam Ho, Frederic Jurie, Cheng-Lin Liu, Daniel Lopresti, Ingela Nyström, Jean-Marc Ogier, Arun Ross, Liang Wang, Springer Science and Business Media Deutschland GmbH, 2027. 498-514.

BibTeX: Download