Publications
2026

Software for dataset-wide XAI: From local explanations to global insights with Zennit, CoRelAy, and ViRelAy
Authors: Anders, C. J., Neumann, D., Samek, W., Müller, K.-R., & Lapuschkin, S.; Published in: PLoS One 21(1): e0336683
Abstract:
The predictive capabilities of Deep Neural Networks (DNNs) are well-established, yet the underlying mechanisms driving these predictions often remain opaque. The advent of Explainable Artificial Intelligence (XAI) has introduced novel methodologies to explore the reasoning behind complex model predictions of complex models. Among post-hoc attribution methods, Layer-wise Relevance Propagation (LRP) has demonstrated notable adaptability and performance for explaining individual predictions – provided the method is used to its full potential. For deeper dataset-wide and quantitative analyses, however, the manual inspection of individual attribution maps remains unnecessarily labor-intensive and time consuming. While several approaches for dataset-wide XAI-analyses have been proposed, unified and accessible implementations of such tools are still lacking. Furthermore, there is a notable absence of dedicated visualization and analysis software to support stakeholders in interpreting both local and global XAI results effectively. This gap underscores the need for comprehensive software tools that facilitate both granular and holistic understanding of model behavior, as well as easing the adaptability of XAI in applications and the sciences. To address these challenges, we present three software packages designed to facilitate the exploration of model reasoning using attribution approaches and beyond: (1) Zennit – a highly customizable and intuitive attribution framework implementing LRP and related methods in PyTorch, (2) CoRelAy – a framework to easily and quickly construct quantitative analysis pipelines for dataset-wide analyses of explanations, and (3) ViRelAy – an interactive web-application for exploring data, attributions, and analysis results. By providing a standardized implementation for XAI, we aim to promote reproducibility in our field and empower scientists and practitioners to uncover the intricacies of complex model behavior.
Access to: full publication

FLIPR: Flexible and interpretable prediction regions for time series
Authors: English, E., & Lippert, C.; Published in: CPAL 2026 (Proceedings Track) Poster
Abstract:
Conformal prediction provides a model-agnostic framework for uncertainty quantification with finite-sample validity guarantees, making it an attractive tool for constructing reliable prediction sets. However, existing approaches commonly rely on residual-based conformity scores, which impose geometric constraints and struggle when the underlying distribution is multimodal. In particular, they tend to produce overly conservative prediction areas centred around the mean, often failing to capture the true shape of complex predictive distributions. In this work, we introduce JAPAN (Joint Adaptive Prediction Areas with Normalising-Flows), a flow-based framework that uses density estimates for several conformal scores. By leveraging flow-based models, JAPAN estimates the (predictive) density and constructs prediction areas by thresholding on the estimated density scores, enabling compact, potentially disjoint, and context-adaptive regions that retain finite-sample coverage guarantees. We theoretically motivate the efficiency of JAPAN and empirically validate it across multivariate regression and forecasting tasks, demonstrating good calibration and tighter prediction areas compared to existing baselines. Furthermore, the several density-based conformity scores showcase the flexibility of our proposed framework.
Access to: full publication

JAPAN: JOINT ADAPTIVE PREDICTION AREAS WITH NORMALISING-FLOWS
Authors: English, E., & Lippert, C.; Published in: ICLR 2026
Abstract:
Conformal prediction provides a model-agnostic framework for uncertainty quantification with finite-sample validity guarantees, making it an attractive tool for constructing reliable prediction sets. However, existing approaches commonly rely on residual-based conformity scores, which impose geometric constraints and struggle when the underlying distribution is multimodal. In particular, they tend to produce overly conservative prediction areas centred around the mean, often failing to capture the true shape of complex predictive distributions. In this work, we introduce JAPAN (Joint Adaptive Prediction Areas with Normalising-Flows), a flow-based framework that uses density estimates for several conformal scores. By leveraging flow-based models, JAPAN estimates the (predictive) density and constructs prediction areas by thresholding on the estimated density scores, enabling compact, potentially disjoint, and context-adaptive regions that retain finite-sample coverage guarantees. We theoretically motivate the efficiency of JAPAN and empirically validate it across multivariate regression and forecasting tasks, demonstrating good calibration and tighter prediction areas compared to existing baselines. Furthermore, several density-based conformity scores showcase the flexibility of our proposed framework.
Access to: full publication

Circuit Insights: Towards Interpretability Beyond Activations
Authors: Golimblevskaia, E., Jain, A., Puri, B., Ibrahim, A., Samek, W., & Lapuschkin, S.; Published in: ICLR 2026
Abstract:
The fields of explainable AI and mechanistic interpretability aim to uncover the internal structure of neural networks, with circuit discovery as a central tool for understanding model computations. Existing approaches, however, rely on manual inspection and remain limited to toy tasks. Automated interpretability offers scalability by analyzing isolated features and their activations, but it often misses interactions between features and depends strongly on external LLMs and dataset quality. Transcoders have recently made it possible to separate feature attributions into input-dependent and input-invariant components, providing a foundation for more systematic circuit analysis. Building on this, we propose WeightLens and CircuitLens, two complementary methods that go beyond activation-based analysis. WeightLens interprets features directly from their learned weights, removing the need for explainer models or datasets while matching or exceeding the performance of existing methods on context-independent features. CircuitLens captures how feature activations arise from interactions between components, revealing circuit-level dynamics that activation-only approaches cannot identify. Together, these methods increase interpretability robustness and enhance scalable mechanistic analysis of circuits while maintaining efficiency and quality.
Access to: full publication

Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
Authors: Hufe, L., Venhoff, C., Purelku, E., Dreyer, M., Lapuschkin, S., & Samek, W.; Published in: ICLR 2026
Abstract:
Typographic attacks exploit multi-modal systems by injecting text into images, leading to targeted misclassifications, malicious content generation and even Vision-Language Model jailbreaks. In this work, we analyze how CLIP vision encoders behave under typographic attacks, locating specialized attention heads in the latter half of the model's layers that causally extract and transmit typographic information to the cls token. Building on these insights, we introduce Dyslexify - a method to defend CLIP models against typographic attacks by selectively ablating a typographic circuit, consisting of attention heads. Without requiring finetuning, dyslexify improves performance by up to 22.06\% on a typographic variant of ImageNet-100, while reducing standard ImageNet-100 accuracy by less than 1\%, and demonstrate its utility in a medical foundation model for skin lesion diagnosis. Notably, our training-free approach remains competitive with current state-of-the-art typographic defenses that rely on finetuning. To this end, we release a family of dyslexic CLIP models which are significantly more robust against typographic attacks. These models serve as suitable drop-in replacements for a broad range of safety-critical applications, where the risks of text-based manipulation outweigh the utility of text recognition.
Access to: full publication

Better than Average: Spatially-Aware Aggregation of Segmentation Uncertainty Improves Downstream Performance
Authors: Guarino, V. E., Winklmayr, C., Franzen, J., Rumberger, J. L., Pfeuffer, M., Greven, S., Maier-Hein, K., Lüth, C. T., Karg, C., & Kainmüller, D.; Published in: CVPR, 2026
Abstract:
Uncertainty Quantification (UQ) is crucial for ensuring the reliability of automated image segmentations in safety-critical domains like biomedical image analysis or autonomous driving. In segmentation, UQ generates pixel-wise uncertainty scores that must be aggregated into image-level scores for downstream tasks like Out-of-Distribution (OoD) or failure detection. Despite routine use of aggregation strategies, their properties and impact on downstream task performance have not yet been comprehensively studied. Global Average is the default choice, yet it does not account for spatial and structural features of segmentation uncertainty. Alternatives like patch-, class- and threshold-based strategies exist, but lack systematic comparison, leading to inconsistent reporting and unclear best practices. We address this gap by (1) formally analyzing properties, limitations, and pitfalls of common strategies; (2) proposing novel strategies that incorporate spatial uncertainty structure and (3) benchmarking their performance on OoD and failure detection across ten datasets that vary in image geometry and structure. We find that aggregators leveraging spatial structure yield stronger performance in both downstream tasks studied. However, the performance of individual aggregators depends heavily on dataset characteristics, so we (4) propose a meta-aggregator that integrates multiple aggregators and performs robustly across datasets.
Access to: full publication

Extending the Generalised Covariance Measure for Conditional Independence Testing with Repeated Measurements
Authors: Huo, Q., Simnacher, M., Xu, X., & Greven, S.; Published in: IWSM, 2026
Abstract:
The Generalised Covariance Measure (GCM) is a regression-based regression-based test of conditional independence that was originally developed for independent and identically distributed observations. However, directly apply- ing the standard GCM to longitudinal data, where repeated measurements within a subject are correlated, can inflate Type I error by ignoring this within-subject dependence. We propose an extension of the GCM that uses mixed-effects mod- els to capture within-subject correlation and thereby recover valid Type I error control in repeated-measures settings.
Access to: full publication

X-SYS: A Reference Architecture for Interactive Explanation Systems
Authors: Labarta, T., Hoang, N., Dreyer, M., Berend, J., Hein, O., Ma, J., Samek, W., & Lapuschkin, S.; Published in: xAI 2026
Abstract:
The explainable AI (XAI) research community has proposed numerous technical methods, yet deploying explainability as systems remains challenging: Interactive explanation systems require both suitable algorithms and system capabilities that maintain explanation usability across repeated queries, evolving models and data, and governance
constraints. We argue that operationalizing XAI requires treating explainability as an information systems problem where user interaction demands induce specific system requirements. We introduce X-SYS, a reference architecture for interactive explanation systems, that guides (X)AI researchers, developers and practitioners in connecting interactive explanation user interfaces (XUI) with system capabilities. X-SYS organizes around four quality attributes named STAR (scalability, traceability, responsiveness, and adaptability), and specifies a five-component de composition (XUI Services, Explanation Services, Model Services, Data Services, Orchestration and Governance). It maps interaction patterns to system capabilities to decouple user interface evolution from backend computation. We implement X-SYS through SemanticLens, a system for semantic search and activation steering in vision-language models. SemanticLens demonstrates how contract-based service boundaries enable
independent evolution, offline/online separation ensures responsiveness, and persistent state management supports traceability. Together, this work provides a reusable blueprint and concrete instantiation for interactive explanation systems supporting end-to-end design under operational constraints.
Access to: full publication

FADE: Why bad descriptions happen to good features
Authors: Puri, B., Jain, A., Golimblevskaia, E., Kahardipraja, P., Wiegand, T., Samek, W., & Lapuschkin, S.; Published in: (ACL), 17138--17160.
Abstract:
Recent advances in mechanistic interpretability have highlighted the potential of automating interpretability pipelines in analyzing the latent representations within LLMs. While this may enhance our understanding of internal mechanisms, the field lacks standardized evaluation methods for assessing the validity of discovered features. We attempt to bridge this gap by introducing FADE: Feature Alignment to Description Evaluation, a scalable model-agnostic framework for automatically evaluating feature-to-description alignment. FADE evaluates alignment across four key metrics – Clarity, Responsiveness, Purity, and Faithfulness – and systematically quantifies the causes of the misalignment between features and their descriptions. We apply FADE to analyze existing open-source feature descriptions and assess key components of automated interpretability pipelines, aiming to enhance the quality of descriptions. Our findings highlight fundamental challenges in generating feature descriptions, particularly for SAEs compared to MLP neurons, providing insights into the limitations and future directions of automated interpretability. We release FADE as an open-source package at: github.com/brunibrun/FADE.
Access to: full publication

Concept activation vectors: A unifying view and adversarial attacks
Authors: Schnoor, E., Tiomoko, M., Said, J., Jung, A., & Samek, W.; Published in: ICASSP - 2026, 3481--3485
Abstract:
The field of mechanistic interpretability aims to study the role of individual neurons in Deep Neural Networks. Single neurons, however, have the capability to act polysemantically and encode for multiple (unrelated) features, which renders their interpretation difficult. We present a method for disentangling polysemanticity of any Deep Neural Network by decomposing a polysemantic neuron into multiple monosemantic "virtual" neurons. This is achieved by identifying the relevant sub-graph ("circuit") for each "pure" feature. We demonstrate how our approach allows us to find and disentangle various polysemantic units of ResNet models trained on ImageNet. While evaluating feature visualizations using CLIP, our method effectively disentangles representations, improving upon methods based on neuron activations.
Access to: full publication

Deep nonparametric conditional independence tests for images
Authors: Simnacher, M., Xu, X., Park, H., Lippert, C., & Greven, S.; Published in: Journal of Machine Learning Research, 27(96), 1--73.
Abstract:
Conditional independence tests (CITs) test for conditional dependence between random variables given a vector of conditioning or confounder variables. As existing CITs are limited in their applicability to complex, high-dimensional variables such as images, we introduce deep nonparametric CITs (DNCITs). The DNCITs combine embedding maps, which extract feature representations of high-dimensional variables, with nonparametric CITs applicable to these feature representations. For the embedding maps, we derive general properties on their parameter estimators to obtain valid DNCITs and show that these properties include embedding maps learned through (conditional) unsupervised or transfer learning. For the nonparametric CITs, appropriate tests are selected and adapted to be applicable to feature representations. Through simulations, we investigate the performance of the DNCITs for different embedding maps and nonparametric CITs under varying confounder dimensions and confounder relationships. We apply the DNCITs to brain MRI scans and behavioral traits, given confounders, of healthy individuals from the UK Biobank, confirming null results from a number of ambiguous personality neuroscience studies, now with a larger data set and with our more powerful tests. In addition, in a confounder control study, we apply the DNCITs to brain MRI scans and a confounder set to test for sufficient confounder control. We provide an R package implementing the proposed DNCITs.
Access to: full publication

Elastic Full Procrustes Analysis of Plane Curves via Hermitian Covariance Smoothing
Authors: Stöcker, A., Pfeuffer, M., Steyer, L., & Greven, S.; Published in: Journal of Computational and Graphical Statistics (2026)
Abstract:
For shapes of plane curves, the coordinate systems and parametrizations used are often arbitrary and not of interest. In statistical shape analysis, curves are thus frequently considered as equivalence classes of parameterized curves with respect to the shape invariances translation, rotation and scale, as well as reparameterization (warping), based on the square-root-velocity (SRV) framework. We propose a novel elastic full Procrustes mean for samples of plane curve shapes. Identifying the real plane with the complex numbers, we establish a connection to covariance estimation in irregular/sparse functional data analysis. We introduce Hermitian covariance smoothing and employ it for mean estimation, thereby newly covering the sparse case and improving robustness to outlier contamination compared to existing methods. Necessary for our approach but also of independent interest, we characterize the covariance structure of rotation-invariant bivariate stochastic processes via complex representations, and identify sampling schemes that allow for observing derivatives/SRV transforms of sparsely sampled curves. In addition, we develop one- and two-way ANOVA for sparse curve shape data, where exact distance computation is not feasible. We demonstrate the performance of our approach in different realistic simulation settings and use it for an ANOVA of tongue shapes during speech production. Proposed methods are implemented in the R package elastes.
Access to: full publication

Explaining deep learning models predicting antimicrobial resistance for tuberculosis via genomic concepts
Authors: Yanez Sarmiento, P., Klein, N., & Renard, B. Y.; Published in: ECML PKDD 2026 Workshop on eXplainable Artificial Intelligence (XAI).

The case for model science: Verify, explore, steer, refine
Authors: Biecek, P., Longo, L., Zhou, J., Fel, T., Holzinger, A., & Samek, W.; Published in: arXiv:2606.01189.
Abstract:
We argue that the AI community is now ready to move beyond benchmarking and consolidate scattered efforts in model analysis into a systematic discipline, a direction we term Model Science. Complex AI models now serve billions of users, yet our understanding of how they work lags far behind our ability to deploy them. Decades of benchmark-driven research have delivered remarkable progress: extensive leaderboards, a wide range of performance metrics, tracking capability gains across diverse tasks; yet this success has also revealed the limits of benchmarks as they tell us whether models perform but not why they succeed or fail, they miss critical failure modes, such as hallucinations or shortcuts. Precedents from established sciences point the way forward: cognitive science shows that understanding complex systems requires complementary levels of analysis; neuroscience demonstrates that deep study of single cases reveals what population studies miss; medicine teaches that specialised training must develop alongside research practice; and agriculture models how shared infrastructure and principles enable cumulative progress. These lessons inform three foundations for Model Science. First, we propose to consolidate research around four functional perspectives: Verify, Explore, Steer, and Refine that address complementary questions about model behaviour. Second, we discuss the required infrastructure for cumulative knowledge: catalogues of datasets, models and findings. Third, we highlight the need for deep analysis of individual model instances, not just model families, because single cases can reveal what population studies miss.
Access to: full publication

Bayesian meta-learning for modeling Alzheimer\'s disease progression
Authors: Hoffmann, C., & Klein, N.; Published in: arXiv:2606.02228.
Abstract:
Predicting whether an individual with Alzheimer's disease will experience mild or severe disease progression is essential for personalized treatment. Typically, practitioners seek to predict the distribution of a discrete disease score, conditional on an individual's current MRI volume and their historical disease trajectory. Classical statistical regression models and single-task neural networks are not well-suited for this purpose because fitting separate models is infeasible (since each individual typically has few observations), while ignoring individual-level correlation leads to poor generalization. Meta-learning, in contrast, provides a natural avenue to dynamically predict distributions without retraining and model nonlinear relationships between the outcome and covariates. Motivated by this, we propose a Bayesian meta-learner that is trained on multiple individuals but tailors the predictive disease score distribution to each individual's historical data. Our model predicts on unseen individuals without retraining, scales linearly with the number of historical observations, and is guaranteed to be less overconfident when predicting long-term disease scores compared to its deterministic counterpart. On real-world data from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database, our model achieves performance competitive with both single-task models and deterministic meta-learners, while substantially improving performance when predicting long-term disease progression.
Access to: full publication

Towards Visually Explaining Statistical Tests with Applications in Biomedical Imaging
Authors: Javanbakhat, M., Komorowski, P., Bareeva, D., Lai, W.-C., Samek, W., & Lippert, C.; Published in: arXiv:2601.13899.
Abstract:
Deep neural two-sample tests have recently shown strong power for detecting distributional differences between groups, yet their black-box nature limits interpretability and practical adoption in biomedical analysis. Moreover, most existing post-hoc explainability methods rely on class labels, making them unsuitable for label-free statistical testing settings. We propose an explainable deep statistical testing framework that augments deep two-sample tests with sample-level and feature-level explanations, revealing which individual samples and which input features drive statistically significant group differences. Our method highlights which image regions and which individual samples contribute most to the detected group difference, providing spatial and instance-wise insight into the test's decision. Applied to biomedical imaging data, the proposed framework identifies influential samples and highlights anatomically meaningful regions associated with disease-related variation. This work bridges statistical inference and explainable AI, enabling interpretable, label-free population analysis in medical imaging.
Access to: full publication

A Review of Vision Transformer Explainability: Methods, Evaluation, and Future Research Directions
Authors: Koreuber, N., Winklmayr, C., Berend, J., Samek, W., & Kainmüller, D.; Published in: University of Potsdam preprint, doi:10.25932/publishup-70991.
Abstract:
Vision transformers have emerged as a powerful alternative to Convolutional Neural Networks in computer vision, increasingly serving as their successor in many applications. However, their black-box nature hinders their deployment in critical domains, where understanding model decisions is essential. Understanding how models work internally is also crucial for detecting biases and model debugging. This need has led to a growing number of explainability methods developed for or adapted to vision transformers in recent years, making it challenging to gain a clear understanding of the current state of the art. This comprehensive review explores the landscape of vision transformer explainability methods. We provide a structured overview of 81 methods, categorized in a novel fine-grained taxonomy, and illustrate its application through two exemplary use case scenarios for method selection. We further perform a comparative analysis of how these methods are evaluated in terms of models, datasets, baselines, and metrics. We observe a shift from early attention- and gradient-based attribution visualizations toward more holistic techniques, a growing use of concepts and embeddings as explanation targets, a gap in textual explanations of pure vision models, and the lack of integration into interactive explanation systems. Our findings reveal a highly fragmented evaluation landscape, with only partly overlapping use of metrics, over 130 distinct baseline methods, and a scarcity of human-centered and actionable evaluation.
Access to: full publication

Counterfactual Explanations for Deep Two-Sample Testing
Authors: Lai, W.-C., Simnacher, M., & Lippert, C.; Published in: arXiv:2606.04009.
Abstract:
Two-sample testing is a fundamental tool for detecting distributional differences across scientific domains, but classical tests (including kernel-based tests) can be ineffective on high-dimensional structured data such as images. Recent deep two-sample tests improve sensitivity in these settings by learning informative representations, yet they provide limited insight into which data features drive rejection of the null hypothesis
Access to: full publication

Calibration-aware relevance estimation for structured filter pruning
Authors: Mitra, P., Schwalbe, G., Loyal, A., Wirth, C., & Klein, N.; Published in: TMLR Paper 10582
Abstract:
Structured pruning is widely used to reduce the computational and memory requirements of deep neural networks (DNNs), enabling efficient deployment in resource-constrained computer vision systems. However, in safety-critical applications, predictive confidence must also be well calibrated, such that confidence accurately reflects the probability of correctness. Existing structured pruning methods primarily optimize predictive performance or feature importance, while largely overlooking uncertainty calibration (UC) as a pruning objective. To address this gap, we propose, to the best of our knowledge, the first calibration-aware sample selection strategy for relevance-based structured pruning. Instead of estimating filter importance over the entire validation set, the proposed framework identifies calibration-critical spillover samples from reliability diagrams and performs Layer-wise Relevance Propagation only on this informative subset, enabling pruning decisions that explicitly account for predictive reliability. We evaluate the proposed approach on convolutional neural networks (VGG-16, ResNet-34/56, and DenseNet-121) and a vision transformer (DeiT-Tiny) for the CIFAR-10, CIFAR-100, and ImageNet datasets. Experimental results demonstrate that the proposed method improves UC while maintaining competitive predictive accuracy under comparable sparsity levels, with reductions in expected calibration error (ECE) of over 50\% relative to conventional structured pruning in the best-performing settings. Furthermore, the resulting pruned models remain amenable to post-hoc recalibration, reducing calibration error below 2\% ECE in several pruning settings while preserving predictive performance. These findings establish UC not only as an evaluation metric but also as an effective optimization objective for structured model compression, improving the reliability of compressed deep neural networks without sacrificing deployment efficiency.
Access to: full publication

Controlling for Omitted Variable Bias in Deep Neural Networks
Authors: Pfeuffer, M., Rane, R. P., Ritter, K., & Greven, S.; Published in: arXiv:2608.25930.
Abstract:
Control variables are widely used in statistical modelling to account for omitted variable bias of known confounders. However, they have largely been underexplored in deep learning. This is surprising, given that deep learning models encode image-inferable covariates, such as demographic variables, into their predictions when these covariates are correlated with the outcome---a form of omitted variable bias referred to as 'shortcut learning'. While many existing confound-control or fairness methods try to restrict the correlation of such covariates with model predictions, we show that this fails to correct for omitted variable bias. We therefore propose a control variable approach for deep learning models, based on generalised additive modelling of the effects of model inputs and covariates. As flexible additive models can suffer from concurvity, we introduce an estimation procedure that refits the final layer of a pre-trained network to include covariate effects, using cross-fitting with ridge penalisation. We show how these effects can be orthogonalised with respect to covariates to exclude their mediated effects and that model predictions can be marginalised over the covariate distribution to control for their effect. This yields unbiased, interpretable predictions and offers flexibility to model the desired effects depending on the scientific or fairness objective. We verify our approach using simulated images, and demonstrate consistent estimation of true effects. Existing methods either require more data or fail to recover the true effects. We apply our method to real neuroimaging data with experimentally induced confounding, where it recovers prediction performance to near the level of a model trained on unconfounded data. Code is available at this https URL.
Access to: full publication

Deep Shape Regression for Planar Curves with Multimodal Covariates
Authors: Pfeuffer, M., Rane, R. P., Yassin, H., Ritter, K., & Greven, S.; Published in: arXiv:2607.19600.
Abstract:
The shape of a planar curve is the geometric information that remains once translation, rotation, scale and reparametrisation are removed and is of interest in many health applications, e.g. in neuroimaging. We propose a deep shape regression model for open planar curves that admits multimodal and high-dimensional covariates. Representing curves as complex-valued functions, we show that the conditional full Procrustes mean is the leading eigenfunction of the conditional covariance. To estimate this covariance surface, we propose a novel deep conditional covariance smoother with modality-specific encoders - e.g. splines for scalar covariates and convolutional networks for images, which classical spline smoothers cannot accommodate. Our model is by construction invariant to the translation, rotation and scaling of the input curves and handles sparsely and irregularly sampled curves. We further provide an algorithm for elastic mean estimation that also removes parametrisation by iterating covariance smoothing, rotational alignment and parametrisation alignment. We illustrate the method on simulated outlines with known conditional mean and multimodal covariates, and give a first application to hippocampal outlines from the ADNI cohort, recovering covariate effects consistent with the literature. Code is available at this https URL.
Access to: full publication

ICON Decomposition: Multivariate Concept-Level Explanations of Deep Representations for Model Auditing
Authors: Rane, R. P., Simnacher, M., Pfeuffer, M., Schulz, M.-A., Siegel, N. T., Dreyer, M., Pahde, F., Samek, W., Greven, S., & Ritter, K.; Published in: arXiv:2608.26083.
Abstract:
Deep neural networks often exploit spurious associations, a failure known as shortcut learning. Auditing for shortcuts requires testing many candidate concepts, such as acquisition settings or demographics. Current concept-based explainability methods test one concept at a time, asking whether it is decodable from a layer. These methods flag concepts that are merely correlated with the outcome or with each other, and their scores are not comparable across layers or concept types. We introduce ICON decomposition, which quantifies how much of a layer's variance each concept explains given all other concepts and the outcome, and how much none of them explains, yielding layer-comparable, calibrated scores that suppress false positives. On simulated data, ICON recovers concept importance more accurately than seven existing methods. On skin-cancer models with inserted artifacts, it detects induced shortcuts where baselines report false positives. On two brain-imaging models, ICON's sparse explanations are validated by retraining and out-of-distribution tests.
Access to: full publication

The algorithm is not the behavior: Learned priors override look-ahead in a chess-playing neural network
Authors: Sandmann, E., Lapuschkin, S., & Samek, W.; Published in: arXiv:2508.21380.
Abstract:
Recent mechanistic work has uncovered learned algorithms within neural networks, from modular arithmetic to search and planning in game-playing agents. But does algorithmic structure guarantee algorithmic behavior? We investigate this in Leela Chess Zero, the strongest neural chess engine, where prior work identified learned look-ahead. By extending the logit lens to its move-selecting policy network, we discover that correct puzzle solutions-including immediate checkmates-often appear in intermediate layers but are systematically overridden in the final output, a phenomenon we term "forgotten puzzles". Replicating prior analyses on these positions, we find that look-ahead operates normally-future moves of the correct continuation are represented, causally important, and linearly decodable-ruling out a failure of the algorithm itself. Instead, late layers increasingly shift toward prioritizing safe play over aggression. To test whether this shift drives the override, we steer the model against these preferences and recover 61.7% of forgotten puzzles, providing causal evidence that safety priors override algorithmically computed solutions. These findings demonstrate that algorithmic structure does not guarantee algorithmic behavior: a model can internally solve a problem and still output the wrong answer.
Access to: full publication

Uncertainty-Aware Deep Learning for Genomics Applications: Insights from an Empirical Study
Authors: Saran, S., Ghanbari, M., & Ohler, U.; Published in: arXiv:2608.11054.
Abstract:
Deep learning models have emerged as the standard computational tool for a wide range of applications in genomics. Yet, uncertainty quantification (UQ) -- and more specifically, the reliability of different uncertainty estimates in this domain -- has received little systematic attention. This work presents an empirical analysis of UQ in deep learning models, focusing on genomics applications. In a series of experiments, we contrast Deep Ensembles, Bayesian Neural Networks, and Monte Carlo-dropout methods. We assess their ability to quantify uncertainty in different scenarios, accounting for common dataset characteristics in two genomic application areas and modalities: sequence-to-activity models, and single-cell expression analysis. Our systematic comparison framework provides guidelines for the applicability and reliability of UQ methods in genomics, highlighting their strengths and limitations in different scenarios. We show that Bayesian Neural Networks are better at capturing uncertainty caused by strong class imbalance and out-of-distribution data in genomics, despite their computational disadvantages. Moreover, we show how uncertainty scores can be used to select high-quality predictions in protein-RNA interactions.
Access to: full publication

A Unified Risk View of Uncertainty: Posterior Risk for Disentanglement and Evaluation Beyond Proxies
Authors: Wizgall, F., Tirpitz, G., Seiler, M., Ritter, K., & Mucsányi, B.; Published in: arXiv:2608.05995.
Abstract:
Reliable uncertainty estimates are critical in safety-sensitive applications, where understanding the sources of predictive uncertainty is essential. This often requires disentangling epistemic uncertainty from aleatoric uncertainty, yet these uncertainty types are not defined consistently across the literature, making it difficult to assess whether a method produces accurate uncertainty estimates. Evaluation is further complicated by the fact that ground-truth epistemic uncertainty is typically unavailable. Existing benchmarks therefore mostly rely on proxy tasks such as out-of-distribution detection, which do not provide complete ground-truth uncertainty targets and offer limited insight into the structure and quality of uncertainty estimates. We propose a unified definition of uncertainty as pointwise posterior risk, the expected loss of a predictor under the distribution of plausible ground-truth functions given the data. This view combines Bayesian uncertainty over functions with estimator-dependent deviations from the posterior mean, capturing effects such as misspecification and optimization error. This formulation constitutes the foundation of a theory-backed benchmark that enables direct computation of oracle epistemic and aleatoric uncertainty using semi-synthetic datasets with real covariates and known generative processes. By avoiding proxy evaluations, the benchmark enables fine-grained analysis of uncertainty estimates. Empirically, we find that accurate prediction does not guarantee reliable uncertainty disentanglement. The benchmark reveals practically useful differences between methods, identifying approaches with meaningful alignment to oracle uncertainty targets while exposing sensitivity to datasets and modeling choices.
Access to: full publication

Interpreting Protein Language Model Embeddings via Orthogonal Projection for Protein Fitness Prediction
Authors: Yanez Sarmiento, P., Rissom, P. F., Pfeuffer, M., Simnacher, M., Safer, J. F., Iqbal, S., Heyne, H. O., Klein, N., & Renard, B. Y.; Published in: arXiv:2608.25548.
Abstract:
Recently, there has been a growing adoption of protein language models (PLMs) in biomedical science. Their embeddings provide a rich numerical representation of protein sequences which achieve state-of-the-art performance on several downstream tasks including protein fitness prediction. However, PLM embeddings are not directly interpretable and, thereby, it remains unclear what features they encode. To gain insight into which biochemical properties of the protein are driving the prediction, we leverage an orthogonal projection technique that removes linear effects of known tabular features from embeddings and extend it to high-order and interaction effects. In this way, we remove the effects of interpretable biochemical features from PLM embeddings. In an ablation study, we show that this leads to a decrease in performance for a downstream classifier trained only on the embeddings to predict protein fitness. In an additional evaluation, we find that these biochemical features explain a substantial part of the variance in the predictions of this classifier. Hence, we can show that PLM embeddings encode patterns correlated with biochemical properties and quantify their contribution to predicting protein fitness. This computationally efficient approach is not limited to the features or embeddings considered here and is readily transferable to problem settings beyond protein fitness prediction.
Access to: full publication

Unsupervised Semantic Segmentation Facilitates Model Understanding
Authors: Yu, X., Mais, L., Franzen, J., Hirsch, P., Lechtenbörger, N., Mardt, A., & Kainmüller, D.; Published in: arXiv:2605.29691.
Abstract:
Self-supervised learning (SSL) has produced a diverse landscape of vision transformers (ViTs) whose pretrained representations support a wide range of downstream tasks. Towards a better understanding of these models, a body of work has assessed the mechanics of their self-attention as well as the types of information captured across their representations, revealing, for example, stark differences between models trained with contrastive learning (CL) and masked image modeling (MIM). However, the total of these advances on model understanding has to date not yet fully permeated a larger community, where, e.g., insights that are specific to CL models are still at times generalized to MIM models. To make model understanding straightforward and intuitive for a broad community, we propose a simple and easily interpretable visualization protocol.
Our protocol is based on visualizing unsupervised semantic segmentation results, yet by no means do we focus on top segmentation performance. Instead, our protocol allows us to easily convey model behavior that consistently emerges across images. Benchmarked on a diverse set of SSL models across layers and representations, our protocol allows us to gain novel insights into distinct positional biases and scaling behaviors, including, e.g., strong boundary artifacts in DINOv3-Large model tokens. These novel insights come on top of more easily conveying a range of previous findings.
Our protocol further allows us to clearly visually convey and distinguish between positional effects and the closely related but distinct locality bias, the latter being much more extensively studied in the literature so far. Our protocol is publicly available, serving to catalyze further model understanding for a broad community.
Access to: full publication

Model-based Fréchet regression in (quotient) metric spaces with a focus on elastic curves
Authors: Steyer, L., Stöcker, A., Greven, S.; Published in: Journal of Multivariate Analysis (2026)
Abstract:
We introduce model-based Fréchet regression in metric spaces. Instead of starting from point-wise conditional Fréchet means, our approach is defined as a constrained minimization problem over a model class of functions. The approach is then applied to develop a general framework of regression for quotient metric spaces with distances induced by isometric group actions. Such spaces arise naturally in applications where objects are considered equivalent up to transformations. We first establish general existence and consistency results for model-based Fréchet regression, with our quotient space regression model as a special case. As an important example we consider regression for elastic curves in the square-root velocity framework. This addresses data such as handwritten letters, movement paths, or outlines of objects, where only the image but not the parametrization of the curves is of interest. To handle sparsely or irregularly sampled curves, we model smooth conditional mean curves using splines. We validate our approach through simulations and an application to hippocampal outlines extracted from Magnetic Resonance Imaging scans. Here we model how the shape of the irregularly sampled hippocampus is related to age, Alzheimer’s disease and sex, to disentangle the shrinking effects of Alzheimer’s from normal aging.
Access to: full publication

A framework for peptide identification on commercial nanopore sequencing platforms
Authors: Beslic, D., Kucklick, M., Graap, E., Sedaghatjoo, S., Renard, B., Fuchs, S., Engelmann, S., Koerber, N.; Published in: BioRxiv
Abstract:
Direct single-molecule peptide analysis could in principle enable rapid and sensitive identification of pathogen-derived or disease-associated biomarkers without reliance on mass spectrometry. However, existing nanopore peptide sensing methods are typically constrained by limited throughput and lack of accessibility beyond specialized setups.
Here, we present an integrated experimental-computational framework for DNA-linked peptide translocation on a commercially available, high-throughput nanopore sequencing platform, the MinION. Synthetic peptides were covalently bound to oligonucleotides at both termini. The resulting peptide-DNA constructs were then translocated through the CsgG-CsgF pores using a DNA motor protein. Current traces were segmented using the known DNA sequences to extract peptide-associated signal regions. From these segments, we extracted signal features and trained feature-based and deep-learning classifiers to distinguish peptides, balancing interpretability and classification performance.
We establish a framework for peptide identification using standard nanopore sequencing hardware. Across a diverse panel of synthetic peptides, our approach resolves single-amino-acid substitutions, maintains performance across independent sequencing runs, and correctly identifies peptides in blind mixtures. Interpretable model analyses connect classifier decisions and common errors to specific signal motifs. By combining commercially available instrumentation with a reproducible experimental and computational workflow, this framework lowers the barrier to nanopore-based proteomics and enables broader adoption across laboratories. It provides a foundation for future developments in amino acid modification detection and sequence analysis.
Access to: full publication

BaGGLS: A Bayesian Shrinkage Framework for Interpretable Modeling of Interactions in High-Dimensional Biological Data
Authors: Lemanczyk, M., Kock, L., Schlimme, J., Klein, N., Renard, B.; Published in: Proceedings ECCB 2026/Bioinformatics
Abstract:
Biological data sets are often high-dimensional, noisy, and governed by complex interactions among sparse signals. This poses major challenges for interpretability and reliable feature selection. Tasks such as identifying motif interactions in genomics exemplify these difficulties, as only a small subset of biologically relevant features (e.g., motifs) are typically active, and their effects are often non-linear and context-dependent. While statistical approaches often result in more interpretable models, deep learning models have proven effective in modeling complex interactions and prediction accuracy, yet their black-box nature limits interpretability. We introduce BaGGLS, a flexible and interpretable probabilistic binary regression model designed for high-dimensional biological inference involving feature interactions. BaGGLS incorporates a Bayesian group global-local shrinkage prior, aligned with the group structure introduced by interaction terms. This prior encourages sparsity while retaining interpretability, helping to isolate meaningful signals and suppress noise. To enable scalable inference, we employ a partially factorized variational approximation that captures posterior skewness and supports efficient learning even in large feature spaces. In extensive simulations, we can show that BaGGLS outperforms the other methods with regard to interaction detection and is many times faster than MCMC sampling under the horseshoe prior. We also demonstrate the usefulness of BaGGLS in the context of interaction discovery from motif scanner outputs and noisy attribution scores from deep learning models. This shows that BaGGLS is a promising approach for uncovering biologically relevant interaction patterns, with potential applicability across a range of high-dimensional tasks in computational biology.
Access to: full publication

Teaching Mishaps With Mistakes: A Peer-Led Seminar Using Student Case Studies to Enhance Retention
Authors: Yuu, E., Renard, B.; Published in: Teaching Statistics 48, no. 3: 227–236.
Abstract:
Traditional teaching often emphasizes correct methods, limiting opportunities to explore analytical errors and biases. We introduce a seminar framework that integrates peer-to-peer teaching with intentional exposure to statistical and machine learning mishaps through flawed, student-designed case studies. Students delivered two presentations: one teaching a chosen mishap and another presenting an original case study embedding errors. Personalized feedback before and after each presentation supported iterative improvement. This structure created a safe environment to intentionally produce, analyze, and discuss errors, helping students understand how mishaps arise, affect results, and can be addressed. Class discussions and peer analysis encouraged deeper engagement and critical thinking. A follow-up questionnaire administered 1 year later showed that seminar participants significantly outperformed peers in identifying and correcting data analysis errors. Our results suggest that peer-led, error-focused learning with repeated feedback enhances engagement, retention, and preparedness for real-world data challenges, offering a valuable complement to traditional, instructor-led teaching.
Access to: full publication

Evaluating Post-hoc Explanations of the Transformer-based Genome Language Model DNABERT-2
Authors: Kurth, I., Yanez Sarmiento, P., Renard, B.; Published in: Proceedings of World Conference on Explainable Artificial Intelligence 2026
Abstract:
Explaining deep neural network predictions on genome sequences enables biological insight and hypothesis generation-often of greater interest than predictive performance alone. While explanations of convolutional neural networks (CNNs) have been shown to capture relevant patterns in genome sequences, it is unclear whether this transfers to more expressive Transformer-based genome language models (gLMs). To answer this question, we adapt AttnLRP, an extension of layer-wise relevance propagation to the attention mechanism, and apply it to the state-of-the-art gLM DNABERT-2. Thereby, we propose strategies to transfer explanations from token and nucleotide level. We evaluate the adaption of AttnLRP on genomic datasets using multiple metrics. Further, we provide an extensive comparison between the explanations of DNABERT-2 and a baseline CNN. Our results demonstrate that AttnLRP yields reliable explanations corresponding to known biological patterns. Hence, like CNNs, gLMs can also help derive biological insights. This work contributes to the explainability of gLMs and addresses the comparability of relevance attributions across different architectures.
Access to: full publication

Deterministic access to global viral sequence data enables robust agentic scientific discovery
Authors: Nasri, F., Gurev, S., Varilly, P., Ramesh, K., O'Leary, N., Cool, J., Renard, B., Sabeti, P., Luebbert, L.; Published in: arXiv preprint (under review at Nature Machine Intelligence)
Abstract:
Public viral genome resources such as the National Center for Biotechnology Information (NCBI) Virus database are central to outbreak response, evolutionary analysis, vaccine design, and genomic surveillance. Yet many high-value retrieval workflows remain optimized for interactive use rather than deterministic, reproducible programmatic interfaces. This creates a challenge for Large Language Model (LLM)-based scientific agents, where errors in metadata interpretation, filtering logic, or retrieval can propagate into incorrect datasets. To evaluate agentic viral data retrieval, we built VirBench, a manually curated benchmark of 120 queries spanning diverse pathogens, taxonomic levels, and metadata filters. When autonomous AI systems, including Biomni, Claude, GPT, and Edison Analysis, were tasked with these queries without a dedicated retrieval layer, performance varied widely: mean accuracy ranged from 16.9% for Claude Sonnet 4 to 91.3% for GPT-5.5, with newer frontier models showing progress but residual errors remaining consequential. To address this, we built gget virus, a deterministic query framework that formalizes NCBI Virus-style filtering as a reproducible programmatic system. By staging retrieval, applying metadata constraints before sequence download, and retrieving structured GenBank records, gget virus reduces data transfer by more than 98% for high-volume queries while preserving exact-match semantics. Instructing autonomous AI systems to use gget virus increased accuracy to at least 90.0% across all evaluated systems and up to 99.7% for GPT-5.5, improved response stability to 0.92-1.00, reduced error magnitude, and generally decreased runtime and tool calls. Together, this work establishes deterministic data access as critical infrastructure for reliable agentic science and provides a reproducible retrieval layer for robust human- and AI-driven viral genomics workflows.
Access to: full publication

Motif Interactions Affect Post-Hoc Interpretability of Genomic Convolutional Neural Networks
Authors: Lemanczyk, M. S., Bartoszewicz, J. M., and Renard, B. Y. ; Published in:
Abstract:
Post-hoc interpretability methods are commonly used to understand decisions of genomic deep learning models and reveal new biological insights. However, interactions between sequence regions (e.g. regulatory elements) impact the learning process as well as interpretability methods that are sensitive to dependencies between features. Since deep learning models learn correlations between the data and output that do not necessarily represent a causal relationship, it is difficult to say how well interacting motif sets are fully captured. Here, we investigate how genomic motif interactions influence model learning and interpretability methods by formalizing possible scenarios where interaction effects appear. This includes the choice of negative data and non-additive effects on the outcome. We generate synthetic data containing interactions for those scenarios and evaluate how they affect the performance of motif detection. We show that post-hoc interpretability methods can miss motifs if interactions are present depending on how negative data is defined. Furthermore, we observe differences in interpretability between additive and non-additive effects as well as between post-hoc interpretability methods.
Access to: full publication

Decoding protein language models: insights from embedding space analysis
Authors: Rissom, P. F., Sarmiento, P. Y., Safer, J., Coley, C. W., Renard, B. Y., Heyne, H. O., & Iqbal, S.; Published in: biorxiv (conditionally accepted at Cell Patterns)
Abstract:
Foundation models, which encode patterns in large, high-dimensional data as embeddings, show promise in many machine learning related applications in molecular biology. Embeddings learned by the models provide informative features for downstream prediction tasks, however, the information captured by the model is often not interpretable. One approach to understanding the captured information is through the analysis of their learned embeddings, which in molecular biology so far has mainly focused on visualizing individual embedding spaces. This study introduces a quantitative framework for cross-space comparison, enabling intuitive exploration and comparison of embedding spaces in molecular biology. The framework emphasizes analyzing the distribution of known biological information within embedding space neighborhoods and provides insights into relationships between multiple embedding spaces. Comparison techniques include global pairwise distance measurements as well as local nearest neighbor analyses. By applying our framework to embeddings from protein language models, we demonstrate how embedding space analysis can serve as a valuable pre-filtering step for task-specific supervised machine learning applications and for the recognition of differential patterns in data encoded within and across different embedding spaces. To support a wide usability, we provide a Python library that implements all analysis methods, available at https://github.com/broadinstitute/EmmaEmb.
Access to: full publication

Embedded Conditional Independence Tests for Large Language Model Generated Text with an Application to German Parliament Speeches
Authors: Simnacher, M., Keilbar, G.,König, B., Lippert, Ch., Greven, S.; Published in: arXiv:2609.00946
Abstract:
Conditional independence tests (CITs) test for conditional dependence between two random objects
Access to: full publication
2025

Benchmarking uncertainty and its disentanglement in multi-label chest X-ray classification
Authors: Baur, S., Samek, W., & Ma, J.; Published in: MICCAI\'25 UNSURE Workshop, 193--203. Springer.
Abstract:
Reliable uncertainty quantification is crucial for trustworthy decision-making and the deployment of AI models in medical imaging. While prior work has explored the ability of neural networks to quantify predictive, epistemic, and aleatoric uncertainties using an information-theoretical approach in synthetic or well-defined data settings like natural image classification, its applicability to real-life medical diagnosis tasks remains underexplored. In this study, we provide an extensive uncertainty quantification benchmark for multi-label chest X-ray classification using the MIMIC-CXR-JPG dataset. We evaluate 13 uncertainty quantification methods for convolutional (ResNet) and transformer-based (Vision Transformer) architectures across a wide range of tasks. Additionally, we extend Evidential Deep Learning, HetClass NNs, and Deep Deterministic Uncertainty to the multi-label setting. Our analysis provides insights into uncertainty estimation effectiveness and the ability to disentangle epistemic and aleatoric uncertainties, revealing method- and architecture-specific strengths and limitations.
Access to: full publication

Mechanistic understanding and validation of large AI models with SemanticLens
Authors: Dreyer, M., Berend, J., Labarta, T., Vielhaben, J., Wiegand, T., Lapuschkin, S., & Samek, W.; Published in: Nature Machine Intelligence, 7, 1572--1585.
Abstract:
Unlike human-engineered systems, such as aeroplanes, for which the role and dependencies of each component are well understood, the inner workings of artificial intelligence models remain largely opaque, which hinders verifiability and undermines trust. Current approaches to neural network interpretability, including input attribution methods, probe-based analysis and activation visualization techniques, typically provide limited insights about the role of individual components or require extensive manual interpretation that cannot scale with model complexity. This paper introduces SemanticLens, a universal explanation method for neural networks that maps hidden knowledge encoded by components (for example, individual neurons) into the semantically structured, multimodal space of a foundation model such as CLIP. In this space, unique operations become possible, including (1) textual searches to identify neurons encoding specific concepts, (2) systematic analysis and comparison of model representations, (3) automated labelling of neurons and explanation of their functional roles, and (4) audits to validate decision-making against requirements. Fully scalable and operating without human input, SemanticLens is shown to be effective for debugging and validation, summarizing model knowledge, aligning reasoning with expectations (for example, adherence to the ABCDE rule in melanoma classification) and detecting components tied to spurious correlations and their associated training data. By enabling component-level understanding and validation, the proposed approach helps mitigate the opacity that limits confidence in artificial intelligence systems compared to traditional engineered systems, enabling more reliable deployment in critical applications.
Access to: full publication

WeightLens: Input-independent interpretability for LLM transcoders
Authors: Golimblevskaia, E., Puri, B., Jain, A., Samek, W., & Lapuschkin, S.; Published in: Github
Abstract:
Automated Interpretability for Large Language Models Using Weight-Based Feature Analysis
WeightLens provides a low-cost framework for analyzing the meaning of LLM features using input-invariant components (weights) instead of repeated activation probing. By leveraging transcoders, it reduces reliance on external LLM explainers while producing meaningful interpretations for token-level features.
Access to: Access publication

Evaluating Interpretable Methods via Geometric Alignment of Functional Distortions: A Unifying and Geometric Perspective on Interpretability Evaluation
Authors: Hedström, A., Bommer, P. L., Burns, T. F., Lapuschkin, S., Samek, W., & Höhne, M. M.-C.; Published in: Transactions on Machine Learning Research, 1--48.
Abstract:
Interpretability researchers face a universal question: without access to ground truth labels, how can the faithfulness of an explanation to its model be determined? Despite immense efforts to develop new evaluation methods, current approaches remain in a pre-paradigmatic state: fragmented, difficult to calibrate, and lacking cohesive theoretical grounding. Observ- ing the lack of a unifying theory, we propose a novel evaluative criterion entitled Generalised Explanation Faithfulness (GEF) which is centered on explanation-to-model alignment, and integrates existing perturbation-based evaluations to eliminate the need for singular, task-specific evaluations. Complementing this unifying perspective, from a geometric point of view, we reveal a prevalent yet critical oversight in current evaluation practice: the failure to account for the learned geometry, and non-linear mapping present in the model, and explanation spaces. To solve this, we propose a general-purpose, threshold-free faithfulness evaluator GEF that incorporates principles from differential geometry, and facilitates evaluation agnostically across tasks, and interpretability approaches. Through extensive cross-domain benchmarks on natural language processing, vision, and tabular tasks, we provide first-of-its-kind insights into the comparative performance of various interpretable methods. This includes local linear approximators, global feature visualisation methods, large language models as post-hoc explainers, and sparse autoencoders. Our contributions are important to the interpretability and AI safety communities, offering a principled, unified approach for evaluation.
Access to: full publication

The atlas of in-context learning: How attention heads shape in-context retrieval augmentation
Authors: Kahardipraja, P., Achtibat, R., Wiegand, T., Samek, W., & Lapuschkin, S.; Published in: Advances in Neural Information Processing Systems, 38, 118164--118208.
Abstract:
Large language models are able to exploit in-context learning to access external knowledge beyond their training data through retrieval-augmentation. While promising, its inner workings remain unclear. In this work, we shed light on the mechanism of in-context retrieval augmentation for question answering by viewing a prompt as a composition of informational components. We propose an attribution-based method to identify specialized attention heads, revealing in-context heads that comprehend instructions and retrieve relevant contextual information, and parametric heads that store entities' relational knowledge. To better understand their roles, we extract function vectors and modify their attention weights to show how they can influence the answer generation process. Finally, we leverage the gained insights to trace the sources of knowledge used during inference, paving the way towards more safe and transparent language models.
Access to: full publication

PathoCellBench: A Comprehensive Benchmark for Cell Phenotyping
Authors: Lüscher, J., Koreuber, N., Franzen, J., Reith, F. H., Winklmayr, C., Baumann, E., Schürch, C. M., Kainmüller, D., & Rumberger, J. L.; Published in: MICCAI 2025, 411--420. Springer Nature Switzerland.
Abstract:
Digital pathology has seen the advent of a wealth of foundational models (FMs), yet to date their performance on cell phenotyping has not been benchmarked in a unified manner. We therefore propose PathoCellBench: A comprehensive benchmark for cell phenotyping on Hematoxylin and Eosin (H&E) stained histopathology images. We provide both PathoCell , a new H&E dataset featuring 14 cell types identified via multiplexed imaging, and ready-to-use fine-tuning and benchmarking code that allows the systematic evaluation of multiple prominent pathology FMs in terms of dense cell phenotype predictions in a range of generalization scenarios. We perform extensive benchmarking of existing FMs, providing insights into their generalization behavior under technical vs. medical domain shifts. Furthermore, while FMs achieve macro F1 scores > 0.70 on previously established benchmarks such as Lizard and PanNuke, on PathoCell , we observe scores as low as 0.20. This indicates a much more challenging task not captured by previous benchmarks, establishing PathoCell as a prime asset for future benchmarking of FMs and supervised models alike. Code and data are available on GitHub.
Access to: full publication

Ensuring medical AI safety: Explainable AI-driven detection and mitigation of spurious model behavior and associated data
Authors: Pahde, F., Wiegand, T., Lapuschkin, S., & Samek, W.; Published in: Machine Learning, 114(9), 206.
Abstract:
Deep neural networks are increasingly employed in high-stakes medical applications, despite their tendency for shortcut learning in the presence of spurious correlations, which can have potentially fatal consequences in practice. Whereas a multitude of works address either the detection or mitigation of such shortcut behavior in isolation, the Reveal2Revise approach provides a comprehensive bias mitigation framework combining these steps. However, effectively addressing these biases often requires substantial labeling efforts from domain experts. In this work, we review the steps of the Reveal2Revise framework and enhance it with semi-automated interpretability-based bias annotation capabilities. This includes methods for the sample- and feature-level bias annotation, providing valuable information for bias mitigation methods to unlearn the undesired shortcut behavior. We show the applicability of the framework using four medical datasets across two modalities, featuring controlled and real-world spurious correlations caused by data artifacts. We successfully identify and mitigate these biases in VGG16, ResNet50, and contemporary Vision Transformer models, ultimately increasing their robustness and applicability for real-world medical tasks. Our code is available at https://github.com/frederikpahde/medical-ai-safety.
Access to: full publication

PD-L1 expression assessment in angiosarcoma improves with artificial intelligence support
Authors: Reith, F. H., Jarosch, A., Albrecht, J. P., Ghoreschi, F., Flörcken, A., Dörr, A., Roohani, S., Schäfer, F. M., Öllinger, R., Märdian, S., Tielking, K., Bischoff, P., Frühauf, N., Brandes, F., Horst, D., Sers, C., & Kainmüller, D.; Published in: Journal of Pathology Informatics, 18, 100447.
Abstract:
Access to: full publication

Automated classification of cellular expression in multiplexed imaging data with Nimbus
Authors: Rumberger, J. L., Greenwald, N. F., Ranek, J. S., Boonrat, P., Walker, C., Franzen, J., Varra, S. R., Kong, A., Sowers, C., Liu, C. C., et al.; Published in: Nature Methods, 22(10), 2161--2170.
Abstract:
Multiplexed imaging offers a powerful approach to characterize the spatial topography of tissues in both health and disease. To analyze such data, the specific combination of markers that are present in each cell must be enumerated to enable accurate phenotyping, a process that often relies on unsupervised clustering. We constructed the Pan-Multiplex (Pan-M) dataset containing 197 million distinct annotations of marker expression across 15 different cell types. We used Pan-M to create Nimbus, a deep learning model to predict marker positivity from multiplexed image data. Nimbus is a pretrained model that uses the underlying images to classify marker expression of individual cells as positive or negative across distinct cell types, from different tissues, acquired using different microscope platforms, without requiring any retraining. We demonstrate that Nimbus predictions capture the underlying staining patterns of the full diversity of markers present in Pan-M, and that Nimbus matches or exceeds the accuracy of previous approaches that must be retrained on each dataset. We then show how Nimbus predictions can be integrated with downstream clustering algorithms to robustly identify cell subtypes in image data. We have open-sourced Nimbus and Pan-M to enable community use at https://github.com/angelolab/Nimbus-Inference.
Access to: full publication

Pioneering new paths: the role of generative modelling in neurological disease research
Authors: Seiler, M., & Ritter, K.; Published in: Pflügers Archiv -- European Journal of Physiology, 477(4), 571--589.
Abstract:
Recently, deep generative modelling has become an increasingly powerful tool with seminal work in a myriad of disciplines. This powerful modelling approach is supposed to not only have the potential to solve current problems in the medical field but also to enable personalised precision medicine and revolutionise healthcare through applications such as digital twins of patients. Here, the core concepts of generative modelling and popular modelling approaches are first introduced to consider the potential based on methodological concepts for the generation of synthetic data and the ability to learn a representation of observed data. These potentials will be reviewed using current applications in neuroimaging for data synthesis and disease decomposition in Alzheimer’s disease and multiple sclerosis. Finally, challenges for further research and applications will be discussed, including computational and data requirements, model evaluation, and potential privacy risks.
Access to: full publication

Exploring the clinical value of concept-based AI explanations in gastrointestinal disease detection
Authors: Storas, A. M., Dreyer, M., Pahde, F., Lapuschkin, S., Samek, W., et al.Storas, A. M., Dreyer, M., Pahde, F., Lapuschkin, S., Samek, W., et al.; Published in: Scientific Reports, 15(1), 28860.
Abstract:
Complex artificial intelligence models, like deep neural networks, have shown exceptional capabilities to detect early-stage polyps and tumors in the gastrointestinal tract. These technologies are already beginning to assist gastroenterologists in the endoscopy suite. To understand how these complex models work and their limitations, model explanations can be useful. Moreover, medical doctors specialized in gastroenterology can provide valuable feedback on the model explanations. This study explores three different explainable artificial intelligence methods for explaining a deep neural network detecting gastrointestinal abnormalities. The model explanations are presented to gastroenterologists. Furthermore, the clinical applicability of the explanation methods from the healthcare personnel’s perspective is discussed. Our findings indicate that the explanation methods are not meeting the requirements for clinical use, but that they can provide valuable information to researchers and model developers. Higher quality datasets and careful considerations regarding how the explanations are presented might lead to solutions that are more welcome in the clinic.
Access to: full publication

Beyond scalars: Concept-based alignment analysis in vision transformers
Authors: Vielhaben, J., Bareeva, D., Berend, J., Samek, W., & Strodthoff, N.; Published in: Advances in Neural Information Processing Systems, 38, 71379--71414.
Abstract:
Measuring the alignment between representations lets us understand similarities between the feature spaces of different models, such as Vision Transformers trained under diverse paradigms. However, traditional measures for representational alignment yield only scalar values that obscure how these spaces agree in terms of learned features. To address this, we combine alignment analysis with concept discovery, allowing a fine-grained breakdown of alignment into individual concepts. This approach reveals both universal concepts across models and each representation’s internal concept structure. We introduce a new definition of concepts as non-linear manifolds, hypothesizing they better capture the geometry of the feature space. A sanity check demonstrates the advantage of this manifold-based definition over linear baselines for concept-based alignment. Finally, our alignment analysis of four different ViTs shows that increased supervision tends to reduce semantic organization in learned representations.
Access to: full publication

Sparse, Efficient and Explainable Data Attribution with DualXDA
Authors: Yolcu, G. Ü., Weckbecker, M., Wiegand, T., Samek, W., & Lapuschkin, S.; Published in: Transactions on Machine Learning Research.
Abstract:
Data Attribution (DA) is an emerging approach in the field of eXplainable Artificial Intelligence (XAI), aiming to identify influential training datapoints which determine model outputs. It seeks to provide transparency about the model and individual predictions, e.g. for model debugging, identifying data-related causes of suboptimal performance. However, existing DA approaches suffer from prohibitively high computational costs and memory demands when applied to even medium-scale datasets and models, forcing practitioners to resort to approximations that may fail to capture the true inference process of the underlying model. Additionally, current attribution methods exhibit low sparsity, resulting in non-negligible attribution scores across a high number of training examples, hindering the discovery of decisive patterns in the data. In this work, we introduce DualXDA, a framework for sparse, efficient and explainable DA, comprised of two interlinked approaches, Dual Data Attribution (DualDA) and eXplainable Data Attribution (XDA): With DualDA, we propose a novel approach for efficient and effective DA, leveraging Support Vector Machine theory to provide fast and naturally sparse data attributions for AI predictions. In extensive quantitative analyses, we demonstrate that DualDA achieves high attribution quality, excels at solving a series of evaluated downstream tasks, while at the same time improving explanation time by a factor of up to 4,100,000 x compared to the original Influence Functions method, and up to 11,000 x compared to the method's most efficient approximation from literature to date. We further introduce XDA, a method for enhancing Data Attribution with capabilities from feature attribution methods to explain why training samples are relevant for the prediction of a test sample in terms of impactful features, which we showcase and verify qualitatively in detail. Taken together, our contributions in DualXDA ultimately point towards a future of eXplainable AI applied at unprecedented scale, enabling transparent, efficient and novel analysis of even the largest neural architectures -- such as Large Language Models -- and fostering a new generation of interpretable and accountable AI systems. The implementation of our methods, as well as the full experimental protocol, is available on github.
Access to: full publication

From what to how: Attributing CLIP\'s latent components reveals unexpected semantic reliance
Authors: Dreyer, M., Hufe, L., Berend, J., Wiegand, T., Lapuschkin, S., & Samek, W.; Published in: arXiv:2505.20229.
Abstract:
Transformer-based CLIP models are widely used for text-image probing and feature extraction, making it relevant to understand the internal mechanisms behind their predictions. While recent works show that Sparse Autoencoders (SAEs) yield interpretable latent components, they focus on what these encode and miss how they drive predictions. We introduce a scalable framework that reveals what latent components activate for, how they align with expected semantics, and how important they are to predictions. To achieve this, we adapt attribution patching for instance-wise component attributions in CLIP and highlight key faithfulness limitations of the widely used Logit Lens technique. By combining attributions with semantic alignment scores, we can automatically uncover reliance on components that encode semantically unexpected or spurious concepts. Applied across multiple CLIP variants, our method uncovers hundreds of surprising components linked to polysemous words, compound nouns, visual typography and dataset artifacts. While text embeddings remain prone to semantic ambiguity, they are more robust to spurious correlations compared to linear classifiers trained on image embeddings. A case study on skin lesion detection highlights how such classifiers can amplify hidden shortcuts, underscoring the need for holistic, mechanistic interpretability. We provide code at https://github.com/maxdreyer/attributing-clip.
Access to: full publication

Relevance-driven Input Dropout: An Explanation-guided Regularization Technique
Authors: Gururaj, S., Grüne, L., Samek, W., Lapuschkin, S., & Weber, L.; Published in: arXiv:2505.21595.
Abstract:
Explainable AI has brought transparency into complex ML blackboxes, enabling, in particular, to identify which features these models use Overfitting is a well-known issue extending even to state-of-the-art (SOTA) Machine Learning (ML) models, resulting in reduced generalization, and a significant train-test performance gap. Mitigation measures include a combination of dropout, data augmentation, weight decay, and other regularization techniques. Among the various data augmentation strategies, occlusion is a prominent technique that typically focuses on randomly masking regions of the input during training. Most of the existing literature emphasizes randomness in selecting and modifying the input features instead of regions that strongly influence model decisions. We propose Relevance-driven Input Dropout (RelDrop), a novel data augmentation method which selectively occludes the most relevant regions of the input, nudging the model to use other important features in the prediction process, thus improving model generalization through informed regularization. We further conduct qualitative and quantitative analyses to study how Relevance-driven Input Dropout (RelDrop) affects model decision-making. Through a series of experiments on benchmark datasets, we demonstrate that our approach improves robustness towards occlusion, results in models utilizing more features within the region of interest, and boosts inference time generalization performance. Our code is available at this https URL.
Access to: full publication

Attribution-Guided Pruning for Compression, Circuit Discovery, and Targeted Correction in LLMs
Authors: Hatefi, S. M. V., Dreyer, M., Achtibat, R., Kahardipraja, P., Wiegand, T., Samek, W., & Lapuschkin, S.; Published in: arXiv:2506.13727.
Abstract:
Large Language Models (LLMs) are widely deployed in real-world applications, yet their internal mechanisms remain difficult to interpret and control, limiting our ability to diagnose and correct undesirable behaviors. Mechanistic interpretability addresses this challenge by identifying circuits -- subsets of model components responsible for specific behaviors. However, discovering such circuits in LLMs remains difficult due to their scale and complexity. We frame circuit discovery as identifying parameters that contribute most to model outputs on task-specific inputs, and use Layer-wise Relevance Propagation (LRP) with reference samples to attribute and extract these components via pruning. Building on this, we introduce contrastive relevance to isolate circuits associated with undesired behaviors while preserving general capabilities, enabling targeted model correction. On OPT-125M, we show that pruning as little as ~0.3% of neurons substantially reduces toxic outputs, while pruning approximately 0.03% of weight elements mitigates repetitive text generation without degrading general performance. These results establish attribution-guided pruning as an effective mechanism for identifying and intervening on behavior-specific circuits in LLMs. We further validate our findings on additional small-scale language models, demonstrating that the proposed approach transfers across architectures. Our code is publicly available at this https URL.
Access to: full publication

Guided active learning for medical image segmentation
Authors: Föllmer, B., Serafimoski, V., Schulze, K., Biavati, F., Stober, S., Samek, W., & Dewey, M.; Published in: In Human-AI Collaboration, 58--67. Springer Nature Switzerland.
Abstract:Active learning has the potential to reduce labeling costs in medical image segmentation by selecting only the most informative samples. However, conventional approaches typically rely on model-based informativeness measures, limiting the expert’s role to passively annotating pre-selected images. This restricts expert-driven prioritization of segmentation targets aligned with clinical objectives. To address this limitation, we propose GALMIS (Guided Active Learning for Medical Image Segmentation), a novel framework that integrates expert-driven guidance into the informative sample selection process. By leveraging submodular subset selection, GALMIS ensures that selected samples are not only informative but also clinically relevant to predefined segmentation targets. We evaluate our approach in both simulated and real active learning scenarios on: (1) foreground-foreground class imbalance in abdominal CT, and (2) clinical targets for coronary artery segmentation in cardiac CT. Our results demonstrate improved labeling efficiency on clinically relevant targets compared to conventional active learning methods. Code is available at https://github.com/Berni1557/TAL.
Access to: full publication

Multimodal Outcomes in N-of-1 Trials: Combining Unsupervised Learning and Statistical Inference
Authors: Schneider, J., Gärtner, T., & Konigorski, S.; Published in: arXiv:2309.06455v3.
Abstract:
N-of-1 trials are within-person crossover trials allowing both personalized and population-level inference on the effect of health interventions. Using the full potential of modern technologies, multimodal N-of-1 trials can integrate multimedia data for measuring health outcomes. However, methodology required for automated applications in large multimodal trials is not available yet. Here, we present an unsupervised approach for modeling multimodal N-of-1 trials, bypassing the need for expensive outcome labeling by medical experts. First, an autoencoder is trained on the outcome medical images. Then, the dimensionality of embeddings is reduced by extracting the first principal component, which is finally tested for its association with the treatment. Results from imaging simulation studies show high power in detecting a treatment effect while controlling type I error rates. An application to imaging N-of-1 trials of acne severity identifies individual treatment effects and supports that our methodology can enable large clinical multimodal N-of-1 trials.
Access to: full publication

Boosting Causal Additive Models
Authors: Kertel, M., Klein, N.; Published in: Journal of Machine Learning Research 26(169):1−49, 2025.
Abstract:
We present a boosting-based method to learn additive Structural Equation Models (SEMs) from observational data, with a focus on the theoretical aspects of determining the causal order among variables. We introduce a family of score functions based on arbitrary regression techniques, for which we establish sufficient conditions that guarantee consistent identification of the true causal ordering. Our analysis reveals that boosting with early stopping meets these criteria and thus offers a consistent score function for causal orderings. To address the challenges posed by high-dimensional data sets, we adapt our approach through a component-wise gradient descent in the space of additive SEMs. Our simulation study supports the theoretical findings in low-dimensional settings and demonstrates that our high-dimensional adaptation is competitive with state-of-the-art methods. In addition, it exhibits robustness with respect to the choice of hyperparameters, thereby simplifying the tuning process.
Access to: full publication

Explaining predictive uncertainty by exposing second-order effects
Authors: Bley, F.,Lapuschkin, S.,Samek, W.,Montavon, G.; Published in: Pattern Recognition, 160, 111171
Abstract:
Explainable AI has brought transparency to complex ML black boxes, enabling us, in particular, to identify which features these models use to make predictions. So far, the question of how to explain predictive uncertainty, i.e., why a model ‘doubts’, has been scarcely studied. Our investigation reveals that predictive uncertainty is dominated by second-order effects, involving single features or product interactions between them. We contribute a new method for explaining predictive uncertainty based on these second-order effects. Computationally, our method reduces to a simple covariance computation over a collection of first-order explanations. Our method is generally applicable, allowing for turning common attribution techniques (LRP, GradientInput, etc.) into powerful second-order uncertainty explainers, which we call CovLRP, CovGI, etc. The accuracy of the explanations our method produces is demonstrated through systematic quantitative evaluations, and the overall usefulness of our method is demonstrated through two practical showcases.
Access to: full publication
2024

AttnLRP: Attention-aware layer-wise relevance propagation for transformers
Authors: Achtibat, R., Hatefi, S. M. V., Dreyer, M., Jain, A., Wiegand, T., Lapuschkin, S., & Samek, W.; Published in: Proceedings of Machine Learning Research, 235, 135--168.
Abstract:
Large Language Models are prone to biased predictions and hallucinations, underlining the paramount importance of understanding their model-internal reasoning process. However, achieving faithful attributions for the entirety of a black-box transformer model and maintaining computational efficiency is an unsolved challenge. By extending the Layer-wise Relevance Propagation attribution method to handle attention layers, we address these challenges effectively. While partial solutions exist, our method is the first to faithfully and holistically attribute not only input but also latent representations of transformer models with the computational efficiency similar to a single backward pass. Through extensive evaluations against existing methods on LLaMa 2, Mixtral 8×7b, Flan-T5 and vision transformer architectures, we demonstrate that our proposed approach surpasses alternative methods in terms of faithfulness and enables the understanding of latent representations, opening up the door for concept-based explanations. We provide an LRP library at https://github.com/rachtibat/LRP-eXplains-Transformers.
Access to: full publication

Quanda: An Interpretability Toolkit for Training Data Attribution Evaluation and Beyond
Authors: Bareeva, D., Yolcu, G. Ü., Hedström, A., Schmolenski, N., Wiegand, T., Samek, W., & Lapuschkin, S.; Published in: arXiv:2410.07158
Abstract:
In recent years, training data attribution (TDA) methods have emerged as a promising direction for the interpretability of neural networks. While research around TDA is thriving, limited effort has been dedicated to the evaluation of attributions. Similar to the development of evaluation metrics for traditional feature attribution approaches, several standalone metrics have been proposed to evaluate the quality of TDA methods across various contexts. However, the lack of a unified framework that allows for systematic comparison limits trust in TDA methods and stunts their widespread adoption. To address this research gap, we introduce Quanda, a Python toolkit designed to facilitate the evaluation of TDA methods. Beyond offering a comprehensive set of evaluation metrics, Quanda provides a uniform interface for seamless integration with existing TDA implementations across different repositories, thus enabling systematic benchmarking. The toolkit is user-friendly, thoroughly tested, well-documented, and available as an open-source library on PyPi and under this https URL.
Access to: full publication

Pruning by explaining revisited: Optimizing attribution methods to prune CNNs and transformers
Authors: Hatefi, S. M. V., Dreyer, M., Achtibat, R., Wiegand, T., Samek, W., & Lapuschkin, S.; Published in: ECCV 2024 Workshops, Lecture Notes in Computer Science, 15643, 152--169. Springer Nature Switzerland.
Abstract:
To solve ever more complex problems, Deep Neural Networks are scaled to billions of parameters, leading to huge computational costs. An effective approach to reduce computational requirements and increase efficiency is to prune unnecessary components of these often over-parameterized networks. Previous work has shown that attribution methods from the field of eXplainable AI serve as effective means to extract and prune the least relevant network components in a few-shot fashion. We extend the current state by proposing to explicitly optimize hyperparameters of attribution methods for the task of pruning, and further include transformer-based networks in our analysis. Our approach yields higher model compression rates of large transformer and convolutional architectures (VGG, ResNet, ViT) compared to previous works, while still attaining high performance on ImageNet classification tasks. Here, our experiments indicate that transformers have a higher degree of over-parameterization compared to convolutional neural networks. Code is available at https://github.com/erfanhatefi/Pruning-by-eXplaining-in-PyTorch.
Access to: full publication

Evaluation of machine learning-based classification of clinical impairment and prediction of clinical worsening in multiple sclerosis
Authors: Noteboom, S., Seiler, M., Chien, C., Rane, R. P., Barkhof, F., Strijbis, E. M., Paul, F., Schoonheim, M. M., & Ritter, K.; Published in: Journal of Neurology, 271(8), 5577--5589.
Sparse Explanations of Neural Networks Using Pruned Layer-Wise Relevance Propagation
Authors:
Paulo Yanez Sarmiento
Simon Witzke
Nadja Klein
Bernhard Y. Renard
Published in:
Machine Learning and Knowledge Discovery in Databases. Research Track. ECML PKDD 2024. Lecture Notes in Computer Science
Abstract:
Explainability is a key component in many applications involving deep neural networks (DNNs). However, current explanation methods for DNNs commonly leave it to the human observer to distinguish relevant explanations from spurious noise. This is not feasible anymore when going from easily human-accessible data such as images to more complex data such as genome sequences. To facilitate the accessibility of DNN outputs from such complex data and to increase explainability, we present a modification of the widely used explanation method layer-wise relevance propagation. Our approach enforces sparsity directly by pruning the relevance propagation for the different layers. Thereby, we achieve sparser relevance attributions for the input features as well as for the intermediate layers. As the relevance propagation is input-specific, we aim to prune the relevance propagation rather than the underlying model architecture. This allows to prune different neurons for different inputs and hence, might be more appropriate to the local nature of explanation methods. To demonstrate the efficacy of our method, we evaluate it on two types of data: images and genome sequences. We show that our modification indeed leads to noise reduction and concentrates relevance on the most important features compared to the baseline.
Access:
Access publication

Cell-type specific prediction of RNA stability from RNA-protein interactions
Authors: Saran, S., Lebedeva, S., Hirsekorn, A., & Ohler, U.; Published in: bioRxiv preprint, doi:10.1101/2024.11.19.624283.
Abstract:
RNA-binding proteins (RBPs) are important contributors to post-transcriptional regulatory processes. The combinatorial action of expressed RBPs and non-coding factors bound to the same transcript determines post-transcriptional properties of the mRNA in a context-dependent manner. To gain a better understanding of RNA stability and translational activity across different conditions, we have compiled and analyzed a set of ribosome profiling datasets for four human cell lines and used existing and newly generated metabolic labeling data to determine matching RNA degradation rates. We then used machine learning methods to predict RNA degradation rate and translation level from RBP binding information, which comprised existing in vivo binding datasets and computationally predicted binding sites. Utilizing this new RNA stability resource, we predicted RNA degradation rate and translation level from RBP binding alone. In vivo binding sites had higher importance for prediction compared to computationally predicted binding sites, likely due to confounding effects. We further explored the feature importance of different RBPs for stability prediction in the context of differential stability conferred by 3’UTR isoforms. Taken together, an RNA stability machine learning model trained on one context successfully generalizes but is impacted by the availability and reliability of current data.
Access to: full publication

DualView: Data Attribution from the Dual Perspective
Authors: Yolcu, G. Ü., Wiegand, T., Samek, W., Lapuschkin, S.
Abstract:
Local data attribution (or influence estimation) techniques aim at estimating the impact that individual data points seen during training have on particular predictions of an already trained Machine Learning model during test time. Previous methods either do not perform well consistently across different evaluation criteria from literature, are characterized by a high computational demand, or suffer from both. In this work we present DualView, a novel method for post-hoc data attribution based on surrogate modelling, demonstrating both high computational efficiency, as well as good evaluation results. With a focus on neural networks, we evaluate our proposed technique using suitable quantitative evaluation strategies from the literature against related principal local data attribution methods. We find that DualView requires considerably lower computational resources than other methods, while demonstrating comparable performance to competing approaches across evaluation metrics. Futhermore, our proposed method produces sparse explanations, where sparseness can be tuned via a hyperparameter. Finally, we showcase that with DualView, we can now render explanations from local data attributions compatible with established local feature attribution methods: For each prediction on (test) data points explained in terms of impactful samples from the training set, we are able to compute and visualize how the prediction on (test) sample relates to each influential training sample in terms of features recognized and by the model. We provide an Open Source implementation of DualView online, together with implementations for all other local data attribution methods we compare against, as well as the metrics reported here, for full reproducibility.
Access to: full publication

Navigating Neural Space: Revisiting Concept Activation Vectors to Overcome Directional Divergence
Authors: Pahde, F., Weber, L., Anders,Ch. J., Samek, W., Lapuschkin, S.; Published in: arXiv:2202.03482
Abstract:
With a growing interest in understanding neural network prediction strategies, Concept Activation Vectors (CAVs) have emerged as a popular tool for modeling human-understandable concepts in the latent space. Commonly, CAVs are computed by leveraging linear classifiers optimizing the separability of latent representations of samples with and without a given concept. However, in this paper we show that such a separability-oriented computation leads to solutions, which may diverge from the actual goal of precisely modeling the concept direction. This discrepancy can be attributed to the significant influence of distractor directions, i.e., signals unrelated to the concept, which are picked up by filters (i.e., weights) of linear models to optimize class-separability. To address this, we introduce pattern-based CAVs, solely focussing on concept signals, thereby providing more accurate concept directions. We evaluate various CAV methods in terms of their alignment with the true concept direction and their impact on CAV applications, including concept sensitivity testing and model correction for shortcut behavior caused by data artifacts. We demonstrate the benefits of pattern-based CAVs using the Pediatric Bone Age, ISIC2019, and FunnyBirds datasets with VGG, ResNet, ReXNet, EfficientNet, and Vision Transformer as model architectures.
Access to: full publication

Understanding the (Extra-)Ordinary: Validating Deep Model Decisions with Prototypical Concept-based Explanation
Authors: Dreyer, M., Achtibat, R., Samek, W., Lapuschkin, S.; Published in: CVPRW, pp. 3491-3501
Abstract:
Ensuring both transparency and safety is critical when deploying Deep Neural Networks (DNNs) in high-risk applications, such as medicine. The field of explainable AI (XAI) has proposed various methods to comprehend the decision-making processes of opaque DNNs. However, only few XAI methods are suitable of ensuring safety in practice as they heavily rely on repeated labor-intensive and possibly biased human assessment. In this work, we present a novel post-hoc concept-based XAI framework that conveys besides instance-wise (local) also class-wise (global) decision-making strategies via prototypes. What sets our approach apart is the combination of local and global strategies, enabling a clearer understanding of the (dis-)similarities in model decisions compared to the expected (prototypical) concept use, ultimately reducing the dependence on human long-term assessment. Quantifying the deviation from prototypical behavior not only allows to associate predictions with specific model sub-strategies but also to detect outlier behavior. As such, our approach constitutes an intuitive and explainable tool for model validation. We demonstrate the effectiveness of our approach in identifying out-of-distribution samples, spurious model behavior and data quality issues across three datasets (ImageNet, CUB-200, and CIFAR-10) utilizing VGG, ResNet, and EfficientNet architectures. Code is available at https://github.com/maxdreyer/pcx.
Access to: full publication

From Hope to Safety: Unlearning Biases of Deep Models via Gradient Penalization in Latent Space
Authors: Dreyer, M., Pahde, F., Anders, Ch.J., Samek, W.,Lapuschkin, S.; Published in: Vol. 38 No. 19: AAAI-24 Special Track Safe, Robust and Responsible AI Track
Abstract:
Deep Neural Networks are prone to learning spurious correlations embedded in the training data, leading to potentially biased predictions. This poses risks when deploying these models for high-stake decision-making, such as in medical applications. Current methods for post-hoc model correction either require input-level annotations which are only possible for spatially localized biases, or augment the latent feature space, thereby hoping to enforce the right reasons. We present a novel method for model correction on the concept level that explicitly reduces model sensitivity towards biases via gradient penalization. When modeling biases via Concept Activation Vectors, we highlight the importance of choosing robust directions, as traditional regression-based approaches such as Support Vector Machines tend to result in diverging directions. We effectively mitigate biases in controlled and real-world settings on the ISIC, Bone Age, ImageNet and CelebA datasets using VGG, ResNet and EfficientNet architectures. Code and Appendix are available on https://github.com/frederikpahde/rrclarc.
Access to: full publication

PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits
Authors: Dreyer, M., Purelku, E., Vielhaben, J., Samek, W., Lapuschkin, S.; Published in: arXiv:2404.06453
Abstract:
The field of mechanistic interpretability aims to study the role of individual neurons in Deep Neural Networks. Single neurons, however, have the capability to act polysemantically and encode for multiple (unrelated) features, which renders their interpretation difficult. We present a method for disentangling polysemanticity of any Deep Neural Network by decomposing a polysemantic neuron into multiple monosemantic "virtual" neurons. This is achieved by identifying the relevant sub-graph ("circuit") for each "pure" feature. We demonstrate how our approach allows us to find and disentangle various polysemantic units of ResNet models trained on ImageNet. While evaluating feature visualizations using CLIP, our method effectively disentangles representations, improving upon methods based on neuron activations. Our code is available at this https URL.
Access to: full publication

Model guidance via explanations turns image classifiers into segmentation models
Authors: Yu, X.,Franzen, J., Samek, W.,Höhne, M., Kainmüller, D. Published in: Longo, L., Lapuschkin, S., Seifert, C. (eds) Explainable Artificial Intelligence. xAI 2024. Communications in Computer and Information Science, vol 2154. Springer, Cham
Abstract:
Heatmaps generated on inputs of image classification networks via explainable AI methods like Grad-CAM and LRP have been observed to resemble segmentations of input images in many cases. Consequently, heatmaps have also been leveraged for achieving weakly supervised segmentation with image-level supervision. On the other hand, losses can be imposed on differentiable heatmaps, which has been shown to serve for (1) improving heatmaps to be more human-interpretable, (2) regularization of networks towards better generalization, (3) training diverse ensembles of networks, and (4) for explicitly ignoring confounding input features. Due to the latter use case, the paradigm of imposing losses on heatmaps is often referred to as “Right for the right reasons”. We unify these two lines of research by investigating semi-supervised segmentation as a novel use case for the Right for the Right Reasons paradigm. First, we show formal parallels between differentiable heatmap architectures and standard encoder-decoder architectures for image segmentation. Second, we show that such differentiable heatmap architectures yield competitive results when trained with standard segmentation losses. Third, we show that such architectures allow for training with weak supervision in the form of image-level labels and small numbers of pixel-level labels, outperforming comparable encoder-decoder models. Code is available at https://github.com/Kainmueller-Lab/TW-autoencoder.
Access to: full publication

Explainable Artificial Intelligence (XAI) 2.0: A manifesto of open challenges and interdisciplinary research directions
Authors: Longo, L., et al.; Published in: Information Fusion, Volume 106 June 2024, 102301
Abstract:
Understanding black box models has become paramount as systems based on opaque Artificial Intelligence (AI) continue to flourish in diverse real-world applications. In response, Explainable AI (XAI) has emerged as a field of research with practical and ethical benefits across various domains. This paper highlights the advancements in XAI and its application in real-world scenarios and addresses the ongoing challenges within XAI, emphasizing the need for broader perspectives and collaborative efforts. We bring together experts from diverse fields to identify open problems, striving to synchronize research agendas and accelerate XAI in practical applications. By fostering collaborative discussion and interdisciplinary cooperation, we aim to propel XAI forward, contributing to its continued success. We aim to develop a comprehensive proposal for advancing XAI. To achieve this goal, we present a manifesto of 28 open problems categorized into nine categories. These challenges encapsulate the complexities and nuances of XAI and offer a road map for future research. For each problem, we provide promising research directions in the hope of harnessing the collective intelligence of interested stakeholders.
Access to: full publication

Reactive Model Correction: Mitigating Harm to Task-Relevant Features via Conditional Bias Suppression
Authors: Bareeva, D., Dreyer, M.,Pahde, F., Samek, W., Lapuschkin, S.; Published in: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)
Abstract:
Deep Neural Networks are prone to learning and relying on spurious correlations in the training data, which, for high-risk applications, can have fatal consequences. Various approaches to suppress model reliance on harmful features have been proposed that can be applied post-hoc without additional training. Whereas those methods can be applied with efficiency, they also tend to harm model performance by globally shifting the distribution of latent features. To mitigate unintended overcorrection of model behavior, we propose a reactive approach conditioned on model-derived knowledge and eXplainable Artificial Intelligence (XAI) insights. While the reactive approach can be applied to many post-hoc methods, we demonstrate the incorporation of reactivity in particular for P-ClArC (Projective Class Artifact Compensation), introducing a new method called R-ClArC (Reactive Class Artifact Compensation). Through rigorous experiments in controlled settings (FunnyBirds) and with a real-world dataset (ISIC2019), we show that introducing reactivity can minimize the detrimental effect of the applied correction while simultaneously ensuring low reliance on spurious features.
Access to: full publication

DeepRepViz: Identifying Potential Confounders in Deep Learning Model Predictions
Authors: Rane, R.P., Kim, J., Umesha, A., Stark, D., Schulz, M., Ritter, K.; Published in: MICCAI 2024. Lecture Notes in Computer Science, vol 15010. Springer, Cham.
Abstract:
Deep Learning (DL) has emerged as a powerful tool in neuroimaging research. DL models predicting brain pathologies, psychological behaviors, and cognitive traits from neuroimaging data have the potential to discover the neurobiological basis of these phenotypes. However, these models can be biased by spurious imaging artifacts or by the information about age and sex encoded in the neuroimaging data. In this study, we introduce a lightweight and easy-to-use framework called ‘DeepRepViz’ designed to detect such potential confounders in DL model predictions and enhance the transparency of predictive DL models. DeepRepViz comprises two components - an online visualization tool (available at https://deep-rep-viz.vercel.app/) and a metric called the ‘Con-score’. The tool enables researchers to visualize the final latent representation of their DL model and qualitatively inspect it for biases. The Con-score, or the ‘concept encoding’ score, quantifies the extent to which potential confounders like sex or age are encoded in the final latent representation and influences the model predictions. We illustrate the rationale of the Con-score formulation using a simulation experiment. Next, we demonstrate the utility of the DeepRepViz framework by applying it to three typical neuroimaging-based prediction tasks (n = 12000). These include (a) distinguishing chronic alcohol users from controls, (b) classifying sex, and (c) predicting the speed of completing a cognitive task known as ‘trail making’. In the DL model predicting chronic alcohol users, DeepRepViz uncovers a strong influence of sex on the predictions (Con-score = 0.35). In the model predicting cognitive task performance, DeepRepViz reveals that age plays a major role (Con-score = 0.3). Thus, the DeepRepViz framework enables neuroimaging researchers to systematically examine their model and identify potential biases, thereby improving the transparency of predictive DL models in neuroimaging studies.
Access to: full publication

Arctique: An artificial histopathological dataset unifying realism and controllability for uncertainty quantification
Authors: Franzen, J., Winklmayr, C., Guarino, V.E., Karg, Ch., Yu, X., Koreuber, N., Albrecht, J.P., Bischoff, P., Kainmueller, D.; Published in: Advances in Neural Information Processing Systems 37 (NeurIPS 2024)
Abstract:
Uncertainty Quantification (UQ) is crucial for reliable image segmentation. Yet, while the field sees continual development of novel methods, a lack of agreed-upon benchmarks limits their systematic comparison and evaluation: Current UQ methods are typically tested either on overly simplistic toy datasets or on complex real-world datasets that do not allow to discern true uncertainty. To unify both controllability and complexity, we introduce Arctique, a procedurally generated dataset modeled after histopathological colon images. We chose histopathological images for two reasons: 1) their complexity in terms of intricate object structures and highly variable appearance, which yields challenging segmentation problems, and 2) their broad prevalence for medical diagnosis and respective relevance of high-quality UQ. To generate Arctique, we established a Blender-based framework for 3D scene creation with intrinsic noise manipulation. Arctique contains up to 50,000 rendered images with precise masks as well as noisy label simulations. We show that by independently controlling the uncertainty in both images and labels, we can effectively study the performance of several commonly used UQ methods. Hence, Arctique serves as a critical resource for benchmarking and advancing UQ techniques and other methodologies in complex, multi-object environments, bridging the gap between realism and controllability. All code is publicly available, allowing re-creation and controlled manipulations of our shipped images as well as creation and rendering of new scenes.
Access to: full publication

Metadata-guided Feature Disentanglement for Functional Genomics
Authors: Rakowski, A., Monti, R.,Huryn, V., Lemanczyk, M., Ohler, U., Lippert, Ch.; Published in: Bioinformatics Volume 40 Supplementary Issue - ECCB 2024 Conference Proceedings
Abstract:
With the development of high-throughput technologies, genomics datasets rapidly grow in size, including functional genomics data. This has allowed the training of large Deep Learning (DL) models to predict epigenetic readouts, such as protein binding or histone modifications, from genome sequences. However, large dataset sizes come at a price of data consistency, often aggregating results from a large number of studies, conducted under varying experimental conditions. While data from large-scale consortia are useful as they allow studying the effects of different biological conditions, they can also contain unwanted biases from confounding experimental factors. Here, we introduce Metadata-guided Feature Disentanglement (MFD)—an approach that allows disentangling biologically relevant features from potential technical biases. MFD incorporates target metadata into model training, by conditioning weights of the model output layer on different experimental factors. It then separates the factors into disjoint groups and enforces independence of the corresponding feature subspaces with an adversarially learned penalty. We show that the metadata-driven disentanglement approach allows for better model introspection, by connecting latent features to experimental factors, without compromising, or even improving performance in downstream tasks, such as enhancer prediction, or genetic variant discovery. The code will be made available at https://github.com/HealthML/MFD.
Access to: full publication

TransferGWAS of T1-weighted Brain MRI Data from UK Biobank
Authors: Rakowski, A., Monti, R., Lippert, Ch.; Published in: medRxiv preprint
Abstract:
Genome-wide association studies (GWAS) traditionally analyze single traits, e.g., disease diagnoses or biomarkers. Nowadays, large-scale cohorts such as the UK Biobank (UKB) collect imaging data with sample sizes large enough to perform genetic association testing. Typical approaches to GWAS on high-dimensional modalities extract predefined features from the data, e.g., volumes of regions of interest. This limits the scope of such studies to predefined traits and can ignore novel patterns present in the data. TransferGWAS employs deep neural networks (DNNs) to extract low-dimensional representations of imaging data for GWAS, eliminating the need for predefined biomarkers. Here, we apply transferGWAS on brain MRI data from the UKB. We encoded 36, 311 T1-weighted brain magnetic resonance imaging (MRI) scans using DNN models trained on MRI scans from the Alzheimer’s Disease Neuroimaging Initiative, and on natural images from the ImageNet dataset, and performed a multivariate GWAS on the resulting features. Furthermore, we fitted polygenic scores (PGS) of the deep features and computed genetic correlations between them and a range of selected phenotypes. We identified 289 independent loci, associated mostly with bone density, brain, or cardiovascular traits, and 14 regions having no previously reported associations. We evaluated the PGS in a multi-PGS setting, improving predictions of several traits. By examining clusters of genetic correlations, we found novel links between diffusion MRI traits and type 2 diabetes.
Access to: full publication

Sparse Explanations of Neural Networks Using Pruned Layer-Wise Relevance Propagation
Authors: Yanez Sarmiento, P., Witzke, S., Klein, N., Renard, B.; Published in: ECML PKDD 2024. Lecture Notes in Computer Science
Abstract:
Explainability is a key component in many applications involving deep neural networks (DNNs). However, current explanation methods for DNNs commonly leave it to the human observer to distinguish relevant explanations from spurious noise. This is not feasible anymore when going from easily human-accessible data such as images to more complex data such as genome sequences. To facilitate the accessibility of DNN outputs from such complex data and to increase explainability, we present a modification of the widely used explanation method layer-wise relevance propagation. Our approach enforces sparsity directly by pruning the relevance propagation for the different layers. Thereby, we achieve sparser relevance attributions for the input features as well as for the intermediate layers. As the relevance propagation is input-specific, we aim to prune the relevance propagation rather than the underlying model architecture. This allows to prune different neurons for different inputs and hence, might be more appropriate to the local nature of explanation methods. To demonstrate the efficacy of our method, we evaluate it on two types of data: images and genome sequences. We show that our modification indeed leads to noise reduction and concentrates relevance on the most important features compared to the baseline.
Access to: full publication
2023

Uncontrolled eating and sensation-seeking partially explain the prediction of future binge drinking from adolescent brain structure
Authors: Rane, R. P., Musial, M. P. M., Beck, A., Rapp, M., Schlagenhauf, F., Banaschewski, T., Bokde, A. L., Paillère Martinot, M.-L., Artiges, E., Nees, F., Lemaitre, H., Hohmann, S., Schumann, G., Walter, H., Heinz, A., & Ritter, K.; Published in: NeuroImage: Clinical, 40, 103520.
Abstract:
Binge drinking behavior in early adulthood can be predicted from brain structure during early adolescence with an accuracy of above 70%. We investigated whether this accurate prospective prediction of alcohol misuse behavior can be explained by psychometric variables such as personality traits or mental health comorbidities in a data-driven approach.
Access to: full publication

Reveal to Revise: An Explainable AI Life Cycle for Iterative Bias Correction of Deep Models
Authors: Pahde, F., Dreyer, M., Samek,W., Lapuschkin, S.; Published in: Medical Image Computing and Computer Assisted Intervention – MICCAI 2023, LNCS, 14221:596-606, Springer, Cham.
Abstract:
State-of-the-art machine learning models often learn spurious correlations embedded in the training data. This poses risks when deploying these models for high-stake decision-making, such as in medical applications like skin cancer detection. To tackle this problem, we propose Reveal to Revise (R2R), a framework entailing the entire eXplainable Artificial Intelligence (XAI) life cycle, enabling practitioners to iteratively identify, mitigate, and (re-)evaluate spurious model behavior with a minimal amount of human interaction. In the first step (1), R2R reveals model weaknesses by finding outliers in attributions or through inspection of latent concepts learned by the model. Secondly (2), the responsible artifacts are detected and spatially localized in the input data, which is then leveraged to (3) revise the model behavior. Concretely, we apply the methods of RRR, CDEP and ClArC for model correction, and (4) (re-)evaluate the model’s performance and remaining sensitivity towards the artifact. Using two medical benchmark datasets for Melanoma detection and bone age estimation, we apply our R2R framework to VGG, ResNet and EfficientNet architectures and thereby reveal and correct real dataset-intrinsic artifacts, as well as synthetic variants in a controlled setting. Completing the XAI life cycle, we demonstrate multiple R2R iterations to mitigate different biases. Code is available on https://github.com/maxdreyer/Reveal2Revise.
Access to: full publication

Explainable AI for Time Series via Virtual Inspection Layers
Authors: Vielhaben, J., Lapuschkin, S., Montavon, G., Samek, W.; Published in: NeurIPS'23 Workshop on Machine Learning for Audio, 2023
Abstract:
The field of eXplainable Artificial Intelligence (XAI) has greatly advanced in recent years, but progress has mainly been made in computer vision and natural language processing. For time series, where the input is often not interpretable, only limited research on XAI is available. In this work, we put forward a virtual inspection layer, that transforms the time series to an interpretable representation and allows to propagate relevance attributions to this representation via local XAI methods like layer-wise relevance propagation (LRP). In this way, we extend the applicability of a family of XAI methods to domains (e.g. speech) where the input is only interpretable after a transformation. Here, we focus on the Fourier transformation which is prominently applied in the interpretation of time series and LRP and refer to our method as DFT-LRP. We demonstrate the usefulness of DFT-LRP in various time series classification settings like audio and electronic health records. We showcase how DFT-LRP reveals differences in the classification strategies of models trained in different domains (e.g., time vs. frequency domain) or helps to discover how models act on spurious correlations in the data.
Access to: full publication

Scalable Estimation for Structured Additive Distributional Regression Through Variational Inference
Authors: Kleinemeier, J., Klein, N.; Published in: arXiv:2311.07371
Abstract:
Structured additive distributional regression models offer a versatile framework for estimating complete conditional distributions by relating all parameters of a parametric distribution to covariates. Although these models efficiently leverage information in vast and intricate data sets, they often result in highly-parameterized models with many unknowns. Standard estimation methods, like Bayesian approaches based on Markov chain Monte Carlo methods, face challenges in estimating these models due to their complexity and costliness. To overcome these issues, we suggest a fast and scalable alternative based on variational inference. Our approach combines a parsimonious parametric approximation for the posteriors of regression coefficients, with the exact conditional posterior for hyperparameters. For optimization, we use a stochastic gradient ascent method combined with an efficient strategy to reduce the variance of estimators. We provide theoretical properties and investigate global and local annealing to enhance robustness, particularly against data outliers. Our implementation is very general, allowing us to include various functional effects like penalized splines or complex tensor product interactions. In a simulation study, we demonstrate the efficacy of our approach in terms of accuracy and computation time. Lastly, we present two real examples illustrating the modeling of infectious COVID-19 outbreaks and outlier detection in brain activity.
Access to: full publicaton

