| CircuitKIT : Circuit Discovery, Evaluation, and Application Toolkit for Mechanistic Interpretability |
2026-07-21 |
|
| CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs |
2026-07-21 |
9 pag...9 pages, 7 figures, 1 table. Code: https://github.com/LLM-Interp/CLT-Forge. Demonstration video: https://youtu.be/6ptrrLawTl8 |
| EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures |
2026-07-21 |
This ...This manuscript is a 80-page hybrid survey and conceptual framework on LLM evaluation and AI-safety failures. It includes 8 figures and multiple evidence-synthesis tables, covering literature from 2018 to 2026. The paper introduces the EvalSafetyGap framework and reports a structured audit of 10 LLMs. It is submitted as a review/survey article and is not currently under consideration elsewhere |
| Breaking the Block: Preserving Data Continuity to Train Superior SAEs for Instruct Models |
2026-07-20 |
|
| Measuring Monosemanticity in Sparse Autoencoders via Latent Activation Coherence |
2026-07-20 |
This ...This is a preprint version. A shorter version of this paper has been accepted for presentation and publication in the post-workshop proceedings of the 8th International Workshop on eXplainable Knowledge Discovery in Data Mining (XKDD 2026), co-located with ECML PKDD 2026. The appendix is included only in this preprint and is not part of the peer-reviewed proceedings paper |
| Every Component is a Lookup: Token Attribution and Composition from a Single Decomposition |
2026-07-19 |
|
| What Do They See? Interpreting Complex Road Scenarios Through the Eyes of Vision-Language-Action Models for Safe and Trustworthy Autonomous Vehicle Learning |
2026-07-18 |
|
| Are Arithmetic Heuristic Neurons Form-Invariant? A Mechanistic Analysis of Symbols, Text, and Code in LLMs |
2026-07-18 |
Under Review |
| Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy |
2026-07-17 |
Prepr...Preprint. Under review |
| Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability |
2026-07-17 |
Prepr...Preprint. Under review |
| Hide and Seek in Embedding Space: Geometry-based Steganography and Detection in Large Language Models |
2026-07-17 |
|
| Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control |
2026-07-16 |
|
| Transcoders for Investigating Deception in Language Models |
2026-07-16 |
|
| Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions |
2026-07-15 |
|
| From Observation to Insight: Mechanistic World Models and the Quest for Autonomous Discovery |
2026-07-15 |
|
| Hierarchical Latent Structures in Data Generation Process Unify Mechanistic Phenomena across Scale |
2026-07-14 |
|
| Same Compression Principle, Different Geometry: Rate-Distortion Signatures Dissociate Biological and Artificial Visual Systems |
2026-07-13 |
|
| Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias |
2026-07-13 |
58 pa...58 pages, 13 figures, 30 tables; project page: https://xzx34.github.io/unfair-judge/ |
| Emergent Hierarchical Monosemantic Neurons from the Group-Contrastive Forward-Forward Algorithm |
2026-07-13 |
|
| Which Neurons Detect Malicious Code? A Probing Study of LLM Security Knowledge |
2026-07-11 |
The p...The paper has been peer reviewed and accepted for publication in the 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026) |
| MLPs are Hebbians: Constructing Efficient Fact-Storing MLPs for Transformers |
2026-07-10 |
|
| XAI and Statistical Analysis for Reliable Intrusion Detection in the UAVIDS-2025 Dataset: From Tree to Hybrid and Tabular DNN Ensembles |
2026-07-10 |
Accep...Accepted at IEEE CITS 2026, Greece |
| When Structured Sparse Autoencoders Learn Consistent Concepts Across Modalities |
2026-07-09 |
|
| Cross-seed explainability using Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoders |
2026-07-09 |
17 pa...17 pages, 4 figures, 6 tables |
| Towards Isolated Interventions via Almost Orthogonal Features in Language Models |
2026-07-09 |
Accep...Accepted as a conference paper at the Conference on Language Modeling (COLM) 2026 |
| Certified Interventional Fidelity: Anytime-Valid, Adaptive Evaluation of Causal Claims in Mechanistic Interpretability |
2026-07-09 |
Accep...Accepted at UAI 2026 (Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence). Code: https://github.com/AsiaeeLab/certified-interventional-fidelity |
| Diagnosing Shape-Prior Shortcuts in Long-Range Single-Shot Fringe Projection Profilometry |
2026-07-09 |
21 pages, 13 figures |
| Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs |
2026-07-08 |
|
| Temporal Preference Concepts and their Functions in a Large Language Model |
2026-07-08 |
|
| How Learning Dynamics Drive Adversarially Robust Generalization? |
2026-07-08 |
Accep...Accepted at the 42nd Conference on Uncertainty in Artificial Intelligence (UAI 2026) |
| Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning |
2026-07-08 |
20 pages |
| Latent Programming Horizons in Coding Agents |
2026-07-06 |
|
| Beyond the Black Box: Interpretability of Agentic AI Tool Use |
2026-07-05 |
12 pa...12 pages, 4 figures, 17 tables |
| Interpretability and Generalization Bounds for Learning Spatial Physics |
2026-07-05 |
To ap...To appear in ICML 2026. 18 pages, 13 figures |
| Cultural Binding Heads in Language Models |
2026-07-04 |
|
| Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models |
2026-07-03 |
|
| Individual Parameters in Weight-Sparse Transformers Appear Interpretable |
2026-07-03 |
20 pa...20 pages, 19 figures, 3 tables. Project website: https://weightpedia.org/individual-parameters-in-sparse-transformers/ |
| MatPhaseBench: A Semantics-Guided Benchmark for Materials Phase Diagrams Understanding |
2026-07-03 |
|
| Induction Heads Interpolate N-Grams |
2026-07-02 |
Publi...Published as a conference paper at ICML 2026. OpenReview: https://openreview.net/forum?id=BSY7jhBxM1 |
| Towards Robustness against Typographic Attack with Training-free Concept Localization |
2026-07-02 |
15 pa...15 pages main text, provisionally accepted to ECCV 2026 |
| Conditional Co-Ablation: Recovering Self-Repair Backups in Transformer Circuits |
2026-07-02 |
|
| Shared Semantics, Divergent Mechanisms: Unsupervised Feature Discovery by Aligning Semantics and Mechanisms |
2026-07-02 |
40 pa...40 pages; accepted as an ICML 2026 Spotlight; project page: https://merenova.github.io/distribution-level-feature-discovery/ |
| Expander Sparse Autoencoders: Parameter-Efficient Dictionaries for Mechanistic Interpretability |
2026-07-02 |
|
| Mechanistic Interpretability and Causal Feature Steering of Neural Quantum States via Sparse Autoencoders |
2026-07-01 |
15 pa...15 pages, 7 figures. Comments welcome! |
| Muon as a Residual Connection |
2026-07-01 |
|
| Interpreting Global Perturbation Robustness of Image Models using Axiomatic Spectral Importance Decomposition |
2026-07-01 |
Accep...Accepted by Transactions on Machine Learning Research (TMLR 2024) |
| MetaOthello: A Controlled Study of Multiple World Models in Transformers |
2026-07-01 |
Camer...Camera-ready version. Accepted to the 43rd International Conference on Machine Learning (ICML 2026) |
| Representation as a Bottleneck for Mechanistic Interpretability: The Manifestation Unit Protocol |
2026-06-30 |
65 pa...65 pages. Interactive demos: https://manifestation-xai.github.io/manifestation-transformers/ , https://manifestation-xai.github.io/manifestation-cnn |
| Surrogate Fidelity: When Can Open LLMs Explain Closed Ones? |
2026-06-30 |
|
Please check the Github page for a better reading experience and more papers.
LLM Reasoning
9 pag...
9 pages. Dataset and reproduction code: https://github.com/senthex-security/senthex-research
13 pa...
13 pages, 2 figures, 8 tables. Code: https://github.com/AmGarfield/OracleGap
18 pa...
18 pages, Accepted by ECML-PKDD 2026
The a...
The authors have decided to withdraw this manuscript due to concerns regarding its current scope, framing, and presentation. Please do not cite this version
Prese...
Presented at COLIEE 2026
Accep...
Accepted by ECCV 2026, project page https://1ranguan.github.io/VST/
EKAW ...
EKAW 2026 (https://ekaw2026.di.unito.it/accepted-posters-and-demos)
Under...
Under Review, preprint
32 pa...
32 pages, 8 figures, 10 tables
This ...
This paper has been withdrawn by the authors because the current version requires substantial revision and further validation before it can be considered a reliable representation of the work
7 pag...
7 pages, 3 figures, 6 tables
SIGCO...
SIGCOMM'26(19 pages, 6 figures, 6 tables)
42 pa...
42 pages, 14 figures, 12 tables
The a...
The article has been accepted by Frontiers of Computer Science (FCS), with the DOI: {10.1007/s11704-026-51673-0}
37 pa...
37 pages, 16 figures, accepted to 3rd AI for Math Workshop at ICML 2026
Accep...
Accepted to ICML 2026
The c...
The code is available at https://github.com/linhh29/Interactive-Learning-for-LLM-Reasoning
Accep...
Accepted as a short paper at IEEE VIS 2026. 5 pages, 2 figures
Chain of Thought
Accep...
Accepted by ACL 2026 Main Conference. 30 pages, 6 figures
12 pa...
12 pages, 7 figures. Zhongyao Yang and Haoyu Li contributed equally to this work
101 p...
101 pages, 66 figures
Accep...
Accepted for 1st Workshop on Small Language Models for Education (SLM4ED '26) at AIED 2026
23 pa...
23 pages, 7 figures, 3 tables
Accep...
Accepted at the Conference on Language Modeling (COLM) 2026. 45 pages, including appendices; 24 figures and 12 tables. Code: https://github.com/withmartian/mi-cot
Accep...
Accepted to ICML Workshop on High-Dimensional Learning Dynamics
9 pag...
9 pages, 5 figures. This version substantially revises the previous preprint with a new method, updated experiments, and rewritten analysis. Code available at the GitHub project repository https://anonymous.4open.science/r/sca-B666
Accep...
Accepted at ICLR 2026. Camera-ready version
Accep...
Accepted to AdaptFM Workshop, ICML 2026
The a...
The authors have decided to withdraw this manuscript due to concerns regarding its current scope, framing, and presentation. Please do not cite this version
Prese...
Presented at COLIEE 2026
accep...
accepted at the 37th IEEE International Symposium on Software Reliability Engineering (ISSRE 2026)
Proje...
Project Page: https://avflamingo.pages.dev/
Accep...
Accepted by ECCV 2026, project page https://1ranguan.github.io/VST/
LLM Interpretability
full ...
full version of this short paper is accepted at Frontiers in Robotics and AI Journal
Paper...
Paper accepted at IEEE Intelligent Transportation Systems Conference (ITSC) 2026
Under...
Under review at ACL Rolling Review (ARR), May 2026 cycle. Also available at https://doi.org/10.5281/zenodo.20125267
Accep...
Accepted at IAQVEC 2026
26 pa...
26 pages, 11 figures, 8 tables
12 pa...
12 pages, 4 figures, 8 tables
Accep...
Accepted by CVPR 2026. Supplementary material included. Project page: https://jjkislele.github.io/LaMoGen/
Accep...
Accepted by International Conference on Learning Representations (ICLR) 2026
Accep...
Accepted by IEEE ICASSP 2026
10 pa...
10 pages, 6 tables, 3 figures, accepted by the 2nd ACM AIware Conference
Accep...
Accepted by AAAI 2026 Workshop on Agentic AI Benchmarks and Applications for Enterprise Tasks
ACL 2...
ACL 2025 main. Our Speech-IQ leaderboard is hosted at huggingface.co/spaces/nvidia/Speech-IQ-leaderboard. Speech-IQ Calculator: https://github.com/YukinoWan/SpeechIQ
Accep...
Accepted by AAAI-2026
5 pag...
5 pages, 1 figure, Accepted for publication at the Demonstration Track of the 40th AAAI Conference on Artificial Intelligence (AAAI 26)
Prepr...
Preprint. Project page: https://davidhalladay.github.io/diysink_demo
Explainable AI
v2: r...
v2: revised after peer review. Evaluation expanded from 3 to 7 datasets, per-dataset Wasserstein-threshold calibration added, effect-size CIs and cross-dataset statistics reported, and analytical zk-SNARK/MPC/TEE baselines added. Minor errors corrected
Accep...
Accepted for AALTD workshop at ECML-PKDD 2026
25 pa...
25 pages, 2 figures, no tables, 16 references. Conceptual and methodological framework for monitoring representational adequacy and detecting explanatory insufficiency in learned representations
Accep...
Accepted for presentation at the AISE Workshop @ IJCAI-ECAI 2026
This ...
This preprint has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this contribution is published in Artificial Intelligence in HCI (HCII 2026), Lecture Notes in Computer Science, vol. 16745, and is available online at https://doi.org/10.1007/978-3-032-30849-8_13
Accep...
Accepted at CVPRW 2026
Invit...
Invited survey paper. 10 pages, 5 figures, 2 tables
Submi...
Submitted to Imaging Neuroscience. This all-in-one version includes supplementary materials. 34 pages, 145 figures, 4 tables
21 pa...
21 pages, 12 figures. Accepted to the 4th World Conference on eXplainable Artificial Intelligence. Method implementation: https://research-software-directory.org/software/distance-explainer
Prepr...
Preprint: Under Review; Updated experiments & Figures
30 pa...
30 pages main text, 8 figures, 4 tables. Supplementary material is included in the appendix
Proc....
Proc. Joint European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML-PKDD 2026), LNCS, Naples, Italy, 7-11 September 2026
Accep...
Accepted for publication in Empirical Software Engineering (EMSE) Journal, 2026
Accep...
Accepted by IEEE Transactions on Information Forensics & Security (TIFS)
Accep...
Accepted for 2026 36th IEEE International Conference on Robot and Human Interactive Communication
6 pag...
6 pages, 2 figures, 4 tables. Accepted at the 2026 IEEE International Conference on Artificial Intelligence (CAI), 8-10 May 2026, Granada, Spain. Code: https://github.com/darianfgUgr/CoNa
Mechanistic Interpretability
9 pag...
9 pages, 7 figures, 1 table. Code: https://github.com/LLM-Interp/CLT-Forge. Demonstration video: https://youtu.be/6ptrrLawTl8
This ...
This manuscript is a 80-page hybrid survey and conceptual framework on LLM evaluation and AI-safety failures. It includes 8 figures and multiple evidence-synthesis tables, covering literature from 2018 to 2026. The paper introduces the EvalSafetyGap framework and reports a structured audit of 10 LLMs. It is submitted as a review/survey article and is not currently under consideration elsewhere
This ...
This is a preprint version. A shorter version of this paper has been accepted for presentation and publication in the post-workshop proceedings of the 8th International Workshop on eXplainable Knowledge Discovery in Data Mining (XKDD 2026), co-located with ECML PKDD 2026. The appendix is included only in this preprint and is not part of the peer-reviewed proceedings paper
Prepr...
Preprint. Under review
Prepr...
Preprint. Under review
58 pa...
58 pages, 13 figures, 30 tables; project page: https://xzx34.github.io/unfair-judge/
The p...
The paper has been peer reviewed and accepted for publication in the 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)
Accep...
Accepted at IEEE CITS 2026, Greece
17 pa...
17 pages, 4 figures, 6 tables
Accep...
Accepted as a conference paper at the Conference on Language Modeling (COLM) 2026
Accep...
Accepted at UAI 2026 (Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence). Code: https://github.com/AsiaeeLab/certified-interventional-fidelity
Accep...
Accepted at the 42nd Conference on Uncertainty in Artificial Intelligence (UAI 2026)
12 pa...
12 pages, 4 figures, 17 tables
To ap...
To appear in ICML 2026. 18 pages, 13 figures
20 pa...
20 pages, 19 figures, 3 tables. Project website: https://weightpedia.org/individual-parameters-in-sparse-transformers/
Publi...
Published as a conference paper at ICML 2026. OpenReview: https://openreview.net/forum?id=BSY7jhBxM1
15 pa...
15 pages main text, provisionally accepted to ECCV 2026
40 pa...
40 pages; accepted as an ICML 2026 Spotlight; project page: https://merenova.github.io/distribution-level-feature-discovery/
15 pa...
15 pages, 7 figures. Comments welcome!
Accep...
Accepted by Transactions on Machine Learning Research (TMLR 2024)
Camer...
Camera-ready version. Accepted to the 43rd International Conference on Machine Learning (ICML 2026)
65 pa...
65 pages. Interactive demos: https://manifestation-xai.github.io/manifestation-transformers/ , https://manifestation-xai.github.io/manifestation-cnn