Skip to content

Latest 50 Papers - July 23, 2026 #124

Description

@github-actions

Please check the Github page for a better reading experience and more papers.

LLM Reasoning

Title Date Comment
They'll Verify. They Just Won't Act. How Authority Framing and Laundered Code Turn a Trusted Agentic CI/CD Pipeline Into an Attack Surface 2026-07-21
9 pag...

9 pages. Dataset and reproduction code: https://github.com/senthex-security/senthex-research

Reasoning Error from Known Fact: Step-Level Self-Consistency Group Relative Policy Optimization for LLM 2026-07-21
Tracing the Shadows: Automatic Tracking and Analysis of Crypto Money Laundering via Transaction Semantic Analysis 2026-07-21
Not All Errors Are Created Equal: ASCoT Addresses Late-Stage Fragility in Efficient LLM Reasoning 2026-07-21
LLM-Grounded Explainable AI for Supply Chain Risk Early Warning via Temporal Graph Attention Networks 2026-07-21
QuArch: A Benchmark for Evaluating LLM Reasoning in Computer Architecture 2026-07-21
Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration 2026-07-20
13 pa...

13 pages, 2 figures, 8 tables. Code: https://github.com/AmGarfield/OracleGap

LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning 2026-07-20
AIGB-R1: Self-Evolving Generative Auto-Bidding via Hierarchical Planner-Executor Optimization 2026-07-19
Debate-on-Graph: Reliable and Adaptive Reasoning of Large Language Model on Uncertain Knowledge Graph 2026-07-19
18 pa...

18 pages, Accepted by ECML-PKDD 2026

Constrained Path Reasoning: Measuring When Committed Stages Earn Their Cost 2026-07-19
Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM Reasoning 2026-07-19
Logic-Guided Socially-aware Robot Navigation World Model 2026-07-18
The a...

The authors have decided to withdraw this manuscript due to concerns regarding its current scope, framing, and presentation. Please do not cite this version

NOWJ@COLIEE 2026: Adaptive Pipelines for Legal Retrieval and Reasoning 2026-07-18
Prese...

Presented at COLIEE 2026

Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously 2026-07-17
Accep...

Accepted by ECCV 2026, project page https://1ranguan.github.io/VST/

REST: Receding Horizon Explorative Steiner Tree for Zero-Shot Object-Goal Navigation 2026-07-16 Accepted to IROS'26
MARS: Multi-hop Adaptive Retrieval and SPARQL Generation for KGQA 2026-07-16
EKAW ...

EKAW 2026 (https://ekaw2026.di.unito.it/accepted-posters-and-demos)

ProfMalPlus: Agent-Coordinated Detection of Malicious NPM Packages via Static-Dynamic Analysis Synergy 2026-07-15
Policy of Thoughts: Scaling Test-Time Training for LLM Reasoning via Online Policy Evolution 2026-07-15
Under...

Under Review, preprint

Assert, don't describe: Linguistic features that shift LLM reasoning about animal welfare 2026-07-12
Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction 2026-07-11
32 pa...

32 pages, 8 figures, 10 tables

Can AI Reason Like an Urban Planner? Benchmarking Large Language Models Against Professional Judgment 2026-07-11
This ...

This paper has been withdrawn by the authors because the current version requires substantial revision and further validation before it can be considered a reliable representation of the work

Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization 2026-07-11 ICML 2026
EmoStyle: Affective Conditioning of Style-Specialist Experts for Emotional Image Generation 2026-07-11
Silent Failures in Quantized LLM Reasoning: A Taxonomy-Based Analysis of Hollow Convergence and Failure Mode Shifts 2026-07-10
7 pag...

7 pages, 3 figures, 6 tables

Using LLMs to Adjudicate Static-Analysis Alerts with Error Reduction Techniques 2026-07-10 22 pages
Hierarchical Chain-of-Thought: Enhancing LLM Reasoning Performance and Efficiency 2026-07-10
Practical Source Code Recovery from Binary Functions Using Anchor-Based Retrieval and LLM Reasoning 2026-07-10 12 pages, 5 figures
RepLLM: Toward Automatically Reproducing Network Research Results 2026-07-09
SIGCO...

SIGCOMM'26(19 pages, 6 figures, 6 tables)

Hierarchical Control in Multi-Agent Games: LLM-based Planning and RL Execution 2026-07-09 12 pages, 9 figures
TrajAudit: Automated Failure Diagnosis for Agentic Coding Systems 2026-07-09
Can We Trust LLM's Logic? Quantifying Uncertainty, Coherence, and Robustness via a Graph-Based Framework 2026-07-09
42 pa...

42 pages, 14 figures, 12 tables

Fast, Slow, and Tool-augmented Thinking for LLMs: A Review 2026-07-08
The a...

The article has been accepted by Frontiers of Computer Science (FCS), with the DOI: {10.1007/s11704-026-51673-0}

MILES: Modular Instruction Memory with Learnable Selection for Self-Improving LLM Reasoning 2026-07-08
Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics 2026-07-07
37 pa...

37 pages, 16 figures, accepted to 3rd AI for Math Workshop at ICML 2026

Knowing When to Quit: A Principled Framework for Dynamic Abstention in LLM Reasoning 2026-07-07
Open-Ended Scenario Reasoning for Specialist Model Adaptation 2026-07-07
RS-Agent: Automating Remote Sensing Tasks through Intelligent Agent 2026-07-07
AgentsCAD: Automated Design for Manufacturing of FDM Parts via Multi-Agent LLM Reasoning and Geometric Feature Recognition 2026-07-07
ChargeBD: Character-Aware Heterogeneous Agent Reasoning for Guided Engineering in Battery Development 2026-07-06
CAP-CoT: Cycle Adversarial Prompt for Improving Chain of Thoughts in LLM Reasoning 2026-07-06
Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens 2026-07-06
Accep...

Accepted to ICML 2026

Interactive Learning for LLM Reasoning 2026-07-06
The c...

The code is available at https://github.com/linhh29/Interactive-Learning-for-LLM-Reasoning

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning 2026-07-05
Scalable Semantic Steering of Embedding Projections 2026-07-04
Accep...

Accepted as a short paper at IEEE VIS 2026. 5 pages, 2 figures

rePIRL: Learn PRM with Inverse RL for LLM Reasoning 2026-07-03
LLM-Enhanced Hierarchical Heterogeneous Graph Representation Learning for Malicious Python Package Detection 2026-07-03
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning 2026-07-03
OpenSIR: Open-Ended Self-Improving Reasoner 2026-07-03

Chain of Thought

Title Date Comment
LinguistAgent Technical Report: A Reflective Multi-Model Platform for Automated Linguistic Annotation 2026-07-21
ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D 2026-07-21 51 pages, 11 figures
AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning 2026-07-21
Accep...

Accepted by ACL 2026 Main Conference. 30 pages, 6 figures

Cognitive Dual-Process Planning for Autonomous Driving with Structured Scene Knowledge and Verifiable Reasoning-Action Consistency 2026-07-21
12 pa...

12 pages, 7 figures. Zhongyao Yang and Haoyu Li contributed equally to this work

DAIS: Dependency-Aware Intermediate QA Supervision for Complex Reasoning 2026-07-21
Measuring Reward-Seeking via Contrastive Belief Updates 2026-07-21
101 p...

101 pages, 66 figures

Not All Errors Are Created Equal: ASCoT Addresses Late-Stage Fragility in Efficient LLM Reasoning 2026-07-21
Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models 2026-07-21
CASE: Causal Alignment and Structural Enforcement for Improving Chain-of-Thought Faithfulness 2026-07-21
Bounding Boxes to Improve Small Language Model Performance on Vision-Based Grading Tasks 2026-07-21
Accep...

Accepted for 1st Workshop on Small Language Models for Education (SLM4ED '26) at AIED 2026

Rationale-Guided Knowledge Distillation for Cross-Lingual Stance Detection 2026-07-21
23 pa...

23 pages, 7 figures, 3 tables

Robust Reasoning Benchmark 2026-07-21
LatentMT: Machine Translation with Latent Reasoning 2026-07-21
Reasoning Fine-Tuning Induces Persistent Latent Policy States 2026-07-20
Accep...

Accepted at the Conference on Language Modeling (COLM) 2026. 45 pages, including appendices; 24 figures and 12 tables. Code: https://github.com/withmartian/mi-cot

Measuring and Mitigating Post-hoc Rationalization in Reverse Chain-of-Thought Generation 2026-07-20 ICML 2026
Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values 2026-07-20
A Geometric Perspective on Stabilizing Value Conflict Resolution 2026-07-20
Accep...

Accepted to ICML Workshop on High-Dimensional Learning Dynamics

Benchmarking Resource-Efficient LLMs for Research Topic Ontology Generation in the Biomedical Field 2026-07-20
SCA: Segment-Wise CoT Compression with Answer Alignment 2026-07-20
9 pag...

9 pages, 5 figures. This version substantially revises the previous preprint with a new method, updated experiments, and rewritten analysis. Code available at the GitHub project repository https://anonymous.4open.science/r/sca-B666

ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding 2026-07-20 ICML 2026 - main
When a Name Is Not a Name: A Benchmark Dataset and Distilled Reasoning for Culturally Entangled Bangla Homographs in Low-Resource LLMs 2026-07-20
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos 2026-07-20
Accep...

Accepted at ICLR 2026. Camera-ready version

Reasoning as a Double-Edged Sword: Architecture and Cross-Stage Robustness in Vision-Language-Action Models 2026-07-20
Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration 2026-07-20
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation 2026-07-20
LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce Learning 2026-07-19
Is Your Model Thinking or Just Stagnating? PUMA: Diagnosing Reasoning Pathology via Phase-Momentum Alignment 2026-07-19
RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought 2026-07-19
Training Continuous Chain of Thought Models: A Tale of Two Regimes 2026-07-18
Accep...

Accepted to AdaptFM Workshop, ICML 2026

FUSAR-R1: A Large-Scale Reasoning Model for Intelligent Interpretation of SAR Images 2026-07-18
GRASP: Grounded CoT Reasoning with Dual-Stage Optimization for Multimodal Sarcasm Target Identification 2026-07-18
Logic-Guided Socially-aware Robot Navigation World Model 2026-07-18
The a...

The authors have decided to withdraw this manuscript due to concerns regarding its current scope, framing, and presentation. Please do not cite this version

NOWJ@COLIEE 2026: Adaptive Pipelines for Legal Retrieval and Reasoning 2026-07-18
Prese...

Presented at COLIEE 2026

Visual Access Boundaries in Vision-Language Model Reasoning 2026-07-18
From Discussion to Execution: Replicating Buggy and Correct Data Science Code 2026-07-18
accep...

accepted at the 37th IEEE International Symposium on Software Reliability Engineering (ISSRE 2026)

WeedExpert-R1: Incentivizing Botanical Reasoning in MLLMs with Reinforcement Learning for Precision Weed Grounding 2026-07-17
RIMS: Preference Optimization via Smoothed Multi-pair Aggregation for Small-Scale LLM Retrieval-Augmented Generation 2026-07-17
Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos 2026-07-17
Proje...

Project Page: https://avflamingo.pages.dev/

What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors 2026-07-17
Latency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length 2026-07-17
MoT: Modularization-of-Thought Prompting for Effective Code Generation 2026-07-17
Length Penalties Make Chain-of-Thought Less Monitorable 2026-07-17
LLM-Guided Transportation Hub Capacity Planning with Textual Business Inputs 2026-07-17
Better Starts, Better Ends: Bootstrapped Iterative Self-Reasoning Distillation for Compressed Reasoning 2026-07-17
Brain-CLIPLM: Semantic Compression for EEG-to-Text Decoding 2026-07-17
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously 2026-07-17
Accep...

Accepted by ECCV 2026, project page https://1ranguan.github.io/VST/

ABot-N1: Toward a General Visual Language Navigation Foundation Model 2026-07-17
Neuro-Symbolic AI for LEED compliance: Document-Centric Benchmarking, Deterministic Numeric Checking, and When Multimodal Hurts 2026-07-17
RecGPT-V3 Technical Report 2026-07-17 Technique Report

LLM Interpretability

Title Date Comment
Understanding Interpretation Difficulty in Harmful Online Communication: Insights from Cybercrime Communities 2026-07-08
NeuraDock Visual Cognitive Load Agent Tutorial: A Quality-Gated Open-Source EEG Workflow for Alpha Dynamics and Real-Time Applications 2026-06-25 22 pages, 10 figures
Don't Go Breaking My LLM: The Impact of Pruning Attention Layers on Explanation Faithfulness and Confidence Calibration 2026-06-23 Accepted at TMLR
LLM-Based Generalizable Hierarchical Task Planning and Execution for Heterogeneous Robot Teams with Event-Driven Replanning 2026-06-18
full ...

full version of this short paper is accepted at Frontiers in Robotics and AI Journal

ICA Lens: Interpreting Language Models Without Training Another Dictionary 2026-06-10 Ongoing Project
Language-Driven Cost Optimization for Autonomous Driving 2026-06-09
Paper...

Paper accepted at IEEE Intelligent Transportation Systems Conference (ITSC) 2026

ReflectiChain: Epistemic Grounding in LLM-Driven World Models for Supply Chain Resilience 2026-06-09
MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding 2026-06-08
Analyzing the Correlation Between Hallucinations and Knowledge Conflicts in Large Language Models 2026-06-07
Thinking Through Signs: PEEL as a Semiotic Scaffolding for Epistemically Accountable AI-Enabled Research 2026-06-02 10 pages, 5 figuras
Unveiling the Limits of Large Language Models in Inferring Pragmatic Meaning from Non-Verbal Responses 2026-06-01
The Limits of LLM Forecasting: Parametric Knowledge Gaps Across Conflict Zones 2026-05-29
Hallucination Behavior in Multimodal LLMs Across Agricultural Image Interpretation and Generation Tasks 2026-05-26
How Much Structure Do LLMs Need? Evaluating LLMs for Bibliometric Cluster Description 2026-05-23
Improving Labeling Consistency with Detailed Constitutional Definitions and AI-Driven Evaluation 2026-05-22
Under...

Under review at ACL Rolling Review (ARR), May 2026 cycle. Also available at https://doi.org/10.5281/zenodo.20125267

Patch-Effect Graph Kernels for LLM Interpretability 2026-05-07
LLM-ADAM: A Generalizable LLM Agent Framework for Pre-Print Anomaly Detection in Additive Manufacturing 2026-05-05 24 pages, 10 figures
From Research Question to Scientific Workflow: Leveraging Agentic AI for Science Automation 2026-04-23
Agentic AI-Enabled Framework for Thermal Comfort and Building Energy Assessment in Tropical Urban Neighborhoods 2026-04-23
Accep...

Accepted at IAQVEC 2026

Learning to Draw ASCII Improves Spatial Reasoning in Language Models 2026-04-16
Arbitration Failure, Not Perceptual Blindness: How Vision-Language Models Resolve Visual-Linguistic Conflicts 2026-04-13
Heterogeneous Graph Importance Scoring and Clustering with Automated LLM-based Interpretation 2026-04-09
26 pa...

26 pages, 11 figures, 8 tables

Runtime Execution Traces Guided Automated Program Repair with Multi-Agent Debate 2026-04-03
12 pa...

12 pages, 4 figures, 8 tables

Discovering Decoupled Functional Modules in Large Language Models 2026-03-18 AAAI-26 Oral
LaMoGen: Language to Motion Generation Through LLM-Guided Symbolic Inference 2026-03-12
Accep...

Accepted by CVPR 2026. Supplementary material included. Project page: https://jjkislele.github.io/LaMoGen/

Pneuma-Seeker: A Relational Reification Mechanism to Align AI Agents with Human Work over Relational Data 2026-03-11
PhysLLM: Harnessing Large Language Models for Cross-Modal Remote Physiological Sensing 2026-03-05
Accep...

Accepted by International Conference on Learning Representations (ICLR) 2026

Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs 2026-03-01
Sparse Shift Autoencoders for Identifying Concepts from Large Language Model Activations 2026-02-27 27 pages, 9 figures
Co-Disclosing the Computer: LLM-Mediated Computing through Reflective Conversation 2026-02-27 CHI'26
Neuro-Symbolic Control with Large Language Models for Language-Guided Spatial Tasks 2026-02-21
Stroke Lesions as a Rosetta Stone for Language Model Interpretability 2026-02-03 45 pages, 17 figures
Uncertainty-Aware Gradient Signal-to-Noise Data Selection for Instruction Tuning 2026-01-20 Preprint
A Shared Geometry of Difficulty in Multilingual Language Models 2026-01-19
Causal-SAM-LLM: Large Language Models as Causal Reasoners for Robust Medical Segmentation 2026-01-16
Accep...

Accepted by IEEE ICASSP 2026

Self-reflection in Automated Qualitative Coding: Improving Text Annotation through Secondary LLM Critique 2026-01-14
Word Synchronization Challenge: A Benchmark for Word Association Responses for Large Language Models 2026-01-14
Can LLMs interpret figurative language as humans do?: surface-level vs representational similarity 2026-01-14 17 pages, 5 figures
Neuro-Symbolic Compliance: Integrating LLMs and SMT Solvers for Automated Financial Legal Analysis 2026-01-07
10 pa...

10 pages, 6 tables, 3 figures, accepted by the 2nd ACM AIware Conference

LLM Interpretability with Identifiable Temporal-Instantaneous Representation 2026-01-02 NeurIPS 2025
Trust Me, I Know This Function: Hijacking LLM Static Analysis using Bias 2025-12-18
Verification-Guided Context Optimization for Tool Calling via Hierarchical LLMs-as-Editors 2025-12-15
Accep...

Accepted by AAAI 2026 Workshop on Agentic AI Benchmarks and Applications for Enterprise Tasks

Visualizing token importance for black-box language models 2025-12-12
SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language Models 2025-12-01
ACL 2...

ACL 2025 main. Our Speech-IQ leaderboard is hosted at huggingface.co/spaces/nvidia/Speech-IQ-leaderboard. Speech-IQ Calculator: https://github.com/YukinoWan/SpeechIQ

For Those Who May Find Themselves on the Red Team 2025-11-23
HSKBenchmark: Modeling and Benchmarking Chinese Second Language Acquisition in Large Language Models through Curriculum Tuning 2025-11-19
Accep...

Accepted by AAAI-2026

KnowThyself: An Agentic Assistant for LLM Interpretability 2025-11-05
5 pag...

5 pages, 1 figure, Accepted for publication at the Demonstration Track of the 40th AAAI Conference on Artificial Intelligence (AAAI 26)

Imperfect Language, Artificial Intelligence, and the Human Mind: An Interdisciplinary Approach to Linguistic Errors in Native Spanish Speakers 2025-11-03 12 pages, 3 figures
To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models 2025-10-09
Prepr...

Preprint. Project page: https://davidhalladay.github.io/diysink_demo

Automated Repair of Ambiguous Problem Descriptions for LLM-Based Code Generation 2025-09-24

Explainable AI

Title Date Comment
LLM-Grounded Explainable AI for Supply Chain Risk Early Warning via Temporal Graph Attention Networks 2026-07-21
Automated Data Engineering and Feature Selection for the Case Study of Warpage Detection in Fused Deposition Modeling 2026-07-20
Explaining and Tuning Transformer-based LLMs in Arithmetic Tasks with Human Strategies 2026-07-19
FST.ai 2.5: Explainable and Uncertainty-Aware AI for Olympic and Para-Taekwondo Decision Support, Athlete Digital Twins, and Federation-Scale Analytics 2026-07-18 14 pages, 4 figures
FAIR_XAI: Improving Multimodal Foundation Model Fairness via Explainability for Wellbeing Assessment 2026-07-17 11 pages, 4 figures
VeriX-Anon: A Multi-Layered Framework for Mathematically Verifiable Outsourced Target-Driven Data Anonymization 2026-07-17
v2: r...

v2: revised after peer review. Evaluation expanded from 3 to 7 datasets, per-dataset Wasserstein-threshold calibration added, effect-size CIs and cross-dataset statistics reported, and analytical zk-SNARK/MPC/TEE baselines added. Minor errors corrected

Scaling Time Series Classification via XAI-Driven Data Reduction 2026-07-17
Accep...

Accepted for AALTD workshop at ECML-PKDD 2026

Automated identification of Ichneumonoidea wasps via YOLO-based deep learning: Integrating HiresCam for Explainable AI 2026-07-16 15 pages, 20 figures
Explaining Process Control Optimisation Recommendations via GradientSHAP and Implicit Differentiation 2026-07-16
Diagnosing and Mitigating Domain Shift in Permission-Based Android Malware Detection 2026-07-15
"Trust Junk" Leads to Unjustified Support for Highly Discriminatory Predictive Models 2026-07-14
MLPTR-CC: Multi-label Pathology Test Recommendation using Classifier Chains and SHAP 2026-07-14
Atomic Units of X: The Compression Layer of Intelligence 2026-07-14
Evaluating RE Practices for Explainability: Synthesizing Insights from Daimler Truck into an Explainable RE Framework Proposal 2026-07-13
Detecting Explanatory Insufficiency in Learned Representations: A Framework for Representational Vigilance 2026-07-13
25 pa...

25 pages, 2 figures, no tables, 16 references. Conceptual and methodological framework for monitoring representational adequacy and detecting explanatory insufficiency in learned representations

ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI 2026-07-10
Knowledge Graphs and Explainable AI as Complementary Resources for Urban Mining 2026-07-10
Accep...

Accepted for presentation at the AISE Workshop @ IJCAI-ECAI 2026

All Explanations are Wrong, But Many Are Useful: Exploring the Rashomon Explanation Set with Large Language Models 2026-07-10
Application of machine learning to monster level prediction in tabletop RPG game design 2026-07-10
Unlearning to Protect: A Distilled Reinforcement Learning Framework with Privacy-Preserving Feature Unlearning and XAI for IoT Security 2026-07-09
The Contribution of XAI for the Safe Development and Certification of AI: An Expert-Based Analysis 2026-07-09
This ...

This preprint has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this contribution is published in Artificial Intelligence in HCI (HCII 2026), Lecture Notes in Computer Science, vol. 16745, and is available online at https://doi.org/10.1007/978-3-032-30849-8_13

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning 2026-07-08 20 pages
Why Fake ? Unveiling the Semantic Vocabulary of Deepfake Detectors 2026-07-08
Accep...

Accepted at CVPRW 2026

Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies 2026-07-08
Invit...

Invited survey paper. 10 pages, 5 figures, 2 tables

Reduced NEXI protocol for the quantification of human gray matter microstructure on the Connectome 2.0 scanner 2026-07-07
Submi...

Submitted to Imaging Neuroscience. This all-in-one version includes supplementary materials. 34 pages, 145 figures, 4 tables

Explainable embeddings with Distance Explainer 2026-07-07
21 pa...

21 pages, 12 figures. Accepted to the 4th World Conference on eXplainable Artificial Intelligence. Method implementation: https://research-software-directory.org/software/distance-explainer

Measuring What Matters: A Unified Evaluation Framework for GNN Explainability 2026-07-06
Dynamic Interest Rate Discovery in Decentralized Finance: A Reverse Kelly Automated Market Maker for Risk-Adjusted Lending 2026-07-05
Explainable AI for Screening Abuse-Related Trauma in Bangladeshi Children: A Training-Free Multimodal Framework Evaluated on Noise-Aware Synthetic Data 2026-07-04 6 pages, 5 figures
Activation-Deactivation: A General Framework for Robust Post-hoc Explainable AI 2026-07-04
Prepr...

Preprint: Under Review; Updated experiments & Figures

Better Together? The Role of Explanations in Supporting Novices in Individual and Collective Deliberations about AI 2026-07-03
30 pa...

30 pages main text, 8 figures, 4 tables. Supplementary material is included in the appendix

Algebraic Model Counting for Global Analysis of Optimal Decision Trees 2026-07-02
Proc....

Proc. Joint European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML-PKDD 2026), LNCS, Naples, Italy, 7-11 September 2026

Quantum-Inspired Vision: Leveraging Wave-Particle Duality for Low-Illumination Enhancement 2026-07-02
Reducing Labeling Effort in Architecture Technical Debt Detection through Active Learning and Explainable AI 2026-07-01
Accep...

Accepted for publication in Empirical Software Engineering (EMSE) Journal, 2026

PACE: A Neuro-Symbolic Framework for Plausible and Actionable Counterfactual Explanations 2026-07-01
CPG-PAD: Concept-Informed Prompts Guided Presentation Attack Detection 2026-07-01
Accep...

Accepted by IEEE Transactions on Information Forensics & Security (TIFS)

Understanding Large Language Models 2026-07-01 25 pages, 1 figure
Explainable AI for Cancer Drug Response Prediction: Beyond Univariate Feature Attributions 2026-07-01
FedXDS: Leveraging Model Attribution Methods to counteract Data Heterogeneity in Federated Learning 2026-06-30
Scientific Explanations in Health Sciences: Causality, Trust, and Epistemic Adequacy 2026-06-30
Explainable AI in Speaker Recognition -- Making Latent Representations Understandable 2026-06-29 A working paper
Feature-level Interaction Explanations in Multimodal Transformers 2026-06-28
Understanding LLM Intervention Explanations in Multi-Party Human-Robot Interaction 2026-06-28
Accep...

Accepted for 2026 36th IEEE International Conference on Robot and Human Interactive Communication

Low-cost concept-based localized explanations: How far can we get with training-free approaches? 2026-06-27
6 pag...

6 pages, 2 figures, 4 tables. Accepted at the 2026 IEEE International Conference on Artificial Intelligence (CAI), 8-10 May 2026, Granada, Spain. Code: https://github.com/darianfgUgr/CoNa

Mechanistic Interpretability

Title Date Comment
CircuitKIT : Circuit Discovery, Evaluation, and Application Toolkit for Mechanistic Interpretability 2026-07-21
CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs 2026-07-21
9 pag...

9 pages, 7 figures, 1 table. Code: https://github.com/LLM-Interp/CLT-Forge. Demonstration video: https://youtu.be/6ptrrLawTl8

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures 2026-07-21
This ...

This manuscript is a 80-page hybrid survey and conceptual framework on LLM evaluation and AI-safety failures. It includes 8 figures and multiple evidence-synthesis tables, covering literature from 2018 to 2026. The paper introduces the EvalSafetyGap framework and reports a structured audit of 10 LLMs. It is submitted as a review/survey article and is not currently under consideration elsewhere

Breaking the Block: Preserving Data Continuity to Train Superior SAEs for Instruct Models 2026-07-20
Measuring Monosemanticity in Sparse Autoencoders via Latent Activation Coherence 2026-07-20
This ...

This is a preprint version. A shorter version of this paper has been accepted for presentation and publication in the post-workshop proceedings of the 8th International Workshop on eXplainable Knowledge Discovery in Data Mining (XKDD 2026), co-located with ECML PKDD 2026. The appendix is included only in this preprint and is not part of the peer-reviewed proceedings paper

Every Component is a Lookup: Token Attribution and Composition from a Single Decomposition 2026-07-19
What Do They See? Interpreting Complex Road Scenarios Through the Eyes of Vision-Language-Action Models for Safe and Trustworthy Autonomous Vehicle Learning 2026-07-18
Are Arithmetic Heuristic Neurons Form-Invariant? A Mechanistic Analysis of Symbols, Text, and Code in LLMs 2026-07-18 Under Review
Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy 2026-07-17
Prepr...

Preprint. Under review

Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability 2026-07-17
Prepr...

Preprint. Under review

Hide and Seek in Embedding Space: Geometry-based Steganography and Detection in Large Language Models 2026-07-17
Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control 2026-07-16
Transcoders for Investigating Deception in Language Models 2026-07-16
Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions 2026-07-15
From Observation to Insight: Mechanistic World Models and the Quest for Autonomous Discovery 2026-07-15
Hierarchical Latent Structures in Data Generation Process Unify Mechanistic Phenomena across Scale 2026-07-14
Same Compression Principle, Different Geometry: Rate-Distortion Signatures Dissociate Biological and Artificial Visual Systems 2026-07-13
Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias 2026-07-13
58 pa...

58 pages, 13 figures, 30 tables; project page: https://xzx34.github.io/unfair-judge/

Emergent Hierarchical Monosemantic Neurons from the Group-Contrastive Forward-Forward Algorithm 2026-07-13
Which Neurons Detect Malicious Code? A Probing Study of LLM Security Knowledge 2026-07-11
The p...

The paper has been peer reviewed and accepted for publication in the 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)

MLPs are Hebbians: Constructing Efficient Fact-Storing MLPs for Transformers 2026-07-10
XAI and Statistical Analysis for Reliable Intrusion Detection in the UAVIDS-2025 Dataset: From Tree to Hybrid and Tabular DNN Ensembles 2026-07-10
Accep...

Accepted at IEEE CITS 2026, Greece

When Structured Sparse Autoencoders Learn Consistent Concepts Across Modalities 2026-07-09
Cross-seed explainability using Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoders 2026-07-09
17 pa...

17 pages, 4 figures, 6 tables

Towards Isolated Interventions via Almost Orthogonal Features in Language Models 2026-07-09
Accep...

Accepted as a conference paper at the Conference on Language Modeling (COLM) 2026

Certified Interventional Fidelity: Anytime-Valid, Adaptive Evaluation of Causal Claims in Mechanistic Interpretability 2026-07-09
Accep...

Accepted at UAI 2026 (Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence). Code: https://github.com/AsiaeeLab/certified-interventional-fidelity

Diagnosing Shape-Prior Shortcuts in Long-Range Single-Shot Fringe Projection Profilometry 2026-07-09 21 pages, 13 figures
Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs 2026-07-08
Temporal Preference Concepts and their Functions in a Large Language Model 2026-07-08
How Learning Dynamics Drive Adversarially Robust Generalization? 2026-07-08
Accep...

Accepted at the 42nd Conference on Uncertainty in Artificial Intelligence (UAI 2026)

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning 2026-07-08 20 pages
Latent Programming Horizons in Coding Agents 2026-07-06
Beyond the Black Box: Interpretability of Agentic AI Tool Use 2026-07-05
12 pa...

12 pages, 4 figures, 17 tables

Interpretability and Generalization Bounds for Learning Spatial Physics 2026-07-05
To ap...

To appear in ICML 2026. 18 pages, 13 figures

Cultural Binding Heads in Language Models 2026-07-04
Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models 2026-07-03
Individual Parameters in Weight-Sparse Transformers Appear Interpretable 2026-07-03
20 pa...

20 pages, 19 figures, 3 tables. Project website: https://weightpedia.org/individual-parameters-in-sparse-transformers/

MatPhaseBench: A Semantics-Guided Benchmark for Materials Phase Diagrams Understanding 2026-07-03
Induction Heads Interpolate N-Grams 2026-07-02
Publi...

Published as a conference paper at ICML 2026. OpenReview: https://openreview.net/forum?id=BSY7jhBxM1

Towards Robustness against Typographic Attack with Training-free Concept Localization 2026-07-02
15 pa...

15 pages main text, provisionally accepted to ECCV 2026

Conditional Co-Ablation: Recovering Self-Repair Backups in Transformer Circuits 2026-07-02
Shared Semantics, Divergent Mechanisms: Unsupervised Feature Discovery by Aligning Semantics and Mechanisms 2026-07-02
40 pa...

40 pages; accepted as an ICML 2026 Spotlight; project page: https://merenova.github.io/distribution-level-feature-discovery/

Expander Sparse Autoencoders: Parameter-Efficient Dictionaries for Mechanistic Interpretability 2026-07-02
Mechanistic Interpretability and Causal Feature Steering of Neural Quantum States via Sparse Autoencoders 2026-07-01
15 pa...

15 pages, 7 figures. Comments welcome!

Muon as a Residual Connection 2026-07-01
Interpreting Global Perturbation Robustness of Image Models using Axiomatic Spectral Importance Decomposition 2026-07-01
Accep...

Accepted by Transactions on Machine Learning Research (TMLR 2024)

MetaOthello: A Controlled Study of Multiple World Models in Transformers 2026-07-01
Camer...

Camera-ready version. Accepted to the 43rd International Conference on Machine Learning (ICML 2026)

Representation as a Bottleneck for Mechanistic Interpretability: The Manifestation Unit Protocol 2026-06-30
65 pa...

65 pages. Interactive demos: https://manifestation-xai.github.io/manifestation-transformers/ , https://manifestation-xai.github.io/manifestation-cnn

Surrogate Fidelity: When Can Open LLMs Explain Closed Ones? 2026-06-30

Metadata

Metadata

Assignees

No one assigned

    Labels

    documentationImprovements or additions to documentation

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions