Skip to content

Latest 50 Papers - July 30, 2026 #129

Description

@github-actions

Please check the Github page for a better reading experience and more papers.

LLM Reasoning

Title Date Comment
Detecting Knowledge Inconsistencies Across Text, Tables, and Knowledge Graphs 2026-07-28
A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series 2026-07-28
MARS: Multi-Agent Re-ranking for Repeat-Order Food Delivery Recommendation 2026-07-28
TrajAudit: Automated Failure Diagnosis for Agentic Coding Systems 2026-07-28
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering 2026-07-28
Accep...

Accepted to the ACL 2026 Main

Hint-Guided Diversified Policy Optimization for LLM Reasoning 2026-07-27
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks 2026-07-27
Accep...

Accepted at ICML 2026

MARS: Multi-hop Adaptive Retrieval and SPARQL Generation for KGQA 2026-07-27
EKAW ...

EKAW 2026 (https://ekaw2026.di.unito.it/accepted-posters-and-demos)

HydroAgent: Formalizing Forecaster Expertise into Skill-Orchestrated Flood Forecasting Workflows 2026-07-27
Reasoning or Memorization: Can LLMs Understand and Generate Chinese Xiehouyu Riddles? 2026-07-26
20 pa...

20 pages; exp 3 is work in progress

Traceable LLM Reasoning for Fake-Order Fraud Detection 2026-07-25 15 pages, 3 figures
Not All LLM Reasoning is Visible in the Chain-of-Thought 2026-07-24
Statistical Early Stopping for Reasoning Models 2026-07-24
Math Education Digital Shadows for Investigating Learning with GenAI: Mathematics Performance, Anxiety, and Confidence in LLMs 2026-07-24
Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning 2026-07-24
Silent Failures in Quantized LLM Reasoning: A Taxonomy-Based Analysis of Hollow Convergence and Failure Mode Shifts 2026-07-23
Out of Sight, Still in Mind: Token Compression for Omni-LLMs 2026-07-23 Preprint
From Noise to Diversity: Random Embedding Injection in LLM Reasoning 2026-07-23
Under...

Under review, ICML 2026 Mechanistic Interpretability Workshop (Spotlight)

WaveformQA: Benchmarking LLM Temporal Reasoning on Digital Waveforms 2026-07-22
10 pa...

10 pages; abridged version published in IEEE International Conference on LLM-Aided Design (ICLAD), 2026

Experience Augmented Policy Optimization for LLM Reasoning 2026-07-22
Reference-Free Evaluation of Reasoning in Open-Ended Question Answering 2026-07-22
They'll Verify. They Just Won't Act. How Authority Framing and Laundered Code Turn a Trusted Agentic CI/CD Pipeline Into an Attack Surface 2026-07-21
9 pag...

9 pages. Dataset and reproduction code: https://github.com/senthex-security/senthex-research

Reasoning Error from Known Fact: Step-Level Self-Consistency Group Relative Policy Optimization for LLM 2026-07-21
Tracing the Shadows: Automatic Tracking and Analysis of Crypto Money Laundering via Transaction Semantic Analysis 2026-07-21
Not All Errors Are Created Equal: ASCoT Addresses Late-Stage Fragility in Efficient LLM Reasoning 2026-07-21
LLM-Grounded Explainable AI for Supply Chain Risk Early Warning via Temporal Graph Attention Networks 2026-07-21
QuArch: A Benchmark for Evaluating LLM Reasoning in Computer Architecture 2026-07-21
Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration 2026-07-20
13 pa...

13 pages, 2 figures, 8 tables. Code: https://github.com/AmGarfield/OracleGap

LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning 2026-07-20
AIGB-R1: Self-Evolving Generative Auto-Bidding via Hierarchical Planner-Executor Optimization 2026-07-19
Debate-on-Graph: Reliable and Adaptive Reasoning of Large Language Model on Uncertain Knowledge Graph 2026-07-19
18 pa...

18 pages, Accepted by ECML-PKDD 2026

Constrained Path Reasoning: Measuring When Committed Stages Earn Their Cost 2026-07-19
Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM Reasoning 2026-07-19
Logic-Guided Socially-aware Robot Navigation World Model 2026-07-18
The a...

The authors have decided to withdraw this manuscript due to concerns regarding its current scope, framing, and presentation. Please do not cite this version

NOWJ@COLIEE 2026: Adaptive Pipelines for Legal Retrieval and Reasoning 2026-07-18
Prese...

Presented at COLIEE 2026

Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously 2026-07-17
Accep...

Accepted by ECCV 2026, project page https://1ranguan.github.io/VST/

REST: Receding Horizon Explorative Steiner Tree for Zero-Shot Object-Goal Navigation 2026-07-16 Accepted to IROS'26
ProfMalPlus: Agent-Coordinated Detection of Malicious NPM Packages via Static-Dynamic Analysis Synergy 2026-07-15
Policy of Thoughts: Scaling Test-Time Training for LLM Reasoning via Online Policy Evolution 2026-07-15
Under...

Under Review, preprint

Assert, don't describe: Linguistic features that shift LLM reasoning about animal welfare 2026-07-12
Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction 2026-07-11
32 pa...

32 pages, 8 figures, 10 tables

Can AI Reason Like an Urban Planner? Benchmarking Large Language Models Against Professional Judgment 2026-07-11
This ...

This paper has been withdrawn by the authors because the current version requires substantial revision and further validation before it can be considered a reliable representation of the work

Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization 2026-07-11 ICML 2026
EmoStyle: Affective Conditioning of Style-Specialist Experts for Emotional Image Generation 2026-07-11
Using LLMs to Adjudicate Static-Analysis Alerts with Error Reduction Techniques 2026-07-10 22 pages
Hierarchical Chain-of-Thought: Enhancing LLM Reasoning Performance and Efficiency 2026-07-10
Practical Source Code Recovery from Binary Functions Using Anchor-Based Retrieval and LLM Reasoning 2026-07-10 12 pages, 5 figures
RepLLM: Toward Automatically Reproducing Network Research Results 2026-07-09
SIGCO...

SIGCOMM'26(19 pages, 6 figures, 6 tables)

Chain of Thought

Title Date Comment
Penelope: Localized Latent Recurrence for Efficient Structured Reasoning 2026-07-28 8 pages, 2 figures
Towards Understanding the Cognitive Habits of Large Reasoning Models 2026-07-28
Publi...

Published at Machine Intelligence Research vol.23, no.4, pp.873-886

CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model 2026-07-28
22 pa...

22 pages, 2 figures, 20 tables. Code at https://github.com/BrainJellyPie/CoTinyVLA

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought 2026-07-28
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation 2026-07-28
TabRank: Chain-of-Thought Distillation for Table Re-Rankers 2026-07-28 8 pages, 3 figures
LaRec: Unleashing LLM-based Latent Reasoning for Generative Recommendation 2026-07-27
CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding 2026-07-27
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems 2026-07-27 version 1.3
Qwen-Music Technical Report 2026-07-27
Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding 2026-07-27 Technical Report
What Does Chain-of-Thought Contribute at Probe Time? Evidence for Local Co-Occurrence Activation 2026-07-27
DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning 2026-07-27 16 pages, 7 figures
SLPO: Scaling Latent Reasoning via a Surrogate Policy 2026-07-27
Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers 2026-07-26
CoopReflect: Towards Natural Language Communication for Cooperative Autonomous Driving via Multi-Agent Learning 2026-07-26
Training Language Models to Cooperate with Inference-Time Controllers 2026-07-26
LLM-Ideoplasticity: Measuring Ideological Plasticity in the Political Behavior of LLMs as a Context-Conditioned Distribution 2026-07-26
Under...

Under review, 40 pages, 18 figures, 11 tables

Guiding Language Models to Be More Empathetic: Culturally Sensitive Mental Health Advice Generation Through Human-LLM Collaboration 2026-07-26
Two Regimes of Chain-of-Thought Unfaithfulness: Behavioral Detection Fails Where Models Are Wrong 2026-07-26 14 pages, 6 figures
EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization 2026-07-25
SymStep: Symbolic Step Verification for Logical Reasoning 2026-07-25
SAFE-MEME: Structured Reasoning Framework for Robust Hate Speech Detection in Memes 2026-07-25
44 pa...

44 pages, 27 figures, 6 tables

Reason Popper-ly: Patching In-Context Reasoning with Inductive Logic Programming 2026-07-25
Accep...

Accepted at the 20th Conference on Neurosymbolic Learning and Reasoning

WCM: World-Cognition Model for Generalizable Human-Robot Interaction 2026-07-25
Not All LLM Reasoning is Visible in the Chain-of-Thought 2026-07-24
How Well Can AI Generate Backlogs from App Mockups? 2026-07-24
Accep...

Accepted at the AIRE Workshop, IEEE 34th International Requirements Engineering Conference (RE) 2026

Spatial-IQ: Deconstructing Spatial Intelligence via Hierarchical Capability Tests 2026-07-24
Language-Routed RAG and Direct Option Scoring for Multilingual Financial QA: DS@GT at FinMMEval 2026-07-24
Accep...

Accepted for publication in the CLEF 2026 Working Notes. 15 pages, 3 figures

RecGPT-V3 Technical Report 2026-07-24 Technique Report
Benchmarking Fine-tuning and Retrieval Strategies for a Multimodal Language Model on the NRC Reactor Operator Licensing Examination 2026-07-24
EVL-MCoT: Enhanced Vision-Language Multi-CoT for Harmful Meme Detection 2026-07-24
Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning 2026-07-24
Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning 2026-07-24
J-CoT: Chain-of-Thought in J-Space 2026-07-24 work in progress
PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation 2026-07-24
Accep...

Accepted to CVPR 2026 (highlight)

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning 2026-07-24
Accep...

Accepted at ICML 2026; previously accepted as a non-archival paper at the Efficient Reasoning Workshop at NeurIPS 2025

LeAct: Learning to Reason from Expert Actions 2026-07-23
27 pa...

27 pages, 3 figures, 11 tables

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization 2026-07-23
18 pa...

18 pages, 7 figures, 6 tables. Accepted at COLM 2026

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment 2026-07-23
Token Budget Saturation and Mechanistic Early Detection of Reasoning Non-Convergence in Chain-of-Thought Models 2026-07-23
When Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time Moderation 2026-07-23
8 pag...

8 pages, 6 figures, 3 tables. Project page: https://ndb796.github.io/ResponseGuard ; Code: https://github.com/ndb796/ResponseGuard

Silent Failures in Quantized LLM Reasoning: A Taxonomy-Based Analysis of Hollow Convergence and Failure Mode Shifts 2026-07-23
ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues 2026-07-23
LatentChem: From Textual CoT to Latent Thinking in Chemical Reasoning 2026-07-23
Accep...

Accepted at ICML 2026

Post-Hoc Reasoning in Chain of Thought: Decoding and Steering Pre-Committed Answers 2026-07-23
Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad 2026-07-23 16 pages, 6 figures
Tight Sample Complexity of Transformers 2026-07-23 in COLT 2026
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI 2026-07-23
The p...

The paper is available on the project website: https://from-chatbot-to-digital-colleague.github.io/

LLM Interpretability

Title Date Comment
CogEEGAgent: Toward Autonomous Cognitive EEG Analysis with Grounded Execution and Selection-Aware Verification 2026-07-27
16 pa...

16 pages, 5 figures, and 16 tables. The supplementary material is included in the same PDF

AlphaRoute: Large Language Models as Semantic Optimizers for Multi-Objective Routing 2026-07-22
7 pag...

7 pages, 5 figures. Accepted for publication in the IEEE International Conference on LLM-Aided Design, 2026, Stanford University, Stanford, CA, USA. Code available at https://github.com/Kcbir/AlphaRoute

Understanding Interpretation Difficulty in Harmful Online Communication: Insights from Cybercrime Communities 2026-07-08
NeuraDock Visual Cognitive Load Agent Tutorial: A Quality-Gated Open-Source EEG Workflow for Alpha Dynamics and Real-Time Applications 2026-06-25 22 pages, 10 figures
Don't Go Breaking My LLM: The Impact of Pruning Attention Layers on Explanation Faithfulness and Confidence Calibration 2026-06-23 Accepted at TMLR
LLM-Based Generalizable Hierarchical Task Planning and Execution for Heterogeneous Robot Teams with Event-Driven Replanning 2026-06-18
full ...

full version of this short paper is accepted at Frontiers in Robotics and AI Journal

ICA Lens: Interpreting Language Models Without Training Another Dictionary 2026-06-10 Ongoing Project
Language-Driven Cost Optimization for Autonomous Driving 2026-06-09
Paper...

Paper accepted at IEEE Intelligent Transportation Systems Conference (ITSC) 2026

ReflectiChain: Epistemic Grounding in LLM-Driven World Models for Supply Chain Resilience 2026-06-09
MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding 2026-06-08
Analyzing the Correlation Between Hallucinations and Knowledge Conflicts in Large Language Models 2026-06-07
Thinking Through Signs: PEEL as a Semiotic Scaffolding for Epistemically Accountable AI-Enabled Research 2026-06-02 10 pages, 5 figuras
Unveiling the Limits of Large Language Models in Inferring Pragmatic Meaning from Non-Verbal Responses 2026-06-01
The Limits of LLM Forecasting: Parametric Knowledge Gaps Across Conflict Zones 2026-05-29
Hallucination Behavior in Multimodal LLMs Across Agricultural Image Interpretation and Generation Tasks 2026-05-26
How Much Structure Do LLMs Need? Evaluating LLMs for Bibliometric Cluster Description 2026-05-23
Improving Labeling Consistency with Detailed Constitutional Definitions and AI-Driven Evaluation 2026-05-22
Under...

Under review at ACL Rolling Review (ARR), May 2026 cycle. Also available at https://doi.org/10.5281/zenodo.20125267

Patch-Effect Graph Kernels for LLM Interpretability 2026-05-07
LLM-ADAM: A Generalizable LLM Agent Framework for Pre-Print Anomaly Detection in Additive Manufacturing 2026-05-05 24 pages, 10 figures
From Research Question to Scientific Workflow: Leveraging Agentic AI for Science Automation 2026-04-23
Agentic AI-Enabled Framework for Thermal Comfort and Building Energy Assessment in Tropical Urban Neighborhoods 2026-04-23
Accep...

Accepted at IAQVEC 2026

Learning to Draw ASCII Improves Spatial Reasoning in Language Models 2026-04-16
Arbitration Failure, Not Perceptual Blindness: How Vision-Language Models Resolve Visual-Linguistic Conflicts 2026-04-13
Heterogeneous Graph Importance Scoring and Clustering with Automated LLM-based Interpretation 2026-04-09
26 pa...

26 pages, 11 figures, 8 tables

Runtime Execution Traces Guided Automated Program Repair with Multi-Agent Debate 2026-04-03
12 pa...

12 pages, 4 figures, 8 tables

Discovering Decoupled Functional Modules in Large Language Models 2026-03-18 AAAI-26 Oral
LaMoGen: Language to Motion Generation Through LLM-Guided Symbolic Inference 2026-03-12
Accep...

Accepted by CVPR 2026. Supplementary material included. Project page: https://jjkislele.github.io/LaMoGen/

Pneuma-Seeker: A Relational Reification Mechanism to Align AI Agents with Human Work over Relational Data 2026-03-11
PhysLLM: Harnessing Large Language Models for Cross-Modal Remote Physiological Sensing 2026-03-05
Accep...

Accepted by International Conference on Learning Representations (ICLR) 2026

Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs 2026-03-01
Sparse Shift Autoencoders for Identifying Concepts from Large Language Model Activations 2026-02-27 27 pages, 9 figures
Co-Disclosing the Computer: LLM-Mediated Computing through Reflective Conversation 2026-02-27 CHI'26
Neuro-Symbolic Control with Large Language Models for Language-Guided Spatial Tasks 2026-02-21
Stroke Lesions as a Rosetta Stone for Language Model Interpretability 2026-02-03 45 pages, 17 figures
Uncertainty-Aware Gradient Signal-to-Noise Data Selection for Instruction Tuning 2026-01-20 Preprint
A Shared Geometry of Difficulty in Multilingual Language Models 2026-01-19
Causal-SAM-LLM: Large Language Models as Causal Reasoners for Robust Medical Segmentation 2026-01-16
Accep...

Accepted by IEEE ICASSP 2026

Self-reflection in Automated Qualitative Coding: Improving Text Annotation through Secondary LLM Critique 2026-01-14
Word Synchronization Challenge: A Benchmark for Word Association Responses for Large Language Models 2026-01-14
Can LLMs interpret figurative language as humans do?: surface-level vs representational similarity 2026-01-14 17 pages, 5 figures
Neuro-Symbolic Compliance: Integrating LLMs and SMT Solvers for Automated Financial Legal Analysis 2026-01-07
10 pa...

10 pages, 6 tables, 3 figures, accepted by the 2nd ACM AIware Conference

LLM Interpretability with Identifiable Temporal-Instantaneous Representation 2026-01-02 NeurIPS 2025
Trust Me, I Know This Function: Hijacking LLM Static Analysis using Bias 2025-12-18
Verification-Guided Context Optimization for Tool Calling via Hierarchical LLMs-as-Editors 2025-12-15
Accep...

Accepted by AAAI 2026 Workshop on Agentic AI Benchmarks and Applications for Enterprise Tasks

Visualizing token importance for black-box language models 2025-12-12
SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language Models 2025-12-01
ACL 2...

ACL 2025 main. Our Speech-IQ leaderboard is hosted at huggingface.co/spaces/nvidia/Speech-IQ-leaderboard. Speech-IQ Calculator: https://github.com/YukinoWan/SpeechIQ

For Those Who May Find Themselves on the Red Team 2025-11-23
HSKBenchmark: Modeling and Benchmarking Chinese Second Language Acquisition in Large Language Models through Curriculum Tuning 2025-11-19
Accep...

Accepted by AAAI-2026

KnowThyself: An Agentic Assistant for LLM Interpretability 2025-11-05
5 pag...

5 pages, 1 figure, Accepted for publication at the Demonstration Track of the 40th AAAI Conference on Artificial Intelligence (AAAI 26)

Imperfect Language, Artificial Intelligence, and the Human Mind: An Interdisciplinary Approach to Linguistic Errors in Native Spanish Speakers 2025-11-03 12 pages, 3 figures

Explainable AI

Title Date Comment
On the Design and Evaluation of Human-centered Explainable AI Systems: A Systematic Review and Taxonomy 2026-07-28
From Dyad to Triad: Eliciting XAI Requirements in Stroke Rehabilitation 2026-07-28
AI an...

AI and Cognitive Computing for Trustworthy Human Machine Systems Session, IEEE International Conference on Systems, Man, and Cybernetics (IEEE SMC 2026)

Explainable AI for Chronic Kidney Disease Prediction Using Simulated Federated Learning 2026-07-28 13 Pages, 5 Figures
A Unified Framework for Uncertainty-Aware Explainable Artificial Intelligence: A Case Study in Power Quality Disturbance Classification 2026-07-28
JobMatchAI-An Intelligent Job Matching Platform Using Knowledge Graphs, Semantic Search and Explainable AI 2026-07-27
Evaluating the Impact of Explainable AI on Trust in AI-Assisted Code Review 2026-07-27
23 pa...

23 pages, 4 figures, 5 tables. To appear in Proceedings of the ACM on Software Engineering (PACMSE), Vol. 3, No. ISSTA, Article ISSTA093 (ISSTA 2026). Published under CC BY 4.0. Replication package: https://doi.org/10.5281/zenodo.21457282

Beyond Local Inspection: Global, Guideline-Grounded Evaluation of Post-hoc XAI Methods for ECG Classification 2026-07-27
Disentangling Acoustic Cues in Alzheimer's Pathology and Perception: The Roles of Language and Gender 2026-07-27
Accep...

Accepted at Interspeech 2026

Explainable AI through the Lens of Material Agency: Enabling Musical Interface Design with Neural Audio Models 2026-07-25
Under...

Under review for "Explainable AI for the Arts" (N. Bryan-Kinns, Ed.), Springer

Context-Aware Concept Distillation for Trustworthy Flood Prediction 2026-07-25
to be...

to be published in IJCAI 2026 proceedings

Unboxing Diffusion Models for the Arts: Interactive Model Bending and Practice-Based Explainability 2026-07-24
Under...

Under review for "Explainable AI for the Arts" (N. Bryan-Kinns, Ed.), Springer

Scalable Explainability-as-a-Service (XaaS) for Edge AI Systems 2026-07-24
8 pag...

8 pages, 5 figures, 2 tables. This version updates metadata after publication in IEEE Xplore and publication by SoutheastCon 2026

Proceedings of The Fourth International Workshop on eXplainable AI for the Arts (XAIxArts 4) 2026-07-22
Scaling Time Series Classification via XAI-Driven Data Reduction 2026-07-22
Accep...

Accepted for AALTD workshop at ECML-PKDD 2026

LLM-Grounded Explainable AI for Supply Chain Risk Early Warning via Temporal Graph Attention Networks 2026-07-21
Automated Data Engineering and Feature Selection for the Case Study of Warpage Detection in Fused Deposition Modeling 2026-07-20
Explaining and Tuning Transformer-based LLMs in Arithmetic Tasks with Human Strategies 2026-07-19
FST.ai 2.5: Explainable and Uncertainty-Aware AI for Olympic and Para-Taekwondo Decision Support, Athlete Digital Twins, and Federation-Scale Analytics 2026-07-18 14 pages, 4 figures
FAIR_XAI: Improving Multimodal Foundation Model Fairness via Explainability for Wellbeing Assessment 2026-07-17 11 pages, 4 figures
VeriX-Anon: A Multi-Layered Framework for Mathematically Verifiable Outsourced Target-Driven Data Anonymization 2026-07-17
v2: r...

v2: revised after peer review. Evaluation expanded from 3 to 7 datasets, per-dataset Wasserstein-threshold calibration added, effect-size CIs and cross-dataset statistics reported, and analytical zk-SNARK/MPC/TEE baselines added. Minor errors corrected

Automated identification of Ichneumonoidea wasps via YOLO-based deep learning: Integrating HiresCam for Explainable AI 2026-07-16 15 pages, 20 figures
Explaining Process Control Optimisation Recommendations via GradientSHAP and Implicit Differentiation 2026-07-16
Diagnosing and Mitigating Domain Shift in Permission-Based Android Malware Detection 2026-07-15
"Trust Junk" Leads to Unjustified Support for Highly Discriminatory Predictive Models 2026-07-14
MLPTR-CC: Multi-label Pathology Test Recommendation using Classifier Chains and SHAP 2026-07-14
Atomic Units of X: The Compression Layer of Intelligence 2026-07-14
Evaluating RE Practices for Explainability: Synthesizing Insights from Daimler Truck into an Explainable RE Framework Proposal 2026-07-13
Detecting Explanatory Insufficiency in Learned Representations: A Framework for Representational Vigilance 2026-07-13
25 pa...

25 pages, 2 figures, no tables, 16 references. Conceptual and methodological framework for monitoring representational adequacy and detecting explanatory insufficiency in learned representations

ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI 2026-07-10
Knowledge Graphs and Explainable AI as Complementary Resources for Urban Mining 2026-07-10
Accep...

Accepted for presentation at the AISE Workshop @ IJCAI-ECAI 2026

All Explanations are Wrong, But Many Are Useful: Exploring the Rashomon Explanation Set with Large Language Models 2026-07-10
Application of machine learning to monster level prediction in tabletop RPG game design 2026-07-10
Unlearning to Protect: A Distilled Reinforcement Learning Framework with Privacy-Preserving Feature Unlearning and XAI for IoT Security 2026-07-09
The Contribution of XAI for the Safe Development and Certification of AI: An Expert-Based Analysis 2026-07-09
This ...

This preprint has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this contribution is published in Artificial Intelligence in HCI (HCII 2026), Lecture Notes in Computer Science, vol. 16745, and is available online at https://doi.org/10.1007/978-3-032-30849-8_13

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning 2026-07-08 20 pages
Why Fake ? Unveiling the Semantic Vocabulary of Deepfake Detectors 2026-07-08
Accep...

Accepted at CVPRW 2026

Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies 2026-07-08
Invit...

Invited survey paper. 10 pages, 5 figures, 2 tables

Reduced NEXI protocol for the quantification of human gray matter microstructure on the Connectome 2.0 scanner 2026-07-07
Submi...

Submitted to Imaging Neuroscience. This all-in-one version includes supplementary materials. 34 pages, 145 figures, 4 tables

Explainable embeddings with Distance Explainer 2026-07-07
21 pa...

21 pages, 12 figures. Accepted to the 4th World Conference on eXplainable Artificial Intelligence. Method implementation: https://research-software-directory.org/software/distance-explainer

Measuring What Matters: A Unified Evaluation Framework for GNN Explainability 2026-07-06
Dynamic Interest Rate Discovery in Decentralized Finance: A Reverse Kelly Automated Market Maker for Risk-Adjusted Lending 2026-07-05
Explainable AI for Screening Abuse-Related Trauma in Bangladeshi Children: A Training-Free Multimodal Framework Evaluated on Noise-Aware Synthetic Data 2026-07-04 6 pages, 5 figures
Activation-Deactivation: A General Framework for Robust Post-hoc Explainable AI 2026-07-04
Prepr...

Preprint: Under Review; Updated experiments & Figures

Better Together? The Role of Explanations in Supporting Novices in Individual and Collective Deliberations about AI 2026-07-03
30 pa...

30 pages main text, 8 figures, 4 tables. Supplementary material is included in the appendix

Algebraic Model Counting for Global Analysis of Optimal Decision Trees 2026-07-02
Proc....

Proc. Joint European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML-PKDD 2026), LNCS, Naples, Italy, 7-11 September 2026

Quantum-Inspired Vision: Leveraging Wave-Particle Duality for Low-Illumination Enhancement 2026-07-02

Mechanistic Interpretability

Title Date Comment
Phase Structure in Rotary Attention: A Spectral Framework for Semantic Continuity and Execution-Boundary Governance 2026-07-28
14 pa...

14 pages; theoretical framework and proposed experimental program

Emergent Latent-State Computation under Stochastic Volatility 2026-07-28
EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures 2026-07-27
74 pa...

74 pages, 2 figures, 4 tables. Hybrid systematic survey and conceptual framework on LLM evaluation and AI-safety failures, synthesizing 373 primary studies (2018-2026). Introduces the EvalSafetyGap framework (Instability Decomposition, Alignment Trilemma) and reports an exploratory ten-model audit. Submitted as a review/survey article; not currently under consideration elsewhere

Do LLMs Know Their Vulnerable Scenarios? 2026-07-26
16 pa...

16 pages, 10 Figures, Under Review

Continuous surrogates versus threshold Boolean networks for modeling Arabidopsis ISR gene regulation 2026-07-25
To be...

To be published in IEEE CIBCB 2026

Towards Isolated Interventions via Almost Orthogonal Features in Language Models 2026-07-25
Accep...

Accepted as a conference paper at the Conference on Language Modeling (COLM) 2026

Through the Bottleneck: How Multi-head Latent Attention Separates Content from Position in Language Models 2026-07-25
Toward Mechanistic Interpretability of an AI Foundation Model Fine-Tuned for Atmospheric Chemistry 2026-07-22
CircuitKIT : Circuit Discovery, Evaluation, and Application Toolkit for Mechanistic Interpretability 2026-07-21
CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs 2026-07-21
9 pag...

9 pages, 7 figures, 1 table. Code: https://github.com/LLM-Interp/CLT-Forge. Demonstration video: https://youtu.be/6ptrrLawTl8

Breaking the Block: Preserving Data Continuity to Train Superior SAEs for Instruct Models 2026-07-20
Measuring Monosemanticity in Sparse Autoencoders via Latent Activation Coherence 2026-07-20
This ...

This is a preprint version. A shorter version of this paper has been accepted for presentation and publication in the post-workshop proceedings of the 8th International Workshop on eXplainable Knowledge Discovery in Data Mining (XKDD 2026), co-located with ECML PKDD 2026. The appendix is included only in this preprint and is not part of the peer-reviewed proceedings paper

Every Component is a Lookup: Token Attribution and Composition from a Single Decomposition 2026-07-19
What Do They See? Interpreting Complex Road Scenarios Through the Eyes of Vision-Language-Action Models for Safe and Trustworthy Autonomous Vehicle Learning 2026-07-18
Are Arithmetic Heuristic Neurons Form-Invariant? A Mechanistic Analysis of Symbols, Text, and Code in LLMs 2026-07-18 Under Review
Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy 2026-07-17
Prepr...

Preprint. Under review

Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability 2026-07-17
Prepr...

Preprint. Under review

Hide and Seek in Embedding Space: Geometry-based Steganography and Detection in Large Language Models 2026-07-17
Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control 2026-07-16
Transcoders for Investigating Deception in Language Models 2026-07-16
Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions 2026-07-15
From Observation to Insight: Mechanistic World Models and the Quest for Autonomous Discovery 2026-07-15
Hierarchical Latent Structures in Data Generation Process Unify Mechanistic Phenomena across Scale 2026-07-14
Same Compression Principle, Different Geometry: Rate-Distortion Signatures Dissociate Biological and Artificial Visual Systems 2026-07-13
Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias 2026-07-13
58 pa...

58 pages, 13 figures, 30 tables; project page: https://xzx34.github.io/unfair-judge/

Emergent Hierarchical Monosemantic Neurons from the Group-Contrastive Forward-Forward Algorithm 2026-07-13
Which Neurons Detect Malicious Code? A Probing Study of LLM Security Knowledge 2026-07-11
The p...

The paper has been peer reviewed and accepted for publication in the 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)

MLPs are Hebbians: Constructing Efficient Fact-Storing MLPs for Transformers 2026-07-10
XAI and Statistical Analysis for Reliable Intrusion Detection in the UAVIDS-2025 Dataset: From Tree to Hybrid and Tabular DNN Ensembles 2026-07-10
Accep...

Accepted at IEEE CITS 2026, Greece

When Structured Sparse Autoencoders Learn Consistent Concepts Across Modalities 2026-07-09
Cross-seed explainability using Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoders 2026-07-09
17 pa...

17 pages, 4 figures, 6 tables

Certified Interventional Fidelity: Anytime-Valid, Adaptive Evaluation of Causal Claims in Mechanistic Interpretability 2026-07-09
Accep...

Accepted at UAI 2026 (Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence). Code: https://github.com/AsiaeeLab/certified-interventional-fidelity

Diagnosing Shape-Prior Shortcuts in Long-Range Single-Shot Fringe Projection Profilometry 2026-07-09 21 pages, 13 figures
Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs 2026-07-08
Temporal Preference Concepts and their Functions in a Large Language Model 2026-07-08
How Learning Dynamics Drive Adversarially Robust Generalization? 2026-07-08
Accep...

Accepted at the 42nd Conference on Uncertainty in Artificial Intelligence (UAI 2026)

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning 2026-07-08 20 pages
Latent Programming Horizons in Coding Agents 2026-07-06
Beyond the Black Box: Interpretability of Agentic AI Tool Use 2026-07-05
12 pa...

12 pages, 4 figures, 17 tables

Interpretability and Generalization Bounds for Learning Spatial Physics 2026-07-05
To ap...

To appear in ICML 2026. 18 pages, 13 figures

Cultural Binding Heads in Language Models 2026-07-04
Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models 2026-07-03
Individual Parameters in Weight-Sparse Transformers Appear Interpretable 2026-07-03
20 pa...

20 pages, 19 figures, 3 tables. Project website: https://weightpedia.org/individual-parameters-in-sparse-transformers/

MatPhaseBench: A Semantics-Guided Benchmark for Materials Phase Diagrams Understanding 2026-07-03
Induction Heads Interpolate N-Grams 2026-07-02
Publi...

Published as a conference paper at ICML 2026. OpenReview: https://openreview.net/forum?id=BSY7jhBxM1

Towards Robustness against Typographic Attack with Training-free Concept Localization 2026-07-02
15 pa...

15 pages main text, provisionally accepted to ECCV 2026

Conditional Co-Ablation: Recovering Self-Repair Backups in Transformer Circuits 2026-07-02
Shared Semantics, Divergent Mechanisms: Unsupervised Feature Discovery by Aligning Semantics and Mechanisms 2026-07-02
40 pa...

40 pages; accepted as an ICML 2026 Spotlight; project page: https://merenova.github.io/distribution-level-feature-discovery/

Metadata

Metadata

Assignees

No one assigned

    Labels

    documentationImprovements or additions to documentation

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions