Skip to content

Latest 50 Papers - August 06, 2026 #134

Description

@github-actions

Please check the Github page for a better reading experience and more papers.

LLM Reasoning

Title Date Comment
Interpretable Adaptive Sampling for LLM Test-Time Scaling 2026-08-04
EffiHolmes: Differential Profiling-Guided Repository Level Time Inefficiency Fix Localization 2026-08-04
Accep...

Accepted at the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026). 13 pages, 3 figures, and 4 tables

Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Explicit and Latent Thoughts for MLLMs 2026-08-04
Accep...

Accepted by ACM MM 2026. 10 pages, 6 figures, 5 tables

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR 2026-08-04
25 pa...

25 pages, 16 figures, and 9 tables

ConFL: Explainable Concurrent Fault Localization via Hierarchy-Guided LLM Reasoning 2026-08-04
Accep...

Accepted at ISSTA 2026

Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks 2026-08-03 working in progress
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning 2026-08-03
Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey 2026-08-03
Accep...

Accepted for publication in Artificial Intelligence Review

HPFA: Hypergraph-Based Paired Failure Attribution for LLM Reasoning 2026-08-03
TCPO: Turn-Level Credit Policy Optimization 2026-08-03
Cognitive Demand Steering for Adaptive Meta-Reasoning in Large Language Models 2026-08-02 21 pages, 3 figures
The Graph Language: How Knowledge Graphs Speak to Large Language Models 2026-08-02
Accep...

Accepted to ISWC 2025

Cloud-ScPO: Hidden-State Geometry for Semi-Supervised Preference Optimization in LLM Reasoning 2026-08-02
14 pa...

14 pages, 2 figures, 7 tables. Preprint

On the Wings of Imagination: Conflicting Script-based Multi-role Framework for Humor Caption Generation 2026-08-02
Paper...

Paper published as a conference paper at ICLR 2026

TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning 2026-08-01
Accep...

Accepted to Transactions on Machine Learning Research, 2026

Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning 2026-08-01
LLM-Assisted Coalition Formation for Cooperative Perception in Autonomous Driving 2026-08-01
Accep...

Accepted for presentation at IEEE Global Communications Conference (GLOBECOM 2026), Cognitive Radio and AI-Enabled Networks Symposium

RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems 2026-07-31 9 pages, 2 figures
ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning 2026-07-31
24 pa...

24 pages, including 13 pages of main text and 11 pages of appendix

ELISA: An Interpretable Hybrid Generative AI Agent for Expression-Grounded Discovery in Single-Cell Genomics 2026-07-31
Think2Go: Generative Next POI Recommendation with LLM Reasoning 2026-07-31
Accep...

Accepted by KDD 2026 Research Track Cycle 1 (Oral presentation)

BLADE: Boundary-Expanded and Layer-Adaptive Dynamic Exit for Efficient LLM Reasoning 2026-07-31 8 pages
Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation 2026-07-30
AgentMap: Joint Equivalence and Subsumption Discovery for Ontology Matching 2026-07-29 21 pages, 5 figures
Dual-Path LLM Reasoning for Multimodal Few-Shot Knowledge Graph Completion 2026-07-29 10 pages, 4 figures
Detecting Knowledge Inconsistencies Across Text, Tables, and Knowledge Graphs 2026-07-29
LLMAR: A Tuning-Free Recommendation Framework for Sparse and Text-Rich Industrial Domains 2026-07-29 10 pages, 3 figures
Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings? 2026-07-29
Accep...

Accepted to the AID-Wild workshop at CAIS 2026

GoGoTB: Agentic RTL Verification with Specification-Grounded Coverage Closure 2026-07-28
9 pag...

9 pages, 7 figures, 3 tables. Xin Xin and Jincheng Lou contributed equally to this work and share first authorship

A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series 2026-07-28
MARS: Multi-Agent Re-ranking for Repeat-Order Food Delivery Recommendation 2026-07-28
TrajAudit: Automated Failure Diagnosis for Agentic Coding Systems 2026-07-28
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering 2026-07-28
Accep...

Accepted to the ACL 2026 Main

Hint-Guided Diversified Policy Optimization for LLM Reasoning 2026-07-27
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks 2026-07-27
Accep...

Accepted at ICML 2026

MARS: Multi-hop Adaptive Retrieval and SPARQL Generation for KGQA 2026-07-27
EKAW ...

EKAW 2026 (https://ekaw2026.di.unito.it/accepted-posters-and-demos)

HydroAgent: Formalizing Forecaster Expertise into Skill-Orchestrated Flood Forecasting Workflows 2026-07-27
Reasoning or Memorization: Can LLMs Understand and Generate Chinese Xiehouyu Riddles? 2026-07-26
20 pa...

20 pages; exp 3 is work in progress

Traceable LLM Reasoning for Fake-Order Fraud Detection 2026-07-25 15 pages, 3 figures
Not All LLM Reasoning is Visible in the Chain-of-Thought 2026-07-24
Statistical Early Stopping for Reasoning Models 2026-07-24
Math Education Digital Shadows for Investigating Learning with GenAI: Mathematics Performance, Anxiety, and Confidence in LLMs 2026-07-24
Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning 2026-07-24
Silent Failures in Quantized LLM Reasoning: A Taxonomy-Based Analysis of Hollow Convergence and Failure Mode Shifts 2026-07-23
Out of Sight, Still in Mind: Token Compression for Omni-LLMs 2026-07-23 Preprint
From Noise to Diversity: Random Embedding Injection in LLM Reasoning 2026-07-23
Under...

Under review, ICML 2026 Mechanistic Interpretability Workshop (Spotlight)

WaveformQA: Benchmarking LLM Temporal Reasoning on Digital Waveforms 2026-07-22
10 pa...

10 pages; abridged version published in IEEE International Conference on LLM-Aided Design (ICLAD), 2026

Experience Augmented Policy Optimization for LLM Reasoning 2026-07-22

Chain of Thought

Title Date Comment
Separating quantum circuits from classical LLMs 2026-08-04 60 pages, 6 figures
Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models 2026-08-04
Evaluating LLM-Based Goal Extraction in Requirements Engineering: Prompting Strategies and Their Limitations 2026-08-04
11 pa...

11 pages, 1 figure. This contribution will be published in the conference proceedings of EASE 2026 Conference (https://conf.researchr.org/home/ease-2026/prompt-se-2026)

Sensitivity, Causality, and Repair Dissociate: A Layer-Wise Analysis of Perturbation Robustness and Its Scaling 2026-08-04
29 pa...

29 pages, 18 figures, 11 tables

TDVR: Joint Text Disambiguation and Viewpoint Reasoning for Zero-Shot 3D Visual Grounding 2026-08-04
10 pa...

10 pages, 5 figures, 7 tables

Risky Business: Measuring The Faithfulness-Safety Tension 2026-08-04
Soft Guidance Starts to Outperform CoT Prompting as LLMs Improve 2026-08-04 10 pages, 3 tables
Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Explicit and Latent Thoughts for MLLMs 2026-08-04
Accep...

Accepted by ACM MM 2026. 10 pages, 6 figures, 5 tables

Automated Visualization Code Synthesis via Multi-Path Reasoning and Feedback-Driven Optimization 2026-08-04
Accep...

Accepted by International Conference on Pattern Recognization (ICPR 2026)

The Tell-Tale Trace: Detecting Reasoning Failures in LLMs Using Chain-of-Thought Dynamics 2026-08-04
States Hidden in Hidden States: Implicit Discrete State Representations Emerge in LLMs' Hidden States 2026-08-04
12 pa...

12 pages, 12 figures. Revised manuscript with public code and reproducibility artifacts; clarified the evaluation protocol; corrected figure labels and chance baselines, the attention-bridge description, cross-references, and probe notation. No new experiments

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs 2026-08-04
Aligned in Form, Not in Meaning: The Comprehension - Containment Decoupling of LLM Safety in Low-Resource Bangla Derogatory Speech 2026-08-03 15 pages, 6 figures
WCM: World-Cognition Model for Generalizable Human-Robot Interaction 2026-08-03
CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning 2026-08-03
Evading Chain-of-Thought Monitoring Through Model Poisoning 2026-08-03 15 pages, 2 figures
The Illusion of Superposition? A Principled Analysis of Latent Thinking in Language Models 2026-08-03 9 pages
GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning 2026-08-03
MoRAL: Sensor-Grounded BEV Reasoning for Compact VLMs toward Edge-Oriented Autonomous Driving 2026-08-03
7 pag...

7 pages, 5 figures, 6 tables. Accepted to the 14th IEEE International Conference on Intelligent Mobile Computing (IEEE IMC 2026), Fukuoka, Japan, July 27-30, 2026

Can Foundation Models Hear What Made That Sound? A Tiered Benchmark of Audio-Language Models and Traditional Classifiers for Closed-Set Sound Source Identification 2026-08-03
AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning 2026-08-03
Recompute or Reuse? Diagnosing and Mitigating Textual Shortcuts in VLM Self-Reflection 2026-08-03 preprint
Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving 2026-08-03
Comparative Validation of GPT-4o-mini and Teacher Mean Scores for Automated Scoring of Music Analysis Responses: Single-Pass Deployment, Repeatability, and Strategy-Specific Bias 2026-08-03
Proce...

Proceedings of the International Meeting of the Psychometric Society: The 91st Annual Meeting, Seoul, Republic of Korea, 2026

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models 2026-08-03
PICTURE: Enhancing Theory-of-Mind in Large Language Models by Revealing, Not Hiding, Characters' Lack of Knowledge 2026-08-03
Latent Thought Credit: Multi-Answer Credit Assignment for Latent Reasoning 2026-08-03
Emergence Invariance: From Symbolized Thought to Interface Refinement 2026-08-03 14 pages, 3 figures
Length Penalties Make Chain-of-Thought Less Monitorable 2026-08-02
Cognitive Demand Steering for Adaptive Meta-Reasoning in Large Language Models 2026-08-02 21 pages, 3 figures
Remember-R1: Mitigating Long-Context Visual Forgetting through Reinforcement Learning 2026-08-02
Accep...

Accepted by ACM Multimedia 2026

It's the Decoding Format, Not the Perturbation: Auditing Consistency-Based Selection for Vision-Language Test-Time Scaling 2026-08-02
CoT-Edit: Let CoT Guide Instruction Video Editing 2026-08-02
Attend to Your Own Thoughts: Breaking the Barrier for Post-Training Quantization of Reasoning LLMs through the Lens of 1.58-Bit Quantization 2026-08-02
This ...

This research work was completed and submitted for publication in early May 2026. The project page: https://github.com/IntelChina-AI/BitTern

AdaThink-Med: Optimizing Inference-Time Compute for Medical Reasoning via Uncertainty Quantification 2026-08-02
CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts 2026-08-02
Accep...

Accepted by ECCV2026. Code is available at https://github.com/hulianyuyy/CoLT

Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy 2026-08-02
21 pa...

21 pages, 7 figures, 14 tables

MonitorVLM-v2: A Deployed Vision-Language Framework for Real-Time Safety Violation Detection 2026-08-02
Gaokerena: A Small Persian Medical Language Model Family 2026-08-02 29 pages, 9 figures
When Does LLM Orchestration Pay Off? A Controlled Evaluation of Accuracy, Cost, and Task Difficulty 2026-08-01
A False Average: Chain-of-Thought Monitors Collapse Where They Are the Only Defense 2026-08-01
Native Multilingual Chain-of-Thought Reasoning in Low-Resource Southeast Asian Languages 2026-08-01
22 pa...

22 pages, 16 figures, 12 tables

Knowledge Graph Augmented Large Language Models for Disease Prediction 2026-08-01
DeceptionX: From Multimodal Evidence to Explainable Deception Detection 2026-07-31
Cross-Task Dissociation in Frontier Vision-Language Model Theory of Mind 2026-07-31
AI and Its Impact on Creativity and Diversity: An Empirical Study of LLM-Generated Product Ideas 2026-07-31
RePaCA: Leveraging Reasoning Large Language Models for Static Automated Patch Correctness Assessment 2026-07-31
Final...

Final published version in the Neurocomputing journal. Volume 701, 7 November 2026, 134583. DOI: https://doi.org/10.1016/j.neucom.2026.134583

Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation 2026-07-31
34 pa...

34 pages, 17 figures, and 21 tables. Accepted to ACL 2026

Data Turnstile: A Scalable Open Framework for Function-Calling Data Generation 2026-07-31 16 pages
Semantics of Subterfuge: Benchmarking Legal Deception Detection Against General-domain State-of-the-Art 2026-07-31 5 pages paper

LLM Interpretability

Title Date Comment
PartInteractor: Intent-Driven Part-Aware 3D Authoring for Continuous Co-Creation in XR 2026-08-02
Accep...

Accepted to ACM UIST 2026

Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth 2026-07-30
How memory can affect collective and cooperative behaviors in an LLM-Based Social Particle Swarm 2026-07-29
11 pa...

11 pages, 4 figures and 2 tables

CogEEGAgent: Toward Autonomous Cognitive EEG Analysis with Grounded Execution and Selection-Aware Verification 2026-07-27
16 pa...

16 pages, 5 figures, and 16 tables. The supplementary material is included in the same PDF

AlphaRoute: Large Language Models as Semantic Optimizers for Multi-Objective Routing 2026-07-22
7 pag...

7 pages, 5 figures. Accepted for publication in the IEEE International Conference on LLM-Aided Design, 2026, Stanford University, Stanford, CA, USA. Code available at https://github.com/Kcbir/AlphaRoute

Understanding Interpretation Difficulty in Harmful Online Communication: Insights from Cybercrime Communities 2026-07-08
NeuraDock Visual Cognitive Load Agent Tutorial: A Quality-Gated Open-Source EEG Workflow for Alpha Dynamics and Real-Time Applications 2026-06-25 22 pages, 10 figures
Don't Go Breaking My LLM: The Impact of Pruning Attention Layers on Explanation Faithfulness and Confidence Calibration 2026-06-23 Accepted at TMLR
LLM-Based Generalizable Hierarchical Task Planning and Execution for Heterogeneous Robot Teams with Event-Driven Replanning 2026-06-18
full ...

full version of this short paper is accepted at Frontiers in Robotics and AI Journal

ICA Lens: Interpreting Language Models Without Training Another Dictionary 2026-06-10 Ongoing Project
Language-Driven Cost Optimization for Autonomous Driving 2026-06-09
Paper...

Paper accepted at IEEE Intelligent Transportation Systems Conference (ITSC) 2026

ReflectiChain: Epistemic Grounding in LLM-Driven World Models for Supply Chain Resilience 2026-06-09
MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding 2026-06-08
Analyzing the Correlation Between Hallucinations and Knowledge Conflicts in Large Language Models 2026-06-07
Thinking Through Signs: PEEL as a Semiotic Scaffolding for Epistemically Accountable AI-Enabled Research 2026-06-02 10 pages, 5 figuras
Unveiling the Limits of Large Language Models in Inferring Pragmatic Meaning from Non-Verbal Responses 2026-06-01
The Limits of LLM Forecasting: Parametric Knowledge Gaps Across Conflict Zones 2026-05-29
Hallucination Behavior in Multimodal LLMs Across Agricultural Image Interpretation and Generation Tasks 2026-05-26
How Much Structure Do LLMs Need? Evaluating LLMs for Bibliometric Cluster Description 2026-05-23
Improving Labeling Consistency with Detailed Constitutional Definitions and AI-Driven Evaluation 2026-05-22
Under...

Under review at ACL Rolling Review (ARR), May 2026 cycle. Also available at https://doi.org/10.5281/zenodo.20125267

Patch-Effect Graph Kernels for LLM Interpretability 2026-05-07
LLM-ADAM: A Generalizable LLM Agent Framework for Pre-Print Anomaly Detection in Additive Manufacturing 2026-05-05 24 pages, 10 figures
From Research Question to Scientific Workflow: Leveraging Agentic AI for Science Automation 2026-04-23
Agentic AI-Enabled Framework for Thermal Comfort and Building Energy Assessment in Tropical Urban Neighborhoods 2026-04-23
Accep...

Accepted at IAQVEC 2026

Learning to Draw ASCII Improves Spatial Reasoning in Language Models 2026-04-16
Arbitration Failure, Not Perceptual Blindness: How Vision-Language Models Resolve Visual-Linguistic Conflicts 2026-04-13
Heterogeneous Graph Importance Scoring and Clustering with Automated LLM-based Interpretation 2026-04-09
26 pa...

26 pages, 11 figures, 8 tables

Runtime Execution Traces Guided Automated Program Repair with Multi-Agent Debate 2026-04-03
12 pa...

12 pages, 4 figures, 8 tables

Discovering Decoupled Functional Modules in Large Language Models 2026-03-18 AAAI-26 Oral
LaMoGen: Language to Motion Generation Through LLM-Guided Symbolic Inference 2026-03-12
Accep...

Accepted by CVPR 2026. Supplementary material included. Project page: https://jjkislele.github.io/LaMoGen/

Pneuma-Seeker: A Relational Reification Mechanism to Align AI Agents with Human Work over Relational Data 2026-03-11
PhysLLM: Harnessing Large Language Models for Cross-Modal Remote Physiological Sensing 2026-03-05
Accep...

Accepted by International Conference on Learning Representations (ICLR) 2026

Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs 2026-03-01
Sparse Shift Autoencoders for Identifying Concepts from Large Language Model Activations 2026-02-27 27 pages, 9 figures
Co-Disclosing the Computer: LLM-Mediated Computing through Reflective Conversation 2026-02-27 CHI'26
Neuro-Symbolic Control with Large Language Models for Language-Guided Spatial Tasks 2026-02-21
Stroke Lesions as a Rosetta Stone for Language Model Interpretability 2026-02-03 45 pages, 17 figures
Uncertainty-Aware Gradient Signal-to-Noise Data Selection for Instruction Tuning 2026-01-20 Preprint
A Shared Geometry of Difficulty in Multilingual Language Models 2026-01-19
Causal-SAM-LLM: Large Language Models as Causal Reasoners for Robust Medical Segmentation 2026-01-16
Accep...

Accepted by IEEE ICASSP 2026

Self-reflection in Automated Qualitative Coding: Improving Text Annotation through Secondary LLM Critique 2026-01-14
Word Synchronization Challenge: A Benchmark for Word Association Responses for Large Language Models 2026-01-14
Can LLMs interpret figurative language as humans do?: surface-level vs representational similarity 2026-01-14 17 pages, 5 figures
Neuro-Symbolic Compliance: Integrating LLMs and SMT Solvers for Automated Financial Legal Analysis 2026-01-07
10 pa...

10 pages, 6 tables, 3 figures, accepted by the 2nd ACM AIware Conference

LLM Interpretability with Identifiable Temporal-Instantaneous Representation 2026-01-02 NeurIPS 2025
Trust Me, I Know This Function: Hijacking LLM Static Analysis using Bias 2025-12-18
Verification-Guided Context Optimization for Tool Calling via Hierarchical LLMs-as-Editors 2025-12-15
Accep...

Accepted by AAAI 2026 Workshop on Agentic AI Benchmarks and Applications for Enterprise Tasks

Visualizing token importance for black-box language models 2025-12-12
SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language Models 2025-12-01
ACL 2...

ACL 2025 main. Our Speech-IQ leaderboard is hosted at huggingface.co/spaces/nvidia/Speech-IQ-leaderboard. Speech-IQ Calculator: https://github.com/YukinoWan/SpeechIQ

For Those Who May Find Themselves on the Red Team 2025-11-23

Explainable AI

Title Date Comment
Paris as a 15-Minute City: An Explainable AI Perspective 2026-08-04
17 pa...

17 pages, 16 figures. Extended report of a poster presented at the NetMob 2025 conference on 8 October 2025

A Unified Framework for Human AI Collaboration in Security Operations Centers with Trusted Autonomy 2026-08-04 Accept
Trustworthy AI in Digital Health: A Comprehensive Review of Robustness and Explainability 2026-08-03
Prepr...

Preprint of the paper published in Progress in Biomedical Engineering. 26 pages, 5 figures

Explainable AI for the EU Right to Explanation: A Systematic Review of the Law-XAI Translation Gap 2026-08-03
Accep...

Accepted at the 9th AAAI/ACM Conference on AI, Ethics and Society (AIES-26)

Crushing the Evidence: A Dual-Penalty Evasion Framework for Fooling White-Box Explainable AI Auditors 2026-08-01 10 pages
Explaining AI-Image Detection: What the Heatmap Actually Shows 2026-07-31
8 pag...

8 pages of main text; 27 pages including references and appendix. 9 figures, 21 tables

Shaping Scientific Explanations to Expert Perspectives with Persona-Conditioned Reinforcement Learning 2026-07-31
Application of machine learning to monster level prediction in tabletop RPG game design 2026-07-31
From Large Language Model Predicates to Logic Tensor Networks: Neurosymbolic Offer Validation in Regulated Procurement 2026-07-30
17 pa...

17 pages, 2 figures, 4 tables, extended version, with appendix

The Case for Vibe Modeling: A Missing Step in AI-Based Trustworthy Software Development 2026-07-30 7 pages, 1 figure
On the Design and Evaluation of Human-centered Explainable AI Systems: A Systematic Review and Taxonomy 2026-07-28
From Dyad to Triad: Eliciting XAI Requirements in Stroke Rehabilitation 2026-07-28
AI an...

AI and Cognitive Computing for Trustworthy Human Machine Systems Session, IEEE International Conference on Systems, Man, and Cybernetics (IEEE SMC 2026)

Explainable AI for Chronic Kidney Disease Prediction Using Simulated Federated Learning 2026-07-28 13 Pages, 5 Figures
A Unified Framework for Uncertainty-Aware Explainable Artificial Intelligence: A Case Study in Power Quality Disturbance Classification 2026-07-28
JobMatchAI-An Intelligent Job Matching Platform Using Knowledge Graphs, Semantic Search and Explainable AI 2026-07-27
Evaluating the Impact of Explainable AI on Trust in AI-Assisted Code Review 2026-07-27
23 pa...

23 pages, 4 figures, 5 tables. To appear in Proceedings of the ACM on Software Engineering (PACMSE), Vol. 3, No. ISSTA, Article ISSTA093 (ISSTA 2026). Published under CC BY 4.0. Replication package: https://doi.org/10.5281/zenodo.21457282

Beyond Local Inspection: Global, Guideline-Grounded Evaluation of Post-hoc XAI Methods for ECG Classification 2026-07-27
Disentangling Acoustic Cues in Alzheimer's Pathology and Perception: The Roles of Language and Gender 2026-07-27
Accep...

Accepted at Interspeech 2026

Explainable AI through the Lens of Material Agency: Enabling Musical Interface Design with Neural Audio Models 2026-07-25
Under...

Under review for "Explainable AI for the Arts" (N. Bryan-Kinns, Ed.), Springer

Context-Aware Concept Distillation for Trustworthy Flood Prediction 2026-07-25
to be...

to be published in IJCAI 2026 proceedings

Unboxing Diffusion Models for the Arts: Interactive Model Bending and Practice-Based Explainability 2026-07-24
Under...

Under review for "Explainable AI for the Arts" (N. Bryan-Kinns, Ed.), Springer

Scalable Explainability-as-a-Service (XaaS) for Edge AI Systems 2026-07-24
8 pag...

8 pages, 5 figures, 2 tables. This version updates metadata after publication in IEEE Xplore and publication by SoutheastCon 2026

Proceedings of The Fourth International Workshop on eXplainable AI for the Arts (XAIxArts 4) 2026-07-22
Scaling Time Series Classification via XAI-Driven Data Reduction 2026-07-22
Accep...

Accepted for AALTD workshop at ECML-PKDD 2026

LLM-Grounded Explainable AI for Supply Chain Risk Early Warning via Temporal Graph Attention Networks 2026-07-21
Automated Data Engineering and Feature Selection for the Case Study of Warpage Detection in Fused Deposition Modeling 2026-07-20
Explaining and Tuning Transformer-based LLMs in Arithmetic Tasks with Human Strategies 2026-07-19
FST.ai 2.5: Explainable and Uncertainty-Aware AI for Olympic and Para-Taekwondo Decision Support, Athlete Digital Twins, and Federation-Scale Analytics 2026-07-18 14 pages, 4 figures
FAIR_XAI: Improving Multimodal Foundation Model Fairness via Explainability for Wellbeing Assessment 2026-07-17 11 pages, 4 figures
VeriX-Anon: A Multi-Layered Framework for Mathematically Verifiable Outsourced Target-Driven Data Anonymization 2026-07-17
v2: r...

v2: revised after peer review. Evaluation expanded from 3 to 7 datasets, per-dataset Wasserstein-threshold calibration added, effect-size CIs and cross-dataset statistics reported, and analytical zk-SNARK/MPC/TEE baselines added. Minor errors corrected

Automated identification of Ichneumonoidea wasps via YOLO-based deep learning: Integrating HiresCam for Explainable AI 2026-07-16 15 pages, 20 figures
Explaining Process Control Optimisation Recommendations via GradientSHAP and Implicit Differentiation 2026-07-16
Diagnosing and Mitigating Domain Shift in Permission-Based Android Malware Detection 2026-07-15
"Trust Junk" Leads to Unjustified Support for Highly Discriminatory Predictive Models 2026-07-14
MLPTR-CC: Multi-label Pathology Test Recommendation using Classifier Chains and SHAP 2026-07-14
Atomic Units of X: The Compression Layer of Intelligence 2026-07-14
Evaluating RE Practices for Explainability: Synthesizing Insights from Daimler Truck into an Explainable RE Framework Proposal 2026-07-13
Detecting Explanatory Insufficiency in Learned Representations: A Framework for Representational Vigilance 2026-07-13
25 pa...

25 pages, 2 figures, no tables, 16 references. Conceptual and methodological framework for monitoring representational adequacy and detecting explanatory insufficiency in learned representations

ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI 2026-07-10
Knowledge Graphs and Explainable AI as Complementary Resources for Urban Mining 2026-07-10
Accep...

Accepted for presentation at the AISE Workshop @ IJCAI-ECAI 2026

All Explanations are Wrong, But Many Are Useful: Exploring the Rashomon Explanation Set with Large Language Models 2026-07-10
Unlearning to Protect: A Distilled Reinforcement Learning Framework with Privacy-Preserving Feature Unlearning and XAI for IoT Security 2026-07-09
The Contribution of XAI for the Safe Development and Certification of AI: An Expert-Based Analysis 2026-07-09
This ...

This preprint has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this contribution is published in Artificial Intelligence in HCI (HCII 2026), Lecture Notes in Computer Science, vol. 16745, and is available online at https://doi.org/10.1007/978-3-032-30849-8_13

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning 2026-07-08 20 pages
Why Fake ? Unveiling the Semantic Vocabulary of Deepfake Detectors 2026-07-08
Accep...

Accepted at CVPRW 2026

Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies 2026-07-08
Invit...

Invited survey paper. 10 pages, 5 figures, 2 tables

Reduced NEXI protocol for the quantification of human gray matter microstructure on the Connectome 2.0 scanner 2026-07-07
Submi...

Submitted to Imaging Neuroscience. This all-in-one version includes supplementary materials. 34 pages, 145 figures, 4 tables

Mechanistic Interpretability

Title Date Comment
Sparse Weight Decomposition for Efficient Circuit Extraction 2026-08-04
Disentangling MLP Neuron Weights in Vocabulary Space 2026-08-04
Accep...

Accepted at COLM 2026

Word Recovery in Large Language Models Enables Character-Level Tokenization Robustness 2026-08-04
The production of meaning in the processing of natural language 2026-08-03
Accep...

Accepted to QNLP 2026, 9 pages, 3 figures, 2 tables

Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models 2026-08-02
Accep...

Accepted as a conference paper at the Conference on Language Modeling (COLM) 2026

DiffuseAgent-MI: Distributionally-Grounded,Tool-Integrated Self-Evolving Agents for Faithful Visual Reasoning 2026-08-01
11 pa...

11 pages, 9 figures, accepted by ICML 2026 manitrack

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions 2026-08-01
EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures 2026-07-31
74 pa...

74 pages, 2 figures, 4 tables. Hybrid systematic survey and conceptual framework on LLM evaluation and AI-safety failures, synthesizing 373 primary studies (2018-2026). Introduces the EvalSafetyGap framework (Instability Decomposition, Alignment Trilemma) and reports an exploratory ten-model audit. Submitted as a review/survey article; not currently under consideration elsewhere

HYVINT: Intensity-Driven Hypergraph Generation with Variational Embeddings 2026-07-29
8 pag...

8 pages, 1 figure, 9 tables

Towards Verifiable Transformers: Solver-Checkable Circuit Explanations 2026-07-29
23 pa...

23 pages. v2: adds GPT-2-scale verified distillation (three-edge verified quote circuit), LayerNorm removal for sparsemax models, and gated localization protocols

Phase Structure in Rotary Attention: A Spectral Framework for Semantic Continuity and Execution-Boundary Governance 2026-07-28
14 pa...

14 pages; theoretical framework and proposed experimental program

Emergent Latent-State Computation under Stochastic Volatility 2026-07-28
Do LLMs Know Their Vulnerable Scenarios? 2026-07-26
16 pa...

16 pages, 10 Figures, Under Review

Continuous surrogates versus threshold Boolean networks for modeling Arabidopsis ISR gene regulation 2026-07-25
To be...

To be published in IEEE CIBCB 2026

Through the Bottleneck: How Multi-head Latent Attention Separates Content from Position in Language Models 2026-07-25
Toward Mechanistic Interpretability of an AI Foundation Model Fine-Tuned for Atmospheric Chemistry 2026-07-22
CircuitKIT : Circuit Discovery, Evaluation, and Application Toolkit for Mechanistic Interpretability 2026-07-21
CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs 2026-07-21
9 pag...

9 pages, 7 figures, 1 table. Code: https://github.com/LLM-Interp/CLT-Forge. Demonstration video: https://youtu.be/6ptrrLawTl8

Breaking the Block: Preserving Data Continuity to Train Superior SAEs for Instruct Models 2026-07-20
Measuring Monosemanticity in Sparse Autoencoders via Latent Activation Coherence 2026-07-20
This ...

This is a preprint version. A shorter version of this paper has been accepted for presentation and publication in the post-workshop proceedings of the 8th International Workshop on eXplainable Knowledge Discovery in Data Mining (XKDD 2026), co-located with ECML PKDD 2026. The appendix is included only in this preprint and is not part of the peer-reviewed proceedings paper

Every Component is a Lookup: Token Attribution and Composition from a Single Decomposition 2026-07-19
What Do They See? Interpreting Complex Road Scenarios Through the Eyes of Vision-Language-Action Models for Safe and Trustworthy Autonomous Vehicle Learning 2026-07-18
Are Arithmetic Heuristic Neurons Form-Invariant? A Mechanistic Analysis of Symbols, Text, and Code in LLMs 2026-07-18 Under Review
Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy 2026-07-17
Prepr...

Preprint. Under review

Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability 2026-07-17
Prepr...

Preprint. Under review

Hide and Seek in Embedding Space: Geometry-based Steganography and Detection in Large Language Models 2026-07-17
Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control 2026-07-16
Transcoders for Investigating Deception in Language Models 2026-07-16
From Observation to Insight: Mechanistic World Models and the Quest for Autonomous Discovery 2026-07-15
Hierarchical Latent Structures in Data Generation Process Unify Mechanistic Phenomena across Scale 2026-07-14
Same Compression Principle, Different Geometry: Rate-Distortion Signatures Dissociate Biological and Artificial Visual Systems 2026-07-13
Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias 2026-07-13
58 pa...

58 pages, 13 figures, 30 tables; project page: https://xzx34.github.io/unfair-judge/

Emergent Hierarchical Monosemantic Neurons from the Group-Contrastive Forward-Forward Algorithm 2026-07-13
Which Neurons Detect Malicious Code? A Probing Study of LLM Security Knowledge 2026-07-11
The p...

The paper has been peer reviewed and accepted for publication in the 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)

MLPs are Hebbians: Constructing Efficient Fact-Storing MLPs for Transformers 2026-07-10
XAI and Statistical Analysis for Reliable Intrusion Detection in the UAVIDS-2025 Dataset: From Tree to Hybrid and Tabular DNN Ensembles 2026-07-10
Accep...

Accepted at IEEE CITS 2026, Greece

When Structured Sparse Autoencoders Learn Consistent Concepts Across Modalities 2026-07-09
Cross-seed explainability using Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoders 2026-07-09
17 pa...

17 pages, 4 figures, 6 tables

Certified Interventional Fidelity: Anytime-Valid, Adaptive Evaluation of Causal Claims in Mechanistic Interpretability 2026-07-09
Accep...

Accepted at UAI 2026 (Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence). Code: https://github.com/AsiaeeLab/certified-interventional-fidelity

Diagnosing Shape-Prior Shortcuts in Long-Range Single-Shot Fringe Projection Profilometry 2026-07-09 21 pages, 13 figures
Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs 2026-07-08
Temporal Preference Concepts and their Functions in a Large Language Model 2026-07-08
How Learning Dynamics Drive Adversarially Robust Generalization? 2026-07-08
Accep...

Accepted at the 42nd Conference on Uncertainty in Artificial Intelligence (UAI 2026)

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning 2026-07-08 20 pages
Latent Programming Horizons in Coding Agents 2026-07-06
Beyond the Black Box: Interpretability of Agentic AI Tool Use 2026-07-05
12 pa...

12 pages, 4 figures, 17 tables

Interpretability and Generalization Bounds for Learning Spatial Physics 2026-07-05
To ap...

To appear in ICML 2026. 18 pages, 13 figures

Cultural Binding Heads in Language Models 2026-07-04

Metadata

Metadata

Assignees

No one assigned

    Labels

    documentationImprovements or additions to documentation

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions