Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
16 results for “interpretability”
Beyond Routing Weights: Faithful Response-Level Interpretation of Mixture-of-Experts Reward Models via Contribution Contrast
Researchers introduced CoCo, a new response-level interpretation method for Mixture-of-Experts reward models that improves interpretability by analyzing contribution contrasts between chosen and rejected responses, rather than relying solely on routing weights.
Aug 10, 2026
Interpreting Black-Box Large Language Models with Sentence-Level Energy Landscapes
Researchers propose a model-agnostic, post-hoc sentence-level attribution method for proprietary LLMs using an Energy-Based Model surrogate to quantify prompt influence without repeated API calls.
Aug 5, 2026
LLMs Can Annotate Attribution Graphs
Researchers propose using LLMs to automate the manual grouping of neural features into supernodes for circuit tracing—a step toward scalable interpretability of language models.
Aug 5, 2026
LAWFUL: Law-Aligned Witness for Faithful Use of Latents
Researchers introduced LAWFUL, a new interpretability framework to assess whether neural networks learn and internally use formal physics laws—specifically testing if a Mocap2Radar transformer encodes the Doppler frequency law—addressing four key gaps in causal and domain-validity analysis for continuous-variable physical systems.
Aug 3, 2026
Probabilistic Concept-Aware Steering for Trustworthy LLM Inference
A new research paper introduces Probabilistic Concept-Aware Steering (PCS), a method to improve interpretability and fine-grained control in LLM inference by replacing binary steering evaluation with probabilistic, continuous semantic alignment.
Jul 22, 2026
Are Arithmetic Heuristic Neurons Form-Invariant? A Mechanistic Analysis of Symbols, Text, and Code in LLMs
A mechanistic interpretability study finds that arithmetic reasoning in Llama-3 models relies on a small, shared set of neurons across symbolic math, word problems, and Python code — suggesting failures stem from inconsistent activation states rather than format-specific circuitry.
Jul 21, 2026
From hyperplanes to hyperellipsoids: characterizing the inherent interpretability of linear and single-qubit mixed-state binary classification models
A new arXiv preprint introduces a conceptual mapping between classical linear binary classifiers and single-qubit mixed-state quantum classifiers, framing the latter as a geometric generalization (hyperellipsoid vs. hyperplane) to lower the barrier for teaching quantum ML concepts.
Jul 20, 2026
Heterogeneous Element-Aware Cross-Version Differencing of Scientific Documents via Layout-Aware Alignment and Structure-Aware Reasoning
A new AI framework for comparing different versions of scientific documents (e.g., manuscript revisions) by jointly modeling layout, structure, and semantic element types — improving accuracy in detecting and localizing changes across text, tables, formulas, and figures.
Jul 17, 2026
MAPS: Modeling Co-Existing Subjective Perspectives and Shared Meaning in Multi-Agent Cognitive Dialogue
A new AI research paper introduces MAPS, a framework for multi-agent dialogue systems that preserves subjective perspectives while enabling shared meaning — advancing interpretability and cognitive grounding in conversational AI.
Jul 17, 2026
Mechanistic interpretability researchers applying causality theory to LLMs
A Hacker News thread titled 'Mechanistic interpretability researchers applying causality theory to LLMs' contains user comments discussing early-stage academic efforts to use causal inference frameworks to understand internal mechanisms of large language models.
Jul 13, 2026
From Approximation to Emergence: A Theory of Deep Learning
A new arXiv monograph proposes a unified theoretical framework for deep learning, positioning emergence—not just approximation—as the central organizing principle of modern AI theory.
Published Jul 3, 2026 · Analyzed Jul 6, 2026
Domain Knowledge Based Temporal-Spatial Graph Convolution Network for ECG Recognition
A new graph convolutional neural network architecture incorporating domain-specific ECG landmarks and temporal-spatial graph structures achieves 88.1% average F1 score on a nine-class Chinese ECG dataset, improving rare-class detection by embedding clinical knowledge into model design.
Published Jul 3, 2026 · Analyzed Jul 6, 2026
Multilayer Q-Matrix-Embedded Neural Network for Cognitive Diagnosis (M-QCDNet): Structure-Aware Deep Learning Architecture for Psychometric Interpretability
A new neural network architecture (M-QCDNet) embeds psychometric Q-matrix structure into deep learning to preserve cognitive interpretability while maintaining predictive power for educational assessment.
Published Jul 3, 2026 · Analyzed Jul 6, 2026
TokenScope: Token-Level Explainability and Interpretability for Code-Oriented Tasks in Large Language Models
TokenScope is a new open-source tool for token-level interpretability in code-generation LLMs, enabling real-time inspection of attention, uncertainty, and structural program behavior during decoding.
Published Jul 3, 2026 · Analyzed Jul 6, 2026
Profit-Based Counterfactual Explanations for Product Improvement: A Case Study of Manga Sales in Japan
Researchers propose 'profit-based counterfactual explanations' (PBCE) — a new AI interpretability method that reframes counterfactual generation as profit maximization, replacing arbitrary target values and abstract distance metrics with economically grounded cost and revenue terms.
Published Jul 3, 2026 · Analyzed Jul 6, 2026
Representation as a Bottleneck for Mechanistic Interpretability: The Manifestation Unit Protocol
Researchers propose a new protocol to improve the reusability of neural-network component-level analyses.
Published Jul 2, 2026 · Analyzed Jul 5, 2026