Find a story

Search Spins

Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.

0 results for “self-attention”

SPIN Processed News Frame: The Hype

BCMT: Blockwise Causal Memory Transformer

BCMT is a new Transformer architecture that replaces dense global self-attention with blockwise local attention plus an exponential causal memory mechanism to improve efficiency for long-context language modeling.

Spin 45% Claim Present in Source AI Risk Moderate
arXiv Computation and Language

Aug 17, 2026

SPIN Processed News Frame: The Hype

Position Encoding in Transformers: From Absolute and Relative Methods to Rotary Position Embeddings and Long-Context Scaling

A technical survey paper on position encoding methods in Transformers synthesizes and compares absolute, relative, and rotary embedding techniques, with emphasis on long-context scaling strategies and empirical evaluation criteria.

Spin 40% Claim Present in Source AI Risk Moderate
arXiv Computation and Language

Aug 13, 2026

SPIN Processed News Frame: The Hype

Designing a Good Virtual Node: Addressable and Cardinality-Preserving Global Memory for Message Passing Architectures

A new research paper proposes an 'addressable and cardinality-preserving' virtual node design for graph neural networks that improves global memory representation without self-attention, enabling injective multiset encoding for tasks like motif counting and link prediction.

Spin 40% Claim Present in Source AI Risk Moderate
arXiv Machine Learning

Aug 5, 2026

SPIN Processed News Frame: The Halo

Hierarchical Grading in Large Language Models

Researchers propose Graded Large Language Models (GLLMs), a theoretical extension of transformer architecture using algebraic grading to improve statistical efficiency for level-stratified prediction tasks, with claims of provable risk separation and pre-certified optimization.

Spin 65% Claim Present in Source AI Risk Moderate
arXiv Machine Learning

Jul 28, 2026

SPIN Processed News Frame: The Cushion

Convolution for Large Language Models

Researchers propose integrating lightweight depthwise convolutions into Qwen3 Transformer blocks to improve local token interaction modeling without meaningfully increasing parameter count, reporting accuracy gains across seven downstream benchmarks.

Spin 22% Claim Present in Source AI Risk Moderate
arXiv Computation and Language

Jul 22, 2026