Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
0 results for “self-attention”
BCMT: Blockwise Causal Memory Transformer
BCMT is a new Transformer architecture that replaces dense global self-attention with blockwise local attention plus an exponential causal memory mechanism to improve efficiency for long-context language modeling.
Aug 17, 2026
Position Encoding in Transformers: From Absolute and Relative Methods to Rotary Position Embeddings and Long-Context Scaling
A technical survey paper on position encoding methods in Transformers synthesizes and compares absolute, relative, and rotary embedding techniques, with emphasis on long-context scaling strategies and empirical evaluation criteria.
Aug 13, 2026
Designing a Good Virtual Node: Addressable and Cardinality-Preserving Global Memory for Message Passing Architectures
A new research paper proposes an 'addressable and cardinality-preserving' virtual node design for graph neural networks that improves global memory representation without self-attention, enabling injective multiset encoding for tasks like motif counting and link prediction.
Aug 5, 2026
Hierarchical Grading in Large Language Models
Researchers propose Graded Large Language Models (GLLMs), a theoretical extension of transformer architecture using algebraic grading to improve statistical efficiency for level-stratified prediction tasks, with claims of provable risk separation and pre-certified optimization.
Jul 28, 2026
Convolution for Large Language Models
Researchers propose integrating lightweight depthwise convolutions into Qwen3 Transformer blocks to improve local token interaction modeling without meaningfully increasing parameter count, reporting accuracy gains across seven downstream benchmarks.
Jul 22, 2026