Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
19 results for “Transformers”
Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
Hugging Face announced new multi-vector (late interaction) embedding models built with Sentence Transformers, enabling more precise semantic search by representing queries and documents as multiple vectors rather than single embeddings.
Aug 18, 2026
I compiled Doom's renderer into a 21B-parameter transformer -- no training anywhere [P]
A researcher compiled the Doom game renderer into a transformer model without training, using a custom compiler to convert the algorithm into transformer weights, resulting in a functional but extremely slow implementation.
Aug 14, 2026
Dual-Flow Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation
Researchers propose Dual-Flow Transformers, a novel architecture that decouples prompt prefill and autoregressive decode computation to reduce cumulative inference cost without increasing prefill overhead.
Aug 14, 2026
Position Encoding in Transformers: From Absolute and Relative Methods to Rotary Position Embeddings and Long-Context Scaling
A technical survey paper on position encoding methods in Transformers synthesizes and compares absolute, relative, and rotary embedding techniques, with emphasis on long-context scaling strategies and empirical evaluation criteria.
Aug 13, 2026
ChronoSSM: Training for Temporally Aware Representations in Autoregressive State Space Models
ChronoSSM is a new autoregressive State Space Model that jointly trains on both event tokens and timestamps to improve temporal reasoning in sequence modeling, addressing a gap where timing is typically treated as secondary to event prediction.
Aug 12, 2026
I’m Researching Leo — a byte-native learning architecture that tries to move beyond Transformers
A solo researcher introduces Leo/PSCLS, an experimental byte-native neural architecture emphasizing persistent state and sparse recurrence over Transformer-style attention, positioning it as a biologically inspired alternative still in early development.
Aug 8, 2026
Recursive transformers for semiconductor thermo-mechanical reliability
A new recursive transformer architecture is proposed to improve parameter efficiency and computational cost for surrogate modeling in semiconductor thermo-mechanical reliability analysis, where training data is scarce and first-principles simulation is prohibitively expensive.
Jul 31, 2026
Multimodal Surface EMG Hand Gesture Recognition Using Query-Based Transformers for Prosthetic Control
A new hybrid convolutional-transformer model called EMG-CrossFormer improves hand gesture recognition accuracy from surface electromyography (sEMG) signals—especially when fused with inertial data—across four benchmark prosthetic control datasets.
Jul 28, 2026
Hierarchical Grading in Large Language Models
Researchers propose Graded Large Language Models (GLLMs), a theoretical extension of transformer architecture using algebraic grading to improve statistical efficiency for level-stratified prediction tasks, with claims of provable risk separation and pre-certified optimization.
Jul 28, 2026
Convolution for Large Language Models
Researchers propose integrating lightweight depthwise convolutions into Qwen3 Transformer blocks to improve local token interaction modeling without meaningfully increasing parameter count, reporting accuracy gains across seven downstream benchmarks.
Jul 22, 2026
The AI data center boom has led to surging demand for power transformers, with average lead times for orders, once measured in months, now stretching into years (Financial Times)
The AI data center boom is causing unprecedented demand for power transformers, extending average order lead times from months to years and straining global supply chains.
Jul 10, 2026
Native-speed vLLM transformers modeling backend
Hugging Face announced integration of vLLM as a native backend for Transformers, enabling faster inference for large language models without requiring users to rewrite code.
Jul 9, 2026
Induction Heads Interpolate N-Grams
A new arXiv preprint formally links induction heads in transformer models to classical statistical smoothing techniques—specifically Jelinek-Mercer and Dirichlet-style smoothing—by analyzing their behavior on order-k Markov chains.
Jul 8, 2026
Train and run transformers directly on Apple's Neural Engine
A Hacker News thread discusses the technical feasibility and implications of running transformer models directly on Apple's Neural Engine, reflecting community interest in on-device AI acceleration.
Published Jul 5, 2026 · Analyzed Jul 8, 2026
From Approximation to Emergence: A Theory of Deep Learning
A new arXiv monograph proposes a unified theoretical framework for deep learning, positioning emergence—not just approximation—as the central organizing principle of modern AI theory.
Published Jul 3, 2026 · Analyzed Jul 6, 2026
Has anyone tried this approach with Fast Byte Latent Transformers ? [R]
A Reddit user asks whether replacing the transformer architecture in a Fast Byte Latent Transformer entropy model with a Mamba architecture is feasible, citing Mamba's computational efficiency (O(n) complexity) and popularity.
Published Jul 2, 2026 · Analyzed Jul 6, 2026
Accelerating Transformers Fine-Tuning with NVIDIA NeMo AutoModel
Hugging Face announced integration with NVIDIA NeMo AutoModel to speed up transformer fine-tuning, positioning it as a performance optimization for developers.
Published Jun 24, 2026 · Analyzed Jul 3, 2026
Experimenting with the proposed Cross-Origin Storage API in Transformers.js
Hugging Face announced experimental integration of the proposed Cross-Origin Storage API into Transformers.js to enable browser-based AI model caching across domains.
Published Jun 23, 2026 · Analyzed Jul 3, 2026
NVIDIA Launches Cosmos 3, the Open Frontier Foundation Model for Physical AI
NVIDIA announced Cosmos 3, an open foundation model for physical AI that integrates vision reasoning, world generation, and action prediction using a novel mixture-of-transformers architecture.
Published Jun 1, 2026 · Analyzed Jul 4, 2026