Curated coverage, model releases, technical papers, and engineering updates tagged with #Transformer.
Autonomy-1 will have a small, transformer-based AI model taking charge of a space probe.
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
An interactive visualization tool showing you how transformer models work in large language models (LLM) like GPT.
Hacker News discussion about "A Mathematical Framework for Transformer Circuits (2021)".
Loop Transformer vs Recursive Transformer: why looping blocks is not what next-gen Transformers need, and how composable context turns CoT into thinking about thinking.
A Look at Recurrent Depth, Hidden Chains of Thought, and Recent Research on Looping Transformer Blocks
Pathway's Baby Dragon Hatchling (BDH) is a brain-inspired, post-transformer architecture that reasons in latent space instead of emitting chain-of-thought tokens. See how Pathway develops and scales BDH on Amazon SageMaker HyperPod, and how BDH-CQ set a new cost-efficiency mark on the ARC-AGI-1 benchmark.
Hacker News discussion about "I trained a small transformer in 1.5hrs and it beats many LLMs".
Peter Cullen, the voice actor who played Optimus Prime in Transformers films and shows and Eeyore in Winnie the Pooh projects, has died. He was 85.
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
A 3.16M-parameter INT4 transformer running entirely in the on-chip memory of a Xilinx Kria KV260. Zero DRAM in the token loop, 59,965 tok/s on the fabric, bit-exact. Chat with it live.
Transformer Transformer: A Unified Model for Motion-Conditioned Robot Co-design. Project website.
I run a small AI lab and playground and got super excited about Anthropics paper "Verbalizable Representations Form a Global Workspace in Language Models" (https://transformer-circ...
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Abstract page for arXiv paper 2607.01232: Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training
A Blog post by Ai2 on Hugging Face
A Blog post by NVIDIA on Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.