A field guide to the 2027 Transformer bet

Attention,
after all?

A five-year bet. A moving definition.
An architecture that keeps changing shape.

Make your own call
The proposition, in plain English

Will Transformer-like models still lead most NLP benchmarks on January 1, 2027?

Yes Jonathan Frankle

No Sasha Rush

Read the original wager ↗

Original site: Yes · checked 28 Sep 2026

01 / Make the call

Your definition. Your verdict.

Historical evidence · 2024

What counts as a Transformer? Change the rules and watch the same published results tell a different story.

What counts as “Transformer-like”?

Also count models that mix conventional attention with state-space, recurrent or linear-attention layers. A little attention is enough to qualify.

Jamba paper, Table 2 · March–July 2024

Under your rules

Yes, in this comparison.

13 / 13

tasks led by a qualifying architecture

“Most” means more than half. This is a result within one paper’s comparison, not a verdict on the 2027 wager.

Try long-context QA, then switch between “Attention backbone” and “Include hybrids.” Jamba’s three task wins change sides.

Inspect the evidence 13 tasks · all reported scores

Five models in the Jamba paper’s 2024 academic comparison. Each reported task or aggregate receives one vote; this selection includes coding and math. Models retain their published scores when a rule excludes their architecture; they do not disappear from the competition.

General language · 13 tasks — Reported score. Higher is better; bold indicates the row maximum.
TaskLlama 2 13BLlama 2 70BGemma 7BMixtral 8×7BJambaLeader qualifies?
HellaSwagSentence completion · 10-shot80.785.381.286.787.1Yes
WinoGrandePronoun resolution · 5-shot72.880.272.381.282.5Yes
ARC-EasyScience questions · 0-shot77.380.281.577.673.5Yes
ARC-ChallengeScience questions · 25-shot59.467.353.266.064.4Yes
PIQAPhysical commonsense · 0-shot80.582.881.283.083.2Yes
Natural QuestionsClosed-book QA · 5-shot37.746.932.644.845.9Yes
TruthfulQATruthfulness · 0-shot37.444.944.846.846.4Yes
BoolQYes/no questions · 10-shot81.785.087.288.488.2Yes
QuACConversational reading · 0-shot42.742.439.240.940.9Yes
GSM8KGrade-school math · 3-shot CoT34.755.354.560.459.9Yes
HumanEvalCode generation · pass@118.329.932.334.829.3Yes
MMLUKnowledge aggregate · 5-shot54.869.864.370.667.4Yes
BBHReasoning aggregate · 3-shot39.451.255.150.345.4Yes

Author-reported, rounded scores; no claim of statistical significance or equal training budgets. Read the source ↗ · Download CSV ↓ · How we count

02 / The interesting part

The goalposts have layers.

Meet the models ↗
01

Replacement

State-space and recurrent models can process language without conventional attention. The open question is how far that advantage travels across tasks, scales and training budgets.

Explore attention-free models →
02

Hybridization

Combine a few attention layers with a different sequence model. If that wins, did the Transformer survive—or did its replacement keep the useful bits?

Explore hybrid models →
03

Definition

A leaderboard tells you who scored highest. It does not tell you where an architectural family ends. “Uses attention” and “is a Transformer” are different claims.

Read our working definitions →

03 / What to watch

What would change the answer?

Read the field notes ↗

Three kinds of evidence worth following. A new release matters when it changes one of these arguments.

The attention-free case

A replacement that wins broadly.

Inspect the challengers →

In the record. Mamba, Hyena and recurrent models document alternatives to conventional attention. Source ↗

What is still needed. A clearly defined task basket, comparable evaluations against strong contemporaries, and enough wins to establish a majority. Efficiency alone does not settle a quality-based wager.

The hybrid case

The winner has a little of both.

See the boundary change the result →

In the record. Jamba already provides a historical example of a hybrid leading some tasks in a published comparison. Source ↗

What is still needed. An agreed rule about whether the hybrid counts. New benchmark scores cannot resolve a disagreement about the category itself.

The definition case

A new mechanism, an old family name.

Explore the broadest definition →

In the record. Linear attention and state-space duality show why an architecture’s name is not a sufficient description. Source ↗

What is still needed. A rule about the mechanism being counted, plus documentation of the actual model. Mathematical equivalence and practical implementation are different kinds of evidence.

A good question outlives a countdown

How much can a Transformer change
and still be a Transformer?

* The original wager names a date, not a time zone. This site uses midnight UTC for the countdown. On 28 September 2026, the original page displayed “Yes.” That is a dated status, not a final adjudication. Read the resolution policy.