DataForge 2026 · Pathway track · Explain the Frontier
We wrote a hall allotment into a fixed-size matrix, the same way the Dragon Hatchling architecture writes a conversation into its synapses[1]: one outer-product update per fact, nothing else. Then we asked it where each student lives. The answers below are computed in your browser as you load this page.
Nothing overflowed and nothing was deleted. The matrix holds exactly 2 KB before the first student is written into it and exactly 2 KB after the twentieth. A Transformer would have kept twenty separate key–value entries and answered every question perfectly, using 5 KB to do it. That gap is small here and it is the entire point: it widens with every student you add, and the matrix never moves.
Look at what the mistakes are. The matrix does not return noise when it fails. It returns another student's real room, confidently. That failure mode has a name in the literature, and three different architectures published in the last two years each pick a different way to live with it.
Reproduce the claim in sixty seconds
That is the whole claim: fidelity is traded against how much you store in a fixed space. Everything below this point is why, and what three architectures do about it.
The one sentence this page has to defend
Spelled out properly: a constant-size matrix updated by outer-product writes stores any number of key–value associations without growing, but the accuracy of what comes back out falls as more associations share the same dimensions. Swapping in a cleverer write rule redistributes that accuracy across the stored items. It does not create more of it. A Transformer's key–value cache is the only option here that never has to choose, and it pays for that in memory that grows with every token.
That is a falsifiable statement and the controls on this page are the test. If a write rule existed that raised recall everywhere at once, the comparison in the finding would show it. Push the load up and watch which curve moves. We were wrong about which rule wins, which is the most useful thing that happened while building this.
Who this is for, and what you should be able to do afterwards
Written for someone comfortable with a dot product and a matrix–vector product, and curious about what sits behind the phrase “fast weights”. No machine-learning background is assumed and nothing here needs a GPU. If you have met a KV cache and want to know what the alternative actually is, that is the gap this fills.
Mechanism
The whole trick is a single line of arithmetic. Given a key k and a value v, add their outer product to the state.
If the keys are close to perpendicular, each one reads back mostly its own value and only a little of everyone else's. That "only a little" is the entire story. It accumulates. With n associations in d dimensions, the crosstalk a query has to see past grows like (n−1)/d, which puts the expected cosine between what you get and what you stored at about √(d/(d+n−1)). We derive that in docs/derivation.md and check it against simulation rather than asking you to take it on faith.
Drag the slider. The matrix on the left never changes size. The bars on the right are what you stored against what comes back.
The finding
All three of these keep exactly the same d×d matrix. They differ only in what they do at write time.
Add and never subtract. This is BDH's synaptic update[1], and it is order-independent: shuffle the writes and you get the identical matrix.
Read the key first, then write only the difference. Gated DeltaNet-2 puts it in one line: "delta-rule models subtract the current read before writing a new value."[4] It is also one gradient step on Titans' memory loss ‖M(k) − v‖²[3].
Fade everything slightly before each write. This is Titans' forgetting gate[3] with λ held fixed instead of learned. At λ = 1 it is exactly Hebbian, which is why the slider goes that far.
What we expected, and what we got
We assumed the delta rule would beat plain Hebbian everywhere, because that is what you would conclude from skimming the abstracts. Averaged over everything stored, it does not. At four times capacity our measurements put Hebbian at 0.449 and the delta rule at 0.241. The chart above shows why. Hebbian spreads its fidelity perfectly evenly over all of history. The delta rule spends almost all of it on recent writes, because subtracting the current read before writing actively erases older content to make room. Neither is better. They are answering different questions.
The same roster, all three rules
Watch where the crosses land rather than how many there are. Hebbian scatters its mistakes through the list. The delta rule and the decay gate get the students at the bottom right and lose the ones at the top, because those were written first. Raise "names taught" to 20 and drop d to 12 to make it obvious.
Inside BDH
The Dragon Hatchling paper defines working memory as a synapse strength σ(i,j) between two neurons, changed during inference by a rule it states in words: co-presence of Y(i) followed by X(j) increases σ(i,j) by Y(i)X(j) (§1.2)[1]. Written out, that is:
Same operation, different notation. Y takes the role of the key and X the role of the value, and the paper itself calls the resulting state "fast weights" and connects it to a linear-attention view. Our demo swaps neuron indices for vector components and stops there.
The honest gap is the activations. BDH's Y and X are sparse and non-negative, roughly 5% of neurons active at a time (§6.4)[1]. Ours are dense Gaussian vectors whose signs cancel. Rather than list that as a caveat and move on, we measured it.
Does sparsity change the answer?
Mean recall at d=64, Hebbian, 40 seeds per cell. Run it yourself with npm run experiments.
| n | dense | 3/64 active (~5%) | 16/64 active (25%) |
|---|---|---|---|
| 16 | 0.905 | 0.909 | 0.738 |
| 32 | 0.823 | 0.833 | 0.634 |
| 64 | 0.712 | 0.716 | 0.560 |
| 128 | 0.579 | 0.578 | 0.517 |
At BDH's reported sparsity the two regimes agree to within noise, so the dense toy is a fair stand-in. The reason is in the geometry: two 3-of-64 sparse keys share an active coordinate only 14% of the time, and when they miss entirely their dot product is exactly zero. That protection from disjoint supports roughly cancels the penalty of losing sign cancellation. Push the density to 25% and the protection is gone, and recall falls apart. You can switch the whole page into sparse mode with the "key geometry" control above.
What is ours and what is theirs
From the BDH paper: the outer-product Hebbian update, the "fast weights" framing, and the ~5% activation sparsity. Those are quoted and cited.
Not from the BDH paper: every curve on this page. Pathway does not publish a recall-against-capacity measurement for BDH's synapse matrix, and we have not run BDH. This is our own toy model of the mechanism, at a scale small enough to check by hand.
Independently corroborated: the interference phenomenon itself. Variational Linear Attention[5] (2026) opens by noting that a linear attention state's norm grows with sequence length, "causing progressive interference between stored associations." Gated DeltaNet-2[4] (2026) exists to edit that compressed memory "without scrambling existing associations." We are demonstrating a known problem, not discovering one.
Where BDH-CQ comes in: its contextual memory accumulates additively as each demonstration arrives, the same write-don't-replace pattern one level up from tokens. Its headline result (29.5% pass@2 on ARC-AGI-1 at $0.0007 per task, 150M parameters)[2] is reported by its authors and has not been independently reproduced, by us or as far as we know by anyone else.
Why the state norm matters, and why our metric hides it
Measured at d=32: the Hebbian state's Frobenius norm tracks √n almost exactly (22.51 measured against 22.63 predicted at n=512), while the delta rule saturates near 5.6 and stops growing. That unbounded growth is exactly what Variational Linear Attention targets[5]. Our recall metric is cosine, which is scale-invariant, so it is blind to it by construction. In a real network the growing norm interacts with normalisation layers and finite precision and causes trouble that this page cannot show you. We think that is the most important limitation here, so we measured the norm separately instead of quietly leaving it out.
BDH-CQ
Everything up to here has treated piling associations into one matrix as the thing that goes wrong. It is also the thing that makes learning from demonstrations work.
BDH-CQ's contextual memory accumulates additively as each demonstration arrives. Its own paper connects that to fast-weight and linear-attention views of contextual association[2]. Additive accumulation has a property worth seeing directly: write the same association several times with different noise on it each time, and the signal is identical on every write so it adds coherently, while the noise is different on every write so it partly cancels.
Below, a rule is written repeatedly into a state that is also holding 24 unrelated associations. Nothing is trained. Nothing is told the rule. Turn the noise up and add demonstrations, and a fixed-size state sharpens up on something it only ever saw corrupted copies of.
Why this is the same claim, not a new one
The budget has not changed. Consistent evidence spends less of it, because ten writes of one association cost roughly what one write costs, while ten unrelated associations cost ten times as much. That is why in-context learning from many demonstrations is cheap for this kind of memory and holding many unrelated facts is expensive. Same fixed state, same arithmetic, read the other way round.
What this is and is not
This is a toy of the accumulation mechanism, not an implementation of BDH-CQ. BDH-CQ ingests real ARC-style demonstrations, has a separate latent workspace that iterates on the query, and reports 29.5% pass@2 on ARC-AGI-1 at $0.0007 per task with 150M parameters[2]. We have implemented none of that and reproduced none of it; the figure is its authors' and we cite it as theirs. What we can show is the one property the contextual state gets from being additive, at a scale where you can check it against a closed form: the conflicting-evidence table lands on a/√(a²+b²) because that is the angle between a weighted sum and one of its terms.
Landscape
The first three rows are running on this page. The fourth is here for context and we have not implemented it.
| Mechanism | Footprint | Exactness | Who keeps the fidelity | State norm | Seen in |
|---|---|---|---|---|---|
| Full KV cache | Grows with every token | Exact, always | Everyone, equally, forever | n/a | Standard Transformer decoding |
| Hebbian | d² · 8 bytes | Approximate | Spread evenly over all of history | grows as √n | BDH synaptic state |
| Delta rule | d² · 8 bytes | Exact for the newest write | Concentrated on recent writes | saturates | DeltaNet, Gated DeltaNet-2, Titans' objective |
| Decay gate | d² · 8 bytes | Approximate | Recent writes, geometrically | saturates lower | Titans' forgetting gate |
| Compressed recurrent state not implemented | Fixed by design | Approximate | Whatever the gating learns | design-dependent | Mamba-family state-space models |
Limits
Our cache is a lookup, not attention.
Real attention softmaxes over all stored keys and has its own failure modes. We index straight into the cache so that the only variable under test is whether the memory grows. That makes the cache look flawless in a way real attention is not.
Cosine hides the norm problem.
Scale-invariance is why Hebbian looks good here at high load despite its state norm growing without bound. Measured separately in the section above.
λ is fixed; Titans learns it.
Our decay gate uses one constant λ for every write. Titans computes αt per step from the input, along with a momentum term we have not implemented at all. Treat our decay row as the idea of forgetting, not as Titans.
One layer, no training.
BDH is a full architecture with learned projections feeding these synapses. We hand it random vectors. Learned keys would be spread out deliberately rather than randomly, which should help, and we have not measured by how much.
The theory curve is first-order.
√(d/(d+n−1)) assumes the crosstalk terms behave like independent isotropic noise. Good for dense Gaussian keys, visibly worse at small d, and it is not the right model for the sparse non-negative regime at all. The dashed line stays on the chart in sparse mode so you can watch it stop fitting.
Twenty facts is not a context window.
We shrank everything until the state fits on a screen. The mechanism is the same at d=16 and d=2048; the numbers are not, and nothing here should be read as a prediction about a production model's context length.
Check yourself
If these are easy, the page worked. If not, every one of them is answerable from the panels above.
And in your own words