Spots

GitHub - Sparticle62ops/pssa: A custom AI architecture being developed in rust

PSSA is a small language model that is not a transformer. It reads text one token at a time through a recurrent state-space layer, keeps a bank of episodic memories it can look things up in, and rewrites part of its own weights while it runs. It is written in Rust from scratch, with no PyTorch, no TensorFlow, and no ML framework of any kind underneath it.

At matched parameters and on the same corpus

At matched parameters and on the same corpus, it learns faster than a transformer and generates text about twelve times quicker on the same CPU. How it differs from a transformer

A transformer scores every pair of tokens in

A transformer scores every pair of tokens in the context, so its cost per step grows with the square of the sequence length and the whole context is re-read at every step. PSSA carries one fixed-size state along the sequence in a single left-to-right pass, and looks things up in a memory bank instead of re-reading the context, so cost grows linearly with length.

Two models, same corpus, same tokenizer, same optimizer

Two models, same corpus, same tokenizer, same optimizer schedule, same seed, same number of parameters. One is PSSA, one is a standard transformer. Over 12.7M tokens of cleaned WikiText-103:

PSSA finished at 3.98 training cross-entropy, the transformer

PSSA finished at 3.98 training cross-entropy, the transformer at 4.43. That is a gap of 0.45 nats, perplexity 53.7 against 83.7. The transformer spent its entire 12.7M-token budget to reach a loss PSSA had already passed around 2M tokens in. The two curves never cross, and they never touch: It holds on text neither model has seen

Training loss only says a model fit the

Training loss only says a model fit the stream it was fed. So both checkpoints were scored on a 198,939-token slice cut from a part of the corpus neither run ever touched:

Every checkpoint of both runs, 64 PSSA links

Every checkpoint of both runs, 64 PSSA links and 43 transformer links, scored on a bounded 9,934-token window of that unseen slice. The curves never cross: PSSA is ahead from the first link and finishes 0.51 nats lower. The table below is the final checkpoint of each run on the full slice. The held-out gap, 0.43 nats, is essentially the training gap. PSSA is not memorizing harder, it is generalizing better. And it is much faster to run Generating 200 tokens on the same CPU, same prompt, same sampler:

A recurrent model carries a fixed-size state, so

A recurrent model carries a fixed-size state, so the cost of each new token does not grow with the length of what came before. A transformer re-reads its whole context every step. What is actually different about it

A recurrent state-space core. Learned continuous state matrices

A recurrent state-space core. Learned continuous state matrices carry information forward in a fixed-size state, instead of attention over the full context window.

An episodic memory bank. 512 slots with hyperbolic

An episodic memory bank. 512 slots with hyperbolic (Poincare-style) retrieval and bounded top-4 search, written to and read from during the run.

News

GitHub - Sparticle62ops/pssa: A custom AI architecture being developed in rust

PSSA is a small language model that is not a transformer.

@spots
Source: Hacker News
See more like this