Lucas Sun

Lucas Sun

I write about math, machine learning, and whatever else I find interesting.

3 posts shown

An Associative Introduction to Deep Learning

One sum, values weighted by key-query brackets, accounts for almost every parameter in a modern network. Starting from the correspondence between a single neuron and a rank-one outer product, this post uses bra-ket notation to derive MLPs, softmax attention, effective rank, linear attention, the delta rule, Gated DeltaNet and PaTH attention as one associative memory wearing different coefficients.