A few months ago, I saw this excellent video about mechanistic interpretability of large language models. The video goes over a language model trained on an exceedingly simple task and inspects its weights and activations in detail to understand exactly what is happening under the hood. This appealed to me for two reasons: 1. I […]