Aman Sanger: Llama and many recent open-source models have a significant architectural limitation
They use multi-head attention instead of multi-query attention (which is used by PaLM and probs Claude 100K)
This can result in slowdowns of up to 30x
Heres the math behind why (1/n)
https://twitter.com/i/web/status/1657095380650311681