The LLM itself is deterministic. It outputs a vector that we interpret as a probability distribution over the set of tokens. It’s the program using the output of the LLM, such as a chatbot program, that selects an individual token using (or not using) these vector elements as weights.
The weights do not change once the model is trained. This is why I am suggesting they could be incorporated directly into the structure of an ASIC for a specific model, rather than storing them in memory.
Edit: Of course another major factor could be that the models are just to big to be wholly implemented in a single IC by any currently existing manufacturer.
The LLM itself is deterministic. It outputs a vector that we interpret as a probability distribution over the set of tokens. It’s the program using the output of the LLM, such as a chatbot program, that selects an individual token using (or not using) these vector elements as weights.
The weights do not change once the model is trained. This is why I am suggesting they could be incorporated directly into the structure of an ASIC for a specific model, rather than storing them in memory.
Edit: Of course another major factor could be that the models are just to big to be wholly implemented in a single IC by any currently existing manufacturer.