They’re actually not because they have randomness built in. Since we are working with probabilities, it won’t always pick the next token that has the highest probability and the randomness can be tuned via a “temperature” setting to make it more or less likely that it will choose the most probable token.
The weights for the model could be stored in a firmware chip but you still need RAM because it pulls all the weights into RAM in order to perform the calculations.
The LLM itself is deterministic. It outputs a vector that we interpret as a probability distribution over the set of tokens. It’s the program using the output of the LLM, such as a chatbot program, that selects an individual token using (or not using) these vector elements as weights.
The weights do not change once the model is trained. This is why I am suggesting they could be incorporated directly into the structure of an ASIC for a specific model, rather than storing them in memory.
Edit: Of course another major factor could be that the models are just to big to be wholly implemented in a single IC by any currently existing manufacturer.
They’re actually not because they have randomness built in. Since we are working with probabilities, it won’t always pick the next token that has the highest probability and the randomness can be tuned via a “temperature” setting to make it more or less likely that it will choose the most probable token.
The weights for the model could be stored in a firmware chip but you still need RAM because it pulls all the weights into RAM in order to perform the calculations.
The LLM itself is deterministic. It outputs a vector that we interpret as a probability distribution over the set of tokens. It’s the program using the output of the LLM, such as a chatbot program, that selects an individual token using (or not using) these vector elements as weights.
The weights do not change once the model is trained. This is why I am suggesting they could be incorporated directly into the structure of an ASIC for a specific model, rather than storing them in memory.
Edit: Of course another major factor could be that the models are just to big to be wholly implemented in a single IC by any currently existing manufacturer.