Post by Ethan Mollick on X
Ethan Mollick@emollick
XThis may end up being a big deal:
Usually LLMs just predict the next token in a sequence, one at a time, but if you have them predict the next several tokens at once you get significantly better performance, faster, and with no added costs. The gains are better for bigger models

1.5K likes26 repliesPosted May 1, 2024