Mercury 2.5 LLM hits 770 tokens per second

hackernews8.0

Mercury 2.5, a new large language model, achieves a generation speed of 770 tokens per second according to Artificial Analysis. This throughput makes it one of the fastest LLMs benchmarked on the platform.

Read full story at hackernews →