google / gemma.cpp

lightweight, standalone C++ inference engine for Google's Gemma models.
Apache License 2.0
5.94k stars 502 forks source link

Integrate matmul into FFW: 4.3x prefill speedup #243

Closed copybara-service[bot] closed 3 months ago

copybara-service[bot] commented 3 months ago

Integrate matmul into FFW: 4.3x prefill speedup

before, bf16:
27.2929 prefill tokens / sec
17.2114 tokens / sec

after, bf16
116.496 prefill tokens / sec
17.5391 tokens / sec