LLM Inference Engine
GPT-2 inference written from scratch in C++, optimised as a measured ladder from a naive forward pass to int4 quantisation.
C++Python
Each one has a writeup: how it works, what I tried that didn’t, and where the numbers came from.
2 of 6 projects
GPT-2 inference written from scratch in C++, optimised as a measured ladder from a naive forward pass to int4 quantisation.
A pipelined RV32I core on an FPGA, plus the constrained-random rig that proves it correct.