Reads the published checkpoint straight from the Hugging Face CDN, in your browser, and checks that a projection really does hold only {−s, 0, +s}. No server, no inference — just the bytes.
A safetensors file begins with a JSON header listing every tensor, its dtype, its shape and its byte offsets. This page fetched that header with an HTTP range request, then fetched only the few megabytes belonging to the projection you picked — never the full 1.2 GB — decoded the bf16 values, and counted them.
Every linear projection in this model is balanced ternary: one scale s per layer, and every weight is −s, 0 or +s. Against 8-bit activations that makes each product −x, 0 or +x, so the matrix multiplies collapse into subtract, skip and add — no multiplier, zero DSP slices. Roughly a third of the weights are exact zeros the array simply skips. The 2-bit packed form the FPGA actually consumes is under hardware/ in the model repo, alongside a NumPy reference for element-by-element checking.