Can a language model run in Linux eBPF?
kept by hikaruchan
Running Qwen3-0.6B’s 28 decoder layers in Linux eBPF shows token inference possible but slow (~1.2 s), highlighting tradeoffs in fixed‑point math, memory, and INT4 quantization.