▲ 1 Fusing a 27B ternary LLM's whole decode step into one CUDA kernel (twitter.com) by Jr23_xd | Jul 16, 2026 | 0 comments on HN Visit Link