Intel ships programmable SHAVE cores inside its NPUs, but the public stack exposes only graph-level programming. npunlock reconstructs the missing path from custom C code to a runnable NPU kernel. The current implementation has been verified on Windows x64 with Meteor Lake / NPU3720. Latest breakthrough — 2026-09-23: One native graph can execute independent FP32-unary and FP16-binary custom branches; explicit ACT-group preflight handles the compiler's branch reordering. Evidence and limits . Quick example This complete FP32 GELU example embeds the C kernel in Python, places it in an NPU graph, and checks the result against NumPy. The bundled npunlock/npu3720_kernel.h target header supplies the NPU3720 invocation and tensor-address helpers. The tested MoviTools toolchain makes most conventional libm functions available to kernels without including <math.h> ; this example calls tanhf directly. See the mlibm.a symbol inventory for the observed candidates.…