deependu
Deependu

Deependu Jha

Open source @ Lightning AI

I write CUDA kernels, poke at PyTorch compiler internals, and try to make models run faster than they have any right to. This is where I write about how that stuff actually works, and share the projects that come out of it.

LLMsCUDATorch CompilersPerformance

ghost-rfdetr

flagship

Custom CUDA kernels and operator fusion for RF-DETR — faster inference out of the box, no retraining or config required.

CUDAKernel FusionRF-DETRInference