Articles
Notes on kernels, alignment and the systems in between.
Deep dives, experiments and things I worked out the hard way.

Can we explain most of the world with less information?
A reflection on information, perception, AI, molecular dynamics, and whether complex systems can be understood through much smaller latent representations.

Positional Encoding in Transformers: From Sinusoidal to RoPE
Understand positional encoding in Transformers, from sinusoidal embeddings to RoPE, and learn how models capture token order and relative position.

PyTorch Compiler Explained: TorchDynamo, AOTAutograd & TorchInductor
Understand how PyTorch compilation works under the hood: bytecode capture with TorchDynamo, forward/backward staging with AOTAutograd, and kernel optimization with TorchInductor.

Understanding Attention: The Idea Behind Modern AI
One paper changed everything: Attention Is All You Need. This post breaks down the foundations behind attention, starting from embeddings. A simple journey from words to meaning.

What Are Transformers in AI?
Learn what Transformers are, how they work, and why they power models like GPT and modern AI systems. A clear, beginner-friendly introduction.