June 29 – July 2, 2026 · San Francisco, CA · imported from ai.engineer's public schedule feed

AI Engineer World's Fair 2026 — unofficial import demo

Unofficial demo. This programme was imported from the AI Engineer World's Fair's own public schedule feed to show vibeboard at real conference scale. Not affiliated with, or endorsed by, the organisers.

All sessions
InferenceSession

Weight Folding, CUDA Streams, and the Bug That Made My Model Speak Backwards

Filip Makraduli

When
Thursday, July 23:45 PM – 4:05 PM · 20 min
Where
Track 9San Francisco, CA · imported from ai.engineer's public schedule feed
Google Calendar

About this session

A talk about contributing GPU benchmarks to an open-source research paper (FlashNorm). I'll walk through the engineering journey: folding norm weights into projections, writing Triton kernels, accidentally making attention bidirectional (oops), and ultimately proving a 33-35% speedup on the norm+project operation. Practical lessons for anyone trying to optimize transformer inference.

Speaker

Filip Makraduli
Filip Makraduli

Founding Member of Technical Staff, Superlinked

Filip Makraduli is an applied AI researcher and founding ML Developer Relations engineer at Superlinked, where he designs and ships small‑LLM inference systems for search, retrieval, and agents in production. He holds a master’s degree in Biomedical Data Science from Imperial College London. Before Superlinked, Filip worked in machine learning, data science, and developer relations roles across early‑stage AI startups and larger enterprises, building language understanding, retrieval‑augmented generation (RAG), and LLM pipeline tooling while partnering closely with product and platform teams. He is a frequent open‑source contributor, with contributions to kernel libraries, model‑inference providers, and hands‑on demos used by practitioners. Filip is a co‑author of several publications on efficient transformer architectures and inference, including work on faster normalization for LLMs. He is an experienced speaker at meetups and conferences such as AI Engineer Europe and Berlin Buzzwords, sharing practical lessons on efficient transformers, retrieval systems, and embedding inference for production AI teams.

More in Inference