Weight Folding, CUDA Streams, and the Bug That Made My Model Speak Backwards
Filip Makraduli
- When
- Thursday, July 23:45 PM – 4:05 PM · 20 min
- Where
- Track 9San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
A talk about contributing GPU benchmarks to an open-source research paper (FlashNorm). I'll walk through the engineering journey: folding norm weights into projections, writing Triton kernels, accidentally making attention bidirectional (oops), and ultimately proving a 33-35% speedup on the norm+project operation. Practical lessons for anyone trying to optimize transformer inference.
Speaker
Founding Member of Technical Staff, Superlinked
Filip Makraduli is an applied AI researcher and founding ML Developer Relations engineer at Superlinked, where he designs and ships small‑LLM inference systems for search, retrieval, and agents in production. He holds a master’s degree in Biomedical Data Science from Imperial College London. Before Superlinked, Filip worked in machine learning, data science, and developer relations roles across early‑stage AI startups and larger enterprises, building language understanding, retrieval‑augmented generation (RAG), and LLM pipeline tooling while partnering closely with product and platform teams. He is a frequent open‑source contributor, with contributions to kernel libraries, model‑inference providers, and hands‑on demos used by practitioners. Filip is a co‑author of several publications on efficient transformer architectures and inference, including work on faster normalization for LLMs. He is an experienced speaker at meetups and conferences such as AI Engineer Europe and Berlin Buzzwords, sharing practical lessons on efficient transformers, retrieval systems, and embedding inference for production AI teams.
More in Inference
- Operating Distributed Inference Systems at ScaleThursday, July 2 · 10:45 AM – 11:05 AM · Track 9
- Routing LLM Inference in Production: From Engine Signals to PolicyThursday, July 2 · 11:10 AM – 11:30 AM · Track 9
- Are LLM Performance Benchmarks Reliable?Thursday, July 2 · 11:40 AM – 12:00 PM · Track 9
- All the Things We Have to Do to Satisfy Your Insatiable Need for TokensThursday, July 2 · 11:40 AM – 12:00 PM · Leadership 1
- Vertical Mobility: Building an AI Inference Platform That Scales from MVP to Trillion-Parameter WorkloadsThursday, July 2 · 12:05 PM – 12:25 PM · Track 9
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.