June 29 – July 2, 2026 · San Francisco, CA · imported from ai.engineer's public schedule feed

AI Engineer World's Fair 2026 — unofficial import demo

Unofficial demo. This programme was imported from the AI Engineer World's Fair's own public schedule feed to show vibeboard at real conference scale. Not affiliated with, or endorsed by, the organisers.

All sessions
Workshops Day 1Workshop

Local LLMs and workstation agents: Part 1

Ahmad Osman

When
Monday, June 2911:05 AM – 12:05 PM · 60 min
Where
Track 6San Francisco, CA · imported from ai.engineer's public schedule feed
Google Calendar

About this session

Have you heard "Buy a GPU," "Opensource AI Must Win," or "Local AI FTW" before? This workshop will be a practical window into that confusing world and a practical map for understanding what different Local AI hardware is actually capable of and which models make sense on each class of machine.

Whether you are just getting started or already running models every day, we will demo and work through why a Mac mini, M4 Pro MacBook Pro, M5 Max MacBook Pro, RTX 5070 8GB laptop, Strix Halo box, DGX Spark, and 2x RTX PRO 6000 Blackwell machine should not be configured, benchmarked, or used the same way.

What are you trying to run? How much VRAM or Unified Memory do you actually need? When does a small machine make sense? When do you need a real GPU box? When does long context, tensor parallelism, or serving infrastructure start to matter?

This should be useful to everyone: people curious about local AI, people buying their first capable machine, people already running models, and people trying to use local inference for scalable agentic workflows.

We will close by showing how Codex can automate the boring part: give it my Inference Engine article, the hardware target, and the model of your choice, then ask it to propose the engine, environment, flags, batch settings, KV-cache settings, and benchmark and evaluation plan.

Speaker

Ahmad Osman
Ahmad Osman

Founder & CEO, Osmantic

Ahmad M. Osman is an AI researcher, systems engineer, and moderator of r/LocalLLaMA, where he helps a fast-growing community make local AI practical. A lifelong builder, he started coding at age 7, and by age 12 was running a private C++ MMORPG server from a Pentium 4 desktop. Today, Ahmad’s work sits at the intersection of LLMs, inference, hardware, infrastructure, and full-stack ownership. He holds dual degrees in Computer Science and Data Science, and is a prominent voice in modern AI infrastructure and self-hosted artificial intelligence. When not on stage, Ahmad can be found building Osmantic, his sovereign AI lab, where he serves as Founder and CEO.

More in Workshops Day 1