Local LLMs and workstation agents: Part 1
Ahmad Osman
- When
- Monday, June 2911:05 AM – 12:05 PM · 60 min
- Where
- Track 6San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
Have you heard "Buy a GPU," "Opensource AI Must Win," or "Local AI FTW" before? This workshop will be a practical window into that confusing world and a practical map for understanding what different Local AI hardware is actually capable of and which models make sense on each class of machine.
Whether you are just getting started or already running models every day, we will demo and work through why a Mac mini, M4 Pro MacBook Pro, M5 Max MacBook Pro, RTX 5070 8GB laptop, Strix Halo box, DGX Spark, and 2x RTX PRO 6000 Blackwell machine should not be configured, benchmarked, or used the same way.
What are you trying to run? How much VRAM or Unified Memory do you actually need? When does a small machine make sense? When do you need a real GPU box? When does long context, tensor parallelism, or serving infrastructure start to matter?
This should be useful to everyone: people curious about local AI, people buying their first capable machine, people already running models, and people trying to use local inference for scalable agentic workflows.
We will close by showing how Codex can automate the boring part: give it my Inference Engine article, the hardware target, and the model of your choice, then ask it to propose the engine, environment, flags, batch settings, KV-cache settings, and benchmark and evaluation plan.
Speaker
Founder & CEO, Osmantic
Ahmad M. Osman is an AI researcher, systems engineer, and moderator of r/LocalLLaMA, where he helps a fast-growing community make local AI practical. A lifelong builder, he started coding at age 7, and by age 12 was running a private C++ MMORPG server from a Pentium 4 desktop. Today, Ahmad’s work sits at the intersection of LLMs, inference, hardware, infrastructure, and full-stack ownership. He holds dual degrees in Computer Science and Data Science, and is a prominent voice in modern AI infrastructure and self-hosted artificial intelligence. When not on stage, Ahmad can be found building Osmantic, his sovereign AI lab, where he serves as Founder and CEO.
More in Workshops Day 1
- Cooking with CodexMonday, June 29 · 9:00 AM – 11:00 AM · Track 3
- The best SDLC is the one you build yourself: Why orchestration changes everythingMonday, June 29 · 9:00 AM – 11:00 AM · Track 4
- AI Security Engineer Foundations + CertificateMonday, June 29 · 9:00 AM – 11:00 AM · Track 5
- Total Recall: Agent Memory and Harness EngineeringMonday, June 29 · 9:00 AM – 11:00 AM · Track 6
- Open-Source Inference Engineering for the Agentic EraMonday, June 29 · 9:00 AM – 11:00 AM · Track 8
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.