Skill issue: stop deploying vision language models, use them with Skills to build e2e vision apps on edge
Merve Noyan
- When
- Tuesday, June 3011:40 AM – 12:00 PM · 20 min
- Where
- Track 2San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
With the boom of vision language models barrier of entry to build vision apps are much lower so developers tend to use them right away. However, these models are very large and inefficient in production. In this talk, I will go through combining vision language models with Skills to build end-to-end vision apps from training to deployment using HF Skills, on top of showing the state-of-the-art in small computer vision/multimodal models.
Speaker
MLE, Hugging Face
Works at Hugging Face open-source team, author of the book Vision Language Models with Hugging Face published by O'Reilly.
More in Vision & OCR
- The State of VisionTuesday, June 30 · 10:45 AM – 11:05 AM · Track 2
- Building the Document Context Layer for AI AgentsTuesday, June 30 · 11:10 AM – 11:30 AM · Track 2
- Modality Misalignment and Originality Attribution in Short-Form Video: A Multi-Agent Approach at Platform ScaleTuesday, June 30 · 12:05 PM – 12:25 PM · Track 2
- From Ingestion to Agents: How Leading AI Teams Build on Document IntelligenceTuesday, June 30 · 1:30 PM – 1:50 PM · Track 2
- The Best Models Still Reason Like ToddlersTuesday, June 30 · 1:55 PM – 2:15 PM · Track 2
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.