Your Moat Is Your Data Model
Mike Phipps
- When
- Thursday, July 211:40 AM – 12:00 PM · 20 min
- Where
- Track 5San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
Every enterprise AI team faces the same strategic question: where in the stack should a small team focus its effort? Models, frontends, and agent frameworks evolve rapidly and are increasingly commoditized. But regardless of how these layers mature, AI in enterprise settings remains bottlenecked by the same underlying problem: structured data is siloed across systems of record with domain-specific schemas, and the unstructured data needed to contextualize it sits in entirely separate systems, with its own systematic complexities. The durable work is cleaning, curating, and semantically modeling this data in an AI-first manner so that any client — chat, workflow, or otherwise — can query across it. That's the moat. At the Gates Foundation, my team built and deployed our foundation-wide knowledge graph on Neo4j that unifies structured and unstructured data behind a single MCP server. The graph itself is modeled for agentic consumption: natural hierarchies are projected as traversable paths rather than flattened tables, and unstructured documents are semantically chunked, tagged, and mapped to structured entities at ingestion time using AI-driven ETL. The result is a semantic layer where an agent can express a complex cross-system question as a concise graph query and receive an accurate answer. This talk is an architectural walkthrough covering the end-to-end pipeline: AI-based extraction and semantic chunking of unstructured documents, the agent-first data modeling decisions, design considerations for our MCP server, and how we handle graph-based retrieval evals. We'll walk through real query sessions showing Claude interacting with the graph through both chat and workflow integrations. The intended takeaway is a practical framework for where a small enterprise team's investment compounds — and why that investment is the data model, not the layers above it.
Speaker
Lead AI Engineer, Gates Foundation
Mike is the lead AI engineer in business operations at the Gates Foundation, where he built and deployed SIP (Strategic Intelligence Platform), the foundation's enterprise-wide knowledge graph. Built on Neo4j and served to Claude through MCP, SIP unifies siloed structured records and unstructured documents into one semantic layer that agents can query directly. Before AI engineering, he earned his PhD at CERN in experimental high-energy nuclear physics, working with some of the largest datasets in science. His work now centers on AI-first data modeling, driven by the conviction that for most enterprise teams the durable moat comes from properly modeling and expressing their data assets and domain knowledge — not the commoditizing layers above.
More in Graphs
- Thinner Agents on a Smarter Substrate: The Ontology-based Semantic LayerThursday, July 2 · 10:20 AM – 10:30 AM · Main Stage
- CrabRAG: Why Automated Assistants Need Graph Memory, Not More TokensThursday, July 2 · 10:45 AM – 11:05 AM · Track 5
- Active Graph Agent Runtime (BabyAGI 4)Thursday, July 2 · 11:10 AM – 11:30 AM · Track 5
- From Systems of Record to Systems of ContextThursday, July 2 · 12:05 PM – 12:25 PM · Track 5
- AI : Learned Execution Graphs for Real-Time Anomaly Detection & Drift Classification in APIsThursday, July 2 · 1:30 PM – 1:50 PM · Track 5
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.