Prompt, Memory, Weights: The Architecture Decisions Most AI Teams Make by Accident
Anant Srivastava
- When
- Wednesday, July 112:05 PM – 12:25 PM · 20 min
- Where
- Expo Stage 4 SESan Francisco, CA · imported from ai.engineer's public schedule feed
About this session
The interesting engineering in production AI isn't in the model. Your knowledge lives in files, databases, and APIs: docs, runbooks, conversations, code. The model just reads tokens. So the real architectural question is which path that knowledge takes to inference: into the prompt directly, into memory for retrieval on demand, or into the weights through fine-tuning. Most teams treat these as a ladder. Start with prompts, escalate to RAG, eventually fine-tune, as if each step is a more advanced version of the last. The field is converging on a different answer: they solve different problems. The prompt shapes behavior and constraints. Memory grounds the model in current, citable knowledge. Weights harden specialized reasoning and format. They're not substitutes you graduate between; they're complementary, and the failures come from using one to do another's job. Fine-tuning to teach the model facts it should have retrieved is the classic trap: you bake in knowledge that's stale the day it ships, and you still can't cite it. This is an opinionated take on all three: when each is the right call, when each is a trap, and the part most teams never build, the circulation between them. Memory that captures what the agent does becomes the dataset you fine-tune on; fine-tuning changes what's worth retrieving; the loop compounds. Get the three paths right and they stop being a pipeline you climb and start being an architecture that learns.
Speaker
Principal Technologist - Data and AI Platforms, Oracle
Anant Srivastava is a Principal Technologist for Data and AI Platforms at Oracle, focused on modern data architecture and AI platform decisions for production AI systems.
More in Context Engineering
- Build-Time vs. Run-Time: Why Your Dev Tools Will Fail in ProductionWednesday, July 1 · 10:45 AM – 11:05 AM · Track 8
- It’s Tokens All The Way Down: How RLMs are DifferentWednesday, July 1 · 11:10 AM – 11:30 AM · Track 8
- 500 Skills, Zero Fine-Tuning: LinkedIn's Playbook for AI Agents That Actually Know Your CodebaseWednesday, July 1 · 11:40 AM – 12:00 PM · Track 8
- Your agents lack context: Here's how to fix "You're absolutely right!"Wednesday, July 1 · 12:05 PM – 12:25 PM · Track 8
- How long can your skills be before your agent forgets what you told it?Wednesday, July 1 · 1:30 PM – 1:50 PM · Track 8
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.