Evaling Video Slop
Maor Bril
- When
- Wednesday, July 11:55 PM – 2:15 PM · 20 min
- Where
- Track 5San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
Everyone is shipping video models. Almost no one is evaling them honestly. CLIP score doesn't catch temporal incoherence. Vibes-based human review doesn't scale. And every "AI judge" you wire up will quietly drift away from human preference unless you measure the drift. This is a tactical talk on building real multimodal eval, using JudgeJudy (open-sourced at Character.ai) as the working example. You'll leave with: Why video is different from text. Temporal consistency, shot continuity, narrative coherence, and the metrics that actually capture each (clip_temporal, temporal_consistency, and friends). AI judges, the real version. Custom rubrics, when they work, when they hallucinate, when they collapse to a single dimension and pretend they didn't. The calibration loop. Pearson/Spearman correlation against human scores, automated rubric improvement, detecting systematic judge bias before it costs you a release. Pairwise preference models for video. Training a Qwen3-VL backbone with Bradley-Terry loss to score "is this slop?" before it ships. Regression gates in CI. How every AgentX release at Character.ai passes through an eval wall before it reaches users. Closing the loop with JudgeJudy. Correlating eval scores against real telemetry (Amplitude, Statsig) and feeding validated gates back into the runtime. If you're shipping any multimodal output and your eval strategy is still "the team watches some clips on Friday," this is the upgrade. github.com/character-ai/judgejudy
Speaker
Chaos Catalyst, Character.ai
Maor is a Principal Software Engineer at Character.ai, where he builds the agentic platform behind Stories, Streams, and the AI Social Feed (16M+ MAU). He open-sourced claude-agent-sdk-go and JudgeJudy, the multimodal eval harness that gates every AgentX release. Before Character.ai he led the Datastores org at Coinbase and shipped infrastructure at Netflix, Google, and VMware (via the Arkin acquisition). Twenty years of building systems that run in production. He writes about agentic systems and AI engineering on LinkedIn.
More in Evals
- Vending-Bench: Long-Horizon Agent Evals for a Simulated Vending BusinessWednesday, July 1 · 10:45 AM – 11:05 AM · Track 5
- From Signal to PR: Anatomy of a Self-Improving AgentWednesday, July 1 · 11:10 AM – 11:30 AM · Track 5
- Building Closed-Loop Evals for a Multimodal Agent at Uber ScaleWednesday, July 1 · 11:40 AM – 12:00 PM · Track 5
- From Agent Traces to Agent Simulations: The next era of agent evaluationWednesday, July 1 · 12:05 PM – 12:25 PM · Track 5
- Model Whisperers: How Evals and Prompts Shape Agent BehaviorWednesday, July 1 · 1:30 PM – 1:50 PM · Track 5
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.