June 29 – July 2, 2026 · San Francisco, CA · imported from ai.engineer's public schedule feed

AI Engineer World's Fair 2026 — unofficial import demo

Unofficial demo. This programme was imported from the AI Engineer World's Fair's own public schedule feed to show vibeboard at real conference scale. Not affiliated with, or endorsed by, the organisers.

All sessions
EvalsSponsor Session

Evaling Video Slop

Maor Bril

When
Wednesday, July 11:55 PM – 2:15 PM · 20 min
Where
Track 5San Francisco, CA · imported from ai.engineer's public schedule feed
Google Calendar

About this session

Everyone is shipping video models. Almost no one is evaling them honestly. CLIP score doesn't catch temporal incoherence. Vibes-based human review doesn't scale. And every "AI judge" you wire up will quietly drift away from human preference unless you measure the drift. This is a tactical talk on building real multimodal eval, using JudgeJudy (open-sourced at Character.ai) as the working example. You'll leave with: Why video is different from text. Temporal consistency, shot continuity, narrative coherence, and the metrics that actually capture each (clip_temporal, temporal_consistency, and friends). AI judges, the real version. Custom rubrics, when they work, when they hallucinate, when they collapse to a single dimension and pretend they didn't. The calibration loop. Pearson/Spearman correlation against human scores, automated rubric improvement, detecting systematic judge bias before it costs you a release. Pairwise preference models for video. Training a Qwen3-VL backbone with Bradley-Terry loss to score "is this slop?" before it ships. Regression gates in CI. How every AgentX release at Character.ai passes through an eval wall before it reaches users. Closing the loop with JudgeJudy. Correlating eval scores against real telemetry (Amplitude, Statsig) and feeding validated gates back into the runtime. If you're shipping any multimodal output and your eval strategy is still "the team watches some clips on Friday," this is the upgrade. github.com/character-ai/judgejudy

Speaker

Maor Bril
Maor Bril

Chaos Catalyst, Character.ai

Maor is a Principal Software Engineer at Character.ai, where he builds the agentic platform behind Stories, Streams, and the AI Social Feed (16M+ MAU). He open-sourced claude-agent-sdk-go and JudgeJudy, the multimodal eval harness that gates every AgentX release. Before Character.ai he led the Datastores org at Coinbase and shipped infrastructure at Netflix, Google, and VMware (via the Arkin acquisition). Twenty years of building systems that run in production. He writes about agentic systems and AI engineering on LinkedIn.

More in Evals