In this episode of convos.dev, Jess brings a blog post that got almost no attention online and deserved way more: Henry Pan's "What 1,000+ Harness Experiments Taught Me About Self-Improving Agents." She walks Chris through Pan's three-loop framework — inner, middle, and outer — and the messy reality of letting an agent improve its own harness: outer loops that cheat, written instructions that don't stop them, and the deterministic supervisor that finally did.
They dig into when a new harness actually deserves promotion (binomial tests), the gambler's fallacy in evals, why you can never trust a single run, and the caching bug you only catch by reading your own logs. Plus life updates: Chris turns a $20 Casio into a semi-smart watch, and Jess is standing again after six weeks non-weight-bearing.
0:00 Intro
0:17 Jess's pick: Henry Pan's 1,000+ harness experiments
1:32 The framework: inner, middle, and outer loops
2:44 When do you promote a harness? Binomial tests
4:09 1,200 runs, fewer than 100 kept
4:42 The outer loop likes to cheat
5:40 The deterministic supervisor: constraints in code, not prompts
6:13 Why 200 isolated rules don't work
7:19 A shared learnings file (the Ralph Wiggum pattern)
7:52 Cracking self-improvement: scary but awesome
9:01 What tasks and models was he actually running?
9:43 You're improving the harness, not the LLM
10:14 The gambler's fallacy in evals
11:30 How many runs until you can trust the number?
11:57 Jess's webinar scare: never trust one run
13:00 The 72% cache-hit bug: go read your logs
14:13 "Where can things go wrong": ten bottlenecks
15:54 Could this train small, specialized models?
16:42 Stronger models compose rules into one big brain
17:34 Data over hype: who deserves the attention
18:55 Life updates
19:04 Chris's hacked Casio: the Ollee Watch board
22:15 Jess can stand again
23:48 A first-ever PT session (and a startup pitch)
24:23 Aging athletes: bodies need maintenance now
26:16 Three minutes vs. a month: the monitor research story
27:16 Outro
Links
- Henry Pan — "What 1,000+ Harness Experiments Taught Me About Self-Improving Agents" : https://www.henrypan.com/blog/2026-05-25-self-improvement-harness/
- Henry Pan — harness-experiment (project repo) : https://github.com/workofart/harness-experiment
- Ollee Watch — smart mainboard for the Casio F-91W : https://www.olleewatch.com
- Submit questions and feedback : https://convos.dev