JAY ZENITH
I build post-training and evaluation systems for tool-using agents: data, environments, verifiable rewards, and RL training infrastructure. My work below.
- predict
Early-stage experiment: can a coding agent predict its own patch's fate before running it, grounded only in what the environment verifies, never its own guess? [full write-up] [code]