JAY ZENITH

I build post-training and evaluation systems for tool-using agents: data, environments, verifiable rewards, and RL training infrastructure. My work below.

  • predict

    Early-stage experiment, inspired by ECHO's world-model result: can a coding agent's prediction of its own patch's fate be grounded by the environment alone, trained only against the verified outcome, never the model's own sampled guess? [full write-up] [code]