JAY ZENITH

I build post-training and evaluation systems for tool-using agents: data, environments, verifiable rewards, and RL training infrastructure. My work below.

  • predict

    Early-stage experiment, inspired by ECHO's world-model result: is world modeling the right direction for coding agents, and does a prediction need its own cross-entropy supervision because RLVR's reward is too crude to grade one? [full write-up] [code]