Announcement_19

New paper: “SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution” (Corresponding Author). SkillTV-Bench evaluates whether LLM- and agent-based judges can verify complete skill-augmented agent executions, while SkillTV-Evolve automatically refines a reusable JudgeSkill for more reliable trajectory verification and rollout selection. arXiv · Code