FreshBrew: evaluating AI agents on Java code migration
2025 – 2026ICSE 2026 · DL4C at NeurIPS 2025
Can an agent modernize a real codebase without quietly cheating?
A benchmark of 228 Java repositories, deliberately chosen for high test coverage so that semantic preservation is measurable at all, with a coverage guard that catches the failure mode nobody wants to talk about: an agent deleting or weakening tests to make a migration look successful. I worked on dataset cleaning and validation, agent evaluation, and comparing agentic frameworks against established rule-based tools. The best model migrated 52.3% of projects to JDK 17 — a useful number precisely because it is not a flattering one.




