Open-weight reasoning models close the gap, a new eval harness, and what changed in evaluation methodology this week.
This is placeholder seed content for the Phase 3 build. Real issues are only published here once the newsletter provider is connected (Phase 5) and the issue has actually been sent.
