Claude 3.7 Sonnet Thinking scores 33.5 (4th place after o1, o3-mini, and DeepSee...

zone411 70 days ago | parent | context | favorite | on: Claude 3.7 Sonnet and Claude Code

Claude 3.7 Sonnet Thinking scores 33.5 (4th place after o1, o3-mini, and DeepSeek R1) on my Extended NYT Connections benchmark. Claude 3.7 Sonnet scores 18.9. I'll run my other benchmarks in the upcoming days.

https://github.com/lechmazur/nyt-connections/