I finally tested GPT-6 Sol on a real work. 2 repos. 105 hidden bugs. Find and fixed what you can. It looks like a huge…
Summary
Testing GPT-6 Sol on bug-finding tasks across 2 repositories with 105 hard bugs revealed significant performance degradation compared to previous models, though it offers the lowest API cost at $9.93. When accounting for cost-efficiency, GPT-6 Sol performs comparably to medium-effort variants of better models, while newer variants like GPT-6 Luna show similarly disappointing results.
Summarized by ThreadOut AI from the full thread. May miss nuance — read the thread below.
- #1
- #2
- #3
- #4
- #5
- #6
- #7