Skip to content
Paweł Huryn@PawelHuryn · Sep 23, 2026

I finally tested GPT-6 Sol on a real work. 2 repos. 105 hidden bugs. Find and fixed what you can. It looks like a huge…

7 tweets1 min read♥ 2Koriginal

Summary

Testing GPT-6 Sol on bug-finding tasks across 2 repositories with 105 hard bugs revealed significant performance degradation compared to previous models, though it offers the lowest API cost at $9.93. When accounting for cost-efficiency, GPT-6 Sol performs comparably to medium-effort variants of better models, while newer variants like GPT-6 Luna show similarly disappointing results.

Summarized by ThreadOut AI from the full thread. May miss nuance — read the thread below.

  1. #1

    I finally tested GPT-6 Sol on a real work. 2 repos. 105 hidden bugs. Find and fixed what you can. It looks like a huge degradation. The results: - GPT-6 Astra (max): 45 - GPT-5.6 Sol (max): 43.5 - Opus 5.5 (max): 41.7 - Muse Spark 1.3 (max): 32.2 - GPT-6 Sol (max): 29.3 Until you measure the cost (API-equivalent): - GPT-6 Astra (max): $33.04 - GPT-5.6 Sol (max): $95.25 - Opus 5.5 (max): $58.53 - Muse Spark 1.3 (max): $18.11 - GPT-6 Sol (max): $9.93 More effort levels (xhigh, high, medium, low) dropping in this thread today 🧵

  2. #2

    BTW you may also like this thread about Opus 5.5: Notes: - Those are not 105 random bugs but hard problems frontier models missed at the beginning or 2026. - Unplanted bugs aren't counted, as OpenAI bugmaxes and reports many real, theoretical, and irrelevant issues solving which would complicate the solution. - Judges from different model families score the solution against the secret answer key. Identified unfixed bugs or incomplete solutions score 0. - More models, effort levels, and data available on the live website: bughunt[.]productcompass[.]pm

    Paweł Huryn@PawelHuryn · Sep 22, 2026

    So I tested Opus 5.5 (max) on real work. 2 repos. 105 planted bugs. Find and fix what you can. The results: GPT-6 Astra (max): 45 for $33.03 (n=3) Fable 5.1 (max): 43 for $77.55 Opus 5.5 (max): 43 for $60.49 Opus 5.0 (max): 27 for $51.33 Muse Spark 1.3 (max): 32.2 for $18.11 Anthropic is back in the game. More effort levels and attempts (n=3) dropping in this thread tonight 🧵

  3. #3

    More data: - GTP-6 Sol (max) ~= GPT-5.6 Sol (medium) - GPT-6 Sol (max) < Opus 5.5 (medium) Live benchmarks: bughunt.productcompass.pm

  4. #4

    All effort levels for GPT-6 Sol. n=3 for max, n=1 for the other effort levels. It draws almost a straight line in the chart. Coming next, starting with max: - GPT-6 Luna - GPT-6 Terra

  5. #5

    GPT-6 Luna (max) after n=1. Slightly better than Qwen3.8-27B you can run locally.

  6. #6

    Confirmed. GPT-6 Luna (max): 18.3 for n=3. The degradation was real.

  7. #7

    Situation today: GPT-6 Sol is like GPT-5.6 Terra GPT-6 Luna is like GPT-5.6 Asteroid