Skip to content
Paweł Huryn@PawelHuryn · Sep 29, 2026

So I tested GPT-6.1 Sol on real work: 2 repos, 105 planted bugs. Find and fix what you can. Unlike GPT-6 Sol, which…

2 tweets1 min read♥ 2.4Koriginal

Summary

A developer tested GPT-6.1 Sol on finding and fixing planted bugs across two repositories, comparing it against other models. GPT-6.1 Sol achieved competitive results (44/105 bugs) at a fraction of the cost ($6.56) compared to more expensive alternatives, making it a cost-effective option for code analysis tasks.

Summarized by ThreadOut AI from the full thread. May miss nuance — read the thread below.

  1. #1

    So I tested GPT-6.1 Sol on real work: 2 repos, 105 planted bugs. Find and fix what you can. Unlike GPT-6 Sol, which was just a nerfed GPT-5.6 Terra, this one is real. The results: - GPT-6 Astra (max): 45 for $33 - GPT-6.1 Sol (max): 44 for $6.56 - GPT-5.6 Sol (max): 43.5 for $95.35 - Opus 5.5 (max): 41.7 for $58.53 - GPT-6 Sol (max): 29.3 for $9.33 With a model like this, the 50% cut to the $200 plan doesn't matter. n=1. More runs and results at more effort levels dropping in this thread over the next few hours 🧵

  2. #2

    This will take a few more hours. The results for the xhigh below. GPT-6.1 Sol is virtually free.