Skip to content
Paweł Huryn@PawelHuryn · Sep 4, 2026

So, I finally tested GPT-6 Astra. 2 real repos, 105 bugs, find and fix what you can. Max effort: GPT-6 Astra: 48/105…

7 tweets1 min read♥ 2.5Koriginal

Summary

Testing GPT-6 Astra on a bug-finding benchmark with 105 bugs across 2 real repositories, it found 48 bugs while being 2x faster and cheaper than competing models like GPT-5.6 Sol and Fable 5.1. The author recommends a hybrid routing approach using Astra for implementation work and Fable or Sol for code review, as each model has different strengths and blind spots.

Summarized by ThreadOut AI from the full thread. May miss nuance — read the thread below.

  1. #1

    So, I finally tested GPT-6 Astra. 2 real repos, 105 bugs, find and fix what you can. Max effort: GPT-6 Astra: 48/105 Fable 5.1: 43/105 GPT-5.6 Sol: 42/105 Gemini 3.8 Flash: 20/105 And it's crazy efficient. Look at the time and cost 🧵

  2. #2

    GPT-6 Astra was 2x faster than GPT-5 6 Sol. It was also 2x faster and cheaper than GPT-5.6 Sol and Fable 5.1. Still Astra alone seems not enough.

  3. #3

    GPT-6 Astra has completely different blind spots than other models. My new routing: - Implementation: Astra - Review: Fable or Sol

  4. #4

    Data, including .json, and live stats as always: bughunt.productcompass.pm More effort levels soon. Notes / FAQ: - I don't count unplanted bugs. OpenAI models report many real irrelevant issues ("bug-maxing") - Those are real bugs frontier model missed at the beginning of 2026 - Answer keys aren't public. That's why the banchmark works - I will soon share my harness, so you can try it with your own repos

  5. #5

    BTW, you can download my AI infographics by subscribing here: claudepdfs.io (Most are not specific to Claude)

  6. #6

    Other effort levels are in: bughunt.productcomoass.pm

  7. #7

    New: you can star the benchmark on GitHub. Just published our website and the data. We start from 0. github.com/phuryn/bug-hun…