Skip to content
Thread
Out
Recently unrolled
Topics
How it works
Unroll
Benchmark
1 unrolled thread about benchmark, each one readable on a single page.
Popular
Recent
Paweł Huryn
@PawelHuryn
So, I finally tested GPT-6 Astra. 2 real repos, 105 bugs, find and fix what you can. Max effort: GPT-6 Astra: 48/105 Fable 5.1: 43/105 GPT-5.6 Sol: 42/105 Gemini 3.8 Flash: 20/105 And it's crazy e...
Summary
Testing GPT-6 Astra on a bug-finding benchmark with 105 bugs across 2 real repositories, it found 48 bugs while being 2x faster and cheaper than competing models like GPT-5.6 Sol and Fable 5.1. The author recommends a hybrid routing approach using Astra for implementation work and Fable or Sol for code review, as each model has different strengths and blind spots.
Sep 4, 2026
7 tweets
♥ 2.5K