Skip to content
Paweł Huryn@PawelHuryn · Sep 10, 2026

So I tested DeepSeek V4.1 Flash on real tasks. 2 repos, 105 hidden bugs, find and fix what you can. Opus 5 (max): 27…

7 tweets1 min read♥ 2.2Koriginal
  1. #1

    So I tested DeepSeek V4.1 Flash on real tasks. 2 repos, 105 hidden bugs, find and fix what you can. Opus 5 (max): 27 Grok 4.6 (max): 27 DeepSeek V4.1 Flash (max): 24 GPT-5.6 Luna (xhigh): 23 Opus 5 (high): 21 A really strong model for everyday tasks. And look at the cost: Opus 5 (max): $51.33 Grok 4.6 (max): $16.96 DeepSeek V4.1 Flash (max): $1.80 GPT-5.6 Luna (xhigh): $2.50 Opus 5 (high): $38.77 It's at the Pareto frontier. Other effort levels (high) and extra tests for DeepSeek V4 Pro dropping every hour in this thread 🧵

  2. #2

    FAQ: - Those are real, tough bugs frontier models struggled with at the beginning of 2026. Don't read it as "the best model solved only 45/105 issues" - Unplanted bugs aren't counted, OpenAI models report many real but irrelevant issues. Solving them would often complicate the solution. That would inflate their scores. - Every model is judged by another model family. I confirmed judges calibration. - Answer keys aren't public. Anonymized runs are linked in GitHub. - Yes, Luna (max) is a crazy good and cost-effective model. - I can't test Muse Spark 1.3, payments don't work in Europe + OpenRouter has unreliable effort levels (ignored). A live benchmark is always here: bughunt.productcompass.pm If you like it, star our repo on GitHub: github.com/phuryn/bug-hun…

  3. #3

    DeepSeek-V4.1-Flash (high) is in: max effort: 24/105, $1.80, 42.6 min high effort: 19/105, $0.31, 26.1 min For comparision: Gemini 3.8 Flash (high), 20/105, $9.78, 29.8 min GLM-5.3 (default): 19/105, $19.73, 66.7 min Coming next: deepseek-v4-pro effort=max deepseek-v4-pro effort=high Updates are published also on GitHub: github.com/phuryn/bug-hun…

  4. #4

    I retested the previous DeepSeek model with effort=max. The jump is signifficant: DeepSeek V4 Pro (max): 16/105, $1.89, 37 min DeepSeekV4.1 Flash (max): 24/105, $1.08, 42.6 min

  5. #5

    Update: V4.1 Flash is clearly superior to V4 Pro:

  6. #6

    BTW want to learn how to build AI products? 8 AI PM cohorts in the price of one, with people like Rohan Varma (OpenAI), Shubham Saboo (Google), Reah Miyara (Google), and Amit Fulay (Uber). As one of the trainers, I got a special $1500 discount code: agentfellow.ai/pawel-kbgef All cohorts, without limits, for the entire year. $22K value. You read this right. And with the link above, you can also attend my Claudathon cohort for free (Sep 19). About: go.productcompass.pm/claudathon - just send me an email pawel@productcompass.pm

  7. #7

    @AIatMeta, I'm considering removing Muse from this benchmark altogether. According to my probes, @OpenRouter doesn't support your effort levels reliably. I see no correlation between tokens and the requested effort level. An alternative is your API. I've been writing to you regarding your payments that do not work in the EU (Visa, MasterCard) but got no response even though ~2 weeks have passed. At the same time, I was able to purchase @Kimi_Moonshot or @Zai_org credits without issues.