Skip to content
SemiAnalysis@SemiAnalysis_ · Sep 8, 2026

Gemini 3.8 Flash and Muse Spark 1.3 are two of the most clearly benchmaxxed models we've seen yet. Despite being…

5 tweets1 min read♥ 2.3Koriginal

Summary

Gemini 3.8 Flash and Muse Spark 1.3 perform well on public benchmarks like Terminal Bench 2.1 but poorly on newer Terminal Bench 4.0, suggesting these models are being trained on benchmark-mimicking data rather than developing genuine capabilities that generalize. The author argues that public benchmarks inevitably become gamed through indirect training data optimization, making private benchmarks necessary for meaningful evaluation.

Summarized by ThreadOut AI from the full thread. May miss nuance — read the thread below.

  1. #1

    Gemini 3.8 Flash and Muse Spark 1.3 are two of the most clearly benchmaxxed models we've seen yet. Despite being comparable to both GPT-6 and Fable 5.1 on Terminal Bench 2.1, their Terminal Bench 4.0 performance is markedly worse. (1/5)🧵

  2. #2

    How is this possible? All of the tasks in Terminal Bench 2.1 are fully public. Though Meta and Google would never train on the tasks directly, they absolutely will buy data from RL env startups that’s designed to mimic TB 2.1 tasks as closely as possible. The net effect is the same. You’d typically expect improved TB 2.1 performance to generalize to other agentic tasks, but Gemini and Muse don't even generalize to TB 4.0. (2/5)

  3. #3

    Ultimately, this is the fate of all good public benchmarks. TB 4.0 is no exception. It’s only useful signal now because it was released 2 weeks ago. Since all the tasks are similarly public, it won't be long until it's hillclimbed by all the aspiring “frontier” labs. (3/5)

  4. #4

    Gemini 3.8 Flash on DeepSWE is another good example. Clearly, Datacurve made a ton of money selling Google DeepSWE-shaped tasks. (4/5)

  5. #5

    The solution is more high quality private benchmarks. If you have a great private benchmark, reach out to @maxkan. We'd love to chat! (5/5)