Gemini 3.8 Flash and Muse Spark 1.3 are two of the most clearly benchmaxxed models we've seen yet. Despite being…
Summary
Gemini 3.8 Flash and Muse Spark 1.3 perform well on public benchmarks like Terminal Bench 2.1 but poorly on newer Terminal Bench 4.0, suggesting these models are being trained on benchmark-mimicking data rather than developing genuine capabilities that generalize. The author argues that public benchmarks inevitably become gamed through indirect training data optimization, making private benchmarks necessary for meaningful evaluation.
Summarized by ThreadOut AI from the full thread. May miss nuance — read the thread below.
- #1
- #2
- #3
- #4
- #5