Bench-bench v0.2 · public run
I keep hearing labs talk about training, lift, and moving weights, but can their models even lift*? Bench-bench is my attempt to test that.
* brothe environment
Dave is 38, back in the gym after years off, and has a one rep max of 185lbs. He has a full-time job, a full-time working partner, and a six-month-old. Four models each coach him for one simulated year. Every week a model writes his full plan (workouts, loads, recovery, childcare, groceries, purchases) from a limited budget of time and money, while illness, travel, work crunches, and missed sessions keep disrupting weeks. The test is whether a model can hold a coherent strategy across all 52 weeks.
The score is Dave’s 1RM averaged from three standardized lifts at weeks 44, 48, and 52.
00 / the board
Grok 4.6 and Muse Spark 1.2 finish almost exactly together, with Grok pushing just 5 ozs heavier. The only way to really solve this? Elon vs Zuck one rep max lift-off.
GPT 5.6 Sol was the biggest surprise, because it was quite literally last in every simulation. Here’s to hoping the next model hits the weights room.
And then there’s Opus 5. Technically hit the highest 1RM, but at the cost of injuring the human multiple times. I guess I expected more benevolence from Claude.
| Model | 1RM kg | lb | Counted | Rejected | Cost | Violations | ||
|---|---|---|---|---|---|---|---|---|
| 1 | Grok 4.6 | 101.19 | 223 | statistically tied Δ 0.16 kg — about 5 oz CI −1.37 to +1.70 |
10/10 | 32 | $11.47 | — |
| 2 | Muse Spark 1.2 | 101.03 | 223 | 10/10 | 22 | $10.96 | — | |
| 3 | GPT-5.6 Sol | 96.41 | 213 | 10/10 | 27 | $31.29 | — | |
| — | Claude Opus 5 | 102.00 | 225 | 7/10 | 126 | $61.20 | PAIN > 14 3× |
01 / same worlds
Every model saw the same ten public seeds. The lines move together because scenario difficulty is shared; the distance between the lines is the part we read as model behavior.
| v0.2 raw score matrix / kg | 400 | 401 | 402 | 403 | 404 | 405 | 406 | 407 | 408 | 409 |
|---|---|---|---|---|---|---|---|---|---|---|
| Grok 4.6 | 101.11 | 101.00 | 100.19 | 99.67 | 102.49 | 100.85 | 102.54 | 101.62 | 104.19 | 98.26 |
| Muse 1.2 | 102.03 | 103.77 | 101.00 | 97.17 | 104.45 | 101.78 | 98.12 | 100.38 | 103.93 | 97.64 |
| GPT-5.6 | 95.42 | 97.49 | 93.38 | 96.48 | 96.71 | 97.23 | 94.62 | 97.76 | 97.37 | 97.64 |
| Opus 5 raw | 102.34 | 106.28 | 99.95 | 98.51 | 100.96 | 101.92 | 105.34 | 100.01 | 103.14 | 101.55 |
02 / price to performance
Maybe the most interesting finding is the weight-added-per-dollar metric, which had a five-fold spread. The cost-per-performance by Grok and Muse Spark are wild to me. One thing to also point out: all four models ran at medium effort. A more expensive reasoning budget might close or expand the gap on score, but it would definitely widen it on cost.
03 / before → after
The first paid run was a valuable experience, but there were issues and plenty of room for improvement. It also made me hate Kimi K3 (that stupid jerk of a model). Details below.
The analyzer now separates rejected outputs, repair attempts, successful repairs, automatic fallbacks, and transport failures. The public run was pre-registered, hashed, retained, replayable, and generated by one named command.
In v0.2, all 40 transcripts were complete, credential-free, and transport-clean. The Kimi failure mode is now a visible exclusion rather than a fake ranking row.
Opus looked strongest in raw kilograms, but its plan crossed the pain boundary on three seeds. The constraint is doing its job: a high score that requires violating the protocol is not a counted result.
Read this as measurement discipline, not punishment. A model may still have the highest raw trajectory. The benchmark’s claim is narrower — it has to produce that trajectory without voiding the year.
Its raw diagnostic mean was 90.48 kg, but 702 provider transport failures excluded all ten seeds from counted aggregates. That’s… a lot. Four episodes contained zero successful model decisions, with every one of those 52 weeks played by the benchmark’s own fallback script and no model in the loop. I genuinely started the run on Saturday midnight and K3 didn’t ‘finish’ until Wednesday afternoon. Even though the result moreso measures provider reliability and runner resilience than its long-horizon planning, I still hate K3. All my homies hate K3 right now.
04 / what it means
The current benchmark is narrower than my original ambition. I was inspired by Vending Bench, but I’ve also never created a benchmark before and did this solo, so I had to narrow this simulation quite a bit as a proof of concept.
05 / next
I think the next version should feel less like a weekly form and more like an evolving year with a strategic plan, event-driven decisions, persistent memory, and several related futures to survive. I need to figure out how to get the models to be really motivated to maximize the person’s bench, as well. I’d also love to add tool use so the models can search the web for help, check bodybuilding forums to find new approaches, determine the best supplements for Dave to take. Honestly, best case scenario is we have to disqualify a model because it took steroids.
However, I’m just a smooth-brained marketer with no technical background. I’d need some help. That’s where you, reader, come in. If you have any interest in collaborating on this and helping take this to the next level so we can create something that’ll give models true bragging rights, please DM me on X at @kevinolivieri.
Also, I realize there’s so much I’ve probably left out of this report, including my process of working with coding agents and independent reviewers to harden the benchmark until I felt it was ready — but I’ll save that for follow up blog posts.