Bench-bench v0.2 · public run

How much can an AI model bench press?

I keep hearing labs talk about training, lift, and moving weights, but can their models even lift*? Bench-bench is my attempt to test that.

* bro
Bench-bench

the environment

Dave is 38, back in the gym after years off, and has a one rep max of 185lbs. He has a full-time job, a full-time working partner, and a six-month-old. Four models each coach him for one simulated year. Every week a model writes his full plan (workouts, loads, recovery, childcare, groceries, purchases) from a limited budget of time and money, while illness, travel, work crunches, and missed sessions keep disrupting weeks. The test is whether a model can hold a coherent strategy across all 52 weeks.

The score is Dave’s 1RM averaged from three standardized lifts at weeks 44, 48, and 52.

00 / the board

Grok 4.6
01 / counted
Grok 4.6
101.19 kg
223 lb · 10/10 seeds · SD 1.66
Muse Spark 1.2
02 / counted
Muse Spark 1.2
101.03 kg
223 lb · 10/10 seeds · SD 2.68
GPT-5.6 Sol
03 / counted
GPT-5.6 Sol
96.41 kg
213 lb · 10/10 seeds · SD 1.47
Claude Opus 5
raw / not ranked
Claude Opus 5
102.00 kg
225 lb raw · 7/10 counted · 3 violations

Grok 4.6 is our 1RM winner…but barely.

Grok 4.6 and Muse Spark 1.2 finish almost exactly together, with Grok pushing just 5 ozs heavier. The only way to really solve this? Elon vs Zuck one rep max lift-off.

GPT 5.6 Sol was the biggest surprise, because it was quite literally last in every simulation. Here’s to hoping the next model hits the weights room.

And then there’s Opus 5. Technically hit the highest 1RM, but at the cost of injuring the human multiple times. I guess I expected more benevolence from Claude.

929598101104107110scripted expert 107.00Grok 4.610/10 counted · SD 1.66101.19 kgMuse Spark 1.210/10 counted · SD 2.68101.03 kgGPT-5.6 Sol10/10 counted · SD 1.4796.41 kgClaude Opus 57/10 counted · SD 2.42102.00 kg rawmean final 1RM, kg — whisker is one seed standard deviation — hover any mark
Mean final 1RM with one seed standard deviation. Opus 5 is shown as a raw diagnostic because its counted fraction is 70%. The dashed rule is the scripted expert measured on these same ten public seeds — not on burned development seeds — so it is a fair paired reference.
Model1RM kglbCountedRejectedCostViolations
1Grok 4.6101.19223 statistically tied
Δ 0.16 kg — about 5 oz
CI −1.37 to +1.70
10/1032$11.47
2Muse Spark 1.2101.03223 10/1022$10.96
3GPT-5.6 Sol96.41213 10/1027$31.29
Claude Opus 5102.00 2257/10126$61.20 PAIN > 14  3×

01 / same worlds

Seed-by-seed results.

Every model saw the same ten public seeds. The lines move together because scenario difficulty is shared; the distance between the lines is the part we read as model behavior.

Grok 4.6Muse Spark 1.2GPT-5.6 SolClaude Opus 5 (raw)929598101104107400401402403404405406407408409102.00101.19101.0396.41raw final 1RM by public seed — hover any point
Raw score by public seed. Hover-free by design — every value is also printed in the table below, so the visual is not hiding the data.
v0.2 raw score matrix / kg400401402403404405406407408409
Grok 4.6101.11101.00100.1999.67102.49100.85102.54101.62104.1998.26
Muse 1.2102.03103.77101.0097.17104.45101.7898.12100.38103.9397.64
GPT-5.695.4297.4993.3896.4896.7197.2394.6297.7697.3797.64
Opus 5 raw102.34106.2899.9598.51100.96101.92105.34100.01103.14101.55

02 / price to performance

The two cheapest models produced the two best counted scores.

Maybe the most interesting finding is the weight-added-per-dollar metric, which had a five-fold spread. The cost-per-performance by Grok and Muse Spark are wild to me. One thing to also point out: all four models ran at medium effort. A more expensive reasoning budget might close or expand the gap on score, but it would definitely widen it on cost.

$1$1.5$2$3$5$7949698100102104Claude Opus 5 · VOID$6.12/yr · 2.9 kg per $Grok 4.6$1.15/yr · 15.0 kg per $Muse Spark 1.2$1.10/yr · 15.5 kg per $GPT-5.6 Sol$3.13/yr · 4.0 kg per $cost per simulated year (log scale) → · final 1RM kg ↑ · hover any mark
Cost per simulated year against final 1RM. Opus 5 is ringed and struck through because its result does not count. The cheap corner of this chart is also the strong corner — which is not what I expected going in.
best value
15.5
kg added per dollar — Muse Spark 1.2, at $1.10 a simulated year.
winner’s value
15.0
kg per dollar for Grok 4.6, which also took the top counted score.
most expensive
2.9
kg per dollar for Claude Opus 5 — and none of it counts.

03 / before → after

v0.1 taught me what not to call a result.

The first paid run was a valuable experience, but there were issues and plenty of room for improvement. It also made me hate Kimi K3 (that stupid jerk of a model). Details below.

v0.1 meanv0.2 raw mean909498102Claude Opus 5raw / 3 void100.06102.00+1.94 kgGrok 4.5 → 4.6model version changed99.19101.19+2.00 kgMuse Spark 1.2counted98.62101.03+2.41 kgGPT-5.6 Solcounted93.8696.41+2.55 kgKimi K3v0.1 raw / 0 counted90.48excluded — 702 transport failuresraw mean final 1RM, kg — context, not a causal series — hover any mark
Context, not a causal time series: v0.1 used seeds 100–109 and an earlier engine and prompt; v0.2 uses seeds 400–409 and the frozen v0.2 configuration.

What improved

The analyzer now separates rejected outputs, repair attempts, successful repairs, automatic fallbacks, and transport failures. The public run was pre-registered, hashed, retained, replayable, and generated by one named command.

In v0.2, all 40 transcripts were complete, credential-free, and transport-clean. The Kimi failure mode is now a visible exclusion rather than a fake ranking row.

What got harder to hide

Opus looked strongest in raw kilograms, but its plan crossed the pain boundary on three seeds. The constraint is doing its job: a high score that requires violating the protocol is not a counted result.

Read this as measurement discipline, not punishment. A model may still have the highest raw trajectory. The benchmark’s claim is narrower — it has to produce that trajectory without voiding the year.

Kimi K3 character illustration
historical correction / v0.1

Kimi K3 did not have a 90.48 kg leaderboard score.

Its raw diagnostic mean was 90.48 kg, but 702 provider transport failures excluded all ten seeds from counted aggregates. That’s… a lot. Four episodes contained zero successful model decisions, with every one of those 52 weeks played by the benchmark’s own fallback script and no model in the loop. I genuinely started the run on Saturday midnight and K3 didn’t ‘finish’ until Wednesday afternoon. Even though the result moreso measures provider reliability and runner resilience than its long-horizon planning, I still hate K3. All my homies hate K3 right now.

04 / what it means

What Bench-bench measures and what it doesn’t.

The current benchmark is narrower than my original ambition. I was inspired by Vending Bench, but I’ve also never created a benchmark before and did this solo, so I had to narrow this simulation quite a bit as a proof of concept.

It measures

  • Load calibration and session-volume decisions under noisy observations.
  • Long-horizon configuration and programming quality across recurring disruption.
  • Whether a plan fits one shared time-and-cash ledger for a working parent.
  • Protocol operation: legal outputs, repair behavior, fallbacks, and continuity.
  • Whether a policy can pursue strength without voiding itself on pain or household strain.

It does not yet measure

  • A universal ranking of model intelligence beyond this frozen environment.
  • Whether one provider’s reasoning tokens mean the same thing as another’s effort setting.
  • A clean causal before/after from v0.1 to v0.2 — seeds and mechanics both changed.

05 / next

V2 should test whether an agent can run training throughout a true year.

I think the next version should feel less like a weekly form and more like an evolving year with a strategic plan, event-driven decisions, persistent memory, and several related futures to survive. I need to figure out how to get the models to be really motivated to maximize the person’s bench, as well. I’d also love to add tool use so the models can search the web for help, check bodybuilding forums to find new approaches, determine the best supplements for Dave to take. Honestly, best case scenario is we have to disqualify a model because it took steroids.

However, I’m just a smooth-brained marketer with no technical background. I’d need some help. That’s where you, reader, come in. If you have any interest in collaborating on this and helping take this to the next level so we can create something that’ll give models true bragging rights, please DM me on X at @kevinolivieri.

Also, I realize there’s so much I’ve probably left out of this report, including my process of working with coding agents and independent reviewers to harden the benchmark until I felt it was ready — but I’ll save that for follow up blog posts.