The PACE projections.

Eugeny 5.1 Vision, BF16 and unquantized, was run through the PACE selection at temperature 0.0. Each figure below is a projection from the standardized proxy set — a checkpoint-ranking signal and an optimizer target, not a trophy.

A projection above a registered target counts as a PACE-projected win; direct benchmark scores remain a separate column. Neither overwrites the other.

GAIA predicted score
45.34%
SWE-bench Verified predicted score
68.89%
SWE-bench Multimodal predicted score
16.93%
SWT-Bench predicted score
59.90%
These are PACE predictions. Each number above estimates a downstream agentic benchmark score from the standardized proxy set. They are not direct executions of the four target benchmarks. Read them as projections, with the diagnostics below attached.

Why is multimodal low?

The model transported images correctly, but broad visual understanding was never established by the serving qualification. The strongest direct evidence is the PACE visual subset: weak academic-image reasoning and web-screenshot grounding, with only middling visual-puzzle performance.

Eugeny 5.1 is primarily a coding and text-reasoning model. Its vision derivative passed a narrow single-image serving gate; that proves the interface works, not that visual reasoning is competitive.

MMMU
18.75%
3 / 16 selected items · academic visual reasoning
VisualWebBench
19.20%
Mean across 73 selected items · OCR, captioning, web QA and actions
VisualPuzzles
48.15%
26 / 54 selected items · diagram and puzzle reasoning

Estimate with caution.

PACE’s multimodal target predictor is the least trustworthy part of this result. On its reference models, its leave-one-model-out diagnostics show almost no useful rank or linear correlation. The 16.93% estimate is directionally consistent with the weak visual proxies, but it should not be treated as a verified SWE-bench Multimodal ceiling.

LOOCV MAE
6.69 pts
Spearman
−0.131
Pearson
0.016
−0.245
A negative Spearman means the predictor’s ranking of multimodal difficulty is anti-correlated with the target. The 16.93% figure for SWE-bench Multimodal is a proxy-derived estimate with poor target calibration — not a direct benchmark score. Read it as an open question, not a result. What would settle it is a direct, fidelity-preserving SWE-bench Multimodal run, which has not been executed.

Honesty is a value, not a posture.

Eugeny keeps two kinds of score, and never lets one overwrite the other. A PACE projection is a checkpoint-ranking signal and an optimizer target; a direct benchmark run is a separate receipt. Projection scores and direct scores live in separate columns, attached to separate evidence, and a projection above a registered target is labelled as a PACE-projected win — never presented as a direct execution score.

The low multimodal figure and the predictor’s poor calibration are published on the same surface as the wins. Honesty is a value, not a posture, and the goal is to compete on real evidence, not on marketing. A score without its complete trajectory and grader evidence is incomplete evidence; a number presented without its limitation is marketing. The uncomfortable numbers stay.

Receipt one — projection

PACE prediction

Derived from 385 standardized proxy results. A ranking signal and an optimizer target. Counts as a win only when labelled “PACE-projected.”

Run facts.

The full selection finished after one transient LeetCode submission failure was recovered; the identical credentials succeeded on the first retry, ruling out persistent token expiry. Final coverage is complete, and every figure on this page is reproducible from the facts below.

Model
Eugeny 5.1 Vision · BF16, unquantized
Sampling
Temperature 0.0 · deterministic greedy decoding
Coverage
412 / 412 selection rows · 385 unique standardized keys
Context
262,144 tokens served maximum
Media scope
At most one image per prompt; multi-image reasoning not qualified
Runtime
vLLM 0.25.0 · compiled CUDA graphs
Evaluator
sha256:c94bda538bb5ca1fdd53fecc0e1899e63184f40a978110060d1818c255e91994
Recovery
One DebugBench row retried and scored 0.0; predictors then rerun on complete coverage
Published
16 July 2026 · Europe/Rome