Honesty is a value. The goal is real evidence.
These are PACE predictions derived from 385 standardized proxy results, published alongside their diagnostics and their limits. They are not direct executions of the four target benchmarks. They are published because honesty is one of our values, and because the goal is to compete on real evidence — evidence the reader can check.
Complete run · 16 July 2026 · Europe/Rome
The PACE projections.
Eugeny 5.1 Vision, BF16 and unquantized, was run through the PACE selection at temperature 0.0. Each figure below is a projection from the standardized proxy set — a checkpoint-ranking signal and an optimizer target, not a trophy.
A projection above a registered target counts as a PACE-projected win; direct benchmark scores remain a separate column. Neither overwrites the other.
Why is multimodal low?
The model transported images correctly, but broad visual understanding was never established by the serving qualification. The strongest direct evidence is the PACE visual subset: weak academic-image reasoning and web-screenshot grounding, with only middling visual-puzzle performance.
Eugeny 5.1 is primarily a coding and text-reasoning model. Its vision derivative passed a narrow single-image serving gate; that proves the interface works, not that visual reasoning is competitive.
Estimate with caution.
PACE’s multimodal target predictor is the least trustworthy part of this result. On its reference models, its leave-one-model-out diagnostics show almost no useful rank or linear correlation. The 16.93% estimate is directionally consistent with the weak visual proxies, but it should not be treated as a verified SWE-bench Multimodal ceiling.
Honesty is a value, not a posture.
Eugeny keeps two kinds of score, and never lets one overwrite the other. A PACE projection is a checkpoint-ranking signal and an optimizer target; a direct benchmark run is a separate receipt. Projection scores and direct scores live in separate columns, attached to separate evidence, and a projection above a registered target is labelled as a PACE-projected win — never presented as a direct execution score.
The low multimodal figure and the predictor’s poor calibration are published on the same surface as the wins. Honesty is a value, not a posture, and the goal is to compete on real evidence, not on marketing. A score without its complete trajectory and grader evidence is incomplete evidence; a number presented without its limitation is marketing. The uncomfortable numbers stay.
PACE prediction
Derived from 385 standardized proxy results. A ranking signal and an optimizer target. Counts as a win only when labelled “PACE-projected.”
Target benchmark run
A direct, fidelity-preserving execution of GAIA, SWE-bench Verified, SWE-bench Multimodal, or SWT-Bench. Separate column. Has not been run for the multimodal target.
Run facts.
The full selection finished after one transient LeetCode submission failure was recovered; the identical credentials succeeded on the first retry, ruling out persistent token expiry. Final coverage is complete, and every figure on this page is reproducible from the facts below.
- Model
- Eugeny 5.1 Vision · BF16, unquantized
- Sampling
- Temperature 0.0 · deterministic greedy decoding
- Coverage
- 412 / 412 selection rows · 385 unique standardized keys
- Context
- 262,144 tokens served maximum
- Media scope
- At most one image per prompt; multi-image reasoning not qualified
- Runtime
- vLLM 0.25.0 · compiled CUDA graphs
- PACE source
dc2ef80e00addd519e7d8479f875cc3ecb46c6cb- Evaluator
sha256:c94bda538bb5ca1fdd53fecc0e1899e63184f40a978110060d1818c255e91994- Recovery
- One DebugBench row retried and scored 0.0; predictors then rerun on complete coverage
- Published
- 16 July 2026 · Europe/Rome