The bet

One dense policy, disciplined above its weight class.

v6 is a single program with one thesis: that a dense policy of roughly nine billion parameters, trained under a verifier-first capability foundry and promoted only on verified evidence, can compete with frontier-lab systems on agentic and coding work — while remaining open-weight and sovereign. The values it runs under are the site's contract: sovereignty, EU constitutional alignment, honesty, open-weight integrity, and capability without compromise. The canonical artifact is one shared dense ~9B policy — not an ensemble, and not a quietly swapped 27B substitute.

The goal is concrete and falsifiable: beat Sonnet 5 on four PACE projections plus ten of twelve capability families, and beat comparable open models (Ornith 9B, Qwen 27B) wherever the comparison is valid. Architecture alternatives such as Qwen 3.6 27B/35B are matched bakeoffs, never silent replacements.

The mechanism for competing above weight class is not scale but operating discipline. The supporting evidence is not yet ours — it is the observation, from a multi-agent panel session, that under identical harnesses the strongest orchestrator beat a larger model purely on behavioral deltas: questioning representations earlier, choosing discriminating experiments, converting observations to code sooner. The bet is that such discipline, internalized into weights as reflexes rather than bolted on as prompt scaffolding, is the variable a small sovereign model can move.

Shared dense policy
~9B
One canonical artifact; alternatives are matched bakeoffs.
Stated goal
10/12
Capability families, plus four PACE projections, versus Sonnet 5.
Compute ceiling
$10K
Hard ceiling with tranched authorization.
Approved spend
$0
As of reconciled metadata; training not yet authorized.
Verifier-first, defined

Every capability lane generates, verifies independently, repairs, splits disjointly, and only then measures student lift. Promotion runs from receipts, never from chat state. Sovereignty is operational: the foundry runs on the project's own gateway with no external API dependency, and every source enters weights only through rights and contamination receipts.

Honesty is a value, and it is load-bearing. As of the reconciled metadata, this is a thesis, not a result. Training is not authorized and approved spend is zero, against a $10K hard ceiling with tranched authorization. The dense v5.1 policy at 262K context remains the production baseline, comparison teacher, and rollback. v6 is product-shaped at the core and research-shaped at the edges.
The program

From a preserved parent to a promoted dense release.

The pipeline runs from an exact preserved pre-alignment parent in the Eugeny 5.1 lineage, through a small late-CPT pilot, replay of the SFT and SimPO behavioral program against frozen gates, and then the combined verified agent-RL and instinct stages, before consolidation, vision regraft, and a dense release.

The current promoted Eugeny 5.1 Vision model remains the production baseline, comparison teacher, and rollback until v6 is promoted. Nothing below is a trained result unless marked otherwise.

  1. Phase 1 Proven

    Pre-alignment parent

    An exact preserved checkpoint in the Eugeny 5.1 lineage. The preserved parent is the source of truth for every later stage; nothing is trained against a moving baseline.

  2. Phase 2 Planned

    Late continued pre-training

    A small continued pre-training (CPT) pilot on rights- and contamination-cleared sources. Admission fails closed when provenance is unknown.

  3. Phase 3 Planned

    SFT & SimPO replay against frozen gates

    The behavioral program — supervised fine-tuning and SimPO preference optimization — replayed against frozen qualification gates, so promotion is evaluated against a fixed bar, not a moving one.

  4. Phase 4 Planned

    Verified agent-RL + instinct (Phase 5)

    The combined verified agent-RL and instinct stages. This is where the two Phase-5 tracks — process awareness and mechanism discovery — live. Specialist adapters are never silently merged; verified examples are replayed into one shared dense checkpoint through a single consolidation route after the unified retention gate.

  5. Phase 5 Planned

    Unified consolidation

    A single consolidation route writes retained examples from every adapter lane back into one shared dense checkpoint. There is one canonical dense policy, not a federation of adapters.

  6. Phase 6 Planned

    Vision regraft

    The native visual tower and processor contract are grafted onto the v6 dense policy, non-destructively and reversibly — the same contract that ships on Eugeny 5.1 Vision today. Multi-image vision remains unverified; the qualified modality contract is one image and a narrow two-frame video gate.

  7. Phase 7 Planned

    Dense promotion

    v6 is promoted as the dense release only on completion of the retention gate. Until then the dense v5.1 policy at 262K remains the served baseline, the comparison teacher, and the rollback.

Architecture alternatives (Qwen 3.6 27B/35B) are matched bakeoffs, evaluated on the same gates — never silent replacements for the ~9B canonical policy.

Phase 5 · Two instinct tracks

Discipline as reflex, and the discovery of mechanism.

Two tracks run inside Phase 5. Process awareness internalizes orchestrator discipline as weight-level reflexes that survive prompt ablation. Mechanism discovery asks whether training on self-generated, certified episodes can make a small policy vertically superior at reverse-engineering novel interactive environments into backtestable world models.

These are the goal paths: the route by which a small dense policy competes with the frontier on behavioral discipline rather than on scale. Both are in development; neither is a trained result. The status of each is stated on its column.

Track A Planned

Process-Awareness Instinct

A small model internalizes orchestrator discipline as weight-level reflexes that survive prompt ablation. Eight behavioral families, each with a structural ground-truth signal and a named failure mode it replaces.

  • Seeds from Kimi K3; the bulk of the data from Qwen on the Fornace gateway.
  • Planned as a QLoRA-adapter pilot before any merge.
  • Ground truth is structural — journal state, DAG timing — not LLM-judged.
  • The judge applies hard rules: a trace that guesses process state is an automatic drop.
  • Promotion requires the reflex to survive scaffold ablation on workflows outside the training distribution.

Status: a canonical track plan, designed and ready for a Phase-0 pilot — not a trained result.

Track B Open research

Mechanism Discovery

Self-generated, certified episodes that reverse-engineer novel interactive environments into backtestable world models. The most ambitious track — and the one gated hardest.

  • An episode is certified only if its world program predicts held-out transitions (a train/verify split inside the episode's own timeline), passes an MDL penalty on program length, and its committed plan executes with zero mispredictions.
  • Exact-match replay over the recorded timeline is rejected as lookup-table gameable.
  • ARC-AGI-3 is the permanent out-of-distribution probe: never trained on, never claimed as a result.
  • Any ARC-adjacent public claim goes only through a sanctioned Kaggle verified open-source track, or is not made.
  • No published work has demonstrated this transfer — so this is an experiment, not a program commitment.

Status: open research, gated by the falsification pilot. No commitment precedes its result.

Track A in detail

The eight process-awareness families.

Process awareness decomposes orchestrator discipline into eight behavioral families. The grounding insight is that under an identical harness, behavioral discipline — not raw scale — separated the best orchestrator. Each family is trained as a weight-level reflex against structural ground truth, with a named failure mode it replaces.

The eight families · ground truth is structural, not LLM-judged
Family Ground truth Failure it replaces The reflex
1. Status reporting Journal phase and agent counts Guessing "it's probably still running" Reads state before reporting it
2. Agentic time estimation Workflow DAG plus elapsed time Human-time intuition ("maybe an hour?") Decomposes parallel phases (max) from sequential barriers (sum); returns ranges with the bottleneck named
3. Research → Plan → Delegate Question complexity and live evidence state Answering from memory; panel-everything or answer-everything-solo Calibrates orchestration depth to the actual question
4. Progressive correction Distance between the user's mental model and the truth Dumping the whole answer; an unstructured "actually…" Names what is right, then the specific error, as an investment in the next question
5. Sub-agent design Role coverage, phase architecture, schema quality A missing skeptic, unstructured output, or prescribed conclusions Designs panels with adversarial coverage
6. Epistemic hygiene Source dates, provenance, and verification status Undated claims, single-source assertions, memory treated as evidence Dates and searches when uncertain, with bounded freshness windows
7. Resource reasoning GPU-hours, parallelism, and sequential barriers Dollar estimates without decomposition Decomposes cost and names fallbacks
8. Serendipity flagging A contradiction between a sub-result and an upstream assumption Waiting for the final report and missing lateral signal Flags the implication inline

Promotion for each family requires the reflex to survive scaffold ablation on workflows outside the training distribution. A narrated behavior that was never executed under the judge is rejected at admission.

Track B in detail

Certify the episode, then falsify the track.

A mechanism-discovery episode is admitted only when it is certified. Certification is structural, not narrative — and the track itself is gated by a pre-registered falsification pilot that runs before any Phase-2 compute.

Held-out prediction
The episode's world program must predict held-out transitions — a train/verify split inside the episode's own timeline.
MDL penalty
The program passes a minimum-description-length penalty on program length, so a lookup table does not qualify as a model.
Zero mispredictions
The committed plan must execute against the environment with zero mispredictions. Exact-match replay over the recorded timeline is rejected as gameable.
Permanent OOD probe
ARC-AGI-3 is never trained on and never claimed as a Eugeny result. Any ARC-adjacent public claim goes only through a sanctioned Kaggle verified open-source track, or is not made.
Falsification is first. Kill criteria are pre-registered before any Phase-2 compute: the goal card freezes gate metrics, held-out family lists, and effect sizes. The pilot passes only on a completion-weighted gain of at least +10pp on held-out families, with a paired-bootstrap lower bound above zero, no PACE regression, and gains that survive a bare-Python-REPL arm whose null hypothesis is that a stock model plus a thin scaffold already captures the value. A failed gate closes the track and triggers a pre-committed fallback; the harness and environment factory survive as reusable RL infrastructure. No commitment precedes the result.
Long context

HiLS-Attention — from dense 256K to learned-sparse millions.

HiLS-Attention is end-to-end learned chunk-wise sparse attention. Each query attends to a local window plus the top-k retrieved distant chunks, at constant cost. A landmark token per chunk produces a learnable summary key; chunk mass is estimated by a learnable first-order Taylor surrogate of the true LogSumExp mass — the principled fix for the prior finding that parameter-free and mean-pooled selectors returned 0% needle recall. Attention is factorized into an inter-chunk softmax (allocating mass across chunks) and an intra-chunk softmax; because the surrogate parameterizes the forward pass, the LM loss directly trains chunk selection, with HoPE positional encoding for length extrapolation.

The dense 262,144-token envelope ships today. Planned HiLS is research, not a verified product capability.

HiLS-Attention needle-in-a-haystack accuracy by context length Reported NIAH accuracy stays above 90 percent from the 8K training point out to 4M tokens — a 64x extrapolation. The dense 262K serving envelope is marked as the proven baseline that ships today. trained · 8K dense 262K · proven, ships now 90% NIAH 64× extrapolation 8K 32K 128K 512K 2M 4M context length (tokens, log scale) NIAH accuracy
Needle-in-a-haystack accuracy by context length, as reported by the HiLS authors. The curve stays above 90% from the 8K training point out to 4M tokens — a 64× extrapolation. The dense 262K envelope is the proven baseline that ships today; the extrapolation beyond it is the research candidate, not a Eugeny measurement.
NIAH @ 4M
>90% retrieval, trained on 8K — a 64× extrapolation regime
Prefill speedup
13.5× at 512K tokens
Decode speedup
15.7× at 512K tokens
Quality
LongBench matching or beating full attention, near-zero normal-context loss
Conversion
~50B-token continued pre-training from a full-attention checkpoint (not architecture surgery)
Backbone integration
8 full-attention layers converted; 24 GatedDeltaNet layers left untouched
Reported by the authors, not yet by us

The figures above are from the HiLS paper (arXiv:2607.02980), not from a Eugeny measurement. Our integration is the custom part: apply HiLS to the eight full-attention layers of the hybrid ~9B backbone and leave the twenty-four GatedDeltaNet layers untouched. That modeling file does not exist upstream and is the main engineering task.

Status
Planned. A separate architecture gate, sequenced after the dense policy is frozen — not a v5.1 proof.
Constraints
HiLS is a training-side CUDA method with no Mac or MLX inference path, no released checkpoints, and serving code still pending.
What ships now
The proven dense fallback — the served model at 262,144 tokens, comfortable in memory with zero quality loss. Dense 262K is the release baseline; HiLS is the candidate long-context SKU behind it.
Conciseness as architecture

The shortest correct path, trained at three layers.

Conciseness in v6 is architectural, not stylistic. One principle — the shortest correct path is the goal — is applied at three layers that compound. A model trained on the direct path, decision-critical content kept and noise removed, behaves as if it already knew what to do.

Data layer

Condense the teacher trace

Teacher chain-of-thought is condensed before it reaches the student. The rewriter keeps decision-critical structure — the state-reading action, the DAG decomposition, dated claims, serendipity flags — and removes exploration, backtracking, and theatrical filler. Target: 40–60% shorter traces.

Policy layer

Reinforce the direct path

Short correct paths are reinforced through the Long2Short class of training: preference pairs and outcome RL whose target is the condensed direct path, with rewards for stopping, non-duplication, and a non-empty parseable final answer — a direct response to PACE failures where models exhausted their budget in reasoning and never emitted a final channel.

Code layer

Edit minimally

The simplify reflex governs editing: minimal correct edits, the smallest helper justified by a no-tool baseline, tests preserved, regressions repaired rather than papered over.

Conciseness is judged by shorter trajectories that still pass the verifier — a capability, not a style preference. Planned

Posture

What v6 is not.

v6 states its boundaries plainly. Honesty is a value, and saying what is not yet true is how it is kept. The boundaries below hold while the result is still open.

Research never blocks the product
The dense SKU ships on its own schedule. Long-context and architecture conversion are sequenced behind it, not gating it. Proof search runs in parallel and stays non-blocking.
Proof search stays separate and honest
PACE projections are not direct benchmark scores and are never presented as such. Projection wins and real scores are kept as separate columns and separate receipts.
Specialist adapters are never silently merged
Mechanism discovery, process awareness, and metatooling adapters are never quietly folded into the canonical dense policy. Verified examples are replayed into one shared dense checkpoint through a single consolidation route after the unified retention gate.
No contaminated data in weights
No private, rights-uncleared, or contaminated data enters weights. Every source passes schema, license, dedup, contamination, and tokenizer receipts. Admission fails closed when provenance is unknown.
Multi-image vision remains unverified
The qualified modality contract covers a single image and a narrow two-frame video gate. No broader multimodal claim is made until independently proven.
The evidence boundary. As of the reconciled metadata, v6 training is not authorized and approved spend is zero. The dense v5.1 policy remains the production baseline, the comparison teacher, and the rollback. What is stated on this page is what is proven, what is planned, and what is open research — and the three are kept distinct.