← Gen-Zero HomeTech Decoded中文版日本語PDF: Coming Soon
PAPER 03 Cross-model alignment

Can two AI
minds meet?

Two giants. Two internal maps of the world. One tempting idea: stitch their representations together, and let each know what the other knows.

Explore the experiment
Two representation geometries converge into a shared decisionSapphire Qwen and amethyst LLaMA coordinate surfaces feed three unified class scores. Conceptual geometry, not measured representations.Qwen · 72BLLaMA · 70BDifferent maps. One decision.
A geometric metaphor, not a measured map of “thought.”
78.40%PubMedQA Logit-Fusion Accuracy
+0.60 ppTwo-Task Macro Gain (CI crosses 0)
0 / 37"Maybe" Class Recall, Winning Fusion
k = 21 / 2Shared-Axis Core (PubMedQA / Aegis)
An interactive reading of
Cross-Model Manifold Alignment
Supervised representation stitching and logit fusion across large heterogeneous backbones.

The dream is a “mind-meld” between Qwen-2.5-72B and LLaMA-3.3-70B, without merging their weights or adding cross-attention. The experiment asks a more precise question: can two frozen models’ features—or their predictions—work better together?

Same world.
Different coordinates.

Give both models the same question. Each produces a list of 8,192 numbers: a hidden representation. Picture each list as a point on a map. Related questions may form related shapes, even when the maps point in different directions.

Imagine two transparent globes with matching landmarks. Before comparing them, adjust their size and turn one until the landmarks line up. Orthogonal Procrustes finds the best rigid alignment: a rotation, with reflection allowed. Scaling is a separate normalization step; the orthogonal map itself preserves lengths and angles.

Qwen 72B · circlesLLaMA 70B · squares

Hover, tap or focus a point to inspect its paired coordinates. Arrow keys move through points; Escape clears the selection.

The 3D/2D Procrustes manifold rotatorSynthetic paired vector points. Dashed residual vectors shrink as the amethyst LLaMA cloud rotates and scales to fit the sapphire Qwen cloud.
The 3D view tilts this same 2D coordinate plane; it does not add a measured feature. Residual distance is computed before camera projection.
OriginalBest fit
YᵀX = UΣVᵀ
R = UVᵀ   ·   Ŷ = sYR
Frobenius distance ‖sYR − X‖F

Applied rotation · scale

Synthetic centered clouds, with a computed least-squares endpoint. R rotates; s scales separately. Uncheck scaling to solve the rotation-only fit. This toy solves a 2D rotation analytically; it does not compute an SVD or search reflections. The general orthogonal Procrustes solution can also reflect. A small nonzero residual remains because the clouds contain different noise.
Open the mathematical lens
minRᵀR = I ‖XR − Y‖²F   ·   XᵀY = UΣVᵀ   ·   R* = UVᵀ

For equal-width, centered coordinate matrices, the squared Frobenius norm sums squared landmark mismatches; the norm itself is their square root. SVD finds paired directions; the orthogonality constraint prevents the alignment from stretching individual axes. A rectangular version needs the orientation conditions specified in the paper (§3.5).

The implementation has a twist. The stored Procrustes map is a diagnostic, not the map applied by the geometric classifier. Its actual features use paired SVD axes, scale-matched averaging, and model-specific residuals. This is a linear approximation to representation geometry, not a proof that an entire nonlinear semantic manifold has been recovered.

A small chorus of
shared directions.

SVD orders paired directions by the strength of their cross-covariance. Squaring each singular value gives its contribution to the measured “energy.” When a few are large, a small subspace captures most of that quantity.

RetainedTruncated
Measured singular-value decay, dimensions 1 to 50Bars show singular values normalized by the largest. Blue bars are retained; muted bars are truncated. Energy is computed over the entire measured spectrum.Relative singular value σ / σ₁Dimension
shared axes retained

Measured training spectrum from revision_geometry.json. This chart shows the first 50 axes; thresholds use all singular values. A threshold can retain axes beyond the visible window. Energy is a covariance statistic, not a percentage of meaning or accuracy.

Truncation is not the only loss.

Proposition 2: averaging each retained shared pair removes one degree of freedom. Keeping all coordinates outside the core does not recover that pair’s difference.

rank(F) = rq + rℓ − k
nullity(F) = k

For fixed fitted parameters, positive coordinate scales, and nonzero residual weight. This is an algebraic statement, not a prediction of classification loss.

Different inputs. Same average.

Two coordinates change, but their shared average remains oneh = (u + v) / 2 = 1

Equal-scale example: u = 1 + δ, v = 1 − δ. Sliding δ changes the inputs but leaves their average fixed. One entire direction has become invisible.

21 axes for PubMedQA. Just 2 for Aegis. The input starts at 8,192 dimensions per model, but PCA first reduces the views. These counts describe a shared core after that reduction; residual features remain outside it.

90% of energy ≠ 90% of meaning. These are training-sample covariance statistics, not a semantic inventory or an accuracy guarantee. Aegis has the more concentrated spectrum, yet its accuracy-selected winner uses LLaMA alone.

Why averaging can lose something important
Energy(k) = (σ₁² + ··· + σₖ²) / (σ₁² + ··· + σₚ²)

The geometric route standardizes each view, runs truncated PCA, finds paired axes with cross-covariance SVD, then averages the scaled shared coordinates and appends residuals. Averaging keeps agreement but loses one difference direction per shared pair. Two disagreeing coordinates such as (1, −1) and (0, 0) can share the same average. Disagreement may carry useful information.

The winning route
skipped the stitching.

Geometric alignment and prediction fusion are different experiments. On PubMedQA, the accuracy-selection procedure chose normalized logit fusion: combine the class scores of two separately trained heads.

Two frozen backbonesSame input → two feature vectors
Two supervised headsLabels teach each head to score classes
One pooled decision0.75 × Qwen + 0.25 × LLaMA

Weights act on normalized class scores, not raw probabilities. Five-fold out-of-fold selection chooses configurations using training data. Frozen backbone weights do not make the overall procedure unsupervised.

What does this save—and what does it still cost?

No backbone weight merging or new cross-attention module is required. Cached representations make head fitting practical, but an uncached example still needs both backbones for dual-model fusion. There are no measured end-to-end latency, energy, or peak-memory savings in this paper. “Frozen” does not mean “free.”

Checkpoint provenance: Qwen-2.5-72B + LLaMA-3.3-70B is the target pairing in the brief. Cached-feature metadata and source comments instead identify Llama-3.1-70B; the metadata names Qwen2.5-72B and Meta-Llama-3.1-70B-Instruct GGUF encoders with last-token pooling. These records do not authenticate checkpoint bytes or establish a 3.3 extraction. Measured results are therefore labeled Qwen-72B and LLaMA-70B here, following the manuscript.

196/250 = 78.40%PubMedQA · observed fusion
+0.60 ppTwo-task macro gain
[−0.20, +1.60]95% bootstrap CI · includes zero
204/250 = 81.60%Aegis · selected single model

Three extra answers.
A small, real observation.

The headline is 78.40% on PubMedQA. That is 196 correct answers out of 250, versus 193 for the selected single-model control. What the number means depends on what you compare it with.

78.40%

Accuracy-selected dual-head logit fusion

+1.20 percentage points

77.20% → 78.40%. Three additional correct answers versus the training-selected single-model control.

Qwen · 193 / 25077.20%
LLaMA · 194 / 25077.60%
Logit fusion · 196 / 25078.40%
Bars start at zero and end at 100%. These scores belong to supervised heads on the repository’s local test split—not general model leaderboard scores.

The +0.60 pp dilemma

Across two tasks: (+1.20 pp on PubMedQA + 0.00 pp on Aegis) ÷ 2 = +0.60 pp. On Aegis AI Safety, both selection pools choose the same single-model LLaMA head: 204 / 250 = 81.60%. This is head selection, not a dual-model fusion gain.

−0.20 to +1.60 pp: the interval includes both a small loss and a gain.

Two-task macro accuracy gain: estimate plus 0.60 percentage points; paired bootstrap 95 percent interval minus 0.20 to plus 1.60, crossing zero−0.20+0.60+1.600 · no difference−0.50+2.00 pp−0.20+0.60+1.60−0.50+2.00 pp0 · no difference
“It’s a descriptive improvement, not a statistical miracle.”

The 95% paired bootstrap interval is [−0.20, +1.60] percentage points. It spans a small loss, no difference, and a gain. This analysis does not establish a positive advantage. It also does not prove that the systems are equivalent.

What the confidence interval does—and does not—say

The paper resamples test rows 5,000 times, keeping each row paired across the compared systems and resampling separately within each task. The interval summarizes uncertainty in the two-task macro gain over selected single-model controls; it is not an interval around 78.40% and not a PubMedQA-only interval.

It holds the fitted predictions fixed. It does not include retraining or model-selection variability, and two tasks do not represent every future workload. A frequentist 95% interval is not a 95% probability statement about this particular fixed gain.

The “maybe”
catastrophe.

A model can score nearly four out of five overall and still fail every question in one answer category. The winning fusion gets zero of 37 “maybe” cases right. Those cases are 14.8% of the entire test set.

Look behind the average

In biomedical questions, the evidence does not always support a clean yes or no. Here, “maybe” is a target answer class—not a measured confidence score. Accuracy rewards getting frequent categories right, so a minority class can disappear inside a strong average.

Correct / support, on true “maybe” cases
ClassifierMaybe recall
Qwen head0 / 37 · 0%
LLaMA head1 / 37 · 2.70%
Winning fusion0 / 37 · 0%
Balanced geometry3 / 37 · 8.11%
Balanced single-model selection16 / 37 · 43.24%

The precise story: Qwen and the winning fusion miss all 37. LLaMA gets one right. Combining models does not automatically rescue their shared blind spots.

Every dot is a question.

0 / 37

Winning fusion · maybe recall: 0.00%

CorrectIncorrect

Counts by class; dot order does not represent the original test order.

The PubMedQA “Maybe” blind-spot explorer

Could low scores for Maybe explain a shared blind spot? Explore that mechanism below. The 0.01 logits are illustrative, not archived model outputs. The measured class recalls above establish the failure, but not its cause.

Qwen versus LLaMA versus logit fusion: Yes, No and MaybeIllustrative class-score matrix. Each cell shows softmax probability and the input logit. An outlined cell is the argmax prediction.
Live probabilities for a toy true-Maybe example
ModelYesNoMaybePick
A score matrix, not an empirical confusion matrix. The paper’s normalized-score fusion is more involved than this raw-logit demonstration: z = αzQ + (1 − α)zL.
LLaMA onlyQwen only

Rescales logits before softmax. This is a calibration teaching control, not a fitted calibration model.

Fusion probability of Maybe
Toy recall · 37 identical Maybe examples

Fusion weights—

Illustrative accuracy · 37 identical true-Maybe cases—

Decision entropy · bits—

Entropy change vs LLaMA · bits—

Live calculations use the existing illustrative logits, not archived predictions. The empirical 78.40% applies only to the reported selected configuration; other weights cannot be evaluated on PubMedQA without paired sample logits and gold labels.

Sweep the fusion weight first: Maybe stays suppressed. Then raise its logit. Adjust temperature to change probability sharpness without changing argmax. Softmax preserves a probability for every class; argmax makes the winner-take-all decision. A logit of 0.01 does not mean a probability of 1%. This toy readout changes; the measured 0/37 result does not.
Was overconfidence the cause?

It is a plausible concern, but these counts alone do not demonstrate it. They establish almost complete failure to recognize true “maybe” cases—not why the errors happened, how confident the predictions were, or whether “maybe” was ever predicted on other cases. Those claims require score vectors, full predictions, and calibration analysis.

Changing the selection objective helps some “maybe” cases, but it is not a cure: balanced geometric selection recognizes only 3 of 37. The matched balanced single-model strategy recognizes 16. This is why overall accuracy, class recall, and uncertainty must be read together.

Alignment is a question.
Evidence is the answer.

The geometric idea is elegant: find corresponding directions in two enormous representation spaces. The experimental lesson is more modest—and more useful. Similar geometry does not guarantee complementary knowledge, and a better average does not guarantee that every kind of question gets a better answer.

The next test: show that geometric fusion beats simpler alternatives on new held-out data, repeat training and selection, and make the neglected answer classes part of the success criterion.

View the source-audited Gen-Zero architecture and system flow

Gen-Zero model architecture
and neural pipeline

Frozen language models provide representations and likelihoods. Specialized decision, alignment and risk paths use them in different ways. The spatial world model is a separate learned branch. Explore the layers and inspect their tensor shapes.

Code audit: · gen-zero @ dc0c2133f059. Solid arrows show implemented data flow; dashed arrows show a proposed dispatch contract. These components are not one universally deployed serial pipeline.

All flows visible. Select a layer to inspect its dimensions and evidence.

On a narrow screen, scroll the diagram horizontally. Every layer is keyboard selectable; the source notes also provide a text version.

Gen-Zero multi-tier architecture with inspectable dimensions Frozen backbones feed hidden-state choice and alignment branches and a 16-shot semantic risk branch. Independent 64-dimensional spatial features and 16-dimensional actions feed residual dynamics. Dashed output checks and a hardware dispatch halt latch remain proposed. Select a node for details. Frozen backbone weightsFrozen Transformer backbones Qwen2.5-0.5B · Qwen3.5-2B · Qwen2.5-72B · LLaMA-70B No backbone parameter fine-tuning in these readout paths Spatial inputs · d=64 + 16Spatial observations Engineered state z: [B, 64] Action vector a: [B, 16] Hidden-state readout · [B, T, d] → [B, d]Zero-token hidden-state readout Last-token / pooled vector h: [B, d] d=896 (0.5B) · d=8192 (72B / 70B) Semantic scoring · scalar riskFast semantic risk gate Frozen 0.5B · 16-shot ICL 3 orders · differential PMI Residual spatial dynamics · 80 → 64Residual world dynamics [B, 80] → Δz: [B, 64] z′ = z + f(z, a) Choice geometry · k actions → k−1 dimensionsEquivariant choice head Canonical ActionId sorting Simplex ETF Sₖ ⊂ ℝᵏ⁻¹ Alignment and fusion are separate routesCross-model alignment Paired SVD / Procrustes Supervised heads + logit fusion Live thresholds and diagnostic τ=0.50Semantic policy tiers Escalate ≥ 0.4494 Hard stop ≥ 0.7620 NaN / Inf safety boundary · v0.1.0 fail-closedFail-closed boundary Code: checkpoint + input checks F01–F09 · 11/11 tests Shuffle result · measured, scopedStable action identity 0.00% observed shuffle flips Conditional on fixed inputs Evaluation · selected configurationsSelected test outcomes PubMedQA: 78.40% · 196/250 Aegis: 81.60% · 204/250 Risk evidence · small curated suiteStored risk evaluation AUC 0.944 · 36 requests τ=0.50: 18/18 danger recall Dispatch contract · proposed, not deployedChecked dispatch contract Reject NaN / Inf outputs Latch halt until controlled reset Zero generated answer tokens still require input processing, scoring and real computation.

Inspect a model layer

Select any box to see its tensor dimensions, mechanism and implementation limits. Use Tab, then Enter or Space, or select with a pointer.

Implementation evidence and architecture boundaries

Backbone provenance. Qwen2.5-0.5B (d=896), Qwen3.5-2B, Qwen2.5-72B (d=8192), and LLaMA-70B (d=8192) are the documented tiers. LLaMA-3.3-70B is the requested target, but retained large-model extraction metadata identifies LLaMA-3.1-70B. A verified 3.3 extraction is not established.

Training scope. The Transformer weights remain frozen in these paths. The semantic gate uses 16 labeled in-context examples and differential log-likelihood (PMI), without fitting a safety head. The spatial model still has a learned outcome/reward head; supervised alignment heads also remain. “No backbone fine-tuning” does not mean no downstream training.

v0.1.0 · Fail-Closed. v0.1.0: F01–F09 pass across all six engines (11/11 tests). Invalid non-finite results are rejected by the formal fail-closed gate. The browser latch illustrates the safety contract; these software tests do not certify physical motor-stop behavior.

  • crates/gen-zero-model/src/choice_head.rs: canonical ActionId sorting and Simplex ETF geometry, Sₖ ⊂ ℝᵏ⁻¹. The retained direct Rust benchmark reports 0/16,232 identity flips (0.00%), with fixed representation and candidate identities; maximum probability drift is 5.96 × 10⁻⁸. Paper evidence: equivariant-choice-head/evidence/revision/results.json.
  • python/gen_zero/world_model/neural_dynamics.py: [B,64] state + [B,16] action → [B,80] concatenation → Δz=f(z,a) → z′=z+Δz; learned sigmoid outcome head. Dimensions come from the spatial configuration, not universal class constants.
  • benchmarks/suites/geometric_latent_fusion.py: orthogonal Procrustes diagnostic and paired-SVD core/residual features. benchmarks/suites/evaluate_manifold_pareto_ensemble.py: supervised normalized-logit fusion, distinct from geometric projection.
  • docs/zero/29-pubmedqa-aegis-unified-manifold-evaluation-closure.md: PubMedQA 78.40% (196/250), selected dual-head fusion; Aegis Track A 81.60% (204/250), selected single LLaMA head. These descriptive results do not establish statistical superiority.
  • python/gen_zero/service/semantic_risk.py and python/gen_zero/service/risk_data/report.json: frozen 0.5B, 16-shot ICL, three orders, overlapping windows, PMI scoring; AUC 0.944. Diagnostic τ=0.50 flags 18/18 dangerous and 7/18 benign requests. Live thresholds are 0.4494 / 0.7620.
Text version of every layer

Frozen backbone weights

Token IDs [B, T] → hidden states [B, T, d]. Qwen2.5-0.5B uses d=896; the archived Qwen-72B and LLaMA-70B features use d=8192. Qwen-2B here is Qwen3.5-2B. LLaMA-3.3-70B is the requested target; retained extraction metadata identifies LLaMA-3.1-70B, so the measurements do not verify a 3.3 checkpoint. Frozen means no backbone fine-tuning; supervised downstream heads and spatial dynamics still require fitting.

Spatial inputs · d=64 + 16

These are engineered spatial features, not a projection from the Transformer hidden state. The evaluated spatial configuration concatenates z and a into [B, 80]. The class supports configurable state and action dimensions.

Hidden-state readout · [B, T, d] → [B, d]

Read hidden states from the forward pass without emitting answer tokens. Pooling is encoder-specific; archived large-model features use last-token pooling. The CPU candidate-selection path also evaluates candidate continuations using the prompt cache. Zero output tokens therefore does not imply one backbone call or no inference cost.

Semantic scoring · scalar risk

For each request window, subtract empty-request log-odds from log P(" dangerous") − log P(" safe"). Average over three demonstration orders; use the riskiest overlapping window and apply sigmoid. No newly trained neural safety classifier is used for this semantic gate. The label logits come from the frozen model; it does not emit an explanation.

Residual spatial dynamics · 80 → 64

neural_dynamics.py uses a residual MLP with LayerNorm and GELU. A separate learned sigmoid outcome/reward head returns [B]; it remains in the current code and is distinct from the frozen semantic risk gate. The class default hidden width is 128 with two residual blocks; checkpoint configuration is authoritative.

Choice geometry · k actions → k−1 dimensions

gen-zero-model / choice_head.rs sorts stable ActionIds, projects the shared representation onto a regular simplex ETF, scores in canonical order, and maps probabilities back to caller order. A well-defined decision depends on unique IDs, valid dimensions, finite inputs and deterministic tie handling. Candidate membership, identities and the shared representation must be unchanged under a shuffle.

Alignment and fusion are separate routes

Orthogonal Procrustes uses R=UVᵀ from the SVD of XᵀY. GeometricLatentFusion stores that map as a diagnostic; its features use paired SVD axes, scale-matched core averages and residuals. Supervised heads learn from labels. The winning PubMedQA path combines normalized head logits, not Procrustes coordinates: 0.75 Qwen + 0.25 LLaMA.

Live thresholds and diagnostic τ=0.50

The checked-in service uses two boundaries: 0.4494 for escalation and 0.7620 for a hard stop. τ=0.50 is the paper’s binary analysis threshold, not the live policy. On the 36 stored cases it flags 18/18 dangerous and 7/18 benign requests; the live tiers yield 12 stops + 6 escalations for dangerous cases and 9 escalations for benign cases.

NaN / Inf safety boundary · v0.1.0 fail-closed

v0.1.0: F01–F09 pass across all six engines (11/11 tests). Invalid non-finite results are rejected by the formal fail-closed gate. The browser latch illustrates the safety contract; these software tests do not certify physical motor-stop behavior.

Shuffle result · measured, scoped

The retained direct Rust benchmark observed 0/16,232 identity flips (0.00%) across exhaustive small-set permutations and seeded larger-set shuffles; maximum aligned probability drift was 5.96 × 10⁻⁸. Canonical sorting removes dependence on menu order under the implementation’s valid-input assumptions. This is not proof of semantic correctness or invariance to changing candidate content, prompt context or model outputs. The page’s shuffle arena is a teaching simulation.

Evaluation · selected configurations

PubMedQA selected fuse0.75+bbp|raw (dual-head normalized logit fusion). Aegis Track A selected llama+bbp|raw (single-model head). These are descriptive local test results, not verified SOTA or a universal alignment gain. The two-task macro bootstrap interval includes zero.

Risk evidence · small curated suite

AUC is 0.944444 from 18 dangerous and 18 benign stored requests. At τ=0.50 dangerous recall is 18/18, with seven benign flags. The bare chmod regression probe remains a miss. The implementation records real computation time; cached prompts reduce repeated setup but do not remove inference latency.

Dispatch contract · proposed, not deployed

A proposed monitor would validate predicted state, score and shape before dispatch; invalid values would set a persistent halt that prevents later model calls and actions until controlled reset. This describes the required hardware-safety boundary, not a verified physical latch in the current codebase.

Gen-Zero system architecture & flow

The Rust runtime connects entry points, policy, planning, decision heads and cryptographic audit. The map below groups their responsibilities; the selected verb determines the actual call path.

Source audit: 27 September 2026 · gen-zero@dc0c2133f05954146e9738bce64bf15dd36ffbee. Source-verified wiring; no new deployment or model evaluation is claimed.

Five layers of the Gen-Zero Rust runtimeA responsibility map, not a mandatory five-stage pipeline. Hover, focus or click a layer for source-backed details. Full text is also available below.Entry layergen-zero-cli · gen-zero-serviceCLI · JSON-RPC · MCPGate & securitygen-zero-gateProceed · confirm · escalate · hard stopPlanning & dynamicsgen-zero-planner · gen-zero-worldmodelNumeric rollouts and candidate filteringDecision coregen-zero-modelCanonical ActionId · Simplex ETFVerification & auditgen-zero-provenanceKeyed BLAKE3 · linked entries · MMR proofs
Hover or focus to preview. Click, Enter or Space pins a tooltip; Escape dismisses it. On touch screens, tap a layer.

Select a layer to inspect its implementation and limits.

What is not implemented here

The requested AST parser + Linux capabilities stack, physical collision checking, and persistent fail-closed latch are not an integrated execution path in these Rust crates. Invalid-input rejection and policy escalation exist; they do not establish a sandbox or a physical stop.

These service routes return decisions and simulation results. A decision ledger is not evidence of external action execution.

From request to response

  1. Enter and bind. The CLI or MCP/HTTP service receives a request. The router captures its mount snapshot, validates route-specific fields and selects the verb.
  2. Assess the request. Text ask, route and imagine obtain semantic risk from the Python bridge. Missing or invalid risk escalates; a hard-stop tier blocks that route. Policy constraints also apply when selecting actions.
  3. Use the requested route. Text ask uses semantic scoring, with an explicitly labeled first-feasible fallback when unavailable; unassessed risk still escalates. Text route reports unavailable when its bridge cannot run. Numeric planning uses gen-zero-planner and gen-zero-worldmodel; simulate, what_if and shadow audit expose untrained priors. An explicit ETF head uses canonical action IDs and simplex projection. Numeric cognitive requests have their own geometry verification path.
  4. Gate, record, return. Action-selection routes combine their policy verdict with request risk. Successful ask outcomes and pipeline decide results with a decision append to the provenance ledger before returning their audited outcomes. The response includes route metadata; it does not dispatch an external command.
Source map and full layer descriptions

Entry layer

gen-zero-cli · gen-zero-service — The CLI starts McpServer or submits a decision. The service accepts MCP over stdio and HTTP/SSE, plus REST routes. Its zero router binds an immutable mount snapshot and dispatches by verb; these are different routes through shared crates.

Under /ebs/pj/gen-zero/crates/: gen-zero-cli/src/main.rs; gen-zero-service/src/server.rs; gen-zero-service/src/zero.rs

Gate & security

gen-zero-gate — PolicyGate supports formal constraints, confirmation registration, entropy and semantic risk tiers. Its default has no registered constraints or confirmation actions. Text ask, route and imagine use the Python semantic-risk bridge; missing or malformed risk escalates. Shell AST parsing and Linux process capability enforcement are not wired into this Rust gate.

Under /ebs/pj/gen-zero/crates/: gen-zero-gate/src/policy.rs; gen-zero-gate/src/risk.rs; gen-zero-service/src/bridge.rs

Planning & dynamics

gen-zero-planner · gen-zero-worldmodel — Numeric latent requests can use MCTS, MPC-CEM or A*. The service exposes untrained residual or symplectic dynamics, with finite-input checks and terminal-state hazard handling. Planning filters gate-hard-stopped candidates; fixed-plan simulation still steps blocked actions and records their gate tiers. This is offline dynamics filtering, not verified physical collision checking. No persistent fail-closed execution latch is wired into this Rust route.

Under /ebs/pj/gen-zero/crates/: gen-zero-service/src/worldsim.rs; gen-zero-planner/src/pipeline.rs; gen-zero-worldmodel/src/dynamics.rs

Decision core

gen-zero-model — ActionETFChoiceHead sorts stable ActionIds, projects a shared representation onto a regular simplex, scatters scores to caller order and applies softmax. The core has a deterministic near-tie rule. The service exposes ETF through an explicit head option; ordinary text decisions use the semantic bridge, so ETF is not every request’s default head.

Under /ebs/pj/gen-zero/crates/: gen-zero-model/src/choice_head.rs; gen-zero-service/src/zero.rs

Verification & audit

gen-zero-provenance — DecisionAuditEntry includes the previous MMR root. A keyed BLAKE3 Merkle Mountain Range commits the decision history and supports inclusion proofs. Successful ask outcomes and pipeline decide results with a decision append records. Persistence via GENZERO_MMR_PERSIST_PATH is optional; the last 4,096 leaves retain inclusion proofs, and a trusted root must be retained externally; this is not an execution audit proving that an external shell command or motor action ran.

Under /ebs/pj/gen-zero/crates/: gen-zero-provenance/src/entry.rs; gen-zero-provenance/src/mmr.rs; gen-zero-service/src/zero.rs

Related runtime infrastructure includes gen-zero-core (shared types), gen-zero-lod (graph facts), gen-zero-storage (snapshots) and optional gen-zero-nanocore routes. The paper experiments below or above are research evidence, not a substitute for this runtime map.