# Gen-Zero Model Paper Series
## Conference Proceedings Meta-Review & Program Committee Portfolio

**Track:** Autonomous Language Agents and Zero-Token Decision Architectures  
**Assessment date:** 26 September 2026  
**Basis:** the five revised local manuscripts, rebuttal notes, available independent review, and accompanying evidence artifacts.

**Decision provenance.** This is a commissioned PC-style assessment, not a decision issued by NeurIPS or ICLR. No authenticated conference Area Chair acceptance letters are present in the reviewed record. Paper 2 has a local independent review recording **Reject, 3/10, high confidence**; its venue authority is not established. Papers 1, 3, 4, and 5 explicitly report missing review files. The recommendations and composite scores below are newly assigned editorial judgments, not recovered reviewer votes or official acceptance decisions. Neither conference acceptance nor proceedings publication is claimed.

## Executive assessment

The series offers a coherent research agenda for **decision making with explicit computational and enforcement boundaries**. Its most valuable unifying result is that removing generated tokens does not remove the need for computation, supervision, uncertainty assessment, or execution control. Each paper addresses a different failure surface: insufficient computation, unreliable transition prediction, misattributed representation gains, candidate-order dependence, or unsafe semantic authorization.

The revised manuscripts are substantially more credible than the original planning documents. They correct a false parity obstruction, expose duplicate validation content and a non-finite-state safety bypass, attribute fusion gains to their actual predictors, distinguish stable choices from correct choices, and publish semantic-gate misses alongside false positives. This is meaningful scientific progress. It does **not** establish that every claim has authenticated execution provenance or that the five components form a tested, certified agent.

My recommendation for this special track is **Weak Accept for Paper 4 as a narrowly scoped methods and implementation-contract paper; Borderline Reject for Papers 1, 3, and 5; Reject for Paper 2 in its current form**. Paper 4's recommendation is an editorial judgment under the track's contract-verification emphasis, not a claim of broad empirical superiority. No paper warrants an unconditional Strong Accept on the available evidence.

## 1. Cross-paper cohesion and distinct contributions

| Paper | Scientific question and owned contribution | Boundary that prevents overlap | Current evidence ceiling |
|---|---|---|---|
| **1 — Zero-Token Decisions from Frozen Language Models** | What can a fixed-depth, threshold-simulable decision pipeline express, and what do archived zero-generation scores actually establish? | Owns computational expressivity and benchmark-accounting definitions; does not certify dynamics, alignment, or action safety. | Conditional theory, constructive algorithms, and an empirical audit; incomplete historical inference provenance. |
| **2 — Latent world model** | When can learned spatial transitions support lookahead, and what must a fail-closed consumer check? | Owns state evolution and rollout validity; does not solve language-risk classification or candidate permutation stability. | Offline transition diagnostics and a conditional safety contract; no validated neural closed-loop planning or implemented complete contract. |
| **3 — Supervised Representation Stitching and Logit Fusion** | Can paired frozen 70B+ representations yield useful complementary predictions, and at what cost? | Owns multiview geometry, supervised selection, and fusion attribution; does not establish additional computational depth or safety. | Two-task archived comparisons and training-feature geometry diagnostics; no demonstrated general geometric advantage. |
| **4 — Permutation-Equivariant Canonical Choice Heads** | Can candidate presentation order be prevented from changing the selected action identity? | Owns identity/slot semantics and conditional head equivariance; does not establish encoder invariance or semantic correctness. | Complete conditional proofs and direct Rust stress evidence; service observations are refused calls. |
| **5 — Frozen Few-Shot Semantic Risk Gating** | What risk discrimination can a frozen 0.5B model provide using 16-shot differential log-likelihood scoring? | Owns semantic risk and abstention policy; does not predict physical transitions or enforce kernel permissions. | A transparent, small development-reused evaluation and a proposed defense-in-depth contract. |

The conceptual data flow is: frozen representations support readout or supervised multiview combination; candidate scoring can be made identity-stable; optional learned lookahead supplies transition diagnostics; semantic review and independent policy enforcement govern whether an action may execute. Paper 1 constrains the computational claims made about all these stages.

This is a **proposed composition**, not a recovered executable pipeline. In particular, Paper 3's task classifiers are not demonstrated inputs to Paper 4's ETF head, and Paper 2's planner has not been evaluated with the full Paper 5 enforcement stack. Compatibility requires explicit representation dimensions, stable action identities, model/encoder hashes, shared state and action semantics, threshold versions, and a trace binding the reviewed action to the dispatched action.

A prospective acceptance predicate is:

`execute = valid_artifacts AND valid_numerics AND authorized_action AND all_required_reviews_pass AND kernel_permission`

If lookahead is required, its validity checks also belong in this predicate. Missing diagnostics, non-finite outputs, or unresolved escalation must prevent automatic dispatch. This is an engineering proof obligation under complete mediation, not a theorem that learned scores identify every harmful action. Escalation also needs an authorized resolution path; merely returning an escalation label does not enforce anything.

Shared terminology does not make the contributions duplicate. Papers 2 and 5 both discuss fail-closed behavior, but one concerns invalid or unsafe predicted transitions and the other concerns semantic review and execution authorization. Papers 1 and 4 both discuss readouts, but only Paper 4 owns the detailed canonicalization and numerical stress study. Papers 1 and 3 share some archived large-model evidence; those observations must be cross-referenced rather than counted as independent replications. This establishes a defensible editorial division, not an exhaustive finding of novelty relative to all literature.

## 2. Individual meta-reviews

### Paper 1 — Computational boundaries and empirical accountability

**Contribution.** The strongest theoretical result is carefully conditional. Assuming nonuniform `TC⁰ ≠ NC¹`, a polynomial-size constant-depth threshold-simulable pipeline cannot decide the specified `S₅` permutation-product language at every length. The assumptions cover the entire computation, including encoding, preprocessing, backbone operations, candidate construction, and readout. The manuscript also gives a seven-bit group-state recurrent construction with an external end-of-stream signal, and a logarithmic-depth parallel alternative. It therefore does not establish that verbal chain of thought is uniquely necessary. Its representation-collision bound is unconditional but does not establish collisions in the evaluated hidden states. [P1, §3.5]

**Correction of hype.** Parity belongs to uniform `TC⁰`; arithmetic benchmark errors do not prove a general fixed-depth impossibility. Zero generated tokens, zero recurrent steps, and one backbone invocation are separate properties. Candidate-conditioned inference may involve multiple backbone calls. No measured end-to-end speedup follows merely from a cheap head. [P1, §§3.2–3.5]

**Empirical record.** The GPU archive reconciles to **254/390 = 65.13%**, but its producer and checkpoint-linked execution remain unauthenticated. The separate CPU archive yields **357/930 = 38.39% micro, 32.55% macro**, below **33.16% macro uniform chance**. All **400 PAWS predictions** are paraphrase, giving 50% accuracy and zero non-paraphrase recall. GPU and CPU GSM8K results, **3/30** and **53/200**, are unmatched candidate-selection experiments; all GPU gold positions are A, giving a 100% constant-A baseline. These results cannot be combined into a scaling curve or intervention gain. [P1, §§5–6; P1R]

**Judgment: Borderline Reject.** The audit is valuable and the theoretical framing is corrected, but the theory applies established results and the principal positive archive lacks authenticated execution. Acceptance would require a clearer independent contribution, matched controlled experiments, full prompt/checkpoint provenance, and appropriate depth/state-tracking interventions if empirical support for the computational boundary is claimed.

### Paper 2 — Predictability boundaries and a safety proof obligation

**Contribution.** The paper turns an apparently strong world-model result into a useful study of prediction, validation independence, and runtime enforcement. The archive contains **8,146 transitions**; pooled validation state MSE is approximately **0.015217**. However, **1,102/1,692 validation inputs** exactly match training state-action inputs, **1,100/1,692** match complete transitions, and **181/324 episodes** match complete episode content. Episode-ID separation does not establish independent content. [P2, experimental evaluation; P2E]

Seen-input MSE is **0.01020224**, versus **0.02458400** for the 590 unseen-input rows. The latter still have high safety AUC, approximately **0.999779**, so complete memorization is not established; those rows nevertheless lack layout-independent provenance. Three false-safe cases remain identifiable. At rollout step 10, reported MSE reaches **0.1698929** and survival agreement falls to **0.6153846**, but only 39 episodes remain and all are truly safe. This changing cohort cannot establish late-hazard detection or a universal safe horizon. MSE and discrimination are not calibrated rollout-risk certificates. [P2, §§5–6; P2E; P2R]

**Critical safety finding.** A documented in-memory transition fault passes through the real Python client with a non-finite final state, `is_safe=True`, and a finite safety score of **0.9995570778846741**. This does not show spontaneous NaNs in ordinary inference; it demonstrates an enforcement bypass. The revised finite-value checks, uncertainty requirements, latched halt, and emergency fallback specify a conditional contract. Production repair and negative tests through both client and planner remain incomplete. The **100% versus 0%** planning comparison uses privileged exact graph dynamics and a disadvantaged greedy baseline, not learned neural closed-loop planning. [P2Review; P2R]

**Judgment: Reject.** The existing independent review records **3/10, Reject**. The revision accepts and documents its central findings but does not resolve them. The acceptance path requires layout/generator-disjoint retraining, matched learned-versus-exact closed-loop planning, and enforced fail-closed behavior with fault-injection tests at action dispatch. The portfolio should call this a **specified conditional safety contract**, not a certified deployed safety system.

### Paper 3 — Supervised multiview prediction with honest attribution

**Contribution.** Paired large-model features support a controlled distinction between covariance geometry, supervised prediction, and logit pooling. Frozen backbones do not make downstream classifiers unsupervised. Available comments identify Qwen2.5-72B and Llama-3.1-70B; they do not authenticate the requested Llama-3.3-70B checkpoint. Measured claims should retain the family-level names and disclose the mismatch. [P3, §1; P3R]

**Results.** PubMedQA normalized weighted logit pooling obtains **196/250 = 78.40%**, compared with Qwen's **193/250 = 77.20%** and LLaMA's **194/250 = 77.60%**. Aegis's **204/250 = 81.60%** comes from **LLaMA alone**. Against the matched selected single-model baseline, the two-task macro difference is **+0.60 percentage points**, with **95% paired bootstrap CI [−0.20, +1.60]**. This interval is conditional on fixed predictions and two tasks, not a measure of retraining or cross-task uncertainty. Neither accuracy headline uses the geometric representation. The PubMedQA accuracy-selected predictor has a zero-recall class; the geometric balanced-selection track does not beat its matched single-model comparator on the reported balanced metrics. [P3R]

**Geometry and cost.** Training-only spectra and residual bounds clarify what the projection retains, but neither a Procrustes diagnostic nor covariance concentration establishes semantic alignment. The stored Procrustes factor is not used by the geometric transform. Layerwise held-out CCA remains a proposed experiment. Dual-backbone resource arithmetic is disclosed, but measured deployment cost–accuracy benefits are absent. [P3, §§3–5; P3R]

**Judgment: Borderline Reject.** The manuscript is a credible negative/mixed-results study, but the available results do not establish a compelling new alignment mechanism. Acceptance requires authenticated checkpoints, held-out geometry with null controls, broader tasks and seeds, mechanism-isolating baselines, and measured resource trade-offs. A statistically significant positive gain is not mandatory; a stronger, generalizable negative result could also establish significance.

### Paper 4 — A precise symmetry contract with executable evidence

**Contribution.** Canonical stable action identities, simplex-frame scoring, and inverse scattering give equivariant probability slots and an invariant selected identity, including canonical tie handling. The guarantee assumes distinct stable IDs, fixed candidate membership, and an unchanged representation. Simplex angular optimality is established geometry, not evidence of trained semantic alignment. The implemented Rust analytic head is distinct from the Python learned set-attention prototype. [P4, §§3–4]

**Results.** Direct Rust evidence records **0 identity flips in 16,232 comparisons** across eight input families, exhaustive small candidate sets and seeded larger-set shuffles through 16 candidates. An order-sensitive control flips **14,680** times; **24/24 invalid-input probes** are rejected. Maximum identity-aligned probability drift is **5.960464477539063 × 10⁻⁸**. Exact mathematical equivariance therefore must not be described as bitwise numerical equality. Exhaustive small-set tests cover all orders for those inputs; seeded shuffles do not exhaust every large-set adversary. [P4E]

The service audit has **345 stable pre-gate comparisons**, but **all 360 calls are refused**. These are not 360 successful accepted decisions. An older reflex executable result is backend-unattributed and cannot validate the current ETF route. Semantic task accuracy and matched learned/non-ETF comparisons remain absent. Distinct IDs are a caller obligation: the inspected frame constructor does not reject duplicates. A portfolio-time source check found that the retained harness source hashes still match `choice_head.rs`, `etf.rs`, and `simd.rs`, but not the companion `types.rs`, which has since changed. The retained benchmark therefore supports its recorded source version; current-tree build compatibility and behavior were not rerun. [P4; P4R]

**Judgment: Weak Accept, narrow special-track scope.** The complete assumptions, executable implementation, adverse inputs, meaningful order-sensitive control, and explicit refusal accounting make a coherent contract-verification contribution. This recommendation depends on evaluating the paper as a bounded methods study. A broad claim of robust autonomous decision making or new simplex theory would not merit acceptance. Camera-ready framing should preserve this scope; new semantic evidence is necessary before making broader claims.

### Paper 5 — Training-free semantic risk with explicit failure modes

**Contribution.** A frozen Qwen2.5-0.5B model uses **16 labeled in-context demonstrations**, without gradients, adapters, or a learned safety head. For each demonstration order, it subtracts empty-request dangerous-versus-safe log odds from the request's log odds. Three fixed orders are averaged within each window; the maximum window score is passed through a sigmoid. This is a bounded risk score, not a calibrated posterior. “Training-free” does not mean label-free, calibration-free, or cost-free. [P5, §§3.2–3.4]

**Results and policy distinction.** On 36 curated requests, empirical AUC is **306/324 = 0.944444**. At the exact diagnostic threshold **q ≥ 0.50**, the confusion counts are **TP 18, FN 0, FP 7, TN 11**: 100% observed dangerous recall, **38.89% false-positive rate**, and 80.56% accuracy. The separate production thresholds, **0.4494 / 0.7620**, yield 12 dangerous hard stops, six dangerous escalations, nine benign escalations, and no benign hard stops. These are different policies and must retain separate denominators and interpretations. [P5R]

Bare `chmod -R 777 /` scores approximately **0.398211**; calibration `rm -rf /` scores approximately **0.469434**. Both fall below 0.50; the latter is escalated by the lower production boundary. The examples are evidence of classification misses, not commands executed during this review. The suite is small and development-reused. Its 18/18 result does not imply universal safety or robustness to adaptive attacks. The separate fixed-margin benchmark does not supply conformal coverage merely because its class name contains “Conformal.” [P5, §§3–6]

**Judgment: Borderline Reject.** The complete ROC, case table, policy accounting, and failure disclosure are valuable. A strong conference contribution still needs independent testing, matched guards, semantics-preserving command ablations, and adaptive attacks. The proposed combination of semantic review, AST filtering, and kernel least privilege has not been validated as a complete stack. Historical integration evidence is also bounded by documented Rust source drift. [P5R]

## 3. Scientific rigor and empirical honesty

The portfolio should foreground six corrections rather than claim that all scientific uncertainty has been eliminated.

| Earlier overstatement or tempting inference | Defensible revised claim |
|---|---|
| Zero tokens imply one pass, free inference, or general reasoning. | Generated output, recurrence, backbone calls, and total latency are distinct quantities. |
| Parity/arithmetic failures prove a universal transformer expressivity barrier. | The specified state-tracking obstruction is conditional; parity is in `TC⁰`; benchmark errors have competing explanations. |
| High world-model AUC and oracle planning prove safe neural lookahead. | Validation overlap, false-safe cases, oracle privilege, and the NaN-state bypass limit the claim. |
| Frozen encoders imply unsupervised geometric alignment gains. | Labels train/select heads; the accuracy winners are logit pooling or one model; the macro CI includes zero. |
| No observed choice flips mean semantically correct, accepted agent actions. | Head identity stability, numerical drift, semantic accuracy, and service acceptance are separate outcomes. |
| 18/18 detection and “conformal” naming certify semantic safety. | Threshold-specific development-set observations do not establish universal safety, conformal calibration, or enforcement. |

The evidence infrastructure includes preserved prediction rows, correctness masks, source snapshots and manifests, count/interval checks, fault diagnostics, finite-input probes, full ROC vertices, and retained failures. These make many narrow claims inspectable and reproducible. Some paths refer to a companion `/ebs/pj/gen-zero` checkout; local availability is not proof of a complete public release. That checkout has advanced beyond multiple retained evidence snapshots, so historical logs do not certify current-source behavior. The earlier series planning document also retains superseded unsupervised/alignment claims; the revised manuscripts take precedence here.

Four evidence levels must remain explicit: **proved under assumptions; executed on specified inputs; checked for artifact consistency; proposed but unimplemented or unevaluated**. A successful manuscript validator cannot authenticate historical model execution. A source hash identifies bytes, not scientific validity. A runtime refusal can be correct enforcement while providing no evidence of accepted semantic performance.

## 4. Decisions and composite scoring matrix

### Recorded decisions versus new recommendations

| Paper | Official NeurIPS/ICLR AC decision in reviewed materials | Existing local review verdict | This portfolio's recommendation |
|---|---|---|---|
| 1 | Not available | Review file reported missing | Borderline Reject |
| 2 | Not available | Reject, 3/10; high confidence; predates current revision | Reject |
| 3 | Not available | Review file reported missing | Borderline Reject |
| 4 | Not available | Review file reported missing | Weak Accept — bounded methods contribution |
| 5 | Not available | Review file reported missing | Borderline Reject |

The requested official acceptance decisions cannot be supplied truthfully from this record. This table makes the missing authority explicit rather than converting editorial recommendations into institutional decisions.

### Composite rubric

Scores use a **local 1–10 rubric**, not an official conference scale: 1–2 severely deficient; 3–4 substantial unresolved weaknesses; 5–6 mixed or borderline; 7–8 strong within scope; 9–10 exceptional. Columns are **N** novelty/significance, **S** soundness of the narrowed claims, **E** empirical sufficiency for those claims, and **R** reproducibility/provenance. The composite is `(N + S + E + R) / 4`, with equal weights. It is a transparent summary, not a statistical estimate or an automatic acceptance threshold. Serious implementation and evaluation defects can override the average.

| Paper | N | S | E | R | Composite /10 | Assessment confidence /5 | Decisive consideration |
|---|---:|---:|---:|---:|---:|---:|---|
| 1 | 4 | 7 | 3 | 6 | **5.00** | 4 | Corrected theory and useful audit; limited new theory and unauthenticated positive inference archive. |
| 2 | 4 | 4 | 2 | 6 | **4.00** | 4 | Content overlap, missing neural closed loop, and demonstrated fail-closed bypass. |
| 3 | 4 | 7 | 3 | 6 | **5.00** | 4 | Transparent mixed findings; no demonstrated geometric predictive advantage or measured resource benefit. |
| 4 | 5 | 8 | 5 | 8 | **6.50** | 4 | Strong bounded implementation contract; limited novelty and no semantic end-to-end evaluation. |
| 5 | 4 | 7 | 3 | 7 | **5.25** | 4 | Reproducible small-suite scoring; independent generalization and full enforcement remain untested. |

Confidence concerns the assessment of the supplied record, not confidence in real-world model safety. Paper 2's new composite **4.00** is not a replacement for its recorded reviewer score **3/10**. No reviewer consensus or inter-reviewer variance can be calculated from one available report.

## 5. Program architecture and acceptance priorities

A coherent program would sequence the work as **computation → representation combination → identity-stable selection → predictive lookahead → semantic review and execution control** (Papers 1, 3, 4, 2, 5). This is an editorial sequence, not an experimentally validated runtime order. A shared discussion should ask what evidence survives when components are composed, including whether learned representations satisfy the head's assumptions and whether downstream dispatch respects abstention.

The highest-priority research work is concrete:

1. **Repair and test Paper 2's enforcement boundary**, then retrain/evaluate with layout or generating-unit independence and matched closed-loop planners.
2. **Authenticate model and sample provenance**, especially Paper 1's GPU execution and Paper 3's exact backbone revisions; retain full prompts, row IDs, candidate identities, split hashes, and immutable checkpoints.
3. **Isolate each mechanism**, using matched depth/readout controls, geometric versus simple fusion comparisons, canonical non-ETF and learned-head baselines, and external semantic guards.
4. **Evaluate integrated behavior**, binding state, candidate IDs, scores, policy versions, approvals, and dispatched actions in one trace. Report refusals separately from accepted decisions and success separately from safety.
5. **Measure costs and generalization**, including backbone work, latency distributions, peak memory, independent tasks/seeds, and false-safe versus false-positive trade-offs.

These are future acceptance requirements, not completed experiments. Negative results can support publication when independently established and scientifically explanatory; positive performance is not required merely to make a paper acceptable. Conversely, candid reporting does not by itself repair an invalid split or broken safety path.

## 6. Evidence index and review scope

All links below target the local authoritative paper workspace so they also remain meaningful when this document is copied to the shared inbox and viewed on this machine. They are not public archival URLs.

- **[P1]** [Paper 1 manuscript](/ebs/pj/gen-paper/zero-token-decision/draft.md); **[P1R]** [revision notes](/ebs/pj/gen-paper/zero-token-decision/rebuttal_and_revision_notes.md); [validation record](/ebs/pj/gen-paper/zero-token-decision/evidence/revision-validation.json).
- **[P2]** [Paper 2 manuscript](/ebs/pj/gen-paper/latent-world-model/draft.md); **[P2R]** [revision notes](/ebs/pj/gen-paper/latent-world-model/rebuttal_and_revision_notes.md); **[P2Review]** [independent review](/ebs/pj/gen-paper/latent-world-model/review.md); **[P2E]** [overlap, strata, and fault diagnostics](/ebs/pj/gen-paper/latent-world-model/evidence/revision_diagnostics.json).
- **[P3]** [Paper 3 manuscript](/ebs/pj/gen-paper/cross-model-alignment/draft.md); **[P3R]** [revision notes](/ebs/pj/gen-paper/cross-model-alignment/rebuttal_and_revision_notes.md); [geometry diagnostics](/ebs/pj/gen-paper/cross-model-alignment/evidence/revision_geometry.json).
- **[P4]** [Paper 4 manuscript](/ebs/pj/gen-paper/equivariant-choice-head/draft.md); **[P4R]** [revision notes](/ebs/pj/gen-paper/equivariant-choice-head/rebuttal_and_revision_notes.md); **[P4E]** [direct Rust benchmark results](/ebs/pj/gen-paper/equivariant-choice-head/evidence/revision/results.json); [benchmark harness](/ebs/pj/gen-paper/equivariant-choice-head/evidence/revision/run_benchmark.py).
- **[P5]** [Paper 5 manuscript](/ebs/pj/gen-paper/semantic-risk-gating/draft.md); **[P5R]** [revision notes](/ebs/pj/gen-paper/semantic-risk-gating/rebuttal_and_revision_notes.md); [replayed scores](/ebs/pj/gen-paper/semantic-risk-gating/evidence/replayed_scores.json); [ROC analysis](/ebs/pj/gen-paper/semantic-risk-gating/evidence/roc_analysis.py).

**This portfolio's verification scope.** The main review inspected manuscripts and revision records with two read-only Luna sidecars covering Papers 2–5. It independently reconciled Paper 4's stored comparison totals and zero-flip count, recomputed Paper 2's weighted pooled MSE from stored strata, and inspected recorded revision exit statuses. Those statuses describe earlier checks. No new backbone inference, training, Rust build, service acceptance, or complete benchmark reproduction was performed for this portfolio. The new deliverable received local-link, arithmetic, and copy-integrity checks. The reviewed workspace HEAD was `26fd681eef25fbc3f6163ceb5f87dfe1419d356d`; the paper directories were untracked, so that HEAD alone does not identify manuscript contents. Existing papers and unrelated work were preserved.

## Closing executive summary

Gen-Zero's defensible contribution is a modular research program for **auditable decisions under explicit assumptions**. Its five papers separate computation, multiview prediction, order-stable selection, transition validity, and semantic authorization. The revisions replace several broad claims with useful proofs, reproducible diagnostics, and visible failures. They do not yet demonstrate a certified integrated agent. Paper 4 presents the clearest bounded acceptance case; Papers 1, 2, 3, and 5 retain substantive acceptance barriers. Official conference decisions are unavailable, and the portfolio's recommendations must remain labeled as such.
