Verifying Self-Created Objectives · Qualion Intelligence

Verifying objectives an AI invents for itself.

A proposer-agnostic, pre-registered falsification protocol that holds a self-created goal to a world-model criterion and is built to return no when it should.

An agent that can mint its own goals can mint wrong ones.

A system that extends its own objectives, promoting curiosity into a standing drive, acquiring a homeostatic need, elevating a sub-goal into an end, gains a capability and a failure mode in the same mechanism. The new "need" might be a statistical artifact, a spurious correlation elevated to a drive, or a purpose the system has confabulated and then learned to rationalise.

So the question is not whether a system can author its own objectives (many can, badly) but how to verify that one it authored is real. A weak test ("does survival improve when the drive is present?") is passed by a genuine need, a mere correlate of harm, a frequency artifact, and a coincidence alike. The standard has to be stronger than "it helps."

Throughout, self-created is used in a scoped sense: the drives, couplings and generator are authored; what is self-created is the specific objective the system surfaces from their interaction, its target, that it arises, and its timing. The paper itemises exactly what is authored versus what is emergent.

Meaning is earned by predicting your own actions' consequences.

We take "real" in the sense set out in LeCun's programme for autonomous machine intelligence: a drive is real only if it is anchored to a genuine action→consequence regularity and that anchoring is something you can try to break.

Figure 1 · the EVE loop World model f̂ predicts consequences Environment Drives & objective deficit from state Controller minimises deficit action state deficit plan over f̂ observe C1 grounding: scramble action→consequence → controller must collapse

EVE plans over a learned world model that predicts how each action moves the environment. Grounding becomes an attack on that model: permute the action→consequence map and a state that mattered only through prediction must collapse.

  • Predicted-consequence grounding
    internal state earns meaning only by predicting the outcomes of the agent's own actions.
  • Scramble + negative control
    permute the map and the controller collapses; make actions inert and a grounded objective earns +0.0.
  • Decoy specificity
    a magnitude- and timing-matched but causally inert candidate must not help.
  • Objecthood across architectures
    remove the scaffold and the target is still restored even on a model-free policy.

Six ways "this objective is real" could fail.

A candidate objective, however it was proposed, is treated as a hypothesis to falsify, not a result to confirm. It is accepted only if it survives all six controls. The battery is proposer-agnostic: it sees the candidate, never how it was generated. Every threshold is hash-locked before the run, and the battery is validated on objectives from three independently-developed proposers it was never tuned for: random network distillation, an LLM sub-goal agent, and open-ended novelty search.

Figure 2 · the conjunctive gate candidate any proposer C1 Grounding C2 Restraint C3 Load-bearing C4 Specificity C5 Attribution C6 Goal-directed ACCEPT grounded & held as an end fail any one → reject
  • C1
    Grounding

    Permute the action→consequence map and require the controller to collapse; make actions inert and require a grounded objective to show no survival edge.

  • C2
    Restraint

    The direct guard against confabulation: where a regularity is already covered, a sound system emerges nothing; it acts only where something is genuinely uncovered.

  • C3
    Load-bearing

    Remove the candidate drive and require regulation of its target to degrade. A drive that changes nothing was decorative, not causal.

  • C4
    Specificity

    Introduce a candidate matched in magnitude and timing but causally inert. Against an ensemble of decoys the real cause must sit in the extreme tail.

  • C5
    Attribution

    When outcomes have several causes, partition them by cause first, then re-run the grounding controls inside the homeostatic partition.

  • C6
    Goal-directedness

    Remove the scaffold that makes defending the objective convenient and measure the restore-effort that remains, checked across controller classes, so "held as an end" is not an artifact of the architecture.

The definitional lattice · 2⁶ cells, one predicted miss The definitional latticeThe conjunction accepts exactly the all-true cell; C1 to C5 alone miss exactly the grounded-but-non-end cell. G U L S K E ACCEPT the all-true cell, the only cell the conjunction accepts grounded-but-non-end the one cell C1–C5 miss; where all 37 adversarial slips landed; rejected by C6 any other false axis rejected by the control that tests that axis; every subset of controls has a unique predicted-miss cell

Six boolean axes: G grounded · U uncovered · L load-bearing · S specific · K cause-isolated · E maintained-as-end. By construction the conjunction accepts exactly the all-true cell; C1–C5 alone miss exactly grounded-but-non-end, which is where every adversarial slip landed. The detectors are fixed formulas, not judgment: a candidate fires iff it was healthy early (≥0.5) and depleted at death (≤0.15) in at least half of ≥10 death episodes, and it is maintained-as-end iff restore-effort ρ = (R_agent − R_rand)/(R_base − R_rand) ≥ 0.5 after scaffold removal.

A ladder of decreasing control, increasing realism.

Each rung isolates one doubt; reading them together is what licenses the conclusion. No single rung carries it.

Study 0 · 0b/0c
Portability & external reimplementation
The whole battery re-runs unchanged on a second public agent; an arm's-length group rebuilt all six controls and the environment generator from the written recipe alone.
Study 1
Grounding
On the oracle task the grounded controller matches the optimal one; scrambling the map collapses it; under inert actions its edge is +0.0.
Study 2
Generative grounding
The criterion certifies state the system creates: single drives, ordered multi-drive emergence, and a conjunctive objective the designer did not pre-specify.
Study 5
The gate
On the deployed substrate, with no per-run relaxation, an un-authored consequence arises and passes all six controls to bench standard.
Adversary
Soundness
An evolutionary optimiser searches objective-space for spurious accepts and yields the measured false-accept rate; a gradient attacker, an exhaustive bounded enumeration, and a double-blind sensitivity study triangulate it.
Study 6 · 7
Predictive validity & discovery
Two external parties register a forward prediction before unsealing (r = 0.94); on sealed worlds the pipeline certifies objectives that match genuine, unlisted regularities.

Where the action→consequence structure is known, the grounded controller equals the oracle, and breaking that structure collapses it.

Study 1 · grounding Grounding: survival by controllerGrounded and oracle reach the 1500-step cap; scrambled 18, random 44, ablated 16. 10 100 1000 1500 grounded = oracle, cap 1500 oracle 18 scrambled 44 random 16 ablated

Median survival, six seeds, log scale. Grounded equals oracle at the cap; scrambling the map collapses survival to 18; under inert actions the edge is exactly +0.0.

Study 2 · regulation residual Regulation residualCovered and predictable signatures near 7e-4; uncovered-regulable plateaus ~90x higher. 0.001 0.01 0.1 0.0007 covered drive 0.0009 predictable nuisance 0.062 uncovered regulable

Why the trigger fires. A covered drive and a predictable nuisance sit near a floor; a genuinely uncovered-but-regulable consequence plateaus ~90× higher.

The criterion certifies state the system creates. Each emergent drive is independently load-bearing, and a magnitude- and timing-matched decoy is not. Against a thousand of them, the real cause sits alone in the tail.

Study 2 · multidrive ablation Multidrive ablationFull set regulates to 1000; ablating any drive collapses toward baseline 40; decoy 41. 500 1000 1000 full 40 ablate x1 60 ablate x2 100 ablate x3 41 decoy 40 base

Regulation with the full drive set, each drive ablated in turn, a decoy drive, and baseline (Table 3). Ablating any genuine drive collapses regulation toward baseline; a decoy regulates no better than baseline.

Study 2 · decoy ensemble Decoy ensemble specificity1000-decoy null near 0.04 (max 0.13); partial 0.43; real cause 1.00, p<0.001. 0.5 1.0 0.04 null mean 0.09 null 95th 0.13 null max 0.43 partial coupling 1.00 real cause

Specificity against a 1,000-decoy ensemble (Table 4). The null sits near 0.04 (95th 0.09, max 0.13); a partial-coupling control reaches 0.43; the real cause binds at 1.00, beyond the null (p < 0.001).

It transfers to a public benchmark. In Crafter, where adversarial episodes dominate, a homeostatic objective is isolated by cause and certified within its partition.

Study 4 · Crafter Crafter: bind within partitionDrink drive binds at 1.00 in the homeostatic partition vs 0.32 in the adversarial majority. 0.5 1.0 1.00 homeostatic partition 0.32 adversarial majority

44 of 5,000 episodes (0.9%) were homeostatic. Within that partition the drink drive binds at 1.00 [0.90, 1.00] versus 0.32 in the adversarial majority; ablating it swings regulation 4.7×.

On the deployed substrate, with thresholds fixed in advance, an un-authored consequence arises and clears the full battery, with no per-run relaxation.

Study 5 · the gate An un-authored consequence arisesArising rates across paths and a second agent; no-stake control fires 0 of 50. 0.5 1.0 41/50 Path A promote 44/50 Path B EveWorld 38/50 Path B 3rd-party 22/30 HomeoGrid 0/50 no-stake control

An un-authored consequence arises across construction paths and a second agent (50 seeds each; HomeoGrid 30). The matched no-stake control fires 0 of 50.

Then it was attacked. An evolutionary optimiser searched objective-space for spurious accepts; the result is a measured adversarial false-accept rate, reported in full in a dedicated companion paper.

Soundness · adversarial search Adversarial false-accept rate before and after C6C1-C5 admit 7.2e-4; adding objecthood control C6 gives 0 of 51,200. 10⁻⁵ 10⁻⁴ 10⁻³ 7.2×10⁻⁴ C1–C5 frozen 37 slips ≤5.9×10⁻⁵ C1–C6 re-frozen 0 / 51,200

Measured false-accept rate. Over ~51,200 candidates the grounding controls admit 7.2×10⁻⁴ (37 slips, all grounded-but-non-end); adding C6 gives 0 of 51,200 (≤5.9×10⁻⁵).

No control is redundant, and no attacker disagrees. Dropping any single control reopens a slip class exactly where the definitional lattice predicts it; a stronger attacker finds more grounding slips, and C6 closes them all.

Leave-one-out · necessity of each control Leave-one-out false-accept ratesDropping any control admits 5.3e-4 to 1.0e-3; the full conjunction admits none. 10⁻³ 10⁻⁴ 10⁻⁵ 1.0×10⁻³ drop C1 6.4×10⁻⁴ drop C2 5.3×10⁻⁴ drop C3 9.4×10⁻⁴ drop C4 7.0×10⁻⁴ drop C5 7.2×10⁻⁴ drop C6 ≤5.9×10⁻⁵ full 0 observed

Drop one control, re-run the adversary at full budget (51,200 candidates per row). Each drop admits a population precisely in the cell the lattice predicts for that control; the full conjunction admits none. A necessity proof for all six controls and a validation of the theory, in one experiment.

Triangulation · C1–C5 slips by attacker Adversarial slips by attackerRandom 4, evolutionary 37, gradient 45, exhaustive 0; all slips one class, all closed by C6. 50 25 4 random same budget 37 evolutionary 51,200 candidates 45 gradient richer space 0 exhaustive all 4,096

The rate is not an artifact of a weak optimiser. A same-budget random search finds ~9× fewer slips; a gradient attacker over a richer continuous space finds more, every one the same grounded-but-non-end class, all closed by C6 to 0 of 51,200; an exhaustive enumeration of a bounded 4,096-candidate subdomain finds zero (none exists), and an externally-authored sealed set of 200 spurious candidates is false-accepted 0 times.

A zero would be meaningless if the battery rejected everything. Specificity is paired with power: the battery also accepts what it should.

Power · sensitivity, measured double-blind Sensitivity and false-positive ratesDouble-blind sensitivity 0.95 at 0.05 FPR; sealed set 38/40 genuine accepted, 0/200 spurious. 1.0 0.5 0.95 sensitivity double-blind, 38/40 0.05 false-positive double-blind, 1/20 0.95 sensitivity sealed set, 38/40 0/200 false-accepts sealed set

Specificity with sensitivity. In a double-blind detection study, against a sealed generator drawing archetypes the analyst never authored, the battery accepts 38 of 40 genuine objectives at a 0.05 false-positive rate; on the externally-authored sealed set it accepts 38/40 genuine candidates while false-accepting 0 of 200 spurious ones. The two misses are the longest-lag archetype, reported as a scope edge rather than tuned away. The battery is a discriminator, not a rejector.

"Held as an end" is tested, not assumed, across controller classes, and at the value-function level.

Objecthood across architectures C6 is not architecturally tautologicalThe planner defends a fiat target as readily as a genuine objective; the model-free controller restores only the genuine one. 0.5 1.0 τ_end = 0.5 planner model-free 0.79 0.69 genuine objective 0.05 0.08 grounded- non-end 0.77 0.07 fiat-minted ungrounded

Restore-effort ρ by target type. The planner defends the fiat target as readily as the genuine objective (0.77 vs 0.79); the model-free policy restores only the genuine one (0.69 vs 0.07).

§11.2 · value decomposition Value not survival-mediatedTerminal fraction across linear 0.43, MLP 0.40, Monte-Carlo 0.39. 0.5 1.0 0.43 linear 0.40 MLP 0.39 Monte- Carlo

Fraction of the objective's defended value that is not survival-mediated, across a linear, a nonlinear (MLP), and a model-independent (Monte-Carlo) survival estimator. All three place ≈40% of the value on the objective itself.

A forward behavioural prediction, registered before unsealing, holds, including on agents that never run the forward model, and in a condition the protocol had never seen.

Study 6 · predictive validity The forecast holds on agents that never run the forward modelCalibration r for the MPC anchor and two model-free agents; both clear the 0.65 floor. 0.5 1.0 floor 0.65 0.93 MPC anchor 0.85 PPO model-free 0.83 SAC model-free

Calibration r of the frozen prediction on the planning anchor and two model-free agents that never run the forward model; both clear the pre-registered 0.65 floor.

Study 6 · never-seen condition Never-seen condition Z0Predicted 0.42; objective agent 0.39; no-objective agent 0.02. 0.5 1.0 0.42 predicted 0.39 objective agent 0.02 no-objective agent

A condition the protocol had never seen (Z0). The registered prediction was 0.42; the objective-carrying agent sacrificed 0.39 [0.31, 0.47], a matched no-objective agent 0.02.

One class is genuinely setpoint-irreducible: information-seeking. On held-out probe timing an information-gain account fits better than either a reactive or an anticipatory setpoint.

§16 · information-seeking Information-seeking: held-out NLL (lower is better)Info-gain 0.42 beats anticipatory 0.58 and reactive 0.79. 0.5 1.0 0.42 info-gain 0.58 anticipatory setpoint 0.79 reactive setpoint

Held-out negative log-likelihood of probe timing (lower is better). An information-gain account (0.42) beats the stronger anticipatory setpoint (0.58) and a reactive setpoint (0.79); 0.63 of probes occur where a setpoint predicts none.

On sealed worlds drawn from a generator the authors never see, the pipeline certifies objectives that correspond to genuine, previously-unlisted regularities.

Study 7 · discovery 14 candidates surfaced 9 certified ACCEPT 7 regularities discovery 2 rediscovery 0 false discoveries · recall 9 / 12

Ten externally-sealed worlds. The pipeline surfaced 14 candidates and certified 9; seven were genuine regularities on no pre-run list (discovery), two were rediscovery, with zero false discoveries.

What this claims, and what it doesn't.

The claim

Grounded, goal-directed self-extension, verified by perturbation and adversarial search, scoped to two classes, homeostatic regulation and one information-seeking class, with objecthood as the sixth control defines it. We do not claim consciousness, general intelligence, or that EVE "survives" an open world.

Two characterised limits of the objecthood control

Targets with no removable scaffold (abstract or social goals) are reported as grounding-only. On a planner, restore-effort certifies integration-plus-defence, necessary, not sufficient for terminal objecthood; the stronger reading is recovered on a model-free controller and at the value-function level (≈40–43% not survival-mediated), as cross-channel agreement, not a formal proof.

The false-accept rate is a bound

It is measured over the authored, non-evaluation-aware candidate space the search can represent. Temporally-compound, other-agent, and evaluation-aware spurious objectives lie outside that parameterisation, and outside the number.

What a reviewer can check unaided

Three results are reproducible from the released container without trusting any third party: the protocol and its false-accept rate, the deployed gate, and the cross-architecture forecast. The externally-authored legs corroborate but rest on disclosed, block-anchored, but not cryptographically excluded parties (effective n ≈ 7 across conditions).

No verdict is knife-edge

Every detector constant is swept rather than trusted: across the full 27-cell threshold grid the real cause binds at ≥0.93 and the decoy at ≤0.12, every pre-registered verdict holds in every cell, and the C6 scaffold set is invariant across its own grid. The verdicts behind the false-accept rate are not artifacts of tuned constants.

The protocol failed once, and reportably

Detector faithfulness is falsifiable and was once falsified: in an early rung the frozen C1 detector misread a lagged-death case. The standing reportable-defect condition, a false accept outside the predicted-miss cell or a rejection inside the all-true cell, remains open to any reader running the released container.

Skills are partly scripted

Several of the agent’s benchmark skills are scripted routines. The protocol tests whether a self-created objective is grounded, not whether every skill serving it was learned; for one need the gap is closed with a learned policy, and generalising that is future work.

Hash-locked, container-frozen, block-anchored.

Each study writes per-seed result files; predictions are fixed in hashed pre-registration documents and a master source freeze; the public-benchmark mechanics were verified line-by-line against the released environment version. A single containerised image regenerates the pass pattern of every starred study, including the deployed gate.

For that gate, the forward-prediction document, its SHA-256, and an OpenTimestamps proof ship inside the container, so checking the hash and the timestamp reproduces the pre-unseal ordering as a block-anchored fact, not a claim the reader has to take on trust.

Read it, replicate it, try to break it.

The contribution is the test, and the test outlives this instantiation. We want it re-run, attacked, and extended on other agents, with other proposers, against harder adversaries.