Verifying objectives an AI invents for itself.
A proposer-agnostic, pre-registered falsification protocol that holds a self-created goal to a world-model criterion and is built to return no when it should.
An agent that can mint its own goals can mint wrong ones.
A system that extends its own objectives, promoting curiosity into a standing drive, acquiring a homeostatic need, elevating a sub-goal into an end, gains a capability and a failure mode in the same mechanism. The new "need" might be a statistical artifact, a spurious correlation elevated to a drive, or a purpose the system has confabulated and then learned to rationalise.
So the question is not whether a system can author its own objectives (many can, badly) but how to verify that one it authored is real. A weak test ("does survival improve when the drive is present?") is passed by a genuine need, a mere correlate of harm, a frequency artifact, and a coincidence alike. The standard has to be stronger than "it helps."
Throughout, self-created is used in a scoped sense: the drives, couplings and generator are authored; what is self-created is the specific objective the system surfaces from their interaction, its target, that it arises, and its timing. The paper itemises exactly what is authored versus what is emergent.
Meaning is earned by predicting your own actions' consequences.
We take "real" in the sense set out in LeCun's programme for autonomous machine intelligence: a drive is real only if it is anchored to a genuine action→consequence regularity and that anchoring is something you can try to break.
EVE plans over a learned world model that predicts how each action moves the environment. Grounding becomes an attack on that model: permute the action→consequence map and a state that mattered only through prediction must collapse.
- Predicted-consequence groundinginternal state earns meaning only by predicting the outcomes of the agent's own actions.
- Scramble + negative controlpermute the map and the controller collapses; make actions inert and a grounded objective earns +0.0.
- Decoy specificitya magnitude- and timing-matched but causally inert candidate must not help.
- Objecthood across architecturesremove the scaffold and the target is still restored even on a model-free policy.
Six ways "this objective is real" could fail.
A candidate objective, however it was proposed, is treated as a hypothesis to falsify, not a result to confirm. It is accepted only if it survives all six controls. The battery is proposer-agnostic: it sees the candidate, never how it was generated. Every threshold is hash-locked before the run, and the battery is validated on objectives from three independently-developed proposers it was never tuned for: random network distillation, an LLM sub-goal agent, and open-ended novelty search.
- C1Grounding
Permute the action→consequence map and require the controller to collapse; make actions inert and require a grounded objective to show no survival edge.
- C2Restraint
The direct guard against confabulation: where a regularity is already covered, a sound system emerges nothing; it acts only where something is genuinely uncovered.
- C3Load-bearing
Remove the candidate drive and require regulation of its target to degrade. A drive that changes nothing was decorative, not causal.
- C4Specificity
Introduce a candidate matched in magnitude and timing but causally inert. Against an ensemble of decoys the real cause must sit in the extreme tail.
- C5Attribution
When outcomes have several causes, partition them by cause first, then re-run the grounding controls inside the homeostatic partition.
- C6Goal-directedness
Remove the scaffold that makes defending the objective convenient and measure the restore-effort that remains, checked across controller classes, so "held as an end" is not an artifact of the architecture.
Six boolean axes: G grounded · U uncovered · L load-bearing · S specific · K cause-isolated · E maintained-as-end. By construction the conjunction accepts exactly the all-true cell; C1–C5 alone miss exactly grounded-but-non-end, which is where every adversarial slip landed. The detectors are fixed formulas, not judgment: a candidate fires iff it was healthy early (≥0.5) and depleted at death (≤0.15) in at least half of ≥10 death episodes, and it is maintained-as-end iff restore-effort ρ = (R_agent − R_rand)/(R_base − R_rand) ≥ 0.5 after scaffold removal.
A ladder of decreasing control, increasing realism.
Each rung isolates one doubt; reading them together is what licenses the conclusion. No single rung carries it.
Where the action→consequence structure is known, the grounded controller equals the oracle, and breaking that structure collapses it.
Median survival, six seeds, log scale. Grounded equals oracle at the cap; scrambling the map collapses survival to 18; under inert actions the edge is exactly +0.0.
Why the trigger fires. A covered drive and a predictable nuisance sit near a floor; a genuinely uncovered-but-regulable consequence plateaus ~90× higher.
The criterion certifies state the system creates. Each emergent drive is independently load-bearing, and a magnitude- and timing-matched decoy is not. Against a thousand of them, the real cause sits alone in the tail.
Regulation with the full drive set, each drive ablated in turn, a decoy drive, and baseline (Table 3). Ablating any genuine drive collapses regulation toward baseline; a decoy regulates no better than baseline.
Specificity against a 1,000-decoy ensemble (Table 4). The null sits near 0.04 (95th 0.09, max 0.13); a partial-coupling control reaches 0.43; the real cause binds at 1.00, beyond the null (p < 0.001).
It transfers to a public benchmark. In Crafter, where adversarial episodes dominate, a homeostatic objective is isolated by cause and certified within its partition.
44 of 5,000 episodes (0.9%) were homeostatic. Within that partition the drink drive binds at 1.00 [0.90, 1.00] versus 0.32 in the adversarial majority; ablating it swings regulation 4.7×.
On the deployed substrate, with thresholds fixed in advance, an un-authored consequence arises and clears the full battery, with no per-run relaxation.
An un-authored consequence arises across construction paths and a second agent (50 seeds each; HomeoGrid 30). The matched no-stake control fires 0 of 50.
Then it was attacked. An evolutionary optimiser searched objective-space for spurious accepts; the result is a measured adversarial false-accept rate, reported in full in a dedicated companion paper.
Measured false-accept rate. Over ~51,200 candidates the grounding controls admit 7.2×10⁻⁴ (37 slips, all grounded-but-non-end); adding C6 gives 0 of 51,200 (≤5.9×10⁻⁵).
No control is redundant, and no attacker disagrees. Dropping any single control reopens a slip class exactly where the definitional lattice predicts it; a stronger attacker finds more grounding slips, and C6 closes them all.
Drop one control, re-run the adversary at full budget (51,200 candidates per row). Each drop admits a population precisely in the cell the lattice predicts for that control; the full conjunction admits none. A necessity proof for all six controls and a validation of the theory, in one experiment.
The rate is not an artifact of a weak optimiser. A same-budget random search finds ~9× fewer slips; a gradient attacker over a richer continuous space finds more, every one the same grounded-but-non-end class, all closed by C6 to 0 of 51,200; an exhaustive enumeration of a bounded 4,096-candidate subdomain finds zero (none exists), and an externally-authored sealed set of 200 spurious candidates is false-accepted 0 times.
A zero would be meaningless if the battery rejected everything. Specificity is paired with power: the battery also accepts what it should.
Specificity with sensitivity. In a double-blind detection study, against a sealed generator drawing archetypes the analyst never authored, the battery accepts 38 of 40 genuine objectives at a 0.05 false-positive rate; on the externally-authored sealed set it accepts 38/40 genuine candidates while false-accepting 0 of 200 spurious ones. The two misses are the longest-lag archetype, reported as a scope edge rather than tuned away. The battery is a discriminator, not a rejector.
"Held as an end" is tested, not assumed, across controller classes, and at the value-function level.
Restore-effort ρ by target type. The planner defends the fiat target as readily as the genuine objective (0.77 vs 0.79); the model-free policy restores only the genuine one (0.69 vs 0.07).
Fraction of the objective's defended value that is not survival-mediated, across a linear, a nonlinear (MLP), and a model-independent (Monte-Carlo) survival estimator. All three place ≈40% of the value on the objective itself.
A forward behavioural prediction, registered before unsealing, holds, including on agents that never run the forward model, and in a condition the protocol had never seen.
Calibration r of the frozen prediction on the planning anchor and two model-free agents that never run the forward model; both clear the pre-registered 0.65 floor.
A condition the protocol had never seen (Z0). The registered prediction was 0.42; the objective-carrying agent sacrificed 0.39 [0.31, 0.47], a matched no-objective agent 0.02.
One class is genuinely setpoint-irreducible: information-seeking. On held-out probe timing an information-gain account fits better than either a reactive or an anticipatory setpoint.
Held-out negative log-likelihood of probe timing (lower is better). An information-gain account (0.42) beats the stronger anticipatory setpoint (0.58) and a reactive setpoint (0.79); 0.63 of probes occur where a setpoint predicts none.
On sealed worlds drawn from a generator the authors never see, the pipeline certifies objectives that correspond to genuine, previously-unlisted regularities.
Ten externally-sealed worlds. The pipeline surfaced 14 candidates and certified 9; seven were genuine regularities on no pre-run list (discovery), two were rediscovery, with zero false discoveries.
What this claims, and what it doesn't.
The claim
Grounded, goal-directed self-extension, verified by perturbation and adversarial search, scoped to two classes, homeostatic regulation and one information-seeking class, with objecthood as the sixth control defines it. We do not claim consciousness, general intelligence, or that EVE "survives" an open world.
Two characterised limits of the objecthood control
Targets with no removable scaffold (abstract or social goals) are reported as grounding-only. On a planner, restore-effort certifies integration-plus-defence, necessary, not sufficient for terminal objecthood; the stronger reading is recovered on a model-free controller and at the value-function level (≈40–43% not survival-mediated), as cross-channel agreement, not a formal proof.
The false-accept rate is a bound
It is measured over the authored, non-evaluation-aware candidate space the search can represent. Temporally-compound, other-agent, and evaluation-aware spurious objectives lie outside that parameterisation, and outside the number.
What a reviewer can check unaided
Three results are reproducible from the released container without trusting any third party: the protocol and its false-accept rate, the deployed gate, and the cross-architecture forecast. The externally-authored legs corroborate but rest on disclosed, block-anchored, but not cryptographically excluded parties (effective n ≈ 7 across conditions).
No verdict is knife-edge
Every detector constant is swept rather than trusted: across the full 27-cell threshold grid the real cause binds at ≥0.93 and the decoy at ≤0.12, every pre-registered verdict holds in every cell, and the C6 scaffold set is invariant across its own grid. The verdicts behind the false-accept rate are not artifacts of tuned constants.
The protocol failed once, and reportably
Detector faithfulness is falsifiable and was once falsified: in an early rung the frozen C1 detector misread a lagged-death case. The standing reportable-defect condition, a false accept outside the predicted-miss cell or a rejection inside the all-true cell, remains open to any reader running the released container.
Skills are partly scripted
Several of the agent’s benchmark skills are scripted routines. The protocol tests whether a self-created objective is grounded, not whether every skill serving it was learned; for one need the gap is closed with a learned policy, and generalising that is future work.
Hash-locked, container-frozen, block-anchored.
Each study writes per-seed result files; predictions are fixed in hashed pre-registration documents and a master source freeze; the public-benchmark mechanics were verified line-by-line against the released environment version. A single containerised image regenerates the pass pattern of every starred study, including the deployed gate.
For that gate, the forward-prediction document, its SHA-256, and an OpenTimestamps proof ship inside the container, so checking the hash and the timestamp reproduces the pre-unseal ordering as a block-anchored fact, not a claim the reader has to take on trust.
Read it, replicate it, try to break it.
The contribution is the test, and the test outlives this instantiation. We want it re-run, attacked, and extended on other agents, with other proposers, against harder adversaries.