z21 loop · cycle 1 · betting record

Foresightgated spike — cheap world-model lookahead for live agents

Bet · gated Small-batch · one session · R&D mode: expect learning, not shipping

Pitch foresight-world-model  ·  Direction Lacuna dir 29561

The bet · stated so it can be lost
Shallow simulated foresight — cheap world-model, depth ≤2, ≤4 branches — beats reactive ReAct on a real, non-sandbox GUI/web benchmark at tolerable token cost.
Start condition · the gate

The session opens with a ≤20-minute go/no-go probe

Confirm a concrete non-sandbox GUI/web benchmark + agent harness we can drive today, off the shelf. The probe decides whether any cycle is spent at all.

≤20-min probe drivable benchmark? Proceed → build the spike foresight-agent vs. reactive ReAct Stop → back to shaping "drivable benchmark" is its own pitch
Circuit breaker fires before the cycle is spent — the unknown surfaces cheaply.
The five-question record

How the table voted

Does the problem matter?Pass Myopia is the wall for long-horizon agents (BALROG, Why Reasoning Fails to Plan); the commercially hot segment.
Is the appetite right?Pass One session is honest because the rabbit-hole caps (coarse deltas, depth ≤2, branch ≤4) force it.
Is the solution attractive?Pass Falsifiable — failure is a result (bet dies cheap), not a debugging spiral.
Is this the right time?Strong pass Dir 29561 early (111 limitations / 16 findings) and accelerating; sandbox-proven, in-the-wild open. The window is now.
Is the capacity there?Flag No visible Mila author gravity — bet rests on velocity + whitespace, not studio home turf. And the reusable-benchmark assumption is unverified → that's what the gate tests.
Definition of shipped
Done = a measured yes/no on the bet, deployed as a runnable comparison — foresight-agent vs. reactive-ReAct on the chosen benchmark, reporting two numbers: task-success delta and tokens-per-task. Not "code written." A clean no that kills the bet counts as shipped.
Carried flags · for portfolio review