A Pawn Answers
Alvin's first live decision became a completed movement job. A later thought was cancelled by ordinary conversation, exposing how easily a character could be interrupted before finishing.
An answer, then an outcome
We ran the first bounded native Claude Code Max deliberation using fixed claude-sonnet-4-6 with only StructuredOutput available—no shell, file, admin, game, or MCP tools. The synthetic decision completed in 3133ms. In the real game, Alvin accepted a nearby move in 3448ms and actually reached (90,80), verified independently. His partial rationale: "I'm just wandering anyway, and that waypoint's close enough."
A thought interrupted
A second live game call triggered event-reflection, but new native Chitchat memory arrived mid-thought. The coordinator cancelled the old deliberation, applied no answer, and dispatched no second job. An implausible second proposal remained pending. The router currently treats all acquired memories as significant, so mundane repeated conversation could repeatedly interrupt a thought before it finishes.
A limit in the evidence
Full two-call acceptance did not pass. During recovery, an older deployed runner reloaded its fixture; the coordinator rejected this mismatched timeline. The operator runner now checks that its deployed entry point matches the local build before touching the game. SQLite reopening verifies the saved original action and interruption, but paired or cold restore of this specific live trial remains unproven.
Knowing when to keep thinking
All 44 automated tests pass, and staging is stopped cleanly. The three-attempt trial cap was reached with no retry and no Jev ledger changes. The concrete lesson: we must refine interruption significance to retain urgent interruption and pawn authority while filtering mundane noise, rather than assuming more calls will automatically help.
Records
Where this entry comes from. Follow these before trusting the prose.
- Live deliberation controlsgithub.com
- Mixed real-game result and saved auditgithub.com