houserules/.sisyphus/plans/rete-rules-engine.md

3215 lines
149 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# TypeScript Rete-Based Rules Engine + Custom-Rules Chess Demo
## TL;DR
> **Quick Summary**: Build `@paratype/rete` (a Doorenbos-style Rete II rules engine in TypeScript with Immer-backed time-travel) and `@paratype/chess` (a browser chess game where every rule — including FIDE rules and 15 preset custom rules — is expressed as Rete productions) served by a `@paratype/chess-server` authoritative Bun WebSocket server for multiplayer.
>
> **Deliverables**:
> - `packages/rete`: Rete engine with alpha/beta nodes, joins, negation, NCC, existential, aggregation, derived facts, cycle detection, event log, N-tick Immer snapshots, TS builder API, JSON serialization (handler-registry pattern)
> - `packages/chess`: Browser chess game with FIDE rules expressed as Rete productions, 15 preset custom rules (toggle between games), React/Vite UI, localStorage persistence, JSON ruleset/game export-import
> - `packages/chess-server`: Bun WebSocket authoritative server with rooms, reconnection, move validation, deterministic broadcast
> - `docs/PHASES.md`, `packages/rete/SPEC.md`, `packages/chess/RULES.md`, `packages/server/PROTOCOL.md` (Phase 0 specification locks)
> - Full CI (typecheck, lint, Vitest with coverage, Playwright E2E, bundle-size, security audit)
>
> **Estimated Effort**: XL (5-phase plan, 70+ tasks)
> **Parallel Execution**: YES — heavy parallelism within waves, strict sequencing between phases
> **Critical Path**: Phase 0 specs → monorepo scaffold → Phase 1 alpha/beta network → Phase 2 advanced nodes → Phase 3 chess engine-as-rules → Phase 4 multiplayer server → Final QA
---
## Context
### Original Request
Build a TypeScript Rete-based rules engine for browser games (inspired by [paranim/pararules](https://github.com/paranim/pararules) in Nim), with immutable data structures enabling time-rewind and debugging. Use it to power a browser chess game with user-customizable rules at runtime (inspired by [chess.dougdoug.com](https://chess.dougdoug.com)).
### Interview Summary
**Key Decisions (all confirmed by user)**:
- **Repo**: Monorepo with `packages/rete` (engine), `packages/chess` (browser game), `packages/server` (WebSocket server)
- **Rule authoring**: Typed TS builder API + JSON serialization via handler-registry pattern (no `eval`, no function-to-string)
- **Fact model**: Strict EAV `(id, attr, value)` — pararules parity
- **Time-travel**: Event log + Immer snapshots every N ticks
- **Chess integration**: Chess rules ARE Rete productions (no chess.js)
- **Feature scope**: Full Doorenbos-style Rete II (alpha, beta, joins, negation, NCC, existential, aggregation, derived facts, cycle detection)
- **Tooling**: Bun workspaces + Vitest + tsc --noEmit + tsup (engine) + Vite (chess demo)
- **Testing**: TDD for engine core; tests-after + Playwright for chess demo; Playwright QA for all tasks
- **Chess v1**: 15 preset rules with toggle UI (between-games toggle only in v1)
- **Persistence**: localStorage auto-save + JSON export/import
- **Play mode**: Networked multiplayer via WebSocket
- **Networking**: Authoritative server — engine on server, clients send intents, server broadcasts events
- **Packaging**: `@paratype/rete`, `@paratype/chess`, `@paratype/chess-server`, MIT license
**Research Findings** (condensed):
- **Pararules uses strict EAV, alpha+beta networks, joins, conditions, derived facts via `thenFinally`, cycle detection via recursion limit** — no native negation/aggregation/NCC; emulated via derived facts. User chose to extend to full Rete II.
- **No production-ready TS Rete engine exists** — Nools is dead (2019), Rete.js is a node editor (not a rules engine), node-rules is not true Rete. Must build greenfield.
- **Immer is the best immutable fit** — structural sharing, draft-based mutations, minimal API overhead.
- **Chess.js hardcodes FIDE rules** — no hook system; validates the "chess-as-rules" architectural choice.
- **Runtime rule injection is THE feature** for chess.dougdoug.com-style play.
### Metis Review (Gaps Addressed)
**Metis identified 36 ambiguity points and classified combined risk as non-linear.** Key resolutions now baked into this plan:
- **"Rete II" is fixed**: Doorenbos 1995 thesis (unlinking, right/left activation, enumerated node types: alpha, beta, join, negation, NCC, existential, aggregation). No other interpretation accepted.
- **Phasing is mandatory**: 5 phases (0: specs → 1: pararules parity → 2: Rete II + local chess → 3: time-travel + presets + UI → 4: multiplayer). Phase N+1 tasks do not start until Phase N acceptance gate is green.
- **RHS serialization = handler-registry pattern**: JSON rules store `{conditions: [...], handler: "registeredName", args: [...]}`. Handlers are registered TS functions in a per-package registry. Zero `eval`, zero function-to-string, zero arbitrary JS in saved JSON.
- **Fact ID authority = server-minted** in multiplayer; client only references positions/piece-ids opaquely; deterministic across clients.
- **Conflict resolution = deterministic**: `salience desc → specificity desc → insertion order asc`. Documented in SPEC.md.
- **Match refraction = once per unique match** (CLIPS-style). Re-fires only when fact identity or bound variables change.
- **RHS purity contract**: No `Date.now()`, `Math.random()`, or I/O in rule RHS. Enforced via ESLint rule + dev-mode runtime guard that wraps globals in engine package.
- **Rule set immutability during active game** — in v1, rules cannot be added/removed mid-game; toggle only between games (simplifies time-travel + multiplayer determinism).
- **Preset rules are dev-authored TS** in v1 — no user JS upload (would be RCE vector on server).
- **Protocol versioning**: Every WebSocket message includes `v: 1`; mismatch = hard disconnect.
---
## Work Objectives
### Core Objective
Deliver a production-quality Rete-based rules engine in TypeScript (with time-travel via Immer) and a fully functional browser chess demo (networked multiplayer, 15 custom rule presets) that proves out the engine as a game-logic substrate.
### Concrete Deliverables
**Packages**:
- `packages/rete/` — Published as `@paratype/rete` (ESM + CJS + .d.ts via tsup)
- `packages/chess/` — Published as `@paratype/chess` (chess UI, Vite dev server, bundled static)
- `packages/server/` — Published as `@paratype/chess-server` (Bun HTTP+WebSocket server)
**Specifications** (Phase 0 locks):
- `packages/rete/SPEC.md` — Engine semantics (fact shape, ID authority, conflict resolution, refraction, iteration order, truth maintenance, cycle limit, RHS purity, JSON schema, named Rete II reference)
- `docs/PHASES.md` — 5-phase plan with gates, non-goals, perf budgets, demo scenarios
- `packages/chess/RULES.md` — 15 concrete preset rules with compat matrix and test scenarios
- `packages/server/PROTOCOL.md` — WebSocket message schemas, reconnection flow, rate limits
**Infrastructure**:
- `.github/workflows/ci.yml` — Typecheck, lint, test+coverage, Playwright, bundle-size, `bun audit`
- Root `bunfig.toml`, `tsconfig.base.json`, `eslint.config.js`, `vitest.workspace.ts`, `playwright.config.ts`
- Pre-commit hook enforcing `bun run check` (lefthook or simple-git-hooks)
**Acceptance Gates** (per phase): executable verification commands (see each phase's wave).
### Definition of Done
Run from repo root:
```bash
bun install # → 0 errors
bun run check # tsc --noEmit + eslint + vitest → all green
bun run test:coverage # engine ≥90% line, chess ≥70%, server ≥80%
bun run playwright test # all E2E scenarios pass
bun run build # all three packages build → dist/ populated
bun run size-limit # engine < 50KB min+gz, chess < 200KB min+gz
bun run replay-determinism # hash of replayed state == recorded hash, 100% match
bun run server & sleep 1 && bun run test:integration # real WebSocket handshake, move exchange
```
### Must Have
- Doorenbos-style Rete II: alpha network, beta network, join nodes, negation nodes, NCC nodes, existential nodes, aggregation nodes, derived-fact production with `thenFinally`-equivalent semantics, cycle detection with configurable recursion limit (default 64)
- Strict EAV fact model with typed attributes; type-safe TS builder API with autocompleted attr names
- JSON serialization of all rules via handler-registry pattern (round-trip equivalence tested per rule)
- Deterministic tick execution: documented conflict resolution (salience → specificity → insertion-order); iteration of Set/Map replaced with sorted arrays everywhere; no `Date.now`/`Math.random`/I/O in RHS
- Immer-backed working memory snapshots at configurable interval N (default 30 ticks); append-only event log with monotonic sequence numbers; replay produces byte-identical state (verified via state hash)
- Full FIDE chess rules expressed as Rete productions in `@paratype/chess`: piece placement, legal move generation per piece, turn order, captures, check detection, castling, en passant, promotion, checkmate, stalemate, 50-move rule, threefold repetition, insufficient material
- 15 concrete preset custom rules in `@paratype/chess` with compatibility matrix; toggleable between games; each with unit tests and at least one Playwright scenario
- Chess UI (React + Vite): 8×8 board with drag-drop moves, legal-move highlighting, rule-toggle screen, save/load UI, JSON export/import, undo via time-travel (to previous turn boundary)
- localStorage auto-save (per tick end) with schema-versioned payload; restore on page load; JSON export-import with validation
- Bun WebSocket server with: room create/join/leave (6-char room codes, 60s reconnect window), authoritative move validation, fact-delta broadcast, protocol versioning (`v` field), rate limit (100 msg/sec/client), 64KB message cap, origin allow-list, structured logging (pino)
- CI green on ubuntu-latest with Bun latest; bundle-size enforced; `bun audit` green
- ≥90% line coverage for `@paratype/rete`; ≥70% for `@paratype/chess`; ≥80% for `@paratype/chess-server`
- Conventional Commits; phase boundaries tagged (`v0.1.0-phase1`, etc.); pre-commit hook runs `bun run check`
### Must NOT Have (Guardrails)
**Scope exclusions (v1)**:
- NO chess AI, puzzles, tutorials, opening books, ELO, matchmaking, tournaments, leaderboards
- NO social features: chat, emotes, friends, profiles, avatars
- NO rule marketplace, remote rule sharing, user-authored JS rule upload
- NO mid-game rule toggle (toggle only between games in v1)
- NO spectators in v1 (2-player rooms only)
- NO server-side game persistence across restart (in-memory rooms only)
- NO mobile-native clients (responsive web only)
- NO accounts, OAuth, email, password, analytics, telemetry, i18n
- NO additional games on top of the engine in this plan
- NO visual rule editor / node graph editor (toggle-only UI in v1)
- NO pararules' Nim macro equivalents via runtime code-gen or `eval`
- NO external TS Rete library dependency (greenfield build)
- NO chess.js dependency (chess rules ARE Rete productions)
- NO Stockfish or other chess engines
- NO persistent user data beyond localStorage
**Code-quality exclusions**:
- NO `as any`, `as unknown as X`, `@ts-ignore`, `@ts-expect-error` in engine package (ESLint-enforced)
- NO `Date.now()`, `Math.random()`, `performance.now()`, `setTimeout`, `setInterval`, `fetch`, `console.log` inside engine RHS code paths (ESLint override on engine package)
- NO raw `Set<object>` or `Map<object, …>` iteration in engine hot paths (must sort to array first)
- NO circular package dependencies (`@paratype/chess` may import `@paratype/rete`; reverse forbidden)
- NO internal JSDoc (public API only); NO over-validation inside module boundaries
- NO premature abstraction / "framework" layer between engine and chess
- NO generic names in engine code: `data`, `result`, `item`, `temp`, `obj`, `foo`
---
## Verification Strategy (MANDATORY)
> **ZERO HUMAN INTERVENTION** — ALL verification is agent-executed. No exceptions.
### Test Decision
- **Infrastructure exists**: NO (fresh repo; infrastructure built in Phase 0 scaffold)
- **Automated tests**: YES (TDD for engine, tests-after for chess/server)
- **Framework**: Vitest (unit/integration), Playwright (E2E browser), custom Bun scripts (WebSocket integration, replay-determinism)
- **TDD workflow**: For engine tasks, each task follows RED (failing Vitest) → GREEN (minimal impl) → REFACTOR (clean up while tests remain green)
### QA Policy
Every task MUST include agent-executable QA scenarios. Evidence saved to `.sisyphus/evidence/task-{N}-{slug}.{ext}`.
- **Engine unit tests**: `bun test <path> -t "<name>"` with exact expected PASS/FAIL line; evidence = stdout log
- **Chess UI**: Playwright (playwright skill) — specific `[data-square="e2"]`, `[data-piece="white-pawn"]` selectors; evidence = screenshot + trace
- **Server integration**: scripted Bun WebSocket client against running server process; evidence = transcript JSON
- **Determinism**: `bun run scripts/hash-state.ts <log>` produces sha256; evidence = hash file
- **Bundle size**: `bun run size-limit`; evidence = stdout showing kb count
- **Build**: `bun run build` → inspect `packages/*/dist/`; evidence = `ls -la` output
- **CI**: `gh run list --limit 1 --json conclusion` → "success"; evidence = run URL
### Mandatory QA Scenario Requirements
Every task MUST have:
- At least 1 happy-path scenario with exact commands, inputs, and assertions
- At least 1 failure/edge-case scenario (invalid input, missing dep, rejected move, protocol mismatch, etc.)
- Evidence path: `.sisyphus/evidence/task-{N}-{scenario-slug}.{ext}`
- Specific selectors/data, not vague descriptions
- Binary pass/fail result (no "looks correct")
---
## Execution Strategy
### Phase Structure (Metis-directed)
5 phases, strictly sequential. Phase N+1 cannot begin until Phase N acceptance gate (see each phase's final wave) is green.
- **Phase 0** — Specification Lock (Wave P0.1 parallel spec authoring, Wave P0.2 scaffold)
- **Phase 1** — Engine Parity with Pararules (alpha/beta, joins, conditions, derived facts, cycle detection, builder API, JSON handler-registry, basic Immer state)
- **Phase 2** — Rete II Extensions + Chess Engine (negation, NCC, existential, aggregation; full FIDE chess as Rete productions; local 2-player via hot-seat for internal validation only)
- **Phase 3** — Time-Travel + Presets + UI (event log + snapshots; replay determinism; 15 preset custom rules; React UI; localStorage; JSON import/export)
- **Phase 4** — Authoritative Multiplayer Server (WebSocket server, rooms, reconnection, protocol v1; client networking layer; end-to-end multiplayer scenarios)
- **Final Wave** — 4 parallel review agents (plan compliance, code quality, manual QA, scope fidelity) → user okay → DONE
### Parallel Execution Waves
```
Phase 0 — Specification Lock
Wave P0.1 (parallel spec authoring — 4 tasks):
├── P0.1 SPEC.md (engine semantics) [deep]
├── P0.2 PHASES.md (phase gates) [writing]
├── P0.3 RULES.md (15 preset custom rules) [deep]
└── P0.4 PROTOCOL.md (WS protocol v1) [deep]
Wave P0.2 (after P0.1, sequential foundation):
├── P0.5 Monorepo scaffold (bun workspaces, tsconfig, eslint, vitest, playwright) [unspecified-high]
└── P0.6 CI pipeline + pre-commit hook [unspecified-high]
GATE: SPEC/PHASES/RULES/PROTOCOL reviewed; bun install + bun run check green; CI green
Phase 1 — Engine Pararules Parity (TDD)
Wave P1.1 (parallel engine primitives — 6 tasks):
├── P1.1 Schema + Fact type with typed attrs [deep]
├── P1.2 Working memory (WM) storage + retrieval [deep]
├── P1.3 Alpha network (fact indexing by (id,attr))[deep]
├── P1.4 Session + lifecycle (init, add, fireRules)[deep]
├── P1.5 TS builder API + handler registry [deep]
└── P1.6 JSON serialization (round-trip) [deep]
Wave P1.2 (parallel join mechanics — 4 tasks):
├── P1.7 Beta network (memory + token propagation) [deep]
├── P1.8 Join nodes with variable binding [deep]
├── P1.9 Condition filters (`cond` analog) [deep]
└── P1.10 Query API (query / queryAll) [deep]
Wave P1.3 (parallel advanced parity — 3 tasks):
├── P1.11 Derived facts (thenFinally equivalent) [deep]
├── P1.12 Cycle detection (recursion limit) [deep]
└── P1.13 Deterministic conflict resolution [deep]
Wave P1.4 (parity validation):
└── P1.14 Pararules golden-file test port [unspecified-high]
GATE: Engine v0.1.0-phase1 tag; 90% coverage; all golden tests green; `bun run check` green
Phase 2 — Rete II Extensions + Chess Engine
Wave P2.1 (parallel Rete II nodes — 4 tasks):
├── P2.1 Negation nodes (NOT) [deep]
├── P2.2 Existential nodes (EXISTS) [deep]
├── P2.3 NCC nodes (not-count-condition) [deep]
└── P2.4 Aggregation nodes (count/sum/collect/min/max) [deep]
Wave P2.2 (chess foundation — parallel 4 tasks):
├── P2.5 Chess attribute schema & piece fact shape [deep]
├── P2.6 Starting-position fact generator [quick]
├── P2.7 Square coordinate & color helpers [quick]
└── P2.8 Piece movement primitive rules (directions/steps) [deep]
Wave P2.3 (chess legal-move rules — parallel 6 tasks):
├── P2.9 Pawn move/capture rules [deep]
├── P2.10 Knight move rules [deep]
├── P2.11 Bishop/Rook/Queen sliding rules [deep]
├── P2.12 King move rules [deep]
├── P2.13 Turn order + move legality integration [deep]
└── P2.14 Capture resolution rules [deep]
Wave P2.4 (chess special rules — parallel 4 tasks):
├── P2.15 Castling (kingside + queenside with history flags) [deep]
├── P2.16 En passant (single-tick capture window) [deep]
├── P2.17 Promotion (to Q/R/B/N) [deep]
└── P2.18 Check detection rule [deep]
Wave P2.5 (chess endgames — parallel 4 tasks):
├── P2.19 Checkmate detection [deep]
├── P2.20 Stalemate detection [deep]
├── P2.21 50-move rule + threefold repetition (aggregation-based) [deep]
└── P2.22 Insufficient material draw [deep]
Wave P2.6 (integration):
└── P2.23 End-to-end FIDE game replay test [unspecified-high]
GATE: Engine v0.2.0-phase2 tag; full FIDE game playable via rules only; `bun run check` green
Phase 3 — Time-Travel + Presets + UI
Wave P3.1 (time-travel — parallel 3 tasks):
├── P3.1 Event log (append-only, monotonic seq) [deep]
├── P3.2 Immer snapshot every N ticks [deep]
└── P3.3 Replay engine + determinism hash verifier [deep]
Wave P3.2 (15 preset rules — parallel 5 tasks x 3 rules each):
├── P3.4 Presets 1-3 (pawn-focused variants) [deep]
├── P3.5 Presets 4-6 (knight/bishop variants) [deep]
├── P3.6 Presets 7-9 (rook/queen/king variants) [deep]
├── P3.7 Presets 10-12 (board/geometry variants) [deep]
└── P3.8 Presets 13-15 (meta rules: HP/heal/immune) [deep]
Wave P3.3 (UI — parallel 5 tasks):
├── P3.9 React + Vite scaffold for chess app [visual-engineering]
├── P3.10 Chessboard component (drag-drop, highlights) [visual-engineering]
├── P3.11 Rule-toggle screen (list with compat warnings) [visual-engineering]
├── P3.12 Save/Load panel + undo via time-travel [visual-engineering]
└── P3.13 JSON export/import + validation [visual-engineering]
Wave P3.4 (persistence + integration):
├── P3.14 localStorage auto-save + restore [unspecified-high]
└── P3.15 End-to-end UI scenario (play game, toggle rule, save, restore) [unspecified-high]
GATE: Engine v0.3.0-phase3 tag; chess UI fully playable locally with presets; `bun run check` green
Phase 4 — Authoritative Multiplayer
Wave P4.1 (server core — parallel 4 tasks):
├── P4.1 Bun HTTP+WS server scaffold + config [unspecified-high]
├── P4.2 Message schemas + validation [deep]
├── P4.3 Room model (create/join/leave, 6-char codes) [deep]
└── P4.4 Rate limiting + origin allow-list + 64KB cap [unspecified-high]
Wave P4.2 (server game logic — parallel 4 tasks):
├── P4.5 Authoritative session per room [deep]
├── P4.6 Move-intent validation + fact-delta broadcast [deep]
├── P4.7 Reconnection flow (60s window, snapshot resume) [deep]
└── P4.8 Structured logging (pino) + metrics [unspecified-high]
Wave P4.3 (client networking — parallel 3 tasks):
├── P4.9 WebSocket client with reconnect + seq ack [deep]
├── P4.10 Client prediction + server reconciliation [deep]
└── P4.11 Room lobby UI (create/join screens) [visual-engineering]
Wave P4.4 (integration):
└── P4.12 E2E multiplayer scenario (two Playwright contexts play a full game) [unspecified-high]
GATE: Engine v0.4.0-phase4 tag; two-browser multiplayer working end-to-end; `bun run check` green
Final Verification Wave (4 parallel reviews)
├── F1 Plan compliance audit (oracle)
├── F2 Code quality review (unspecified-high)
├── F3 Real manual QA via Playwright + scripted WS client (unspecified-high)
└── F4 Scope fidelity check (deep)
→ Present results → Wait for explicit user okay → Tag v1.0.0
```
### Dependency Matrix (abbreviated — full matrix embedded in each task's "Blocked By")
- **Phase 0 tasks**: No external deps; P0.5 blocks ALL Phase 1+ tasks; P0.6 depends on P0.5
- **P1.1-P1.6**: parallel within Wave P1.1, block P1.7-P1.10
- **P1.7-P1.10**: parallel within Wave P1.2, block P1.11-P1.13
- **P1.11-P1.13**: parallel within Wave P1.3, block P1.14
- **P1.14**: Phase 1 gate; blocks all Phase 2
- **P2.1-P2.4**: Rete II nodes, parallel, block P2.21 (aggregation-dependent)
- **P2.5-P2.8**: chess foundation, parallel, block P2.9-P2.14
- **P2.9-P2.14**: legal-move rules, parallel, block P2.15-P2.18
- **P2.15-P2.18**: special rules, parallel, block P2.19-P2.22
- **P2.19-P2.22**: endgames, parallel, block P2.23
- **P2.23**: Phase 2 gate; blocks all Phase 3
- **P3.1-P3.3**: time-travel, parallel, block P3.14 (restore requires replay)
- **P3.4-P3.8**: presets, parallel, block P3.11 (UI needs presets listed)
- **P3.9-P3.13**: UI tasks, mostly parallel (P3.10 depends on P3.9; others parallel with P3.10)
- **P3.14**: localStorage, depends on P3.3 + P3.12
- **P3.15**: Phase 3 gate; blocks all Phase 4
- **P4.1-P4.4**: server core, parallel, block P4.5-P4.8
- **P4.5-P4.8**: server game logic, parallel, block P4.9-P4.11
- **P4.9-P4.11**: client networking, parallel, block P4.12
- **P4.12**: Phase 4 gate; blocks Final Wave
- **F1-F4**: parallel; all must APPROVE before user-okay
### Agent Dispatch Summary
- **Phase 0 (6)**: P0.1-P0.4 → `deep`+`writing`; P0.5-P0.6 → `unspecified-high`
- **Phase 1 (14)**: All `deep` (TDD engine work); P1.14 → `unspecified-high`
- **Phase 2 (23)**: All `deep`; P2.23 → `unspecified-high`
- **Phase 3 (15)**: P3.1-P3.8 → `deep`; P3.9-P3.13 → `visual-engineering`; P3.14-P3.15 → `unspecified-high`
- **Phase 4 (12)**: P4.1 → `unspecified-high`; P4.2-P4.3 → `deep`; P4.4 → `unspecified-high`; P4.5-P4.7 → `deep`; P4.8 → `unspecified-high`; P4.9-P4.10 → `deep`; P4.11 → `visual-engineering`; P4.12 → `unspecified-high`
- **Final (4)**: F1 → `oracle`; F2 → `unspecified-high`; F3 → `unspecified-high`; F4 → `deep`
---
## TODOs
> Implementation + Test = ONE Task. Never separate.
> EVERY task has: Recommended Agent Profile + Parallelization info + QA Scenarios.
> **A task WITHOUT QA Scenarios is INCOMPLETE. No exceptions.**
### Phase 0 — Specification Lock
- [x] P0.1. **Author `packages/rete/SPEC.md` — engine semantics specification**
**What to do**:
- Create directory `packages/rete/`
- Write `packages/rete/SPEC.md` with sections (exactly these, `## ` headings):
1. `## Fact Model` — strict EAV (id, attr, value); id minted by Session (auto-increment), opaque to users; attr is branded string literal type; value is typed per attr via schema
2. `## ID Authority` — Session owns counter; in multiplayer, only server increments; clients receive facts with server-assigned ids
3. `## Conflict Resolution` — deterministic order: salience desc → specificity (# of conditions) desc → rule insertion order asc
4. `## Match Refraction` — each unique match fires once; re-fires only on fact change affecting bindings
5. `## Iteration Order` — all Session iteration uses sorted arrays (sort keys documented per structure); no raw `Set<object>` iteration in hot paths
6. `## Truth Maintenance` — derived facts (thenFinally) retract when any supporting fact retracts; logical dependency tracked per derived fact
7. `## Cycle Detection` — configurable recursion limit (default 64); exceeded → `RecursionLimitExceededError` with cycle trace
8. `## RHS Purity Contract` — RHS may NOT call Date.now, Math.random, performance.now, setTimeout, setInterval, fetch, or any I/O; enforced via ESLint rule `no-impure-rhs` (custom rule) + dev-mode runtime global wrapping
9. `## JSON Rule Schema` — handler-registry pattern: `{name, salience, conditions: [...], handler: "registeredName", args: JsonValue[]}`; NO function-to-string, NO eval, NO arbitrary JS
10. `## Rete II Reference Target` — Doorenbos 1995 thesis; enumerate node types in scope: AlphaNode, BetaMemory, JoinNode, NegationNode, NccNode, ExistentialNode, AggregationNode, DerivedFactProduction
**Must NOT do**:
- Do NOT include implementation code in SPEC.md
- Do NOT reference specific library versions
- Do NOT leave any section as TBD
**Recommended Agent Profile**:
- **Category**: `deep` — Requires careful semantic reasoning about Rete and distributed determinism
- **Skills**: [`context7`, `web-search`]
- `context7`: Look up canonical Rete references (Forgy 1982, Doorenbos 1995)
- `web-search`: Find CLIPS/Drools/Jess documentation for conflict resolution conventions
**Parallelization**:
- **Can Run In Parallel**: YES
- **Parallel Group**: Wave P0.1 (with P0.2, P0.3, P0.4)
- **Blocks**: P0.5, ALL engine implementation tasks
- **Blocked By**: None — start immediately
**References**:
**Pattern References**:
- Pararules README semantics: https://github.com/paranim/pararules#overview
**API/Type References**:
- Will be source of truth — no prior file
**External References**:
- Doorenbos 1995: "Production Matching for Large Learning Systems" — canonical Rete II thesis
- CLIPS reference manual — conflict resolution strategies
- Pararules source: https://github.com/paranim/pararules/blob/master/src/pararules/engine.nim
**WHY Each Reference Matters**:
- Doorenbos is the only authoritative source for "Rete II"; without it, scope ambiguity persists
- CLIPS refraction/salience is the de facto industry standard
- Pararules defines our baseline behavior to match
**Acceptance Criteria**:
- [ ] File `packages/rete/SPEC.md` exists
- [ ] `[ "$(grep -c '^## ' packages/rete/SPEC.md)" -ge "10" ]` → true (exactly 10 `## ` sections)
- [ ] `grep -q 'Doorenbos' packages/rete/SPEC.md` → 0 exit
- [ ] `grep -q 'handler-registry' packages/rete/SPEC.md` → 0 exit
- [ ] `grep -q 'no-impure-rhs' packages/rete/SPEC.md` → 0 exit
**QA Scenarios**:
```
Scenario: SPEC.md exists with required structure
Tool: Bash
Preconditions: clean repo
Steps:
1. Run: test -f packages/rete/SPEC.md
2. Run: grep -c '^## ' packages/rete/SPEC.md
3. Run: for term in "Fact Model" "ID Authority" "Conflict Resolution" "Match Refraction" "Iteration Order" "Truth Maintenance" "Cycle Detection" "RHS Purity Contract" "JSON Rule Schema" "Rete II Reference Target"; do grep -q "^## $term" packages/rete/SPEC.md || echo "MISSING: $term"; done
Expected Result: Step 1 exit 0; Step 2 outputs exactly 10; Step 3 outputs nothing (no MISSING lines)
Failure Indicators: missing file, section count != 10, any MISSING line
Evidence: .sisyphus/evidence/task-P0.1-spec-exists.log
Scenario: SPEC.md forbids eval in JSON schema section
Tool: Bash
Preconditions: SPEC.md written
Steps:
1. Run: awk '/^## JSON Rule Schema/,/^## /' packages/rete/SPEC.md | grep -qiE 'eval|function-to-string|arbitrary JS' && echo "OK" || echo "MISSING_FORBID_EVAL"
Expected Result: stdout "OK"
Evidence: .sisyphus/evidence/task-P0.1-json-forbid.log
```
**Commit**: YES
- Message: `docs(rete): author engine specification (SPEC.md)`
- Files: `packages/rete/SPEC.md`
- Pre-commit: none (doc-only commit; hook runs `bun run check` which no-ops on empty repo)
- [x] P0.2. **Author `docs/PHASES.md` — phase gates + non-goals + perf budgets**
**What to do**:
- Create `docs/PHASES.md` with sections (exact headings):
- `## Phase 0 — Specification Lock`
- `## Phase 1 — Pararules Parity`
- `## Phase 2 — Rete II + Chess Engine`
- `## Phase 3 — Time-Travel + Presets + UI`
- `## Phase 4 — Authoritative Multiplayer`
- `## Non-Goals (v1)`
- `## Performance Budgets`
- `## Demo Scenarios`
- Each phase section: bullet-listed in-scope deliverables + executable acceptance-gate commands + explicit Must-NOT-Have exclusions
- Non-Goals: copy the plan's "Must NOT Have" list
- Performance Budgets: `insert(fact)` < 0.5ms @ 10k facts; `fireRules()` < 5ms for chess ruleset; replay 1000 events < 500ms; engine bundle < 50KB min+gz; chess bundle < 200KB min+gz; server tick broadcast < 50ms p99
- Demo Scenarios: one per phase, each a scripted flow (e.g., "Phase 1 demo: run `bun test packages/rete` — all pararules golden tests pass")
**Must NOT do**:
- Do NOT duplicate SPEC.md content; link to it
- Do NOT set unrealistic budgets (these are contractual)
**Recommended Agent Profile**:
- **Category**: `writing` — Documentation authoring, prose-heavy
- **Skills**: [`web-search`]
- `web-search`: Reference typical WebSocket server perf budgets and bundle-size norms
**Parallelization**:
- **Can Run In Parallel**: YES
- **Parallel Group**: Wave P0.1 (with P0.1, P0.3, P0.4)
- **Blocks**: All subsequent task enumeration validation
- **Blocked By**: None
**References**:
**Pattern References**: None (new doc)
**External References**:
- bundlephobia.com for bundle size norms
- Pino docs for server-logging perf norms
**Acceptance Criteria**:
- [ ] File `docs/PHASES.md` exists
- [ ] `[ "$(grep -c '^## Phase' docs/PHASES.md)" -eq "5" ]` → true (Phase 0..4)
- [ ] `grep -q 'Non-Goals' docs/PHASES.md` → 0 exit
- [ ] `grep -q 'Performance Budgets' docs/PHASES.md` → 0 exit
- [ ] `grep -q '< 50KB' docs/PHASES.md` → 0 exit
**QA Scenarios**:
```
Scenario: PHASES.md has all required sections
Tool: Bash
Preconditions: none
Steps:
1. Run: test -f docs/PHASES.md
2. Run: for h in "Phase 0 — Specification Lock" "Phase 1 — Pararules Parity" "Phase 2 — Rete II + Chess Engine" "Phase 3 — Time-Travel + Presets + UI" "Phase 4 — Authoritative Multiplayer" "Non-Goals (v1)" "Performance Budgets" "Demo Scenarios"; do grep -qF "## $h" docs/PHASES.md || echo "MISSING: $h"; done
Expected Result: step 1 exit 0; step 2 outputs nothing
Evidence: .sisyphus/evidence/task-P0.2-sections.log
Scenario: Perf budgets are numeric and concrete (failure path)
Tool: Bash
Preconditions: PHASES.md written
Steps:
1. Run: awk '/^## Performance Budgets/,/^## /' docs/PHASES.md | grep -E '(TBD|TODO|FIXME)' && echo "FAIL: placeholder found" || echo "OK"
Expected Result: stdout "OK"
Evidence: .sisyphus/evidence/task-P0.2-budgets.log
```
**Commit**: YES
- Message: `docs(root): author PHASES.md with phase gates and perf budgets`
- Files: `docs/PHASES.md`
- Pre-commit: none
- [x] P0.3. **Author `packages/chess/RULES.md` — 15 concrete preset custom rules**
**What to do**:
- Create directory `packages/chess/`
- Write `packages/chess/RULES.md` listing exactly 15 preset custom rules
- Each rule has `### {rule-name}` heading, plus bullet subsections: `**ID**`, `**Description**`, `**Base Rule Affected**` (which FIDE production it modifies, or "additive"), `**Mode**` (additive | override), `**Incompatible With**` (list of other rule IDs), `**Test Scenarios**` (≥3 concrete scenarios describing input board state + expected behavior), `**Edge Cases**` (interaction with en passant, castling, promotion as relevant)
- Propose 15 concrete rules; include at least 3 from each category: movement-modifier (e.g., "Pawns may move backward"), piece-ability (e.g., "King heals +1HP when not in check"), win-condition (e.g., "Capture any piece to win"), board-geometry (e.g., "Board wraps horizontally"), meta-state (e.g., "Pieces have 3 HP; captures deal 1 damage")
**Must NOT do**:
- Do NOT leave any rule as "TBD" or "example rule"
- Do NOT allow two rules to be mutually required (circular dependency)
- Do NOT define rules requiring user-authored JS (v1 preset-only constraint)
**Recommended Agent Profile**:
- **Category**: `deep` — Game design + rule-interaction reasoning
- **Skills**: [`web-search`]
- `web-search`: Survey chess variants (Fairy chess, Pocket chess, Really Bad Chess) for rule inspiration
**Parallelization**:
- **Can Run In Parallel**: YES
- **Parallel Group**: Wave P0.1 (with P0.1, P0.2, P0.4)
- **Blocks**: P3.4-P3.8 (preset implementations)
- **Blocked By**: None
**References**:
- Fairy chess variants: https://en.wikipedia.org/wiki/Fairy_chess_piece
- Chess.dougdoug.com (concept inspiration)
**Acceptance Criteria**:
- [ ] File `packages/chess/RULES.md` exists
- [ ] `[ "$(grep -c '^### ' packages/chess/RULES.md)" -eq "15" ]` → true (exactly 15 rule headings)
- [ ] `grep -c '\*\*ID\*\*:' packages/chess/RULES.md` == 15
- [ ] `grep -c '\*\*Incompatible With\*\*:' packages/chess/RULES.md` == 15
- [ ] Rule IDs unique: `grep -oE '\*\*ID\*\*: [a-z-]+' packages/chess/RULES.md | sort -u | wc -l` == 15
**QA Scenarios**:
```
Scenario: Exactly 15 unique preset rules defined
Tool: Bash
Steps:
1. Run: test -f packages/chess/RULES.md
2. Run: grep -c '^### ' packages/chess/RULES.md
3. Run: grep -oE '\*\*ID\*\*: [a-z0-9-]+' packages/chess/RULES.md | sort -u | wc -l
Expected Result: step 1 exit 0; step 2 outputs 15; step 3 outputs 15
Evidence: .sisyphus/evidence/task-P0.3-rules-count.log
Scenario: No TBD placeholders (failure path)
Tool: Bash
Steps:
1. Run: grep -E '(TBD|TODO|FIXME|example rule|placeholder)' packages/chess/RULES.md && echo "FAIL" || echo "OK"
Expected Result: stdout "OK"
Evidence: .sisyphus/evidence/task-P0.3-no-placeholders.log
```
**Commit**: YES
- Message: `docs(chess): author RULES.md with 15 concrete preset custom rules`
- Files: `packages/chess/RULES.md`
- Pre-commit: none
- [x] P0.4. **Author `packages/server/PROTOCOL.md` — WebSocket protocol v1**
**What to do**:
- Create directory `packages/server/`
- Write `packages/server/PROTOCOL.md` defining WebSocket protocol v1
- Include `## Overview` explaining: all messages have top-level `v: 1`; all include `seq: number` (monotonic); all include `ts: number` (unix ms); mismatched `v` → hard disconnect; max message 64KB; rate limit 100 msg/sec/client; origin allow-list
- Enumerate at least 8 message types, each as `### Message: {name}` with subsections: `**Direction**` (C→S | S→C | bidir), `**Purpose**`, `**JSON Schema**` (fenced zod-like pseudo-schema or JSON example), `**Example**` (fenced json), `**Error Cases**` (listed)
- Required message types: `room.create`, `room.join`, `room.leave`, `game.move` (C→S intent), `game.state` (S→C full snapshot on join/reconnect), `game.delta` (S→C fact changes per tick), `game.end`, `error`
- Include `## Reconnection Flow` — client disconnects, 60s window, reconnect with last seen `seq`, server replays deltas since that seq
- Include `## Auth` — room code 6 chars [A-Z0-9]; optional room token (UUID v4) returned on create; every subsequent message includes token
- Include `## Rate Limiting` — token bucket per connection, 100 msg/sec, burst 20; over-limit → disconnect with `error` code `RATE_LIMIT`
**Must NOT do**:
- Do NOT define message types requiring session persistence across server restart (v1 in-memory only)
- Do NOT define spectator-related messages (v1 2-player only)
- Do NOT define rule-mutation-during-game messages (v1 between-games only)
**Recommended Agent Profile**:
- **Category**: `deep` — Protocol design requires precision and failure-mode reasoning
- **Skills**: [`web-search`, `code-search`]
- `web-search`: Look at lichess/chess.com WebSocket patterns
- `code-search`: Find battle-tested WebSocket protocols (e.g., y-websocket, automerge)
**Parallelization**:
- **Can Run In Parallel**: YES
- **Parallel Group**: Wave P0.1 (with P0.1, P0.2, P0.3)
- **Blocks**: Phase 4 (all server tasks)
- **Blocked By**: None
**References**:
- y-websocket protocol docs (well-designed minimal WS protocol)
- RFC 6455 (WebSocket) for base protocol
**Acceptance Criteria**:
- [ ] File `packages/server/PROTOCOL.md` exists
- [ ] `[ "$(grep -c '^### Message: ' packages/server/PROTOCOL.md)" -ge "8" ]` → true
- [ ] `grep -q 'Reconnection Flow' packages/server/PROTOCOL.md` → 0 exit
- [ ] `grep -q 'Rate Limiting' packages/server/PROTOCOL.md` → 0 exit
- [ ] `grep -q 'v: 1' packages/server/PROTOCOL.md` → 0 exit
**QA Scenarios**:
```
Scenario: Protocol defines all 8+ required message types
Tool: Bash
Steps:
1. Run: test -f packages/server/PROTOCOL.md
2. Run: grep -c '^### Message: ' packages/server/PROTOCOL.md
3. Run: for m in "room.create" "room.join" "room.leave" "game.move" "game.state" "game.delta" "game.end" "error"; do grep -qF "### Message: $m" packages/server/PROTOCOL.md || echo "MISSING: $m"; done
Expected Result: step 1 exit 0; step 2 ≥ 8; step 3 outputs nothing
Evidence: .sisyphus/evidence/task-P0.4-messages.log
Scenario: Rate-limit and auth sections present (failure path)
Tool: Bash
Steps:
1. Run: grep -qc 'Rate Limiting' packages/server/PROTOCOL.md && grep -qc 'Auth' packages/server/PROTOCOL.md && echo "OK" || echo "FAIL"
Expected Result: stdout "OK"
Evidence: .sisyphus/evidence/task-P0.4-sections.log
```
**Commit**: YES
- Message: `docs(server): author PROTOCOL.md defining WebSocket protocol v1`
- Files: `packages/server/PROTOCOL.md`
- Pre-commit: none
- [x] P0.5. **Scaffold monorepo skeleton (Bun workspaces + tsconfig + eslint + vitest + playwright)**
**What to do**:
- Root: `package.json` with `"workspaces": ["packages/*"]`, `"private": true`, `"packageManager": "bun@latest"`
- Root scripts: `check` (runs typecheck + lint + test), `typecheck` (bun x tsc -b), `lint` (bun x eslint), `test` (bun x vitest run), `test:coverage` (vitest run --coverage), `build` (bun run --filter '*' build), `size-limit` (placeholder)
- Root `tsconfig.base.json`: target ES2022, module ESNext, moduleResolution Bundler, strict: true, noImplicitAny, exactOptionalPropertyTypes, noUncheckedIndexedAccess, verbatimModuleSyntax
- Root `tsconfig.json`: references to all packages
- Root `eslint.config.js` (flat): typescript-eslint strict preset; `no-restricted-globals` ban `Date`, `Math.random`, `performance`, `setTimeout`, `setInterval`, `fetch` within `packages/rete/src/**/rhs/**` and engine RHS paths (override-based); `@typescript-eslint/no-explicit-any` error
- Root `vitest.workspace.ts` listing all packages
- Root `playwright.config.ts` with chess app base URL (http://localhost:5173)
- Packages: `packages/rete/package.json` (`"name": "@paratype/rete"`, type module, main dist/index.js, types dist/index.d.ts), `tsconfig.json` extending base, empty `src/index.ts` with `export {}`, `README.md` (one-line description)
- Same skeleton for `packages/chess` (`"name": "@paratype/chess"`) and `packages/server` (`"name": "@paratype/chess-server"`)
- Add `.gitignore`: `node_modules/`, `dist/`, `.sisyphus/evidence/`, `*.log`, `.DS_Store`, `coverage/`, `playwright-report/`, `test-results/`
- Add `LICENSE` (MIT) with paratype org name
- Add root `README.md`: project overview, link to SPEC/PHASES/RULES/PROTOCOL
**Must NOT do**:
- Do NOT install production dependencies beyond what's needed for scaffolding (TypeScript, Vitest, ESLint, Playwright, tsup)
- Do NOT add Immer/React/Vite yet (Phase 3 concern)
- Do NOT add WebSocket / pino yet (Phase 4 concern)
- Do NOT write any engine/chess/server source code beyond `export {}`
**Recommended Agent Profile**:
- **Category**: `unspecified-high` — Tooling setup with many moving parts
- **Skills**: [`context7`]
- `context7`: Look up Bun workspace, Vitest workspace, Playwright, ESLint flat config docs
**Parallelization**:
- **Can Run In Parallel**: NO (sole foundation task)
- **Parallel Group**: Wave P0.2 (sequential)
- **Blocks**: P0.6 and ALL implementation tasks
- **Blocked By**: P0.1, P0.2 (need SPEC to know package boundaries)
**References**:
**Pattern References**: None (greenfield)
**External References**:
- Bun workspaces: https://bun.sh/docs/install/workspaces
- Vitest workspace: https://vitest.dev/guide/workspace
- Playwright config: https://playwright.dev/docs/test-configuration
- typescript-eslint flat config: https://typescript-eslint.io/packages/typescript-eslint/#flat-config
**Acceptance Criteria**:
- [ ] `bun install` exits 0
- [ ] `bun run check` exits 0 (zero tests OK; zero lint errors)
- [ ] `bun run build` exits 0 (emits dist/ for each package OR exits 0 with skip — depends on tsup wiring; at minimum `tsc -b` passes)
- [ ] Files exist: `package.json`, `tsconfig.base.json`, `tsconfig.json`, `eslint.config.js`, `vitest.workspace.ts`, `playwright.config.ts`, `.gitignore`, `LICENSE`, `README.md`
- [ ] Directory tree: `packages/rete/{package.json,tsconfig.json,src/index.ts,README.md}`, same for `chess` and `server`
**QA Scenarios**:
```
Scenario: Fresh clone installs and checks clean
Tool: Bash
Preconditions: repo on fresh checkout; Bun installed
Steps:
1. Run: bun install 2>&1 | tee /tmp/p05-install.log
2. Run: bun run check 2>&1 | tee /tmp/p05-check.log
3. Run: bun run build 2>&1 | tee /tmp/p05-build.log
Expected Result: step 1 exits 0; step 2 exits 0; step 3 exits 0; no errors in logs
Failure Indicators: any non-zero exit, "error" token in logs
Evidence: .sisyphus/evidence/task-P0.5-install-check-build.log
Scenario: ESLint rejects Math.random in engine RHS path (failure path validating config correctness)
Tool: Bash
Preconditions: scaffold complete
Steps:
1. Create temp file: mkdir -p packages/rete/src/rhs && printf 'export const x = () => Math.random();\n' > packages/rete/src/rhs/_temp.ts
2. Run: bun run lint 2>&1 | tee /tmp/p05-lint-fail.log
3. Capture exit: echo "exit=$?"
4. Cleanup: rm packages/rete/src/rhs/_temp.ts
Expected Result: step 2 outputs ESLint error referencing Math.random and exits non-zero
Evidence: .sisyphus/evidence/task-P0.5-lint-rejects-random.log
Scenario: Workspace package names are correct
Tool: Bash
Steps:
1. Run: jq -r .name packages/rete/package.json
2. Run: jq -r .name packages/chess/package.json
3. Run: jq -r .name packages/server/package.json
Expected Result: outputs "@paratype/rete", "@paratype/chess", "@paratype/chess-server" respectively
Evidence: .sisyphus/evidence/task-P0.5-pkg-names.log
```
**Commit**: YES
- Message: `chore(root): scaffold monorepo with Bun workspaces, TypeScript, Vitest, ESLint, Playwright`
- Files: `package.json`, `tsconfig.base.json`, `tsconfig.json`, `eslint.config.js`, `vitest.workspace.ts`, `playwright.config.ts`, `.gitignore`, `LICENSE`, `README.md`, `packages/*/package.json`, `packages/*/tsconfig.json`, `packages/*/src/index.ts`, `packages/*/README.md`, `bun.lockb`
- Pre-commit: `bun run check` (hook installed next task)
- [x] P0.6. **CI pipeline (`.github/workflows/ci.yml`) + pre-commit hook (lefthook)**
**What to do**:
- Create `.github/workflows/ci.yml`:
- Trigger: pull_request, push to main
- Jobs: `check` (typecheck, lint, test with coverage upload), `build` (build all packages, upload dist artifacts), `e2e` (Playwright headless), `size` (bundle size check), `audit` (`bun audit`)
- All on ubuntu-latest with `oven-sh/setup-bun@v1` pinning to stable
- Cache: `~/.bun/install/cache`
- Upload Playwright traces on failure
- Create `lefthook.yml` at root with pre-commit hook running `bun run check` (fast — typecheck + lint + unit tests only, not Playwright)
- Install lefthook as dev dep; add `postinstall` script running `bunx lefthook install`
- Add `.github/workflows/README.md` explaining CI status badges
- Add size-limit config to root `package.json` (size-limit dev dep; initial budget: engine 50KB, chess 200KB — both placeholders until dist exists; the CI job passes when empty)
**Must NOT do**:
- Do NOT add Node.js matrix (Bun only, per decision)
- Do NOT add deployment workflows (out of scope)
- Do NOT skip `bun audit` (security requirement)
**Recommended Agent Profile**:
- **Category**: `unspecified-high`
- **Skills**: [`context7`, `code-search`]
- `context7`: Look up current `oven-sh/setup-bun` action options
- `code-search`: Find production CI workflows for Bun monorepos on grep.app
**Parallelization**:
- **Can Run In Parallel**: NO
- **Parallel Group**: Wave P0.2 (after P0.5)
- **Blocks**: All subsequent commits (CI becomes a required status check)
- **Blocked By**: P0.5
**References**:
- setup-bun action: https://github.com/oven-sh/setup-bun
- lefthook: https://github.com/evilmartians/lefthook
- size-limit: https://github.com/ai/size-limit
**Acceptance Criteria**:
- [ ] `.github/workflows/ci.yml` exists and passes `actionlint` (`bun x @action-validator/cli action-validator .github/workflows/ci.yml` OR `gh workflow view` after push)
- [ ] `lefthook.yml` exists at root
- [ ] `bun run check` is wired as pre-commit (running `bunx lefthook run pre-commit` executes check)
- [ ] First push triggers CI; all jobs green
- [ ] `gh run list --limit 1 --json conclusion -q '.[0].conclusion'` returns `"success"`
**QA Scenarios**:
```
Scenario: CI green on first push
Tool: Bash
Preconditions: remote configured; push enabled
Steps:
1. Run: git add -A && git commit -m "ci: verify pipeline" --allow-empty
2. Run: git push
3. Wait: sleep 120 (or poll with gh run watch)
4. Run: gh run list --limit 1 --json conclusion,databaseId,url -q '.[0]'
Expected Result: stdout contains `"conclusion":"success"` and a URL
Evidence: .sisyphus/evidence/task-P0.6-ci-success.json
Scenario: Pre-commit hook blocks bad commit (failure path)
Tool: Bash
Preconditions: hook installed
Steps:
1. Run: echo 'const x: any = 1;' > packages/rete/src/_bad.ts
2. Run: git add packages/rete/src/_bad.ts
3. Run: git commit -m "bad" 2>&1 | tee /tmp/p06-hook.log; echo "exit=$?"
4. Cleanup: git reset HEAD && rm packages/rete/src/_bad.ts
Expected Result: commit fails; log shows ESLint "no-explicit-any" error
Evidence: .sisyphus/evidence/task-P0.6-hook-blocks.log
Scenario: actionlint accepts workflow
Tool: Bash
Steps:
1. Run: bun x @action-validator/cli action-validator .github/workflows/ci.yml
Expected Result: exit 0
Evidence: .sisyphus/evidence/task-P0.6-actionlint.log
```
**Commit**: YES
- Message: `ci(root): add GitHub Actions pipeline and lefthook pre-commit hook`
- Files: `.github/workflows/ci.yml`, `.github/workflows/README.md`, `lefthook.yml`, `package.json` (size-limit config + lefthook dep), `bun.lockb`
- Pre-commit: `bun run check`
### Phase 1 — Engine Pararules Parity (TDD)
- [x] P1.1. **Schema + Fact type with typed attributes (TDD)**
**What to do**:
- RED: In `packages/rete/src/schema.test.ts`, write failing tests:
- `defineSchema({ Health: 'number', Position: 'Vec2' })` returns object with keyed attrs typed correctly
- Attempting to create a `Fact` with wrong value type for an attr produces a TypeScript type error (type-level test via `@ts-expect-error` comments in a `.type-test.ts` file)
- Runtime fact creation: `fact(id, attr, value)` returns `{ id, attr, value }` with branded types
- GREEN: Implement in `packages/rete/src/schema.ts`:
- `export function defineSchema<S extends Record<string, unknown>>(defs: S)` returning typed schema object
- `export type Fact<S>` as tagged union discriminated by `attr` key
- `export function fact<S, K extends keyof S>(id: EntityId, attr: K, value: S[K]): Fact<S>`
- `EntityId` as branded `number` via `type EntityId = number & { readonly __brand: 'EntityId' }`
- REFACTOR: Extract type utilities to `schema.types.ts` if file exceeds 150 LOC; add JSDoc on public exports only
- Export from `packages/rete/src/index.ts`
**Must NOT do**:
- Do NOT use `any` or `unknown as X` casts
- Do NOT expose Immer (Phase 3 concern)
- Do NOT allow runtime attr name collisions silently — error-throw on duplicate
**Recommended Agent Profile**:
- **Category**: `deep`
- **Skills**: [`context7`]
- `context7`: Look up TypeScript branded types and discriminated unions best practices
**Parallelization**:
- **Can Run In Parallel**: YES
- **Parallel Group**: Wave P1.1 (with P1.2, P1.3, P1.4, P1.5, P1.6)
- **Blocks**: P1.7-P1.13 (beta network, derived facts all depend on Fact type)
- **Blocked By**: P0.5, P0.6 (scaffold + CI); P0.1 (SPEC.md defines fact shape)
**References**:
**Pattern References**:
- `packages/rete/SPEC.md` §Fact Model — canonical fact shape
**External References**:
- Branded types: https://egghead.io/blog/using-branded-types-in-typescript
- Discriminated unions: https://www.typescriptlang.org/docs/handbook/2/narrowing.html#discriminated-unions
**Acceptance Criteria**:
- [ ] `bun test packages/rete/src/schema.test.ts` → all green
- [ ] `bun x tsc --noEmit -p packages/rete/tsconfig.json` → 0 errors
- [ ] Type-level tests in `schema.type-test.ts` compile (failures are intentional via `@ts-expect-error`)
- [ ] Coverage of `schema.ts` ≥ 95% line
**QA Scenarios**:
```
Scenario: Schema + fact round-trip with correct types
Tool: Bash
Steps:
1. Run: bun test packages/rete/src/schema.test.ts 2>&1 | tee /tmp/p11-test.log
2. Run: grep -E '(PASS|FAIL|Tests )' /tmp/p11-test.log
Expected Result: output contains "PASS" and final line "Tests {N} passed" with 0 failures
Evidence: .sisyphus/evidence/task-P1.1-schema-tests.log
Scenario: Type-level rejection of invalid value (failure path)
Tool: Bash
Steps:
1. Run: bun x tsc --noEmit -p packages/rete/tsconfig.json 2>&1 | tee /tmp/p11-tsc.log
2. Run: grep -c 'error TS' /tmp/p11-tsc.log
Expected Result: step 1 exits 0; step 2 outputs 0 (all @ts-expect-error annotations consumed cleanly)
Evidence: .sisyphus/evidence/task-P1.1-tsc.log
```
**Commit**: YES
- Message: `feat(rete): add schema and typed Fact primitives (P1.1)`
- Files: `packages/rete/src/schema.ts`, `packages/rete/src/schema.types.ts`, `packages/rete/src/schema.test.ts`, `packages/rete/src/schema.type-test.ts`, `packages/rete/src/index.ts`
- Pre-commit: `bun run check`
- [x] P1.2. **Working-memory (WM) storage + retrieval (TDD)**
**What to do**:
- RED: `packages/rete/src/wm.test.ts` — failing tests:
- `WM.insert(id, attr, value)` stores fact; duplicate `(id, attr)` replaces value (update semantics per SPEC)
- `WM.retract(id, attr)` removes fact; returns true if existed, false if not
- `WM.contains(id, attr)` returns boolean
- `WM.get(id, attr)` returns value or undefined
- `WM.allFacts()` returns sorted stable array (sort key: `[id, attr]`) — iteration determinism per SPEC §Iteration Order
- GREEN: `packages/rete/src/wm.ts` — `class WorkingMemory<S>` using `Map<EntityId, Map<AttrKey, FactValue>>`; `allFacts()` flattens and sorts
- REFACTOR: Add internal change-subscription hook (array of listener callbacks) called on every insert/retract — used later by alpha network. Document the subscription API in JSDoc.
**Must NOT do**:
- Do NOT emit events during iteration (mutation-during-iteration = undefined behavior)
- Do NOT expose raw Map objects (encapsulation)
**Recommended Agent Profile**:
- **Category**: `deep`
- **Skills**: []
**Parallelization**:
- **Can Run In Parallel**: YES
- **Parallel Group**: Wave P1.1 (with P1.1, P1.3-P1.6)
- **Blocks**: P1.3 (alpha consumes WM events), P1.7 (beta), P1.10 (query)
- **Blocked By**: P0.5, P0.6, P0.1
**References**:
- `packages/rete/SPEC.md` §Fact Model, §Iteration Order
**Acceptance Criteria**:
- [ ] `bun test packages/rete/src/wm.test.ts` all green
- [ ] Coverage ≥ 95%
- [ ] No raw Map/Set exposed in public API (`grep -E 'export (const|function|class).*(Map|Set)' packages/rete/src/wm.ts` empty)
**QA Scenarios**:
```
Scenario: WM insert/get/retract/contains semantics
Tool: Bash
Steps:
1. Run: bun test packages/rete/src/wm.test.ts 2>&1 | tee /tmp/p12.log
Expected Result: "Tests {N} passed, 0 failed"
Evidence: .sisyphus/evidence/task-P1.2-wm.log
Scenario: allFacts() returns deterministic order (failure path for non-determinism)
Tool: Bash
Steps:
1. Run: bun test packages/rete/src/wm.test.ts -t "allFacts deterministic order" 2>&1 | tee /tmp/p12-order.log
Expected Result: test named "allFacts deterministic order" passes; verifies same order across multiple invocations with Map insertion-order permutation
Evidence: .sisyphus/evidence/task-P1.2-wm-order.log
```
**Commit**: YES
- Message: `feat(rete): add WorkingMemory with deterministic iteration (P1.2)`
- Files: `packages/rete/src/wm.ts`, `packages/rete/src/wm.test.ts`, `packages/rete/src/index.ts`
- Pre-commit: `bun run check`
- [x] P1.3. **Alpha network: fact indexing by (id, attr) pattern (TDD)**
**What to do**:
- RED: `packages/rete/src/alpha.test.ts` — failing tests:
- AlphaNode matches facts by optional id-wildcard + required attr key; stores matched facts in AlphaMemory
- Inserting a fact dispatches it to all matching AlphaNodes
- Retracting a fact removes it from AlphaMemories
- A condition like `(Player, X, ?x)` creates one alpha node indexed by `(attr=X, id=Player)`; `(?id, X, ?x)` indexed by `(attr=X)`
- GREEN: `packages/rete/src/alpha.ts`:
- `class AlphaNetwork` subscribes to `WorkingMemory` events
- `class AlphaNode` with `condition: { id?: EntityId, attr: AttrKey }`
- `class AlphaMemory` holds `Fact[]` sorted by (id, attr)
- `AlphaNetwork.buildNode(cond)` — memoized: same condition → same node (sharing)
- Emits change events (`activate(fact)`, `deactivate(fact)`) to downstream (beta) subscribers
- REFACTOR: Extract indexing (attr → AlphaNode[]) as inverted index; ensure O(1) dispatch per fact
**Must NOT do**:
- Do NOT scan all alpha nodes per fact (must use index)
- Do NOT retain references to retracted facts
**Recommended Agent Profile**:
- **Category**: `deep`
- **Skills**: [`context7`]
- `context7`: Rete alpha network implementation patterns
**Parallelization**:
- **Can Run In Parallel**: YES
- **Parallel Group**: Wave P1.1
- **Blocks**: P1.7 (beta network), P1.8 (joins)
- **Blocked By**: P0.5, P0.6, P0.1, P1.2 (WM events)
**References**:
- `packages/rete/SPEC.md` §Fact Model
- Doorenbos thesis §2.2 (Alpha Network)
**Acceptance Criteria**:
- [ ] `bun test packages/rete/src/alpha.test.ts` all green
- [ ] Coverage ≥ 90%
- [ ] Dispatch is O(1) per fact: benchmark test asserting 10k inserts in <50ms
**QA Scenarios**:
```
Scenario: Alpha network dispatches to matching nodes only
Tool: Bash
Steps:
1. Run: bun test packages/rete/src/alpha.test.ts 2>&1 | tee /tmp/p13.log
Expected Result: "Tests {N} passed, 0 failed"
Evidence: .sisyphus/evidence/task-P1.3-alpha.log
Scenario: Alpha dispatch performance (failure path if slow)
Tool: Bash
Steps:
1. Run: bun test packages/rete/src/alpha.test.ts -t "dispatch 10000 facts in under 50ms" 2>&1 | tee /tmp/p13-perf.log
Expected Result: test passes; log includes timing assertion under 50ms
Evidence: .sisyphus/evidence/task-P1.3-alpha-perf.log
```
**Commit**: YES
- Message: `feat(rete): add AlphaNetwork with inverted-index dispatch (P1.3)`
- Files: `packages/rete/src/alpha.ts`, `packages/rete/src/alpha.test.ts`, `packages/rete/src/index.ts`
- Pre-commit: `bun run check`
- [x] P1.4. **Session lifecycle: init, add rule, fire (TDD)**
**What to do**:
- RED: `packages/rete/src/session.test.ts` — failing tests:
- `const session = new Session(schema, { autoFire: false })` creates session
- `session.add(rule)` registers a rule (rule definition opaque for now; covered by P1.5)
- `session.insert(id, attr, value)` / `session.retract(id, attr)` delegate to WM
- `session.fireRules()` returns number of rules that fired
- With `autoFire: true`, insert/retract auto-calls fireRules
- `session.fireRules({ recursionLimit: 64 })` — cycle detection (covered by P1.12, stub throws)
- GREEN: `packages/rete/src/session.ts`:
- `class Session<S>` holding `WorkingMemory<S>`, `AlphaNetwork`, `ProductionNode[]`, config `{ autoFire, recursionLimit }`
- Public API: `add(prod)`, `insert`, `retract`, `fireRules`, `contains`, `get`, `allFacts`
- Fire: iterate pending activations in deterministic order (per SPEC conflict resolution), call RHS, repeat until fixed-point or recursion limit
**Must NOT do**:
- Do NOT leak internal AlphaNetwork / beta / production types to public API
- Do NOT implement conflict resolution yet (P1.13) — stub with insertion-order
**Recommended Agent Profile**:
- **Category**: `deep`
- **Skills**: []
**Parallelization**:
- **Can Run In Parallel**: YES
- **Parallel Group**: Wave P1.1
- **Blocks**: P1.7-P1.14 (all downstream engine tasks need Session)
- **Blocked By**: P1.2, P1.3 (WM + Alpha ready)
**References**:
- `packages/rete/SPEC.md` §Conflict Resolution (stub per insertion-order), §RHS Purity Contract
**Acceptance Criteria**:
- [ ] `bun test packages/rete/src/session.test.ts` all green
- [ ] Public API surface locked via `type` export; `tsd` or `expect-type` verifies no `any` leaks
- [ ] Coverage ≥ 90%
**QA Scenarios**:
```
Scenario: Session lifecycle (insert, fire, retract)
Tool: Bash
Steps:
1. Run: bun test packages/rete/src/session.test.ts 2>&1 | tee /tmp/p14.log
Expected Result: "Tests {N} passed, 0 failed"
Evidence: .sisyphus/evidence/task-P1.4-session.log
Scenario: autoFire flag controls behavior (failure path)
Tool: Bash
Steps:
1. Run: bun test packages/rete/src/session.test.ts -t "autoFire=false does not fire on insert" 2>&1 | tee /tmp/p14-af.log
Expected Result: named test passes
Evidence: .sisyphus/evidence/task-P1.4-session-autofire.log
```
**Commit**: YES
- Message: `feat(rete): add Session lifecycle (P1.4)`
- Files: `packages/rete/src/session.ts`, `packages/rete/src/session.test.ts`, `packages/rete/src/index.ts`
- Pre-commit: `bun run check`
- [x] P1.5. **Typed TS builder API + handler registry (TDD)**
**What to do**:
- RED: `packages/rete/src/builder.test.ts` — failing tests:
- `rule('name').what((Player, X, v('x'))).what((Player, Y, v('y'))).then('moveHandler', ['x', 'y'])` produces a `RuleDefinition` object with conditions and `handler` ref
- `HandlerRegistry.register('moveHandler', (session, match) => { ... })` stores the function
- Attempting to build a rule referencing an unregistered handler throws (at build time, not fire time)
- Variable bindings use `v('name')` helper; unbound variables cause type error
- GREEN: `packages/rete/src/builder.ts` — fluent builder returning `RuleDefinition`
- GREEN: `packages/rete/src/registry.ts` — `HandlerRegistry` (Map-backed, with `register`, `get`, `has`, `verify`)
- Session.add validates all referenced handlers exist via `registry.verify(rule)`
**Must NOT do**:
- Do NOT allow function references directly in conditions (must be via registry name) — this enforces JSON serializability from day 1
- Do NOT use `eval` or `new Function`
**Recommended Agent Profile**:
- **Category**: `deep`
- **Skills**: [`context7`]
- `context7`: TypeScript builder-pattern type inference
**Parallelization**:
- **Can Run In Parallel**: YES
- **Parallel Group**: Wave P1.1
- **Blocks**: P1.6 (JSON serialization needs builder output), P1.7+ (all rule tests use builder)
- **Blocked By**: P1.1 (schema types)
**References**:
- `packages/rete/SPEC.md` §JSON Rule Schema (handler-registry pattern)
**Acceptance Criteria**:
- [ ] `bun test packages/rete/src/builder.test.ts` green
- [ ] `bun test packages/rete/src/registry.test.ts` green
- [ ] Coverage ≥ 90%
- [ ] `grep -r "new Function\|eval(" packages/rete/src` → empty
**QA Scenarios**:
```
Scenario: Builder produces serializable rule definitions
Tool: Bash
Steps:
1. Run: bun test packages/rete/src/builder.test.ts 2>&1 | tee /tmp/p15.log
Expected Result: all pass
Evidence: .sisyphus/evidence/task-P1.5-builder.log
Scenario: No eval/Function anywhere in engine src (failure path)
Tool: Bash
Steps:
1. Run: grep -rE "new Function|eval\(" packages/rete/src 2>&1 | tee /tmp/p15-grep.log; echo "exit=$?"
Expected Result: grep exits 1 (no matches); log empty
Evidence: .sisyphus/evidence/task-P1.5-no-eval.log
```
**Commit**: YES
- Message: `feat(rete): add typed rule builder + handler registry (P1.5)`
- Files: `packages/rete/src/builder.ts`, `packages/rete/src/registry.ts`, `packages/rete/src/builder.test.ts`, `packages/rete/src/registry.test.ts`, `packages/rete/src/index.ts`
- Pre-commit: `bun run check`
- [x] P1.6. **JSON serialization round-trip (TDD)**
**What to do**:
- RED: `packages/rete/src/serialize.test.ts` — failing tests:
- `serialize(rule)` produces JSON conforming to SPEC §JSON Rule Schema
- `deserialize(json, registry)` produces a `RuleDefinition` equivalent (deep-equal after normalization)
- Round-trip: `deserialize(serialize(rule)) ≡ rule` for every shape (conditions, variables, handler refs, salience)
- Deserialization with unknown handler throws `UnknownHandlerError`
- Schema validation (zod or hand-rolled) rejects malformed JSON
- GREEN: `packages/rete/src/serialize.ts` with `serialize`, `deserialize`, exported JSON schema (as `RULE_SCHEMA_V1` constant)
**Must NOT do**:
- Do NOT support "v0" or back-compat (there is no prior version)
- Do NOT serialize runtime function references
**Recommended Agent Profile**:
- **Category**: `deep`
- **Skills**: []
**Parallelization**:
- **Can Run In Parallel**: YES
- **Parallel Group**: Wave P1.1
- **Blocks**: P3.13 (JSON import/export UI), P4.2 (server protocol)
- **Blocked By**: P1.5 (builder types)
**References**:
- `packages/rete/SPEC.md` §JSON Rule Schema
**Acceptance Criteria**:
- [ ] `bun test packages/rete/src/serialize.test.ts` green
- [ ] Round-trip test covers ≥10 distinct rule shapes
- [ ] Coverage ≥ 95%
**QA Scenarios**:
```
Scenario: Round-trip 10 distinct rule shapes
Tool: Bash
Steps:
1. Run: bun test packages/rete/src/serialize.test.ts 2>&1 | tee /tmp/p16.log
2. Run: grep -c "round-trip shape" /tmp/p16.log
Expected Result: all tests pass; step 2 outputs ≥ 10
Evidence: .sisyphus/evidence/task-P1.6-roundtrip.log
Scenario: Malformed JSON rejected (failure path)
Tool: Bash
Steps:
1. Run: bun test packages/rete/src/serialize.test.ts -t "malformed JSON throws" 2>&1 | tee /tmp/p16-bad.log
Expected Result: named test passes
Evidence: .sisyphus/evidence/task-P1.6-bad-json.log
```
**Commit**: YES
- Message: `feat(rete): add JSON serialize/deserialize round-trip (P1.6)`
- Files: `packages/rete/src/serialize.ts`, `packages/rete/src/serialize.test.ts`, `packages/rete/src/index.ts`
- Pre-commit: `bun run check`
- [x] P1.7. **Beta network: memory + token propagation (TDD)**
**What to do**:
- RED: `packages/rete/src/beta.test.ts` — failing tests covering single-condition rule (beta reduces to alpha), two-condition rule (one join), three-condition chain
- GREEN: `packages/rete/src/beta.ts` — `BetaMemory`, `Token` (parent + fact chain), activation/deactivation propagation; each production node accumulates full matches
**Must NOT do**:
- Do NOT allocate new Tokens on every fact change if shared chains unchanged (reuse via parent reference)
**Recommended Agent Profile**:
- **Category**: `deep`
- **Skills**: [`context7`]
**Parallelization**:
- **Can Run In Parallel**: YES (Wave P1.2 with P1.8, P1.9, P1.10)
- **Blocks**: P1.11-P1.14, P2.*
- **Blocked By**: P1.3 (alpha), P1.4 (session)
**References**: `packages/rete/SPEC.md` §Iteration Order; Doorenbos §2.4
**Acceptance Criteria**:
- [ ] `bun test packages/rete/src/beta.test.ts` green
- [ ] Coverage ≥ 90%
**QA Scenarios**:
```
Scenario: Multi-condition rule produces join matches
Tool: Bash
Steps:
1. Run: bun test packages/rete/src/beta.test.ts 2>&1 | tee /tmp/p17.log
Expected Result: all pass
Evidence: .sisyphus/evidence/task-P1.7-beta.log
Scenario: Retraction removes join matches (failure path)
Tool: Bash
Steps:
1. Run: bun test packages/rete/src/beta.test.ts -t "retraction removes dependent tokens" 2>&1 | tee /tmp/p17-ret.log
Expected Result: pass
Evidence: .sisyphus/evidence/task-P1.7-retract.log
```
**Commit**: YES
- Message: `feat(rete): add BetaMemory + Token propagation (P1.7)`
- Files: `packages/rete/src/beta.ts`, `packages/rete/src/beta.test.ts`, `packages/rete/src/index.ts`
- Pre-commit: `bun run check`
- [x] P1.8. **Join nodes with variable binding (TDD)**
**What to do**:
- RED: `packages/rete/src/join.test.ts` — join on shared variable `?id` (e.g., `(?id, X, ?x)` ∧ `(?id, Y, ?y)` must match when id is the same), numeric equality tests
- GREEN: `packages/rete/src/join.ts` — JoinNode with tests[] (equality constraints between left token's binding and right fact's field)
- Handle many-to-many, many-to-one, and cross-product cases
**Must NOT do**:
- Do NOT implement inequality tests yet (those go in P1.9 conditions)
**Recommended Agent Profile**:
- **Category**: `deep`
**Parallelization**: YES — Wave P1.2
- **Blocks**: P1.11-P1.14, P2.*
- **Blocked By**: P1.7
**References**: Doorenbos §2.5
**Acceptance Criteria**:
- [ ] `bun test packages/rete/src/join.test.ts` green
- [ ] Coverage ≥ 90%
- [ ] Benchmark: 100 entities × 3-condition join < 10ms
**QA Scenarios**:
```
Scenario: Multi-variable join matches entities consistently
Tool: Bash
Steps:
1. Run: bun test packages/rete/src/join.test.ts 2>&1 | tee /tmp/p18.log
Expected Result: all pass
Evidence: .sisyphus/evidence/task-P1.8-join.log
Scenario: Join perf benchmark (failure path)
Tool: Bash
Steps:
1. Run: bun test packages/rete/src/join.test.ts -t "3-condition join on 100 entities under 10ms" 2>&1 | tee /tmp/p18-perf.log
Expected Result: pass
Evidence: .sisyphus/evidence/task-P1.8-join-perf.log
```
**Commit**: YES
- Message: `feat(rete): add JoinNode with variable-binding equality tests (P1.8)`
- Files: `packages/rete/src/join.ts`, `packages/rete/src/join.test.ts`, `packages/rete/src/index.ts`
- Pre-commit: `bun run check`
- [x] P1.9. **Condition filters (cond equivalent) (TDD)**
**What to do**:
- RED: `packages/rete/src/condition.test.ts` — filter predicates applied after join; predicates are registered (via registry, for JSON serializability)
- GREEN: `packages/rete/src/condition.ts` — `FilterNode` holding `predicate: string` (registry key) + `args: JsonValue[]`; applies to incoming tokens
**Must NOT do**:
- Do NOT allow inline arrow functions in conditions (must use registry)
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P1.2
- **Blocks**: P1.11-P1.14, P2.*
- **Blocked By**: P1.8 (join produces tokens to filter)
**References**: `packages/rete/SPEC.md` §JSON Rule Schema
**Acceptance Criteria**:
- [ ] `bun test packages/rete/src/condition.test.ts` green
- [ ] Coverage ≥ 90%
**QA Scenarios**:
```
Scenario: Filter predicate correctly rejects tokens
Tool: Bash
Steps:
1. Run: bun test packages/rete/src/condition.test.ts 2>&1 | tee /tmp/p19.log
Expected Result: all pass
Evidence: .sisyphus/evidence/task-P1.9-cond.log
Scenario: Unregistered predicate throws (failure path)
Tool: Bash
Steps:
1. Run: bun test packages/rete/src/condition.test.ts -t "unknown predicate throws" 2>&1
Expected Result: pass
Evidence: .sisyphus/evidence/task-P1.9-unknown.log
```
**Commit**: YES
- Message: `feat(rete): add FilterNode with registered predicates (P1.9)`
- Files: `packages/rete/src/condition.ts`, `packages/rete/src/condition.test.ts`, `packages/rete/src/index.ts`
- Pre-commit: `bun run check`
- [x] P1.10. **Query API: query / queryAll (TDD)**
**What to do**:
- RED: `packages/rete/src/query.test.ts` — `session.query(rule)` returns first match or throws; `session.queryAll(rule)` returns all; `session.query(rule, { bindings })` filters by binding value
- GREEN: `packages/rete/src/query.ts` — wraps production node's accumulated matches; deterministic order per SPEC
**Must NOT do**:
- Do NOT allow query on rules without registered production
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P1.2
- **Blocks**: P2.*
- **Blocked By**: P1.8 (beta produces tokens)
**References**: `packages/rete/SPEC.md` §Iteration Order
**Acceptance Criteria**:
- [ ] `bun test packages/rete/src/query.test.ts` green
- [ ] Coverage ≥ 95%
**QA Scenarios**:
```
Scenario: query returns deterministic ordering
Tool: Bash
Steps:
1. Run: bun test packages/rete/src/query.test.ts 2>&1 | tee /tmp/p110.log
Expected Result: all pass
Evidence: .sisyphus/evidence/task-P1.10-query.log
Scenario: query on missing rule throws (failure path)
Tool: Bash
Steps:
1. Run: bun test packages/rete/src/query.test.ts -t "query on unknown rule throws" 2>&1
Expected Result: pass
Evidence: .sisyphus/evidence/task-P1.10-unknown.log
```
**Commit**: YES
- Message: `feat(rete): add query/queryAll API (P1.10)`
- Files: `packages/rete/src/query.ts`, `packages/rete/src/query.test.ts`, `packages/rete/src/index.ts`
- Pre-commit: `bun run check`
- [x] P1.11. **Derived facts via thenFinally-equivalent (TDD)**
**What to do**:
- RED: `packages/rete/src/derived.test.ts` — `rule.thenFinally('aggregateHandler', [])` fires after all activations of a tick; derived facts auto-retract when supporting matches disappear (truth maintenance)
- GREEN: `packages/rete/src/derived.ts` — `ProductionNode.thenFinally` handler; tracks derived facts per match chain; on match removal, retracts corresponding derived fact
**Must NOT do**:
- Do NOT allow derived fact id collision with user facts (derived facts use negative EntityIds)
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P1.3
- **Blocks**: P1.14, P2.21 (repetition detection uses derived facts)
- **Blocked By**: P1.7-P1.10
**References**: `packages/rete/SPEC.md` §Truth Maintenance
**Acceptance Criteria**:
- [ ] `bun test packages/rete/src/derived.test.ts` green
- [ ] Coverage ≥ 90%
**QA Scenarios**:
```
Scenario: thenFinally aggregates after tick
Tool: Bash
Steps:
1. Run: bun test packages/rete/src/derived.test.ts 2>&1 | tee /tmp/p111.log
Expected Result: all pass
Evidence: .sisyphus/evidence/task-P1.11-derived.log
Scenario: Derived fact retracts when support retracts (failure path)
Tool: Bash
Steps:
1. Run: bun test packages/rete/src/derived.test.ts -t "derived retracts on support loss" 2>&1
Expected Result: pass
Evidence: .sisyphus/evidence/task-P1.11-retract.log
```
**Commit**: YES
- Message: `feat(rete): add derived facts with thenFinally + truth maintenance (P1.11)`
- Files: `packages/rete/src/derived.ts`, `packages/rete/src/derived.test.ts`, `packages/rete/src/index.ts`
- Pre-commit: `bun run check`
- [x] P1.12. **Cycle detection with recursion limit (TDD)**
**What to do**:
- RED: `packages/rete/src/cycle.test.ts` — rule A inserts fact triggering rule B inserting fact triggering A (cycle); `fireRules({ recursionLimit: 4 })` throws `RecursionLimitExceededError` with cycle trace; `recursionLimit: 0` disables (for advanced use)
- GREEN: wire recursion counter into Session.fireRules; build cycle trace (last N activations); error includes rule names
**Must NOT do**:
- Do NOT silently skip cycles (error must be loud)
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P1.3
- **Blocks**: P1.14
- **Blocked By**: P1.4 (session)
**References**: `packages/rete/SPEC.md` §Cycle Detection
**Acceptance Criteria**:
- [ ] `bun test packages/rete/src/cycle.test.ts` green
- [ ] Coverage ≥ 95%
**QA Scenarios**:
```
Scenario: Cycle exceeding limit throws with trace
Tool: Bash
Steps:
1. Run: bun test packages/rete/src/cycle.test.ts 2>&1 | tee /tmp/p112.log
Expected Result: all pass; trace includes rule names
Evidence: .sisyphus/evidence/task-P1.12-cycle.log
Scenario: recursionLimit 0 allows unlimited (failure path for infinite loop detection)
Tool: Bash
Steps:
1. Run: timeout 5 bun test packages/rete/src/cycle.test.ts -t "recursionLimit 0 runs to natural fixpoint" 2>&1
Expected Result: pass within 5s (natural fixpoint reached)
Evidence: .sisyphus/evidence/task-P1.12-unlimited.log
```
**Commit**: YES
- Message: `feat(rete): add cycle detection with recursionLimit (P1.12)`
- Files: `packages/rete/src/cycle.ts`, `packages/rete/src/cycle.test.ts`, `packages/rete/src/session.ts`
- Pre-commit: `bun run check`
- [x] P1.13. **Deterministic conflict resolution (TDD)**
**What to do**:
- RED: `packages/rete/src/conflict.test.ts` — given N matching activations, firing order is: salience desc → specificity (# conditions) desc → insertion order asc; deterministic across runs
- GREEN: `packages/rete/src/conflict.ts` — `orderActivations(activations)` pure function; integrate into Session.fireRules
**Must NOT do**:
- Do NOT use `Math.random` for tiebreaking
- Do NOT sort by rule name lexicographically (that hides bugs via alphabetization)
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P1.3
- **Blocks**: P1.14
- **Blocked By**: P1.4
**References**: `packages/rete/SPEC.md` §Conflict Resolution
**Acceptance Criteria**:
- [ ] `bun test packages/rete/src/conflict.test.ts` green
- [ ] Fuzz test: 100 random rule sets, 2 identical runs → identical fire order
- [ ] Coverage ≥ 95%
**QA Scenarios**:
```
Scenario: Fire order matches spec for mixed salience
Tool: Bash
Steps:
1. Run: bun test packages/rete/src/conflict.test.ts 2>&1 | tee /tmp/p113.log
Expected Result: all pass
Evidence: .sisyphus/evidence/task-P1.13-conflict.log
Scenario: Determinism fuzz (failure path for non-det)
Tool: Bash
Steps:
1. Run: bun test packages/rete/src/conflict.test.ts -t "fuzz 100 rule sets yield identical fire order" 2>&1
Expected Result: pass
Evidence: .sisyphus/evidence/task-P1.13-fuzz.log
```
**Commit**: YES
- Message: `feat(rete): add deterministic conflict resolution (P1.13)`
- Files: `packages/rete/src/conflict.ts`, `packages/rete/src/conflict.test.ts`, `packages/rete/src/session.ts`
- Pre-commit: `bun run check`
- [x] P1.14. **Pararules golden-file test port**
**What to do**:
- Port 5-10 representative pararules tests from `paranim/pararules/tests/*.nim` to TS/Vitest under `packages/rete/tests/golden/`
- Each test = fixture (fact insertion script + rule definitions) + expected query results snapshot
- Add Vitest snapshots for derived-fact cases
- Document each golden's pararules-origin line reference in a `GOLDEN-MAP.md`
**Must NOT do**:
- Do NOT skip tests that exercise derived facts / multi-condition joins
**Recommended Agent Profile**: `unspecified-high`
- **Skills**: [`repo-analysis`]
- `repo-analysis`: Retrieve pararules tests from GitHub
**Parallelization**: NO — Wave P1.4 (parity gate)
- **Blocks**: Phase 2 start
- **Blocked By**: P1.1-P1.13
**References**:
- https://github.com/paranim/pararules/blob/master/tests/test1.nim
- https://github.com/paranim/pararules/blob/master/tests/test2.nim
- https://github.com/paranim/pararules/blob/master/tests/test3.nim
**Acceptance Criteria**:
- [ ] `bun test packages/rete/tests/golden` → all green
- [ ] `packages/rete/tests/golden/GOLDEN-MAP.md` lists each ported test with origin line
- [ ] Coverage of engine src ≥ 90%
- [ ] Tag `v0.1.0-phase1`
**QA Scenarios**:
```
Scenario: Golden suite passes end-to-end
Tool: Bash
Steps:
1. Run: bun test packages/rete/tests/golden 2>&1 | tee /tmp/p114.log
2. Run: bun run test:coverage -- --coverage.reporter=text packages/rete 2>&1 | tee /tmp/p114-cov.log
3. Run: grep -oE 'All files.*[0-9.]+' /tmp/p114-cov.log | head -1
Expected Result: all tests pass; line coverage ≥ 90%
Evidence: .sisyphus/evidence/task-P1.14-golden.log, .sisyphus/evidence/task-P1.14-cov.log
Scenario: Phase 1 tag exists
Tool: Bash
Steps:
1. Run: git tag v0.1.0-phase1
2. Run: git tag | grep v0.1.0-phase1
Expected Result: tag output present
Evidence: .sisyphus/evidence/task-P1.14-tag.log
```
**Commit**: YES
- Message: `test(rete): port pararules golden tests; tag Phase 1 parity (P1.14)`
- Files: `packages/rete/tests/golden/*.test.ts`, `packages/rete/tests/golden/GOLDEN-MAP.md`
- Pre-commit: `bun run check`
- Post-commit: `git tag v0.1.0-phase1`
### Phase 2 — Rete II Extensions + Chess Engine
- [x] P2.1. **Negation nodes (NOT) (TDD)**
**What to do**:
- RED: `packages/rete/src/negation.test.ts` — `rule.whatNot((Player, Dead, v(true)))` matches only when no fact satisfies the negated pattern; activation toggles when blocking fact inserted/retracted
- GREEN: `packages/rete/src/negation.ts` — `NegationNode` per Doorenbos §2.6; counts matching facts; token passes iff count is zero
**Must NOT do**: implement unsafe NOT (unbound vars in NOT) — reject at build time
**Recommended Agent Profile**: `deep`; Skills: [`context7`]
**Parallelization**: YES — Wave P2.1 (with P2.2, P2.3, P2.4)
**Blocks**: P2.13, P2.18 (check detection uses NOT)
**Blocked By**: P1.14 (Phase 1 gate)
**References**: Doorenbos §2.6.1
**Acceptance Criteria**:
- [ ] `bun test packages/rete/src/negation.test.ts` green
- [ ] Coverage ≥ 90%
**QA Scenarios**:
```
Scenario: NOT fires when pattern absent; retracts when inserted
Tool: Bash
Steps: 1. Run: bun test packages/rete/src/negation.test.ts 2>&1 | tee /tmp/p21.log
Expected Result: all pass
Evidence: .sisyphus/evidence/task-P2.1-not.log
Scenario: Unsafe NOT (unbound var) rejected at build (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/rete/src/negation.test.ts -t "unsafe NOT rejected" 2>&1
Expected Result: pass
Evidence: .sisyphus/evidence/task-P2.1-unsafe.log
```
**Commit**: YES — `feat(rete): add negation nodes (NOT) (P2.1)` — files: `packages/rete/src/negation.ts`, `packages/rete/src/negation.test.ts`
- [x] P2.2. **Existential nodes (EXISTS) (TDD)**
**What to do**: `rule.whatExists((Attacker, AttacksSquare, v('sq')))` — EXISTS is negation-of-negation; propagate token if ≥1 matching fact. GREEN: `packages/rete/src/existential.ts`
**Must NOT do**: double-count (increment on same fact twice)
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P2.1
**Blocks**: P2.18, P2.19
**Blocked By**: P1.14
**References**: Doorenbos §2.6.2
**Acceptance Criteria**:
- [ ] `bun test packages/rete/src/existential.test.ts` green
- [ ] Coverage ≥ 90%
**QA Scenarios**:
```
Scenario: EXISTS toggles correctly
Tool: Bash
Steps: 1. Run: bun test packages/rete/src/existential.test.ts 2>&1 | tee /tmp/p22.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P2.2-exists.log
Scenario: Multiple supporting facts do not re-activate (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/rete/src/existential.test.ts -t "single activation despite multiple supports" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P2.2-single.log
```
**Commit**: YES — `feat(rete): add existential nodes (EXISTS) (P2.2)`
- [x] P2.3. **NCC nodes: not-count-condition (TDD)**
**What to do**: Subconjunction negation — "no matching combination of N conditions exists". GREEN: `packages/rete/src/ncc.ts`. Per Doorenbos §2.6.3, NCC is a sub-network whose top-level production feeds a negation partner.
**Must NOT do**: collapse NCC into NOT (NCC is strictly more powerful)
**Recommended Agent Profile**: `deep`; Skills: [`context7`]
**Parallelization**: YES — Wave P2.1
**Blocks**: P2.19
**Blocked By**: P1.14, P2.1 (reuses negation machinery)
**References**: Doorenbos §2.6.3
**Acceptance Criteria**:
- [ ] `bun test packages/rete/src/ncc.test.ts` green
- [ ] Coverage ≥ 90%
**QA Scenarios**:
```
Scenario: NCC rejects when combination exists
Tool: Bash
Steps: 1. Run: bun test packages/rete/src/ncc.test.ts 2>&1 | tee /tmp/p23.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P2.3-ncc.log
Scenario: NCC partner cleanup on retract (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/rete/src/ncc.test.ts -t "NCC partner cleans up on retract" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P2.3-retract.log
```
**Commit**: YES — `feat(rete): add NCC nodes (P2.3)`
- [x] P2.4. **Aggregation nodes: count/sum/collect/min/max (TDD)**
**What to do**: `rule.whatAggregate(count, (?id, Health, v('h')))` returns count bound to variable. Support `count`, `sum`, `min`, `max`, `collect` (array). Incremental update: maintain running total rather than full recompute. GREEN: `packages/rete/src/aggregate.ts`
**Must NOT do**: full-recompute on every change (performance); operate on raw Set iteration
**Recommended Agent Profile**: `deep`; Skills: [`context7`]
**Parallelization**: YES — Wave P2.1
**Blocks**: P2.21 (50-move + threefold use aggregation)
**Blocked By**: P1.14
**References**: Drools aggregation patterns
**Acceptance Criteria**:
- [ ] `bun test packages/rete/src/aggregate.test.ts` green
- [ ] Benchmark: 1000 facts × 5 aggregators < 20ms per full re-run
- [ ] Coverage ≥ 90%
**QA Scenarios**:
```
Scenario: All 5 aggregators produce correct values
Tool: Bash
Steps: 1. Run: bun test packages/rete/src/aggregate.test.ts 2>&1 | tee /tmp/p24.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P2.4-agg.log
Scenario: Incremental sum on retract (failure path for full recompute)
Tool: Bash
Steps: 1. Run: bun test packages/rete/src/aggregate.test.ts -t "sum updates incrementally on retract" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P2.4-incr.log
```
**Commit**: YES — `feat(rete): add aggregation nodes (count/sum/collect/min/max) (P2.4)`
- [x] P2.5. **Chess attribute schema + piece fact shape**
**What to do**: `packages/chess/src/schema.ts` — define attrs: `PieceType` (pawn|knight|bishop|rook|queen|king), `Color` (white|black), `Square` (a1..h8 as number 0..63), `Position` (id→Square), `HasMoved` (bool for castling), `Turn` (color), `HalfmoveClock` (number), `FullmoveNumber` (number), `EnPassantTarget` (Square?); piece entity convention (each piece = one entity with multiple attrs)
- TDD the schema types (compile-time only test via `expect-type`)
**Must NOT do**: use strings for squares (numeric 0..63 for perf)
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P2.2
**Blocks**: P2.8-P2.22
**Blocked By**: P1.14
**References**: `packages/chess/RULES.md`, `packages/rete/SPEC.md`
**Acceptance Criteria**:
- [ ] `packages/chess/src/schema.ts` exports typed schema
- [ ] `bun x tsc -b packages/chess` → 0 errors
- [ ] `bun test packages/chess/src/schema.test.ts` green
**QA Scenarios**:
```
Scenario: Chess schema compiles with strict types
Tool: Bash
Steps:
1. Run: bun x tsc -b packages/chess 2>&1 | tee /tmp/p25.log
2. Run: bun test packages/chess/src/schema.test.ts 2>&1 | tee /tmp/p25-test.log
Expected: step 1 exit 0; step 2 all pass
Evidence: .sisyphus/evidence/task-P2.5-schema.log
Scenario: Square is numeric 0..63 (failure path for string squares)
Tool: Bash
Steps: 1. Run: grep -E "type Square = .*0..63|type Square = .*number" packages/chess/src/schema.ts
Expected: match present
Evidence: .sisyphus/evidence/task-P2.5-square.log
```
**Commit**: YES — `feat(chess): add attribute schema and piece fact shape (P2.5)`
- [x] P2.6. **Starting-position fact generator**
**What to do**: `packages/chess/src/starting-position.ts` — `generateStartingPosition(session)` inserts 32 piece facts for FIDE start. TDD via snapshot of `session.allFacts()` sorted output.
**Must NOT do**: hardcode as JSON fixture (must be generated deterministically)
**Recommended Agent Profile**: `quick`
**Parallelization**: YES — Wave P2.2
**Blocks**: P2.8+
**Blocked By**: P2.5
**Acceptance Criteria**:
- [ ] `bun test packages/chess/src/starting-position.test.ts` green
- [ ] Facts match FIDE snapshot
**QA Scenarios**:
```
Scenario: Starting position snapshot matches FIDE
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/starting-position.test.ts 2>&1 | tee /tmp/p26.log
Expected: all pass; snapshot file `__snapshots__/starting-position.test.ts.snap` exists
Evidence: .sisyphus/evidence/task-P2.6-start.log
Scenario: Exactly 32 pieces (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/starting-position.test.ts -t "inserts exactly 32 piece entities" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P2.6-count.log
```
**Commit**: YES — `feat(chess): add starting-position fact generator (P2.6)`
- [x] P2.7. **Square + color helpers**
**What to do**: `packages/chess/src/coord.ts` — pure functions: `fileOf(square)`, `rankOf(square)`, `squareFromFileRank(f, r)`, `colorOf(square)` (light/dark), `oppositeColor(c)`, `isOnBoard(f, r)`; TDD each
**Must NOT do**: use string representations internally
**Recommended Agent Profile**: `quick`
**Parallelization**: YES — Wave P2.2
**Blocks**: P2.9-P2.12
**Blocked By**: P2.5
**Acceptance Criteria**:
- [ ] `bun test packages/chess/src/coord.test.ts` green
- [ ] Coverage ≥ 100%
**QA Scenarios**:
```
Scenario: Coord helpers pure + total
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/coord.test.ts 2>&1 | tee /tmp/p27.log
Expected: all pass; 100% line coverage
Evidence: .sisyphus/evidence/task-P2.7-coord.log
Scenario: Off-board rejection (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/coord.test.ts -t "isOnBoard rejects out-of-range" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P2.7-off.log
```
**Commit**: YES — `feat(chess): add coordinate + color helpers (P2.7)`
- [x] P2.8. **Piece movement primitive rules (directions + steps)**
**What to do**: `packages/chess/src/rules/primitives.ts` — rule-level primitives that legal-move rules build on: `StraightLineMoves`, `DiagonalMoves`, `SingleStepMoves`, `KnightOffsets`, `PawnSingleAdvance`, `PawnDoubleAdvance`, `PawnDiagonalCapture`. Each primitive is one Rete production generating candidate moves as derived facts (e.g., `CandidateMove(pieceId, targetSquare)`).
**Must NOT do**: embed legality checks (check/pin/etc) in primitives (those layer in P2.13+)
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P2.2
**Blocks**: P2.9-P2.14
**Blocked By**: P2.5, P2.7
**Acceptance Criteria**:
- [ ] `bun test packages/chess/src/rules/primitives.test.ts` green
- [ ] Each primitive is a registered rule (listed in a primitives manifest)
**QA Scenarios**:
```
Scenario: Primitives generate candidate moves
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/primitives.test.ts 2>&1 | tee /tmp/p28.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P2.8-prims.log
Scenario: Primitives do NOT generate captures (separation of concerns, failure path)
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/primitives.test.ts -t "primitives produce only non-capture candidates" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P2.8-sep.log
```
**Commit**: YES — `feat(chess): add movement primitive rules (P2.8)`
- [x] P2.9. **Pawn move/capture rules**
**What to do**: `packages/chess/src/rules/pawn.ts` — productions: `PawnSingleMove`, `PawnDoubleMoveFromHome`, `PawnDiagonalCapture`. Use primitives + filters. Color-aware (white advances +rank, black -rank). Emit `LegalMove(pieceId, from, to)` derived facts. TDD each case including blocked paths.
**Must NOT do**: handle en passant yet (P2.16)
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P2.3 (with P2.10-P2.14)
**Blocks**: P2.13, P2.16, P2.17
**Blocked By**: P2.8
**Acceptance Criteria**:
- [ ] `bun test packages/chess/src/rules/pawn.test.ts` green
- [ ] Tests cover: single move, double from home, blocked by own piece, blocked by enemy, diagonal capture, no diagonal without capture
**QA Scenarios**:
```
Scenario: All pawn movement and capture cases
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/pawn.test.ts 2>&1 | tee /tmp/p29.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P2.9-pawn.log
Scenario: Pawn cannot move diagonally without capture (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/pawn.test.ts -t "pawn diagonal without capture rejected" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P2.9-diag.log
```
**Commit**: YES — `feat(chess): add pawn move/capture rules (P2.9)`
- [x] P2.10. **Knight move rules**
**What to do**: `packages/chess/src/rules/knight.ts` — 8 L-offsets; leap over other pieces; `LegalMove` emission
**Must NOT do**: filter path squares (knight leaps)
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P2.3
**Blocks**: P2.13
**Blocked By**: P2.8
**Acceptance Criteria**:
- [ ] `bun test packages/chess/src/rules/knight.test.ts` green
**QA Scenarios**:
```
Scenario: Knight L-moves from all positions
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/knight.test.ts 2>&1 | tee /tmp/p210.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P2.10-knight.log
Scenario: Knight leaps over pieces (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/knight.test.ts -t "knight ignores intervening pieces" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P2.10-leap.log
```
**Commit**: YES — `feat(chess): add knight move rules (P2.10)`
- [x] P2.11. **Bishop/Rook/Queen sliding rules**
**What to do**: `packages/chess/src/rules/sliding.ts` — `SlidingMove` production parameterized by directions (diagonal, orthogonal, both); uses aggregation or sequential tokens to stop at first blocker (own = stop before; enemy = capture then stop)
**Must NOT do**: generate moves beyond blocker
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P2.3
**Blocks**: P2.13, P2.15 (castling reads rook moves)
**Blocked By**: P2.8
**Acceptance Criteria**:
- [ ] `bun test packages/chess/src/rules/sliding.test.ts` green
- [ ] Tests cover all three pieces × blocker scenarios
**QA Scenarios**:
```
Scenario: Sliding moves stop correctly
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/sliding.test.ts 2>&1 | tee /tmp/p211.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P2.11-sliding.log
Scenario: Sliding piece cannot jump (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/sliding.test.ts -t "bishop stops at blocker" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P2.11-stop.log
```
**Commit**: YES — `feat(chess): add bishop/rook/queen sliding rules (P2.11)`
- [x] P2.12. **King move rules (basic)**
**What to do**: `packages/chess/src/rules/king.ts` — 8 adjacent squares; excludes squares occupied by own piece. Castling deferred to P2.15; check-aware rejection deferred to P2.13.
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P2.3
**Blocks**: P2.13, P2.15, P2.18, P2.19
**Blocked By**: P2.8
**Acceptance Criteria**:
- [ ] `bun test packages/chess/src/rules/king.test.ts` green
**QA Scenarios**:
```
Scenario: King single-step moves
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/king.test.ts 2>&1 | tee /tmp/p212.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P2.12-king.log
Scenario: King blocked by own piece (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/king.test.ts -t "king blocked by own piece" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P2.12-block.log
```
**Commit**: YES — `feat(chess): add king basic move rules (P2.12)`
- [x] P2.13. **Turn order + move legality integration**
**What to do**: `packages/chess/src/rules/turn.ts` — only pieces of current turn's color generate legal moves; after move, turn flips; move-intent fact (`AttemptedMove`) validated vs `LegalMove` set; on success, update piece positions + retract old `LegalMove` facts. Uses negation to reject intents with no matching LegalMove.
**Must NOT do**: allow movement into check (that's P2.18, but at least queue the integration point here)
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P2.3
**Blocks**: P2.14, all further rules
**Blocked By**: P2.9, P2.10, P2.11, P2.12, P2.1 (negation)
**Acceptance Criteria**:
- [ ] `bun test packages/chess/src/rules/turn.test.ts` green
- [ ] Full single move validated and applied
**QA Scenarios**:
```
Scenario: Legal move applied; turn switches
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/turn.test.ts 2>&1 | tee /tmp/p213.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P2.13-turn.log
Scenario: Illegal move rejected, turn unchanged (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/turn.test.ts -t "illegal move rejected" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P2.13-illegal.log
```
**Commit**: YES — `feat(chess): add turn order + move integration (P2.13)`
- [x] P2.14. **Capture resolution rules**
**What to do**: `packages/chess/src/rules/capture.ts` — when a LegalMove targets an enemy-occupied square, applying the move retracts the captured piece's facts (Position, PieceType, Color) via the RHS handler.
**Must NOT do**: modify captured piece's facts (they retract entirely in FIDE; other presets may vary — handled in presets)
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P2.3
**Blocks**: P2.16, P2.19-P2.22
**Blocked By**: P2.13
**Acceptance Criteria**:
- [ ] `bun test packages/chess/src/rules/capture.test.ts` green
**QA Scenarios**:
```
Scenario: Capture removes enemy piece
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/capture.test.ts 2>&1 | tee /tmp/p214.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P2.14-cap.log
Scenario: Cannot capture own piece (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/capture.test.ts -t "cannot capture own piece" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P2.14-own.log
```
**Commit**: YES — `feat(chess): add capture resolution (P2.14)`
- [x] P2.15. **Castling (kingside + queenside)**
**What to do**: `packages/chess/src/rules/castling.ts` — productions requiring: King has not moved (HasMoved=false), relevant Rook has not moved, no pieces between, king not in check, transit squares not attacked. Two-piece move: king + rook positions updated atomically.
**Must NOT do**: allow castling through check
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P2.4 (with P2.16-P2.18)
**Blocks**: P2.23 (integration test)
**Blocked By**: P2.11, P2.12, P2.18 (check detection for transit squares)
**References**: FIDE §3.8.2
**Acceptance Criteria**:
- [ ] `bun test packages/chess/src/rules/castling.test.ts` green
- [ ] Tests: kingside, queenside, rejected after king moves, rejected through check
**QA Scenarios**:
```
Scenario: Both castling directions + rejection cases
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/castling.test.ts 2>&1 | tee /tmp/p215.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P2.15-castle.log
Scenario: Castling rejected through check (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/castling.test.ts -t "castling rejected when king passes attacked square" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P2.15-through.log
```
**Commit**: YES — `feat(chess): add castling rules (P2.15)`
- [x] P2.16. **En passant (single-tick capture window)**
**What to do**: `packages/chess/src/rules/enpassant.ts` — after a pawn's double-advance, set `EnPassantTarget(turn, square)` fact for one turn; eligible-pawn rule emits LegalMove that captures via adjacent target; target fact retracts on next turn.
**Must NOT do**: allow en passant beyond one turn window
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P2.4
**Blocks**: P2.23
**Blocked By**: P2.9, P2.14
**References**: FIDE §3.7.3
**Acceptance Criteria**:
- [ ] `bun test packages/chess/src/rules/enpassant.test.ts` green
**QA Scenarios**:
```
Scenario: En passant capture works within 1-turn window
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/enpassant.test.ts 2>&1 | tee /tmp/p216.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P2.16-ep.log
Scenario: En passant disallowed after window closes (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/enpassant.test.ts -t "en passant window closes after one turn" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P2.16-window.log
```
**Commit**: YES — `feat(chess): add en passant rule (P2.16)`
- [x] P2.17. **Promotion**
**What to do**: `packages/chess/src/rules/promotion.ts` — when a pawn reaches final rank, retract pawn PieceType fact and insert new PieceType (Q/R/B/N). The choice is specified in the `AttemptedMove` fact via `promoteTo` field; default to Q if missing.
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P2.4
**Blocks**: P2.23
**Blocked By**: P2.9
**References**: FIDE §3.7.5
**Acceptance Criteria**:
- [ ] `bun test packages/chess/src/rules/promotion.test.ts` green
- [ ] Tests: promotion to Q/R/B/N, default-to-queen
**QA Scenarios**:
```
Scenario: Pawn promotion to each valid piece
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/promotion.test.ts 2>&1 | tee /tmp/p217.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P2.17-promo.log
Scenario: Invalid promotion target rejected (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/promotion.test.ts -t "promotion to king rejected" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P2.17-invalid.log
```
**Commit**: YES — `feat(chess): add pawn promotion rule (P2.17)`
- [x] P2.18. **Check detection**
**What to do**: `packages/chess/src/rules/check.ts` — derived fact `InCheck(color)` when any enemy piece has a LegalMove targeting that color's king. Uses EXISTS node. Rules that would leave own king in check are filtered out of LegalMove (self-check filter).
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P2.4
**Blocks**: P2.15 (castling through check), P2.19, P2.23
**Blocked By**: P2.2 (exists), P2.9-P2.14
**Acceptance Criteria**:
- [ ] `bun test packages/chess/src/rules/check.test.ts` green
**QA Scenarios**:
```
Scenario: Check detected; self-check prevented
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/check.test.ts 2>&1 | tee /tmp/p218.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P2.18-check.log
Scenario: Move leaving own king in check rejected (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/check.test.ts -t "move exposing own king rejected" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P2.18-self.log
```
**Commit**: YES — `feat(chess): add check detection + self-check filter (P2.18)`
- [x] P2.19. **Checkmate detection**
**What to do**: `packages/chess/src/rules/checkmate.ts` — derived fact `GameOver(result, reason)` when: `InCheck(turn)` AND no LegalMove exists for any piece of `turn`. Uses NCC.
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P2.5 (with P2.20-P2.22)
**Blocks**: P2.23
**Blocked By**: P2.3 (NCC), P2.18
**Acceptance Criteria**:
- [ ] `bun test packages/chess/src/rules/checkmate.test.ts` green
- [ ] Tests: Fool's Mate, Scholar's Mate, back-rank mate
**QA Scenarios**:
```
Scenario: Checkmate positions detected
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/checkmate.test.ts 2>&1 | tee /tmp/p219.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P2.19-mate.log
Scenario: Check without mate is not mate (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/checkmate.test.ts -t "check with escape is not mate" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P2.19-noesc.log
```
**Commit**: YES — `feat(chess): add checkmate detection (P2.19)`
- [x] P2.20. **Stalemate detection**
**What to do**: `packages/chess/src/rules/stalemate.ts` — `GameOver('draw', 'stalemate')` when: NOT `InCheck(turn)` AND no LegalMove exists for `turn`.
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P2.5
**Blocks**: P2.23
**Blocked By**: P2.3 (NCC), P2.18
**Acceptance Criteria**:
- [ ] `bun test packages/chess/src/rules/stalemate.test.ts` green
**QA Scenarios**:
```
Scenario: Stalemate detected
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/stalemate.test.ts 2>&1 | tee /tmp/p220.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P2.20-stale.log
Scenario: Checkmate not mistaken for stalemate (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/stalemate.test.ts -t "checkmate distinguished from stalemate" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P2.20-dist.log
```
**Commit**: YES — `feat(chess): add stalemate detection (P2.20)`
- [x] P2.21. **50-move rule + threefold repetition (aggregation-based)**
**What to do**: `packages/chess/src/rules/draws.ts` — track halfmove clock (resets on pawn move or capture) via a rule; 50-move rule fires at 100 halfmoves. For threefold, maintain a `PositionHash` fact per tick; aggregation counts occurrences of each hash; threshold of 3 → draw claim available.
**Must NOT do**: auto-claim (threefold is claimable, but plan keeps it auto-triggered on 3rd occurrence for simplicity; documented)
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P2.5
**Blocks**: P2.23
**Blocked By**: P2.4 (aggregation), P2.14
**References**: FIDE §5.2.2, §5.2.3
**Acceptance Criteria**:
- [ ] `bun test packages/chess/src/rules/draws.test.ts` green
**QA Scenarios**:
```
Scenario: 50-move + threefold detected
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/draws.test.ts 2>&1 | tee /tmp/p221.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P2.21-draws.log
Scenario: Clock reset on capture (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/draws.test.ts -t "halfmove clock resets on capture" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P2.21-reset.log
```
**Commit**: YES — `feat(chess): add 50-move and threefold repetition rules (P2.21)`
- [x] P2.22. **Insufficient material draw**
**What to do**: `packages/chess/src/rules/insufficient.ts` — draw when material sets are: KvK, KvK+N, KvK+B, K+BvK+B (same color bishop). Uses aggregation count over piece types.
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P2.5
**Blocks**: P2.23
**Blocked By**: P2.4 (aggregation)
**Acceptance Criteria**:
- [ ] `bun test packages/chess/src/rules/insufficient.test.ts` green
**QA Scenarios**:
```
Scenario: All 4 insufficient-material configurations detected
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/insufficient.test.ts 2>&1 | tee /tmp/p222.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P2.22-insuf.log
Scenario: Bishops on opposite colors NOT draw (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/rules/insufficient.test.ts -t "opposite-color bishops is not insufficient" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P2.22-opp.log
```
**Commit**: YES — `feat(chess): add insufficient material draw (P2.22)`
- [x] P2.23. **End-to-end FIDE game replay integration test**
**What to do**: `packages/chess/tests/fide-games/`: include 5 famous games as PGN fixtures (Immortal, Opera, Evergreen, Kasparov vs Topalov 1999, Deep Blue vs Kasparov G6 1997). Write a runner that parses PGN, drives moves through the engine, asserts each move accepted, asserts terminal state (mate/draw/resign). Resigns are not a chess rule — handled as UI-only terminal state for now; filter those from fixtures.
**Must NOT do**: add a PGN parser dependency (write minimal hand-rolled parser for SAN within `packages/chess/src/pgn.ts`)
**Recommended Agent Profile**: `unspecified-high`
**Parallelization**: NO — Wave P2.6 (gate)
**Blocks**: Phase 3
**Blocked By**: P2.1-P2.22
**Acceptance Criteria**:
- [ ] `bun test packages/chess/tests/fide-games` → all 5 games replay to completion
- [ ] Phase 2 tag: `git tag v0.2.0-phase2`
**QA Scenarios**:
```
Scenario: 5 classic games replay end-to-end
Tool: Bash
Steps:
1. Run: bun test packages/chess/tests/fide-games 2>&1 | tee /tmp/p223.log
2. Run: grep -c 'PASS.*\.pgn' /tmp/p223.log
Expected: all tests pass; grep >= 5
Evidence: .sisyphus/evidence/task-P2.23-games.log
Scenario: Phase 2 tag created
Tool: Bash
Steps: 1. Run: git tag v0.2.0-phase2 && git tag | grep v0.2.0-phase2
Expected: tag present
Evidence: .sisyphus/evidence/task-P2.23-tag.log
```
**Commit**: YES — `test(chess): replay 5 classic FIDE games; tag Phase 2 (P2.23)`; post-commit: `git tag v0.2.0-phase2`
### Phase 3 — Time-Travel + Presets + UI
- [x] P3.1. **Event log: append-only, monotonic sequence numbers (TDD)**
**What to do**: `packages/rete/src/eventlog.ts` — `class EventLog` records every `insert(id, attr, value)`, `retract(id, attr)`, and rule-fire as `{ seq, ts, kind, payload }`. Append-only; `getSince(seq)` returns entries after seq. Session integrates: every state-mutating call appends to log (if log attached). Tests cover monotonic seq, replay-safe encoding, payload determinism.
**Must NOT do**: allow out-of-order writes
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P3.1 (with P3.2, P3.3)
**Blocks**: P3.3, P3.14, P4.7
**Blocked By**: P2.23
**Acceptance Criteria**:
- [ ] `bun test packages/rete/src/eventlog.test.ts` green
- [ ] Coverage ≥ 95%
**QA Scenarios**:
```
Scenario: Log records every mutation monotonically
Tool: Bash
Steps: 1. Run: bun test packages/rete/src/eventlog.test.ts 2>&1 | tee /tmp/p31.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P3.1-eventlog.log
Scenario: Out-of-order append rejected (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/rete/src/eventlog.test.ts -t "out-of-order append throws" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P3.1-order.log
```
**Commit**: YES — `feat(rete): add append-only event log with monotonic sequence (P3.1)`
- [x] P3.2. **Immer snapshot every N ticks (TDD)**
**What to do**: `packages/rete/src/snapshot.ts` — on every Nth `fireRules()` call (configurable, default N=30), capture full WM state via Immer's `produce`. Structural sharing minimizes copies. `getSnapshotAt(seq)` returns nearest snapshot ≤ seq. Add `Session` option `snapshotInterval: number`.
**Must NOT do**: snapshot mid-tick (must be at tick boundary only)
**Recommended Agent Profile**: `deep`; Skills: [`context7`]
**Parallelization**: YES — Wave P3.1
**Blocks**: P3.3, P3.14
**Blocked By**: P2.23
**References**: Immer docs
**Acceptance Criteria**:
- [ ] `bun test packages/rete/src/snapshot.test.ts` green
- [ ] Memory test: 1000 ticks with N=30 produces ~33 snapshots, total memory < 10MB for chess-sized WM
- [ ] Coverage ≥ 90%
**QA Scenarios**:
```
Scenario: Snapshots captured at expected interval
Tool: Bash
Steps: 1. Run: bun test packages/rete/src/snapshot.test.ts 2>&1 | tee /tmp/p32.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P3.2-snap.log
Scenario: Memory bound with structural sharing (failure path if Immer misused)
Tool: Bash
Steps: 1. Run: bun test packages/rete/src/snapshot.test.ts -t "1000 ticks under 10MB" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P3.2-mem.log
```
**Commit**: YES — `feat(rete): add Immer snapshots at tick boundaries (P3.2)`
- [x] P3.3. **Replay engine + determinism hash verifier (TDD)**
**What to do**: `packages/rete/src/replay.ts` — `replayFromLog(log, schema, handlers): Session` reconstructs WM by replaying events on a fresh session. `stateHash(session): string` produces sha256 over sorted facts. Determinism test: recording a random fact/rule sequence, replaying, comparing hashes — must match byte-for-byte. Add `scripts/replay-determinism.ts` runner for CI.
**Must NOT do**: depend on Map/Set iteration order (sort before hashing)
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P3.1
**Blocks**: P3.14, P4.7
**Blocked By**: P3.1, P3.2
**Acceptance Criteria**:
- [ ] `bun test packages/rete/src/replay.test.ts` green
- [ ] `bun run scripts/replay-determinism.ts packages/chess/tests/fide-games/*.pgn` → 5/5 hash match
- [ ] Coverage ≥ 95%
**QA Scenarios**:
```
Scenario: Replay hash matches recording hash across 5 games
Tool: Bash
Steps:
1. Run: bun test packages/rete/src/replay.test.ts 2>&1 | tee /tmp/p33.log
2. Run: bun run scripts/replay-determinism.ts 2>&1 | tee /tmp/p33-run.log
3. Run: grep -c 'MATCH' /tmp/p33-run.log
Expected: step 1 pass; step 3 >= 5
Evidence: .sisyphus/evidence/task-P3.3-replay.log, .sisyphus/evidence/task-P3.3-hashes.log
Scenario: Injected non-determinism detected (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/rete/src/replay.test.ts -t "non-deterministic RHS produces MISMATCH" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P3.3-mismatch.log
```
**Commit**: YES — `feat(rete): add replay engine + state-hash determinism verifier (P3.3)`
- [x] P3.4. **Preset rules 1-3 (pawn-focused variants)**
**What to do**: Implement 3 of the 15 presets from `packages/chess/RULES.md` (assume first 3 are pawn-focused: e.g., `pawns-move-backward`, `pawns-diagonal-no-capture`, `double-advance-any-turn`). Each preset = one or more rule definitions in `packages/chess/src/presets/{id}.ts`, a registered toggle in `packages/chess/src/presets/registry.ts`, unit tests, compatibility declarations.
**Must NOT do**: implement presets outside the first 3
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P3.2 (with P3.5-P3.8)
**Blocks**: P3.11 (UI needs presets registered)
**Blocked By**: P2.23, P0.3 (RULES.md)
**Acceptance Criteria**:
- [ ] `bun test packages/chess/src/presets/{preset-1,2,3}.test.ts` green
- [ ] 3 presets registered; registry has 3 entries in this task
**QA Scenarios**:
```
Scenario: Presets 1-3 toggle on/off correctly
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/presets 2>&1 | tee /tmp/p34.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P3.4-presets.log
Scenario: Incompatible presets flag conflict (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/presets -t "incompatible pair flagged" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P3.4-incompat.log
```
**Commit**: YES — `feat(chess): add preset rules 1-3 (P3.4)`
- [x] P3.5. **Preset rules 4-6 (knight/bishop variants)**
**What to do**: Implement presets 4-6 from RULES.md. Same structure as P3.4.
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P3.2
**Blocks**: P3.11
**Blocked By**: P2.23, P0.3
**Acceptance Criteria**:
- [ ] `bun test packages/chess/src/presets` includes 6 preset files green; registry has 6 entries
**QA Scenarios**:
```
Scenario: Presets 4-6 functional
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/presets/knight*.test.ts packages/chess/src/presets/bishop*.test.ts 2>&1 | tee /tmp/p35.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P3.5-presets.log
Scenario: Registry expanded to 6 (failure path)
Tool: Bash
Steps: 1. Run: bun -e "import {REGISTRY} from './packages/chess/src/presets/registry.ts'; console.log(Object.keys(REGISTRY).length)"
Expected: stdout "6"
Evidence: .sisyphus/evidence/task-P3.5-count.log
```
**Commit**: YES — `feat(chess): add preset rules 4-6 (P3.5)`
- [x] P3.6. **Preset rules 7-9 (rook/queen/king variants)**
**What to do**: Implement presets 7-9 from RULES.md.
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P3.2
**Blocks**: P3.11
**Blocked By**: P2.23, P0.3
**Acceptance Criteria**: registry has 9 entries; all tests green
**QA Scenarios**:
```
Scenario: Presets 7-9 functional
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/presets/rook*.test.ts packages/chess/src/presets/queen*.test.ts packages/chess/src/presets/king*.test.ts 2>&1 | tee /tmp/p36.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P3.6.log
Scenario: Registry has 9 entries (failure path)
Tool: Bash
Steps: 1. Run: bun -e "import {REGISTRY} from './packages/chess/src/presets/registry.ts'; console.log(Object.keys(REGISTRY).length)"
Expected: stdout "9"
Evidence: .sisyphus/evidence/task-P3.6-count.log
```
**Commit**: YES — `feat(chess): add preset rules 7-9 (P3.6)`
- [x] P3.7. **Preset rules 10-12 (board/geometry variants)**
**What to do**: Implement presets 10-12 from RULES.md — board-geometry changes (e.g., horizontal wrap). These modify coord helpers via override or interception rule.
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P3.2
**Blocks**: P3.11
**Blocked By**: P2.23, P0.3
**Acceptance Criteria**: registry has 12 entries; all tests green
**QA Scenarios**:
```
Scenario: Geometry presets alter legal moves correctly
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/presets/wrap*.test.ts packages/chess/src/presets/geometry*.test.ts 2>&1 | tee /tmp/p37.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P3.7.log
Scenario: Wrap preset enables horizontal movement across board edge (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/presets/wrap-horizontal.test.ts -t "rook crosses file-a to file-h" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P3.7-wrap.log
```
**Commit**: YES — `feat(chess): add preset rules 10-12 (P3.7)`
- [x] P3.8. **Preset rules 13-15 (meta rules: HP/heal/immunity)**
**What to do**: Implement presets 13-15 from RULES.md — introduce HP/cooldown/immunity attributes in chess schema extensions (within chess package only, not engine). These require adding extended attrs to chess schema (via `extendChessSchema` helper), supporting facts (HP defaults to 1 for FIDE).
**Must NOT do**: leak chess-schema extensions into engine core
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P3.2
**Blocks**: P3.11
**Blocked By**: P2.23, P0.3
**Acceptance Criteria**: registry has 15 entries; HP-aware rules tested
**QA Scenarios**:
```
Scenario: HP preset: captures deal 1 damage; piece with 2 HP survives first hit
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/presets/hp*.test.ts packages/chess/src/presets/heal*.test.ts packages/chess/src/presets/immune*.test.ts 2>&1 | tee /tmp/p38.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P3.8.log
Scenario: Full registry has 15 entries (failure path)
Tool: Bash
Steps: 1. Run: bun -e "import {REGISTRY} from './packages/chess/src/presets/registry.ts'; console.log(Object.keys(REGISTRY).length)"
Expected: stdout "15"
Evidence: .sisyphus/evidence/task-P3.8-count.log
```
**Commit**: YES — `feat(chess): add preset rules 13-15 (P3.8)`
- [x] P3.9. **React + Vite scaffold for chess app**
**What to do**: Wire up Vite + React 19 (or latest) in `packages/chess/`: `index.html`, `src/app/main.tsx`, `src/app/App.tsx` (root, routes: Home, Game, Rules, Save), Vite config with base URL, Tailwind for styling (or CSS modules if Tailwind explicitly disliked by user; default Tailwind). Bundle size placeholder — checked by size-limit later.
**Must NOT do**: install UI libraries beyond React, Tailwind, minimal dnd (react-dnd) if needed — no Material UI, no Ant Design
**Recommended Agent Profile**: `visual-engineering`; Skills: [`interface-design`, `context7`]
**Parallelization**: YES — Wave P3.3 (with P3.10-P3.13 — P3.10 depends on P3.9)
**Blocks**: P3.10-P3.13
**Blocked By**: P2.23
**References**: Vite React guide; Tailwind setup
**Acceptance Criteria**:
- [ ] `cd packages/chess && bun run dev` starts at localhost:5173
- [ ] `bun run build` produces static dist/
- [ ] Playwright can open the home route and see an app root element
**QA Scenarios**:
```
Scenario: Dev server starts; home page renders root
Tool: Playwright
Preconditions: bun run dev started in background on port 5173
Steps:
1. Navigate to http://localhost:5173/
2. Wait for selector '[data-testid="app-root"]'
3. Screenshot
Expected: root element visible
Evidence: .sisyphus/evidence/task-P3.9-home.png
Scenario: Build produces static bundle (failure path for missing build script)
Tool: Bash
Steps: 1. Run: cd packages/chess && bun run build && ls -la dist/
Expected: dist/ contains index.html
Evidence: .sisyphus/evidence/task-P3.9-build.log
```
**Commit**: YES — `feat(chess): scaffold Vite + React app (P3.9)`
- [x] P3.10. **Chessboard component with drag-drop + legal-move highlights**
**What to do**: `packages/chess/src/ui/Board.tsx` — 8×8 grid, piece SVG icons (inline or public/), drag-drop via HTML5 DnD or react-dnd; on drag-start, query engine for that piece's LegalMoves and highlight target squares; on drop, dispatch AttemptedMove fact. Uses `useSession()` hook providing reactive fact subscriptions (implemented via tick-subscription observer on session).
**Must NOT do**: load piece images from external CDN (bundle locally)
**Recommended Agent Profile**: `visual-engineering`; Skills: [`interface-design`]
**Parallelization**: NO (depends on P3.9) — Wave P3.3
**Blocks**: P3.15
**Blocked By**: P3.9
**Acceptance Criteria**:
- [ ] Board renders with 32 pieces in starting position
- [ ] Drag pawn e2→e4: piece moves on board; engine fact updated
- [ ] Illegal move: piece snaps back; no fact change
**QA Scenarios**:
```
Scenario: Legal drag-drop move applied
Tool: Playwright
Steps:
1. Navigate to http://localhost:5173/game
2. Locator '[data-square="e2"]' → dragTo '[data-square="e4"]'
3. Wait selector '[data-square="e4"] [data-piece="white-pawn"]'
4. Screenshot
Expected: pawn on e4
Evidence: .sisyphus/evidence/task-P3.10-e2e4.png
Scenario: Illegal move rejected (failure path)
Tool: Playwright
Steps:
1. Navigate to /game
2. Locator '[data-square="e2"]' dragTo '[data-square="e5"]' (illegal double-plus)
3. Wait selector '[data-square="e2"] [data-piece="white-pawn"]' (pawn still home)
4. Screenshot
Expected: pawn returned
Evidence: .sisyphus/evidence/task-P3.10-reject.png
```
**Commit**: YES — `feat(chess): add interactive Chessboard with drag-drop (P3.10)`
- [x] P3.11. **Rule-toggle screen (preset list + compatibility warnings)**
**What to do**: `packages/chess/src/ui/Rules.tsx` — list all 15 presets with description, toggle switch, compat-warning banner when incompatibility detected; "Apply and start new game" button; toggles only between games (disabled during active game — grayed state).
**Recommended Agent Profile**: `visual-engineering`
**Parallelization**: YES — Wave P3.3
**Blocks**: P3.15
**Blocked By**: P3.4-P3.8, P3.9
**Acceptance Criteria**:
- [ ] 15 toggle rows render; enabling two incompatibles shows warning
- [ ] Starting new game applies enabled presets
**QA Scenarios**:
```
Scenario: Toggle preset, start new game, effect observable
Tool: Playwright
Steps:
1. Navigate to /rules
2. Click '[data-preset="pawns-move-backward"] [data-role="toggle"]'
3. Click '[data-action="start-new-game"]'
4. Navigate to /game
5. Locator '[data-square="e2"]' dragTo '[data-square="e1"]' (backward move; normally illegal)
6. Wait selector '[data-square="e1"] [data-piece="white-pawn"]'
Expected: pawn moved backward
Evidence: .sisyphus/evidence/task-P3.11-back.png
Scenario: Incompatible presets show warning (failure path)
Tool: Playwright
Steps:
1. Navigate to /rules
2. Enable two presets listed as incompatible in RULES.md
3. Expect '[data-testid="compat-warning"]' visible
Expected: warning shown
Evidence: .sisyphus/evidence/task-P3.11-warn.png
```
**Commit**: YES — `feat(chess): add rule-toggle UI with compatibility warnings (P3.11)`
- [x] P3.12. **Save/Load panel + undo via time-travel**
**What to do**: `packages/chess/src/ui/SavePanel.tsx` + undo button in Game view; undo uses time-travel to rewind to previous `Turn`-changed fact boundary (one full move back); save panel lists slots from localStorage (schema-versioned JSON).
**Recommended Agent Profile**: `visual-engineering`
**Parallelization**: YES — Wave P3.3
**Blocks**: P3.14, P3.15
**Blocked By**: P3.3, P3.9
**Acceptance Criteria**:
- [ ] Undo rewinds one full move
- [ ] Save to slot, reload page, load — same position
**QA Scenarios**:
```
Scenario: Undo reverts one move
Tool: Playwright
Steps:
1. Navigate to /game
2. Drag e2→e4; drag e7→e5
3. Click '[data-action="undo"]'
4. Assert '[data-square="e5"] [data-piece]' is NOT black-pawn (reverted)
5. Assert turn indicator shows 'black'
Expected: state reverted
Evidence: .sisyphus/evidence/task-P3.12-undo.png
Scenario: Save/load round-trip (failure path)
Tool: Playwright
Steps:
1. Play 4 moves
2. Click '[data-action="save"]' into slot "test"
3. page.reload()
4. Click '[data-action="load"]' slot "test"
5. Assert board state matches pre-reload
Evidence: .sisyphus/evidence/task-P3.12-saveload.png
```
**Commit**: YES — `feat(chess): add Save/Load panel + time-travel undo (P3.12)`
- [x] P3.13. **JSON export/import + validation**
**What to do**: `packages/chess/src/ui/ImportExport.tsx` + `packages/chess/src/persist/io.ts` — export button produces a downloadable JSON file (schema: `{ version: 1, rules: [...], facts: [...] }`); import button accepts file, validates against schema (via `@paratype/rete`'s exported schema + chess extension schema), applies rules + facts.
**Must NOT do**: allow importing from untrusted URL (file-upload only)
**Recommended Agent Profile**: `visual-engineering`
**Parallelization**: YES — Wave P3.3
**Blocks**: P3.15
**Blocked By**: P1.6, P3.9
**Acceptance Criteria**:
- [ ] Export downloads valid JSON parseable by the importer
- [ ] Invalid JSON shows user-facing error, no crash
**QA Scenarios**:
```
Scenario: Export then re-import round-trip
Tool: Playwright
Steps:
1. Navigate to /game; make 3 moves
2. Click '[data-action="export"]'; Playwright captures download as /tmp/export.json
3. Click '[data-action="import"]'; upload /tmp/export.json
4. Assert board state matches pre-import
Evidence: .sisyphus/evidence/task-P3.13-export.json, .sisyphus/evidence/task-P3.13-import.png
Scenario: Malformed JSON rejected with user message (failure path)
Tool: Playwright
Steps:
1. Click '[data-action="import"]'; upload fixture with `{"bad":"data"}`
2. Assert '[data-testid="import-error"]' visible with descriptive message
Evidence: .sisyphus/evidence/task-P3.13-bad.png
```
**Commit**: YES — `feat(chess): add JSON export/import with validation (P3.13)`
- [x] P3.14. **localStorage auto-save + restore**
**What to do**: `packages/chess/src/persist/autosave.ts` — subscribe to session tick end; on every turn boundary, write serialized state + event log to localStorage key `paratype-chess:v1:autosave`. On app load, if key present, restore via `replayFromLog`. Include schema version in payload.
**Must NOT do**: write on every tick (too noisy); write to sessionStorage (lost on close)
**Recommended Agent Profile**: `unspecified-high`
**Parallelization**: YES — Wave P3.4 (with P3.15)
**Blocks**: P3.15
**Blocked By**: P3.3, P3.12
**Acceptance Criteria**:
- [ ] After 3 moves, localStorage has `paratype-chess:v1:autosave`
- [ ] Reload page → game resumes in same position
**QA Scenarios**:
```
Scenario: Autosave persists across reload
Tool: Playwright
Steps:
1. Navigate to /game; play 5 moves
2. localStorage.getItem('paratype-chess:v1:autosave') not null
3. Reload
4. Assert board state matches
Evidence: .sisyphus/evidence/task-P3.14-autosave.png
Scenario: Schema version mismatch discards silently (failure path)
Tool: Playwright
Steps:
1. Set localStorage to stale payload with version 0
2. Reload
3. Assert new game started (no crash)
Evidence: .sisyphus/evidence/task-P3.14-stale.png
```
**Commit**: YES — `feat(chess): add localStorage auto-save and restore (P3.14)`
- [x] P3.15. **End-to-end UI scenario (gate)**
**What to do**: Playwright scenario at `packages/chess/e2e/full-flow.spec.ts` — open app → toggle 2 presets → start game → play 5 moves → save → reload → game restored → export → import in fresh context → play 3 more moves → undo → play until checkmate (scripted sequence) → assert Game Over banner.
**Must NOT do**: use timing-based waits (`waitForTimeout` is banned; use selector waits)
**Recommended Agent Profile**: `unspecified-high`; Skills: [`playwright`]
**Parallelization**: NO — Wave P3.4 (gate)
**Blocks**: Phase 4
**Blocked By**: P3.1-P3.14
**Acceptance Criteria**:
- [ ] `bun x playwright test packages/chess/e2e/full-flow.spec.ts` green
- [ ] Video + trace artifacts captured
- [ ] Phase 3 tag: `git tag v0.3.0-phase3`
**QA Scenarios**:
```
Scenario: Full flow end-to-end
Tool: Playwright
Preconditions: bun run dev serving packages/chess
Steps: (executed by the spec file; evidence is trace + video)
Expected Result: spec passes; video shows full flow
Evidence: .sisyphus/evidence/task-P3.15-full-flow.webm, .sisyphus/evidence/task-P3.15-trace.zip
Scenario: Phase 3 tag present
Tool: Bash
Steps: 1. Run: git tag v0.3.0-phase3 && git tag | grep v0.3.0-phase3
Expected: present
Evidence: .sisyphus/evidence/task-P3.15-tag.log
```
**Commit**: YES — `test(chess): e2e full-flow scenario; tag Phase 3 (P3.15)`; post-commit: `git tag v0.3.0-phase3`
### Phase 4 — Authoritative Multiplayer
- [x] P4.1. **Bun HTTP+WS server scaffold + config**
**What to do**: `packages/server/src/index.ts` — `Bun.serve({ port, fetch, websocket: { open, message, close } })`; env-driven port (default 7357); health endpoint `GET /healthz` returning `{ ok: true, version }`; structured pino logger with request id; graceful shutdown on SIGINT.
**Recommended Agent Profile**: `unspecified-high`; Skills: [`context7`]
**Parallelization**: YES — Wave P4.1 (with P4.2-P4.4)
**Blocks**: P4.5-P4.11
**Blocked By**: P3.15
**References**: Bun.serve docs, pino
**Acceptance Criteria**:
- [ ] `bun run packages/server/src/index.ts` starts; `curl localhost:7357/healthz` returns 200
- [ ] Logs emit JSON lines
**QA Scenarios**:
```
Scenario: Server responds to health check
Tool: Bash
Steps:
1. Run: bun run packages/server/src/index.ts &
2. Sleep 2
3. Run: curl -sS -o /tmp/p41.json -w "%{http_code}" http://localhost:7357/healthz
4. Kill %1
Expected: status 200; body has {"ok":true}
Evidence: .sisyphus/evidence/task-P4.1-health.log
Scenario: SIGINT shuts down gracefully (failure path)
Tool: Bash
Steps:
1. Run: bun run packages/server/src/index.ts &
2. SIGINT; wait; echo $?
Expected: exit 0
Evidence: .sisyphus/evidence/task-P4.1-shutdown.log
```
**Commit**: YES — `feat(server): scaffold Bun HTTP+WS server with health + logging (P4.1)`
- [x] P4.2. **Message schemas + validation (TDD)**
**What to do**: `packages/server/src/protocol.ts` — zod schemas per PROTOCOL.md message type; `validateMessage(raw): Result`; top-level `v` version check; round-trip tested.
**Must NOT do**: use JSON.parse without validation
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P4.1
**Blocks**: P4.5, P4.6
**Blocked By**: P0.4 (PROTOCOL.md), P3.15
**Acceptance Criteria**:
- [ ] `bun test packages/server/src/protocol.test.ts` green
- [ ] Coverage ≥ 95%
**QA Scenarios**:
```
Scenario: All 8+ message types round-trip
Tool: Bash
Steps: 1. Run: bun test packages/server/src/protocol.test.ts 2>&1 | tee /tmp/p42.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P4.2-proto.log
Scenario: Malformed message rejected (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/server/src/protocol.test.ts -t "invalid v rejected" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P4.2-bad.log
```
**Commit**: YES — `feat(server): add protocol schemas + validation (P4.2)`
- [x] P4.3. **Room model (create/join/leave, 6-char codes)**
**What to do**: `packages/server/src/rooms.ts` — `class RoomRegistry` with `createRoom()` → 6-char [A-Z0-9] code + uuid-v4 token; `joinRoom(code, token)`; 2-player max; token-authenticated per message; TDD.
**Must NOT do**: persist across restart (v1 constraint)
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P4.1
**Blocks**: P4.5
**Blocked By**: P3.15
**Acceptance Criteria**:
- [ ] `bun test packages/server/src/rooms.test.ts` green
- [ ] Code generation uniqueness fuzz (1000 codes, 0 collisions expected)
**QA Scenarios**:
```
Scenario: Room create, join, duplicate-join-rejected
Tool: Bash
Steps: 1. Run: bun test packages/server/src/rooms.test.ts 2>&1 | tee /tmp/p43.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P4.3-rooms.log
Scenario: Third player rejected (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/server/src/rooms.test.ts -t "third join rejected" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P4.3-third.log
```
**Commit**: YES — `feat(server): add room registry with codes + tokens (P4.3)`
- [x] P4.4. **Rate limiting + origin allow-list + 64KB cap**
**What to do**: `packages/server/src/middleware.ts` — per-connection token bucket (100 msg/sec, burst 20); WebSocket upgrade rejects non-allow-list origins (configurable via env `ALLOWED_ORIGINS`); reject payloads > 64KB with disconnect.
**Recommended Agent Profile**: `unspecified-high`
**Parallelization**: YES — Wave P4.1
**Blocks**: P4.12
**Blocked By**: P3.15
**Acceptance Criteria**:
- [ ] `bun test packages/server/src/middleware.test.ts` green
- [ ] Stress test: 200 msg/sec triggers RATE_LIMIT disconnect
**QA Scenarios**:
```
Scenario: Rate-limit trips on over-limit
Tool: Bash
Steps: 1. Run: bun test packages/server/src/middleware.test.ts 2>&1 | tee /tmp/p44.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P4.4-rl.log
Scenario: Origin disallowed rejected (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/server/src/middleware.test.ts -t "origin not in allow-list rejected" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P4.4-origin.log
```
**Commit**: YES — `feat(server): add rate-limit, origin allow-list, message-size cap (P4.4)`
- [x] P4.5. **Authoritative session per room**
**What to do**: `packages/server/src/game-session.ts` — each room holds a `Session` from `@paratype/rete` + chess rules; server is the only one that calls `insert/retract/fireRules`. Fact IDs minted here only.
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P4.2
**Blocks**: P4.6, P4.12
**Blocked By**: P4.1, P4.2, P4.3
**Acceptance Criteria**:
- [ ] `bun test packages/server/src/game-session.test.ts` green
**QA Scenarios**:
```
Scenario: Each room has isolated session state
Tool: Bash
Steps: 1. Run: bun test packages/server/src/game-session.test.ts 2>&1 | tee /tmp/p45.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P4.5-sess.log
Scenario: Fact IDs do not collide across rooms (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/server/src/game-session.test.ts -t "room fact ids distinct" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P4.5-ids.log
```
**Commit**: YES — `feat(server): add authoritative game session per room (P4.5)`
- [x] P4.6. **Move-intent validation + fact-delta broadcast**
**What to do**: `packages/server/src/broadcast.ts` — on `game.move` intent: insert `AttemptedMove` fact; fire rules; diff pre/post WM; broadcast `game.delta` with added/removed facts to both clients. Assigned `seq` per delta for reconnection.
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P4.2
**Blocks**: P4.12
**Blocked By**: P4.5
**Acceptance Criteria**:
- [ ] Integration test: send legal move → both clients receive delta with updated Position fact
- [ ] Illegal move → `error` message; no broadcast
**QA Scenarios**:
```
Scenario: Legal move broadcast to both clients
Tool: Bash (WS client script)
Steps:
1. Launch server
2. Run: bun run scripts/ws-client.ts --script fixtures/two-client-legal-move.json
Expected: both clients receive matching game.delta with Position change
Evidence: .sisyphus/evidence/task-P4.6-delta.json
Scenario: Illegal move rejected; no broadcast (failure path)
Tool: Bash
Steps: 1. Run: bun run scripts/ws-client.ts --script fixtures/illegal-move.json
Expected: error to sender; zero delta messages
Evidence: .sisyphus/evidence/task-P4.6-illegal.json
```
**Commit**: YES — `feat(server): add move validation + fact-delta broadcast (P4.6)`
- [x] P4.7. **Reconnection flow (60s window, snapshot resume)**
**What to do**: `packages/server/src/reconnect.ts` — on disconnect, start 60s timer; during grace, incoming (code, token) matches → resume and send `game.state` (full snapshot) + all deltas since client's last `seq`. After 60s, room aborts with `game.end` broadcast to remaining client.
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P4.2
**Blocks**: P4.12
**Blocked By**: P4.5, P4.6, P3.3 (replay for determinism)
**Acceptance Criteria**:
- [ ] Integration test: disconnect, reconnect within 30s, resume state exactly
**QA Scenarios**:
```
Scenario: Reconnect within grace resumes game
Tool: Bash
Steps: 1. Run: bun run scripts/ws-client.ts --script fixtures/reconnect-within-grace.json
Expected: client B reconnects, receives state, continues game
Evidence: .sisyphus/evidence/task-P4.7-reconnect.json
Scenario: Reconnect after grace fails with game.end (failure path)
Tool: Bash
Steps: 1. Run: bun run scripts/ws-client.ts --script fixtures/reconnect-after-grace.json
Expected: rejected; remaining client received game.end
Evidence: .sisyphus/evidence/task-P4.7-expired.json
```
**Commit**: YES — `feat(server): add reconnection with 60s grace + snapshot resume (P4.7)`
- [x] P4.8. **Structured logging + metrics**
**What to do**: `packages/server/src/logging.ts` — pino logger with request-scoped `roomId`, `clientId`, `seq`; per-tick duration metric; `/metrics` endpoint (Prometheus text format) with counters: `rooms_active`, `messages_received_total`, `moves_validated_total{result}`, tick duration histogram.
**Recommended Agent Profile**: `unspecified-high`
**Parallelization**: YES — Wave P4.2
**Blocks**: P4.12
**Blocked By**: P4.1
**Acceptance Criteria**:
- [ ] `curl localhost:7357/metrics` returns text/plain with expected series
**QA Scenarios**:
```
Scenario: Metrics endpoint exposes required series
Tool: Bash
Steps:
1. Run: bun run packages/server/src/index.ts &
2. Sleep 2
3. Run: curl -sS http://localhost:7357/metrics | grep -E 'rooms_active|messages_received_total|moves_validated_total|tick_duration'
4. Kill %1
Expected: all 4 series present
Evidence: .sisyphus/evidence/task-P4.8-metrics.log
Scenario: Log lines are valid JSON (failure path)
Tool: Bash
Steps:
1. Run: bun run packages/server/src/index.ts 2>&1 | head -20 | jq -e .
Expected: exit 0 for each line (jq parses)
Evidence: .sisyphus/evidence/task-P4.8-logs.log
```
**Commit**: YES — `feat(server): add pino logging and Prometheus metrics (P4.8)`
- [x] P4.9. **WebSocket client library with reconnect + seq ack**
**What to do**: `packages/chess/src/net/client.ts` — `class GameClient` with `connect(code, token)`, exponential-backoff reconnect, sequence-ack tracking, event emitter for `game.state`, `game.delta`, `error`. Client owns a local engine session but only applies deltas received from server (no self-validation of moves).
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P4.3 (with P4.10, P4.11)
**Blocks**: P4.12
**Blocked By**: P4.2 (protocol schemas)
**Acceptance Criteria**:
- [ ] `bun test packages/chess/src/net/client.test.ts` green
- [ ] Reconnect after drop succeeds within 30s
**QA Scenarios**:
```
Scenario: Client handshake + delta application
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/net/client.test.ts 2>&1 | tee /tmp/p49.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P4.9-client.log
Scenario: Reconnect after forced disconnect (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/net/client.test.ts -t "reconnect restores state" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P4.9-recon.log
```
**Commit**: YES — `feat(chess): add WebSocket client library with reconnect (P4.9)`
- [x] P4.10. **Client prediction + server reconciliation**
**What to do**: `packages/chess/src/net/prediction.ts` — on user drag-drop, client locally applies move optimistically to engine session; sends intent to server; on `game.delta`, reconciles (replaces predicted state with authoritative state). On `error` response, rolls back.
**Must NOT do**: drift — always re-hash local state against server snapshot on receipt; mismatch → resync from server full state
**Recommended Agent Profile**: `deep`
**Parallelization**: YES — Wave P4.3
**Blocks**: P4.12
**Blocked By**: P4.9
**Acceptance Criteria**:
- [ ] `bun test packages/chess/src/net/prediction.test.ts` green
- [ ] Simulated latency (100ms artificial delay) doesn't cause desync
**QA Scenarios**:
```
Scenario: Optimistic prediction matches authoritative result
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/net/prediction.test.ts 2>&1 | tee /tmp/p410.log
Expected: all pass
Evidence: .sisyphus/evidence/task-P4.10-pred.log
Scenario: Rejected prediction rolls back (failure path)
Tool: Bash
Steps: 1. Run: bun test packages/chess/src/net/prediction.test.ts -t "rejected intent rolls back" 2>&1
Expected: pass
Evidence: .sisyphus/evidence/task-P4.10-rollback.log
```
**Commit**: YES — `feat(chess): add client prediction + server reconciliation (P4.10)`
- [x] P4.11. **Room lobby UI (create/join screens)**
**What to do**: `packages/chess/src/ui/Lobby.tsx` — home route with two buttons: "Create Room" (shows generated code, share link) and "Join Room" (input for code). After join, redirect to `/game` with active session.
**Recommended Agent Profile**: `visual-engineering`; Skills: [`interface-design`]
**Parallelization**: YES — Wave P4.3
**Blocks**: P4.12
**Blocked By**: P4.9
**Acceptance Criteria**:
- [ ] Playwright: create room in ctx A, join in ctx B, both see game
- [ ] Invalid code shows error
**QA Scenarios**:
```
Scenario: Two contexts join same room
Tool: Playwright
Steps:
1. Context A navigates /; clicks [data-action="create-room"]; notes [data-testid="room-code"] value (CODE)
2. Context B navigates /; types CODE in [data-testid="room-code-input"]; clicks [data-action="join-room"]
3. Both reach /game; both see starting position
Expected: both boards render
Evidence: .sisyphus/evidence/task-P4.11-create.png, .sisyphus/evidence/task-P4.11-join.png
Scenario: Invalid code errors (failure path)
Tool: Playwright
Steps:
1. Navigate /; type "XXXXXX"; click join
2. Assert [data-testid="lobby-error"] visible
Evidence: .sisyphus/evidence/task-P4.11-bad.png
```
**Commit**: YES — `feat(chess): add lobby UI for create/join rooms (P4.11)`
- [x] P4.12. **E2E multiplayer scenario (Phase 4 gate)**
**What to do**: `packages/chess/e2e/multiplayer.spec.ts` — launches server + client (via Playwright webServer config); two contexts create/join room, play 10-move game alternating sides; ctx A disconnects at move 6, reconnects at move 7; game completes to checkmate; assert both clients see identical final state.
**Must NOT do**: use fixed sleeps; use selector waits
**Recommended Agent Profile**: `unspecified-high`; Skills: [`playwright`]
**Parallelization**: NO — Wave P4.4 (gate)
**Blocks**: Final Wave
**Blocked By**: P4.1-P4.11
**Acceptance Criteria**:
- [ ] `bun x playwright test packages/chess/e2e/multiplayer.spec.ts` green
- [ ] Phase 4 tag: `git tag v0.4.0-phase4`
**QA Scenarios**:
```
Scenario: Two-browser full multiplayer game with mid-game reconnect
Tool: Playwright (see spec)
Expected: spec passes; video captured
Evidence: .sisyphus/evidence/task-P4.12-mp.webm, .sisyphus/evidence/task-P4.12-trace.zip
Scenario: Phase 4 tag present
Tool: Bash
Steps: 1. Run: git tag v0.4.0-phase4 && git tag | grep v0.4.0-phase4
Expected: present
Evidence: .sisyphus/evidence/task-P4.12-tag.log
```
**Commit**: YES — `test(root): E2E multiplayer with reconnect; tag Phase 4 (P4.12)`; post-commit: `git tag v0.4.0-phase4`
---
## Final Verification Wave (MANDATORY — after ALL implementation tasks)
> 4 review agents run in PARALLEL. ALL must APPROVE. Present consolidated results to user and get explicit "okay" before marking work complete.
> **Do NOT auto-proceed after verification. Wait for user's explicit approval.**
> **Never mark F1-F4 as checked before getting user's okay.** Rejection or user feedback → fix → re-run → present again → wait for okay.
- [x] F1. **Plan Compliance Audit** — `oracle`
Read this plan end-to-end. For each "Must Have": verify implementation exists (read file, run command, inspect built artifact). For each "Must NOT Have": search codebase for forbidden patterns (e.g., `grep -r "as any" packages/rete/src`), reject with file:line if found. Check evidence files exist in `.sisyphus/evidence/`. Verify all 5 phase tags exist (`git tag | grep phase`). Compare deliverables against plan.
Output: `Must Have [N/N] | Must NOT Have [N/N] | Phase tags [5/5] | Tasks [N/N] | VERDICT: APPROVE/REJECT`
- [x] F2. **Code Quality Review** — `unspecified-high`
Run `bun run typecheck` + `bun run lint` + `bun run test:coverage` + `bun run size-limit`. Review all changed files for: `as any` / `@ts-ignore` / `@ts-expect-error`, empty catches, `console.log` in prod code, commented-out code, unused imports, `Date.now()`/`Math.random()` in engine RHS paths, raw `Set<object>` iteration in engine hot paths. Check AI slop: excessive comments, over-abstraction, generic names (data/result/item/temp/obj). Audit bundle sizes against budgets (engine < 50KB min+gz, chess < 200KB min+gz).
Output: `Build [PASS/FAIL] | Lint [PASS/FAIL] | Tests [N pass/N fail, coverage X%/Y%/Z%] | Bundle [engine Xkb / chess Ykb] | Files [N clean/N issues] | VERDICT`
- [x] F3. **Real Manual QA via Playwright + Scripted Clients** — `unspecified-high` (+ `playwright` skill)
Start from clean state: `rm -rf node_modules && bun install && bun run build`. Launch chess server. Execute EVERY QA scenario from EVERY task — follow exact steps, capture evidence. Test cross-task integration: play a full FIDE game; toggle 3 presets between games; play a custom-rules game; save via localStorage; reload browser; verify state persisted; export JSON; import into fresh browser; play a multiplayer game across two browser contexts with reconnect mid-game. Test edge cases: illegal move rejected, rate-limit trip, protocol version mismatch hard-disconnect, 60s reconnect boundary. Save to `.sisyphus/evidence/final-qa/`.
Output: `Scenarios [N/N pass] | Integration [N/N] | Edge Cases [N tested] | VERDICT`
- [x] F4. **Scope Fidelity Check** — `deep`
For each task: read "What to do", read actual diff (`git log` / `git diff` on that task's commits). Verify 1:1 — everything in spec was built (no missing), nothing beyond spec was built (no creep). Check "Must NOT do" compliance in diff. Detect cross-task contamination: Task N touching Task M's files. Flag unaccounted changes. Verify commit messages follow Conventional Commits with scope (`feat(rete):`, `feat(chess):`, `feat(server):`).
Output: `Tasks [N/N compliant] | Contamination [CLEAN/N issues] | Unaccounted [CLEAN/N files] | Commit format [N/N] | VERDICT`
---
## Commit Strategy
- **Conventional Commits** enforced: `type(scope): description` where `scope ∈ {rete, chess, server, root}`
- **Types**: `feat`, `fix`, `test`, `refactor`, `chore`, `docs`, `perf`, `build`, `ci`
- **Atomic commits**: one logical change per commit. TDD tasks commit test+impl together.
- **Every commit**: passes `bun run check` (tsc + eslint + vitest) — enforced via pre-commit hook AND CI required-status-check
- **Phase boundaries tagged**: `v0.1.0-phase1`, `v0.2.0-phase2`, `v0.3.0-phase3`, `v0.4.0-phase4`, `v1.0.0` (final)
- **No WIP commits on main**; feature work in feature branches (if branching used) or linearly via rebase on main
- **No squash-merge across phases**; each phase is a merge train
Per-task commit details live in each TODO's `Commit:` block.
---
## Success Criteria
### Verification Commands (run from repo root)
```bash
bun install # → 0 errors
bun run typecheck # → 0 errors
bun run lint # → 0 errors
bun run test # → all green
bun run test:coverage # → engine ≥90%, chess ≥70%, server ≥80%
bun run build # → dist/ populated in all 3 packages
bun run size-limit # → engine < 50KB, chess < 200KB
bun run playwright test # → all E2E pass
bun run scripts/replay-determinism.ts fixtures/game-*.log # → hashes match for every fixture
bun run start:server & # server up
sleep 2
bun run test:integration # WebSocket handshake, move exchange, reconnect
kill %1
gh run list --limit 1 --json conclusion -q '.[0].conclusion' # → "success"
git tag --list # → contains v0.1.0-phase1 … v1.0.0
```
### Final Checklist
- [ ] All "Must Have" present (verified by F1)
- [ ] All "Must NOT Have" absent (verified by F1 and F2)
- [ ] All phase tags present (v0.1.0-phase1 … v1.0.0)
- [ ] Engine coverage ≥90% / chess ≥70% / server ≥80%
- [ ] Bundle sizes within budget (engine <50KB, chess <200KB)
- [ ] Playwright scenarios all green
- [ ] Server integration tests all green
- [ ] Replay-determinism hash match 100%
- [ ] CI green on latest commit
- [ ] User has given explicit approval after F1-F4 presentation