wiki-vault/Sources/Dev/2026-06-04-2026-06-03-petal-dispatch-live-experiments.md

268 lines
20 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
created: '2026-06-04'
path: Sources/Dev
project: execution-journal
tags:
- petal-dispatch
- harness
- lovebug
- claude-code
- experiments
- execution-journal
- sse
type: session-notes
---
> **Window:** 2026-06-03 22:14Z 23:07Z (~53 min)
> **Host:** herbydev
> **Dispatch session UUID:** dba5750f-e07f-4dba-8ce2-8f61b80842eb
> **Active chain:** 134c053b-2774-4255-8750-ab91ec91b675.jsonl
> **Raw SSE capture:** /tmp/sse_events.jsonl (host-local — not persisted)
**Workers in window (4):**
| session_id | uuid | title | status |
|---|---|---|---|
| task_cffef42d7c9f | bbe286d0-f34d-4794-b96c-d08c3220903a | harness terminal-event probe | crashed (SIGTERM, exit_code 143) |
| task_81c483b3703c | 6fca83fe-2dce-4c9f-add5-5067a1ab8216 | three-turn streaming demo | idle (3 turns) |
| task_aa44afb48ca6 | 55023fdb-e7db-493f-90a4-1e6e37c3408c | three-turn sequential demo | idle (3 turns, gated on read_transcript) |
| task_4bd577a6d91e | 96538513-8d1a-4108-a89f-dd9bec068961 | dependent-chain demo | idle (3 dependent turns: noun → haiku → Latin) |
---
# petal-dispatch — live experiment session, 2026-06-03
Travis used the petal-dispatch web UI (the v0.5 prototype) to drive a series of experiments against the Dispatch claude. The session was deliberately exploratory: validate the harness emits the event types the design assumes, then poke at multi-turn worker semantics and at how Dispatch-claude streaming actually feels through the UI. A second session (Lovebug-side) was running in parallel and supplying critical feedback on findings, which Travis forwarded into this Dispatch chat verbatim.
## 1. Experiments run
### 1.1 Harness terminal-event probe ("does idle ≠ exit?")
Travis pasted the spec from his other session:
> In the live PoC, spawn a worker, let it finish, and capture the raw event stream the harness receives. Specifically: does an IDLE (or equivalent) event fire when the worker exits/completes, distinct from any event tied to the worker choosing to emit a message? Dump the raw stream-json around a worker's completion — show the event type, whether it carries the worker's session_id, and the timestamp relative to the worker process actually ending.
Dispatch's plan, executed in-order:
1. Read `src/petal_dispatch/tasks.py` and `harness_routes.py` to find which call sites emit which event kinds.
2. Confirm `/events/stream` is reachable (`curl` against `http://127.0.0.1:8001/events/stream`).
3. Drop a python SSE-capture script at `/tmp/sse_capture.py` that records each event with wall-clock + monotonic timestamps.
4. Spawn worker `task_cffef42d7c9f` ("harness terminal-event probe") with prompt `Reply with exactly: DONE-CAPTURE. Nothing else.`.
5. Let it complete its single turn; observe spawn/user_message/message/idle on the SSE stream.
6. SIGTERM the worker subprocess to force the terminal-event path; capture the resulting `task.crashed`.
7. Diff wall-clock kill time vs. crash-event timestamp.
This was the most data-dense experiment of the session. Output captured at `/tmp/sse_events.jsonl` (5 events).
### 1.2 Three-turn streaming demo (fire-and-forget queueing)
Worker `task_81c483b3703c` ("three-turn streaming demo"). Travis prompted:
> task a worker with these questions, in this order: turn 1 pick your favorite color, turn 2 guess at what my favorite color is, turn 3 give me 5 lines why in programming languages python is better than c.
Dispatch called `start_task` with turn 1 and then immediately `send_message` for turns 2 and 3 with no gating — three stdin writes back-to-back before the worker had answered any of them. Worker processed them sequentially from its stdin buffer (the pipe queued at the OS layer). Replies: Blue / Green / 5-line Python advocacy.
### 1.3 Three-turn sequential demo (read-as-gate)
Worker `task_aa44afb48ca6` ("three-turn sequential demo"). Travis:
> ok try this execute the turns 2 and 3 only after getting the response from 1 and 2 respectively.
Dispatch used the `read_transcript format=auto` MCP tool as a serializer — it blocks on the worker's `_turn_done` event before returning. Pattern: `send_message → read_transcript → send_message → read_transcript`. Same final transcript content (Blue / Green / 5-line Python), but on the harness side the three `task.idle` events landed spread out in time instead of bunched at the end.
### 1.4 Dependent-chain demo (true cross-turn dependency)
Worker `task_4bd577a6d91e` ("dependent-chain demo"). Goal: prove the orchestrator-side dependency loop works, where each prompt is *built from* the previous reply, not just sent after it. Dispatch picked the example:
> pick a noun → write a haiku about it → translate its last line to Latin
Trace:
- Turn 1 prompt: pick any single noun. Reply: `Lantern`.
- Turn 2 prompt (built from turn-1 reply): write a 3-line haiku about "Lantern". Reply:
> Paper skin glowing
> A small flame remembers night
> Wind learns how to breathe
- Turn 3 prompt (built from turn-2 reply): translate the last line — "Wind learns how to breathe" — into Latin. Reply: `Ventus spirare discit`.
Dispatch noted explicitly: the worker had no way to interpret "the last line of that haiku" — the orchestrator did that extraction and re-quoted it into the next prompt.
### 1.5 Streaming-UX probes (interspersed conversation, no worker)
Several turns were not experiments per se but back-and-forth where Travis probed why Dispatch's own replies sometimes felt instant and sometimes delayed by ~15s. These ended up surfacing the streaming-vs-agentic-loop distinction (see §3).
## 2. Findings (verbatim where possible)
### 2.1 Five-event signature around worker lifecycle
From the terminal-event probe, the harness emitted these in this order:
| seq | mono ts | kind | payload (trimmed) |
|---|---|---|---|
| 21 | 518824.2655 | `task.user_message` | `{turn_index:1, message_preview:"Reply with…"}` |
| 22 | 518824.2673 | `task.spawned` | `{cwd, uuid, title}` |
| 23 | 518826.7635 | `task.message` | `{turn_index:1, preview:"DONE-CAPTURE", full_message_path}` |
| 24 | 518826.8250 | `task.idle` | `{last_turn_index:1}` |
| 25 | 518853.4358 | `task.crashed` | `{exit_code:143, last_error:"DONE-CAPTURE", duration_seconds:29.356}` |
Every event carries `source: "task_cffef42d7c9f"` (the worker's session_id) and a monotonic `seq`.
### 2.2 idle ≠ exit (the headline finding)
Dispatch's verbatim conclusion:
> Yes — two separate events back-to-back:
> - `task.message` (seq 23, +0ms baseline) — fires when the worker emits an assistant turn (stream-json `assistant` line).
> - `task.idle` (seq 24, +61.5ms after the message) — fires on the stream-json `result` line, i.e. turn boundary, while the subprocess is still alive.
>
> They're emitted from different branches in `tasks.py:278-298`. `task.idle` does not carry the message body, only `last_turn_index`. Both carry session_id via the harness `source` field.
And distinguishing two senses of "exits/completes":
> 1. **Turn complete, process still alive** → `task.idle`. This is the steady state — a v0.5 worker stays running between turns because `--print --input-format stream-json` keeps stdin open. You'll see N `task.idle` events over a worker's life.
> 2. **Process actually exits** → `task.done` (rc 0) or `task.crashed` (anything else). Fires from `_emit_terminal_event` in the stdout-reader's `finally` block, after the OS reaps the process. Exactly one per worker lifetime.
### 2.3 Three latent bugs flagged in passing
Dispatch's own words (these are the findings Travis explicitly wanted captured for the other session):
**Spawn / user_message ordering race:**
> `task.spawned` arrived *after* `task.user_message` (1.8ms gap). Both are scheduled via `asyncio.create_task` without ordering guarantees (`tasks.py:_fire`), and `start_task` calls `start_fresh()` (spawn) then `send()` (user_message) on the same tick. Consumers can't rely on spawn being seq-earlier than the first user_message. Probably fine, but if anything downstream assumes "spawn is always first," it's a latent bug.
**`last_error` misleading on SIGTERM:**
> `task.crashed.last_error` was populated with `"DONE-CAPTURE"` — the assistant's own reply, because `_last_error_from_tail` picks the last text block when no `result.error` is present. On a clean SIGTERM that's misleading; the field name implies a fault.
**`task.done` is unreachable in practice:**
> There is no `task.done` path covered in this capture (would require exit code 0). The v0.5 worker has no natural exit, so `task.done` only fires via the unrouted `manager.kill()` followed by a clean wait — currently unreachable from outside the process.
### 2.4 Implication for constitution principle #5
Travis forwarded the other session's framing — that principle #5 (the completion signal must derive from process state, not tool choice) is violated because process-exit is architecturally unreachable in v0.5. Dispatch's response:
> Honest answer: yes for principle #5 specifically, but it's a different architecture with its own tradeoffs, not a free win. … One-shot fits the "concurrency-only, harness-observed lifecycle" framing better than long-lived. Principle #5 rewrites itself: *completion = subprocess reap, observed by the harness, period.* No conditional sentinels, no /complete endpoint, no flaky-circuit worry.
Tradeoffs Dispatch laid out for one-shot:
1. Startup cost per turn (~13s claude CLI cold-start, MCP wiring).
2. "Worker" stops meaning a long-lived process; session_id ↔ conversation UUID on disk, "running" only describes the brief spawn-to-exit window.
3. No in-memory state between turns (probably fine — `--resume` rehydrates from disk).
4. Task-complete vs. turn-complete distinction moves up to the orchestrator — every reply is a `task.done`, and whether the *task* (user intent) is done becomes Lovebug's call. Dispatch noted this is "exactly the synthesis-layer hook Lovebug's instance-2 wanted, just expressed structurally instead of via a sentinel."
### 2.5 Agreed v0.6 shape (Travis confirmed)
After the back-and-forth, Travis confirmed:
> <10% follow-up + 30-60s tolerance means one-shot default is the right base — the 2-4s cold start is well inside the budget, and you're not paying it 90% of the time anyway.
The shape they aligned on, to be implemented when the "real tool" is built:
- **Default:** spawn-per-turn, `task.done` on real process exit, principle #5 satisfied literally.
- **Toggle (off by default):** keep worker alive between turns + add a `/complete` (or stdin-EOF) path so `task.done` is still reachable cleanly.
- **UI:** per-worker status pill that distinguishes spawn/run/idle states visually so users don't read 4s of nothing as "broken."
### 2.6 `send_message` queueing semantics (incidental finding from §1.2/1.3)
Dispatch confirmed (with code-level specificity):
> `send_message` writes to the worker's stdin immediately and returns `{ok:true}`. If the worker is mid-turn, your message sits in the OS pipe buffer and gets picked up after the current `result` line. So yes, queuing happens — but at the pipe level, silently, with no visibility.
Three options Dispatch sketched for making this explicit:
1. `send_message` waits for idle before writing stdin (one-liner: `await self._turn_done.wait()` at top of `Worker.send`). Honest, but the MCP call can block 3060s on a busy worker.
2. `send_message` refuses if busy, returns `{ok:false, error:"busy"}`. Pushes retry logic onto callers.
3. Add a `wait_for_idle` tool, keep `send_message` fire-and-forget.
Dispatch's lean: **option 3** — "Keeps `send_message` honest (it does one thing, fast), and gives callers an explicit 'I'm choosing to wait' verb that maps to the harness `task.idle` event." Composes across workers too.
## 3. Streaming-UX findings (Dispatch self-observation)
Travis noticed Dispatch sometimes felt instant and sometimes delayed by 15s. Dispatch's diagnosis, verbatim:
> Streaming is on the whole time. What you're feeling is **time-to-first-token**, and it has two main sources, neither of which I control:
>
> 1. **Extended thinking.** For harder prompts … the model engages thinking blocks before any visible text. Those don't render in the UI — they look like dead air, then text starts streaming at full speed.
> 2. **Input processing.** Longer prompt = longer prefill before the first output token.
UX recommendation (Dispatch):
> A "thinking…" indicator that flips on the moment a request is sent and flips off when the first real token arrives covers all three cases (worker spawn, model thinking, prefill) with one mechanism.
And the related streaming-vs-agentic-loop clarification:
> Streaming and the agentic loop are orthogonal: streaming is "within a turn," the loop is "across turns." … What looked like "Claude posting in parts" was actually 6-7 separate turns chained together by tool results.
## 4. Open questions / next moves
- **Decision on principle #5 wording.** Two options framed in §2.4. Travis leaned one-shot-default with opt-in long-lived. Not yet written into the constitution / plan document.
- **`task.done` reachability path.** If long-lived workers stay supported even as a toggle, design the `/complete` endpoint (or stdin-EOF) so the terminal-event path is deliberately reachable, not only via SIGTERM.
- **`last_error` semantics on SIGTERM.** Distinguish "I was killed cleanly" from "I crashed with an error." Today `_last_error_from_tail` falls back to the last assistant text block, which is misleading. Cheap fix: populate `last_error` only when there's an actual `result.error` or the exit code maps to a fault signal; otherwise leave it null.
- **Spawn/user_message ordering.** Either guarantee `task.spawned` is seq-earlier than the first `task.user_message`, or document that consumers must not rely on it. Probably the former — small refactor in `tasks.py:_fire` / `start_task`.
- **Live worker streaming into main chat?** Travis: "I am getting turn three response in the summary, but not separate after it return. no big deal, unless its gated somehow." Dispatch's recommendation: keep the sidebar as the worker's stage, main chat as orchestrator voice — don't pipe worker SSE into the chat pane. Open as a deliberate design call.
- **More testing with the real tool.** Travis: "maybe more testing is needed? probably with the real tool once we have it." Explicit deferral for the long-lived vs. one-shot UX feel.
- **Memory save?** Dispatch asked "Worth saving any of this to memory before you flip back to the other session, or hold it loose until you build?" — Travis did not answer that directly in-session.
## 5. Conversation timeline (UTC)
| time | who | content |
|---|---|---|
| 22:14:16 | user | "hi, are you still there" |
| 22:14:18 | dispatch | "Yep, still here. What's up?" |
| 22:15:06 | user | mentions building a new tool, current UI is a prototype |
| 22:15:28 | user | pastes the harness-event-probe spec from the other session |
| 22:15:35 | dispatch | states plan: read tasks.py + harness.py, then capture SSE around worker exit |
| 22:1522:18 | dispatch | reads code, confirms `/events/stream` reachable, drops `/tmp/sse_capture.py` |
| 22:19:08 | dispatch | starts capture script in background |
| 22:19:45 | harness | `task.spawned` + `task.user_message` for `task_cffef42d7c9f` |
| 22:19:47 | harness | `task.message` "DONE-CAPTURE" + `task.idle` |
| 22:20:14 | harness | `task.crashed` exit_code=143 after SIGTERM (163ms post-signal) |
| 22:20:58 | dispatch | full answer landed: 5-event table, idle≠exit explanation, 3 latent bugs flagged |
| 22:35:50 | user | "great, that was helpful, i have a follow up question…" |
| 22:37:00 | user | pastes Lovebug-instance-2's principle-#5 / idle-≠-exit critique |
| 22:37:55 | dispatch | answers: one-shot fixes #5 literally; lays out 4 tradeoffs; lean = one-shot if 2-4s latency OK |
| 22:42:37 | user | "1-3s is fine. … toggle, off by default … more testing needed" |
| 22:42:47 | dispatch | confirms: one-shot default + opt-in long-lived; defer until real tool |
| 22:44:53 | user | "follow-up < 10% 30-60s tolerance" |
| 22:45:05 | dispatch | confirms shape: one-shot default + UI status pill |
| 22:47:25 | user | "30s tolerance ok but I want instant feedback! are you controlling streaming?" |
| 22:47:46 | dispatch | not gating; TTFT = thinking + prefill; UX fix = "thinking…" indicator on send |
| 22:49:09 | user | "then how did the one chat come in parts" |
| 22:49:24 | dispatch | streaming-vs-agentic-loop distinction; perceived "parts" = 6-7 chained turns |
| 22:51:14 | user | proposes a three-turn riff (color/guess/python) |
| 22:51:40 | dispatch | answers all three in one streamed turn (teal / green / 5 lines), notes single-turn-single-stream |
| 22:52:41 | user | "task a worker with these questions, in this order …" |
| 22:52:49 | dispatch | spawns `task_81c483b3703c`, sends turns 1+2+3 back-to-back (fire-and-forget) |
| 22:53:08 | dispatch | reports results; explains stdin pipe queueing |
| 22:54:36 | user | "execute the turns 2 and 3 only after getting the response from 1 and 2" |
| 22:54:42 | dispatch | spawns `task_aa44afb48ca6`, gates each send on `read_transcript(auto)` |
| 22:55:07 | dispatch | reports results; explains temporal-shape difference from previous run |
| 22:58:12 | user | "do you have to queue the messages? is there a way to send after getting the response?" |
| 22:58:48 | dispatch | 3-layer answer; recommends option 3 (`wait_for_idle` tool + fire-and-forget send) |
| 23:02:19 | user | "does that mean I can give you three tasks that need answers from the previous before the next?" |
| 23:02:35 | dispatch | yes already demonstrated; offers to run a real dependent chain |
| 23:04:27 | user | "lets try an example of three turns I cant" (cut off) |
| 23:04:32 | dispatch | asks for clarification |
| 23:04:50 | user | "I cant think of a good example, please pick one" |
| 23:05:03 | dispatch | picks: noun haiku Latin translation of last line |
| 23:05:05 | dispatch | spawns `task_4bd577a6d91e`; chains 3 dependent turns |
| 23:05:32 | dispatch | reports: Lantern haiku "Ventus spirare discit"; reflects on orchestrator's role |
| 23:06:47 | user | "interesting, I am getting turn three response in the summary, but not separate after it return. no big deal, unless its gated somehow" |
| 23:07:08 | dispatch | not gated, not streamed live; `read_transcript` returns full reply as one chunk; recommends sidebar as worker stage |
## 6. Tooling observations on petal-dispatch itself
- **`task.crashed` after clean SIGTERM is currently the only way to drive the terminal-event path.** `task.done` (rc 0) is structurally unreachable in v0.5 because long-lived workers never naturally exit.
- **Spawn-time ordering is asyncio-task-creation-order, not wall-clock-or seq order.** `task.user_message` landed 1.8ms before `task.spawned` in the probe capture. Both fired from `_fire()` via `asyncio.create_task` without sequencing.
- **`last_error` on `task.crashed` falls back to the worker's last text block when there's no `result.error`.** Misleading semantics implies a fault when the cause was just SIGTERM.
- **Three-event terminal-latency:** SIGTERM `task.crashed` event landed on the SSE stream ~163ms later. That's the stdout-EOF-propagation + event insert + LISTEN/NOTIFY + SSE fanout path measured end-to-end.
- **`read_transcript format=auto` is a usable serializer.** It blocks on `_turn_done` and returns the rendered transcript including the latest turn. Used implicitly in the sequential and dependent-chain demos to serialize sends works, but conflates "wait for idle" with "read the transcript," which is why Dispatch suggested splitting them in v0.6.
- **Stream-vs-loop confusion is real for the operator.** Travis read multi-turn agentic output as "Claude streaming in parts." Dispatch diagnosed it correctly as 67 chained turns separated by tool executions, but the UI doesn't make that distinction visible.
- **`send_message` returns `{ok:true}` immediately**, regardless of worker busy state. The queueing happens silently in the OS pipe buffer.
- **`/tmp/sse_events.jsonl` is the captured raw stream from the probe** (with wall + monotonic timestamps per event). Useful artifact for any later replay / regression test of the harness event ordering.
## 7. Cross-reference
- Strategic context for the principle-#5 question lives in trellis thread 586 (`execution-journal-synthesis-layer`).
- The "Lovebug-instance-2" feedback Travis pasted in §1 is the source for the rewritten-principle-#5 language proposed in §2.4.
- This session's experiments are the empirical input that should feed back into the v0.6 plan (one-shot default, opt-in long-lived, `wait_for_idle` tool, fixed `last_error` semantics, ordered spawn-vs-user_message).