Attribute the thinking level in effect to each assistant message: from
thinking_level_change branch entries for history, from the live session
level for streamed message.end events and join-time stream snapshots.
The chat metadata row renders it after the model; thinking "off" stays
hidden so non-reasoning bubbles are unchanged.
spawn_session and spawn_subsession now forward the dispatching session's
current thinking level to the new session, clamped by pi to the spawned
model's capabilities, instead of falling back to the configured default.
spawn_session and spawn_subsession accept an optional model parameter as
an exact provider/model-id (strict matching; unknown specs fail listing
available models; omitting it keeps the inherited model). The chat
composer opens a model completion menu on # and inserts a
#provider/model-id reference into the draft, which agents forward as
the model parameter.
Post-review hardening for the extension dialog feature:
- dispose() and closeActive() now settle startup-parked session_start
dialogs before awaiting pending opens, so daemon shutdown or closing a
session whose open is parked on a dialog can no longer block behind the
dialog timeout (infinite with extensionDialogsTimeoutMs: 0).
- A failed create now drops its dead dialog cards, closes the
early-subscribed socket, and ignores late dialog frames instead of
leaving an unanswerable card on the failed row.
- The pending-start status resync checks its staleness guard before
applying the unordered snapshot, so a late response can no longer
clobber the post-swap session state.
- The dialog countdown no longer queues one screen-reader announcement
per second (decorative; the daemon-owned dialog.closed event is the
real signal) and no longer renders "1h 60m" near hour boundaries.
A dialog opened from a session_start hook parks session construction
before the session ever becomes active, but every servable path gated on
readiness: the answer/cancel routes and status 404'd (or parked behind
the in-flight open), and the client only subscribed once its create
request resolved — so the dialog that gated readiness could never be
answered and always rode to the daemon timeout.
Daemon: hold a startupSessions registry for the duration of extension
binding and let status, answerDialog, and cancelDialog resolve active →
startup → getOrOpen fallback. getOrOpen itself is untouched, so prompts
and other mutations still cannot reach a half-constructed session.
status() no longer parks behind an in-flight open; it resolves from the
startup window (intended semantics change, lifecycle test updated).
Client: the pending-start row learns the real session id from the first
token-matched session.startup event, connects its otherwise idle session
socket to the constructing session, and recovers pre-subscription opens
with a merge-based status resync (the unordered HTTP snapshot only adopts
dialog ids the ordered socket channel never reported). The leg-4 dialog
card renders in the startup view with no component changes, and answers
go out under the real id. Readiness proceeds as before once the hook
settles.
A user abort while a tool_call dialog is parked deadlocked the dialog
until its timeout: pi's agent loop waits for the parked dialog handler
before emitting agent_end, and run-scoped dialogs were settled only on
agent_end. Settle them synchronously at abort-request time, before
awaiting the runtime abort, so a hung or failing abort cannot strand
the parked waiter. Keep the agent_end settlement as the run-crash
backstop; the store makes the double settlement a stale no-op.
ctx.ui.confirm()/select()/input() from extensions now open daemon-owned
pending dialog records, publish dialog.opened/dialog.closed, and park a
Promise that settles on the browser's answer or cancel, the extension's
own signal/timeout, the extensionDialogsTimeoutMs daemon default (5 min,
0 = forever; tuning knob, not a gate), agent_end for run-scoped dialogs,
or session-ended on close/replace/dispose. Answers resolve the parked
extension Promise directly via POST /sessions/:id/dialogs/answer|cancel
- never the prompt queue - with first-wins stale semantics across
browsers, and SessionStatus.pendingDialogs rehydrates reloading clients.
Domain layer for issue #106: a daemon-owned, per-session
PendingExtensionDialogStore (multiple open dialogs, no supersede,
kind-validated answers, stale-tolerant closes) plus the shared
PendingExtensionDialog/ExtensionDialogOutcome wire types,
SessionStatus.pendingDialogs, and dialog.opened/dialog.closed events.
Relay: issue-106-extension-dialogs leg 1
An ordinary chat message answers the session's open ask in the user's
own words, so keeping the form open would invite answers to questions
the conversation has already moved past. The form now closes as
cancelled, browsers clear the live card, and the model is told without
being woken so the notice rides into the turn the message itself
starts. Ignored duplicate queued messages skip the void on purpose:
they must not void an ask posted after the queued original.
Five tests failed only on the Windows runner. Their fixtures used bare
POSIX-absolute paths such as /srv/other-worktree and /old-project, which
win32 treats as absolute but drive-relative: resolve() maps them onto the
runner's current drive (D:\...). The code under test canonicalizes stored
cwds by contract, so assertions comparing against the raw fixture string
never matched on Windows:
- parentSessionLocator and crossWorkspace listings annotate
parentSessionCwd with canonicalizeStoredCwd(header.cwd);
- cleanup forgets unread via canonicalizeStoredCwd(record.cwd), which
missed the marker the test seeded with the raw path.
Resolve the fixture paths once at declaration, matching the existing
WORKSPACE_CWD = resolve("/workspace") convention, so fixtures model what
a real Windows Pi would record. Linux behavior is unchanged.
The migration moved ~/.pi-web/archived-sessions* into a custom
PI_WEB_DATA_DIR at session daemon startup. Remove it outright: data
directories are independent, and anyone who sets a custom data
directory can copy archived-sessions.json and archived-sessions/
manually while the daemon is stopped.
Without the migration, runSessionDaemonStartup's only remaining
contract was sequencing create/register/listen, so inline those steps
into sessiond.ts and drop the wrapper module and its tests.
A session spawned into another worktree recorded a parent that no
listing contained, so the row showed only "parent unavailable" and its
parent's row looked childless. Both facts were accurate and useless:
neither said where the related session actually was.
Report both directions from the session store instead. A missing
parent's cwd and id come from its own file header, so one 4 KB read per
distinct missing parent resolves it without listing other workspaces;
children are counted by listing sibling workspaces and matching the
parent path they already recorded, needing no header reads. Reads are
memoized per path because Pi writes headers once, and the cache is
released on dispose. Both directions are best-effort: an unreadable
header or an unlistable worktree leaves a session unannotated rather
than failing the listing.
In the browser, an orphan child keeps the same child marker as a nested
one, dimmed, so it no longer renders as a root; whereabouts are stated
once on the meta line ("parent in feature/foo", "2 children
elsewhere"), where a clamped title cannot hide them. A "Go to parent
session" action switches to the owning workspace and selects the
parent. Live session.created events keep child counts current instead
of leaving them stale until the next listing.
Session and workspace paths reach the browser from two producers: store
enumeration for a listing, and the live runtime for a broadcast. They
are now compared through one normalizing helper, so tree nesting and
child counts cannot silently miss a link when only a trailing separator
differs.
Extract the shared "workspaces of the project containing this cwd"
lookup out of ProjectScopedSpawnTargetResolver so spawn targeting and
child counting share one implementation, and register it regardless of
whether spawning is enabled: children can predate a config change, and
the tree should stay honest about them either way.
Relay unread-ux leg 1: audit verified every archive path (single, bulk,
tree, cleanup, restore, delete, rebind) clears unread state in the right
order with no resurrection vector. Add regression tests for the uncovered
paths: bulk archive, archive with descendants, cleanup, bulk delete, and
archiving an active session with a pending activity latch.
Startup progress rides the per-session activity channel with an "active"
phase, because a startup phase really is in progress. But isSessionActive()
treated any active activity as work, so a session that was merely *opening*
enabled "Stop Active Work", disabled "Reload from disk" with the misleading
"Stop current session activity before reloading" tooltip, showed the row's
active-work indicator, and — for any caller that hands a startup activity to
WorkspaceActivityService — reported the whole workspace as busy. Selecting an
archived, read-only session reported active work while it opened.
Starting is not working. publishStartupProgress now marks its reports with a
new optional SessionActivity.startup field, and isSessionActive() does not
count a marked activity. Every affected consumer — the session list, the core
actions, the app's activity-transition handling, and the server's workspace
aggregation — reads that one helper, so the correction lands in all of them at
once.
The marker is a new field rather than a new phase on purpose: six readers test
phase === "active" directly, including the pending row's "creating · " prefix,
the chat dock's active styling, and the daemon's own heartbeat re-publication.
It can only ever remove the activity-phase reason for being active, so
streaming, bash, compaction, and queued prompts still report as active through
the status even while a startup report is the latest activity. The chat dock
still shows the startup text; this changes what counts as work, not what is
shown.
The browser's own pending-create row keeps its previous appearance: it borrows
only the daemon's phase text and drops the marker, since that row stands for a
create the user is waiting on rather than a session the daemon is opening.
Integrate the pending-ask store into the session service so an open ask is
visible, observable, and closable.
- `statusFromSession` projects `pendingAsk`, so a browser rehydrates an open
ask from `GET /sessions/:sessionId/status` after reload or a web/API restart.
- `openAsk` publishes `ask.opened`, and publishes `ask.closed` first when the
new ask supersedes an unanswered one.
- `submitAsk` / `cancelAsk` close the ask and hand the outcome to the model as
a `pi-web.ask.answers` follow-up custom message (`triggerTurn`,
`deliverAs: "followUp"`), the same delivery subsession notices use. A stale
ask id is reported, not thrown: losing the race against a supersede or
another browser is ordinary. Cancel still reports every question as
unanswered so the model is not left waiting for a promised message.
- The open ask is forgotten when its runtime closes; nothing is left to
receive the answers.
- `POST /sessions/:sessionId/ask/{submit,cancel}` behind the existing
`/api/sessions/*` daemon proxy, allowlisted for machine federation.
Startup progress could still be shown on the wrong session's row. Routing by
known session id first closed the case where the browser knew the other
session, but left open the case where it does not -- which the browser is
designed to produce. While a create is pending for a workspace,
applyCreatedSession deliberately withholds a session.created event for that
workspace and stashes it, to avoid a duplicate row. So during exactly the
window this feature exists for, a session created by an agent's spawn or by
another tab is intentionally absent from the session list. Its startup events
carried an unrecognised id and a matching cwd, and were routed onto the user's
pending create row, showing a phase and a label belonging to another session.
Workspace path was never evidence of identity; it was the only key both sides
happened to share. Give them a real one. The browser already invents a
temporary row id for a pending create, so it now sends that id with the create
request as an opaque startupToken; the daemon carries it through construction,
echoes it on the startup events it publishes for that construction, and the
browser matches it exactly. The token is a throwaway label the daemon never
interprets. It never becomes the session id: activity.sessionId still carries
Pi's SessionManager id, which remains how an open of an already-known session
is routed.
With exact identity available, the guessing is deleted rather than gated.
startupProgressPendingStart goes entirely, and with it the selected-machine
comparison, the cwd filter, and the single-match ambiguity rule: a second
concurrent create carries a different token, and a foreign workspace or
non-selected machine carries no token this browser is waiting on, so those
cases stop existing rather than needing detection. One Map lookup replaces a
filtered scan. cwd comes off the event, since it existed only as the routing
key and nothing else read it.
No compatibility path is needed. session.startup is unreleased -- checked
against the published tarball, not only git tags -- so no deployed daemon
emits these events and no deployed browser parses them. An older daemon
ignores the extra request field; a newer daemon talking to an older browser
degrades to the pre-existing generic wording, as does any unmatched token.
One silent behaviour change to state plainly: startupProgress guarded on
`sessionId === "" || cwd === ""`. Removing cwd from the event removes the
meaningful half of that guard, and that half had no test. The session-id half
is kept, which is the half that actually protects honest reporting.
The replaced ambiguity test is rewritten rather than dropped, so the same three
scenarios still pin the user-visible guarantee -- no match means the generic
wording stays -- now including the reproduced foreign-session case, which fails
against the previous code. Session creation ordering, semantics, and queueing
are unchanged; the token is a passthrough label read only to build an event.
Register a core ask_user custom tool that posts a question set to the user's
browser and terminates the run instead of awaiting an answer. The tool is thin:
it shapes its TypeBox params into domain questions, lets PendingAskStore own
validation, and reports a superseded unanswered ask back to the model.
Gated by the askUser config key, threaded through PiSessionServiceDependencies
and sessiond. Unlike the delegation tools, ask_user is available to tracked
children too: the questions reach the user of the asking session.
Own the one-open-ask-per-session lifecycle in daemon-side domain logic: validate
model-authored question sets, validate submitted answers against them, and
compute the answered-versus-unanswered outcome both the model-facing follow-up
message and the browser record are rendered from.
Also lands the answer half of the shared ask contract alongside its first
consumer.
Two behaviors of the narrowed provider freeze were unprotected: deleting
the baseline rebase, or swallowing Pi's validation error on the accept
path, both left the suite green.
Add a replay test asserting one applied update and one de-duplicated
ignored entry across four identical registrations, which is what a
per-session session_start handler produces. This fails if the rebase is
removed, because every replay then differs from the original catalog and
is accepted forever.
Add a test for a provider whose models carry their own api/baseUrl, so a
refreshed catalog omitting them fails validation. It pins that the error
reaches the extension rather than being silently swallowed, and that the
recorded baseline and previously registered models survive.
Note that a poisoned-baseline reordering is deliberately not asserted:
recording the incoming config early only corrupts `models`, and models
are expected to differ, so no later comparison can observe it.
Startup progress resolved its target row by workspace path first and only
fell back to a known session id, which let one row be shown another row's
phase. While a create is pending in a workspace, an existing session in that
same workspace can also be opened -- by selecting another row, by another
tab, or by a subsession open -- and that open publishes the same cwd. The
cwd-first order rewrote such an event onto the pending create row, so a user
watching a session being created could be told a phase that belonged to a
different session. That is exactly the dishonest attribution this work set
out to avoid.
A known session id is the strongest available proof of the target, so it is
now checked first; workspace routing is used only when the id is unknown,
which is precisely the pre-session case it exists for. No wording changed and
no event changed; only which row an event is applied to.
Two tests were added where behavior was asserted but not proved. The
controller test fails against the previous order, so the misattribution is
now pinned. The service test covers a startup whose extension binding
rejects, proving the window still ends with an idle report rather than
leaving a waiting row labelled with a phase the service has left.
Creating or opening a session could stall for reasons the daemon knew
about and never shared. The browser invented the whole message it showed
while waiting -- "Creating session: Waiting for the backend session to be
ready" -- which says that we are waiting but never what for. A shared
ModelRuntime read during startup can be handed a network refresh that is
already in flight, and extensions may do their own network I/O while
loading, so the wait is real and previously unattributable.
The pre-session gap turned out to be a missing shared key rather than a
missing channel: publishActivity needs the PiAgentSession being built, but
the session id and cwd are both known before the first await. So create()
now publishes a new global session.startup event carrying an ordinary
SessionActivity, routed by cwd -- the one identity a browser row waiting
for a session id can match, since the client-invented pending id is
unknown to the daemon and the daemon's id is unknown to the browser.
Two phases are reported, each published before the await it describes so
the label changes during the wait rather than after it: "Starting the Pi
session" and "Loading session extensions". Both are facts, because the
service awaits exactly one call for each. A concurrent background catalog
refresh is appended as a note ("provider model lists are refreshing"),
never as the cause: the refresher can prove a refresh is running but not
that this startup joined it. ModelCatalogRefresher gains only a read-only
isRefreshInFlight() getter; cadence, timeout, and coalescing are untouched.
Reporting is event-only and synchronous. It writes no activities entry, no
workspace activity, and no unread state, so a failed creation leaves
nothing stranded, no await is added, and creation ordering and semantics
are unchanged. The window-ending idle report is skipped when a real
activity was published during startup, so an extension error survives.
The browser applies startup progress only when it can prove the target:
one non-discarded pending start in that cwd on the selected machine, or a
session whose id it already knows. A foreign workspace, another machine,
or two concurrent starts in one workspace keep today's generic wording
rather than showing one row the phase of another. An idle report restores
that generic wording, including the queued-messages variant.
docs/config.md said nothing a request triggers waits on a catalog fetch.
That is not strictly true for a refresh already in flight, so both it and
the generated docs/config.html now state the exception and say PI WEB
reports it while it happens.
The global provider bootstrap froze all three ModelRuntime mutation
methods after startup, so a provider extension that fetched an updated
model catalog had that work silently discarded.
registerProvider is now applied when the provider ID is already in the
frozen baseline and the incoming config equals the recorded baseline in
every field except `models`. Refreshing extensions re-send a complete
provider config rather than a models-only delta, so the test is
"equal except models", not "contains only models".
Everything else stays a logged no-op: unknown provider IDs, any change
to name/baseUrl/apiKey/api/streamSimple/headers/authHeader/oauth/
refreshModels, native registration, and unregistration. Function-valued
fields compare by reference and so always read as a mismatch, which is
the intended conservative direction.
An accepted update rebases the stored baseline from Pi's merged record,
so repeat refreshes work and an unchanged replay is correctly ignored
rather than re-applied on every session start. The accept path stays
synchronous and never awaits or networks; Pi's own trailing
fire-and-forget local refresh is untouched.
Bump @earendil-works/pi-coding-agent, pi-ai, and pi-agent-core to 0.82.1
together and raise the peer range to >=0.82.1 <0.83. All three must move in
lockstep: bumping only two leaves a duplicate pi-ai copy in the tree, which
surfaces as misleading type-identity errors rather than real API breaks.
Pi 0.82 removed ModelRuntime.reloadConfig() and merged it into refresh(),
which now does ModelConfig.load, configureRadiusProviders, and rebuildProviders
before refreshing. Port the five production call sites literally, passing no
options so refresh() keeps defaulting allowNetwork to modelNetworkEnabled --
which the shared runtime pins to false by constructing under PI_OFFLINE. No
call site passes allowNetwork: true.
The auth tests lose reloadConfig() as an observation seam, so the offline
regression cases now drive removeRuntimeApiKey(), the surviving public mutation
that still forwards the construction-time network flag to refresh().
The reworked assertion ran under the file-level PI_OFFLINE=1 stub, so the
runtime was offline whether or not createOfflineModelRuntime forced it and
the test passed with the fix fully removed. Clear the stub for that case,
and rewrap a docblock line.
Finding 6: serialize createOfflineModelRuntime so overlapping calls cannot
interleave their PI_OFFLINE save/restore pairs and leave the process offline,
and name the process-wide visibility of that window in the docblock.
Finding 7: assert the offline construction through the public refresh seam via
reloadConfig() — the request path that regressed — instead of reading upstream's
private modelNetworkEnabled field.
Finding 8.4/8.5: document the background provider-catalog refresh in
docs/config.md and docs/config.html (cadence, timeout, single retry, offline
opt-out via PI_WEB_OFFLINE / PI_OFFLINE only), and update the changeset to
match the behavior after the earlier fixes.
`dispose()` only cleared timers, so a refresh already in flight kept its
provider fetch alive for the rest of the timeout budget and could delay
daemon shutdown, which is exactly when sessiond disposes the refresher.
A refresher-lifetime AbortController is now combined with the per-run
timeout via `AbortSignal.any`, and `dispose()` aborts it. A
dispose-triggered abort logs as expected shutdown info rather than a
timeout warning or an error, whether the runtime resolves as aborted or
rejects.
`start()` is now idempotent: a second call previously overwrote both
timer handles and leaked the first pair, which kept firing.
Also replaces the `then().catch()` bookkeeping chain in `queueRefresh()`
with an awaited private `runCycle()`, keeping the coalescing, retry, and
dispose semantics unchanged.
Tick the background catalog refresher hourly instead of every four hours:
pi stamps `checkedAt` after a fetch completes, so a tick at exactly its 4h
freshness window always landed a few seconds short and only fetched on
every other tick (~8h effective). Scheduled runs stay unforced, so the
extra ticks are nearly free and pi's gate keeps deciding when to fetch.
Auth-triggered refreshes now pass `force: true` so a re-login of a
provider refreshed within the last four hours actually reaches the
network. A request queued behind an in-flight run keeps the strongest
mode asked for, so a forced request is never downgraded.
Raise the whole-cycle timeout to 60s, since one run covers every
refreshable provider and a background job has no startup budget, and give
a timed-out or errored run exactly one bounded retry. Retries never earn
retries, are superseded by any fresh request, and are cleared by
`dispose()`.
The background model catalog refresher always requested a network refresh,
so sessiond fetched provider catalogs on a schedule even when the operator
set PI_OFFLINE or PI_WEB_OFFLINE. Before the refresher existed, those
settings made every runtime refresh local-only.
Add `offlineModeEnabled()` to the config module and inject the resulting
flag from sessiond's frozen daemon environment, so the refresher schedules
nothing and ignores auth-triggered requests in offline mode. The narrower
PI_SKIP_VERSION_CHECK / PI_WEB_SKIP_VERSION_CHECK keys are deliberately not
included: they only suppress release lookups.
The shared ModelRuntime was constructed with network refreshes enabled, so
reloadConfig()/login()/logout() — called on the model picker, session model
changes, and auth dialogs — performed unbounded provider-catalog fetches.
A single stalled fetch blocked those requests for minutes and, through pi's
coalesced per-provider refresh, dragged session creation along with it.
Construct the runtime with PI_OFFLINE forced so every runtime-driven refresh
stays local, and add ModelCatalogRefresher as the single deliberate network
path: bounded by an abort timeout, serialized through one in-flight run,
scheduled in the background, and triggered after provider auth changes.
Relaxes the provider policy from 'global config only' to 'global sources':
providers registered by agent-dir (global) extensions are learned once at
daemon startup and allowed on the shared runtime; project-extension
registrations are still rejected with a session warning. Global extensions
load identically for every session, so their providers are daemon-consistent
and cannot leak project state (#76).
- Shim now allows allowlisted ids through and also covers Pi 0.81's native
provider path (registerNativeProvider), closing a bypass.
- Startup learning step loads only global extensions against a scratch cwd
and diffs the runtime's registered provider ids.
- Bumps @earendil-works/* dev/peer ranges to >=0.81.1 <0.82; adapts to the
Agent.streamFn -> streamFunction rename.
- Docs, changeset, unit and acceptance tests updated (global-extension allow
path, late re-registration a la pi-tensorx, native provider rule).
Unit tests for the policy shim (swallowed registrations, no-op
unregister, untouched global providers, rejection wording) and
acceptance tests wired as sessiond wires production: load-time
rejections surface as session warnings while extension tools and
commands keep working, late registrations are broadcast to active
sessions' notification inboxes, colliding provider ids across
workspaces cannot affect each other, and a project-level models.json
does not alter the shared runtime.
PI WEB only supports globally configured providers (Pi built-ins,
agent-dir models.json, environment credentials). A daemon-wide shim on
the shared ModelRuntime swallows extension registerProvider calls and
makes unregisterProvider a no-op, so one workspace's extensions can no
longer corrupt the provider set of concurrent sessions (issue #76).
Rejections during a services load surface as session warnings through
the existing diagnostics pipeline; late registrations from session
event handlers broadcast a notification to active sessions. Everything
else extensions register keeps working.
Requires manual restart of pi-web-sessiond.service (daemon wiring changed).