Skip to content

Use Chatbook's reported line number for multi-input evaluations - #256

Merged
rhennigan merged 1 commit into
mainfrom
feature/multiline-wl-tool-inputs
Oct 8, 2026
Merged

rhennigan merged 1 commit into
mainfrom
feature/multiline-wl-tool-inputs

Conversation

@rhennigan

@rhennigan rhennigan commented Oct 8, 2026 •

Copy link
Copy Markdown
Member

Summary

WolframResearch/Chatbook#1677 (Chatbook 2.7.28) changed WolframLanguageToolEvaluate to evaluate each top-level input of the code separately, the way an interactive kernel session does. Each input gets its own line number:

1 + 1
2 + 2
Out[1]= 2

Out[2]= 4

One tool call can now use up several line numbers, so incrementing the session's line counter once per call puts the next call's Out[n] on the wrong line. Chatbook now reports the line number for the next input as a new "Line" property, and this PR uses it.

Changes (Kernel/Tools/WolframLanguageEvaluator.wl)

  • Line counter: chatbookToolEvaluate also requests the "Line" property, continues the session's $line from it, and drops it from the result before returning. Without a valid reported line (older Chatbook, a failed evaluation, or the eval kernel quit), the counter advances by one as before.
  • "Local" bookkeeping: session start/save/resume and the paclet load in the eval subkernel now go through a new localKernelEvaluate. On 2.7.28+ it evaluates with "Line" -> None instead of rolling back $Line. This matters because Chatbook now records In/Out history for held input, so the old rollback would overwrite history entries. As a side effect, %, Out[n] and InString[n] work in Local-method sessions now; they used to return internal ByteArrays.
  • syncEvalKernelLine: this is now a no-op on 2.7.28+, because Chatbook applies an explicit "Line" option in the Local and Cloud evaluator kernels itself.
  • UI notebook (MCP Apps): the input cell holds all of the code, so it's now labeled with the first input's line number, like a notebook input cell. Before, it took the label from the last Out[n]=, so a three-input call starting at line 1 showed In[3]:=. makeEvaluatorUIResult takes that line number as a new third argument.

All of this is gated on the installed Chatbook version, using the same pattern as the existing $CloudSessionMX gate ($linePropertyChatbookVersion = "2.7.28", chatbookLinePropertyQ). With older Chatbook versions (the minimum is still 2.7.0) the behavior is unchanged. I also updated docs/tools.md and the comments that described the old line handling.

Related: WolframResearch/Chatbook#1677, #249. #249 is fixed on the Chatbook side by the same PR, which parses each input in the evaluator kernel.

Test plan

  • New mocked unit tests in Tests/EvaluatorSessions.wlt, independent of the installed Chatbook: version gate, continuing at the reported line (string and property-list requests), fallback without a reported line, old-Chatbook path, localKernelEvaluate call shape for both versions, and the syncEvalKernelLine no-op.
  • New integration tests, run only with Chatbook 2.7.28+: multi-input line numbers across calls, a session switch, and a resume from disk, under the "Session" method and (when a sandbox kernel can start) the "Local" method, including %/Out/InString history.
  • New UI label tests in Tests/WolframLanguageEvaluator-UI.wlt, with cloud deployment mocked.
  • Installed Chatbook 2.7.0 (TestReport): EvaluatorSessions.wlt + WolframLanguageEvaluator-UI.wlt, 88/88 pass; the 2.7.28-only tests skip.
  • Dev Chatbook 2.7.28 (wolframscript with PacletDirectoryLoad): EvaluatorSessions.wlt 71/71 (including the existing cloud-session integration test), WolframLanguageEvaluator-UI.wlt 20/20, Tools.wlt 110/110.
  • End-to-end through the real tool function under "Session", "Local" and "Cloud" with Chatbook 2.7.28: numbering is correct across multi-input calls, session switches, and a resume from disk.
  • The same end-to-end run with Chatbook 2.7.0: output is identical to the previous behavior.
  • UI path against the real cloud: a two-input call deployed a notebook labeled In[1]:= / Out[2]=, and the next call continued at In[3]:= / Out[3]=.
  • CodeInspector: only the two FIXME comments that were already there.

Note: the newest Chatbook on the public paclet server is still 2.7.0, and CI doesn't install a newer one, so the 2.7.28-only integration tests will skip in CI for now.

🤖 Generated with Claude Code

https://claude.ai/code/session_015FJqQoZs86ntnnPqsvkCf6

Chatbook 2.7.28 (WolframResearch/Chatbook#1677) evaluates each top-level
input of the evaluator code separately, so one tool call can use up
several line numbers. Incrementing the session's line counter once per
call no longer matches the Out[n] labels.

- Request the new "Line" property and continue the session's line
  counter from it, falling back to one line per call when it isn't
  available (older Chatbook, failed evaluation, kernel quit).
- Evaluate "Local" eval-kernel bookkeeping with "Line" -> None instead of
  rolling back $Line, so it uses no line number and no longer overwrites
  In/Out history entries.
- Skip syncEvalKernelLine with 2.7.28+, which applies the "Line" option
  in the Local and Cloud evaluator kernels itself.
- Label the UI notebook's input cell with the first input's line number.

Gated on the Chatbook version, so behavior is unchanged with older
Chatbook versions.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015FJqQoZs86ntnnPqsvkCf6
Copilot AI balanced review requested due to automatic review settings October 8, 2026 22:29

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

The core behavior depends on Chatbook 2.7.28 paths that currently skip in CI.

0 open findings

What changed in this PR

Updates evaluator sessions for Chatbook 2.7.28’s per-input line numbering and history behavior.

Changes:

  • Continues session counters from Chatbook’s "Line" property.
  • Preserves Local evaluator history during bookkeeping.
  • Adds compatibility, integration, and UI-label tests.
File Description
Kernel/​Tools/​WolframLanguageEvaluator.wl Implements version-gated line and Local-session handling.
Tests/​EvaluatorSessions.wlt Tests line counters, history, sessions, and compatibility.
Tests/​WolframLanguageEvaluator-UI.wlt Tests multi-input notebook labels.
docs/​tools.md Documents version-dependent line behavior.

🧠 Review effort: Balanced


💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.

@rhennigan
rhennigan merged commit 9c933e7 into main Oct 8, 2026
2 checks passed
@rhennigan
rhennigan deleted the feature/multiline-wl-tool-inputs branch October 8, 2026 23:10
rhennigan added a commit that referenced this pull request Oct 9, 2026
Bring in PR #256 (Chatbook's reported line numbers for multi-input
evaluations) and the 2.2.18 version bump, and rebuild the agent skills
so that they are stamped with version 2.2.18.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HqbPD3YXk9krL9wVG9AACD
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants