Skip to content

Add the wolfram-debugging agent skill - #257

Merged
rhennigan merged 10 commits into
mainfrom
feature/wolfram-debugging-skill
Oct 9, 2026
Merged

rhennigan merged 10 commits into
mainfrom
feature/wolfram-debugging-skill

Conversation

@rhennigan

@rhennigan rhennigan commented Oct 9, 2026 •

Copy link
Copy Markdown
Member

Summary

Adds a wolfram-debugging agent skill: verified recipes, a tested helper package and environment notes for finding out why Wolfram Language code misbehaves. It is added to the Wolfram (Computation Tools), WolframLanguage (Development Tools) and WolframPacletDevelopment built-in toolsets, which is every one except WolframAlpha.

The skill

  • SKILL.md covers:
    • how to tell where code runs (MCP Local, MCP Session, wolframscript, cloud) and what differs there; for example, Trace records nothing in MCP Session or the cloud
    • ten ground rules: check that Trace works, handle held stack frames, mind parse-per-input, scope changes in shared MCP kernels, and so on
    • a symptom → first move → reference routing table
    • inline recipes for the stack at the first message, spies/mocks/overrides, unevaluated calls, stack sampling of slow code and ReadProtected definitions
  • references/ has 12 topic files: messages and stacks, unevaluated calls, tracing and performance, overrides and watchpoints, definitions and source, package loading, testing, parallel and async code, cloud and HTTP, headless kernels and crashes, environments, and the helper API.
  • scripts/WolframDebugging.wl has 55 public symbols: 54 helper functions such as CollectMessages, StackAtMessage, WhyNoMatch, SampleStacks, LogCalls, WithOverrides, WatchChanges, FindSymbolSource and TestFailureSummary, plus $CaptureFunction. They return small, bounded results and are tested in Tests/WolframDebuggingSkill.wlt.

Build and registration

  • Scripts/Resources/AgentSkillsBuilder.wl now copies hand-authored files from directly inside AgentSkills/Skills/<name>/references/ and scripts/ into the built skill.
    • This loosens the previous check, which rejected every file except SKILL.md. A stray file in those two directories is now shipped, for example a leftover generated scripts/Old.wls. The exception is a file whose path matches one the build generates, which fails with SkillFileConflict (case-insensitive). Two hand-authored files whose paths differ only in case fail with SkillFileCaseConflict.
    • Files anywhere else, nested directories and names that don't fit the allowed pattern still fail with UnexpectedSkillFiles.
    • The existing ExtraSkillFile test now puts its stray file at the top level of the skill directory instead of in scripts/.
    • Built skills without generated scripts no longer need references/Scripts.md.
  • The skill is registered in $defaultAgentSkills, added to the toolsets in Kernel/AgentToolsObject.wl, listed in AgentSkills/Manifest.wl and the Claude Code plugin marketplace, and documented in docs/agent-skills.md, docs/agent-tools-objects.md, the specs, README, AGENTS.md and the docs that list skills per toolset.
  • The shared SetUpWolframMCPServer.md reference is in every built skill. It now lists wolfram-debugging in its toolset table and in the expected output of AgentToolsObject["WolframLanguage"]["AgentSkillNames"] (Step 2). That expected output is only right with an AgentTools release that includes this change; the newest on the paclet server is 2.2.0.
  • Existing deployments are not changed. The skill is installed the next time DeployAgentTools runs for one of these toolsets, which includes choosing Computation Tools or Development Tools in the preferences panel.
  • Reviewer question: the Wolfram (Computation Tools) server has no SymbolDefinition or TestReport tool, but the skill mentions both. Is adding the skill to that bundle intended?

Chatbook 2.7.28 (unreleased)

The skill describes how the MCP evaluator behaves with Chatbook 2.7.28, which evaluates tool code one input at a time (WolframResearch/Chatbook#1677, see #256) and fixes #249.

Chatbook 2.7.28 is not on the public paclet server yet. The newest there is 2.7.0, so until it ships, every user gets the older behavior. In MCP Local, tool code is then parsed as one input per call, and typed symbols land in Global` .

Size of the always-loaded text

The latest commit condenses SKILL.md toward the agentskills.io budgets (about 100 tokens of metadata, under 5000 tokens of instructions). The body is now under 5000 tokens, and the description is down to 118 tokens. The references, which agents load on demand, are unchanged.

(LLMTokenize, GPT-2) before after
description 224 118
body 5871 4449

Test plan

  • Tests/WolframDebuggingSkill.wlt: 122/122
  • Tests/AgentSkillsBuild.wlt: 76/76 (built skills match Scripts/BuildAgentSkills.wls)
  • Tests/AgentSkills.wlt: 162/162
  • Tests/AgentToolsObject.wlt: 78/78
  • Tests/DeployAgentToolsSkills.wlt: 169/169
  • Audit of the condensed SKILL.md. An agent checked every claim and helper signature against the references and the helper package source, and ran every snippet in an MCP Local kernel with Chatbook 2.7.0. Claims that depend on Chatbook 2.7.28 were checked only against the references, not run. This includes the MCP Local cells of the environment table (the context of typed symbols, and $MessageList reset per input). The audit's fixes are in the last commit.
  • Three debugging tasks (message origin, unevaluated calls, slow code with a mock) were each run once with the skill before condensing and once after, then graded blind against verified answers. Both versions passed every check. The grader preferred the condensed version on all three tasks, because it reached the answer in fewer evaluator calls.
  • CI (runs the full suite against the MX build)

All local runs used the installed Chatbook 2.7.0. The 2.7.28 behavior comes from commit daa9147 and was not re-checked in this session.

Review fixes

  • HTTPDiagnostics answers debug=1 only with "Debug" -> Automatic; the default is now False. The references say to debug a private deployment as its owner. It no longer reuses cached request data after running outside a request. Checked on private Wolfram Cloud deployments.
  • The skill builder rejects hand-authored files whose paths differ only in case.
  • references/Environments.md now agrees with itself: when MCP Local restarts its subkernel, Print output still arrives, without its label. Checked against a Local-method server running Chatbook 2.7.28.

🤖 Generated with Claude Code

https://claude.ai/code/session_01EmmjJxgtH4Ye2gWta1gjcj

rhennigan and others added 4 commits October 8, 2026 15:20
Add a wolfram-debugging skill to every built-in toolset except
WolframAlpha. It covers message handlers and stack traces, Trace and its
alternatives where Trace records nothing, overrides and spies with
Internal`InheritedBlock, finding definitions and source files, package
loading, testing, headless kernels, parallel and cloud code, and how the
MCP evaluator (Local and Session), wolframscript, wolfram -script and
the cloud differ. Its scripts/WolframDebugging.wl provides 55 helper
functions, tested in Tests/WolframDebuggingSkill.wlt.

Passages that depend on the MCP Local context bug are marked with #249.

The skill builder now copies hand-authored references/ and scripts/
files from AgentSkills/Skills/<name>/ into the built skill, and rejects
other files and paths that collide with generated ones.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FwsDzVwfn44xLSfbeS1sCW
Chatbook 2.7.28 evaluates WL tool code one input at a time in the
evaluator kernel, which fixes AgentTools#249. The skill now describes
that behavior:

- Typed code lives in the session context in MCP Local too, and sessions
  are isolated; packages loaded on an earlier line resolve by short name.
- Each input is parsed just before it runs and gets its own Out[n]; %
  works; Print and messages look the same in both methods; Print output
  survives aborts and time-outs, and the time limit covers the whole call.
- Updated bracket repair, $MessageList, $Pre/$Post, HoldComplete display,
  the Catch rule for $RecursionLimit overflows and the MCP Local restart
  timing.
- Warn that in MCP Local a fresh subkernel's first message fails under a
  low $RecursionLimit and leaves $Context broken.

RunTestsByID no longer needs its MCP Local workaround: "Context" ->
Automatic is now just $Context.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HqbPD3YXk9krL9wVG9AACD
Bring in PR #256 (Chatbook's reported line numbers for multi-input
evaluations) and the 2.2.18 version bump, and rebuild the agent skills
so that they are stamped with version 2.2.18.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HqbPD3YXk9krL9wVG9AACD
Cut the skill's always-loaded text to the agentskills.io budgets (~100
tokens of metadata, under 5000 tokens of instructions) while keeping its
concepts. Measured with LLMTokenize (GPT-2), the description goes from
224 to 118 tokens and the body from 5871 to 4449 (216 lines, was 288).

- Shorten the description to the trigger phrases and drop the list of
  internal function names.
- Merge the symptom table and the references table into one routing
  table, and add TimeCalls to it.
- Keep the stack-at-message, spy/mock, unevaluated-call and stack
  sampling recipes; fold the stack-at-call recipe into the spy section
  and shorten the ReadProtected and "what got loaded" sections. All of
  them are in the references in full.
- Tighten the ground rules. Rule 4 now says to put Get/Needs in its own
  call, since per-line parsing needs Chatbook 2.7.28+ and AgentTools
  supports 2.7.0. The spy section warns that ClearAll inside
  Internal`InheritedBlock wipes the outer definitions.
- Fix two environment-table cells that were stale after the Chatbook
  2.7.28 update: Print is shown in both MCP methods, and $MessageList is
  reset per input in MCP Local (as references/Environments.md says).

An audit checked every claim and snippet against the references and the
helper package, and three debugging tasks run with the old and the new
skill passed all checks with both versions.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EmmjJxgtH4Ye2gWta1gjcj
Copilot AI balanced review requested due to automatic review settings October 9, 2026 13:54

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The build currently fails, and unresolved security, Chatbook compatibility, and documentation issues remain.

7 open findings
What changed in this PR

Adds a comprehensive Wolfram Language debugging skill, supporting helper APIs, environment-specific guidance, packaging, deployment, and documentation.

Changes:

  • Adds the debugging skill, references, helper package, and tests.
  • Extends the skill builder for hand-authored files.
  • Registers the skill across built-in development toolsets and packaging.
File Description
.claude-plugin/​marketplace.json Adds the debugging skill plugin.
AGENTS.md Documents skill source files.
AgentSkills/​Manifest.wl Registers the new skill.
AgentSkills/​References/​SetUpWolframMCPServer.md Updates toolset skill lists.
AgentSkills/​Skills/​wolfram-debugging/​SKILL.md Adds primary debugging guidance.
AgentSkills/​Skills/​wolfram-debugging/​references/​CloudAndHTTP.md Adds cloud and HTTP guidance.
AgentSkills/​Skills/​wolfram-debugging/​references/​DefinitionsAndSource.md Adds definition inspection guidance.
AgentSkills/​Skills/​wolfram-debugging/​references/​Environments.md Documents runtime differences.
AgentSkills/​Skills/​wolfram-debugging/​references/​HeadlessAndCrashes.md Covers headless failures and crashes.
AgentSkills/​Skills/​wolfram-debugging/​references/​HelperFunctions.md Documents helper APIs.
AgentSkills/​Skills/​wolfram-debugging/​references/​MessagesAndStacks.md Covers messages and stacks.
AgentSkills/​Skills/​wolfram-debugging/​references/​OverridesAndWatchpoints.md Covers overrides and watchpoints.
AgentSkills/​Skills/​wolfram-debugging/​references/​PackageLoading.md Covers package-loading diagnosis.
AgentSkills/​Skills/​wolfram-debugging/​references/​ParallelAndAsync.md Covers parallel and asynchronous code.
AgentSkills/​Skills/​wolfram-debugging/​references/​Testing.md Adds test-debugging guidance.
AgentSkills/​Skills/​wolfram-debugging/​references/​TracingAndPerformance.md Covers tracing and profiling.
AgentSkills/​Skills/​wolfram-debugging/​references/​UnevaluatedCalls.md Diagnoses unmatched calls.
AgentSkills/​Skills/​wolfram-debugging/​scripts/​WolframDebugging.wl Implements debugging helpers.
Assets/​AgentSkills/​wolfram-alpha/​SKILL.md Updates generated version.
Assets/​AgentSkills/​wolfram-alpha/​references/​SetUpWolframMCPServer.md Regenerates setup reference.
Assets/​AgentSkills/​wolfram-debugging/​SKILL.md Adds built debugging skill.
Assets/​AgentSkills/​wolfram-debugging/​references/​CloudAndHTTP.md Adds built HTTP reference.
Assets/​AgentSkills/​wolfram-debugging/​references/​DefinitionsAndSource.md Adds built definitions reference.
Assets/​AgentSkills/​wolfram-debugging/​references/​Environments.md Adds built environment reference.
Assets/​AgentSkills/​wolfram-debugging/​references/​GetWolframEngine.md Adds shared engine setup reference.
Assets/​AgentSkills/​wolfram-debugging/​references/​HeadlessAndCrashes.md Adds built headless reference.
Assets/​AgentSkills/​wolfram-debugging/​references/​HelperFunctions.md Adds built helper reference.
Assets/​AgentSkills/​wolfram-debugging/​references/​MessagesAndStacks.md Adds built message reference.
Assets/​AgentSkills/​wolfram-debugging/​references/​OverridesAndWatchpoints.md Adds built override reference.
Assets/​AgentSkills/​wolfram-debugging/​references/​PackageLoading.md Adds built loading reference.
Assets/​AgentSkills/​wolfram-debugging/​references/​ParallelAndAsync.md Adds built concurrency reference.
Assets/​AgentSkills/​wolfram-debugging/​references/​SetUpWolframMCPServer.md Adds built MCP setup reference.
Assets/​AgentSkills/​wolfram-debugging/​references/​Testing.md Adds built testing reference.
Assets/​AgentSkills/​wolfram-debugging/​references/​TracingAndPerformance.md Adds built tracing reference.
Assets/​AgentSkills/​wolfram-debugging/​references/​UnevaluatedCalls.md Adds built call reference.
Assets/​AgentSkills/​wolfram-debugging/​scripts/​WolframDebugging.wl Adds built helper package.
Assets/​AgentSkills/​wolfram-language/​SKILL.md Updates generated version.
Assets/​AgentSkills/​wolfram-language/​references/​SetUpWolframMCPServer.md Regenerates setup reference.
Assets/​AgentSkills/​wolfram-notebooks/​SKILL.md Updates generated version.
Assets/​AgentSkills/​wolfram-notebooks/​references/​SetUpWolframMCPServer.md Regenerates setup reference.
Assets/​AgentSkills/​wolfram-paclets/​SKILL.md Updates generated version.
Assets/​AgentSkills/​wolfram-paclets/​references/​SetUpWolframMCPServer.md Regenerates setup reference.
Kernel/​AgentSkills.wl Registers the built-in skill.
Kernel/​AgentToolsObject.wl Adds it to three toolsets.
README.md Lists the new skill.
Scripts/​Resources/​AgentSkillsBuilder.wl Copies hand-authored skill files.
Specs/​AgentSkills.md Updates the skill specification.
Specs/​AgentToolsObject.md Updates bundle definitions.
Tests/​AgentSkills.wlt Updates built-in skill tests.
Tests/​AgentSkillsBuild.wlt Tests builder enhancements.
Tests/​AgentToolsObject.wlt Tests updated toolsets.
Tests/​DeployAgentToolsSkills.wlt Tests deployment behavior.
Tests/​WolframDebuggingSkill.wlt Tests debugging helpers.
docs/​agent-skills.md Documents the skill and builder.
docs/​agent-tools-objects.md Updates bundle documentation.
docs/​building.md Documents copied source files.
docs/​getting-started.md Updates repository layout.
docs/​mcp-clients.md Updates deployment example.
docs/​preferences-content.md Updates preference toolsets.

🧠 Review effort: Balanced


💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.

Comment thread AgentSkills/Skills/wolfram-debugging/scripts/WolframDebugging.wl
Comment thread AgentSkills/Skills/wolfram-debugging/scripts/WolframDebugging.wl Outdated
Comment thread Scripts/Resources/AgentSkillsBuilder.wl Outdated
Comment thread AgentSkills/References/SetUpWolframMCPServer.md
Comment thread AgentSkills/Skills/wolfram-debugging/SKILL.md
Comment thread AgentSkills/Skills/wolfram-debugging/references/Environments.md
Comment thread AgentSkills/Skills/wolfram-debugging/references/Environments.md Outdated
rhennigan and others added 6 commits October 9, 2026 17:48
On a case-sensitive checkout, references/Notes.md and references/notes.md
could both be copied into a built skill, and they would overwrite each
other when the skill is installed on a case-insensitive file system. The
build now fails with SkillFileCaseConflict instead.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EmmjJxgtH4Ye2gWta1gjcj
The output budget section said that Print output is lost when MCP Local
restarts its subkernel after a non-abortable call outlasts the time
limit, which contradicted the environment matrix. With Chatbook 2.7.28,
the Print output of the input that timed out still arrives, without its
"During evaluation of In[n]:=" label, and Prints of earlier inputs of
the call keep their labels (checked against a Local-method server with
a 3 s limit and AbortProtect[Pause[40]]).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EmmjJxgtH4Ye2gWta1gjcj
HTTPDiagnostics returns its JSON report to whoever made the request,
including the request parameters, messages, stack frames and Print
text. With "Debug" -> Automatic as the default, any caller of a public
deployment could get that report by adding debug=1. "Debug" now
defaults to False, and the references say to debug a private deployment
as its owner and to remove the wrapper before others use the API.

Also stop the kernel's evaluation cache from reusing request data:
HTTPRequestData does not record its dependency on the request, so
after HTTPDiagnostics had run outside of a request (as when it is tried
locally), later requests saw no parameters and debug=1 was ignored.
requestData now calls Update[HTTPRequestData] before each read.

Checked on private deployments in the Wolfram Cloud: debug=1 is
ignored by default and answered with "Debug" -> Automatic, failures
still return the report, and anonymous requests get 401.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EmmjJxgtH4Ye2gWta1gjcj
CheckPaclet flagged the wolfram-debugging skill's helper package
(scripts/WolframDebugging.wl, both the source and the built copy in
Assets/AgentSkills) with CodeInspectionFileIssue/NotPublisherContext,
because its context is WolframDebugging` rather than one under Wolfram`.
That failed the CI build. The package is a standalone script that
agents load by file and call by its full names, not paclet code, so the
hint does not apply.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EmmjJxgtH4Ye2gWta1gjcj
On a time-out, GNU timeout -s KILL exits with 137 (it kills its process
group, itself included), not 124. Only uutils timeout, the default on
Ubuntu 26.04, exits with 124. RunIsolated therefore reported time-outs
as "Killed" on most Linux systems, including the CI image (Ubuntu 24.04,
coreutils 9.4), where the RunIsolated test failed.

Exit code 137 now counts as "TimedOut" when the run lasted the full time
limit, and as "Killed" when it ended earlier (e.g. out of memory).
Environments.md said that 124 is a time-out and 137 a kill from outside;
it now gives both meanings of 137.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EmmjJxgtH4Ye2gWta1gjcj
Where $Notebooks is True, Echo prints a cell with CellPrint instead of
calling Print. That happens with a front end, e.g. under UsingFrontEnd,
and in cloud kernels: CloudEvaluate, deployed APIs and the remote MCP
server. CapturePrints only blocked Print, so there Echo output escaped
it. CI runs the whole test suite under UsingFrontEnd, which is why the
CapturePrints test failed there.

CapturePrints now blocks CellPrint as well and records each cell as
text. Echo cells give the same text as Echo prints without a front end,
e.g. ">> lbl 2". Checked under UsingFrontEnd, in CloudEvaluate and in a
deployed API; the remote MCP server routes Echo through CellPrint too.

The new CapturePrints-Notebooks test runs with $Notebooks False and
True, so both paths are tested locally and in CI.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EmmjJxgtH4Ye2gWta1gjcj
@rhennigan
rhennigan merged commit bf25da5 into main Oct 9, 2026
1 check passed
@rhennigan
rhennigan deleted the feature/wolfram-debugging-skill branch October 9, 2026 20:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Evaluator "Local" method parses tool-call code into the shared Global` context instead of the session context

2 participants