AI Jam Sessions · Research note

Do jam-actions traces train a better tool-grounded musical-QA model?

A preregistered confirmatory evaluation on frozen fine-tuned artifacts.

mcp-tool-shop · AI Jam Sessions

An open dataset + fine-tuning research program · preregistered, LLM-assisted analysis with an external deterministic verifier

On-platform research note & project page — not an arXiv-indexed paper.

AI Jam Sessions — piano keys dissolving into a tool-node network
Verdict · frozen wording class, bar met

Powered win.

The jam-actions v1 recipe trains a model that beats the prompted baseline at tool-grounded musical QA on a preregistered 36-record cohort dominated by held-out material. Five frozen fine-tunes (no retraining, no reselection, artifacts sha-pinned before this cohort existed), scored once each against a fresh sealed baseline, move the primary condition from 0.678 to 0.890.

+0.212
mean Δ, primary condition
29/36
paired wins vs a 24/34 bar
p < 0.0001
exact sign test (p = 0.000039)
10/12
wins on never-trained music

Abstract

jam-actions is an open dataset of executable, tool-grounded reasoning traces over symbolic music. We ask a narrow question: does fine-tuning on these traces produce a model that answers musical questions better than a strong prompted baseline?

Across three preregistered arcs the answer sharpened. v0 was an honest negative. v1 was directionally better but underpowered — 12/16 paired wins, one short of a frozen bar that a 16-record cohort could barely resolve. This arc, B-1, widens the sealed cohort to 36 records dominated by never-trained material, mints a fresh sealed baseline on the published v0.5.0 records, and re-scores the frozen v1 artifacts. Nothing was retrained or reselected. On the primary tool-grounded condition the fine-tunes move accuracy from 0.678 to 0.890 (Δ +0.212; song-cluster CI95 [0.128, 0.305], the interval that grazed zero at n=16), winning 29 of 36 paired records against a preregistered 24/34 bar (p < 0.0001), and clearing significance on the twelve held-out clair-de-lune records the models never saw. They remain below baseline when answering from prose alone — reported here with equal weight.

The result

Four evaluation conditions, all-seeds means over the 36-record cohort (n=3 per condition). The fine-tunes win decisively where their tools are — the tool_inspected surface — and lose on the prose-only and no-tools surfaces.

0.000.250.500.751.000.6780.890Δ +0.212tool_inspectedprimary0.3960.313Δ −0.083fullsecondary0.3810.307Δ −0.074text_onlycontrol0.3970.304Δ −0.093random_midicontrol
Prompted baseline — qwen2.5:7b Fine-tuned — v1 recipe, all-seeds mean
Figure 1. Primary and secondary conditions. The fine-tunes answer better by inspecting (tool_inspected: +0.212) and worse by recalling prose (full, text_only, random_midi all below baseline). Only tool_inspected carries the preregistered claim; the controls are shown to bound it.
ConditionBaselineFine-tunedΔWinsRecord CI95Sign p
tool_inspected (primary) 0.678 0.890 +0.212 29/36 2 ties [+0.148, +0.275] 0.000039
full (secondary) 0.396 0.313 −0.083 9/36 3 ties [−0.136, −0.033] 0.0135
text_only (control) 0.381 0.307 −0.074 13/36 1 tie [−0.123, −0.026] 0.175
random_midi (control) 0.397 0.304 −0.093 9/36 [−0.150, −0.034] 0.0039

Three arcs, one dataset

The claim is the endpoint of a documented sequence, not a single lucky run. Each arc was preregistered before any model was called; each reports all five seeds, no best-of-seeds.

v0
Honest negative
Fine-tuning did not beat the prompted baseline. Reported as-is.
v1
Directionally better, underpowered
12/16 paired wins, Δ +0.202, p = 0.0043 — one win short of a frozen ≥13/16 bar a 16-record cohort could barely resolve.
B-1 · this note
Powered win
Same frozen artifacts, wider preregistered cohort. 29/36 wins, Δ +0.212, p < 0.0001. The effect holds at proper power.

The question B-1 existed to answer — was v1's 12/16 miss a power artifact or a real ceiling? — is answered: power artifact.

Method

The confirmatory cohort (36 records, three strata)

The cohort is frozen by a seeded, byte-replayable derivation; its receipt is the authority (sha 7ea74fc9…). It is dominated at the margin by material the models never trained on.

StratumnWhat it isBaseline → FTΔWinsExact p
CL 12 Every clair-de-lune test record — never in any training corpus, any arc, any form. 0.729 → 0.906 +0.176 10/12 0.039
LG 15 The sealed-history train-song cohort (v0/v1 continuity). 0.711 → 0.887 +0.176 11/15 (1t) 0.057
NW 9 Seeded-blind new train records, drawn before any B-1 output existed. 0.556 → 0.876 +0.320 8/9 (1t) 0.0078
0.000.250.500.751.000.7290.906Δ +0.176CL · n=1210/12 wins0.7110.887Δ +0.176LG · n=1511/15 (1t) wins0.5560.876Δ +0.320NW · n=98/9 (1t) wins
Prompted baseline — qwen2.5:7b Fine-tuned — v1 recipe, all-seeds mean
Figure 2. Primary condition, unpooled by stratum. CL is the confirmatory headline: on twelve clair-de-lune records the models never saw in training — eleven of which had never been evaluated by anything — the frozen fine-tunes win 10/12 and the stratum clears significance on its own. LG reproduces v1's seen-song outcome exactly (11/15, 1 tie) under a different baseline on corrected records — the measurement is stable; the added power came from widening the cohort.

What was pinned, what never moved

What did not improve — reported with equal weight

The prose-only and no-tools surfaces remain below baseline, consistent with both prior arcs: full −0.083, text_only −0.074, random_midi −0.093 (all record-CIs exclude zero; text_only's sign test does not reach significance). The v1 fine-tunes' competence is concentrated where their tools are — they answer better by inspecting and worse by recalling prose. The preregistered claim does not extend to prose surfaces; that is the target of a future prose-surface retrain, not this arc.

Outcome-dependence, disclosed

This arc exists because v1 missed its bar by one win, and a confirmatory re-test designed after seeing a near-miss carries selection risk. The mitigations are all mechanical: the artifacts were frozen before this cohort existed; the added material is dominated by never-trained records (all 12 CL records structurally excluded from every training corpus; the 9 NW records drawn blind by a fixed seed); the bar was frozen ex-ante as an exact-number table at the same α the v1 bar used, recomputed for the new n, not loosened; one eval per model, no reruns, all five seeds report; and both possible outcomes had pre-committed wordings — a miss would have been a publishable result, not a retry trigger.

The claim, and the gate that opened

jam-actions traces, augmented with execution-verified grounding-shaped examples (the v1 recipe), train a model that beats the prompted baseline at tool-grounded musical QA — +0.21 mean, 29/36 paired wins against a preregistered 24/34 bar, p < 0.0001, strongest on never-trained music — while remaining below baseline when answering from prose alone.

Per the preregistered claim class, a passing result opens a director gate before anything ships. The gate fired 2026-07-11 with this result in view: publish all five seed adapters, with the claim tied to the all-seeds mean, per-seed numbers disclosed, no best-of-seeds. The published artifact is mcp-tool-shop/jam-ft-v1-qwen25 — the five selected-epoch PEFT adapters, byte-identical to the frozen artifacts this note evaluated (adapter_model.safetensors sha256s on the card and in the publish receipt).

Reproducibility & receipts

Citation

This note presents an evaluation of the jam-actions dataset. Please cite the dataset:

@misc{jam_actions_2026,
  title        = {jam-actions-v0: executable tool-grounded reasoning traces over symbolic music},
  author       = {{mcp-tool-shop} and {AI Jam Sessions}},
  year         = {2026},
  howpublished = {Hugging Face Datasets},
  doi          = {10.5281/zenodo.21313954},
  url          = {https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v0},
  note         = {Confirmatory fine-tuning evaluation (arc B-1): +0.212 on tool-grounded
                  musical QA, 29/36 paired wins vs a preregistered 24/34 bar, p < 0.0001.}
}