AI Jam Sessions · Research note
A preregistered confirmatory evaluation on frozen fine-tuned artifacts.
An open dataset + fine-tuning research program · preregistered, LLM-assisted analysis with an external deterministic verifier
On-platform research note & project page — not an arXiv-indexed paper.
The jam-actions v1 recipe trains a model that beats the prompted baseline at tool-grounded musical QA on a preregistered 36-record cohort dominated by held-out material. Five frozen fine-tunes (no retraining, no reselection, artifacts sha-pinned before this cohort existed), scored once each against a fresh sealed baseline, move the primary condition from 0.678 to 0.890.
jam-actions is an open dataset of executable, tool-grounded reasoning traces over symbolic music. We ask a narrow question: does fine-tuning on these traces produce a model that answers musical questions better than a strong prompted baseline?
Across three preregistered arcs the answer sharpened. v0 was an honest negative. v1 was directionally better but underpowered — 12/16 paired wins, one short of a frozen bar that a 16-record cohort could barely resolve. This arc, B-1, widens the sealed cohort to 36 records dominated by never-trained material, mints a fresh sealed baseline on the published v0.5.0 records, and re-scores the frozen v1 artifacts. Nothing was retrained or reselected. On the primary tool-grounded condition the fine-tunes move accuracy from 0.678 to 0.890 (Δ +0.212; song-cluster CI95 [0.128, 0.305], the interval that grazed zero at n=16), winning 29 of 36 paired records against a preregistered 24/34 bar (p < 0.0001), and clearing significance on the twelve held-out clair-de-lune records the models never saw. They remain below baseline when answering from prose alone — reported here with equal weight.
Four evaluation conditions, all-seeds means over the 36-record cohort (n=3 per condition). The fine-tunes win decisively where their tools are — the tool_inspected surface — and lose on the prose-only and no-tools surfaces.
tool_inspected carries the preregistered claim; the controls are shown to bound it.| Condition | Baseline | Fine-tuned | Δ | Wins | Record CI95 | Sign p |
|---|---|---|---|---|---|---|
| tool_inspected (primary) | 0.678 | 0.890 | +0.212 | 29/36 2 ties | [+0.148, +0.275] | 0.000039 |
| full (secondary) | 0.396 | 0.313 | −0.083 | 9/36 3 ties | [−0.136, −0.033] | 0.0135 |
| text_only (control) | 0.381 | 0.307 | −0.074 | 13/36 1 tie | [−0.123, −0.026] | 0.175 |
| random_midi (control) | 0.397 | 0.304 | −0.093 | 9/36 — | [−0.150, −0.034] | 0.0039 |
The claim is the endpoint of a documented sequence, not a single lucky run. Each arc was preregistered before any model was called; each reports all five seeds, no best-of-seeds.
The question B-1 existed to answer — was v1's 12/16 miss a power artifact or a real ceiling? — is answered: power artifact.
The cohort is frozen by a seeded, byte-replayable derivation; its receipt is the authority (sha 7ea74fc9…). It is dominated at the margin by material the models never trained on.
| Stratum | n | What it is | Baseline → FT | Δ | Wins | Exact p |
|---|---|---|---|---|---|---|
| CL | 12 | Every clair-de-lune test record — never in any training corpus, any arc, any form. | 0.729 → 0.906 | +0.176 | 10/12 | 0.039 |
| LG | 15 | The sealed-history train-song cohort (v0/v1 continuity). | 0.711 → 0.887 | +0.176 | 11/15 (1t) | 0.057 |
| NW | 9 | Seeded-blind new train records, drawn before any B-1 output existed. | 0.556 → 0.876 | +0.320 | 8/9 (1t) | 0.0078 |
p4-receipt.json. No retraining, no reselection — selection at v1 used only inner-validation records and never saw any cohort record.qwen2.5:7b on the published v0.5.0 records — because erratum-002 corrected the Bach records' prose that the prompts embed. Both arms saw identical prompts.The prose-only and no-tools surfaces remain below baseline, consistent with both prior arcs: full −0.083, text_only −0.074, random_midi −0.093 (all record-CIs exclude zero; text_only's sign test does not reach significance). The v1 fine-tunes' competence is concentrated where their tools are — they answer better by inspecting and worse by recalling prose. The preregistered claim does not extend to prose surfaces; that is the target of a future prose-surface retrain, not this arc.
This arc exists because v1 missed its bar by one win, and a confirmatory re-test designed after seeing a near-miss carries selection risk. The mitigations are all mechanical: the artifacts were frozen before this cohort existed; the added material is dominated by never-trained records (all 12 CL records structurally excluded from every training corpus; the 9 NW records drawn blind by a fixed seed); the bar was frozen ex-ante as an exact-number table at the same α the v1 bar used, recomputed for the new n, not loosened; one eval per model, no reruns, all five seeds report; and both possible outcomes had pre-committed wordings — a miss would have been a publishable result, not a retry trigger.
jam-actions traces, augmented with execution-verified grounding-shaped examples (the v1 recipe), train a model that beats the prompted baseline at tool-grounded musical QA — +0.21 mean, 29/36 paired wins against a preregistered 24/34 bar, p < 0.0001, strongest on never-trained music — while remaining below baseline when answering from prose alone.
Per the preregistered claim class, a passing result opens a director gate before anything ships. The gate fired 2026-07-11 with this result in view: publish all five seed adapters, with the claim tied to the all-seeds mean, per-seed numbers disclosed, no best-of-seeds. The published artifact is mcp-tool-shop/jam-ft-v1-qwen25 — the five selected-epoch PEFT adapters, byte-identical to the frozen artifacts this note evaluated (adapter_model.safetensors sha256s on the card and in the publish receipt).
experiments/finetune-arc-v2/P0-LOCK.md, frozen at commit c9dcd17 before any model call.data/b1-cohort.json (sha 7ea74fc9…, sampler seed 20260712).8bef493d…, eval-logic files byte-identical to the sealed v0/v1 pins.evals/b1-stats.json — seeded RNG, pooled + per-stratum + per-record table.jam-actions-v0-0.5.0-cut-2026-07-11, DOI 10.5281/zenodo.21313954.This note presents an evaluation of the jam-actions dataset. Please cite the dataset:
@misc{jam_actions_2026,
title = {jam-actions-v0: executable tool-grounded reasoning traces over symbolic music},
author = {{mcp-tool-shop} and {AI Jam Sessions}},
year = {2026},
howpublished = {Hugging Face Datasets},
doi = {10.5281/zenodo.21313954},
url = {https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v0},
note = {Confirmatory fine-tuning evaluation (arc B-1): +0.212 on tool-grounded
musical QA, 29/36 paired wins vs a preregistered 24/34 bar, p < 0.0001.}
}